blog cover

Fine-tuning made our offchain intelligence more performant and scalable

The smartest LLMs get the most attention for good reasons, but open models offer compelling value for many practical use cases. One major upside is that they let you specialize through fine-tuning, which can substantially improve task performance at a similar inference cost.

At Cambrian, we maintain agentic data pipelines to power our offchain data API products, including sentiment data derived from tweets on X. High accuracy matters, but the cost of processing millions of tweets also determines how much we can cover. We found fine-tuning gives the best value for processing data with LLMs at scale and expanding coverage for our offchain sentiment API endpoints. In this post, we cover our methodology for scoring the tweets we evaluate for social sentiment and explain why fine-tuning lets us make our product more performant than relying on vanilla LLMs.

Using AI to analyze tweets across various criteria

Cambrian maintains a suite of offchain intelligence endpoints, encompassing token analysis, sentiment shifts, alpha, and others. Powering these endpoints requires tracking and processing cryptocurrency-related tweets. We start by scoring every tweet on a rubric of relevance, sentiment, and alpha, which lets us direct deep research pipelines towards the most promising tweets in subsequent steps. In this blog post we will only focus on a proprietary scoring rubric covering relevance, sentiment, and preliminary alpha scores, which we apply to all tweets:

cambrian-fine-tuned-x-screening-1-model-performance-pareto-frontier.png

Our scoring process begins with matching tweets to a given cryptocurrency through deterministic steps. First, we use an LLM to confirm whether the association is correct. We then evaluate the sentiment the author expressed towards that cryptocurrency and assign a preliminary “alpha” score. The scores obtained during the LLM scoring stage are graded against a minimum threshold, which if met, determines whether a tweet is a good candidate for deeper research. This helps filter out tweets that lack substance.

We have performed the above steps on millions of tweets. Each new model offers a step improvement in our scoring outcome, but it requires reprocessing all tweets. The cost of using frontier models for each pass is prohibitive. So we began fine-tuning with the ultimate goal of achieving frontier-level performance at a fraction of the time and cost, so our users can get more X alpha from a larger pool of tweets. We describe our process and results in the following section.

Fine-tuning enables us to massively scale our coverage

We focused this particular fine-tuning effort on adding token relevance, sentiment, and what we define as alpha to all the tweets we track for our offchain social data API. To observe the value of fine-tuning, we compared the accuracy models could achieve across the three tasks on evaluation sets not seen in training, and we saw the expected tradeoff between price and performance. The chart below compares models by their estimated cost of scoring 1,000 tweets (on a log scale to see the relationships better), and average accuracy on the y-axis:

On one end of the spectrum, Fable 5 scored the highest but would only allow us to score about 535,200 tweets on a budget of $10,000. By comparison, our fine-tuned model can process over 1.3B tweets on the same budget - an over 2,500x volume increase.

The table below shows the models at the Pareto frontier that were available when training on our fine-tuned model was completed:

cambrian-fine-tuning-x-screening-1-model-cost-comparison-table.png

We found that introducing greater specialization through fine-tuning for these tasks gave us a much better tradeoff between accuracy and cost. Our fine-tuned solution that is live in production scored over 93% but at a fraction of the cost.

By using open-weight models, we fine-tuned them to fit our needs and gain much greater control over the inference process, which allowed us to improve performance and increase the number of tweets we can process. We also used affordable cloud GPU providers like SaladCloud that fit our needs and helped us quickly and cost-effectively score all historical data.

Conclusion

Fine-tuning helps us scale our agentic pipelines and provide broader coverage through our offchain API data products. Lower costs at the initial scoring stage let us evaluate more tweets and direct deeper research towards the most promising candidates. This translates to greater business value for our users, who can rely on our social sentiment endpoints knowing they are choosing a data provider that continuously maximizes the net number of tweets processed while optimizing for quality. That means more, better-quality data at no additional cost for users. We are already working on the next iteration of our production models to improve quality further as we continue to expand our coverage. Stay tuned for updates on our fine-tuning efforts.

Join the conversation - follow Ricky on X at @rickydata42 .


About Cambrian

Cambrian is the financial intelligence layer for agents and institutions. Our API delivers real-time and historical blockchain data, covering yield, liquidity positions, risk, trading activity, and market sentiment, for agentic and institutional DeFi applications. Founded in 2024, Cambrian is backed by Polychain Capital, Franklin Templeton, a16z crypto, Flow Traders, Selini Capital, and others.