DataCentreNews India - Specialist news for cloud & data centre decision-makers
India
Olas says new forecasting model matches GPT-4.1 benchmark

Olas says new forecasting model matches GPT-4.1 benchmark

Tue, 22nd Sep 2026 (Today)
Raphael Veloso
RAPHAEL VELOSO News Editor

Olas has released Olas-Predict-R1-14B, a 14-billion-parameter open-weight forecasting model. It says the model matched GPT-4.1 on prediction-market forecasting tests.

Developed by Valory, a core contributor to Olas, the model was fine-tuned on 214,529 prediction records from 5,116 resolved markets. The records came from Olas Predict, where AI agents have been making forecasts and trading on events ranging from elections and court rulings to product launches.

Valory then tested the model on 2,628 separate markets that resolved after the training period. Based on the figures released, it reached 75.8% forecasting accuracy on those unseen markets, up from 71.4% for the base model, DeepSeek-R1-Distill-Qwen-14B.

That represents roughly a 20% reduction in forecasting error relative to the base model. Olas also says the model can run on a single GPU, unlike larger systems that typically require more computing resources.

Training data

The results add to a broader debate in artificial intelligence over whether targeted training data can improve specialist systems, rather than relying on larger model sizes alone. Here, the training data came from real-world forecasts with known outcomes, giving researchers examples they could check against what actually happened.

Prediction markets are one of the few environments where large volumes of probabilistic judgments are routinely recorded and later resolved. That makes them useful for training forecasting systems because each prediction can be compared with an eventual result, creating a feedback loop that is harder to obtain in many other AI applications.

Valory Chief Executive Officer, Dr. David Minarsch made that case in comments accompanying the launch.

"The headline in AI has been that better performance requires bigger, more powerful general-purpose models. Our results point to another path. Prediction markets produce particularly valuable training data because every forecast is ultimately tested against a real-world outcome. In our case, that gives the model hundreds of thousands of examples of what better and worse forecasting looks like across thousands of real events," said Minarsch. 

"By turning those outcomes into a direct learning signal, we can train a relatively lean, specialized model that can run on a single GPU and is highly capable at the task it was built for. And because the model is open-weight, developers can self-host it, integrate it into their own products or agents as a forecasting component, and use it independent of any third party, making user-owned AI a reality," Minarsch added. 

Specialist use

The model does not retrieve information on its own. Instead, it is designed to sit within a broader application, where developers provide the evidence or external data and use the model to produce a probability estimate.

That makes its immediate use case narrower than that of general-purpose chatbots or reasoning models. Developers could use it as the forecasting component in a prediction-market agent or as part of a decision-support tool that needs a numerical estimate of the likelihood of an outcome.

Organisations can run the model on their own infrastructure rather than relying entirely on third-party application programming interfaces. That may appeal to teams that want more control over deployment and cost, particularly when working with private data or needing predictable spending.

The release also reflects the continued growth of open-weight AI models, which developers can download and host themselves. In the commercial market, this has become one of the clearest dividing lines between providers that offer access through hosted services and those trying to broaden adoption through more direct distribution.

Broader push

For Olas, the forecasting model fits into a broader effort to build products around AI agents and market-based systems. Olas Predict has been operating since 2023, generating a large archive of forecasts that can now be reused as training data.

That archive may become more valuable if repeated rounds of fine-tuning produce further gains. Olas plans to continue training new models as more records are collected and to test them against future markets to measure whether forecasting improves.

The benchmark claim is likely to draw attention because GPT-4.1 is a much larger general-purpose system than a 14B model. While strong performance on a specialist benchmark does not translate into broader superiority, it offers another example of how smaller systems can remain competitive on tightly defined tasks when trained on domain-specific data.

Whether that pattern can extend beyond prediction markets will depend on the availability of similarly rich datasets in other fields. What makes this case distinctive is that the underlying task produces a steady stream of outcome-verified examples, which are often scarce in AI training.

In the reported evaluation, Olas-Predict-R1-14B achieved 75.8% accuracy across 2,628 unseen markets after being fine-tuned on 214,529 predictions from 5,116 resolved markets.