FutureEval: Metaculus's AI Forecasting Benchmark
How good are AI forecasting bots, really? FutureEval compiles three years of research and testing across all major models to ground hype in data.
New research in FutureEval from April and May, 2026. We’ve published more analysis on what makes a forecasting bot work, covering model choice, research strategy, and advice from the bot makers themselves. Highlights are at the bottom of the post, with full writeups on our resources page. If you’re already familiar with FutureEval, scroll to Research Highlights section what’s new.
What FutureEval is
Forecasting is an unfalsifiable test of real reasoning: information retrieval, world modeling, judgment under uncertainty. FutureEval measures how accurately AI predicts future outcomes across science, technology, health, geopolitics, AI itself, and more.
Three pillars
Model Leaderboard. We run major AI models on a fixed prompt across most open Metaculus questions, then score and rank them as those questions resolve. Performance over time shows how fast AI forecasting is improving
Bot Tournaments. Developers enter their own bots and compete for a share of $175k in prizes a year. The primary $50k seasonal tournament repeats every four months and is always open. A faster $1k tournament, MiniBench, runs every two weeks
Human Baselines. Tournament questions draw from the Metaculus community, and our hand-picked Pro Forecasters forecast a subset each season. Together they set a high bar to measure the bots against
April-May 2026 Research Highlights
In May we published a synthesis of 11 Metaculus analyses run since 2024, a three-season study of bot search providers, and a roundup of advice from the bot makers themselves. The top takeaways:
Pros still beat bots in comparisons to date. Across four quarterly tournaments and the ongoing leaderboard, head-to-head spot peer scores have put Pros in the lead every season by a large margin. External claims that bots already match the best humans all rest on backtesting against already-resolved questions, and have not been reproduced under live, forward-looking evaluation.
Pick a current frontier reasoning model. Model choice was the largest differentiator in bot performance, and simple one-shot bots on frontier models placed in the top five of the leaderboard. Prompting strategy alone cannot lift a sub-frontier model to frontier level. Good scaffolding is worth around nine months of base model progress, but only on top of the right model.
Breadth is the predictor. No individual search provider showed a consistent advantage across seasons. What correlated with score was the number of distinct research sources a bot used: winners averaged about 1.7, non-winners 1.00. Separately, agentic and iterative search tends to beat one-shot retrieval.
Post-forecast adjustments are worthwhile. Most winners aggregate across multiple forecasts, apply post-hoc calibration, and cap predictions at a min and max. Reading the logs matters too: one maker lost roughly $500 in prize money from a single uncaught bug on a rare question format.
Full writeups are on our FutureEval News page.
Join the current tournament
The $50k Summer Bot Tournament is live, with 300 to 500 questions running through early September. It is always open to new entrants: you can join any time and start at the middle of the leaderboard with 0 points, entering an early version of your bot and improving it over the season.







