A benchmark is a proxy
The score measures performance on the test, not on your work. The moment a number becomes the target, it stops measuring the thing you care about.
Benchmarks
Every week another model claims the top spot. The leaderboards are public, but reading them wrong costs you the right tool. This page links every scoreboard worth checking, and gives you a reading method that does not go stale.
Links checked July 2026.
The score measures performance on the test, not on your work. The moment a number becomes the target, it stops measuring the thing you care about.
When every frontier model scores above 90 percent, the test has stopped discriminating. Progress did not stop; the ruler ran out. Look for the newer, harder successor.
Models may have seen public test questions during training. Boards with private or rotating question sets are harder to game.
The gap between the top models is usually smaller than the gap between a good and a bad prompt. Read trends, not crowns.
Human-preference votes, coding-ticket completion, and agent autonomy measure different things. The best chat model is not automatically the best coding model.
Speed and price decide what you will actually use daily. A slightly weaker model you run on everything beats the perfect one you ration.
Benchmarks now exist for every kind of AI, not just chat. Match the board to your domain and your use case.
Start here
Independent measurements of intelligence, speed and price for every major model in one view. The best first stop, because it forces the three-way trade-off you actually face.
artificialanalysis.ai
Millions of blind human votes on real prompts. Measures which answers people prefer, and preference is not correctness, because style wins votes.
lmarena.ai
A research nonprofit tracking capability trends over time, including the hardest math problems. Read it for trajectory, not for who leads today.
epoch.ai
Coding
Can a model fix real GitHub issues from real repositories on its own. The closest public proxy for completing a developer's ticket end to end.
www.swebench.com
Practical code editing across many programming languages, maintained by the author of a working coding tool. Read it for day-to-day pair-programming strength.
aider.chat/docs/leaderboards/
Questions rotate on a schedule, so models cannot have memorized them. A contamination-resistant second opinion.
livebench.ai
How well agents handle real work in a command-line environment. Read it for agentic coding beyond single-file edits.
www.tbench.ai
Agents and autonomy
Measures how long a task a model can complete on its own, and how fast that horizon grows. The single most important trend line for when AI can carry your multi-hour work.
metr.org
Agents operating a real computer: files, apps, browsers. Read it for how far computer-use agents actually are, versus the demos.
os-world.github.io
A model runs a simulated vending business for a long stretch, and the score is whether it stays coherent over hundreds of steps or drifts into nonsense. Read it for long-horizon reliability, the thing short demos hide.
andonlabs.com/evals/vending-bench-2
Vision and multimodal
Expert-level questions that mix text with images, charts and diagrams across college subjects. Read it for whether a model can reason about what it sees, not just label it.
mmmu-benchmark.github.io
Blind human votes on prompts that include an image. Read it for which model people prefer when the question depends on a picture.
lmarena.ai
Media generation
Blind human votes ranking text-to-image models. Read it for prompt-faithfulness and look, not for speed or price.
artificialanalysis.ai/text-to-image/arena
The same blind-vote method for text-to-video. The fastest-moving board here, so expect the order to shift month to month.
artificialanalysis.ai/text-to-video/arena
Ranks models that change an existing image from an instruction. The capability behind change this, keep the rest.
artificialanalysis.ai/image/leaderboard/editing
Science and medicine
Doctor-written conversations scoring medical accuracy and safety across multiple turns. Read it as a floor for clinical caution, never as medical advice.
openai.com/index/healthbench/
Practical biology-research tasks: reading figures, protocols and databases. Read it for how close AI is to a capable research assistant at the bench.
github.com/Future-House/LAB-Bench
Chemistry knowledge and reasoning scored against expert chemists. Read it for domain competence, not lab safety.
lamalab-org.github.io/chembench/
Standard tasks for predicting a material's properties, the groundwork under AI materials discovery. Read it to compare methods, not products.
matbench.materialsproject.org/
A shared yardstick for data-driven global weather forecasting. Read it for how AI models compare against traditional physics-based forecasts.
sites.research.google/gr/weatherbench/
The long-running community experiment where protein-structure prediction is measured, the arena that made the case for AlphaFold. Read it for the state of the art in structural biology.
predictioncenter.org/
Forecasting and finance
A live test of predicting future real-world events, refreshed automatically so it cannot be memorised. Read it for calibrated judgement under uncertainty, not recall.
www.forecastbench.org/
Bots forecasting real open questions next to human forecasters. Read it for how AI stacks up against skilled human prediction.
www.metaculus.com/aib/
Independent evaluations of AI on real professional work in law, finance and tax. Read it for reliability in domains where a mistake is expensive.
www.vals.ai/
The edge of reasoning
Novel puzzles humans find easy and models find hard, designed to resist memorization. Measures adapting to the genuinely new.
arcprize.org
Expert-written questions at the frontier of human knowledge. Read it as distance to the expert ceiling, not day-to-day usefulness.
artificialanalysis.ai/evaluations/humanitys-last-exam
Trick questions where ordinary people still beat frontier models. Built by the researcher behind AI Explained, one of the channels in our library. A humility check on the hype of the week.
simple-bench.com
Games and play
Models play full games head-to-head, from chess to social games like werewolf. Read it for strategic planning, and in social games for cooperation and bluffing.
www.kaggle.com/benchmarks/kaggle/game-arena
An interactive-reasoning benchmark built as novel little game worlds an agent has to learn to play. Read it as the game-native successor to static puzzle tests.
arcprize.org
Kept for context. These once defined the field and are now saturated: every strong model scores near the top, so they no longer separate models. A saturated test is not a failed one; it means the field moved on.
Broad knowledge across 57 subjects. The default benchmark of the early 2020s, now saturated near the human ceiling and superseded by MMLU-Pro and Humanity's Last Exam.
github.com/hendrycks/test
Grade-school math word problems. Saturated: frontier models cluster at the top, so it no longer separates them.
github.com/openai/grade-school-math
Small Python coding problems. The original code benchmark, now saturated and superseded by SWE-bench and LiveBench.
github.com/openai/human-eval
Commonsense sentence completion. Saturated, kept for historical comparison.
rowanzellers.com/hellaswag/
The language-understanding suite that defined progress before large chat models. Fully saturated, of historical interest.
super.gluebenchmark.com/
The image-classification dataset that kick-started the deep-learning era in 2012. Long saturated, and the reason today's vision models exist.
www.image-net.org/
The boards tell you what models can do. Knowing what to do with them is the other half.