Benchmarks

The scoreboards, and how to read them

Every week another model claims the top spot. The leaderboards are public, but reading them wrong costs you the right tool. This page links every scoreboard worth checking, and gives you a reading method that does not go stale.

Links checked July 2026.

How to read any leaderboard

01

A benchmark is a proxy

The score measures performance on the test, not on your work. The moment a number becomes the target, it stops measuring the thing you care about.

02

Watch for saturation

When every frontier model scores above 90 percent, the test has stopped discriminating. Progress did not stop; the ruler ran out. Look for the newer, harder successor.

03

Ask about contamination

Models may have seen public test questions during training. Boards with private or rotating question sets are harder to game.

04

Rank one changes weekly. Trajectory does not.

The gap between the top models is usually smaller than the gap between a good and a bad prompt. Read trends, not crowns.

05

Match the board to your use case

Human-preference votes, coding-ticket completion, and agent autonomy measure different things. The best chat model is not automatically the best coding model.

06

Score is one axis of three

Speed and price decide what you will actually use daily. A slightly weaker model you run on everything beats the perfect one you ration.

The boards

Benchmarks now exist for every kind of AI, not just chat. Match the board to your domain and your use case.

Start here

Coding

Agents and autonomy

Vision and multimodal

Media generation

Science and medicine

Forecasting and finance

The edge of reasoning

Games and play

Revealed preference

Archive: saturated benchmarks

Kept for context. These once defined the field and are now saturated: every strong model scores near the top, so they no longer separate models. A saturated test is not a failed one; it means the field moved on.

The boards tell you what models can do. Knowing what to do with them is the other half.