Back to writing

Model intelligence benchmarks are flawed

2026·08·054 min

Right now, a lot of people following AI advances are absolutely certain that models are in a recursive self-improvement loop. The evidence for this is that each successive model saturates each benchmark set and the graphs look like hockey-stick charts.

Here's the most famous example that's always shared on social media that supposedly shows model intelligence improving exponentially.

Frontier language model intelligence index increasing across model release dates from November 2022 to July 2026
Frontier Language Model Intelligence, over time. Source: Artificial Analysis.

The number of tasks they can complete in each benchmark is growing exponentially, so model intelligence must be growing exponentially too, right?

This is only true if each task in each benchmark requires progressively more intelligence. If, let's say, a benchmark is made up of 100 tasks (20 require an "IQ"1 of 110, 50 require an "IQ" of 120, and the final 30 require an "IQ" of 130), then when the model reaches these IQ thresholds, it will look like exponential improvement without an exponential increase in intelligence.

Linear intelligence vs. benchmark score

Linear improvement

IQ: 100 → 140 (+5 each release)

Linear improvementIQ: 100 → 140 (+5 each release)100110120130140100M1M2110M3M4120M5M6130M7M8140M9Model release

Benchmark score

Score doubles from 1% to 64%, then saturates at 99%.

Benchmark scoreScore doubles from 1% to 64%, then saturates at 99%.0%25%50%75%100%1%M1M24%M3M416%M5M664%M7M899%M9Model release
An expanded illustration of the same effect: linear gains can produce an exponential-looking benchmark curve before the score saturates.

Am I saying it isn't improving exponentially? No, but I think a lot of the increase in capability could be explained by linear improvements in intelligence too. Fundamentally, the shape of these benchmark curves alone cannot tell us the rate at which model intelligence is improving.2

We will only find out what's happening with the next few model releases. If it's truly increasing in intelligence exponentially, we will have models with 300+ IQ in the next few years. This type of intelligence and ability to reason will truly be alien to everyone, and we will hardly be able to make sense of it.

Finally, we shouldn't underestimate the impact of consistent linear improvements in intelligence and the combination of a superhuman (but not exponentially improving) reasoning model with the entire corpus of knowledge at its, figurative, fingertips.

  1. 1.

    IQ is calibrated relative to the human population and cannot literally measure intelligence this far beyond the human range but I decided to use it anyway because it's easy to understand

  2. 2.

    The problem with model intelligence benchmarks exists with human testing too! Psychometricians established Item Response Theory (IRT) as a way of modeling the probability that a person of a given ability answers a question correctly. This model could then be used to determine the difficulty of a question and to determine the ability of a test taker that answers the question correctly.

Also published on X