Model intelligence benchmarks are flawed
Right now, a lot of people following AI advances are absolutely certain that models are in a recursive self-improvement loop. The evidence for this is that each successive model saturates each benchmark set and the graphs look like hockey-stick charts.
Here's the most famous example that's always shared on social media that supposedly shows model intelligence improving exponentially.

The number of tasks they can complete in each benchmark is growing exponentially, so model intelligence must be growing exponentially too, right?
This is only true if each task in each benchmark requires progressively more intelligence. If, let's say, a benchmark is made up of 100 tasks (20 require an "IQ"1 of 110, 50 require an "IQ" of 120, and the final 30 require an "IQ" of 130), then when the model reaches these IQ thresholds, it will look like exponential improvement without an exponential increase in intelligence.
Linear intelligence vs. benchmark score
Model intelligence
Linear improvement
IQ: 100 → 140 (+5 each release)
Tasks completed
Benchmark score
Score doubles from 1% to 64%, then saturates at 99%.
Am I saying it isn't improving exponentially? No, but I think a lot of the increase in capability could be explained by linear improvements in intelligence too. Fundamentally, the shape of these benchmark curves alone cannot tell us the rate at which model intelligence is improving.2
We will only find out what's happening with the next few model releases. If it's truly increasing in intelligence exponentially, we will have models with 300+ IQ in the next few years. This type of intelligence and ability to reason will truly be alien to everyone, and we will hardly be able to make sense of it.
Finally, we shouldn't underestimate the impact of consistent linear improvements in intelligence and the combination of a superhuman (but not exponentially improving) reasoning model with the entire corpus of knowledge at its, figurative, fingertips.