A study in measured intelligence · 2021–2026
AI Out-Scores 999 of 1,000 People
In April 2026, a frontier AI model out-scores 999 of 1,000 people on the Mensa Norway reasoning test.
Five years ago it out-scored 81. This page shows the crossing — and what the viral charts leave out.
Since 2023, a project called Tracking AI has been giving language models the same 35-question pattern-recognition test that Mensa Norway offers to humans, and converting the results to an IQ-style score. The scores climbed from below-average to the top of the human scale in about five years. That much is real.
But a single number hides what actually happened. So instead of a line on a chart, start with a room.
Prologue · A millennium on one side of the line
For a thousand years, everyone was smarter than AI
Zoom all the way out. The filled shape below is everyone alive, from the year 1000 to today. For 950 of those years the question this page asks is meaningless: there is no AI to be smarter than. Computers don't exist until the 1940s. The first machine to even take this test scores 79, in 2021.
Watch the camera. It sweeps a thousand years, then zooms — twice — because that is the only honest way to show a story whose entire plot happens in the last five years of a millennium. The red share rising from the bottom is the fraction of humanity the best model out-scores.
The compute strip along the bottom is the explanation hiding under the spectacle: a line that is exactly zero for nine and a half centuries, becomes measurable within living memory, and rises so fast at the end that the camera has to zoom twice just to keep it on screen. The red tide you just watched is what that line bought. Whether it caused it is a longer argument — this page only claims they arrived together.
Act I · The crossing
Put 1,000 people in a room. Draw a line down the middle.
On one side stand the people who would still out-score the best AI model on this test. On the other, the people the model now beats. Each time the best score to date rises, more people walk across the line. (The December 2025 peak of 147 and today's leaderboard top of 145 both round to the same 999 of 1,000 — so at this resolution, nobody walks back.)
Watch the middle of the decade. In early 2024 the room is split almost exactly in half — GPT-4's 101 is dead-average, and half the room crosses in a single step. By September 2024 (o1, score 120), 909 people have crossed. After that the crossings get small not because progress stopped, but because almost nobody is left: the walk from 136 to 145 moves only seven people.
That's the honest shape of the story. The frontier is now deep in the human tail, where every additional point buys fewer and fewer people.
Act II · The ladder
The people it passed on the way up
Percentiles are abstract. People are not. Here is the same climb told as a lineup of measured humans. (One note first: IQ is age-normed — a bright ten-year-old scores 100 against other ten-year-olds — so "the IQ of a child" isn't a rung on this ladder. People and professions are.)
Two of the rungs are famous for a reason. Richard Feynman's school IQ test came back 125 — a number his biographer treats as proof of what such tests miss, since Feynman went on to be, by acclamation, one of the smartest humans of his century. And Albert Einstein has no rung at all: he was never tested. Every "Einstein's IQ" figure you have ever seen is invented.
The Feynman rung is the one to sit with. If a paper-and-pencil test could underrate Feynman by that much, it can misjudge a machine in either direction — and that cuts both ways for every number on this page. The ladder is honest about what it measures: performance on one narrow test, climbing past the measured scores of specific people. It is not a claim that any model out-thinks Feynman. Nobody who met him would take that trade.
Act III · The line they drew
The viral chart plots the top of a wide band
The version of this story going around plots one line: the best score at each moment. But at any given moment, many models exist, and they are not close together. Models released in Q1 2025 alone spanned 97 to 126 — a 29-point range, wider than the gap between an average person and a Mensa member. The median model runs 10–25 points below the frontier in every quarter with enough data to check.
Two things this chart refuses to do, on purpose: it does not space the x-axis evenly by data point (which visually smooths the early years into a tidy climb), and it does not average vision-model scores into the line (vision variants score 20–60 points lower on the same test; mixing modalities muddies the series). And note the right edge — the newest points sit slightly below the December 2025 peak. The top of the leaderboard has compressed, with three frontier models within four points of each other. A "2.5 points per month, and not slowing" claim does not survive contact with that.
Act IV · The tail
Where each score lands on the human curve
The room and the line are two views of the same distribution. Here is the third: the bell curve itself, with the shaded area showing the share of people the frontier model out-scores at each milestone.
Data table — every value used on this page
| Period | Frontier model | Score | Out-scored, per 1,000 | Status |
|---|
What this does not say
The trajectory is real. These five caveats are what keep it honest:
- It is one narrow test. The Mensa Norway test is 35 visual pattern-recognition puzzles. It measures nothing about coding, factual reliability, tool use, or professional judgment. "IQ" is a convenient label doing a lot of work; "score on the Mensa Norway reasoning test" is the defensible phrasing.
- The early anchors are estimates. Tracking AI did not run systematically before ~2023. GPT-3 = 79, Claude-2 = 83, and GPT-4 = 101 are retrospective reconstructions — directionally fine, precisely sourced, no. They are drawn hollow above for that reason.
- Contamination is a live concern. The test has been public online for years. Models may have seen it, or things like it, in training.
- The far tail is extrapolated. IQ norms above ~130 are thin even for humans. "999 of 1,000" is the arithmetic of a normal curve, not a measured census — the room is a model, and it is only as good as the normality assumption in its tail.
- Frontier ≠ fleet. The best model is not the typical model. The median release runs 10–25 points behind the frontier, and vision variants of the same models score 20–60 points lower. Any claim about "AI's IQ" that quotes only the top number is quoting the whisker, not the box.