A study in measured intelligence · 2021–2026

AI Out-Scores 999 of 1,000 People

In April 2026, a frontier AI model out-scores 999 of 1,000 people on the Mensa Norway reasoning test.

Five years ago it out-scored 81. This page shows the crossing — and what the viral charts leave out.

Since 2023, a project called Tracking AI has been giving language models the same 35-question pattern-recognition test that Mensa Norway offers to humans, and converting the results to an IQ-style score. The scores climbed from below-average to the top of the human scale in about five years. That much is real.

But a single number hides what actually happened. So instead of a line on a chart, start with a room.

Prologue · A millennium on one side of the line

For a thousand years, everyone was smarter than AI

Zoom all the way out. The filled shape below is everyone alive, from the year 1000 to today. For 950 of those years the question this page asks is meaningless: there is no AI to be smarter than. Computers don't exist until the 1940s. The first machine to even take this test scores 79, in 2021.

Watch the camera. It sweeps a thousand years, then zooms — twice — because that is the only honest way to show a story whose entire plot happens in the last five years of a millennium. The red share rising from the bottom is the fraction of humanity the best model out-scores.

1000
Out-score the best AI AI scores higher
Filled area: world population, UN & HYDE historical estimates. Red share: fraction of people the best Mensa Norway score to date out-scores, assuming IQ ~ N(100, 15). Bottom strip: training compute of the largest AI run, log scale, order-of-magnitude estimates in the style of Epoch AI (Theseus 1950 ≈ 104.6 FLOP → AlexNet 2012 ≈ 1017.7 → GPT-4 2023 ≈ 1025.3 → 2026 frontier ≈ 1026.7, est.). Applying today's test-score distribution to the whole population is a visual metaphor, not a claim about medieval psychometrics.

The compute strip along the bottom is the explanation hiding under the spectacle: a line that is exactly zero for nine and a half centuries, becomes measurable within living memory, and rises so fast at the end that the camera has to zoom twice just to keep it on screen. The red tide you just watched is what that line bought. Whether it caused it is a longer argument — this page only claims they arrived together.

Act I · The crossing

Put 1,000 people in a room. Draw a line down the middle.

On one side stand the people who would still out-score the best AI model on this test. On the other, the people the model now beats. Each time the best score to date rises, more people walk across the line. (The December 2025 peak of 147 and today's leaderboard top of 145 both round to the same 999 of 1,000 — so at this resolution, nobody walks back.)

81 of 1,000 crossed
Still out-scores the best AI AI scores higher
Each dot is one person in a normally distributed room (mean 100, SD 15). A dot crosses when the top Mensa Norway score on trackingai.org exceeds its IQ. 2021–2024 anchors are retrospective estimates — see caveats.

Watch the middle of the decade. In early 2024 the room is split almost exactly in half — GPT-4's 101 is dead-average, and half the room crosses in a single step. By September 2024 (o1, score 120), 909 people have crossed. After that the crossings get small not because progress stopped, but because almost nobody is left: the walk from 136 to 145 moves only seven people.

That's the honest shape of the story. The frontier is now deep in the human tail, where every additional point buys fewer and fewer people.

Act II · The ladder

The people it passed on the way up

Percentiles are abstract. People are not. Here is the same climb told as a lineup of measured humans. (One note first: IQ is age-normed — a bright ten-year-old scores 100 against other ten-year-olds — so "the IQ of a child" isn't a rung on this ladder. People and professions are.)

Two of the rungs are famous for a reason. Richard Feynman's school IQ test came back 125 — a number his biographer treats as proof of what such tests miss, since Feynman went on to be, by acclamation, one of the smartest humans of his century. And Albert Einstein has no rung at all: he was never tested. Every "Einstein's IQ" figure you have ever seen is invented.

Human anchors: average adult = 100 by definition; college-graduate mean ≈ 115 (occupational studies); Feynman's school test = 125 (Gleick, Genius); Mensa cutoff = 130; Kasparov ≈ 135 (magazine-administered test, 1987, as reported); "genius" = 140 (Terman's 1916 Stanford-Binet label). Einstein was never tested. AI scores: Tracking AI, Mensa Norway, text models.

The Feynman rung is the one to sit with. If a paper-and-pencil test could underrate Feynman by that much, it can misjudge a machine in either direction — and that cuts both ways for every number on this page. The ladder is honest about what it measures: performance on one narrow test, climbing past the measured scores of specific people. It is not a claim that any model out-thinks Feynman. Nobody who met him would take that trade.

Act III · The line they drew

The viral chart plots the top of a wide band

The version of this story going around plots one line: the best score at each moment. But at any given moment, many models exist, and they are not close together. Models released in Q1 2025 alone spanned 97 to 126 — a 29-point range, wider than the gap between an average person and a Mensa member. The median model runs 10–25 points below the frontier in every quarter with enough data to check.

Frontier (best available) Mensa Norway score over time, on a true time axis. Hollow points are retrospective estimates. The box shows the spread of models released in Q1 2025 (min 97, quartiles 101–121, median 109, max 126); the tick above Q2 2025 marks that quarter's median (122). Source: Tracking AI; Q1 2025 distribution via the June 2025 Visual Capitalist table.

Two things this chart refuses to do, on purpose: it does not space the x-axis evenly by data point (which visually smooths the early years into a tidy climb), and it does not average vision-model scores into the line (vision variants score 20–60 points lower on the same test; mixing modalities muddies the series). And note the right edge — the newest points sit slightly below the December 2025 peak. The top of the leaderboard has compressed, with three frontier models within four points of each other. A "2.5 points per month, and not slowing" claim does not survive contact with that.

Act IV · The tail

Where each score lands on the human curve

The room and the line are two views of the same distribution. Here is the third: the bell curve itself, with the shaded area showing the share of people the frontier model out-scores at each milestone.

Human IQ distribution (mean 100, SD 15). Shaded: the fraction of people scoring below the AI frontier at each milestone. Percentiles above ~130 rely on extrapolated norms — treat the far tail as indicative, not precise.
Data table — every value used on this page
PeriodFrontier modelScoreOut-scored, per 1,000Status

What this does not say

The trajectory is real. These five caveats are what keep it honest:

  1. It is one narrow test. The Mensa Norway test is 35 visual pattern-recognition puzzles. It measures nothing about coding, factual reliability, tool use, or professional judgment. "IQ" is a convenient label doing a lot of work; "score on the Mensa Norway reasoning test" is the defensible phrasing.
  2. The early anchors are estimates. Tracking AI did not run systematically before ~2023. GPT-3 = 79, Claude-2 = 83, and GPT-4 = 101 are retrospective reconstructions — directionally fine, precisely sourced, no. They are drawn hollow above for that reason.
  3. Contamination is a live concern. The test has been public online for years. Models may have seen it, or things like it, in training.
  4. The far tail is extrapolated. IQ norms above ~130 are thin even for humans. "999 of 1,000" is the arithmetic of a normal curve, not a measured census — the room is a model, and it is only as good as the normality assumption in its tail.
  5. Frontier ≠ fleet. The best model is not the typical model. The median release runs 10–25 points behind the frontier, and vision variants of the same models score 20–60 points lower. Any claim about "AI's IQ" that quotes only the top number is quoting the whisker, not the box.