Connect with us

Hi, what are you looking for?

AI Tools Directory

AGI May Have a Measurement Problem, Not an Intelligence Problem

AGI May Have a Measurement Problem, Not an Intelligence Problem
AGI May Have a Measurement Problem, Not an Intelligence Problem

For most of the past decade, the working assumption was that AGI would announce itself a jump in capability so unmistakable that nobody would need a scoreboard. That assumption is quietly dying.

The reason is not that AI has stopped getting smarter. It is that the instruments used to decide “smarter” are breaking faster than they can be replaced. A frontier model can embarrass the state of the art from two years ago and simultaneously reveal the benchmark that crowned it as nearly useless. We are now in the strange position where AGI is declared arrived by one chief executive, disputed by definition, and impossible to independently certify. The sharpest framing of today’s AGI controversy the one behind the recent debate over GPT-6 Astra is therefore not “does the intelligence exist.” It is: can we measure it when it does?

Benchmarks stop working. The field finally measured that.

The new evidence is not anecdotal. In June 2026, the EvalEval Coalition a collaborative evaluation project spanning multiple institutions presented a systematic study of 60 text-based LLM benchmarks at ICML. Their saturation index formalizes a problem researchers have suspected for years: benchmarks become unable to distinguish leading models long before models master them.

The results are blunt. Roughly 29 of the 60 benchmarks about 48 percent already show high or very high saturation. On a saturated benchmark, a model scoring 91 percent, another 92, and another 92.3 are statistically indistinguishable; the gaps are smaller than the measurement noise. A benchmark in that state is no longer a measuring instrument. It is a historical document.

The study also demolished the popular beliefs about what keeps a benchmark alive. Private test sets provide no protective effect once the distributional characteristics become known. Open-ended formats do not outperform multiple choice. Multilinguality does not help once you control for age. The two strongest predictors of saturation are mundane: how old the benchmark is and how large its test set is. Expert-curated benchmarks like ARC-AGI and BIG-Bench Hard lasted longest but none lasts forever.

The saturation treadmill, in three acts

The arc of the ARC-AGI benchmarks designed by François Chollet to test fluid intelligence rather than accumulated knowledge is the clearest demonstration that the measurement problem is accelerating.

Act one: ARC-AGI-1. Released in 2019, it stumped frontier systems for years; the state of the art hovered around 34 percent while humans solved the bulk of it with little effort. In late 2024, a preview of OpenAI’s o3 model finally broke through but the fine print mattered. It scored 76 percent at an estimated 200pertask,or88percentatroughly200pertask,or88percentatroughly20,000 per task. The same model, shipping later at lower compute, scored 53 percent. And high-performing humans, the benchmark’s creators noted, solved more than 97 percent of tasks without much effort. ARC-AGI-1 had saturated below the human ceiling it never managed to measure the upper range of human fluid intelligence at all.

Act two: ARC-AGI-2. Launched in 2025 with a human-calibration study (400-plus participants in San Diego; every task solved by at least two people in under two attempts), the new instrument was deliberately hard. At launch, the best reasoning models managed single-digit scores. It also introduced something the first version lacked: an explicit efficiency metric, because, in the organizers’ words, “intelligence is not solely defined by the ability to solve problems.” From then on, the score reported was not just capability but the cost of intelligence.

The treadmill then sped up. By July 2026, Claude Opus 5 scored 90.4 percent on ARC-AGI-2 at maximum reasoning effort near saturation roughly 18 months after launch.

Act three: ARC-AGI-3. The newest generation widened the gap between instruments again. On ARC-AGI-3, Claude Opus 5 reached 30.16 percent (at high reasoning effort), while ARC Prize reported GPT-6 Astra at 62.7 percent in its standard setup, at a cost of roughly $26,000. Two frontier models, more than double apart, on an instrument that is already several generations ahead of what either company is likely to cite in a launch presentation.

Notice the pattern. The industry doesn’t raise the bar the way a sport does slowly, with rules. It raises it the way a hedge fund raises a wall: each generation buys months, not years. The measurement is losing faster than the intelligence is winning.

Goodhart is doing the measuring

There is an older name for what is happening. Economist Charles Goodhart observed in the 1970s in the phrasing popularized by Marilyn Strathern that “when a measure becomes a target, it ceases to be a good measure.” Every benchmark in the ARC-AGI sequence saturates for the same underlying reason: the moment a metric becomes the objective, the system under test begins optimizing the metric rather than the property the metric was meant to stand for.

The mechanisms are now routine. Models are trained and fine-tuned on the distribution of benchmark items, not just benchmark data. Contamination between training corpora and test sets is common enough to have spawned its own detection tools. And, as the commentary accompanying Humanity’s Last Exam noted, today’s models can search the web during the test meaning a correct answer may reflect retrieval and digestion of a public solution, not reasoning at all.

Humanity’s Last Exam was built to be the answer to all this. Launched in January 2025 by the Center for AI Safety, Scale AI, and a consortium of contributors, it gathered 2,500 expert-level questions across more than a hundred subjects, filtered through more than 70,000 failed model attempts, and protected by a private held-out set and a bug bounty. Its creators billed it as the final closed-ended academic benchmark of its kind.

The trajectory since then tells the real story. GPT-4o opened at 2.7 percent. By the spring of 2025, Gemini 3 Pro had reached 38.3 percent. According to Artificial Analysis’ leaderboard, the current top model sits near 59 percent within shouting distance of the “expert-level closed-ended questions” ceiling, roughly eighteen months after launch. The creators now maintain a rolling version of the exam precisely because they know the static one will not last.

None of this is a failure of the people who built these benchmarks. It is the nature of the task. Every closed-ended, verifiable question has a finite answer space, and the models are becoming excellent at searching it.

Capability is not intelligence and systems are not models

Beneath the benchmark churn sits a deeper measurement error. The numbers we report as “model intelligence” are actually measurements of an entangled system: the training data, the retrieval layer, the test-time compute budget, the scaffolding, and increasingly the money spent.

The ARC-AGI-2 story made this visible. The o3 preview scored 88 percent on ARC-AGI-1 at high compute, but the shipped model scored 53. The first number was never a property of a model it was a snapshot of model plus compute plus harness at a specific price. When ARC-AGI began reporting cost as part of the score, it was conceding formally what was true all along: performance and efficiency are different properties, and treating the first as if it implied the second is how “intelligence” claims inflate.

Chollet’s 2019 essay, On the Measure of Intelligence, argued a similar point before the current generation existed: intelligence is best understood as the efficiency of acquiring new skills, not the accumulated knowledge a system can display. On that definition, a model fluent across 100 domains but helpless in genuinely unfamiliar ones is less intelligent than its benchmark collage suggests regardless of the average score.

Researchers studying what happens after a benchmark saturates have now proposed measuring agents along six dimensions rather than one: construct validity, out-of-distribution generalization, efficiency, reliability, the relative contribution of the model versus its scaffolding, and real-world uplift from human collaboration. That is a more honest list of what a maturity claim actually needs. It is also, notably, a list almost nobody outside the evaluation community publishes.

Which returns to the underlying problem: AGI claims are currently asserted under definitions nobody has ratified, using instruments whose expiry dates they don’t publish.

Measurement is becoming a governance matter

This is no longer an academic argument. Benchmarks are wired directly into high-stakes decision-making the EU AI Act’s Article 51 obligations and the UK’s AI Safety Institute’s Inspect framework both lean on reference evaluations. When nearly half the benchmarks in the ecosystem are saturated, static regulatory lists are quietly becoming stale rulebooks: they create the false confidence that a capability question has been settled when the measuring tool has simply stopped discriminating.

Every AGI claim now moves capital, hiring, and national strategy. When a claim is treated as both finding and marketing at once as was visibly the case in the Astra moment the temptation is to trust the printout rather than interrogate the instrument. That is exactly the pattern of over-trust described in our analysis of what happens when we stop checking what AI tells us. The measurement problem and the trust problem are the same problem viewed from two sides.

There is also the question nobody has answered: who is allowed to certify AGI? Nvidia’s chief executive can declare it arrived; a competitor’s lab can dispute it the same afternoon; an independent evaluator can report that the score depends on which instrument and which budget you chose. There is currently no recognized authority, no accepted definition, and no audit trail exactly the kind of governance vacuum explored in the debate over who should control AI. When a threshold has economic consequences this large and certification authority this weak, the result is the worst possible outcome: claims that cannot be falsified and therefore cannot be trusted.

What a serious measurement agenda would look like

The good news is that the fixes are mostly boring which is why they might actually happen. The EvalEval recommendations are a reasonable start: report uncertainty-aware statistics, maintain adversarial and dynamically updated evaluation sets, and treat each benchmark as a dated instrument with a documented expiry, not a permanent property of the model.

The bigger shift is structural. Measure systems, not just models, and print the whole bill capability, cost, reliability, and the hardware context all at once. Adopt the six-dimension approach for agentic systems rather than chasing a single accuracy number. And most importantly, separate the act of claiming from the act of certifying: an AGI declaration should be an audited measurement, not a press release, and it should name its definition, its instrument, and its uncertainty alongside the headline number.

None of this requires the field to agree on one perfect definition of intelligence. It requires something more modest: that progress claims meet the same evidentiary standard as any scientific claim.

The bottom line

The intelligence may be much closer than the measurements allow us to conclude or much farther. The honest position is that we cannot currently tell the difference, and our instruments are not going to fix that on their own. The saturating benchmarks, the retrieval-era contamination, the cost-adjusted scores, and the missing certification authority all point the same direction.

AGI’s hard problem was never going to be whether machines get smart enough. It is whether the field can build instruments trustworthy enough that an “AGI has arrived” headline means something other than a CEO’s opinion, at a time when the easiest thing to do is to keep moving the finish line.

More reading: where this intersects with AI’s real-world limits the common-sense gap that keeps agents from reliably stopping themselves.

You May Also Like

Tech

On Monday, September 28, 2026, OpenAI confirmed that it would not release GPT-6.1 Astra, a next-generation prototype that was nearing launch. The company had...

Blog

NVIDIA's Southeast Asia strategy goes far beyond chip sales: 1GW AI factories, language models it doesn't charge for, a robot testbed, and a supply-chain...

Blog

OpenAI has spent nine months failing to hire a communications chief. The rejections, the crisis backlog, the Altman problem, and what it means before...

Blog

The old MBC 1 Nilesat frequency 11938 V no longer works. What replaced it on Nilesat 201, where MBC 1 actually broadcasts in 2026...