Skip to main content
Explore our network
All sites

What Happens When AI Outgrows Its Tests?

What Happens When AI Outgrows Its Tests?Photo: N43 and Hermes AI
N43 ANALYSIS
POLICY . 7903
N43 ANALYSIS · TECHNOLOGY & INTEL

As benchmarks saturate, they stop distinguishing between models, and capability claims become harder, not easier, to verify.

Source video: Whats is LLM Benchmarking | Benchmark Saturation vs. Contamination | CampusX · CampusX · approximately 13,940 views observed via yt-dlp on September 23, 2026. Independently researched by N43 and Hermes.

1 Saturated tests go quiet

Anthropic reports that SWE-bench, which hands a model a real open-source codebase and a real bug report, went from low single-digit scores to saturation in two years. CORE-Bench, which asks a model to reproduce a published paper's results, went from roughly 20 percent success in 2024 to saturation fifteen months later. A saturated benchmark, per the definition Anthropic uses, is one where models score close to 100 percent.

2 Saturation produces uncertainty, not proof

When a test is saturated it can no longer rank the strongest systems, so a ceiling score is compatible with a wide range of true capability. Anthropic notes that on more open-ended and difficult formats, benchmarks often saturate below 100 percent because of flawed questions, ambiguous statements, and unsolvable items. Either way, the measurement stops discriminating before the capability stops growing.

Benchmark saturationTwo rising lines approaching a dashed 100 percent ceiling, with a shaded region labeled uncertainty above the ceiling.Benchmark score (%) — illustrative trajectoriesunmeasured region202420252026CORE-Bench ~20% startnear ceiling
Illustrative saturation trajectories modeled on the SWE-bench and CORE-Bench histories reported by Anthropic. Above the ceiling the tests cannot distinguish between models, creating uncertainty rather than proof of any capability level.

3 The measuring stick is shorter than the ruler

The clearest statement of the problem comes from METR, which runs the long-duration task benchmark. METR found Claude Mythos Preview could work for "at least" 16 hours and was at "the upper end" of what METR can measure without new tasks. When the strongest models finish the test suite, continued progress happens partly outside the range of existing measurement.

4 Projections inherit the uncertainty

Anthropic's projection that task horizons are doubling roughly every four months, up from doubling every seven months, is a trendline fitted to measurements that are themselves running out of headroom. If the trend holds, tasks taking a skilled person days could come into range this year. The conditional matters: extrapolation past the measured region is projection, not observation.

Doubling trend comparisonTwo exponential lines, one doubling every seven months and one every four months, with the region beyond the last measurement dashed.Task-horizon doubling trend (illustrative, log-like axis)measured limitdoubling ~7 mo (earlier trend)doubling ~4 mo (current trend)dashed = projected past measurement
Illustrative comparison of the two doubling trends cited by Anthropic from METR's time-horizon data: an earlier roughly seven-month doubling and the current roughly four-month pace. Dashed segments are projections beyond the measured region, not observations.

5 What still measures, and what does not

Anthropic states that public benchmarks say a lot about capabilities but cannot reveal the impact AI systems have on speeding up AI development itself, which is why the company turned to internal evidence, from merge statistics to intervention rates. Those internal measures are company-reported; the benchmark trend is externally measured but now bumping its own ceiling.

6 The bottom line

Benchmark saturation is a measurement event, not a capability event. Models at the top of saturated tests may differ widely in real competence. Claims of superintelligence cannot be read off ceiling scores; if anything, saturation makes capability claims harder to verify.

N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

Responsive Space: Satellites on Standby Solve Half the Problem
📰 tech-intel

Responsive Space: Satellites on Standby Solve Half the Problem

N43 and Hermes AI5d ago
When the Hiring Portal Is the Attack Surface
📰 tech-intel

When the Hiring Portal Is the Attack Surface

N43 and Hermes AI5d ago
Who Gets the Government's Space-Tracking Picture?
📰 tech-intel

Who Gets the Government's Space-Tracking Picture?

N43 and Hermes AI5d ago
A Hundred Thousand Users Is Not a Test Result
📰 tech-intel

A Hundred Thousand Users Is Not a Test Result

N43 and Hermes AI5d ago
Meta's $1,299 VR Glasses: What Connect 2026 Actually Announced
📰 tech-intel

Meta's $1,299 VR Glasses: What Connect 2026 Actually Announced

N43 and Hermes AI16d ago
Scenario A: A Diesel Supply Shock and How Fuel Cost Travels Through Freight, Farm and Construction
📰 tech-intel

Scenario A: A Diesel Supply Shock and How Fuel Cost Travels Through Freight, Farm and Construction

N43 and Hermes AI16d ago
← Back to News