In January 2025, Scale AI and the Center for AI Safety published the first results from Humanity’s Last Exam, a benchmark assembled from expert questions designed to remain difficult even for frontier artificial intelligence systems. Questions had been crowdsourced through fall 2024 from a global pool of specialists, producing an initial public release of 3,000 questions.
The first scores were tiny. GPT-4o managed 2.7 percent, Claude 3.5 Sonnet reached 4.1 percent, and OpenAI’s o1 scored 8 percent. Yet the 2026 Stanford AI Index records the striking result that followed: frontier models gained 30 percentage points on Humanity’s Last Exam in a single year.

A test built by nearly 1,000 experts to sit beyond the machines
Humanity’s Last Exam was created in response to benchmark saturation. Once leading models begin scoring close to the ceiling of a test, that test becomes much less useful for distinguishing one frontier system from another.
The official HLE project says the finalized benchmark contains 2,500 questions across more than 100 subjects, contributed by nearly 1,000 subject experts affiliated with more than 500 institutions in 50 countries. The questions range from mathematics and natural science to ancient languages and specialized academic fields.
The path to that final 2,500-question set matters. Scale AI’s January 2025 release described an initial public exam of 3,000 questions. After community review and a bug-bounty process, HLE was finalized at 2,500 questions on April 3, 2025, with flagged and searchable questions removed and replaced.
The questions were built to have single, verifiable answers and to resist quick retrieval. A Texas A&M account of the project gives examples involving Palmyrene inscriptions, bird microanatomy and Biblical Hebrew pronunciation.
That same account identifies Dr. Tung Nguyen, an instructional associate professor in Texas A&M’s Department of Computer Science and Engineering, as one of the project’s major contributors. Nguyen contributed 73 of the 2,500 public questions, the second-highest total by an author, and wrote more math and computer-science questions than any other contributor.
Why the old benchmarks stopped separating the models
Humanity’s Last Exam was not created because existing benchmarks had become useless in every sense. It was created because some of the best-known academic tests were becoming too easy for frontier models to serve as useful measuring sticks at the top end.
As SciTechDaily’s Texas A&M-sourced coverage noted, evaluations such as MMLU had become less effective at distinguishing cutting-edge systems as scores approached the ceiling. HLE deliberately moved the difficulty frontier outward.
Its question-selection process reinforced that goal. Candidate questions were tested against leading systems, and questions that contemporary models could already answer were removed. The result was not an ordinary exam sampled from an established curriculum. It was constructed around the remaining edge of model capability.
The scoreboard moved quickly
The early numbers made HLE look forbidding: GPT-4o at 2.7 percent, Claude 3.5 Sonnet at 4.1 percent and o1 at 8 percent. Those are the figures behind the benchmark’s reputation as a test built specifically to expose what frontier models still could not do.
By February 2026, the picture had changed sharply. Google’s Gemini 3.1 Pro model card reported 44.4 percent on the full HLE set without tools. With search and code enabled, the score rose to 51.4 percent. The same comparison listed Claude Opus 4.6 at 40.0 percent without tools and 53.1 percent with them.
Those distinctions matter. A no-tools score and a score produced with search and code are not interchangeable measures, even when they use the same underlying benchmark. The headline trend remains dramatic, but the evaluation setup has to travel with the number.
And the frontier continued moving. In April 2026, Anthropic reported results for Claude Mythos Preview of 56.8 percent without tools and 64.7 percent with tools. Anthropic also cautioned that the model’s HLE performance at low effort could indicate some memorization, another reason leaderboard numbers need context rather than being treated as pure measures of intelligence.
Reasoning time is one part of the story
One documented shift behind newer reasoning systems is the use of additional computation while answering. In its technical introduction to o1, OpenAI reported that o1’s performance improved both with more reinforcement learning during training and with more time spent thinking at inference, often called test-time compute.
That offers one concrete reason reasoning-heavy benchmarks began moving. A system that can spend additional computation working through a difficult problem is operating differently from one expected to produce an immediate first-pass response.
But HLE scores alone cannot tell us exactly how much of the year-over-year gain came from test-time compute, new training methods, better data, tool use or other changes inside the frontier labs. The benchmark measures the result. It does not isolate every cause behind that result.
What a 50 percent score actually means
A model reaching 50 percent on Humanity’s Last Exam has become much better at answering a deliberately difficult set of closed-ended expert questions. That is significant. It is not the same thing as demonstrating half of human intelligence, half of expert knowledge or general intelligence.
A University of Sydney publication of an analysis by Kai Riemer and Sandra Peter makes that distinction directly. A benchmark can show whether a model produces correct answers on the tasks in front of it without establishing that the system possesses the professional judgement, lived context or open-ended research abilities a human expert brings to an ambiguous situation.
HLE’s creators make a similarly narrow claim for the benchmark itself. High accuracy would demonstrate strong performance on closed-ended, verifiable questions and cutting-edge academic knowledge. It would not, by itself, demonstrate autonomous research capability or artificial general intelligence.

The public questions and the private holdout
There is another important distinction in HLE’s design. The finalized 2,500-question dataset is publicly released. Separately, the project maintains a private set of held-out questions to help assess overfitting.
That is different from keeping most of the 2,500-question benchmark secret. Public questions let researchers inspect and reproduce the benchmark, while the private holdout gives the organizers an additional check on whether strong performance reflects generalization rather than simple exposure to the released material.
The organizers have also acknowledged that a static benchmark cannot remain untouched forever. On October 8, 2025, they released HLE-Rolling, a dynamic fork intended to keep the evaluation moving as models improve and the original question set ages.
A benchmark can become a target
There is a built-in tension in every successful benchmark. Researchers want a stable test so progress can be compared across systems and across time. But once a leaderboard matters, model developers have a strong incentive to improve on the capabilities the benchmark rewards.
That does not make HLE meaningless. It means the score should be read as a measurement of a specific capability under a specific evaluation setup, not as a universal ranking of everything an AI system can do.
Stanford’s 2026 AI Index makes the broader point in its discussion of benchmark reliability and gaming: as model performance rises, evaluation itself becomes a moving research problem. HLE is unusually visible because the movement has happened so quickly.
The people writing the questions
The human side of the benchmark is easy to miss when attention is fixed on the leaderboard. HLE drew contributions from nearly 1,000 subject experts across hundreds of institutions and more than 100 subject areas.
Nguyen’s 73-question contribution illustrates how much specialized work sits behind the final number. The project combined people working in computer science with contributors from history, physics, linguistics, medicine and many other fields, each trying to identify questions that remained easy to verify but difficult for contemporary models to solve.
In that sense, HLE is not a catalogue of everything humans know. It is a deliberately selected map of places where expert knowledge remained ahead of frontier models when the questions were assembled.
What comes after a benchmark built to be beaten
The speed of HLE’s score increases does not mean the exam has already become worthless. It does show why its creators launched a rolling version rather than assuming one frozen set of questions could define the frontier indefinitely.
Evaluation is also broadening beyond closed-ended academic questions. The 2026 Stanford AI Index tracks agentic systems that operate in software environments and complete longer sequences of actions, alongside traditional reasoning, language and multimodal benchmarks. Those evaluations ask a different question: not simply whether a model knows an answer, but whether it can carry out a task over multiple steps.
No single benchmark captures all of those capabilities. HLE remains useful precisely because its target is narrower: difficult, verifiable, expert-level questions for which correctness can be scored cleanly.
The number to keep in mind
The striking number is still the one in Stanford’s summary: a 30-percentage-point gain in one year on a benchmark intentionally constructed to be difficult for frontier AI.
The surrounding numbers make the pace tangible. GPT-4o began at 2.7 percent. By February 2026, Gemini 3.1 Pro was reported at 44.4 percent without tools. By April, Anthropic reported an unreleased Mythos Preview model at 56.8 percent without tools, while also warning that possible memorization complicated interpretation of that result.
Humanity’s Last Exam has not stopped measuring the frontier. The frontier has simply moved much faster than a benchmark with that name might have suggested. Its creators are already updating the test, and the next difficult question is no longer whether today’s models can improve on it. It is how long any fixed exam can remain ahead of them.