Publications
Opinion

A Student Does Not Grade Their Own Exam

Rodrigo CavalcantiSeptember 2026

A student does not grade their own exam, and a big tech company cannot audit its own models. That said, on September 4, 2026, Artificial Analysis published version 4.2 of its Intelligence Index, which measures different capabilities of large language models. The release was relatively quiet because it touches on a subject the academic field would rather avoid: the impossibility of full reproducibility. Its announcement fit in a single line: tasks closer to real-world use and more private test sets to prevent gaming. Only three days later, on September 7, Artificial Analysis rushed out version 4.3.

In v4.2, private test sets accounted for 40% of the index. Version 4.3 raised that share to 45%, more than twice the weight used in the previous generation of the benchmark, and the plan is to increase it further in v5. This made us proud at IDJÉ because private tests have been part of our design from the start. “Does that make you anti-science?” Of course not. We are against gaming the system. From the outset, our question has been simple: if the “thermometer” used to measure the promises made about AI models is broken, how can we trust future promises about alignment?

What changed at Artificial Analysis

AA-Briefcase now accounts for 15% of the index. It contains 91 tasks across four realistic work scenarios. GDP.pdf, developed by Surge AI, evaluates document reasoning through 100 tasks distributed across 4,592 pages, with responses assessed against 1,275 criteria.

At the other end, GPQA Diamond was removed because it had reached saturation. This is the familiar life cycle of public benchmarks such as MMLU, MATH and GSM8K. A dataset is released, quickly becomes the target of optimization and then undergoes what the field calls benchmaxxing: direct or indirect optimization for a known test. It can be used once or twice. After that, it is over.

The usual sequence is predictable. The test is published, papers on arXiv describe the findings, the items and their solutions enter pre-training or fine-tuning corpora, and scores rise. The instrument stops measuring an ability and starts training it. Then it becomes useless.

Publishing an evaluation set therefore means giving big tech companies annotated data for free. The effect is political as well as methodological: whoever defines the metric defines the problem, and the problem shrinks once it must fit the evaluator’s instrument. A closed test set disrupts this logic. The model encounters items it has not seen before, and the score is more likely to reflect ability rather than memorization.

A recent case shows that protecting the items is not enough: the evaluation environment must also be controlled. On ARC-AGI-3, the same GPT-6 Astra scored 62.7% with ARC Prize’s standardized harness, on a respected evaluation, and 99.9% with OpenAI’s own adapter. To compare models, what matters is a common evaluation condition controlled by an independent evaluator. That is the original purpose of a benchmark: to measure the model, not all the superpowers of the apparatus built around it.

Neither IDJÉ nor Artificial Analysis started this movement to preserve datasets. Humanity’s Last Exam (HLE), created by Scale AI and the Center for AI Safety, was designed with approximately 2,500 public questions and an additional private set for detecting overfitting. A systematic gap between performance on the public and private sets signals rote learning. The hidden questions serve as a detector: if the discrepancy is large, there is evidence that the public material was used in training.

Why IDJÉ was right

A two-layer model is taking shape. The methodology is open. The scoring weights, agent and grading prompts, harness, rubric taxonomy and scored items remain under the evaluator’s custody. The evaluation field should not take local annotation work and export that intelligence in exchange for a publication and another line on an academic CV. Big tech companies already have every incentive to keep scraping the world’s intelligence.

There is a second argument, one of ecological validity. Local items carry context, normative references, practices and linguistic uses that do not travel without loss. When exposed, precisely those situated features become the next optimization target. The test then begins to measure exposure to itself. The paradox is straightforward: the same capacity that helps a model improve at a task also helps it escape the test designed to measure that task.

There is, of course, a serious objection. Private tests concentrate power in the evaluator. Whoever controls the benchmark defines the weights, chooses the judges and sets the scales. Trust in the figures therefore depends on trust in the guardian of the test. Who funds the evaluator? A big tech company, or a shell institute that facilitates capture? Is the methodology sufficiently open to permit criticism? Does the evaluator disclose enough examples to show that the closed set preserves the situated context that gives it meaning and rigor?

IDJÉ is aligned with the direction now being taken by major independent laboratories: an open benchmark becomes “homework” with a very short shelf life. Once the test becomes the target, it stops being a test and becomes training material. Artificial Analysis’s decision to more than double the weight of private evaluations, retire saturated tests and publish only the method and a small sample points in the same direction. For a small Brazilian laboratory that has kept 100% of its datasets closed, this does not mean lagging behind. It means being ahead of a shift that one of the field’s leading references has now validated.

References

ARTIFICIAL ANALYSIS. Announcing Artificial Analysis Intelligence Index v4.2. September 4, 2026. Available at: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2. Accessed on: September 9, 2026.

ARTIFICIAL ANALYSIS. Announcing the Artificial Analysis Intelligence Index v4.3. September 7, 2026. Available at: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3. Accessed on: September 9, 2026.

ARTIFICIAL ANALYSIS. Artificial Analysis Intelligence Benchmarking Methodology. Available at: https://artificialanalysis.ai/methodology/intelligence-benchmarking. Accessed on: September 9, 2026.

ARC PRIZE. OpenAI's GPT-6 Astra on ARC-AGI-3. September 3, 2026. Available at: https://arcprize.org/blog/astra. Accessed on: September 9, 2026.

SCALE AI; CENTER FOR AI SAFETY. Submit Your Toughest Questions for Humanity’s Last Exam. Scale Blog, September 16, 2024. Available at: https://scale.com/blog/humanitys-last-exam. Accessed on: September 9, 2026.

PHAN, L. et al. Humanity’s Last Exam. arXiv:2501.14249, 2025. Available at: https://arxiv.org/abs/2501.14249. Accessed on: September 9, 2026.

DEEPLEARNING.AI. The Problem with Benchmark Contamination in AI. The Batch, October 30, 2024. Available at: https://www.deeplearning.ai/the-batch/the-problem-with-benchmark-contamination-in-ai. Accessed on: September 9, 2026.