Research essay · Benchmark design

What Makes a Benchmark Worth Testing?

How case construction shapes scores, what a smaller test set preserves, and how to build evidence for the decisions we actually care about.

A question has kept coming back in our work on agent evaluation: if a small subset of cases tells us almost as much as the full benchmark, what were the other cases contributing?

That question reaches beyond saving inference calls. A thousand cases can repeat the same failure condition. A difficult suite can give weaker models an uninformative wall of zeros. Two models can obtain the same score while differing sharply in tool use, state tracking, recovery, and efficiency. A smaller suite may preserve the average while erasing precisely those differences.

The ranking example makes the problem tangible: A solves six out of ten cases; B solves five, but all five are difficult. Calling A “better” silently assumes that the ten successes have equal value. Where did that assumption come from? Often, from how many cases happened to be collected or generated.

Case construction is already part of the scoring rule. The production process determines which demands enter a benchmark and how much weight they receive. Compression may faithfully preserve that score without validating its weights or preserving the evidence needed for other evaluation goals.

The practical question is therefore: how do we build a set of cases that provides enough distinct, trustworthy evidence for a stated decision?

01 / What does a smaller test set preserve?

“Evaluate performance” can mean estimating a fixed benchmark’s mean, ordering nearby models, predicting utility on future work, or describing behavior across several goals. These are different statistical targets.

tinyBenchmarks showed that 100 selected examples could estimate scores on roughly 14,000 MMLU questions in its studied settings. Anchor Points similarly exploited relationships between item outcomes across models to reduce evaluation cost.16 17 But How Reliable is Language Model Micro-Benchmarking? found that strong aggregate prediction need not yield reliable comparisons between close models; in its settings, distinguishing such models could require up to about 250 examples, where random sampling became competitive.18

There are at least three reasons a benchmark can be compressible. Its cases may repeat the same challenge. The calibration models may have highly correlated abilities, so they respond similarly even to substantively different cases. Or the prediction target may simply be a mean: recovering one number is easier than recovering an entire behavior profile.

A useful thought experiment exposes the assumption. Two models can behave identically on a selected subset and differently everywhere else. Observations on the subset alone cannot tell them apart. Prediction becomes possible through assumptions about cross-item relationships, model populations, or task structure—and those assumptions need validation on new models.

Intended use Evidence the suite must retain What a low score-estimation error leaves open
Estimate a fixed suite’s mean Calibration on held-out models Whether the original suite represents useful work
Compare nearby models Paired differences and uncertainty Whether a small gap has the correct sign
Diagnose several behaviors Evidence for each goal and interpretable omissions Whether a rare behavior disappeared during selection
Screen costly failures Detection within relevant risk strata Whether a high mean hides an important tail
Choose a deployment model Utility on an independent workload Whether the benchmark’s weights match actual needs

Our development comparisons sharpened a related distinction: more failed cases did not always mean broader coverage of task families, and improvements on individual coverage measures did not imply a better overall trade-off. Those observations concern coverage and cost. Establishing score preservation requires a separate held-out score-estimation experiment.

Compression is evidence about a prediction problem. It becomes evidence about benchmark quality only after we specify what information the benchmark was meant to preserve.

02 / The hidden step: case counts become weights

Partition a benchmark into non-overlapping task families. Family g contains ng cases, with model m achieving a family mean smg. The ordinary mean is:

SD(m) = Σg (ng / N) smg

The factor ng/N is a weight. If a workflow is easy to generate or verify, producing more variants increases its influence on the score. A production convenience has become an evaluation preference.

Equal weighting is well motivated when cases are sampled from the target workload and successes have comparable value. Multiple cases can also improve precision within a family. The problem arises when convenience sampling or generation quotas quietly stand in for workload frequency, error cost, or a declared scientific priority.

Consider the ten-case example. A solves all five routine cases and one of five hard cases. B solves none of the routine cases and all five hard cases. Now copy each hard-case record once, leaving every model response unchanged.

Suite composition A B Leader by the ordinary mean
5 routine + 5 hard 6/10 = 60.0% 5/10 = 50.0% A
5 routine + the same 5 hard records twice 7/15 = 46.7% 10/15 = 66.7% B
Independent task demands added 0 0 Ranking still reverses

This is a constructed counterexample: copying deterministic records adds no independent task evidence, but moves the hard-family weight from 1/2 to 2/3. Repeated stochastic executions are a different operation: they can reduce execution noise.

Let p be the hard-family weight in the intended use. Keeping the two family means fixed gives:

UA(p) = 1 − 0.8p    UB(p) = p    UA = UB ⇔ p = 5/9 ≈ 55.6%
Figure 1 · A ranking is conditional on task weights. Analytic toy example: the curves cross at p = 5/9. The vertical guides show the original mix and the duplicated hard-case mix.
Figure 1 · A ranking is conditional on task weights. Analytic toy example: the curves cross at p = 5/9. The vertical guides show the original mix and the duplicated hard-case mix.

A has the higher observed utility below that threshold; B has the higher utility above it. Equivalently, if each hard-case success is worth r times a routine success, B wins when r > 1.25. Difficulty itself does not choose p or r. An obscure hard task can be unimportant, while a simple everyday task can matter enormously.

The Benchmark Lottery provides an empirical counterpart: choosing four of SuperGLUE’s eight tasks yields 70 combinations and six different first-place models.21 The important design question is what justifies the combination.

This suggests a concrete separation: set family weights from the intended use, and allocate cases within families to improve precision. More variants should not automatically create more importance. For overlapping goals, aggregation needs an explicit overlap rule rather than the partition formula above.

Would item response theory automatically put B ahead?

In the simplest Rasch model, it would not. With the same complete item set, fixed item difficulties, common discrimination, a single ability dimension, and local independence, the total number correct is sufficient for the ability parameter. The likelihood derivative is Σyi − Σσ(θ − bi), so the fitted ability increases with the total correct. A score of 5/10 does not overtake 6/10 just because the correct items look harder. More elaborate models change the assumptions; they still do not determine what a user should value.

03 / What frontier labs choose to measure

Model releases reveal which kinds of work their authors emphasize: engineering changes, professional deliverables, business automation, computer use, and scientific workflows. The following snapshot was checked on October 8, 2026.

Release report Selected benchmarks Settings that define the comparison
OpenAI · GPT-6.1 Sol DeepSWE 1.1; GDP.pdf; AutomationBench 1.0.6; OSWorld 2.0; Terminal-Bench Science 0.1 OSWorld uses partial reward on a specified offline version; reasoning settings and per-task cost are reported.1
Anthropic · Sonnet 5.5 Terminal-Bench 4.0; FrontierCode 1.1; CursorBench 4.0; GDPval-AA v2.1; OSWorld 2.1 OSWorld is partial credit; the highest reasoning setting does not win every task. FrontierCode: 46.2% at Max versus 52.1% at Xhigh.2
Google DeepMind · Gemini 4 Argon Engineering, knowledge work, science, long context, and computer use The OSWorld offline evaluation specifies its version and 500-step budget, taking the best of three runs; self-evaluated and externally reported results are distinguished.3

My reading is that these selections serve three needs: relevance to product workflows, checkable outputs, and room to distinguish strong models at meaningful cost. That explains their usefulness as evaluation choices; the public reports do not disclose the complete internal selection process.

The benchmark name alone is insufficient to compare scores. Offline versus combined online/offline tasks, binary completion versus partial reward, one run versus best-of-three, and different tool permissions or agent frameworks define different measurements. A useful model claim is conditional on the task set, weights, framework, budget, and scorer version.

Adoption by a frontier lab establishes influence. To understand why a benchmark deserves confidence, we need to inspect its construction.

04 / What careful construction actually does

“Human-written,” “expert-designed,” and “algorithmically generated” describe production methods. Their value comes from the checks those methods enable.

Benchmark Construction evidence The next question
GPQA, original 448 expert-written science questions; expert accuracy 65%, non-expert accuracy 34% despite internet access.4 How far does scientific question answering transfer to open-ended research?
DeepSWE 113 tasks across 91 repositories and five languages; prompts, behavioral verifiers, reference solutions, and quality review.5 22 Which engineering workload does the repository and task mix represent?
AutomationBench 1.0.6 Six business functions, 47 tools, isolated company states, and positive and negative end-state assertions.6 Do workflow quotas and interaction rules match the intended deployment?
AppWorld, original Nine apps, 457 APIs, 750 tasks; programmatic state checks, including unwanted changes.7 What does success reveal beyond completing the specified task?
ToolSandbox Stateful tools, implicit dependencies, a simulated user, and intermediate/final milestones.8 Do milestones accommodate valid alternative paths?
OSWorld 2.0 108 long-horizon tasks; end-state grading and trajectory-level challenge exposure.9 Does challenge attribution remain reliable across frameworks and budgets?
HealthBench, original 5,000 conversations; 262 doctors; 48,562 contextual rubric criteria.10 How do rubric judgments relate to outcomes in the intended setting?

DeepSWE makes the acceptance process concrete. Each task includes a prompt, verifier, and reference solution; checks focus on observable behavior, exercise alternative valid implementations, and include repeated runs for flakiness. That supports confidence in individual tasks. Representativeness is a separate question: active open-source repositories above a popularity threshold define a particular slice of engineering. A fixed framework controls conditions but does not establish equal suitability for every model; the native-framework comparison in the construction report used ten tasks.22

AutomationBench deliberately creates ambiguous records, stale information, and business rules embedded in messages. Its checks include required actions and prohibited side effects. In the official example, satisfying five of six assertions gives partial credit; full success requires all six. These choices make “update a record” materially different from “update the correct record and avoid damaging another.” The family quotas still need a rationale if the score is to estimate an actual workload.6

Algorithmic generation can make goals explicit too. AutoBencher optimizes dimensions including difficulty, topic importance, and novel performance patterns.11 The next step is to validate the connection between those construction goals and the judgment the resulting suite supports.

The wider evidence shows why this connection matters. In Measuring what Matters, 29 experts reviewed 445 benchmark papers.12

Figure 2 · Defining a target and validating its measurement are separate steps. Bean et al., 445 papers: 78.2% defined the phenomenon, 53.4% provided construct-validity evidence, and 16.0% used uncertainty estimates or statistical tests. Categories overlap; these are rates of reported practices.
Figure 2 · Defining a target and validating its measurement are separate steps. Bean et al., 445 papers: 78.2% defined the phenomenon, 53.4% provided construct-validity evidence, and 16.0% used uncertainty estimates or statistical tests. Categories overlap; these are rates of reported practices.

05 / A response is a clue; execution makes it interpretable

An agent task labeled “state tracking” can fail before the agent reaches the relevant state change. It can also succeed through a valid path that never requires remembering old feedback. In both cases, the task result is meaningful, but the label overstates the evidence for that specific capability.

OSWorld 2.0 already analyzes whether challenges were encountered, blocked, or left untested.9 This suggests a stronger way to use response examples: inspect the gap between the demand a case claims to test and the demand its execution actually exercises.

Observed response Competing explanation A useful check
Removing old feedback changes nothing The task did not need it; another source supplied it; the model was robust; the intervention failed Inspect information access, verify the intervention, and distinguish valid alternative routes
Removing feedback sharply reduces success The information mattered; or length, formatting, feasibility, or tool behavior also changed Use a matched non-target edit and an information-restoration control
Many failures occur before a target challenge An upstream bottleneck obscures the target Report overall task outcomes alongside challenge-reach counts
Many cases respond in the same way Repeated evidence; or correlated source models Group by task family and test new model families
Partial score changes but completion does not Local progress changed; the completion threshold did not Read intermediate state and final delivery together

A useful audit follows five links: what the task promises to measure; which information and states execution actually reaches; whether the intervention preserves a valid task; whether matched and restoration controls isolate the intended change; and whether the interpretation transfers to another model family.

Legal alternative paths matter. If a task permits either a batch query or several smaller queries, both should pass when they produce the right result. If a purported memory task can be solved entirely from the current request, that is evidence against its memory interpretation, not necessarily against its value as an end-to-end task.

The denominators matter too. Reporting only trajectories that reached the challenge can make a model that fails early look strong on the few surviving runs. Retain both the overall task denominator and the challenge-reach denominator, with missing execution or scoring evidence shown separately.

The research opportunity is to identify systematic mismatches between claimed demands and observed evidence, then show that repairing them improves an independent evaluation decision.

06 / The scorer is part of the experiment

A verifier can reject a correct implementation because it expects an unstated function name, or accept a wrong implementation because it checks only a convenient output. Passing a reference solution establishes one positive example. A stronger acceptance test also exercises valid alternatives and deliberately wrong solutions.

OpenAI’s audit of 138 selected difficult SWE-bench Verified tasks reported material test or specification issues in 59.4% of that audited set.13 This directly identifies a mechanism by which model scores can reflect verifier requirements that the task did not state.

Figure 3 · Correct grading and metric meaning are two different questions. Left: the selected 138-task SWE audit, not the full 500-task suite; 40.6% is the remainder outside the listed issue categories. Right: OSWorld 2.0, Opus 4.8, maximum thinking, batched calls, 500 steps. Binary completion is 20.6%; mean partial reward is 54.8%. Sources: [13] and [9].
Figure 3 · Correct grading and metric meaning are two different questions. Left: the selected 138-task SWE audit, not the full 500-task suite; 40.6% is the remainder outside the listed issue categories. Right: OSWorld 2.0, Opus 4.8, maximum thinking, batched calls, 500 steps. Binary completion is 20.6%; mean partial reward is 54.8%. Sources: [13] and [9].

The right panel illustrates a different issue: a perfectly implemented metric can still be misread. Partial credit measures progress toward delivery; binary completion measures whether the entire requirement was met. A partial score of 54.8% does not mean 54.8% of tasks were completed.

Recent protocol audits also test whether unintended strategies can earn high scores. One causal-discovery example raised a score from 0.018 to 0.639 by exploiting a variable-ordering pattern in the generator.14 An agent-safety audit found that labeling every example “unsafe” achieved F1 = 0.690 on R-Judge, exceeding five of 21 models that made differentiated judgments.15 These are concrete tests of the relationship between the scoring rule and the behavior it is meant to reward.

A benchmark acceptance suite should therefore contain both sides: behaviors that deserve credit and behaviors that should be rejected, including unintended side effects.

07 / Under a fixed budget, optimize the evidence you need

Harder cases help when current models are saturated. When almost every model fails, adding more difficult cases can produce little additional information about their differences. In a one-dimensional Rasch model with unit discrimination, item information is p(1 − p), maximized at p = 0.5.

Figure 4 · Difficulty is relative to the population being measured. Analytic Rasch curve, I(θ) = p(1 − p), with p = σ(θ − b). It describes information for local ability estimation; business importance and failure cost are separate quantities.
Figure 4 · Difficulty is relative to the population being measured. Analytic Rasch curve, I(θ) = p(1 − p), with p = σ(θ − b). It describes information for local ability estimation; business importance and failure cost are separate quantities.

Rare-failure screening has another sample-size logic. With independent cases and a true failure probability of 1%, the probability of observing no failures in n trials is 0.99n.

Figure 5 · Estimating a mean and detecting a rare failure require different designs. In this independent 1% failure model, zero failures remain likely after 20 tests (81.8%) or 100 tests (36.6%). At least 299 tests give roughly a 95% chance of observing one or more failures.
Figure 5 · Estimating a mean and detecting a rare failure require different designs. In this independent 1% failure model, zero failures remain likely after 20 tests (81.8%) or 100 tests (36.6%). At least 299 tests give roughly a 95% chance of observing one or more failures.

For multi-goal agent evaluation, the useful output is often a vector: evidence for tool use, feedback retention, state updates, recovery, and quality-preserving efficiency. A high value on one axis does not fill a missing axis.

Coverage should have a defined unit. Let Ug be the set of audited behavioral conditions for goal g, and Eg(S) the conditions supported by acceptable execution evidence from selected cases S. A simple distinct-condition coverage is:

Cg(S) = |Eg(S) ∩ Ug| / |Ug|

Here the difficult work is defining and auditing the conditions. If a hundred variants all support the same condition, they increase repeated observations, not the number of distinct conditions. Report those counts separately. An empty or unassessed Ug is unknown coverage, not 100%.

CheckList’s capability-by-test-type organization offers a useful precedent.20 Stateful agents add the requirement that the condition must actually occur during execution. Label coverage, tool-call coverage, and interpretable behavioral coverage consequently answer different questions.

A Pareto frontier can preserve trade-offs between coverage, redundancy, quality, and cost when no suite wins on all of them. Its meaning is conditional on the objectives and candidate pool. The objectives themselves need validation against the evaluation decision.

08 / Design a case, then design the suite

Consider an illustrative task: update the correct order using the approval valid at a specified time, notify its owner, and leave other orders unchanged. The environment contains two similarly named orders, an earlier approval, and a later change message.

Calling this a “memory and planning” task is only the beginning. Its construction should specify the following.

Design layer Concrete requirement Acceptance check
Task contract Entity identity, cutoff time, approval precedence, recipient, prohibited changes All required information is available through allowed access
Initial state Resolve the relationship between old approval and later message Reset is reproducible; the conflict is real and resolvable
Valid solutions Batch queries and sequential queries are both permitted Alternative correct implementations pass
Wrong solutions Wrong order, stale approval, missing notification, extra notification Each targeted mistake is rejected
Challenge exposure Which versions were read, and when? Separate failure to reach the message from misuse after reading it
Response controls Change only the intended information availability Preserve feasibility; add matched edits and restoration
Place in the suite Routine workload, rare risk, or deliberate challenge? State its family, aggregation weight, and coverage unit

This brings the interpretation closer to observable behavior: “the agent read the new approval but used the old state” is more actionable than “the agent lacks memory.” A task that supports only end-to-end completion can remain valuable without receiving a finer diagnostic label.

At suite level, the construction order matters. Start with the decision and its goals; identify independent task families and behavioral conditions; assign weights from the intended use; construct and accept instances; use model responses to detect gaps and redundancy; then freeze the suite for independent evaluation. If diagnostic challenges and representative workload samples are both useful, retain both groups with their own interpretation.

Dynamic collection, as in Dynabench, can discover new failures.19 Retaining an anchor set alongside versioned challenge sets helps distinguish model progress from changes to the measuring instrument.

09 / Three claims worth testing

The argument becomes a research program when it specifies results that could prove it wrong. These are the three hypotheses I would prioritize; they remain to be tested on the target agent tasks.

Hypothesis Experiment Result that would weaken it
Production quotas materially affect comparisons Hold model outcomes and independent task content fixed; vary family multiplicities; compare ordinary means with declared family weights Rankings are stable, or the original frequencies already match the intended use
Score fidelity and goal fidelity can separate Compare score-preserving subsets on independently audited goals, near-model comparisons, and new model families Cheap score estimators retain all required evidence equally well
Response-assisted construction improves independent judgments Compare accepted, response-informed suites with random, stratified, structural, difficulty, and score-estimation baselines under matched information and total cost Simple baselines match it, gains stay within source models, or construction cost erases the benefit

For each comparison, freeze the evaluation goals, selectors, acceptance rules, and tolerances before the target run. Use independent evidence for the final judgment: a separate workload for model selection, confirmed behavioral defects for diagnosis, or new task families for transfer. Reusing the selection response as the sole success metric would only show that the optimizer optimized its objective.

Cost accounting should include task construction, model probing, and recurring evaluation. Report both equal-case-count and equal-total-cost comparisons. A costly selection stage can be worthwhile when reused many times; the break-even point is part of the result.

A stopping rule should answer the original question about “how many cases are enough”: enough for which goals, with what uncertainty, on which population? Additional cases become unnecessary for a declared use when their marginal improvement in independent goal satisfaction or decision quality falls below a prespecified tolerance. A flat development score alone cannot establish that condition.

The question every additional case should answer

A benchmark does more than collect tasks. It distributes attention among demands, turns those demands into observable events, decides which events count as success, and aggregates the results into a judgment.

That is why compression, construction, scoring, and ranking belong in the same discussion. If production quotas determine the weights, a compact suite may reproduce an arbitrary preference very efficiently. If task labels are unsupported by execution, broad nominal coverage may hide missing evidence. If the scorer rewards the wrong behavior, adding cases can make a biased judgment look more precise.

The constructive alternative is to make each link inspectable: justify the mix, verify the case, observe the challenge, test the scorer, and validate the resulting decision. For every additional case, ask which judgment it improves—and what new evidence makes that improvement possible.

Sources and figure data

Web sources checked October 8, 2026. Figures 1, 4, and 5 are analytic examples; Figures 2 and 3 redraw public data. Data and sources.

  1. OpenAI. Introducing GPT-6.1 Sol. Model release report, 2026.
  2. Anthropic. Introducing Claude Sonnet 5.5. Model release report, 2026.
  3. Google DeepMind. Gemini 4 Argon: Model evaluation. Official evaluation methodology, 2026.
  4. Rein et al. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. 2023 preprint; COLM 2024.
  5. Datacurve. DeepSWE v1.1. Official project, v1.1.
  6. Zapier. AutomationBench. Official project, leaderboard v1.0.6.
  7. Trivedi et al. AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. ACL 2024; original 750-task release.
  8. Lu et al. ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities. Findings of NAACL 2025.
  9. Yuan et al. OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. 2026 preprint and project; configuration specified in Figure 3.
  10. Arora et al. HealthBench: Evaluating Large Language Models Towards Improved Human Health. 2025 preprint; original release.
  11. Li et al. AutoBencher: Towards Declarative Benchmark Construction. ICLR 2025.
  12. Bean et al. Measuring what Matters: Construct Validity in Large Language Model Benchmarks. NeurIPS 2025 Datasets and Benchmarks; Section 2.
  13. OpenAI. Why we no longer evaluate SWE-bench Verified. February 23, 2026; selected audit of 138 difficult tasks.
  14. Shao et al. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI. 2026 preprint; causal-discovery example in Case 6.
  15. Wang et al. Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks. 2026 preprint.
  16. Maia Polo et al. tinyBenchmarks: evaluating LLMs with fewer examples. ICML 2024.
  17. Vivek et al. Anchor Points: Benchmarking Models with Much Fewer Examples. EACL 2024.
  18. Yauney, Warraich, and Swayamdipta. How Reliable is Language Model Micro-Benchmarking?. ICLR 2026.
  19. Kiela et al. Dynabench: Rethinking Benchmarking in NLP. NAACL 2021.
  20. Ribeiro et al. Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. ACL 2020.
  21. Dehghani et al. The Benchmark Lottery. 2021 preprint.
  22. Datacurve. DeepSWE: Measuring frontier coding agents. Construction and quality-assurance methodology, May 26, 2026.