ResolveBench runs frontier AI agents through realistic, audited support workflows (authenticate, investigate, apply policy, execute the right tool, leave the database correct) and grades them on the actual end-state, not what they claim. Then it shows exactly why they fail. On the hardest cases, 57% of failures reached the correct, safe resolution and still failed, on a skipped check, a missing evidence ID, a prohibited action.
On the hardest cases, 57% of failing runs reached the correct, safe resolution, then failed the bar anyway: a skipped verification read, a missing evidence ID, a prohibited action. So we score the same 6,400 runs two ways. Outcome asks only "did the database end up correct?" Operational asks "correct and done by the book", the ResolveBench bar. The gap between them is reliability that single-run scores never see.
Score on the final outcome alone and pass⁸ jumps by half to nearly double; every failure in that gap got the answer right but the process wrong. That gap, not the leaderboard, is what breaks agents in production.
pass¹ asks "can it ever get this right?"; pass⁸ asks "does it get this right every time?" The drop between them is the reliability tax single-run benchmarks never charge.
Reliability decay: pass^k as k rises from 1 to 8. Hover any point for its value; click a model to toggle it.
Scored at a strict 90/100 composite bar and ranked by pass⁸: reliability, not average quality. Safety is a per-task hard-fail dimension. Reasoning effort is labeled per model: high runs think harder.
Score = mean composite quality 0–100 (a separate axis from reliability) · pass¹ = chance one run succeeds · pass⁸ = solves on all 8 runs · Safety = share of tasks with zero prohibited-tool use (n = tasks with a violation) · spark = pass¹→pass⁸ decay.
Every run's full trajectory is public: golden plan beside the agent's actual tool calls. Across 6,400 runs the failures sort into four mechanisms, and almost none of them are "didn't know the answer."
112 runs across 21 tasks fired a prohibited tool: escalating, refunding, or rebooking when the right move was simply to explain. Doing too much fails the case exactly like doing too little.
The decision is correct but unsubstantiated. The typical miss cites just one of three required IDs; Evidence or tool-use is the weakest competence dimension for every model (best Evidence 4.17/5, Kimi K2.6 bottoms at 2.41).
Skipped verification and authentication. 48.8% of runs miss a required tool call (vs 7.4% redundant), and a run that hits the auth gate without authenticating first fails 73% of the time.
The escalation decision is well-calibrated: 0 of 1,856 trials on escalation tasks substituted a forbidden action. What breaks is completeness: the handoff skips the verification reads and evidence a real one requires.
Not all failures are equal. Here's what actually breaks each config's failing runs: the share caused by a wrong end-state, a missing evidence ID, skipped authentication, other tool-use errors, or an over-action. Almost everything is process, not the answer.
The published benchmark covers airline, hotel, and utility support. The methodology is domain-agnostic: the same end-state scoring, operational checks, and audited goldens apply anywhere, so we're expanding into the workflows enterprises ask about most.
We build audited tasks for your support workflows, evaluate any model on them, and help you close the reliability gap.