Results¶
Everything on this page is measured on our own benchmark. Where a result is a null, it is reported as a null. Nothing here uses a precomputed leaderboard number.
The short version: cheap-first allocation with verified escalation reaches
always-frontier quality for a fraction of the cost, and — new this cycle — it
now does so on a cache-safe strategy you can select by name
(router.strategy: session_cascade) rather than only on a blocked one.
Session-Cascade, which makes one decision per session, costs $23.40
cache-aware at 97.24% on the 181-task scoring path, against Price-Cascade's
$22.27 at the same pass rate and Always-Frontier's $94.37 at 95.03%. That row
is what a default install runs. Its sibling kNN-semantic-cascade
(router.strategy: knn_semantic_cascade, the opt-in routing model) opens the same ladder
on the kNN pick rather than on the cheapest model, and reaches the same 97.24%
for $26.88 cache-aware — on this corpus, dominated. Read that verdict with
the caveat it comes with: the shipped default, and the routing model priced
against it.
On the harder, fully-measured 74-task set Session-Cascade
costs $28.76 at 90.91% against Always-Frontier's $37.63 at 86.36%. Details,
both sets, and every interval: Routing at session
cadence. Every cascade number on this page is an
offline replay whose rungs each start from a fresh tree and a fresh context;
live they do not — the divergence, stated
once.
Two things temper it. The margin is much smaller on measured cells than on imputed ones — the four-fold saving above is largely an artifact of imputation, and we now show both bases side by side rather than let you find that yourself. And the machine learning still contributes nothing: the saving comes from mechanism, not prediction. The escalation signal is real but sits in the recurrence policy at the shipped threshold once the reproduction phase is excluded, not in the prefix risk model (which reads no skill).
We do not claim the project's make-or-break gate is passed — that gate is written about the owner's own coding-agent workflow, and everything here is SWE-bench. The reasons are in Routing results and This is not the make-or-break gate.
How to read this page¶
Two models, two different jobs. Shunt learns two separate things. Conflating them makes every number ambiguous, so they stay apart throughout.
| Routing model | Escalation model | |
|---|---|---|
| Question | Which (model, effort) should start this task? | Is this attempt going to fail, escalate now? |
| When | Once, at the task boundary, before any tokens are spent | Mid-task, at decision boundaries |
| Input | The task text, embedded | The running trajectory: discussion, tool use, verified check results |
| Learns from | Task outcome, pass/fail | Whether this attempt ultimately failed |
| Today | k-nearest-neighbours over task embeddings | A recurrence rule over verified failing-check ids |
| Status | No measurable signal over the base rate yet | Policy: OK_OFFLINE_ONLY — real edge at the shipped threshold once the reproduction phase is excluded (eval-only); prefix model: NO_SKILL |
| Next | bigram / linear models, calibrated classifiers, better selection rules | calibrated risk scoring, structural loop features, late fusion |
Rank and cost are different orderings. We have conflated them ourselves, so, precisely:
- Cost is the dollars a task actually consumed:
real_cost, cache-aware, read from the provider's own usage accounting. Never a price-list estimate. - Rank is a model's position in the registry, which is ordered by price. It is not a capability ordering.
The two come apart, which is the whole reason routing might be worth doing:
| model | price ($/Mtok) | measured pass rate | 95% CI |
|---|---|---|---|
| deepseek-v4-flash | 0.42 | 68.9% | 0.620–0.751 |
| deepseek-v4-pro | 1.30 | 85.0% | 0.794–0.893 |
| qwen3.7-plus | 1.60 | 43.7% | 0.337–0.541 |
| gpt-5-mini | 2.25 | 54.5% | 0.476–0.613 |
| kimi-k2.5 | 3.60 | 49.6% | 0.408–0.584 |
| glm-5.2 | 5.80 | 57.1% | 0.465–0.672 |
| kimi-k3 | 18.00 | 84.5% | 0.766–0.901 |
deepseek-v4-flash costs 5× less than gpt-5-mini and solves more. Price does
not buy capability monotonically. A model is only worth its price if it earns it
on your tasks. (These rates are each model's marginal rate over the tasks it
actually ran; coverage is adaptive, so they are not cross-comparable at face
value. On the tasks where they co-ran (coverage varies by model pair), the ordering
barely survives — and the cheapest model is the most flattered by adaptive
coverage. Below is the historical "common 74" table from the prior run (a fixed
subset for reference); the current run's coverage varies:
| model | pass rate, own coverage | pass rate, common 74 | pooling bias |
|---|---|---|---|
| deepseek-v4-flash | 68.9% | 44.6% | +24.3pp |
| gpt-5-mini | 54.5% | 35.1% | +19.4pp |
| glm-5.2 | 57.1% | 51.4% | +5.8pp |
| kimi-k2.5 | 49.6% | 43.2% | +6.3pp |
| kimi-k3 | 84.5% | 77.0% | +7.5pp |
| qwen3.7-plus | 43.7% | 41.9% | +1.8pp |
That 24.3pp is over four times the 5pp non-inferiority margin the kill gate is judged at, so read any single-model rate as a marginal over its own task subset, never as a comparison.)
The dataset¶
The benchmark is a set of challenges: SWE-bench Verified tasks we run ourselves. Each challenge is attempted by multiple experiments, one per (model, reasoning-effort) arm, and each runs to a verified pass or fail judged by the task's own tests.
One run yields two distinct data products:
- A pass/fail label per (challenge, arm), which supervises the routing model.
- A full trajectory log (discussion, tool calls, verified check results), which supervises the escalation model.
Both come deterministically from real agent runs, so both models can be re-evaluated offline against data already on disk. That is what makes iteration cheap: a new routing rule or a new escalation detector can be tested without spending another cent. Both are supervised learning problems. The benchmark is the label factory.
The assumption that fills the gaps¶
Running every arm on every challenge is expensive, so unmeasured cells are filled under a monotone capability ladder: above a success is success, below a failure is failure.
Our own paired test does not confirm it. Across 478 co-measured within-model arm pair-observations, more reasoning effort is worth +1.7pp, exact McNemar two-sided p = 0.428, which is indistinguishable from no effect. Monotonicity is violated outright on 7.3% of those pairs (35 of 478): the higher-effort arm failed a task the lower-effort arm passed. All nine plotted arm pairs have intervals straddling zero.
We flag this rather than lean on it, because it is load-bearing. It is what
fills every unmeasured cell, and every imputed cell is filled pass=True.
Routing results¶
The embedding numbers below have now been re-measured, and they got worse. The kNN
strategies used to embed the manifest description — <repo>@<commit12> - resolve
<test-node-id>, median 106 characters — while the agent was handed the upstream
problem_statement. The router and the work it routed never saw the same text. The
manifest has been rebuilt with the real statements (median 1185 characters),
routing_text() prefers them, and every row below is recomputed on that basis.
Routing quality fell: kNN went from 81.71% to 77.72%, which is inside noise of Always-Cheap's 75.54%, and kNN-semantic-tier from 67.43% to 65.76%. The 106-character label was not merely uninformative — it was mildly leaky, because the repo name it carried is a weak proxy for task difficulty. Given the correct input, the learned router is not distinguishable from the trivial policy. The zero-ML rows (Oracle, Price-Cascade, Always-Cheap, Always-Frontier) use no embeddings and are unchanged.
Seven router-selection strategies — each one a rule for picking the model a
task starts on — scored on the same 181 tasks (19 unscorable), from
strategy_summary.csv.
Two further scored strategies, Session-Cascade and kNN-semantic-cascade, are not
selection rules at all: they model the escalation layer over whatever base
routing chose, so they get their own
section rather than rows here. Both are selectable
— router.strategy: session_cascade is that layer over an always_cheap base
and is the shipped default, and router.strategy: knn_semantic_cascade is the same
layer over the kNN pick, which you opt into. Costs in this
table are naive per-task sums, cache-blind:
| strategy | passes | pass rate | 95% CI | total cost | avg cost/task | cumulative regret |
|---|---|---|---|---|---|---|
| Oracle (hindsight, not deployable) | 176 | 97.24% | 94.48–99.45 | $14.75 | $0.0815 | 0.00 |
| Price-Cascade (blocked, not deployable) | 176 | 97.24% | 94.48–99.45 | $22.27 | $0.1230 | 0.75 |
| kNN-semantic-cascade (within-task) (blocked, not deployable) | 176 | 97.24% | 94.48–99.45 | $25.47 | $0.1407 | 1.07 |
Session-Cascade (strategy: session_cascade — the shipped default) |
176 | 97.24% | 94.48–99.45 | $28.21 | $0.1559 | 1.35 |
| kNN-difficulty-cascade (blocked, not deployable) | 176 | 97.24% | 94.48–99.45 | $28.51 | $0.1575 | 1.38 |
| Difficulty-Band-cascade (blocked, not deployable) | 176 | 97.24% | 94.48–99.45 | $28.51 | $0.1575 | 1.38 |
| Always-Frontier | 172 | 95.03% | 91.71–97.79 | $94.37 | $0.5214 | 11.96 |
| kNN-semantic | 143 | 77.72% | 71.20–83.70 | $11.79 | $0.0641 | 34.35 |
| Always-Cheap | 139 | 75.54% | 69.02–81.52 | $1.50 | $0.0081 | 37.32 |
| kNN-difficulty (control — never shippable) | 139 | 75.54% | 69.02–81.52 | $1.80 | $0.0098 | 37.35 |
| kNN-semantic-tier (blocked, not deployable) | 121 | 65.76% | 58.70–72.28 | $11.53 | $0.0627 | 56.32 |
Only Always-Frontier and Always-Cheap in this table name a strategy the router
will accept — LIVE_STRATEGIES in src/shunt/router/policy.py is the allowlist, and
router.strategy is validated against it at boot. kNN is not one of them: it is
the selection rule with the escalation ladder removed, which no configuration produces,
and it is kept as the control that isolates what the ladder buys. The selectable
knn_semantic_cascade is that same pick with the ladder — the opt-in routing strategy, not the
shipped default, which is session_cascade and never queries the neighbourhood at all. The remaining rows are real
measurements of things you cannot configure; benchmark/routing/strategy_class.py
carries each one's blocker and its path to live.
The result that matters is which strategy gets there¶
Price-Cascade uses no embeddings, no nearest neighbours, and no training. It
tries the models in ascending price order and stops at the first one whose patch
passes. Setting aside the hindsight Oracle, it is the cheapest row in the table
whose quality interval overlaps Always-Frontier's — and it is not purchasable
today. Stopping at the first passing patch requires a verified outcome
mid-session; that is more than one decision per session and it breaks
cache-safety, so price_cascade is rejected at boot. The $22.27 @ 97.24%
operating point measures a mechanism, not a product capability.
The learned kNN-semantic-cascade (within-task) costs more ($25.47 against $22.27) for the same
97.24%, and is blocked on the same cache-safety ground. The machine learning is
not paying for itself. What buys the quality back is verified escalation —
which is why the shipped router carries an escalation ladder rather than a
cascade, at a lower ceiling (see escalation).
Even the regret ordering is unresolved: Price-Cascade's bootstrap interval on total regret is [0.46, 1.08] and kNN-semantic-cascade (within-task)'s is [0.73, 1.44]. They overlap.
Measured versus projected¶
That table is still part projection. Of the dollars behind it:
| strategy | projected share of cost | projected passes |
|---|---|---|
| Always-Cheap | 3.3% | 10 of 139 |
| Oracle | 3.9% | 14 of 178 |
| kNN-semantic-tier | 24.6% | 36 of 121 |
| Price-Cascade | 27.7% | 25 of 178 |
| kNN-semantic-cascade (within-task) | 28.1% | 31 of 178 |
| kNN-semantic | 35.7% | 25 of 143 |
| Always-Frontier | 44.9% ($43.15 of $96.02) | 89 of 175 |
Every projected cell is filled pass=True. Imputation is not neutral: the
always-frontier baseline is charged full price on tasks a cheaper model
demonstrably solved, which is exactly where a router's apparent saving comes
from.
So we cut to the subset where no cell on either path was projected. On those 87 co-measured tasks:
- Always-Frontier: $46.65 @ 92.0%
- Price-Cascade: $15.61 @ 93.1%
Price-Cascade is $31.03 cheaper, with 1 versus 0 discordant task, McNemar exact p = 1.000. The saving survives measurement, with no quality difference resolved either way.
That 87-task cut is pairwise: it keeps the tasks where the two strategies being compared both landed on measured cells, so it admits a different task set for every pair. A stricter basis — every strategy measured on the same tasks — now exists, and on it the saving is much smaller than a third of the baseline. See the correction.
Why we still do not call the gate passed¶
The gate's own emitted verdict is UNTESTED and its provisional read is a fail — that is stated in full below. Three further reasons specific to the comparison on this page, which we would rather state than have you find.
- The measured subset is opportunistic. Those 87 tasks are what survived after removing every projected cell, not a pre-registered sample. Closing this with a designed run is the top priority.
- The two quality figures are not the same kind of number. A cascade stops at the first attempt whose tests pass and is then scored on that same label, so 92.9% is a best-of-N coverage statistic while Always-Frontier's 91.7% is single-shot. We flagged exactly this pattern as a flaw in published work; it applies to us too. The cost axis is honest, because every attempt in the chain is billed. The quality axis flatters any retry strategy, ours included. We do not yet know how to remove this.
- The stopping oracle is not free in production. In the benchmark, "did it pass" comes from the task's own test suite. On your machine that signal comes from your tests, which is the whole premise, but it is a real dependency.
The plain kNN strategy is weaker still¶
It buys about 2.2pp of pass rate over always-cheapest for roughly 7.9× the cost, and that margin sits far inside both intervals: [71.20, 83.70] against [69.02, 81.52]. Indistinguishable at this sample size.
Routing at session cadence¶
Both cascades above are cheap because they retry inside a task and read a verified outcome after every attempt. That is precisely what the router rejects at boot. So the interesting question was never "how cheap is a cascade" — it was how much of that survives when the ladder is paced the way Shunt actually runs it.
Session-Cascade answers it. One decision per session; the effort rung first
(same model, higher reasoning — the prompt cache survives), the rank rung only
once effort is exhausted; the climbed rank persists into the next session; every
attempt is billed. Nothing switches inside a cached turn, so it is cache-safe by
construction rather than by policy.
It is the row you get. strategy: session_cascade, escalation.enabled: true
and escalation.rank_shortlist: 3 are all defaults in
src/shunt/config/router.yaml, so this operating point is what a default install
runs, under its own name:
router:
strategy: session_cascade # = always_cheap + the escalation ladder
capture:
work_dir: /path/to/your/repo # without a repo there is no verified failure, so nothing climbs
benchmark/routing/strategy_class.py now classifies the row LIVE, derived
from the product's own LIVE_STRATEGIES rather than restated. It was previously
BLOCKED on a narrow technicality — that no router.strategy value named a
layer — which was true and also read, in every figure, as "you cannot run
this". Registering the preset removes the technicality instead of hedging around
it. Note what stays blocked: Price-Cascade and kNN-semantic-cascade (within-task) verify inside
one task, and no configuration will ever do that here.
What the row does not say is whether its rungs are worth buying. It prices the ladder; it does not score the models the ladder steps to. Those are measured one rung at a time, against the cheap base model, in the ladder-rungs figure — and on this corpus the shipped ladder now steps the rung the measurement supports (glm-5.2) and still skips the best-measured rung (kimi-k3) on a price slot. Read the two together.
The shipped default, and the routing model priced against it¶
Both selectable cascades climb the same ladder; they differ in one thing, the rung the
first session opens on. session_cascade opens it on the cheapest model.
knn_semantic_cascade — the opt-in routing strategy — opens it on the kNN pick. Both rows below
are replayed through the same EscalationRunner, the same rank floor and the same billing
loop, so the difference between them is the opening rung and nothing else.
| strategy | passes | pass rate | 95% CI | naive cost | cache-aware cost | avg cost/task | cumulative regret |
|---|---|---|---|---|---|---|---|
Session-Cascade (strategy: session_cascade, the default) |
176 | 97.24% | 94.48–99.45 | $28.21 | $23.40 | $0.1592 | 1.35 [0.81, 1.95] |
kNN-semantic-cascade (strategy: knn_semantic_cascade, opt-in) |
176 | 97.24% | 94.48–99.45 | $31.66 | $26.88 | $0.1749 | 1.69 [1.13, 2.29] |
On this corpus the routing model is dominated by the cheap start. Both reach 97.24% — the ladder walks to a model that solves the task either way — and opening on the kNN pick instead of on the cheapest model costs $3.48 more cache-aware ($26.88 against $23.40) for no measured quality. That is the same verdict the selection-rule rows already carry, arriving by a second route: the embedding buys nothing here the ladder does not already deliver, and it pays for models the ladder would have skipped. It is also why the default is the cheap start, and why the routing model this project is named for is something you switch on. Read it as a difference of totals on one corpus, not as a paired test — the two rows were not contrasted pairwise, and the frontier figure is where the domination is drawn.
Why that number probably flatters the cheap start — an argument, not a result. Both
rows rest on the same modelling assumption spelled out below: results.csv carries no
failing-check identity, so the replay gives each task one stable key and every failure
recurs identically. That produces the fastest climb the policy could ever make — one
recurrence per rung, immediately — and a fast climb is worth far more to the row that
starts at the bottom and has the most rungs to walk. Live, the counter needs two
confirmed same-key failures inside a ten-decision window, and it only ever steps at a
session boundary, so a real climb is slower and every rung it wastes is a whole failed
session of your time. That is exactly the cost a better first pick would avoid, and it is
the part this corpus cannot see: nothing here measures wall-clock, or a climb paced by a
real recurrence counter. Sessions themselves are counted — sessions_mean and
sessions_p95 are published per strategy and the kill gate judges on both — but a session
count is a count of handoffs, not a clock, and it says nothing about how long each one took. We publish the dominance because it is what we
measured. We do not claim it settles the question in production, and we have not
quantified the gap — this paragraph is an argument for why the offline number is the
cheap start's best case, not a correction to it.
Two things the knn_semantic_cascade row does not say. It is not pre-registered: the 5pp
non-inferiority gate named the bare kNN selection rule as its verdict arm before any of
this was measured, and that arm is deliberately left where it is — repointing it after
seeing the data would rewrite the registered test. The
kill-gate figure therefore carries both, the pre-registered arm
and the shipped default, with the labels saying which is which; and the rename exposed
that the pre-registered arm adjudicates a configuration no operator can select, which is a
pre-existing defect rather than one this cycle created. Both rows also carry the
lower-bound disclosure in the next paragraph, and both are offline replays from a fresh
tree — what live does differently.
One modelling assumption, stated up front. The live counter escalates on
recurrence of the same normalized failing-check id, and results.csv records a
per-cell pass/fail with no failure identity. The replay therefore gives each task
one stable key, so every failure of a task recurs identically. That is the
assumption most favourable to escalation — the ladder climbs as fast as the
policy ever could — which makes every cost below a lower bound on live spend,
never an optimistic one.
The two task sets, and why totals do not cross between them¶
The result is reported on two bases, chosen because they are biased in opposite directions.
Read this before comparing anything. Set A has 184 tasks and set B has 74 (the count of tasks where all six enabled models were measured, re-derived from
benchmark/routing/results.csv). Total costs are sums over tasks, so a set-A total and a set-B total are not comparable. Compare only within a set. Pass rates and orderings do carry across; dollar totals do not.
Set A — the shipped scoring path, coverage-completed, 184 tasks. 35% of its cells are monotone-imputed. Its subset guard, verbatim:
scored on 184/200 tasks selected by coverage; deepseek-v4-flash passes 74.1% here vs 12.5% on the 16 dropped (+61.6pp) — difficulty-biased, not a random sample
| strategy | pass rate | 95% CI | naive cost | naive 95% CI | cache-aware cost |
|---|---|---|---|---|---|
| Oracle (hindsight — bound) | 96.74% | — | $18.33 | [11.05, 27.18] | $18.33 |
| Price-Cascade (blocked) | 96.74% | 94.02–98.91 | $27.11 | [17.92, 37.81] | $27.11 |
Session-Cascade, rank_shortlist=3 (strategy: session_cascade) |
96.74% | 94.02–98.91 | $33.56 | [22.53, 46.14] | $28.71 |
| kNN-semantic-cascade (within-task) (blocked) | 96.74% | 94.02–98.91 | $30.44 | [20.39, 42.06] | $30.44 |
Session-Cascade, rank_shortlist=0 (pre-shortlist) |
96.57% | 93.71–98.86 | $48.19 | [29.32, 69.57] | $35.79 |
| Always-Frontier | 95.11% | 91.85–97.83 | $96.02 | [88.73, 104.19] | $96.02 |
| kNN-semantic (control — the pick without the ladder) | 77.72% | 71.20–83.70 | $11.79 | [7.80, 16.13] | $11.79 |
| Always-Cheap | 75.54% | 69.02–81.52 | $1.50 | [1.31, 1.71] | $1.50 |
| kNN-semantic-tier (blocked) | 65.76% | — | $11.53 | [8.85, 14.62] | $11.53 |
Set B — raw, un-imputed, fully-measured tasks only: 74 scorable. Every
cell here was actually run; the scorable count is re-derived from
benchmark/routing/results.csv (tasks where all six enabled models were measured),
never a hardcoded number. Its subset guard, verbatim:
scored on 66/74 tasks selected by coverage; deepseek-v4-flash passes 50.0% here vs 0.0% on the 8 dropped (+50.0pp) — difficulty-biased, not a random sample
| strategy | pass rate | 95% CI | naive cost | naive 95% CI | cache-aware cost |
|---|---|---|---|---|---|
| Price-Cascade | 90.91% | 83.33–96.97 | $27.10 | [19.49, 36.56] | $27.10 |
Session-Cascade, rank_shortlist=3 (strategy: session_cascade) |
90.91% | 83.33–96.97 | $33.75 | [24.39, 44.23] | $28.76 |
Session-Cascade, rank_shortlist=0 |
90.91% | 84.85–96.97 | $66.89 | [47.66, 89.68] | $49.34 |
| Always-Frontier | 86.36% | 80.30–93.94 | $37.63 | [30.57, 45.58] | $37.63 |
| kNN-semantic-cascade (within-task) | 90.91% | 84.85–95.45 | $34.09 | [24.40, 46.95] | $34.09 |
| kNN-semantic | 54.55% | 42.42–66.67 | $12.34 | [6.45, 18.85] | $12.34 |
| Always-Cheap | 49.25% | 37.31–62.69 | $0.72 | [0.57, 0.88] | $0.72 |
Note the guards run in opposite directions. Set A drops the tasks the cheap model almost never solves, so it is biased easy; set B drops tasks the cheap model never solves, from a pool where it solves only half, so it is biased hard. Always-Cheap reads 75.54% on one and 49.25% on the other, which is how far apart the two selections are. The ordering survives both. That is the load-bearing point — not either set's dollar figure.
Cache-aware against naive¶
Both cost columns are reported because they are different quantities, and only one of them is the bill.
Session-Cascade is the only strategy in the two Set tables below the
cache term moves — the comparison table above reports single, cache-blind totals
only. That is not an accident of the model: among the strategies those tables
carry, Session-Cascade is the only one that re-serves the same model on
consecutive attempts, because its first escalation rung raises reasoning effort
rather than rank. Every other row there either never retries or steps to a
different model on each attempt, which forfeits the cached prefix, so its naive
and cache-aware totals are the same number. (The two difficulty cascades —
kNN-difficulty-cascade, Difficulty-Band-cascade — also re-serve one model and
therefore move under the cache term too; they are absent from those tables, and
their numbers are in routing.md.)
The size of the term tracks how much same-model repetition each configuration
does. At rank_shortlist=3 it removes 14% of the naive total on set A ($33.56 →
$28.71) and 15% on set B ($33.75 → $28.76). At rank_shortlist=0, which walks
every rank instead of jumping over the shortlist, the ladder is longer, repeats
more, and the term removes 26% and 26% respectively — a bigger discount on a much
worse total.
Caching is scoped per task. A task is one session, and the discount applies to repeated attempts within it. Summing a whole run as a single cached sequence would be wrong: consecutive tasks are different sessions with different prefixes, and charging them at the cache-read rate would invent a discount no provider offers. The per-model discount and input share behind these numbers are measured from the registry and the corpus; the hit rate is assumed at 90% and flagged as assumed wherever it is reported — see cache economics for how far that assumption can move the figure.
What the paired test says¶
Total-cost intervals in the tables above are per-strategy and overlap freely, which resolves almost nothing. The paired per-task bootstrap does better: it compares strategies on the same task, so the task-difficulty variance that inflates those intervals cancels. Set A, cache-aware, 2000 draws over the 175 shared tasks:
| comparison | Δ cost | 95% CI | reading |
|---|---|---|---|
sl=3 vs Price-Cascade |
+$1.37 | [+0.82, +2.00] | real, and small |
sl=2 vs Price-Cascade |
+$0.76 | [−0.06, +1.85] | not distinguishable |
sl=3 vs kNN-semantic-cascade (within-task) |
−$1.23 | [−3.42, +0.73] | not distinguishable |
sl=3 vs sl=0 |
−$13.30 | [−21.00, −6.42] | the shortlist pays |
sl=3 vs Always-Frontier |
−$66.12 | [−74.19, −57.90] | the headline |
sl=2 vs sl=3 |
+$0.61 | [−0.42, +1.47] | not distinguishable |
So: paying one decision per session instead of one per attempt costs $1.37 over
Price-Cascade on this set, and buys cache-safety and a mechanism you can
actually enable. Against kNN-semantic-cascade (within-task) the two are not distinguishable. Against
Always-Frontier the gap is large and the interval is nowhere near zero.
The last row is why the shipped default stays at 3. rank_shortlist=2 is not
distinguishable from 3 against either Price-Cascade or 3 itself, so there is no
evidence to move the default. We are not tuning on a difference we cannot
resolve.
Quality is a separate axis and it resolves less. On set A, Session-Cascade and
Price-Cascade read an identical 96.74%, and Always-Frontier's 95.11% sits
inside every one of those intervals. On set B the shipped ladder's 90.91%
[83.33, 96.97] point estimate is above Always-Frontier's 86.36% [80.30, 93.94],
but the intervals overlap heavily: the two are not distinguishable on quality,
and we do not claim the ladder passes more tasks. What set B shows is a cost
result at quality that is not resolved as different — and note that no paired
bootstrap was run on set B, so its cost gap is weaker evidence than set A's.
The correction: the saving is much smaller on measured cells¶
This is the part we held back until this measurement existed, and it revises what this page used to lead with.
Read the two sets against each other on the same comparison:
| basis | imputed share | Price-Cascade vs Always-Frontier | Session-Cascade sl=3 (cache-aware) vs Always-Frontier |
|---|---|---|---|
| Set A (184 tasks) | 35% of cells | $27.11 vs $96.02 — 72% cheaper | $28.71 vs $96.02 — 70% cheaper |
| Set B (74 tasks) | none, 100% measured | $27.10 vs $37.63 — 28% cheaper | $28.76 vs $37.63 — 24% cheaper |
A four-fold saving becomes roughly a quarter. Both bases put the cascade family
below the fixed-frontier baseline, so the direction holds on measured data —
but the magnitude is mostly imputation, and it collapses in the direction you
would expect once you see the mechanism: every imputed cell is filled
pass=True, which charges Always-Frontier full price on tasks a cheaper model
demonstrably solved. That fill is exactly where a router's apparent saving comes
from, and set B removes it.
Set B is also a stricter basis than the 87-task subset in Measured versus projected above. That one is pairwise co-measured — it keeps tasks where the two strategies being compared both landed on measured cells, which is a different, more permissive filter per comparison. Set B requires every strategy's cells to be measured on the same tasks, so all seven rows are scored on one honest basis. Where the two disagree about how much survives, prefer set B.
We are stating this plainly rather than reporting the flattering basis and letting a reader discover the other one.
This is not the make-or-break gate¶
It is strong evidence toward it, and it is not it.
What the gate actually asks. It used to ask one thing of one baseline: beat fixed-frontier-with-caching on cost, at equal quality. It now asks, in order:
- Quality is a gate, not a tradeable axis. Paired non-inferiority within 5 percentage points. A router inferior beyond the margin fails whatever else it wins, and a quality win buys nothing — the claim under test is "the cheapest model that can do the job", so crediting quality would let any router pass by escalating more.
- Then tolerance-aware Pareto dominance over four operational axes — the four distinct things a user pays: cache-aware cost, mean sessions, p95 sessions, and the coefficient of variation of per-task cost. All four are minimised. The router must be no worse on any of them and strictly better on at least one, at one relative tolerance of 5% shared by all four. One tolerance rather than a table per axis, because a per-axis knob is a per-axis place to move a result after seeing it; and it cuts both ways, turning wins into ties as readily as losses.
- Against two baselines, independently.
Always-Frontier, the fixed-frontier-with-caching arm — andPrice-Cascade, a zero-ML constant policy: cheapest first, escalate on a verified failure, the same handoff ladder, and no routing decision anywhere in it. Clearing one and losing the other is a failure, not an average.
The second baseline is the one that matters. Published work on agentic-coding routers has run the decisive ablation — always sending every task to one cheap strong model, with the same handoff, matched a learned router at the same cost per solved task — and the old single-baseline gate could not express that falsifier at all. Adding it made the bar harder in exactly the direction the router was winning. The full criterion, and the positive control and null the gate itself had to clear before any of its verdicts could be quoted, are in Benchmark design.
And it is asked of the owner's own coding-agent workflow. Everything above is SWE-bench Verified — a different corpus, a different task distribution, and a harness where the pass signal is the task's own test suite rather than a working repository's. A result on SWE-bench does not discharge a gate written about day-to-day agent traffic, and we are not going to let the two blur together because the numbers came out well.
The gate's current verdict is UNTESTED, and its provisional read is not a pass¶
The gate has been run on the committed corpus. Its emitted verdict is UNTESTED: not "no result yet" and not "a weak result", but not adjudicable. The coverage precondition it inherits from the offline gate is tripped — the baseline arm's cells are 51.6% measured and the router arm's 83.7%, against a floor of 90% — and below that floor the gate refuses to emit PASS or FAIL at all, because a verdict drawn from a corpus that is more than a tenth filled in by assumption would be a statement about the imputation, not about the router.
Underneath that refusal the gate still records what the arithmetic says, labelled
provisional. On kNN-semantic-cascade, at the 5% tolerance:
| Axis | vs Always-Frontier |
vs Price-Cascade |
|---|---|---|
| quality (5pp non-inferiority) | passes, +1.63pp | passes, level |
| cache-aware cost | better, $38.49 against $96.02 | worse, $38.49 against $27.11 |
| mean sessions | worse, 2.076 against 1.000 | worse, 2.076 against 1.592 |
| p95 sessions | worse, 7 against 1 | worse, 7 against 4 |
| cost variability (CV) | worse, 1.894 against 0.558 | better, 1.894 against 2.533 |
| dominance | no | no |
Read plainly: the router clears the quality bar against both baselines and buys its money win with round trips. Against the frontier baseline it is cheaper and worse on every other axis a user pays. Against the constant policy — the one with no routing decision in it — it is worse on cost and sessions and the session tail, and its only win is a steadier bill. The provisional read is therefore a FAIL against both baselines, and the binding falsifier on today's evidence is not the expensive frontier model. It is the policy with no model-choosing in it.
That provisional read may not be quoted as a result, and we are not quoting it as one. It is what the criterion computes on a corpus the criterion itself says it cannot adjudicate; publishing it is how we avoid the alternative, which is holding a number back until it improves. What resolves the UNTESTED is coverage — a designed, measured run that lifts both arms above the floor — not a better router. Until then the honest statement is the narrow one: the gate is untested, and nothing here is evidence that it would pass.
Three further reasons the gate stays open, unchanged by this measurement:
- Both task sets are coverage-selected, and both guards say so. Neither is a pre-registered random sample; they are biased in opposite directions, which is why the ordering surviving both is the claim rather than either total.
- The quality axis still flatters any retry strategy. A ladder scored on whether some attempt passed is a best-of-N statistic; Always-Frontier's is single-shot. The cost axis is honest because every attempt is billed. The quality axis is not, and we still do not know how to remove that.
- The shipped router's own selection still misses the bar. The paired non-inferiority test on the kNN router against fixed-frontier is inferior by more than the pre-registered 5pp margin on all three evidence bases — see the kill-gate figure. The ladder above sits over base routing; it does not repair it.
Why the learned router adds nothing¶
Leave-one-task-out routing pass rate for the kNN router adds nothing over the base rate: at k=2 it is 0.7880, against Always-Cheap's 75.54%. It sits inside the shuffled-outcome permutation null band of [0.7663, 0.8315] (null mean 0.7946, 200 permutations).
Three more readings point the same way:
- Against a fixed model it loses. The best single always-one-model policy on this suite scores 0.9511. The router is 0.1631 below it. A router that cannot beat one fixed model is not routing.
- Its neighbourhoods are no better than random. Observed mean neighbourhood purity 0.919 sits inside a permutation null of [0.8839, 1.0000]. True chance purity is 0.9128 and the majority class alone is 0.9543, so a purity near 1.0 is what a constant router scores for free.
- It is close to a constant function. At k=20 the router sends 167 of 175
tasks to
deepseek-v4-flashand 8 toqwen3.7-plus, differing from always-cheapest on 8 tasks, for +$0.32 (+23%) and +1 pass.
Difficulty itself is present in the data (tasks are solved by 0 to 6 of the enabled models) but is not separable from prompt embeddings here. PCA of the shipped jina vectors shows hard and easy tasks intermixed, with PC1 and PC2 together carrying only 15.1% of the variance.
How weak a routing signal could this suite have seen?¶
The null above is a null on a pipeline that has been shown to recover a signal
it is known to contain — the positive control and destroyed-signal null in
benchmark/routing/instrument_control.py, which both selection rules now clear.
That licenses one claim and no more: nothing was found at the strength
probed. The planted control signal is far stronger than any routing signal a
real corpus would carry, so it cannot tell you whether the null means "there is
nothing here" or "there is something here, below our floor".
benchmark/routing/sensitivity.py measures that floor. It re-assigns the real
outcome rows to the real tasks so that a controlled fraction of them line up
with a direction in the real embedding space, sweeps that fraction downward, and
reports the smallest effect the published test still flags at 80% power. Because
planting only re-assigns rows, the permutation null does not move: the bar is the
one quoted above. The floor is reported as the AUROC a perfect reader of the
planted signal would achieve, which is the same unit the escalation section uses.
The answer is not encouraging, and it is a result in its own right: the floor sits far above any published task-difficulty detector, and under the transfer figure's own selection-corrected best-over-k rule the test does not reach 80% power even against a perfectly predictable corpus. Read the routing null as "this suite cannot resolve a plausible routing signal", not as "no routing signal exists." Run it yourself — it prints and writes nothing:
The one positive signal. Tasks from the same source repository transfer
slightly better than across repos: a same-repo diagonal advantage of +0.0330
(0.7836 versus 0.7506) against a matched shuffled-outcome null of
[−0.0176, 0.0236], z = +2.93 over 200 permutations. Small, real, and
repo-local rather than task-semantic. It is also fully explained by the embedded
string: the repo name was 14% of it (see the caveat opening this section), so this
is a measurement of the label, not of transfer. It does not survive as evidence
until re-measured on problem_statement.
The oracle-gap decomposition¶
We decomposed the gap between the hindsight Oracle and the fixed-frontier baseline to find how much of it is learnable at all. The answer changed the roadmap.
| cost | saving vs always-frontier | what it requires | |
|---|---|---|---|
| Always-Frontier | $96.02 | — | nothing |
| Price-Cascade (blocked) | $27.11 | 71.8% | no prediction — but mid-session verification, which is why it is not deployable, and why session_cascade re-buys most of it at session cadence for +$1.37 |
| Difficulty-only oracle | $14.25* | 83.9%* | perfect difficulty prediction |
| Oracle (exact model) | $18.33 | 80.9% | + hindsight token counts |
A difficulty-only oracle, one that always picks the cheapest model that solves the task and ignores which specific model it is, agrees with the full Oracle on 170 of 177 tasks (96%) and costs only $0.66 more. There is essentially no "one magic model for this task" effect to capture. The Oracle is almost entirely "use cheap when cheap works".
That splits the headroom in two. Both shares below are of the headroom itself, the $77.69 gap between Always-Frontier and the Oracle, not of the baseline's total cost:
- ~90% of the headroom is mechanically available. Collectable today by trying models cheapest-first and stopping at the first verified pass. No model, no features, no training.
- The remaining ~9% of the headroom requires predicting task difficulty, and our kNN router does not predict it: leave-one-out accuracy never beats the base rate at any k. That is a result about the router as it stands, embedding the 106-character label — not about task embeddings, which have not been tested here (see the caveat opening Routing results).
To make the two denominators reconcilable: that ~9% residual is ~7% of Always-Frontier's total cost (about $7.0 of $96.02). Different denominator, same dollars.
The honest conclusion is neither "routing works" nor "there is nothing here". It is that the prize is real and large, almost all of it is mechanical, and the learned part is currently worth nothing.
Provenance caveat. Unlike every other table on this page, this decomposition
came from a one-off analysis and is not yet regenerated by a committed script.
The Always-Frontier / Price-Cascade / Oracle costs above ARE the regenerated
strategy_summary.csv values (re-scored after the registry default_arm
change); the difficulty-only row (marked *) and the agreement count and the
headroom split still carry the pre-change one-off numbers, which are internally
consistent with the old strategy summary and have not been re-derived. Until the
analysis ships as part of the pipeline, treat this section as the one place here
you cannot reproduce with a single command. Making it reproducible is queued.
Escalation results¶
Status: OK_OFFLINE_ONLY — through the recurrence policy, and the shipped threshold itself
once the reproduction phase is excluded.
OK_OFFLINE_ONLY here means a statistical signal on the offline corpus: the best policy cell
clears its permutation null, and its precision interval clears the base rate.
It is not a shippable verdict. The deployability field (below) marks the
number OFFLINE-ONLY UPPER BOUND: the sweep scores one event per step while
production decides once per session, so this is a per-step signal the live
router does not run. "There is a signal" and "you can ship it" are different
sentences, and only the first is being asserted.
This section reports every escalation measurement. The single claim we are willing to make from them, scoped exactly and with its pre-registered falsifiers and their verdicts, is on its own page: the escalation claim.
The old evaluation could not have detected success¶
We rebuilt the escalation evaluation this cycle. The old label was positional, defined as the last few steps of a failed run, so a content-free clock scored AUROC 0.970 while a perfect task-level oracle capped at 0.757. Any detector tuned against that was tuned against a clock.
The shipped threshold is a coin flip — because it counts the reproduction phase¶
The shipped default (escalate_after_n=2, stale_window=10) fires on 723 of
723 trajectories: P(fail | fired) = 0.418 — exactly the
base rate, lift 1.00, no edge. This is not a tuning accident: every run, resolved or not,
starts by reproducing the bug, so the first one or two replayed steps are red on the target
F2P test and the counter trips immediately. The counter is counting the target bug at t=0 —
the agent's normal "fail, read the traceback, fix" loop — as if it were evidence the agent is
stuck. As-shipped, this is literally a coin flip, and no escalate_after_n tuning below the
run length fixes that.
The same mechanism discriminates once the reproduction phase is excluded¶
The eval replays the identical recurrence rule in a second family (count_from_first_edit):
failures before the agent's first edit-like action are treated as the reproduction phase and
not counted. Measured over the same 723 stamped runs, that variant separates immediately
and strongly:
| cell | fires | P(fail|fired) | lift | AUROC | len-only |
|---|---|---|---|---|---|
| n=2 | 431/723 | 0.589 [0.508, 0.655] | 1.41 | 0.710 | 0.568 |
| n=3 | 354/723 | 0.638 [0.554, 0.703] | 1.53 | 0.722 | 0.576 |
| n=5 | 265/723 | 0.694 [0.612, 0.764] | 1.66 | 0.708 | 0.575 |
| n=10 | 182/723 | 0.808 [0.738, 0.870] | 1.93 | 0.702 | 0.560 |
| n=20 | 115/723 | 0.835 [0.756, 0.900] | 2.00 | 0.636 | 0.560 |
These rows are the stale_window=1000 cells — a window wide enough to hold n recurrences, and
the family the status verdict is selected from. They are not the shipped configuration:
Shunt ships escalate_after_n=2, stale_window=10, and every committed figure draws that canonical
cell (edit-gated n=2: fires 431/723, P(fail|fired)=0.589, AUROC 0.710). The n=2 row is
identical across windows. The full 30-cell-per-family grid — every escalate_after_n from 1 to
50 at both windows — is on the sweep-table figures and in the report JSON.
The n=3 cell clears the family-wise permutation null (AUROC 0.722 against [0.5, 0.5499], adjusted p = 0.0005 over 2000 permutations) and the length-stratified null (0.722 against [0.498, 0.563]) — so the edge is recurrence beyond run length, not the length of the runs it selects. It fires on 354 of 723 runs — a useful fraction, not a tail. This is the honest read of the escalation idea: looking for repeated failures is the right approach; the shipped implementation was counting the wrong failures. The edit-gated family is eval-only (production has no per-step action stream to gate on); closing that gap is a design question, not a data one.
The as-shipped (reproduction-counted) family only separates at high thresholds¶
For completeness, the as-shipped family — the counter that counts every same-key failure
including the reproduction phase — over 723 stamped trajectories (152 distinct challenges, base
rate 0.418), the sweep varies escalate_after_n × stale_window (30 cells). The two knobs are
coupled: _in_window admits at most stale_window events, so reaching n
recurrences needs a window at least that wide, and the grid sweeps the window
over {10, 1000}. Its only edge sits at thresholds no one would ship (see the edit-gated family
above for the same mechanism at the shipped threshold):
- The shipped default (
escalate_after_n=2,stale_window=10) fires on 723 of 723 trajectories: P(fail | fired) = 0.418 — exactly the base rate, lift 1.00, no edge. It fires on essentially everything because reproduction failures recur at step 1–2. - As the threshold rises, precision separates from the base rate: n=5 reads 0.423, n=8
0.449, n=10 0.481, n=15 (
stale=1000) 0.534 (lift 1.28), n=20 0.577 (lift 1.38), and n=30 (stale=1000) 0.701 (lift 1.68). Every cell reports its OWN marginal challenge bootstrap (the family-wise maxT correction is applied to the AUROC null only, never to a precision interval — a CI that excluded a cell's own point estimate would not be an interval for it). The n=30 cell's marginal interval is [0.601, 0.782], so the numbers quoted against it are the ones the report prints. - The n=30 cell clears the gate outright: AUROC 0.658 against the
max-over-cells family-wise null 95% [0.5, 0.5523], adjusted p = 0.0005.
The harness's
OK_OFFLINE_ONLYbadge, however, is carried by the edit-gated n=3 cell above (AUROC 0.722 at stale_window=1000 — the best SKILLED cell, selected across both families and maxT-corrected, which is NOT the shipped cell the figures draw), not by this one; the gate itself is described on the offline-eval page. - About 40% of that excess over chance is run-length selection, and the report now says so.
Firing at n=30 requires ≥30 same-key failing steps, which requires a long run, and run length
is outcome-correlated on this corpus. A pure "run length ≥ threshold" predictor scores AUROC
0.561 at the same flag count, and the cell still clears the length-stratified null
(labels permuted within length bins, so the length→failure association survives): AUROC 0.658
against [0.536, 0.582], p = 0.0005. So the recurrence-specific signal is real but roughly
about 40% of the raw 0.658-to-0.5 excess is the length of the runs the threshold selects. Both
references — the length-only baseline and the length-stratified null — are reported on every
swept cell (JSON
length_baseline_auroc/null_auroc_length_stratified, and a column on the sweep table), because a cell that only clears the challenge-block null while matching its length baseline is selection, not recurrence. stale_windowis not inert at high n: with the window at 10 the policy stops firing once n ≥ 12 — it takes a window at least that wide to collect the recurrence. Only thestale=1000rows reach the null-clearing edge.
The prefix risk model is honestly NO_SKILL¶
The other half of the eval fits a continuous risk score from prefix-only features. On this corpus it reads no signal:
- At depth 10 — the shallowest full-rank, leak-safe depth — prefix AUROC is 0.478 and the incremental over the prior floored at chance is −0.022, inside the family-wise null. This is a real negative, not a bug: the escalation signal does not live in a shallow prefix on this corpus.
- Depth 5 is dropped from the reported ladder: its design is rank-deficient (412 of 414 admitted rows carry an identical feature vector, because in the first five replayed steps the agent is still reproducing the bug). Depth 20 is dropped too: at that depth the admission test selects failures (run length is outcome-correlated), so its near-nonzero incremental measured selection, not prefix evidence.
- The minimum detectable effect is ≈ 0.59, so this corpus cannot resolve a weaker detector. The 723 scored runs cluster on only 152 distinct challenges, so settling the prefix question needs roughly four times the distinct challenges (152 → ~640); more runs per existing challenge buy almost nothing, because the clustering already inflates variance ~3×.
The instrument is valid¶
The R0 gate passes: a planted, known-learnable signal is recovered by the
assembled pipeline (AUROC 1.000), and a within-challenge shuffle of the
outcome labels collapses it to chance (0.535, inside the band). A second
positive control proves power at the effect size the claim rests on, not only
detectability: with the fired↔failed link made imperfect (70% of failed runs
fire, 40% of resolved runs fire spuriously) the best cell still clears the
family-wise null at AUROC 0.634. A permanent
test (tests/escalation/test_instrument_validity.py) enforces both controls, so the
NO_SKILL verdict above is a null on an instrument shown to detect a signal,
not a null on a broken one.
The retraction, still on record¶
- Task identity alone, which is what the routing model already knows at t=0, was once published here at AUROC 0.886. We have retracted that number. The prior gave each run the leave-one-out failure rate of its own instance's other runs, while the cross-validation split grouped by instance — so it was scored on labels from its own test fold. A router meeting a new task has no such siblings. That leaked quantity still appears in the harness, now explicitly labelled as not the baseline (it scores 0.804 here). Grouped honestly the deployable prior is 0.497, i.e. no better than chance.
- Earlier prefix increments — +0.144 at 5 decisions, +0.061 floored, and +0.076 at the best depth (20 decisions) — were inflated by the between-fold base-rate accounting artifact in the prior column and by depth selection. Under the corrected methodology the increment is −0.022 at depth
- The comparator is floored at chance,
AUROC(prior + prefix) − max(AUROC(prior), 0.5), so an anti-predictive baseline cannot be beaten into an apparent finding. - The null permutes labels within a challenge, not globally: a global shuffle destroys the challenge-level clustering of outcomes, which collapses the prior to chance under the null while the observation keeps a real one — the two arms then sit in different headroom regimes and the gate has no power. Permuting inside each challenge preserves every challenge's outcome multiset, so the prior is identical in both arms and only the prefix's contribution is nulled.
So the honest verdict has changed: the escalation signal is real, and it lives in
the recurrence rule at the shipped threshold once the reproduction phase is
excluded. As shipped the counter counts every same-key failure including the
reproduction phase — it fires on everything and reads the base rate — and the
prefix model remains NO_SKILL, but the recurrence mechanism, gated on failures
after the agent's first edit, now clears its gate at escalate_after_n=3.
Two caveats that make it harder than it looks¶
A data gap, reduced but not closed. 253 of the committed corpus's trajectories once carried no
per-step outcomes, so the recurrence trigger structurally could not fire on them
— and three models (kimi-k2.5, qwen3.7-plus, glm-5.2) sat at zero
coverage entirely. Because stamping coverage tracked capture date and capture
date correlates with model, model and coverage were confounded. Those runs have
since been re-stamped offline by container replay at zero API cost: 723 of the committed
corpus's trajectories now carry verified per-step outcomes. But coverage is NOT uniform:
the two models that once sat at zero (qwen3.7-plus 47/65, glm-5.2 16/31)
still carry 3× the unstamped share of the other models, so stamping coverage still
tracks the same model-correlated axis as before — reduced, not eliminated. The 99
unstamped trajectories break down as: 23 whose captured state was
lost mid-run and whose steps the state-capture audit therefore marks unmeasured
rather than failed, 30 that carry no state-capture audit record at all (so whether
their capture was lost is unknown, and their stamps cannot be trusted), and 46
that the per-step stamping stage simply never reached. The model/coverage
confound is therefore still present and the prefix NO_SKILL verdict above stands
on the complete corpus with that caveat.
The value is not identified. Our logging policy never escalates, so P(escalate) = 0 and the overlap condition that every off-policy estimator requires fails. No stored trajectory contains an escalation that actually happened. We currently cannot distinguish "escalation helps" from "escalation hurts" from this data, at any confidence. The fix is ε-greedy randomisation at flagged checkpoints with logged propensities.
Figures¶
This page is the narrative — what we found and what it means. The figures themselves are documented one by one, with how to read each axis and what each one cannot support, in Routing → Figures and Escalation → Figures. Each PNG carries only its claim, its sample size, and — where a reader could be actively misled — one red line; see how the figures are built.
Routing¶
Earlier versions of these figures encoded a 106-character description label
rather than the SWE-bench problem statement, and every embedding null was
reported as a coverage gap for that reason. That excuse is gone: the task
manifest was rebuilt with the real problem statements (median 1185 characters),
the router embeds them, and the null did not move. Predicting per-task
solvability over 190 tasks with ≥2 measured models, leave-one-out, real jina
embeddings, k=20:
| Input to the encoder | LOO R² |
|---|---|
| 106-char identifier label (old) | −0.0712 |
| Real problem statement (new) | −0.0662 |
| SWE-bench human difficulty tag | +0.1876 |
| Shuffled-outcome null, 95% | [−0.1128, +0.0172] |
The embedding sits inside the null band on the correct input, while a three-level
human tag clears it on the same pipeline, the same n and the same null. That is a
working positive control beside a negative result, so this is a falsification,
not a coverage gap: on this corpus, embedding similarity carries no per-task
outcome signal. See embedding_signal.png.
Each figure is documented individually — how to read its axes, what to look for, and what it cannot support — beside the mechanism it illustrates:
- Routing (18 figures): routing.md → Figures
- Escalation (6 figures): escalation.md → Figures
- The live router (8 figures): inference.md → Figures
The figures live under docs/assets/figures/, one subdirectory per half
(routing/, escalation/, inference/), inside the published docs tree, which is
why the pages above can link them relatively. A committed figures.json per half — beside
the code that writes it, in benchmark/routing/, benchmark/escalation/ and
src/shunt/inspect/inference/ — records every
figure's full record and its input digest, and a lint gate (SH009) holds that manifest in
bijection with the sections above — so a retired figure cannot leave a stale description
behind it, and a documented figure cannot go missing.
Where this leaves the project¶
Routing: the mechanism works, the model does not, and the gate is untested —
its coverage floor is tripped, so it is not adjudicable, and the provisional read
underneath that refusal is a fail against both baselines rather than a near miss
(the verdict).
The routing null is a null on a suite whose minimum detectable effect sits
far above any plausible routing signal, so it bounds our resolution rather than
the idea. What changed this cycle is that the mechanism no longer has to be
quoted from a strategy the router refuses to run: the escalation ladder at
session cadence is cache-safe, ships enabled, is now selectable by name
(router.strategy: session_cascade), and reaches the blocked cascades'
operating point for $1.37 more on the 175-task set (session
cadence). What also changed is the size of the
prize — on fully-measured tasks the cascade family is ~25% cheaper than
fixed-frontier rather than ~75%, and the difference was imputation. The next
moves are a designed measured run to replace both opportunistic subsets, sized
against that floor, and better routing models (bigram and linear, calibrated
classifiers, better selection rules) evaluated against the same nulls.
Escalation: the recurrence mechanism works, and the shipped implementation was
counting the wrong failures. As shipped, the counter counts the reproduction
phase — every run's first reds are the target bug at t=0 — so it
fires on 723/723 runs and reads the base rate: a coin flip. Gated on
failures after the agent's first edit, the same rule separates: at the shipped
threshold n=2 it reads AUROC 0.710 at P(fail|fired)=0.589 (fires 431/723), and
the family's best cell, n=3, reaches AUROC 0.722 at P=0.638 (354/723), both clearing
the family-wise and the length-stratified nulls. At the session cadence the
ladder's value is large: escalating to a frontier model after a cheap session
failed resolves 3.02× more tasks than a same-cost retry (observational). It does
not, however, beat an always-frontier arm on quality at that cadence, nor firing
at random at the same rate — read the session-value figure's third panel before
quoting the ratio. Nor does a cost result survive as a fallback: the per-arm
USD-per-resolve figures are computed on different task sets, and the
common-task-set read that does exist prices two arms that differ in outcome on
none of the 48 instances, so its money answer comes with no quality axis at all.
What we do and do not assert, with the
pre-registered falsifiers and their verdicts, is stated once:
the escalation claim. The prefix
risk model remains NO_SKILL (the corpus cannot resolve a shallow prefix detector
below AUROC ≈ 0.59), and the edit-gated variant is eval-only — production has no
per-step action stream to gate on. The remaining work is making the post-edit
gate real in the live capture path, more distinct challenges (~640), and ε-greedy
randomisation with logged propensities so the value question becomes identified at
all.
Related reading: Benchmark for how the harness works, Benchmark design for why it is built this way, and Benchmark dataset for what is in it.