Suite crm_ops@1.0.0 · 70 tasks × K=4 · 14 clusters · baseline crm_ops-baseline-15c80d461a · candidate crm_ops-candidate-7130c8451e · policy 37c15e5837 · generated 2026-08-11 21:36 UTC

Gated metrics

VerdictMetricBaselineCandidate Δ (95% CI)marginp rawp BHpowertest
UNDERPOWERED outcome.task_success 1.000 [1.000, 1.000] 0.946 [0.918, 0.974] -0.054 [-0.082, -0.026] 0.030 0.0675 0.0844 63% permutation
UNDERPOWERED reliability.pass_hat_k@2 1.000 [1.000, 1.000] 0.893 [0.837, 0.949] -0.107 [-0.163, -0.051] 0.050 0.0407 0.0679 50% permutation
REGRESSION trajectory.f1 1.000 [1.000, 1.000] 0.926 [0.911, 0.941] -0.074 [-0.089, -0.059] 0.030 3.8e-05 9.6e-05 98% paired_t
PASS trajectory.argument_correctness 1.000 [1.000, 1.000] 0.971 [0.949, 0.994] -0.029 [-0.051, -0.006] 0.050 0.9580 0.9580 99% permutation
REGRESSION efficiency.latency_ms 34.286 [30.200, 38.371] 818.250 [675.329, 961.171] -783.964 [-922.840, -645.089] 8.571 3.1e-08 1.6e-07 6% paired_t
outcome.task_successoutcome.task_success: -0.0536 [-0.0816, -0.0255], margin -0.0300reliability.pass_hat_k@2reliability.pass_hat_k@2: -0.1071 [-0.1633, -0.0510], margin -0.0500trajectory.f1trajectory.f1: -0.0738 [-0.0889, -0.0586], margin -0.0300trajectory.argument_correctnesstrajectory.argument_correctness: -0.0285 [-0.0510, -0.0061], margin -0.0500efficiency.latency_msefficiency.latency_ms: -783.9643 [-922.8399, -645.0887], margin -8.5714-1015.981-457.134+101.713change (candidate − baseline), 95% CI; red tick = margin

Naive threshold vs statistical gate

A fixed −3% threshold gate would have FAILED on efficiency.latency_ms, outcome.task_success, trajectory.f1. AgentGate's verdict: REGRESSION.

The industry-standard rule — fail if any metric drops more than 3% — never consults the sample size. On 70 tasks a 3-point move is well inside sampling noise.

Reliability: pass^k decay

0.000.250.500.751.00baseline pass^1 = 1.000baseline pass^2 = 1.000baseline pass^3 = 1.000baseline pass^4 = 1.000candidate pass^1 = 0.946candidate pass^2 = 0.893candidate pass^3 = 0.839candidate pass^4 = 0.786k=1k=2k=3k=4P(all k repetitions succeed) — grey: baseline, blue: candidate

pass^k is the probability that all k repetitions succeed. A single-run pass rate cannot show this curve, which is why every task runs K times (E2).

Flakiest tasks (succeeded sometimes, not always)
TaskSuccessesK
address_delft-013 4
address_langley-013 4
address_langley-023 4
address_langley-033 4
escalation_delivery-033 4
escalation_delivery-053 4
refund_large_compiler-013 4
refund_large_compiler-033 4
refund_large_compiler-053 4
refund_large_orbital-053 4

Power and minimum detectable effect

Why each metric ruled the way it did

outcome.task_success — UNDERPOWERED
reliability.pass_hat_k@2 — UNDERPOWERED
trajectory.f1 — REGRESSION
trajectory.argument_correctness — PASS
efficiency.latency_ms — REGRESSION

Reproducibility

comparison idb57493ab492a
suite content hash36e21e7f686c09e5
baseline runcrm_ops-baseline-15c80d461a
candidate runcrm_ops-candidate-7130c8451e
policy hash37c15e58378604da
exit code1