Suite crm_ops@1.0.0 · 70 tasks × K=4 · 14 clusters · baseline crm_ops-baseline-15c80d461a · candidate crm_ops-candidate-ea3b8dbb84 · policy 37c15e5837 · generated 2026-08-11 21:36 UTC

Gated metrics

VerdictMetricBaselineCandidate Δ (95% CI)marginp rawp BHpowertest
UNDERPOWERED outcome.task_success 1.000 [1.000, 1.000] 0.786 [0.563, 1.000] -0.214 [-0.437, +0.009] 0.030 0.1276 0.3191 8% permutation
UNDERPOWERED reliability.pass_hat_k@2 1.000 [1.000, 1.000] 0.786 [0.563, 1.000] -0.214 [-0.437, +0.009] 0.050 0.1276 0.3191 11% permutation
PASS trajectory.f1 1.000 [1.000, 1.000] 1.000 [1.000, 1.000] +0.000 [+0.000, +0.000] 0.030 1.0000 1.0000 100% paired_t
PASS trajectory.argument_correctness 1.000 [1.000, 1.000] 1.000 [1.000, 1.000] +0.000 [+0.000, +0.000] 0.050 1.0000 1.0000 100% paired_t
PASS efficiency.latency_ms 34.286 [30.200, 38.371] 34.286 [30.200, 38.371] +0.000 [+0.000, +0.000] 8.571 1.0000 1.0000 100% paired_t
outcome.task_successoutcome.task_success: -0.2143 [-0.4373, +0.0088], margin -0.0300reliability.pass_hat_k@2reliability.pass_hat_k@2: -0.2143 [-0.4373, +0.0088], margin -0.0500trajectory.f1trajectory.f1: +0.0000 [+0.0000, +0.0000], margin -0.0300trajectory.argument_correctnesstrajectory.argument_correctness: +0.0000 [+0.0000, +0.0000], margin -0.0500efficiency.latency_msefficiency.latency_ms: +0.0000 [+0.0000, +0.0000], margin -8.5714-10.286+0.000+10.286change (candidate − baseline), 95% CI; red tick = margin

Naive threshold vs statistical gate

A fixed −3% threshold gate would have FAILED on outcome.task_success. AgentGate's verdict: UNDERPOWERED.

The industry-standard rule — fail if any metric drops more than 3% — never consults the sample size. On 70 tasks a 3-point move is well inside sampling noise.

Reliability: pass^k decay

0.000.250.500.751.00baseline pass^1 = 1.000baseline pass^2 = 1.000baseline pass^3 = 1.000baseline pass^4 = 1.000candidate pass^1 = 0.786candidate pass^2 = 0.786candidate pass^3 = 0.786candidate pass^4 = 0.786k=1k=2k=3k=4P(all k repetitions succeed) — grey: baseline, blue: candidate

pass^k is the probability that all k repetitions succeed. A single-run pass rate cannot show this curve, which is why every task runs K times (E2).

Power and minimum detectable effect

Why each metric ruled the way it did

outcome.task_success — UNDERPOWERED
reliability.pass_hat_k@2 — UNDERPOWERED
trajectory.f1 — PASS
trajectory.argument_correctness — PASS
efficiency.latency_ms — PASS

Reproducibility

comparison ide25fde54011e
suite content hash36e21e7f686c09e5
baseline runcrm_ops-baseline-15c80d461a
candidate runcrm_ops-candidate-ea3b8dbb84
policy hash37c15e58378604da
exit code0