Suite crm_ops@1.0.0 · 70 tasks × K=4 · 14 clusters · baseline crm_ops-baseline-15c80d461a · candidate crm_ops-candidate-15c80d461a · policy 37c15e5837 · generated 2026-08-11 21:35 UTC

Gated metrics

VerdictMetricBaselineCandidate Δ (95% CI)marginp rawp BHpowertest
PASS outcome.task_success 1.000 [1.000, 1.000] 1.000 [1.000, 1.000] +0.000 [+0.000, +0.000] 0.030 1.0000 1.0000 100% paired_t
PASS reliability.pass_hat_k@2 1.000 [1.000, 1.000] 1.000 [1.000, 1.000] +0.000 [+0.000, +0.000] 0.050 1.0000 1.0000 100% paired_t
PASS trajectory.f1 1.000 [1.000, 1.000] 1.000 [1.000, 1.000] +0.000 [+0.000, +0.000] 0.030 1.0000 1.0000 100% paired_t
PASS trajectory.argument_correctness 1.000 [1.000, 1.000] 1.000 [1.000, 1.000] +0.000 [+0.000, +0.000] 0.050 1.0000 1.0000 100% paired_t
PASS efficiency.latency_ms 34.286 [30.200, 38.371] 34.286 [30.200, 38.371] +0.000 [+0.000, +0.000] 8.571 1.0000 1.0000 100% paired_t
outcome.task_successoutcome.task_success: +0.0000 [+0.0000, +0.0000], margin -0.0300reliability.pass_hat_k@2reliability.pass_hat_k@2: +0.0000 [+0.0000, +0.0000], margin -0.0500trajectory.f1trajectory.f1: +0.0000 [+0.0000, +0.0000], margin -0.0300trajectory.argument_correctnesstrajectory.argument_correctness: +0.0000 [+0.0000, +0.0000], margin -0.0500efficiency.latency_msefficiency.latency_ms: +0.0000 [+0.0000, +0.0000], margin -8.5714-10.286+0.000+10.286change (candidate − baseline), 95% CI; red tick = margin

Naive threshold vs statistical gate

A fixed −3% threshold gate would have PASSED this comparison. AgentGate's verdict: PASS.

The industry-standard rule — fail if any metric drops more than 3% — never consults the sample size. On 70 tasks a 3-point move is well inside sampling noise.

Reliability: pass^k decay

0.000.250.500.751.00baseline pass^1 = 1.000baseline pass^2 = 1.000baseline pass^3 = 1.000baseline pass^4 = 1.000candidate pass^1 = 1.000candidate pass^2 = 1.000candidate pass^3 = 1.000candidate pass^4 = 1.000k=1k=2k=3k=4P(all k repetitions succeed) — grey: baseline, blue: candidate

pass^k is the probability that all k repetitions succeed. A single-run pass rate cannot show this curve, which is why every task runs K times (E2).

Power and minimum detectable effect

Why each metric ruled the way it did

outcome.task_success — PASS
reliability.pass_hat_k@2 — PASS
trajectory.f1 — PASS
trajectory.argument_correctness — PASS
efficiency.latency_ms — PASS

Reproducibility

comparison id9ae37f09073e
suite content hash36e21e7f686c09e5
baseline runcrm_ops-baseline-15c80d461a
candidate runcrm_ops-candidate-15c80d461a
policy hash37c15e58378604da
exit code0