Suite crm_ops@1.0.0 · 70 tasks × K=4 · 14 clusters · baseline crm_ops-baseline-15c80d461a · candidate crm_ops-candidate-cf750cb4b9 · policy 37c15e5837 · generated 2026-08-11 21:35 UTC

Gated metrics

VerdictMetricBaselineCandidate Δ (95% CI)marginp rawp BHpowertest
REGRESSION outcome.task_success 1.000 [1.000, 1.000] 0.143 [0.000, 0.333] -0.857 [-1.047, -0.667] 0.030 0.0001 0.0002 9% permutation
REGRESSION reliability.pass_hat_k@2 1.000 [1.000, 1.000] 0.143 [0.000, 0.333] -0.857 [-1.047, -0.667] 0.050 0.0001 0.0002 12% permutation
REGRESSION trajectory.f1 1.000 [1.000, 1.000] 0.486 [0.369, 0.603] -0.514 [-0.631, -0.397] 0.030 0.0001 0.0002 12% permutation
REGRESSION trajectory.argument_correctness 1.000 [1.000, 1.000] 0.330 [0.167, 0.493] -0.670 [-0.833, -0.507] 0.050 0.0001 0.0002 14% permutation
PASS efficiency.latency_ms 34.286 [30.200, 38.371] 26.857 [25.335, 28.379] +7.429 [+4.369, +10.488] 8.571 1.0000 1.0000 100% permutation
outcome.task_successoutcome.task_success: -0.8571 [-1.0474, -0.6669], margin -0.0300reliability.pass_hat_k@2reliability.pass_hat_k@2: -0.8571 [-1.0474, -0.6669], margin -0.0500trajectory.f1trajectory.f1: -0.5143 [-0.6315, -0.3971], margin -0.0300trajectory.argument_correctnesstrajectory.argument_correctness: -0.6701 [-0.8333, -0.5068], margin -0.0500efficiency.latency_msefficiency.latency_ms: +7.4286 [+4.3693, +10.4879], margin -8.5714-10.477+0.958+12.394change (candidate − baseline), 95% CI; red tick = margin

Naive threshold vs statistical gate

A fixed −3% threshold gate would have FAILED on outcome.task_success, trajectory.argument_correctness, trajectory.f1. AgentGate's verdict: REGRESSION.

The industry-standard rule — fail if any metric drops more than 3% — never consults the sample size. On 70 tasks a 3-point move is well inside sampling noise.

Reliability: pass^k decay

0.000.250.500.751.00baseline pass^1 = 1.000baseline pass^2 = 1.000baseline pass^3 = 1.000baseline pass^4 = 1.000candidate pass^1 = 0.143candidate pass^2 = 0.143candidate pass^3 = 0.143candidate pass^4 = 0.143k=1k=2k=3k=4P(all k repetitions succeed) — grey: baseline, blue: candidate

pass^k is the probability that all k repetitions succeed. A single-run pass rate cannot show this curve, which is why every task runs K times (E2).

Power and minimum detectable effect

Why each metric ruled the way it did

outcome.task_success — REGRESSION
reliability.pass_hat_k@2 — REGRESSION
trajectory.f1 — REGRESSION
trajectory.argument_correctness — REGRESSION
efficiency.latency_ms — PASS

Reproducibility

comparison id1de68e23b377
suite content hash36e21e7f686c09e5
baseline runcrm_ops-baseline-15c80d461a
candidate runcrm_ops-candidate-cf750cb4b9
policy hash37c15e58378604da
exit code1