Skip to content

Why naive pilots fail in production.

Four failure modes from a seeded Monte-Carlo simulation - the same patterns we harden against on real workflows. The script is published in our repo. Your signed targets come from your cases in weeks 1-2, not these demo numbers.

14.2%

Naive end-to-end

After 12 workflow steps

0

Wrong actions

On 500-case engineered queue

92.3%

Signed criteria

Same model, hardened

91.4%

Post-upgrade floor

With eval-on-release

Error compounding kills multi-step pilots

Each step in a workflow is a separate bet. A naive 85% per-step success rate compounds to 14.2% over twelve steps - the shape of most real operational chains.

14.2% end-to-end at 12 steps

14.2%

Naive end-to-end

96.8%

Engineered end-to-end

Error compounding kills multi-step pilots

0%25%50%75%100%123456789101112workflow steps85%72.2%61.4%52.2%44.4%37.7%32.1%27.2%23.2%19.7%16.7%14.2%99.8%99.6%99.4%98.9%98.7%98.5%98.3%97.8%97.6%97.4%97.3%96.8%
Naive pilot (85% per step)
Engineered workflow (99%+ per step)
Naive pilot (85% per step)
Engineered workflow (99%+ per step)

Silent wrong actions are the real risk

On a 500-case AP exception queue, the naive pilot produced 40 silent wrong actions. Engineering converts failures into escalations and human gates - not hidden damage.

0 wrong actions on 500 cases

40

Naive wrong actions

0

Engineered wrong actions

Silent wrong actions are the real risk

Naive pilot423 (84.6%)37 (7.4%)40 (8.0%)Engineered workflow396 (79.2%)64 (12.8%)40 (8.0%)
Correct / touchless
Stalled or escalated
Human-gated
Silent wrong action
Correct / touchless
Stalled or escalated
Human-gated
Silent wrong action

Hardening beats swapping models

Pass rate on the same eval suite rises through engineering iterations - integrations, bounded scope, gates, regression cases - not a bigger frontier model.

54% → 92.3% on same model

54%

Baseline pass rate

92.3%

After hardening

Hardening beats swapping models

0%25%50%75%100%≥90% sign-off threshold54%v0.1demo prompt67%v0.3+ real integrations78%v0.5+ bounded scope86%v0.7+ gates & rollback90%v0.9+ regression cases92.3%v1.0signed criteria
Below sign-off threshold
≥90% sign-off threshold
Below sign-off threshold
≥90% sign-off threshold

Model upgrades need eval discipline

Over twelve months and three provider releases, unmonitored accuracy drifts from 92% to 71%. Eval-on-upgrade catches regressions before complaints do.

Floor holds above 90%

71%

Unmonitored floor

91.4%

Eval-on-upgrade floor

Model upgrades need eval discipline

60%70%80%90%100%M1M2M3M4M5M6M7M8M9M10M11M12releasereleaserelease92.3%92.3%92.2%88.2%87.8%75.3%74.7%74.1%71%71%71%71%92.3%92.5%92.2%91.3%91.2%90.6%90.8%90.5%91.3%91.1%91%91.4%
Unmonitored (found via complaints)
Eval-on-upgrade
Provider release
Unmonitored (found via complaints)
Eval-on-upgrade
Provider release

How to read this page

01

Reproducible simulation

Charts are outputs from a seeded Monte-Carlo run on a representative AP-exception archetype - not client results. The simulation script is available on request; it is the same script used to generate these charts.

02

Your signed targets

In a real engagement, targets are set from your historical cases in weeks 1-2 and signed before go-live. That signed number is what we warrant.

03

Why we engineer

These charts explain why we build gates, eval harnesses, and production records - not what your workflow’s final numbers will be.

These charts explain why we engineer gates, eval harnesses, and production records - not what your workflow's final numbers will be. Want to see the same patterns on a live queue?

Watch the live simulation

Run this against your own queue

The simulation becomes real in weeks 1-2 of an engagement: criteria signed with your workflow owner, evals on 50+ of your historical cases, results in a record you own.