Three checks before an automation number enters your model
A printable worksheet for testing what an AI metric measures, where human work remains, and whether the economics survive.
The three checks
Keep the workflow boundary, period and sample size attached to every number.
The P&L bridge
Translate only a surviving capability claim into captured economic value.
Evidence behind the talk
The examples illustrate why the denominator, workflow boundary and human acceptance standard matter.
GitHub Copilot controlled experiment: 55.8% faster, n = 95
The study measured professional developers completing one bounded JavaScript HTTP server task in a controlled setting. Read the paper.
Faros AI: +98% merged PRs, +91% review time
Observational telemetry covered more than 10,000 developers across 1,255 teams. The result is a warning about downstream work, not a direct headcount finding. Read the report.
Uber: more than 70% of pull requests AI-attributed
AI attribution does not by itself disclose end-to-end autonomy. Uber describes workflows that can include human review or escalation. Read Uber Engineering.
METR: automated scores exceeded maintainer decisions by 24.2 points
Four maintainers from three repositories reviewed 296 AI-generated pull requests across 95 unique issues, blind to source. Read METR.
Builder.ai: a public warning did not become a diligence gate
The talk separates the 2019 warning about claimed AI capability from later revenue, financing and governance issues. Sources: Wall Street Journal, Financial Times, and bankruptcy filing.