Soukaina Abelhad AI Engineer New York 2026
THE AI CLAIM DILIGENCE CARD

Three checks before an automation number enters your model

A printable worksheet for testing what an AI metric measures, where human work remains, and whether the economics survive.

Two-page PDF. No signup.
Created by Soukaina Abelhad

The three checks

Keep the workflow boundary, period and sample size attached to every number.

01   ARITHMETIC
Volume versus labor
Request 18 months of workload volume against fully loaded human labor. Count reviewers, trainers, escalations and rework.
02   DEFINITION
Which rung?
Separate AI drafted, human edited, human approved and end-to-end straight-through processing inside a named boundary.
03   THE TELL
Can they measure quality?
Ask who labels ground truth, whether experts agree, what validated the judge, and what bad output escapes.

The P&L bridge

Translate only a surviving capability claim into captured economic value.

Eligible workflow valuestarting pool
× adoption rateactual use
× straight-through successno human intervention
× accepted qualityexpert standard
− AI + review + rework costfully loaded cost
= captured value

Evidence behind the talk

The examples illustrate why the denominator, workflow boundary and human acceptance standard matter.

GitHub Copilot controlled experiment: 55.8% faster, n = 95

The study measured professional developers completing one bounded JavaScript HTTP server task in a controlled setting. Read the paper.

Faros AI: +98% merged PRs, +91% review time

Observational telemetry covered more than 10,000 developers across 1,255 teams. The result is a warning about downstream work, not a direct headcount finding. Read the report.

Uber: more than 70% of pull requests AI-attributed

AI attribution does not by itself disclose end-to-end autonomy. Uber describes workflows that can include human review or escalation. Read Uber Engineering.

METR: automated scores exceeded maintainer decisions by 24.2 points

Four maintainers from three repositories reviewed 296 AI-generated pull requests across 95 unique issues, blind to source. Read METR.

Builder.ai: a public warning did not become a diligence gate

The talk separates the 2019 warning about claimed AI capability from later revenue, financing and governance issues. Sources: Wall Street Journal, Financial Times, and bankruptcy filing.

Ask what must be true for the number to matter