Independent Study | 2026
We put a frontier AI model up against real parcel audit data.
Here's why "just use ChatGPT" doesn't work — and what the evidence actually shows.
Study Details
4-week feasibility study | Independent AI engineering firm | Real FedEx invoices & negotiated agreement terms | Frontier LLM (Claude Sonnet 4.6)
What This Study Found
We tested the "just use AI" pitch rigorously. The results were unambiguous.
Reveel commissioned an independent AI engineering firm to spend four weeks attempting to replicate a parcel rate audit using a frontier large language model — real FedEx invoices, real negotiated agreement terms, 52,289 charge lines scored against known-correct output.
“A coin flip’s accuracy with an average error of $7.80 per charge line is not an audit. On a typical invoice file, that error is larger than the discrepancies being audited.”
Two approaches were tested. The first — prompting the LLM directly — returned wrong answers 53% of the time, with errors biased toward understating problems rather than flagging them. The second — using AI to build deterministic audit software — reached 99.5% accuracy, but required four weeks of specialist engineering for a single carrier, still left material gaps unsolved, and was only possible because the team had Reveel’s verified audit output to test against.
That last point is the one most teams building in-house miss entirely. Without ground truth to compare against, you cannot measure your own accuracy, find your own blind spots, or distinguish a genuine billing discrepancy from a logic error in your own engine. The answer key is the product.
Key Findings
Four findings worth taking into any AI governance conversation.
Here's why "just use ChatGPT" doesn't work — and what the evidence actually shows.
On data quality
Bad data doesn’t just slow AI down. It makes AI confidently wrong.
The model returned plausible-looking results that were wrong more than half the time — biased toward understating problems, not surfacing them. Poor data quality produces clean output you trust, not noise you catch.
On scaling AI
“Good enough” depends entirely on what you’re asking AI to do.
AI excels at reading ambiguity and extracting structure. It fails at applying the same logic verifiably across millions of records. Governance frameworks that treat those two tasks the same will fail at one of them.
On exception-based decisions
An exception is only actionable if you can verify it is real.
Without a validated baseline, a genuine billing discrepancy and a logic error in your own engine look identical. Teams moving to exception-based workflows need verified ground truth before they can trust what they’re acting on.
On governance at scale
Governance gaps invisible at pilot scale become expensive at production scale.
This study covered one carrier, one agreement, two invoices. Real operations span multiple carriers, annual rate changes, amendments, and new surcharges every season. The governance model that survives a pilot rarely survives production.
Independent Study | 2026
Read the full study.
The complete report covers the experiment setup, both tests in full, the accuracy progression that makes Test 2 work, the answer-key problem every DIY team faces, a head-to-head comparison table, and what the findings mean for how shippers should think about AI in parcel spend management.