“Can we pull line items off scanned supplier invoices accurately enough to skip manual entry?”
- The pass mark we agreed
- 95% field-level accuracy across 500 real invoices, with errors flagged rather than silently wrong
- What we built
- An ingestion pipeline against a sample of the client's own scans, two model approaches run side by side, and a scoring harness that compared every extracted field against a hand-keyed answer set.
- What came back
- 91% on the first pass. The gap was almost entirely one supplier's carbon-copy forms, which no model handled. The recommendation was to ship with a review queue for low-confidence extractions rather than chase the last four points.
