Forty Routine Invoices and One Angry Supplier
Open a shared AP mailbox at Keller North America, the continent-wide operations of the world’s largest geotechnical specialist contractor, on any given morning and you'll find the whole range: clean PDF invoices, scanned documents photographed at an angle, monthly statements, credit notes, payment queries — and, every so often, a supplier who has chased the same payment three times and is now writing in a very different tone.
The last one is the email that matters most this morning. It's also the one most likely to sit unread behind forty routine invoices, because a mailbox sorted by arrival time has no idea that message is different.
That observation — that the hard part of AP intake is judgment, not typing — shaped everything about this deployment.
Why Rules Weren't Enough
The team had the classic options on the table. Basic OCR could read clean invoices but couldn't decide what a document was. Rules-based routing worked until a vendor changed their template. And nothing in the rules world could notice an upset supplier, or handle the genuinely messy cases: a single email carrying an invoice, a statement, and a question.
What the AP team actually did all day was a sequence of judgment calls: What is this? Is it urgent? Does it split into multiple things? Who should handle it? Has it been processed before? The automation had to make those same calls — and be right about them reliably enough to run unattended.
What the Pipeline Does
The deployed system polls each regional mailbox on a 30-minute cycle and, for every new email:
- Classifies the message and each attachment — invoice, statement, query, credit note, or other.
- Reads the tone. Sentiment detection flags frustrated or escalating messages for immediate human attention, ahead of the routine queue.
- Splits mixed content. An email with an invoice and a statement and a query becomes three correctly-typed work items, each tracked separately with a threaded note tying them back to the original message.
- Extracts vendor, invoice number, dates, amounts, and references — zero model training, any format.
- Checks for duplicates against a ledger of over 1,200 previously processed invoices, flagging repeats with a note citing when the original was first seen.
- Routes each item to an AP team member by region, on a configured round-robin rotation.
- Files finished documents into SharePoint, mirroring the team's own folder structure.
And when anything in the chain fails — an unreadable document, a verification mismatch — the pipeline rolls the entire email back. No half-processed documents, no orphaned entries, just a clean exception for a human to look at.
Proving It Before Trusting It
The team didn't flip this on and hope. Fifty-seven archived email scenarios — the weird ones included — were replayed through the pipeline until all fifty-seven passed. A six-week pilot processed more than 1,200 real documents, with findings folded back into the system each week. Only after a full source-to-target dry run, exception paths included, did live mailboxes come online.
That's the quiet lesson of the project: autonomy isn't a switch, it's a burden of proof. The mailbox runs itself now — because for six weeks, it had to prove it could.



