AI & Document Processing · Professional Services

Intelligent document reading to automate corporate expense reconciliation

A multinational professional services firm helps companies review corporate expenses, a job that requires cross-checking thousands of documents in varied formats and justifying every decision for audit. In a six-week Discovery, Luby validated with real evidence the biggest technical risk of an automated reconciliation platform: reading financial documents that follow no standard, reliably and at a predictable cost.

~11 days

To build the complete AI reading pipeline

62

Real documents analyzed to choose the technology on data, not opinion

Cents

Per document: reading a receipt costs around $0.01–0.02

Challenge

Documents that arrive in every format

Expense review was heavy on manual work. Automating it depended on a condition no screen solves alone: correctly reading photos of receipts, stamps, reports dozens of pages long, card statements of almost a hundred pages, and spreadsheets with thousands of rows and no standard layout. Before sizing the full solution, the client needed objective answers.

  • Does automated reading work on these documents?
  • With what accuracy, and measured how?
  • How much does it cost per document and per thousand documents?
  • What fails, and why?
  • How do you avoid getting locked into a single AI vendor?

Solution

Answers backed by evidence, not opinion

Luby structured a six-week Discovery, with a Vision Document, a Work Breakdown Structure (WBS), and a Roadmap, followed by a proof of concept (PoC) designed as a measurement lab, not a demo.

Technology choice based on evidence

  • 62 real documents from the sample analyzed locally before writing the engine
  • Official pricing from the three major cloud providers compared side by side
  • Cost per page came out equivalent for traditional OCR and multimodal LLMs, but the LLM returns data already interpreted
  • The team's own initial hypothesis, a hybrid model with parsers plus an LLM, was ruled out in writing: it added a second system to maintain in exchange for savings of a few cents

Multimodal reading pipeline

  • Six input types: text PDF, PDF via vision, image, spreadsheet, HTML, and Word
  • Automatic choice of the best route for each document
  • Three reading strategies, with automatic validation and the reason for every deviation logged

01

Direct reading

Large documents, such as an 89-page statement, are read in full, with no cuts.

02

Calibrated reading

In spreadsheets with thousands of rows, the AI reads only a sample and returns a map of the structure. Code applies that map to 100% of the rows with deterministic rules and automatic validation. The full spreadsheet never goes to the model, which cuts cost and risk.

03

Chunked reading

When validation fails, the document is read in parts and consolidated, and the reason for the deviation is logged.

Results

What the evidence showed

The answers came from the real document sample, with cost, accuracy, and limitations measured.

Technical risk mitigated with evidence

Reading worked on the real sample, including photos, long reports, and spreadsheets with no standard, and the system flags unreadable documents instead of making up data.

0.97 / 0.72

Honest model confidence

In the same batch, the AI reported 0.97 confidence on the tabular part and 0.72 on the visual part, consistent with how hard each part really was.

Architecture refined by experiment

A 4-page summary reproduced the same 67 rows and the same total as a 62-page report, at 86% of the cost. That approach became the recommendation for the next phase.

< $1

Predictable cost

A full round of experiments with 16 reads cost less than $1, and the projection points to about $15–30 per thousand documents, interpretation included.

700+

A base ready to scale

More than 700 automated tests, a catalog of limitations handed to the client as a map of the next phase, and an architecture where the AI engine can be swapped without rewriting the pipeline.

Failure is data, not a bug

That was the founding principle of the PoC. The result is a metric that hides nothing.

Failure catalog

Every failure is categorized in an exportable catalog, and cost is recorded even when a read fails or is canceled. Temporary errors, like network instability or rate limits, are retried; permanent errors are never masked.

Human review as the source of truth

A reviewer checks every extracted field (correct, incorrect, or missing), and that verdict becomes the system's answer key. From it, the dashboard calculates per-field accuracy, read rate, cost per document, the projection per thousand documents, and latency, always with a 95% confidence interval.

Declared confidence versus real accuracy

The dashboard compares the confidence the model declares with its real accuracy, to answer whether that confidence can be used to prioritize human review.

Explainable, auditable, and free of vendor lock-in

Every suggestion comes with its reason, and the AI engine can be swapped without rewriting the pipeline.

Explainable reconciliation, human decision

The matching engine across reports, statements, and receipts is deterministic and auditable. It scores each pair on amount, date, and parties involved, and spells out the reason for each signal in Portuguese. The tool suggests and the analyst decides, one by one or in bulk.

No vendor lock-in

A single data contract defines the extraction format, validates the AI's response, and types the entire application. Switching providers, including to an AI service hosted in the client's own environment, is a setting on a screen.

Engines compared side by side

Different engines can be compared side by side on the same documents, with versioned instructions to keep runs comparable.

AI-accelerated engineering, with governance

The PoC was built by a Tech Lead working with AI agents under explicit rules. An instructions file defines what the AI cannot break, and business decisions and commits stay with the human. The delivery pipeline meets corporate standards:

  • CI with type checking, tests, and migrations applied to a clean database
  • Code quality gate
  • Security scanning
  • Secrets managed in a vault
  • Deployment on Kubernetes

Stack

Technologies

Application, data, and AI

TypeScriptNestJSReactPostgreSQLPrismaRedis/BullMQZodMultimodal LLMs (multi-provider architecture)

Cloud, delivery, and quality

AWS (S3, EKS)GitHub ActionsArgoCDSonarQubeTrivy

Does your process depend on documents no one can standardize?

Talk to Luby and find out how to validate AI with evidence before investing in the full solution.