Case Study 01 — Invoice ingestion, automated end-to-end | SatyaHQ
CASE 01 Finance operations · AP automation

Invoice ingestion, automated end-to-end.

A leading finance company was drowning in 40,000 invoices a month across 200+ supplier formats. We replaced manual keying with an LLM-in-the-loop ingestion pipeline — and lifted throughput by 95% without adding a single headcount.

Productivity
+95%
Invoices per FTE, week over week
Time per invoice
45s
Down from 12 min manual keying
Supplier formats
200+
PDF, scans, email, portals, EDI
Keying errors
<0.2%
Two-stage validation gate
The client
A leading finance company handling receivables across dealers, distributors, and OEM partners in India.
Industry
Non-banking finance
Scope
AP invoice ingestion, master data, ERP posting
Timeline
14-week build · Live in production
The problem

Every supplier arrived in a different shape.

Invoices came in as clean PDFs, scanned images, email bodies, supplier-portal downloads and EDI feeds — with GST layouts that shifted per state and per vendor. Six clerks spent their days retyping the same fields into the ERP, and the queue only ever grew.

Where the time went.

A time-and-motion study of a typical week showed clerks spending 84% of their shift on data entry and correction, with format-specific handling driving the tail. Scanned invoices were the worst — OCR quality made every field a manual re-check.

PainFormat varianceManual keyingDuplicate posting risk

Average minutes per invoice, by format

Pre-automation baseline · 4-week sample · n = 6,842
0 5 10 15 20 min Scanned image 17.0m Email body 13.0m Portal export 12.0m Digital PDF 10.0m EDI feed 5.0m avg 12.0m
Pre-automation, per-format
The approach

One pipeline. Any supplier. Every format.

We collapsed six manual handoffs into a single automated flow with a human review lane for anything the model wasn't sure about. The pipeline learns from every correction.

1

Universal intake

Email drop, SFTP, portal scrapers, and mobile capture all land in one queue with source metadata preserved.

2

Layout-aware extraction

An LLM pre-labels every field against your GST + PO taxonomy. Per-field confidence attached.

3

Human review on edges

Only low-confidence fields hit the reviewer queue. Everything else is posted straight through.

4

ERP post + audit trail

Two-way sync with the ERP, with every change traceable back to the model version and reviewer.

Pipeline architecture

Ingest → Extract → Review → Post
Email drop SFTP Portal scrapers Mobile capture LLM extraction layout-aware · per-field confidence GATE STRAIGHT-THROUGH · 82% HUMAN REVIEW · 18% Auto-post to ERP Reviewer console reviewer corrections retrain the model
The results

Six clerks now supervise the pipeline instead of running it.

The same team clears an invoice queue 20× larger without overtime, and the model's straight-through rate keeps climbing as it learns from every reviewer correction.

Minutes per invoice — before vs after

All formats · steady-state, month 6
0 5 10 15 min BEFORE 12.0 min AFTER 0.75 min ↓ 94% Straight-through rate 82%

Supplier format mix

Share of monthly invoice volume · all handled by one pipeline
40K invoices / month
Digital PDF 45%
Scanned image 22%
Email body 15%
Portal export 12%
EDI feed 6%

Monthly invoice throughput — same team

Manual baseline vs automated pipeline · rolling 12-month view
Volumes in thousands
80K 60K 40K 20K 0 M0 M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 Manual ceiling ≈ 3K / mo Go-live Straight-through 82% 72K / mo · same team
Automated throughput
Manual ceiling
The impact

Same team. Twenty times the throughput.

Peak monthly volume
72K
Invoices processed, month 12
FTEs freed
4
Redeployed to vendor reconciliation
Straight-through rate
82%
And climbing every retrain
Payback
4.2mo
Full return on the build

"The team stopped being data entry clerks. They became the exception handlers — and the pipeline handles everything else. Volume is now a growth question, not an ops question."

— Head of Finance Operations, [Client name]
Deliverables

What they own after go-live.

Every artifact we shipped is theirs — the pipeline, the taxonomy, the model, and the tooling. No lock-in.

Ingestion pipeline

  • Multi-source intake (email, SFTP, portal, mobile)
  • Layout-aware extraction service
  • Confidence-gated routing
  • ERP write-back with audit log

Reviewer tooling

  • Web console with keyboard-first workflow
  • Side-by-side source ↔ extraction view
  • Correction capture into training set
  • Per-reviewer accuracy scoring

Model & taxonomy

  • Fine-tuned extraction model
  • GST + PO + vendor master taxonomy
  • Retrain scripts and gold-set
  • Weekly agreement dashboard
PythonFastAPIOCR + LLM ensembleSAP / Tally connectorsS3 + PostgresAirflowGrafana
Next case study

02 — Reconciliation without the reconcilers

Multi-SKU, multi-tender matching brought to 99.9% auto-match.

Read case 02 →