A new format should not mean a new product.
Enterprise documents arrived in different layouts and levels of quality. A wrong field could undermine the workflow that depended on it. Format-specific parsers made each new variant another engineering commitment.
The long tail mattered more than the happy path.
I analyzed the long tail of failed extractions. The question was how to define a field consistently across unfamiliar documents, rather than keep adding rules for familiar layouts.
The decision
- Chose
- Schema-driven retrieval with evaluation before release.
- Over
- Build 6+ more per-format parsers.
- Evidence
- The long tail of failed extractions exposed the limits of format-specific rules.
- Trade-off
- Generalization across new variants required scrutiny beyond performance on familiar formats.
- Cost
- Upfront schema definition, evaluation work, and reviewer workflow design.
- Golden setsPer-tenant production documents. The reference for release decisions.
- LLM-as-judgeA complementary assessment of extraction quality.
- Human reviewInspect uncertain outputs against their source.
- Online monitoringWatch quality as production inputs change.
Evidence for a human release decision.
A system that knows when to ask a person.
The platform combined schema-driven retrieval, confidence-based routing, and source-linked review. Reviewer corrections stayed scoped to the tenant so that feedback retained its account context.
I built an evaluation framework in four layers, and release decisions were made against it. Per-tenant golden sets, seeded with production documents, were used to judge every release. LLM-as-judge, human review, and online monitoring added complementary checks.
System Details
- Parsing
- Native document text with OCR for scanned regions.
- Retrieval
- Schema-driven retrieval, with field definitions supplied by each tenant.
- Review
- Confidence thresholds route uncertain fields to human review. Values link back to source regions.
- Feedback
- Tenant-scoped reviewer corrections.
- Evaluation
- Per-tenant golden sets, LLM-as-judge, human review, online monitoring.
- Latency
- p50 1.2s, p95 ~3s on the extraction step.
- Deployment
- Single-tenant, inside a global manufacturer's own environment, on self-hosted Llama.
- Stack
- Python, React, Qdrant, Llama, Redis, Postgres, AWS.
Reliability became a release decision.
The platform processed 100K+ documents/month at 95%+ accuracy: field-level exact match, before human review. Turnaround was 70% faster at under 1 cent per page.
Straight-through processing ranged from 70 to 90%+, depending on the account and document complexity. Schema-driven retrieval replaced 6+ planned per-format parsers.
Trust needs evidence before it needs scale.
A capable model is only part of a usable product. The durable work was making uncertainty visible and giving the team shared evidence for deciding what was ready for production.