All work
ProductSquadsLead Product ManagerFeb 2025 to Sep 2025

Document AI that earns its place in production.

The roadmap called for 6+ more format-specific parsers. I chose schema-driven retrieval, with evaluation before release and human review wherever confidence was low.

Trust mechanism: evaluation before release, humans wherever confidence is low.

PE-backed enterprise document AI. Clients included global industrial and chemical manufacturers.

documents/month
100K+
Production volume
accuracy
95%+
Field-level exact match, before human review
faster turnaround
70%
Across document processing
per page
<1¢
Processing cost
Source documentSafety data sheet
Substance identification
Isopropyl alcohol
Composition
67-63-099%
Hazard identification
Flammable liquid, Cat. 2
Revision: 2024-03
Extracted fieldsReview
Product name
Isopropyl alcohol
High confidence
CAS number
67-63-0
High confidence
Concentration
99%
High confidence
Hazard class
Flammable liquid, Cat. 2
Low confidence
re-extracted at 300 DPI
Revision date
2024-03
High confidence

Source linked. Human judgment retained.

Source-linked fields let a reviewer inspect uncertainty in context before accepting an extraction.

My scope

I owned product decisions across extraction, retrieval, evaluation, and reviewer workflows. Delivered with 2 engineering pods and a 50-person human review team.

The Problem

A new format should not mean a new product.

Enterprise documents arrived in different layouts and levels of quality. A wrong field could undermine the workflow that depended on it. Format-specific parsers made each new variant another engineering commitment.

What I Found

The long tail mattered more than the happy path.

I analyzed the long tail of failed extractions. The question was how to define a field consistently across unfamiliar documents, rather than keep adding rules for familiar layouts.

The decision

Chose
Schema-driven retrieval with evaluation before release.
Over
Build 6+ more per-format parsers.
Evidence
The long tail of failed extractions exposed the limits of format-specific rules.
Trade-off
Generalization across new variants required scrutiny beyond performance on familiar formats.
Cost
Upfront schema definition, evaluation work, and reviewer workflow design.
  1. Golden setsPer-tenant production documents. The reference for release decisions.
  2. LLM-as-judgeA complementary assessment of extraction quality.
  3. Human reviewInspect uncertain outputs against their source.
  4. Online monitoringWatch quality as production inputs change.
Measured signals
Field-level exact match
Hallucination rate
Confidence calibration

Evidence for a human release decision.

Release decisions were made against per-tenant golden sets. This was a review framework, not an automated release gate.
What Shipped

A system that knows when to ask a person.

The platform combined schema-driven retrieval, confidence-based routing, and source-linked review. Reviewer corrections stayed scoped to the tenant so that feedback retained its account context.

I built an evaluation framework in four layers, and release decisions were made against it. Per-tenant golden sets, seeded with production documents, were used to judge every release. LLM-as-judge, human review, and online monitoring added complementary checks.

System Details

Parsing
Native document text with OCR for scanned regions.
Retrieval
Schema-driven retrieval, with field definitions supplied by each tenant.
Review
Confidence thresholds route uncertain fields to human review. Values link back to source regions.
Feedback
Tenant-scoped reviewer corrections.
Evaluation
Per-tenant golden sets, LLM-as-judge, human review, online monitoring.
Latency
p50 1.2s, p95 ~3s on the extraction step.
Deployment
Single-tenant, inside a global manufacturer's own environment, on self-hosted Llama.
Stack
Python, React, Qdrant, Llama, Redis, Postgres, AWS.
Outcomes

Reliability became a release decision.

The platform processed 100K+ documents/month at 95%+ accuracy: field-level exact match, before human review. Turnaround was 70% faster at under 1 cent per page.

Straight-through processing ranged from 70 to 90%+, depending on the account and document complexity. Schema-driven retrieval replaced 6+ planned per-format parsers.

What I Learned

Trust needs evidence before it needs scale.

A capable model is only part of a usable product. The durable work was making uncertainty visible and giving the team shared evidence for deciding what was ready for production.

Next case study · Closphere

A shared truth across disconnected inventory.

Read case study
Let’s connect

Let’s talk product.

For senior PM and AI product opportunities, or a complex problem worth working on.