# Document Extraction API

*/Opportunities/Document_Extraction_API*

## Opportunity Overview

**Wedge**: The initial beachhead targets customs brokerages extracting line-item data from commercial invoices and packing lists. This niche faces the highest regulatory penalty for data entry errors, driving immediate adoption for high-accuracy extraction APIs. From this wedge, the product expands into adjacent freight forwarding documents like bills of lading, and eventually into post-delivery financial reconciliation workflows.
**Timing**: Multimodal foundational models now process complex document images natively and accurately without expensive fine-tuning. Concurrently, context windows have expanded enough to ingest 100-page logistics packets in a single pass while adhering to strict JSON schemas.
**Why This I C P**: Logistics operators handle extreme document variability across thousands of international partners, making template-based OCR completely unworkable. They feel acute margin pressure from manual data entry labor, forcing their internal developers to seek immediate automation solutions.
**Size Of Prize**: There are roughly 40,000 mid-market freight forwarders, 3PLs, and customs brokerages in the US and Europe. At an average spend of $15,000 per year on manual data entry and legacy OCR software per firm, the addressable prize is approximately $600M annually.
**Gap Narrative**: Developers in supply chain and logistics process a constant influx of varied, unstructured documents like bills of lading, customs declarations, and commercial invoices. Legacy OCR tools fail on variable layouts and require constant template maintenance, while off-the-shelf LLMs hallucinate data and lack reliable bounding-box citations. This API accepts raw PDFs and images, returning structured JSON with guaranteed schema adherence and confidence scores for every extracted field.
**Defensibility**: The core extraction capability is fundamentally a commodity, as foundational models continue to improve their native vision and extraction features. Defensibility only emerges through workflow lock-in by integrating directly into legacy on-premise systems and building custom fallback routing for human-in-the-loop review. Without deep integrations into the customer's proprietary data pipelines, the API remains highly susceptible to displacement by cheaper models.
**Why This Thesis**: An API thesis fits perfectly because logistics companies already use transportation management systems that require structured data ingestion. Developers need an infrastructure layer they can embed directly into their existing ingestion pipelines, not another standalone web dashboard that breaks their automated workflows.

## Opportunity Linked Thesis

**Thesis**: [Software](/Theses/Software)

## Opportunity Linked I C P

**Icp**: [Commercial Mortgage Lender](/CompanyTypes/Commercial_Mortgage_Lender)

## Opportunity Market Sizing

_Illustrative — target and order-of-magnitude estimate figures, not an achieved track record (this Thing is concept-stage)._

**S A M**: ~$300-500M US mid-market commercial mortgage lenders
**S O M**: ~$10-25M
**T A M**: ~15k commercial lending institutions × ~$60k/yr ≈ ~$900M
**Growth Rate**: ~12-18%/yr, driven by rising underwriting labor costs and increasing commercial document complexity
**Paid Comparable Spend**: ~$70k-120k/yr per lender on outsourced manual data entry teams or legacy template-based OCR licenses

## Opportunity Incumbents

- [Amazon Textract](/Products/Amazon_Textract) — Tool
- [Google Cloud DocumentAI](/Products/Google_Cloud_DocumentAI) — Tool
- [Azure Form Recognizer](/Products/Azure_Form_Recognizer) — Tool
- [Tesseract OCR](/Products/Tesseract_OCR) — Open-Source
- [Abbyy FlexiCapture](/Products/Abbyy_FlexiCapture) — Tool
- [Outsourced Data Entry](/Products/Outsourced_Data_Entry) — Service
- [Sensible API](/Products/Sensible_API) — Tool
- [Apache PDFBox](/Products/Apache_PDFBox) — Open-Source

## Opportunity Win Conditions

**Kill Thresholds**:
- Zero-touch processing rate < 85 percent after 30 days
- Pilot-to-paid conversion < 30 percent at $50k annual contract value
- Customer onboarding requires > 3 weeks of custom engineering
- Field-level accuracy drops below 98 percent on scanned PDFs
**Leading Metrics**:
- Time-to-first-production-API-call
- Zero-touch processing rate for rent rolls
- Human-in-the-loop correction frequency per document
- Average latency per 50-page mortgage package
**What Proves Right**: Mid-market commercial mortgage lenders route at least 40 percent of their inbound rent rolls and operating statements through the API in the first 60 days. Operations teams eliminate outsourced data entry spend because the API returns fully mapped JSON requiring zero human validation. Customers convert to $60k annual contracts following a 14-day proof of concept.
**What Proves Wrong**: Underwriters abandon the API because it hallucinates critical financial values on unstructured operating statements. The integration requires extensive custom mapping for each new broker template, destroying the margin advantage over manual data entry. Lenders refuse to pay a premium over Amazon Textract because the accuracy delta remains below 5 percent.

## Opportunity Build Profile

**Hardest Part**: Reconstructing reading order and correctly associating key-value pairs across page breaks and nested tables in highly variable layouts without human intervention.
**Min Viable Scope**: Extract a fixed schema from exactly one document type like commercial leases or bills of lading. Leave out multi-language support, handwriting recognition, and zero-shot querying across arbitrary document types.
**Cold Start Problem**: Robust models require millions of labeled examples of edge-case layouts that companies keep private. Break this by hand-labeling a narrow, publicly available dataset like SEC filings or county land records to win the first three paid pilots.
**Time To First Value**: Under 5 minutes. The gating step is the developer generating an API key, sending a POST request with a test PDF, and parsing the JSON response.
**Data Moat Available**: true
**Technical Difficulty**: High

## Neighborhood

### Where the gap lives

- [Reading Comprehension](/Skills/Reading_Comprehension) — latent gap · Skills

### Incumbent in

- [Azure Document Intelligence](/Products/Azure_Document_Intelligence) — incumbent in · Products
- [AWS Textract](/Products/AWS_Textract) — incumbent in · Products
- [ABBYY FlexiCapture](/Products/ABBYY_FlexiCapture) — incumbent in · Products
- [Tesseract OCR](/Products/Tesseract_OCR) — incumbent in · Products
- [Apache PDFBox](/Products/Apache_PDFBox) — incumbent in · Products
- [Outsourced Data Entry](/Products/Outsourced_Data_Entry) — incumbent in · Products
- [Sensible API](/Products/Sensible_API) — incumbent in · Products
- [Google Cloud DocumentAI](/Products/Google_Cloud_DocumentAI) — incumbent in · Products

### Applies thesis

- [Commercial Mortgage Lender](/CompanyTypes/Commercial_Mortgage_Lender) — applies thesis · CompanyTypes

### Embodies

- [Software](/Theses/Software) — embodies · Theses

### Similar Opportunities

- [Document Ingestion Service](/Opportunities/Document_Ingestion_Service) — similar · Opportunities
- [Document Ingestion Service](/Skills/Reading_Comprehension/Opportunities/Document_Ingestion_Service) — similar · Opportunities
- [DCP Ingest API](/Opportunities/DCP_Ingest_API) — similar · Opportunities
- [AI Waybill Parsing](/Opportunities/AI_Waybill_Parsing) — similar · Opportunities
- [Layout Semantics Engine](/Opportunities/Layout_Semantics_Engine) — similar · Opportunities
- [Vision Parsing Engine](/Opportunities/Vision_Parsing_Engine) — similar · Opportunities
- [Inbound Material Triage](/Opportunities/Inbound_Material_Triage) — similar · Opportunities
- [Document Routing Engine](/Opportunities/Document_Routing_Engine) — similar · Opportunities
- [Headless Document Clearance](/Opportunities/Headless_Document_Clearance) — similar · Opportunities
- [AI Customs Clearance for Freight Forwarders](/Opportunities/AI_Customs_Clearance_for_Freight_Forwarders) — similar · Opportunities
- [Dynamic Endpoint Aggregator](/Opportunities/Dynamic_Endpoint_Aggregator) — similar · Opportunities
- [Supplier Document Extraction](/Opportunities/Supplier_Document_Extraction) — similar · Opportunities
- [AI Invoice Extraction](/Opportunities/AI_Invoice_Extraction) — similar · Opportunities
- [AI Invoice Reconciliation for Freight Forwarders](/Opportunities/AI_Invoice_Reconciliation_for_Freight_Forwarders) — similar · Opportunities
- [Headless Invoice Extraction for AP](/Opportunities/Headless_Invoice_Extraction_for_AP) — similar · Opportunities
- [Transit Schedule Extraction for 3PLs](/Opportunities/Transit_Schedule_Extraction_for_3PLs) — similar · Opportunities
- [Resilient Ingestion Broker](/Opportunities/Resilient_Ingestion_Broker) — similar · Opportunities
- [Freight Invoice Auditing](/Knowledge/Transportation/Opportunities/Freight_Invoice_Auditing) — similar · Opportunities
- [Headless Invoice Extraction for Accountants](/Opportunities/Headless_Invoice_Extraction_for_Accountants) — similar · Opportunities
- [Demurrage Recovery Automation](/Opportunities/Demurrage_Recovery_Automation) — similar · Opportunities
