Passport Data Extraction
End-to-end pipeline that extracts structured fields from passport images using OpenCV preprocessing and PaddleOCR with MRZ parsing.
- Accuracy
- 97%
- Processing
- <2s
- Uptime
- 99.9%
Problem & Solution
The Problem
Manual passport data entry is slow and error-prone. Organizations need to digitize travel documents at scale with high accuracy across varying image quality and layouts.
The Solution
Built a pipeline that preprocesses passport images (deskew, denoise, contrast enhancement), localizes field regions and MRZ zones, runs OCR via PaddleOCR, and outputs structured JSON with name, passport ID, DOB, nationality, and address fields.
System Architecture
End-to-end flow from intake to outcome
- 01
Image Ingestion
FastAPIAccepts passport uploads via API or batch processing queue.
- 02
Preprocessing
OpenCVDeskews, denoises, and normalizes contrast on the input image.
- 03
Field & MRZ Extraction
PaddleOCRLocalizes field regions and MRZ zones, then runs OCR per region.
- 04
Structured Output
PythonValidates and returns extracted fields as typed JSON.
Related Projects
End-to-end cheque digitization pipeline using OpenCV for image preprocessing, PaddleOCR for text and amount extraction, and a Flask + Docker orchestration layer.
- Field accuracy
- 96.2%
- Throughput
- 120 img/s
- p95 latency
- 850ms
Suite of context-aware document intelligence pipelines: PDF intelligence (PyMuPDF + spaCy + BERT), passport MRZ segmentation, and prescription OCR.
- PDF entity F1
- 0.91
- MRZ accuracy
- 98.5%
- Rx field accuracy
- 94%