Optional add-on module
Turn scanned paper into searchable truth.
The enterprise OCR node extracts text from scanned plans and documents, in Spanish and English, and feeds it straight into full-text search. It never touches the protected originals.
Board_Minutes_2026_Scan.pdf
412 pages · searchable PDF + .txt
Technical_Report_ES.pdf
page 268 of 391 · engine: es
Plant_Overview.pptx
queued · engine: auto-detect
## Page 268 of 391, confidence 0.97
…the design pressure of the vessel shall not exceed 24 bar in accordance with…
invisible text layer injected · original scan untouched
OCR capabilities
Extraction without exposure
Asynchronous extraction
Documents are queued and processed in the background. Status is surfaced in real time and text becomes searchable as soon as a page is complete.
Bilingual EN/ES
Language-aware pipelines for Spanish and English, with normalized tag extraction and locale-safe number and date handling.
Full-text indexing
Extracted text feeds the search index directly, enabling contextual discovery down to the exact page across the entire repository.
Queued, throttled, observable
A dedicated worker pool with predictable throughput, per-task status, and exportable results. No black-box processing.
What it does
From dead pixels to living text
Scanned documents are images — invisible to search, impossible to audit. Aegis OCR turns them back into text your platform can actually use.
Searchable PDF generation
Scanned documents come back as fully searchable PDFs with an invisible OCR text layer injected beneath the original image — the document looks identical, but every word is now selectable, indexable, and findable.
Hybrid per-page extraction
Each page is evaluated individually: native text is used when it passes quality heuristics, and OCR is applied only to the pages that need it. Documents that already contain text are never re-OCRed.
Multi-script UTF-8 recognition
Latin, Cyrillic, Greek, Chinese, Japanese, Korean, Arabic, Devanagari, Thai and more — each page is routed to the recognition engine that reads it best. Administrators choose the active languages, and each script model downloads on demand. Spanish technical documents additionally get a correction layer for accents, acronyms, and domain terminology.
Adaptive capture quality
Short documents — maps, blueprints, forms — are rendered at 400 DPI for maximum glyph fidelity; long documents run at an efficient 200 DPI standard.
Enterprise operation
Built for production loads, not demos
The OCR node manages its own resources, protects itself from overload, and leaves no residue behind.
Bulk queue processing
An asynchronous task queue processes documents in bulk with live progress, persistent history, and cancellation — files up to 3 GB each.
Self-tuning workers
The worker pool scales itself from available CPU and RAM every 30 seconds, and admission control rejects jobs when the node is saturated instead of degrading it.
Authenticated & audited
Signed session cookies, bcrypt-hashed credentials, per-user file isolation, and an audit trail of every upload, download, and deletion with IP and user-agent.
24-hour automatic retention
Processed tasks and files are automatically purged after 24 hours. The OCR node never becomes an uncontrolled copy of your document repository.
Deployment options
Run it in our cloud — or on your hardware
Aegis OCR is offered fully managed in our cloud, with zero infrastructure on your side. If your policy requires on-premise processing, these are the minimum requirements for a dedicated node.
Aegis Cloud Recommended
On-premise node
Minimum requirements when contracting local processing.
No GPU is required — OCR inference runs entirely on CPU. The node can share the same server as your Aegis Secure View deployment and is delivered as a hardened systemd service.
Bring your hardest documents
Send us a sample of your scanned plans and we will show extraction quality on your own files.