Skip to main content

Optional add-on module

Turn scanned paper into searchable truth.

The enterprise OCR node extracts text from scanned plans and documents, in Spanish and English, and feeds it straight into full-text search. It never touches the protected originals.

ocr.your-datacenter.local
Aegis OCR Node
esencyrillic+9

Board_Minutes_2026_Scan.pdf

412 pages · searchable PDF + .txt

Completed

Technical_Report_ES.pdf

page 268 of 391 · engine: es

68%

Plant_Overview.pptx

queued · engine: auto-detect

Pending

## Page 268 of 391, confidence 0.97

…the design pressure of the vessel shall not exceed 24 bar in accordance with…

invisible text layer injected · original scan untouched

3 workers · CPU 41% auto-purge in 24h
Queued, throttled, and observable. Never a black box.

OCR capabilities

Extraction without exposure

Asynchronous extraction

Documents are queued and processed in the background. Status is surfaced in real time and text becomes searchable as soon as a page is complete.

Bilingual EN/ES

Language-aware pipelines for Spanish and English, with normalized tag extraction and locale-safe number and date handling.

Full-text indexing

Extracted text feeds the search index directly, enabling contextual discovery down to the exact page across the entire repository.

Queued, throttled, observable

A dedicated worker pool with predictable throughput, per-task status, and exportable results. No black-box processing.

What it does

From dead pixels to living text

Scanned documents are images — invisible to search, impossible to audit. Aegis OCR turns them back into text your platform can actually use.

Searchable PDF generation

Scanned documents come back as fully searchable PDFs with an invisible OCR text layer injected beneath the original image — the document looks identical, but every word is now selectable, indexable, and findable.

Hybrid per-page extraction

Each page is evaluated individually: native text is used when it passes quality heuristics, and OCR is applied only to the pages that need it. Documents that already contain text are never re-OCRed.

Multi-script UTF-8 recognition

Latin, Cyrillic, Greek, Chinese, Japanese, Korean, Arabic, Devanagari, Thai and more — each page is routed to the recognition engine that reads it best. Administrators choose the active languages, and each script model downloads on demand. Spanish technical documents additionally get a correction layer for accents, acronyms, and domain terminology.

Adaptive capture quality

Short documents — maps, blueprints, forms — are rendered at 400 DPI for maximum glyph fidelity; long documents run at an efficient 200 DPI standard.

Enterprise operation

Built for production loads, not demos

The OCR node manages its own resources, protects itself from overload, and leaves no residue behind.

Bulk queue processing

An asynchronous task queue processes documents in bulk with live progress, persistent history, and cancellation — files up to 3 GB each.

Self-tuning workers

The worker pool scales itself from available CPU and RAM every 30 seconds, and admission control rejects jobs when the node is saturated instead of degrading it.

Authenticated & audited

Signed session cookies, bcrypt-hashed credentials, per-user file isolation, and an audit trail of every upload, download, and deletion with IP and user-agent.

24-hour automatic retention

Processed tasks and files are automatically purged after 24 hours. The OCR node never becomes an uncontrolled copy of your document repository.

Deployment options

Run it in our cloud — or on your hardware

Aegis OCR is offered fully managed in our cloud, with zero infrastructure on your side. If your policy requires on-premise processing, these are the minimum requirements for a dedicated node.

Aegis Cloud Recommended

Zero infrastructure, zero maintenance
We operate, monitor, and scale the OCR nodes for you
Automatic updates and model improvements
Same 24-hour automatic data retention
Reachable through your existing Aegis endpoint

On-premise node

Minimum requirements when contracting local processing.

CPU2 cores minimum · 4+ cores recommended
RAM4 GB minimum · 8+ GB for files over 1 GB
Disk20 GB free · SSD recommended
GPUNot required · CPU inference by design
OSUbuntu 24.04 LTS or similar · Linux x86_64
RuntimePython 3.10+ · Node 20 (build only) · Caddy 2
Servicesystemd unit · co-hosts with the Aegis Secure View node

No GPU is required — OCR inference runs entirely on CPU. The node can share the same server as your Aegis Secure View deployment and is delivered as a hardened systemd service.

Bring your hardest documents

Send us a sample of your scanned plans and we will show extraction quality on your own files.

Request a demo