Twelve languages,
on your own machine.
A scanned page is a picture. OCR puts a text layer under it, so the words can be searched, selected and copied — and here that happens on the machine in front of you, with the language data already in the package. No upload, no queue, no per-page price.
Languages in this document
Tick every script the document uses. A document in two scripts needs both; ticking one it does not use costs accuracy rather than nothing.
An English-only OCR is not an OCR for India.
A land record in Marathi, a caste certificate in Kannada, a court order in Bengali — these are the documents people actually need to make searchable, and they are the ones every general-purpose tool hands back as an empty text layer.
All twelve travel inside the package. Nothing is fetched the first time you use one, which is the usual reason an “offline” OCR turns out not to be.
Fast enough that the progress bar is the boring part.
Measured on a 29-page filing, Release build, this machine. Your hardware will differ; the point of publishing the number is that it is a number rather than an adjective.
English
end to end
uploaded
now and later
in the package
Those pages already carried a text layer, which is the fast case. A 300 dpi scan with nothing underneath it is slower, and that number is not one we have measured yet — so it is not on this page.
The picture is untouched.
The text goes under the page, not over it. What you see afterwards is the scan exactly as it was; what a search finds is the layer beneath.
WHAT CHANGES
A text layer is added. The page image, its resolution and its colour are not re-encoded, so a scan does not lose a generation of quality for having been made searchable.
WHAT DOES NOT
The file you started from. The result is a new file where you chose to put it, and the original is byte-for-byte what it was.
Latin text is set in a standard font that is never embedded. Indic scripts need one whose character map covers them, so those documents embed a subset of a font Windows already ships — a subset, so a Hindi page does not add three megabytes of typeface to your file.
One of forty-three commands.
creasepoint ocr scan.pdf --lang eng+hin --out searchable.pdf 29/29 pages ocr -> searchable.pdf 2.6 MB in 0.9 s revisions 1 (no previous content retained)
Name a language it does not have and it tells you which ones it does, rather than
failing with a code. Run it over a folder with batch.