About text recognition
A scan or a photo of a page is only a picture: you cannot search it, copy a sentence from it or have a screen reader read it aloud. Optical character recognition (OCR) finds the letters in the image and turns them into real text.
There are two tools, for two kinds of files:
- OCR PDF adds an invisible text layer to a scanned PDF. The pages look the same, but search, copy and paste now work. Use it for contracts, invoices, archived letters and anything you will want to find again.
- Image to text reads a photo or a screenshot and gives you plain text to copy or save.
Recognition runs in your browser with Tesseract, the open-source OCR engine, and it reads English and Russian. A page takes a few seconds on a laptop and longer on a phone. For long or very poor scans, OCR PDF offers a server mode that runs OCRmyPDF. Job cleanup is scheduled 60 min after job creation. Technical failures may delay deletion.
For the best result, photograph the page flat, in even light and with all of it in the frame. A sharp scan saves more time than correcting the text afterwards.
Questions and answers
Which languages can be recognized?
English and Russian, including pages that mix the two. Other languages are not supported yet: their text would come out garbled.
How accurate is the recognition?
Clean printed text is recognized very accurately. Handwriting, blurry photos, skewed pages and decorative fonts lower the quality, so check names and numbers in the result.
Does OCR change how my PDF looks?
No. OCR PDF keeps each page image as it is and places the recognized text in an invisible layer under it.
Are my scans uploaded?
Not unless you choose the server mode of OCR PDF. By default recognition runs on your device, in a frame that has no network access.