OCR PDF: recognize text in scanned documents
Turn scans into PDFs you can search and copy from. Recognition runs in your browser by default.
France · Scheduled cleanup: 60 min after job creation · Free to use
Formats & limits
Formats & limits
- Accepted formats
- Result
| On a computer | On a phone or tablet | In the cloud | |
|---|---|---|---|
| Max file size | 200 MB | 50 MB | 100 MB |
| Max pages | 50 | 10 | 200 |
| Max image size | 30 MP | 16 MP | No limit |
Scheduled cleanup: 60 min after job creation · Usually about 20 s
Limits on your device depend on its memory. Larger files may still work; the tool warns you first.
How it works
- 01 Choose or drop a scanned PDF.
- 02 Pick the text language: Russian, English or both.
- 03 Keep “On this device”, or switch to the cloud mode for long or poor scans.
- 04 Press Recognize text and download the searchable PDF.
About this tool
A scanned contract or a photographed report is just a picture: you cannot search it, copy a quote from it or have a screen reader read it aloud. Hushdesk runs Tesseract, the open-source OCR engine, right in this tab. It recognizes the Russian and English text on every scanned page and places it as an invisible layer under the image. The pages look exactly as before and nothing is re-compressed, yet Ctrl+F, copy and paste now work. On-device OCR is in beta: a page takes a few seconds on a laptop and noticeably longer on a phone. For long or very poor scans you can switch to the cloud mode, which runs OCRmyPDF on our server (France).
-
Invisible text layer
Each recognized word sits exactly over its image, so search highlights the right spot while the page looks unchanged.
-
Russian and English
Language data for both is built in. Choosing just one language reduces mix-ups between look-alike Latin and Cyrillic letters.
-
Honest confidence
Every page gets a confidence score, and pages that were hard to read are named in the result so you know what to check.
-
Private by default
In device mode the PDF is processed in an isolated frame that your browser keeps off the network.
How accurate is on-device OCR?
On a clean 300 dpi scan the engine reads almost every character correctly. Skew, blur, speckles and tiny print lower that. Pages with an average confidence below 75% are listed in the result, and symbols such as № or «guillemets» are worth a quick look.
Questions and answers
How long does on-device OCR take?
About 2–5 seconds per page on a modern laptop, longer on phones. Each page has a 60-second limit: a page that needs more is skipped and named in the result, so one bad page never holds up the rest.
Why is the on-device mode marked beta?
It works well on clean scans, but it uses the fast Tesseract models inside your browser and phones have little memory. That is why it is limited to 50 pages on computers and 10 on phones; bigger or harder files are better sent to the cloud mode.
What happens to pages that already contain text?
By default they are left untouched and only image-only pages are recognized. If a page mixes real text with a scanned area, turn off “Skip pages that already have text”.
Does OCR change how my PDF looks?
No. The page content and its images stay as they are; only an invisible text layer is added. Bookmarks, links and form fields are kept.
Which languages can OCR PDF recognize?
Russian and English, alone or together. Other languages and handwriting are not recognized reliably.
When is my PDF uploaded for OCR?
Only if you choose the cloud mode and confirm the upload. OCRmyPDF then processes it on our server (France), and the file is deleted automatically. In device mode nothing leaves your browser.