Extract text from PDF
Get the text of a PDF as a plain UTF-8 .txt file, in reading order, without uploading the document.
Formats & limits
Formats & limits
- Accepted formats
- Result
- TXT
| On a computer | On a phone or tablet | |
|---|---|---|
| Max file size | 300 MB | 100 MB |
| Max pages | 2000 | 300 |
Limits on your device depend on its memory. Larger files may still work; the tool warns you first.
How it works
- 01 Choose the PDF you want to extract text from.
- 02 Optionally list the pages, or turn on “Mark where each page starts” under More options.
- 03 Press Extract text.
- 04 Copy the text from the preview or download the .txt file.
About this tool
Copying text out of a PDF viewer often gives you jumbled lines: two columns woven together, a running header in the middle of a sentence. Hushdesk looks at the position of every character and rebuilds the reading order, so a two-column article comes out column by column while titles and page numbers stay where they belong. You get a plain UTF-8 text file that opens in any editor and can go straight into a translator, a search index or a chatbot. English, Russian and most other scripts come through intact. Pages that are only scanned pictures contain no letters to extract: the tool names them and suggests OCR instead.
-
Real reading order
Columns are detected on every page, so a two-column paper reads left column first, then the right one, instead of line by line across both.
-
Any alphabet
The file is saved as UTF-8: Cyrillic, accented letters, currency signs and symbols arrive exactly as they appear in the document.
-
Scans are flagged
When a page is a picture of text, you learn which page it is instead of silently getting an empty result.
-
Private by design
The text is read by an engine running in your browser tab. Reports, contracts and medical letters are never sent to anyone.
Questions and answers
Why is some text missing from the extracted file?
Most likely those pages are scans: the PDF holds a photo of the page, not the letters themselves. The result warns about such pages. Run the document through OCR first, then extract the text.
Does the text file keep tables and formatting?
No. A .txt file has no fonts, bold type or table cells; each table row becomes a line with its cells separated by spaces. For an editable document that keeps formatting, use PDF to Word.
How are two-column pages handled?
The tool looks for a vertical white gap running down the page between blocks of text and reads the blocks on either side in turn. Full-width headings and footers act as separators, so the order stays natural.
Can I extract text from only a few pages?
Yes. Put pages and ranges such as 3, 7-9 into the Pages field; everything else is skipped.
Does it work with Russian and other languages?
Yes, any script the PDF stores as text is extracted in UTF-8. Right-to-left languages such as Arabic or Hebrew are an exception: their words may come out in reverse order.
Are words that were split by a hyphen joined back together?
No. A hyphen that breaks a word at the end of a line stays as it is in the PDF, because the tool cannot always tell a line-break hyphen from a real one.
Is the PDF uploaded to extract its text?
No. Reading the characters and building the text file both happen on your device, inside a frame that your browser cuts off from the network.