How the text is extracted
Most PDFs made from a word processor, a website or an accounting program store real text: each piece of a line is saved with its font, size and position. This tool opens the PDF with pdf.js (Apache-2.0), the engine behind Firefox's PDF viewer, reads those pieces page by page, and rebuilds lines from their positions:
Same line if the baseline moves less than half the font size. Space between two pieces if the gap is wider than 0.2 × the font size. Blank line (new paragraph) if the next line is more than 1.8 × the font size further down.
Worked example: an invoice line
An invoice prints "Total due" in 12 pt type and "$1,240.00" further right on the same baseline, 40 pt after "due" ends. The baseline doesn't move (0 < 6 pt, half the font size), so both pieces stay on one line; 40 pt is more than 0.2 × 12 = 2.4 pt, so a space goes between them: "Total due $1,240.00". The next line, "Thank you", sits 14 pt lower: that is more than 6 pt, so it starts a new line, but less than 1.8 × 12 = 21.6 pt, so no blank line is added.
Scanned pages have no text to extract
A scanner or phone camera saves each page as a picture. The page looks like text, but the file holds only pixels, so there is nothing for this tool to read. After extraction, the tool lists the pages that came back empty. If all of them did, the PDF is a scan. OCR PDF recognises the letters in those pictures and saves a searchable copy, and Image to text does the same for a photo or screenshot.
Limits to know about
| Content | What you get |
|---|---|
| Normal paragraphs | Text with line breaks where the PDF breaks lines, and blank lines between paragraphs |
| Tables | One line per row, cells separated by spaces; columns are not kept |
| Two-column layouts | Depends on how the file was made: usually one column after the other, sometimes interleaved |
| Headers, footers, page numbers | Included, because they are text on the page |
| Scanned pages | Empty, and listed so you know to use OCR |
| Text drawn as shapes (some logos, outlined fonts) | Not found: there are no letters stored |
Privacy
Contracts, statements and medical letters are the PDFs people most often need text from. Here the file is read by your own browser: nothing is uploaded, and the .txt download is created on your device. Large PDFs are read one page at a time, so a 500-page manual works on an ordinary laptop, and you can cancel at any point.
Questions people ask
Why did it find no text in my PDF?
The PDF is almost certainly a scan or a photo: each page is a picture of text, with no actual letters stored in the file. This tool reads the text layer only, so it finds nothing. Run the file through OCR PDF, which recognises the letters in the pictures and adds a text layer, then come back here, or copy the text OCR PDF shows you.
Why are some words run together or split across lines?
A PDF stores text as positioned pieces, not as sentences. The tool puts pieces on the same baseline onto one line and adds a space where there is a visible gap. Unusual fonts, letter-spaced headings and justified text can fool the gap check, giving "Ac count" or "Thetotal". Hyphenated line breaks stay as they are in the PDF.
Does it keep columns and tables?
No. The output is plain text, so tables become lines with the cells separated by spaces, and two-column pages may come out one column after the other or with lines mixed, depending on how the PDF was made. For tables, copying from the PDF reader into a spreadsheet often works better.
Can it read a password-protected PDF?
Yes, if you know the password. You are asked for it when you add the file, and it is only used in your browser to open the file. Files that open without a password but block copying are read normally.
Is the PDF uploaded?
No. The text is read by pdf.js, Mozilla's open-source PDF engine, running in your browser. The file is not sent to a server, and the text you copy or download is created on your device.
