Add a PDF and download its text as a plain .txt file. The file is named after the PDF, so report.pdf becomes report.txt, and a note tells you how much came out, for example “12,431 characters of text from 8 pages.”
Copying from a PDF reader is fine for a paragraph. For a 40-page contract, a set of meeting minutes or a report you want to search, count or paste into another program, selecting page by page gets tedious and often drops line breaks or picks up stray headers. This pulls the text from every page, or just the pages you name, in one go.
The output is UTF-8, so it holds any language the PDF contains, including Chinese, Japanese, Korean and Arabic. It is plain text: bold, fonts, images and layout are not kept.
How to pDF to Text
- Add your PDF. Drop it on the box above or tap to choose it. One file at a time, up to 100 MB. Password-protected files need Unlock PDF first.
- Pick the pages. Leave Pages (blank = all) empty for the whole document, or enter ranges such as 1-3, 7, 10- where 10- means page 10 to the end.
- Decide on page markers. Mark where each page starts is ticked by default and puts a line reading — Page N — before each page's text. Untick it for one continuous block.
- Extract and download. You get a .txt file that opens in Notepad, TextEdit or any editor, with a note showing the character and page count.
How the text is read
Most PDFs made from a word processor, a web page or accounting software contain a text layer: the actual characters, stored alongside instructions for where to draw them. This tool reads that layer using pdf.js, the same engine Firefox uses to display PDFs, so what you get is the real text from the file rather than a guess made from a picture of it.
Line breaks follow the PDF. Wherever the document ends a line on the page, the text file ends a line too, which means a paragraph that wrapped across six lines in the PDF arrives as six lines. If you are pasting into an email or a word processor and want flowing paragraphs, you may need to join some lines afterwards.
The page markers make long extracts easier to work with. With Mark where each page starts on, you can jump to --- Page 23 --- in a text editor and know exactly where a quote came from, which helps when you are citing a document or checking a clause against the original.
Reading order, tables and other layouts
Text comes out in the order the PDF stores it. For letters, reports, contracts and most single-column documents that is the order you would read it in. Multi-column layouts such as newsletters and academic papers are less predictable, and lines from neighbouring columns can end up interleaved.
Tables arrive as lines of text with their columns no longer lined up. The numbers and labels are all there, but you will need to rebuild the grid yourself if you want it in a spreadsheet. Headers, footers and printed page numbers are included as well, since they are part of each page’s text, so expect “Page 3 of 12” or a company name to repeat through the file.
Common uses include pasting a document into notes or an email, feeding it into another program or an AI tool, running a word count or a search across a long file, giving screen-reader users a version that reads cleanly, and keeping a plain-text archive copy that any computer will open in twenty years.
Scanned PDFs and garbled output
Scans have no text to extract. A PDF made by a scanner or from a phone photo of a page is a set of images. It looks like text, but there are no characters inside it. The tool detects this and tells you: “No text was found. This PDF is probably a scan.” A file like that needs OCR (text recognition) before any text can be pulled out of it.
Gibberish usually means a broken font map. If the output is full of wrong letters or symbols while the PDF looks fine on screen, the fonts inside it lack a proper character map, so the shapes drawn on the page do not say which letters they are. Some older or badly made PDFs have this problem. The workaround is to turn the pages into images with PDF to JPG and run those images through an OCR program elsewhere.
Locked files. A PDF that asks for a password to open cannot be read until it is unlocked. Run it through Unlock PDF with the password, then extract the text from the unlocked copy.
Questions people actually ask
How do I extract all the text from a PDF?
Add the PDF above, leave the pages box empty and press the button. You get a plain .txt file containing the text from every page, with a marker line before each one unless you untick that option. It works for files up to 100 MB, which covers documents of several hundred pages.
Why does PDF to text give me no text from my PDF?
The PDF is almost certainly a scan or a photo of a page. It contains pictures of text rather than characters, so there is nothing to extract, and the tool tells you so. You need to run it through OCR (text recognition) first to turn the images into real text.
Can I copy text from only some pages of a PDF?
Yes. Type the pages into the Pages box using commas and ranges, such as 1-3, 7, 10- where 10- means from page 10 to the end. Only those pages are extracted. Leave the box blank and the whole document is converted instead. Page numbers count from the first page of the file.
Does converting a PDF to TXT keep the formatting?
No. A .txt file holds characters and line breaks only, so bold, italics, fonts, colours, images and page layout are all lost. Line breaks follow where the PDF ends each line, and tables come out as plain lines with the columns no longer aligned. If you need the layout, keep the PDF.
Why is the text from my PDF coming out as random symbols?
The fonts in that PDF are missing a proper character map, so the shapes on the page are not linked to real letters. It happens with some older or badly made files. Convert the pages to images with PDF to JPG, then use an OCR program to read the text from those images.