Guide
Copying Text Out of a PDF: Why It Comes Out Jumbled
You select a paragraph in a PDF, paste it somewhere else, and get its lines in the wrong order, words cut in half, or a string of symbols. It is easy to assume the tool you used is broken. Usually it is not — the PDF simply never stored the text the way you see it.
A PDF is a drawing, not a document
The format was designed to look identical everywhere. To do that, a PDF records instructions such as "draw these letters at these coordinates in this font." There is no built-in notion of a paragraph, a column, or a reading order. Software that extracts text has to reconstruct all of that from positions, and it works by reading the instructions in the order the file lists them.
For a single-column page written top to bottom, that order usually matches what you read. For anything else it can drift. A two-column article may come out with the first line of column one followed by the first line of column two. A table may come out cell by cell in whatever order the exporting program drew them. Headers, footers and page numbers appear on every page, right in the middle of your text.
The usual ways extraction goes wrong
Line-end hyphenation is the most common one. If a word was broken across two lines with a hyphen, the hyphen is real text in the file, so it comes out as "docu-" on one line and "ment" on the next. Ligatures (the combined shapes for letter pairs like fi or fl) can arrive as one odd character instead of two letters, depending on how the font was built. Extra spaces or missing spaces appear because spacing in a PDF is often just a gap between two drawn pieces, not a space character.
The nastiest case is a font without a proper mapping from glyphs back to characters. The page looks perfect, but the text underneath is a set of internal codes, so what you copy is symbols or nonsense. Nothing on the receiving side can repair that, because the real characters were never stored. The fix is to go back to whatever produced the file and export it again.
When there is no text at all
If you select and nothing highlights, or extraction returns nothing, the page is probably a picture — a scan or a photograph of a document. There is no text layer to read. Scanned vs text-based PDFs explains a five-second test for telling the two apart, and what follows for search, size and accessibility. Getting text out of a picture needs OCR, which is a separate technique and something HerePDF does not offer.
What the extract text tool does
HerePDF's extract text tool reads the text layer with PDF.js inside your browser and shows it as one combined view that you can copy or download as a .txt file. Pages are separated by blank lines. A page with no extractable characters is marked inline as "no extractable text" instead of being silently skipped, so you can see exactly which pages were pictures. It does not try to reconstruct columns or tables — the order you get is the order the file stores.
Because the file is read locally and never uploaded, this is safe to use on documents such as contracts or medical letters that you would not want to paste into an online converter. The broader reasoning is in why client-side PDF tools are safer for sensitive documents.
Why selecting text in a viewer behaves the same way
The copy-and-paste problem and the extraction problem are the same problem. When you drag a selection across a page in a PDF viewer, the viewer is doing what an extraction tool does: working out which pieces of text lie under your cursor and in which order to join them. Good viewers guess well on simple pages and badly on complicated ones, which is why a selection across two columns can grab text from both or jump between them. It is not that the file is corrupt, it is that the file never said which piece follows which.
A worked example
Take a two-page newsletter with a headline, two columns of body text, a sidebar and a footer with the page number. The exporting program may have drawn the headline, then the sidebar, then column two, then column one, then the footer, because that was the order the layout software processed its frames. Extract that page and you get precisely that sequence. The text is all there and none of it is wrong — it is just in the order it was drawn. Reading it takes about a minute of cut and paste, far less effort than retyping it.
Text that is present but cannot be trusted
There is one more category: a PDF that looks like a scan but has a hidden text layer added by an OCR program. The picture is what you see, and the text underneath is the OCR program's guess at what the picture says. It is searchable and extractable, but it inherits every misreading the OCR made — a 0 for an O, a rn for an m, a missing line. If the numbers or names matter, check the extracted text against the page itself instead of assuming they match.
Using the result
Extracted text is best for search, quoting, feeding into a spreadsheet or a translation tool, and checking whether a document says what you think it says. It is a poor route to reproducing the layout: tables, footnotes and figures do not survive the trip in any usable form. If you need the layout itself, work from the original source document, or rebuild the pages you need with the organize tool and keep them as a PDF.
Getting cleaner results
Extract first, then tidy: join the hyphenated line breaks and remove repeated headers and footers with a find-and-replace pass in a text editor. If a page has two columns, it often helps to extract, then reorder the blocks manually. If the text is garbled rather than merely disordered, check with the source's author whether a fresh export or the original document is available — that is a far better route than trying to decode the codes.