PDF to Word
Load a PDF and download a .docx file you can really edit: lines become paragraphs again, split words become whole again and headings become Word styles. The PDF never leaves your device.
Why a PDF cannot be edited like a document
A PDF does not know what a paragraph is. Inside there are only instructions like «write these words at this height, with this font»: one line after another, each in its place. That is why, when you copy a piece of a PDF and paste it into Word, you end up with a line break at the end of every line, words split by a hyphen, the page number in the middle of a sentence and the header repeated every forty lines.
This page does the work you would do by hand, only faster: it looks at where every piece of text sits, rebuilds paragraphs, headings, lists and tables and writes them into a real .docx file, which you edit like any other document. The text flows from one page to the next, can be corrected, restyled and printed with the margins you like.
What it rebuilds
Paragraphs. When a line ends before the margin and the first word of the next line would have fitted, the paragraph ended there; otherwise the lines are joined. First line indents, the hanging indent of bibliographies and the extra space between paragraphs count as well.
Words split by a hyphen at the end of a line become whole again: «docu-» and «ment» make «document». The hyphen stays when the compound word appears whole elsewhere in the document, or when the next line starts with a capital letter or a digit, as in «Sub-Saharan».
Headings, recognised by the size of the type (the largest is level 1) and by isolated lines set entirely in bold. They become Word's Heading 1, 2 and 3 styles, so the automatic table of contents and the navigation pane work straight away.
Lists with bullets or numbers, on two levels too, as real Word lists. A number at the start of a line becomes a list only if the list starts at 1 or if it is the number after the previous one; a dash only if there are at least two in a row, because a dash at the start of a line is often dialogue.
Simple tables, the ones with well separated columns, which become Word tables with the first row as the header. And inside sentences bold and italics stay, while footnote numbers become superscripts again.
The things Word does not need
Headers and footers. A line repeated identically at the top or bottom of several pages is removed, even when only the number changes: «Page 3 of 12» and «Page 4 of 12» are the same line. Lone page numbers at the bottom of the sheet go too.
Page break line endings. A paragraph that starts on one page and ends on the next becomes a single paragraph again, even when a footnote sits in between, and the footnote stays right after the paragraph. On two column pages the whole left column is read first and then the right one, as anyone would: reading from edge to edge would glue together half sentences that have nothing to do with each other.
An example with numbers. On the test document we check the tool with, two pages of justified and hyphenated text, sixteen split words become whole again, the four header and footer lines disappear and the paragraph that runs from one page to the next becomes one again. After the conversion the page tells you the same counts for your own file, so you know what was touched.
What cannot stay the same
A Word document and a PDF are built differently, and nobody can honestly promise an identical layout. Those who promise it put every line in a text box: the file looks the same, but it cannot be edited, because fixing one word means moving the boxes below by hand. Here the opposite choice is made: the text flows and can be edited, and the graphics are not copied. So no exact positions, colours, boxes or page breaks in the same spot.
Images do not go into the document: to take them out of the PDF, one by one and at their original resolution, there is Extract images from PDF and Office. Formulas and charts come out as scattered text, because in the PDF they are pieces of text placed in different spots. Tables with merged cells or with text on several lines inside a cell are rebuilt badly: for those there is Extract tables from a PDF, which does only that job.
Scans and protected PDFs
If the PDF is a photograph of a sheet, there is no text inside: there are images, and no words come out of an image without character recognition. The page notices and tells you which pages are like that; if the whole document is, it does not give you an empty Word file, it sends you to Make a scanned PDF searchable, which puts the recognised text under the images, and after that the conversion works.
If the PDF asks for a password to open, nothing is guessed here: if you know it you remove it with Remove a PDF password, and then you come back here. PDFs with restrictions only, the ones that open but cannot be printed or copied, are read without any problem.
The Word file you get
It is a .docx file like any other, A4 sized with the Calibri font, which opens in Word, LibreOffice, Pages and Google Docs. Headings use Word styles, lists are real lists and tables are real tables, with the header row repeated at the top of every page. This page writes it, byte by byte, with the same code as Markdown to Word, which does the same job starting from Markdown text.
Before being published, the tool was tested on PDFs with the truth written alongside: the Word file it produces is read back by two programs other than ours, which must find exactly the same paragraphs, the same headings at their levels, the same lists and the same cells as the original document.
What it does not do
It does not read scans without text, it does not copy the layout, images and charts, and it does not go the other way: from Word to PDF, Word already does it with «Save as». If you only need the plain text, without headings or lists, to paste it into an email, the quicker route is Extract text from PDF.
The PDF does not leave your device: it is read and converted here, in the browser, and the Word file is made here. You can use it for a contract, a payslip or a medical record too.