Make a scanned PDF searchable

It does not give you the text: it gives you back your PDF. Same pages and same look, with an invisible text layer on top lined up with the words, so Ctrl+F and copy and paste start working. All in your browser.

Tap here or drop the scanned PDF
The PDF stays in your browser: it is not sent anywhere, and the recognition does not leave here either.

What comes out of here, and what does not

What comes out of here is a .pdf file, not a text file. The pages are the same, the look is the same, the weight grows by very little: what changes is that an invisible text layer has been laid over the photographs, lined up word by word. From then on the PDF behaves like a normal one, Ctrl+F goes through it, the text can be selected and copied, and the programs that search inside it finally find something.

If what you need instead is the text and nothing else, to paste into an email or a document, the right card is a different one and it is Extract text from an image. And if you do not know which of the two cases is yours, the quickest way to find out is to try opening the file with Extract text from a PDF.

The pages are not remade: something is written on top

The easy road would be turning every page into an image and building a new PDF. That is not what happens here, and it is a choice: the file would bloat, the rendering would get worse at every step and the result would be what you already get by hand with two tools of this site one after the other. The original objects of the page are not touched, and only the text is added at the end.

The hard part is not recognising the letters, it is putting them in the right place. Between the pixels of the image and the points of the sheet there are the scale of the drawing, the rotation of the page (documents that went through a scanner are rotated nearly always) and the origin of the sheet, which is not necessarily zero. Those are the three sums that make a signature or a line of text land in the middle of nowhere, and here they are done by the same conversion already used by Write on a PDF.

The quality is measured, not promised

At the end you get one line per page with how many words were written and how sure the recognition was. It is there because the worst failure here is handing over a file that looks searchable and finds nothing: whoever files a bill notices six months later, when they search for the meter number and the computer says it is not there.

The words the recognition is not sure about are not written, and they are counted. Better one word missing than one word invented inside an archive. When a page comes out badly it is nearly always the sheet being skewed or photographed at an angle, and there you gain far more by straightening it with Straighten a photographed document than by changing any setting.

How long it takes, and what it does not do

The recognition runs inside your browser, so the time depends on your computer: a sheet costs a few seconds, and the estimate you see while it works comes from the pages already done, not promised up front. You can stop whenever you want and keep what has been done. The ceiling is sixty pages, said before you begin, and above that you cut the document with Split PDF and extract pages.

It does not recognise handwriting, it does not put tables back into columns and it does not open PDFs protected by a password, which have to be unlocked first with Remove the password from a PDF. The pages that already have text are left alone, otherwise every search would find the same word twice. Afterwards the file is ready for Search for a word inside many PDFs and for Extract tables from a PDF.