BestPDFInTheWorld
Convert · Premium

OCR a scanned PDF

A scanned PDF is a stack of photographs: you cannot search it, copy from it or select a single line. OCR reads what is in those images and gives you back the same document with an invisible text layer on top, or, if you prefer, just the plain text in a .txt file. Everything runs in your browser: the scan never leaves your computer, which is rare for an OCR tool.

How it works

Three steps, nothing to install.

The whole job happens on your computer, start to finish.

1

Choose the file

Drag the scanned PDF into the tool window, or an image of a document.

2

Choose the language

Portuguese, English or Spanish. This is the language of the document, not of the interface, and choosing wrong degrades the result badly.

3

Download the result

A searchable PDF with the original pages, or the plain text as .txt.

What exactly is a searchable PDF?

It is the original document, pages untouched, plus a text layer on top that the reader does not draw. Nothing new appears on the page, but your PDF reader's search finds the words, the mouse selects lines and copy and paste works. It is also how search engines and internal archives start finding the document.

Are my documents sent to a server?

No, and here that matters more than anywhere else on this site. Text recognition is the service that most often forces you to upload the file to someone else's machine, and a scan is usually exactly what you do not want to upload: a contract, a receipt, a letter from your doctor. Here the recognition engine is downloaded into your browser and works on your machine.

Why is the first run slower?

Because the first time around your browser downloads the recognition engine and the language model you picked, which together are several megabytes. It is then kept in the browser: next time it starts straight away. Switch language and it downloads that language's model, and only that one.

How long does a document take?

It depends on the machine and the quality you pick, but expect a few seconds per page. The tool reports progress page by page so you can see it moving. Long documents genuinely take a while: that is the price of reading letter by letter.

Your files never leave the browser. The recognition engine comes to them, not the other way around.
Honest limits

The result is only as good as the scan.

A straight page, decently lit and at a decent resolution, reads back almost perfectly. A photograph taken at an angle, with a shadow across it, or scanned at low resolution, gives text with errors, sometimes plenty of them. If you can scan it again, do: it pays off more than any option in this tool.

It only reads the three languages it has a model for.

Portuguese, English and Spanish. A document in French or German will be read with whichever model you pick, and the result will be poor. A document mixing two languages is read as the one you pick, and the other comes out worse.

It does not read handwriting.

This recognises printed type. A signature, a handwritten note in the margin or a form filled in by hand are not recognised, and whatever comes out of them is noise.

The text layer is an approximation.

Words land on top of the words in the image, but the selection box can sit slightly off, especially with very small type or on skewed pages. That does not affect search: it affects the rectangle your reader highlights when it finds the word.

Tables and columns can come out scrambled.

Recognition reads lines of text, not the structure of the page. A two column document, or a dense table, often comes out in the wrong order in the .txt file. In the searchable PDF that does not show, because each word stays where it belongs.

Questions

Questions about PDF OCR.

Do I need an account to use OCR?

Yes, this is a Premium tool and it requires a subscription and signing in.

Can I use a photograph instead of a PDF?

You can. The tool accepts PNG, JPEG and WebP as well as PDF. A straight, well lit photograph gives a result; one taken at an angle gives little.

My PDF already has text. Do I need OCR?

No. If you can select a word in your PDF reader, the text is already there: use Extract text instead, which is instant and invents nothing. OCR is for documents where you cannot select anything at all.

Which quality should I pick?

Normal covers most cases. High helps with small type or weak scans, and takes considerably longer. Fast is for checking quickly whether a document is worth processing.

Are my files kept on your servers?

No. Since nothing is sent, nothing is kept outside your browser.

A scan you
cannot search?

Open the tool and try it. Nothing to install, nothing sent anywhere.

Open OCR