Make a scanned PDF searchable

Recognize the text in a scan and write it back invisibly, so the document can be searched and copied from. It looks exactly the same.

Drop a scanned PDF here

A scan or photograph of a document. The kind you cannot select text from.

Nothing is uploaded. Recognition runs on your device. Up to 50 MB.

How it works

Step 1

Open your scan

The tool checks whether there is already a text layer, and says so rather than adding a worse second copy on top.

Step 2

Recognize

Each page is read on your device. The first run downloads a recognition engine of about 9 MB; after that it is cached.

Step 3

Download

The recognized words are written back over the image invisibly. The page looks identical and is now searchable.

What it handles

A picture of a page becomes a document

01

Invisible, not overlaid

Text is drawn in a rendering mode that paints nothing. It is real text to search, selection and screen readers, and shows up nowhere on the page.

02

Positioned to the ink

Each word is sized to the width it occupies, so selecting a line highlights the words you can see rather than drifting across the page.

03

Nothing uploaded

Most OCR services take your document. Scans are usually the most sensitive files people have. Contracts, statements, ID pages.

04

Accessible too

A scanned page is invisible to a screen reader. A text layer makes the document readable by one, which for some readers is the difference between usable and not.

Questions

Will the page look different?

No. The recognized text is drawn in an invisible rendering mode, so the image you see is untouched. The only change is that the words are now selectable.

How accurate is it?

On a clean scan, high, but recognition is a best reading, not a transcript. Confidence is reported after each run. Digits and unusual proper nouns are the most likely things to be misread, so check anything you rely on.

Why is the first run slow?

The recognition engine is about 9 MB and downloads once, on first use. Nobody using the other tools pays for it. After that it is cached by your browser.

What about handwriting?

Poorly. The engine is trained on printed text. Handwritten notes on an otherwise printed page will usually come out as nonsense, which is worth knowing before relying on the result.

Does it work on other languages?

English only at the moment. Other languages need their own recognition data, which would be a further download each.