Recognize the text in a scan and write it back invisibly, so the document can be searched and copied from. It looks exactly the same.
A scan or photograph of a document. The kind you cannot select text from.
Nothing is uploaded. Recognition runs on your device. Up to 50 MB.
The tool checks whether there is already a text layer, and says so rather than adding a worse second copy on top.
Each page is read on your device. The first run downloads a recognition engine of about 9 MB; after that it is cached.
The recognized words are written back over the image invisibly. The page looks identical and is now searchable.
Text is drawn in a rendering mode that paints nothing. It is real text to search, selection and screen readers, and shows up nowhere on the page.
Each word is sized to the width it occupies, so selecting a line highlights the words you can see rather than drifting across the page.
Most OCR services take your document. Scans are usually the most sensitive files people have. Contracts, statements, ID pages.
A scanned page is invisible to a screen reader. A text layer makes the document readable by one, which for some readers is the difference between usable and not.
Recognized text is a best reading, not a transcript. Digits and unusual names are the most likely things to be misread. Check anything you rely on against the original. Full disclaimer.
No. The recognized text is drawn in an invisible rendering mode, so the image you see is untouched. The only change is that the words are now selectable.
On a clean scan, high, but recognition is a best reading, not a transcript. Confidence is reported after each run. Digits and unusual proper nouns are the most likely things to be misread, so check anything you rely on.
The recognition engine is about 9 MB and downloads once, on first use. Nobody using the other tools pays for it. After that it is cached by your browser.
Poorly. The engine is trained on printed text. Handwritten notes on an otherwise printed page will usually come out as nonsense, which is worth knowing before relying on the result.
English only at the moment. Other languages need their own recognition data, which would be a further download each.