Skip to content

OCR that runs on your own device

Seven tools that read text, codes and barcodes out of pictures. Recognition happens inside this page — the reader and the language pack come from this site, and the picture itself never leaves your device.

7 tools, no signup, no watermark

Most of these do one thing to one picture. To copy the words out of a photograph, start with Image to Text; to get at text you cannot select on screen, Screenshot OCR; to make a scan searchable without changing how it looks, OCR PDF. The notes below the list cover the parts that surprise people — why resolution matters more than file size, and where a browser is the wrong tool.

Codes and Barcodes

2 tools

See what a QR code says before you follow it, and get the number off a barcode with its check digit verified.

What OCR actually does, and what it cannot

OCR recognises shapes and returns characters. It does not read. There is no understanding of the words behind it, so a "1" mistaken for an "l" comes back looking exactly as certain as every correct character around it. That is why these pages show the reader's own confidence beside the text instead of handing you a clean-looking result and leaving you to find the errors. On clear printed text, photographed square and in focus, the error rate is a few characters in a thousand. On glare, a shadow or an unusual typeface it can be very much worse, and nothing about the output will look any different.

There are two readers behind these pages. A neural one finds text wherever it is on the page with no idea what script it is in, and reads fifty-odd languages from a single dictionary. Tesseract 5, with a separately trained model per language, sits behind it and covers over 120. Leaving the language set to automatic lets the neural reader try first, which matters because naming the wrong language is the classic way to get nothing back — a Polish menu identified as Cyrillic and read with the Russian model returns an empty page, confidently. Naming a language yourself pins the job to that Tesseract model, which is what you want when you know exactly what you have.

Resolution is what matters, not file size. The engine wants a certain letter height and everything is scaled to give it one: pages of a scanned PDF are rendered at 300 DPI for reading, and a picture is measured from the ink itself rather than from its dimensions, then enlarged or shrunk to suit. Screen text is the case people get wrong — at normal size a letter is about ten pixels tall, roughly a third of what the reader wants, so a screenshot is scaled up before it is read. Below about eight pixels a letter there is no detail left to enlarge, and no amount of processing invents it. In the other direction, a 48-megapixel photograph of a menu is no more accurate than a sharp small one; past a point more pixels cost time and buy nothing.

Skew and uneven lighting destroy more than they look as though they should. A photographed page is a degree or two off square and lit from one side, so Image to Text and Scan Document measure that angle across the whole page and remove it, then estimate the brightness of the paper behind the ink in patches and lift it to a uniform white. That is not cosmetic. Straight ruling lines are what lets a table be found as a grid at all, and even contrast is what stops the shadow of your hand being read as ink. A screenshot gets neither treatment, deliberately: it is already square and evenly lit, and conditioning it as though it were a photograph only adds error.

Tables and multi-column layouts are the genuinely hard case, and the numbers are worth knowing. Read as whole-page prose, a ruled form ran cell text into the next column and dropped a third of the labels — fifteen to twenty per cent of characters wrong. Finding the ruling lines first and reading each cell on its own brought that to about five per cent. So a table with visible rules is found as a grid and comes back as tab-separated rows that paste straight into a spreadsheet, and columns of prose are read in reading order rather than straight across the page. A table laid out with spacing alone has no rules to find, and survives the paste by luck rather than by design.

A searchable PDF and extracted text are different things. OCR PDF and Scan Document leave the picture exactly as it was and lay the recognised words invisibly over it: the page looks identical, Ctrl+F starts finding things, and you can select and copy. Nothing is ever painted onto the page — which also means a misreading hides inside a document that looks perfect. Image to Text and Screenshot OCR do the opposite: you get the characters and nothing else, with no typefaces, no colours, no bold and not the picture. One caveat on the text layer, since it is easy to miss — it can carry only characters a bundled font can write, so Chinese, Japanese and Korean words are counted for you rather than written, until a font for them ships.

Where a browser is genuinely the wrong tool

The picture never leaving your device is the whole argument for reading it here, and it costs accuracy on hard documents. That trade is real, and worth saying plainly rather than burying.

Joined-up handwriting is not recognised — here, or by any general OCR engine. These are printed-type engines: they recognise letter shapes, and cursive has no separate letter shapes to recognise. Hand-printed letters do often come back well enough to correct rather than retype — block capitals, a form filled in by hand, a carefully lettered note — which is why Handwriting Recognition exists and why it puts the confidence figure next to the text. Reading cursive properly needs a model trained on handwriting specifically, and the good ones run on a server rather than in a tab.

Cloud OCR will beat this on hard documents. Google, Amazon and Microsoft run recognisers trained on far more data than fits in a browser tab, and on a poor photograph, a dense multi-column page, an unusual typeface or a filled-in form they will be noticeably more accurate. So will a paid desktop package such as ABBYY FineReader. If accuracy on a difficult document matters more than the document staying on your machine, use one of those — that is the honest comparison, and the machine you should make the choice on.

Batch work belongs on a command line. These read one document when a person clicks. Two thousand scans on a schedule is a job for tesseract's own command-line build, or for OCRmyPDF, which wraps it and adds a text layer to a whole folder of PDFs in a single command. Both are free, and both are the same sort of engine doing the work with none of the clicking.

Some things are simply not done. Perspective is not corrected: a page photographed from the side keeps its trapezoid shape and only the rotation is fixed, so shooting from directly above is worth the extra second. A heavy fold, or a curve near a book's spine, stays curved. And there is no camera here — the code readers decode a picture you already have, not a live feed.

One practical cost of running locally: the reader and the language pack are fetched from this site the first time you use a language, which is a few megabytes and makes the first run slower than the ones after it. Nothing comes from anybody else's CDN, and nothing goes back.

Choosing between the tools that sound alike

Image to Text vs Screenshot OCR. The same reader, conditioned differently, and the difference is not cosmetic. A photograph is straightened and its lighting levelled; a screenshot is enlarged and otherwise left alone, because it has no skew and no uneven lighting and treating it as though it did makes it worse. Screenshot OCR also takes a paste — Ctrl+V, or Cmd+V on a Mac, straight from the clipboard, with nothing saved to disk first.

Image to Text vs Scan Document. Image to Text gives you the words and throws the picture away. Scan Document keeps the picture, straightens it, whitens the paper behind the ink and gives you a PDF that looks scanned — optionally with the text laid invisibly behind it. If you want to paste the text somewhere, the first. If you want to send somebody the document, the second.

Scan Document vs OCR PDF. Scan Document starts from photographs and produces a PDF. OCR PDF starts from a PDF that already exists and adds a text layer to it. Neither changes how the pages look. Photograph a contract and you want the first; receive a scan by email and you want the second.

Handwriting Recognition vs Image to Text. The same engine, with the table finder switched off and the confidence figure given the prominence it deserves, because with handwriting that number is the useful part of the result. If the writing is joined up neither page will read it, and the handwriting page says so rather than returning plausible nonsense.

QR Code Reader vs Barcode Reader. Square codes against striped ones, and it is worth knowing which you are holding. The QR page reads QR, Micro QR, Data Matrix, Aztec and PDF417 — which covers boarding passes and driving licences, whose codes are square-format even when they look like a smear of dots. The barcode page reads the striped retail and logistics formats — EAN, UPC, Code 128, Code 39, ITF — and checks the check digit, the last digit calculated from the others, which is the difference between a number you can trust and one that merely looks right. Neither is OCR: a code is a symbol with a known geometry, so decoding it is arithmetic rather than a guess about letter shapes. They fail differently too. Square codes carry error correction and survive a scratch or a logo printed over the middle; striped ones carry none at all, so glare across the bars means no read rather than a wrong number.

What these tools are built on

Open-source engines doing the actual work, and the specifications they implement.

  • TesseractApache-2.0 — recognises text in 120-plus languages, with the language packs served from this site.
  • ZXingApache-2.0 — reads QR codes and barcodes.
  • QR code specification (Denso Wave)Specification — publishes the version sizes and error-correction levels, from the format's inventor.

Frequently asked questions

Is my picture or document uploaded?

No. Recognition runs in your browser. The reader and the language pack are served from this site the first time you use a language, and the picture itself never leaves the device — so it works with the Wi-Fi off after that first load, and nothing is stored anywhere. It matters most on the code pages, where a QR code can hold a Wi-Fi password, a payment request or a two-factor secret.

How accurate is it?

On clear printed text, photographed square and in focus, a few characters in a thousand. Accuracy falls with blur, glare, low contrast, unusual typefaces and anything printed over a picture. The confidence figure is shown beside the text rather than left for you to guess at, and when it is low the picture is almost always the problem: more light, less angle, and filling the frame with the text fix most of it.

Can it read handwriting?

Separated, hand-printed letters often, joined-up writing no. That is not a limitation of this site: general OCR engines are trained on printed type and recognise letter shapes, and cursive has no separate letter shapes to recognise. Handwriting Recognition is the page that says so plainly and shows you what it managed, which is more use than a confident wrong answer.

Which languages can it read?

Over 120, including English, Bengali, Hindi, Urdu, Arabic, Spanish, French, German, Portuguese, Russian, Chinese, Japanese and Korean. The script is worked out from the picture itself; where several languages share a script — Latin, Cyrillic, Arabic, Devanagari — your browser's own language settles it, and you can always change it.

Does it use my camera?

No, and there is no permission to grant. Every tool here reads a picture you already have — a photo from the camera roll, a screenshot, a file somebody sent you. If the code or the page is in front of you, photograph it first and drop the photo in.

Will a table come out as a table?

If it has visible ruling lines, yes: the grid is found first and each cell is read on its own, so a label does not run into its value, and the result is tab-separated rows that paste straight into a spreadsheet. A table laid out with spacing alone and no rules is read as columns of text in reading order, which usually survives the paste but is not guaranteed.