Skip to content

Image to text (OCR)

Read the text in a photo, scan or screenshot, then copy it or save it as a .txt file. The OCR runs on your device.

    The size shown is a one-time download from this site, cached by your browser for next time. Your file is never uploaded.

    Your result

    Add an image to see how many words were found and how confident the OCR is.

    How the text is read

    This tool uses Tesseract.js, the WebAssembly build of the open-source Tesseract OCR engine, with the tessdata_fast language models (both Apache-2.0). The engine runs in a background worker in your browser. It finds lines of text, splits them into words, and gives each word a confidence score from 0 to 100.

    Before reading, small images are enlarged, because Tesseract needs letters to be a decent number of pixels tall. The rule this page uses:

    Scale = 2 if the longest side is 2,000 px or less, otherwise 1, then reduced if needed to stay under 25 megapixels (8 on phones). Confidence shown = average of all word confidences.

    Worked example: a laptop screenshot

    A 1280 × 720 screenshot of an email has a longest side of 1,280 px, so it is doubled to 2560 × 1440, which is 3.7 megapixels, well under the cap. Body text that was 12 px tall becomes 24 px tall. Say Tesseract finds 180 words, 172 of them at 93% and 8 short ones (like "Re:" and an email address) at 61%. The average is (172 × 93 + 8 × 61) ÷ 180 = 91.6%, and the word-box view shows those 8 in red so you know what to proofread.

    Getting better results

    Most bad OCR comes from the picture, not the engine. The Tesseract documentation recommends at least 300 DPI for scans, straight text lines, and dark text on a light background. In practice:

    ProblemWhat to do
    Phone photo of a pageHold the phone flat above the page, fill the frame, use daylight, and avoid shadows from your hand. Tilted pages read badly because Tesseract's line finding gets much worse when text is skewed.
    Scanner settings300 DPI, grayscale or black and white. 150 DPI scans lose small print.
    White text on a dark backgroundInvert the colours first. Current Tesseract versions expect dark text on a light background.
    Only part of the image has textCrop to the text with a small margin. Logos, photos and borders can be misread as letters.
    Wrong letters in another languageChoose that language. English data reads "ñ" or "ü" poorly, and reads Hindi or Arabic not at all.
    Sideways or upside-down imageRotate it upright first. This tool does not detect page orientation.

    What it is not good at

    Accuracy depends on image quality, so always proofread before you rely on the text. Tables come out as lines of text without columns. Handwriting, stylised logos, very small print and text on busy photos are hit and miss. The fast language models trade a little accuracy for speed; on clean printed text the difference is small.

    Privacy

    Pasted screenshots often contain messages, addresses and account numbers. Here they are read by your own browser: the page downloads the OCR engine and language data from this site, and no image is sent anywhere. If your image is a scanned PDF rather than a picture, use OCR PDF to make the PDF itself searchable.

    Questions people ask

    Is my image uploaded to read the text?

    No. The OCR engine (Tesseract, compiled to WebAssembly) is downloaded from this site into your browser, and it reads the image on your device. The picture never leaves your computer or phone. After the first use you can even turn off Wi-Fi and it keeps working, because the engine and English data are cached.

    How accurate is it?

    On a sharp, straight scan or screenshot of printed text, expect nearly every word to be right, with a confidence score in the 90s. Phone photos at an angle, low light, fancy fonts and handwriting do much worse: handwriting is often unreadable. The confidence score and the coloured word boxes show which words to double-check.

    Can it read handwriting?

    Only neat, printed-style handwriting, and even then with mistakes. Tesseract is trained on printed text. For cursive notes, typing them is usually faster than correcting the OCR output.

    Which languages can it read?

    English, Spanish, French, German, Portuguese, Hindi, Arabic, Chinese (Simplified) and Urdu. Pick the language before reading. Non-English data files are 0.6 to 1.7 MB and download only when you choose them.

    How do I paste a screenshot?

    Take the screenshot (Windows: Win+Shift+S, Mac: Cmd+Ctrl+Shift+4 copies to the clipboard), click anywhere on this page and press Ctrl+V (Cmd+V on a Mac). The image is added and read straight away.

    Why is the first run slow?

    The first time, your browser downloads the OCR engine (about 3.7 MB) and the language data (about 1.9 MB for English). After that they come from the browser cache, and a typical screenshot is read in 1 to 4 seconds on a laptop.