· 8 min read · Updated · By the PDFQuick editorial team
How OCR works — and how to get better results from scans
A scanned PDF looks like a document, but to a computer each page is just a photograph. You can't search it, copy from it or edit it. Optical character recognition (OCR) is the process that looks at those pixels and works out which letters they represent. PDFQuick runs OCR with the open-source Tesseract engine directly in your browser, so the scan never leaves your device.
The four stages of OCR
- Pre-processing: the page is converted to high-contrast black and white so letters stand out from the paper.
- Layout analysis: the engine finds blocks of text, then lines, then words, ignoring photos and borders.
- Recognition: each word is compared against learned letter shapes, using a neural network trained on many fonts.
- Language modelling: a dictionary for the chosen language helps decide between similar shapes, such as “rn” and “m”.
Why results go wrong
- Low resolution: below about 200 dpi, small letters merge into blobs.
- Skewed or curved pages: photographs of book pages bend lines, confusing line detection.
- Shadows and uneven lighting: dark corners turn into noise after contrast is boosted.
- Wrong language selected: accented characters and common words are guessed incorrectly.
- Handwriting: general OCR engines are trained on printed type and rarely read handwriting well.
How to scan for better OCR
- Scan at 300 dpi. Higher rarely helps and makes processing slower.
- Use a flatbed or a scanning app that flattens and crops pages automatically.
- Photograph in even daylight with the phone held parallel to the page.
- Choose greyscale rather than colour for text documents — it gives cleaner contrast.
- Select the correct language in the OCR tool before starting.
What to expect from the output
On a clean, straight, printed page, modern OCR typically gets the vast majority of characters right. On a faded photocopy, accuracy can drop considerably. Always proofread names, numbers and dates — these are the errors that matter and that a dictionary can't correct, because “1O” and “10” are both plausible.
OCR on your own device is slower than a cloud service with dedicated servers. A 20-page document may take a few minutes on a phone. The trade-off is that a confidential scan is never uploaded.
After OCR
Once you have a Word document, restyle headings and fix the few recognition mistakes, then save. If you only need the words — for example to paste into notes — PDF to Text is faster once the document has a text layer.