7 Ways to Get Better OCR Results from Scanned PDFs
OCR can feel like magic or like a coin flip β and the difference is almost always the input. Here are seven things that decide whether a scanned PDF search finds everything or misses half of it.
Optical Character Recognition reads the shapes on a page image and rebuilds them into text. Feed it a sharp, straight, well-lit page and modern engines are remarkably accurate. Feed it a dark, skewed phone photo and even the best engine will guess. The good news: most of what matters is under your control before you ever upload.
1. Scan at 300 DPI when you can
Resolution is the single biggest lever. At 150 DPI, small print starts losing the details that distinguish an 8 from a B. 300 DPI is the sweet spot for documents; going beyond 400 rarely helps and makes files heavier.
2. Keep the page straight
A few degrees of tilt is survivable; a 15Β° skew turns clean lines of text into diagonal noise. If your scanner has auto-deskew, leave it on. Photographing a document? Shoot from directly above, not at an angle.
3. Light the page evenly (phone photos)
Shadows across a photographed page are OCR poison β the engine reads the shadow edge as ink. Natural side light or a plain desk lamp beats an overhead light that puts your phone's shadow on the paper. Flash tends to blow out a stripe of the page; avoid it.
4. Don't crush the file before OCR
Heavy JPEG compression saves space by discarding exactly the fine edges OCR needs. If you plan to search a document, run OCR on the best copy you have first β you can always compress the PDF afterwards.
5. Crop the clutter
Staple shadows, binder rings, your thumb at the page edge β anything that isn't the document adds noise. A quick crop of the margins before OCR gives the engine less to misread.
6. Grayscale beats bad color
A crisp grayscale scan usually outperforms a washed-out color one, because OCR works on contrast between ink and paper. If your scanner offers a "text / black & white document" mode, that's the one built for this job.
7. Use a search that forgives OCR mistakes
Even with a perfect scan, engines confuse look-alikes: Oβ0, Iβ1, Sβ5, Bβ8. This is where the search tool matters as much as the scan. PDF Everyday's OCR search folds those confusable characters into the same class on both sides of the comparison, and ignores dots, dashes and spaces inside codes β so INV-2024-OO87 still matches even if the scanner read it as INV-2024-0087. You can read exactly how that works in our OCR explainer.
π Test your scan right now
Upload any scanned PDF and search it β results stream in live, page by page, with counts.
Search a scanned PDF β free βFrequently asked questions
What DPI should I scan at for OCR?
300 DPI is the sweet spot for typical documents. Below 200 DPI small print degrades quickly; above 400 DPI gains are minimal and files get heavy.
Can I OCR a photo taken with my phone?
Yes β shoot from directly above in even light, avoid flash and shadows, and make sure the text is in focus.
My scan is low quality. Is it hopeless?
Not necessarily. A tolerant search that folds commonly confused characters (O/0, I/1, S/5) can still find codes in imperfect scans.
Does PDF Everyday store my scans?
No. Files are processed in memory and deleted immediately after the result is generated.