Make scanned PDFs searchable and readable
Clean up a copy of your scan, add an OCR text layer, and verify the recognized text. OCR can misread numbers, names and low-quality print.
1. Decide which output you need
Use OCR to text for a plain-text extraction, or Searchable PDF to keep the scanned pages with an invisible text layer.
2. Improve the scan before recognition
Rotate sideways pages, straighten tilted scans, and try scan cleanup on a copy. Aggressive cleanup may remove faint writing, so compare with the original.
3. Run OCR
The OCR engine and language files load before recognition. They are program assets, not your documents. Large scans can take substantial memory and time; keep the tab open and start with a short file.
4. Check the result
- Search for a phrase you can see on the scan.
- Select and copy a paragraph to check reading order.
- Manually check dates, amounts, names and other important values.
- Keep the source scan for comparison.
Engine reference: Tesseract.js documentation. Check first-use downloads and local processing.