What it does
PDF OCR reads text from scanned or image-based PDF pages and produces a searchable PDF or plain text.
- Recognises scanned text
OCR a scanned PDF by uploading it; the tool recognises the text in the page images and produces a searchable PDF (or plain text) with a selectable text layer. Choose the language, then download. Recognition runs in your browser, so the document is never uploaded.
Optical Character Recognition for PDFs — reading text from scanned page images and adding a searchable, selectable text layer or exporting plain text.
Related terms
OCR adds a hidden text layer aligned to the image so the scanned PDF becomes searchable and selectable.
You can choose from a range of recognition languages for best accuracy.
Recognition runs locally with Tesseract in your browser.
PDF OCR reads text from scanned or image-based PDF pages and produces a searchable PDF or plain text.
Pages are rasterised with pdf.js and passed to the Tesseract.js OCR engine, which returns recognised text positioned as a hidden layer.
Make scanned contracts, receipts, books and forms searchable, or extract their text for reuse.
Accuracy depends on scan quality, language and fonts; low-resolution or skewed scans reduce recognition quality.
OCR runs entirely in the browser, so sensitive scans never leave your device.