Extract English Text from an Image
Recognize printed English text in a PNG or JPEG with temporary server-side Tesseract OCR.
Run OCR on up to five PDF pages and download recognized English text in page order.
Selecting a file does not upload it. Upload and convert sends it to temporary server storage. One file at a time; 10 MiB maximum.
Supported formats: .pdf
OCR limit: five PDF pages, four megapixels per page and sixteen megapixels total. OCR results require human review.
Each page is rendered locally inside the server worker at 144 DPI and recognized by Tesseract. The text file separates pages with form-feed markers. This also works on supported PDFs that already contain text, but OCR can introduce errors.
Columns, reading order, tables and punctuation may not be reconstructed correctly. Blank pages contribute no recognized words. Original formatting is not retained. English printed text only. Maximum 10 MiB, five PDF pages, four megapixels per rendered page and sixteen megapixels total. Encrypted, malformed and active-content PDFs are rejected. Handwriting, unusual layouts and poor scans can produce missing or incorrect words. Review the result; no accuracy percentage is promised.
Your file is uploaded only when you choose to convert. An isolated worker processes it without network access. Uploaded files expire after one hour; completed outputs expire one hour after completion. Failed, cancelled or deleted jobs become eligible for immediate cleanup. Service outages can delay physical deletion; a previously issued download link can remain valid briefly. Clear files to request early deletion. Downloads saved on your device remain there.
PDF processing is limited to 200 pages; HTML and reverse conversions have the stricter limits listed above. Complex files may reach time, memory or scratch limits sooner. No malware-free, perfect-fidelity or certified-accessibility claim is made.
No. OCR returns plain text, not reconstructed cells or formulas.
No. Password-protected PDFs are rejected; no password guessing or bypass is attempted.
Recognize printed English text in a PNG or JPEG with temporary server-side Tesseract OCR.
Create a PDF with scanned page images and a searchable English OCR text layer.
Combine multiple PDFs into one document in your chosen order. Preserve page sizes and process your files privately in this browser.
Turn a PDF into individual page files or custom page groups. Download separate PDFs or a ZIP without uploading your document.