Extract English Text from an Image
Recognize printed English text in a PNG or JPEG with temporary server-side Tesseract OCR.
Create a PDF with scanned page images and a searchable English OCR text layer.
Selecting a file does not upload it. Upload and convert sends it to temporary server storage. One file at a time; 10 MiB maximum.
Supported formats: .pdf
OCR limit: five PDF pages, four megapixels per page and sixteen megapixels total. OCR results require human review.
Pages are rendered at 144 DPI in source order. Tesseract produces page images with an invisible recognized-text layer, and the worker validates the combined PDF and page count before download.
This is a rasterized OCR copy: original vector content, bookmarks and document structure are not retained. Cleanup can change contrast, orientation and edges; turn it off to reduce visual changes. Text-layer alignment and recognition are imperfect. This is not PDF/A certification or accessible-PDF tagging. English printed text only. Maximum 10 MiB, five PDF pages, four megapixels per rendered page and sixteen megapixels total. Encrypted, malformed and active-content PDFs are rejected. Handwriting, unusual layouts and poor scans can produce missing or incorrect words. Review the result; no accuracy percentage is promised.
Your file is uploaded only when you choose to convert. An isolated worker processes it without network access. Uploaded files expire after one hour; completed outputs expire one hour after completion. Failed, cancelled or deleted jobs become eligible for immediate cleanup. Service outages can delay physical deletion; a previously issued download link can remain valid briefly. Clear files to request early deletion. Downloads saved on your device remain there.
PDF processing is limited to 200 pages; HTML and reverse conversions have the stricter limits listed above. Complex files may reach time, memory or scratch limits sooner. No malware-free, perfect-fidelity or certified-accessibility claim is made.
The worker checks that the output has the same number of pages and processes them in source order. Rendering and optional cleanup may change visual details.
No. OCR text and its placement can be inaccurate, especially with skew, noise, small text and complex layouts.
Recognize printed English text in a PNG or JPEG with temporary server-side Tesseract OCR.
Run OCR on up to five PDF pages and download recognized English text in page order.
Combine multiple PDFs into one document in your chosen order. Preserve page sizes and process your files privately in this browser.
Turn a PDF into individual page files or custom page groups. Download separate PDFs or a ZIP without uploading your document.