Convert a scanned PDF to text (OCR)

A digitized form, a photocopied contract, a photocopy of an archive: a PDF produced by a scanner or photocopier usually contains nothing but images of pages, with no text actually present in the file — impossible to search for a word, copy an excerpt, or have it read by a screen reader. This tool converts that kind of PDF into plain text, a Word document, or a searchable PDF that keeps the scan's appearance while becoming indexable.

First, a check that avoids unnecessary work

Before doing anything, the tool opens the PDF and checks whether it already contains extractable text: that's the case for a PDF exported directly from office software, or already processed by another OCR tool upstream. If text is found, it's extracted directly from the file, skipping image recognition entirely — faster, and above all more reliable than recognition that would reinterpret an image of text that's already present as-is in the file. A message tells you when this shortcut was taken. Only if no text is found — the case of a genuine scan, image by image — does the optical character recognition engine come into play.

Page by page, with real progress

Each page of a scanned PDF is first rasterized at high resolution (300 dots per inch, the reference for reliable recognition), then processed individually: detecting any 90, 180 or 270 degree rotation (common across a batch of pages scanned in different orientations), fine tilt correction, denoising, thresholding, then character recognition in whichever language(s) you chose. The progress bar advances page by page for real, not arbitrarily: on a document running to dozens of pages, you always know where processing stands.

Three outputs, one content

The plain text gathers every page's content in order, ready to copy. The Word document adds a page break between each source page and one paragraph per detected text block, staying close to the original structure while remaining editable. The searchable PDF reuses each page's image exactly as uploaded — stamps, signatures, letterheads included — overlaying an invisible text layer: the document looks exactly as it did before, but becomes searchable and copyable in any modern PDF reader.

Why the 20-page limit

The 20-page-per-job limit isn't arbitrary: it keeps processing time compatible with interactive use (tens of seconds rather than several minutes) on servers shared among every visitor to the site. For a longer document, splitting it into several PDFs under 20 pages with the PDF splitting tool, then processing each part, remains the simplest solution.

Frequently asked questions

How do I know if my PDF is already searchable?
Open it in a PDF reader and try selecting text with your mouse: if a word highlights, the PDF already has a text layer and OCR wouldn't add anything. This tool checks that automatically for you: if text is detected, it's extracted directly, skipping image recognition entirely, which is faster and more reliable.
How many pages can I process at once?
20 pages maximum per PDF. This limit exists to keep server-side processing time reasonable; for a longer document, split it first with the PDF splitting tool.
How long does processing a multi-page PDF take?
Each page is processed one after another (rasterizing, deskewing, recognition), with progress shown in real time. A 10-page scanned PDF of decent quality typically processes in under a minute.
Does the result keep the original PDF's layout?
The searchable PDF keeps the original scan's visual appearance exactly, page by page. The plain text and Word document only reconstruct the textual content in reading order, without complex layouts (multiple columns, tables, boxes).
My scanned PDF has pages rotated 90° — is that a problem?
No, the tool automatically detects and corrects an upside-down or quarter-turned page, page by page — a common case when a scanner digitizes both sides of a notebook regardless of orientation.
Can I run OCR on a password-protected PDF?
No, remove the password first with the PDF protection tool (the "remove a known password" feature, which requires knowing it): this tool never attempts to bypass protection.
What does the confidence score shown for each page mean?
It's the recognition engine's average confidence over the detected words on that page, from 0 to 100%. Below 60%, the page deserves a manual check: the recognized text may contain errors, especially on a poor-quality scan or an old document.
Is the uploaded PDF file kept?
No, neither the uploaded PDF nor the generated documents are kept beyond 30 minutes: no data is logged or reused, and you can trigger immediate deletion as soon as you've retrieved your result.