Convert a scanned PDF to text (OCR)
A digitized form, a photocopied contract, a photocopy of an archive: a PDF produced by a scanner or photocopier usually contains nothing but images of pages, with no text actually present in the file — impossible to search for a word, copy an excerpt, or have it read by a screen reader. This tool converts that kind of PDF into plain text, a Word document, or a searchable PDF that keeps the scan's appearance while becoming indexable.
First, a check that avoids unnecessary work
Before doing anything, the tool opens the PDF and checks whether it already contains extractable text: that's the case for a PDF exported directly from office software, or already processed by another OCR tool upstream. If text is found, it's extracted directly from the file, skipping image recognition entirely — faster, and above all more reliable than recognition that would reinterpret an image of text that's already present as-is in the file. A message tells you when this shortcut was taken. Only if no text is found — the case of a genuine scan, image by image — does the optical character recognition engine come into play.
Page by page, with real progress
Each page of a scanned PDF is first rasterized at high resolution (300 dots per inch, the reference for reliable recognition), then processed individually: detecting any 90, 180 or 270 degree rotation (common across a batch of pages scanned in different orientations), fine tilt correction, denoising, thresholding, then character recognition in whichever language(s) you chose. The progress bar advances page by page for real, not arbitrarily: on a document running to dozens of pages, you always know where processing stands.
Three outputs, one content
The plain text gathers every page's content in order, ready to copy. The Word document adds a page break between each source page and one paragraph per detected text block, staying close to the original structure while remaining editable. The searchable PDF reuses each page's image exactly as uploaded — stamps, signatures, letterheads included — overlaying an invisible text layer: the document looks exactly as it did before, but becomes searchable and copyable in any modern PDF reader.
Why the 20-page limit
The 20-page-per-job limit isn't arbitrary: it keeps processing time compatible with interactive use (tens of seconds rather than several minutes) on servers shared among every visitor to the site. For a longer document, splitting it into several PDFs under 20 pages with the PDF splitting tool, then processing each part, remains the simplest solution.