Page text

Making a scanned sheet searchable, and repairing text that came out badly.

A PDF that was exported from CAD already carries its text, and everything that reads a sheet - search, title detection, the text markup tools - works straight away. A scan is a picture of a drawing with no text in it at all, and none of that works until the text is recovered.

OCR reads the image and puts a text layer behind it. Find it at Document > Manage Pages > Tools > OCR.

The three passes

PassWhat it doesWhen
Patch Missing TextOCRs only the pages that have no text layer.The usual one. A set that is mostly CAD exports with a few scanned sheets dropped in.
Enhance Poor TextRe-OCRs pages whose text is low-quality or garbled.The sheet has text, but searching it finds nothing and the title detection reads nonsense.
Re-OCR All PagesDiscards existing text and re-runs OCR on every page.Last resort. It throws away good text along with bad.

Patch Missing Text is the one to reach for first: it is the cheapest, it cannot damage text that is already good, and on a typical set it is the only one you need.

After OCR

Recovered text feeds everything else that reads the sheet, so the order matters:

  1. OCR the pages that need it.
  2. Auto Detect the titles and page numbers.
  3. Then work the set.

OCR itself is available on every subscription. Step 2, the automatic reading of title blocks, is part of the Advanced feature set - see Managing pages.

Running detection first on a scanned set finds nothing, and it is easy to conclude the detection is broken rather than that there was no text to read.