How OCR works and why every document needs it

Why each imported document goes through OCR, PDF text versus visual OCR, what the OCR Status column means, and how the OCR quota works.

Every document you import into Blast Audit goes through OCR (optical character recognition). OCR finds each word on the page and records where it sits. Snips, search in the viewer, document matching and extraction all work from that text, so a document has to finish OCR before you can use it. There is no setting to skip OCR.

How a document is read

Blast Audit picks the method for each document when you import it.

DocumentMethodShown as
PDF that already contains text (exported from accounting software, a bank portal, Word)The text is read directly from the PDF. This takes seconds.PDF text
Scan, photo or PDF without textThe page images are read by visual OCR. This takes longer.Visual OCR

For a PDF read from its own text, the Documents page shows: "Document ready: text was read directly from the PDF. Use visual OCR only if some snips look wrong or are missing." Some PDFs carry hidden text that does not match the page. In that case, click the ⋮ button at the end of the document's row, then Run visual OCR.

Follow the progress

When you import documents, the pane goes through Uploading files, Getting ready and Reading documents. Keep Excel open and stay on the Documents page until it finishes. If you leave during the import, the OCR result may be lost and you will need to run OCR again.

The OCR Status column shows where each document is:

StatusMeaning
Not Started / OCR requiredOCR has not run. Click Start OCR.
QueuedWaiting its turn.
ProcessingBeing read.
CompleteReady for snips, matching and extraction.
OCR failedClick Re-run OCR in the document's ⋮ menu.

Re-run OCR reads the document again with visual OCR. If a re-run fails, Blast Audit keeps the previous result.

OCR quota

Visual OCR counts pages against a monthly quota shared by everyone in your organization: about 10,000 pages per user per month, pooled. A PDF whose text is fully readable at import skips OCR and does not count. When a document needs OCR, all of its pages count. When the quota is reached, the pane tells you, and the documents not processed stay on OCR required. Ask your administrator about your organization's usage. The quota comes with your licences: see Choose a plan, buy licences or get a quote.

Where OCR runs

Visual OCR runs on Microsoft Azure. Where your data is stored explains where documents are processed and how long they are kept.

Related problems

If OCR stays on Processing, fails, or finds no text, see Fix OCR that is stuck, slow or failed.

Still stuck?

Email the team with what you tried and, if you can, a screenshot of what you see.

Email support