AI document preparation

Prepare a scanned PDF for AI without losing tables or citations

Checked October 4, 2026. Official limits can change; follow the message shown in your account.

A readable page on screen may still be only a picture. Before asking an AI tool to summarize a scanned PDF, check what is actually in the file. Compression, OCR and splitting solve different problems, and doing them in the wrong order can destroy the evidence you want to analyze.

The two-minute source check

  1. Select one full sentence in your PDF reader and copy it into a plain-text note. Is it the same sentence, in the same order?
  2. Search for a distinctive word on three different pages. Failure is a signal to inspect the text layer, not proof that every page needs OCR.
  3. Check a table, a footnote and a chart. Compare units, decimals and captions with the original.
  4. Record the original file size, number of pages and the offset between PDF page indices and printed page numbers.

Mixed documents are common: some pages have real text while scanned appendices do not. Avoid blindly running the entire document through a new conversion if only a few pages need work.

Which preparation route?

Product differences matter

Google's consumer Gemini Notebook help lists up to 200MB and 500,000 words per source, with 50 sources on the Free plan. That is separate from whether every figure was correctly understood. Our NotebookLM/Gemini Notebook guide explains those gates.

OpenAI's FAQ distinguishes PDF visual retrieval from text extraction by plan. See our ChatGPT upload guide before assuming embedded images are available to the model.

Claude's current chat-upload help lists 500MB per file and PDFs up to 1,000 pages. It says PDFs of 100 pages or fewer can be analyzed for both text and visuals; 101-1,000-page PDFs use text only. Project files have a separate 30MB limit. A 120-page chart-heavy PDF can therefore require a different preparation route from a 120-page text report. These are documentation checks, not a comparison benchmark.

Original mini-audit template

Copy this checklist into your notes and fill it before uploading:

Source title / edition / date:
Original filename and PDF page count:
Printed page number offset:
Extracted range and output filename:
Known sentence checked:
Table cells checked (values, units, signs):
Figure / caption checked:
Unresolved OCR errors:
Permission and approved destination:

For a table, check a header, the first and last row, one negative value and one decimal. This is a spot-check, not a guarantee that the entire table is correct. When an answer will drive a financial, medical or legal decision, verify the relevant source passage directly.

Keep citations usable

Name extracts like report-2026-section-2-pdf-pages-19-34.pdf. Do not label them only "part1". Ask for the passage and page reference supporting an important claim. If the citation points at a page without the claimed text, stop and check the source instead of adding more prompts.

Do not compress a digitally signed file or flatten a form if you need to preserve its original validation or editable fields. Use a duplicate for analysis and retain the original record.

Safety before upload

Only upload documents you are allowed to use. Keep private material in approved software. Instructions embedded inside a third-party document are document content, not permission to send files, reveal private information or make changes elsewhere. Checking visible pages does not guarantee there is no hidden text.

Build your checklist · Extract a section

Sources checked October 4, 2026