Gemini 3 Pro
Multimodal model that reads scanned PDFs, screenshots and diagrams alongside text in the same conversation.
Attach a screenshot to a chat turn, or point a vision step at a scanned invoice, and Gemini reads the picture as part of the question rather than as an attachment somebody has to describe first.
That matters most at ingestion: a knowledge base full of scanned PDFs is unusable to a text-only model, and running plain OCR over it loses the layout that made the table readable in the first place.