DeepMind publishes a white paper defining "Visual General Intelligence"
The August 26 paper stakes out what it would actually mean for a model to understand images the way it understands language.
On August 26, Google DeepMind published 'Visual General Intelligence: A White Paper,' an attempt to define, in concrete and testable terms, what it would mean for a model to reason about visual information with the same depth and flexibility that current frontier models bring to text.
The paper arrives alongside a wave of DeepMind work this quarter targeting two stubborn multimodal weak points: hallucination in vision-language models and coherence across long video generations. Reducing factual errors in visual outputs has been one of the harder unsolved problems in multimodal AI, and this white paper reads as DeepMind laying out its long-term roadmap rather than announcing a single model.
For a field that has spent three years mostly scaling text reasoning, a serious definitional framework for visual intelligence is a signal about where the next big capability jump is expected to come from.