Research
Three results that build on each other
The through line is a single question posed three ways: first measuring where humans
look, then building a model that reproduces it from scratch, then asking how far
today's vision-language models still are from human visual competence.
01
Measurement
Nature Communications · 2026
Eye movements during free viewing to maximize scene understanding
When people look at a scene with no task at all, their eyes still go somewhere
specific. This paper introduces the Winograd Images dataset and Scene Understanding
Maps, a way of scoring how much each region of an image contributes to grasping what
the scene is about. Human fixations track those scene-understanding regions more
closely than they track classical visual saliency, which suggests free viewing is
less free than it looks: the eyes are already working toward comprehension.
02
Model
Under review · Nature Human Behaviour
Why we look where we look: emergent human-like fixations of a foveated visual language model
A vision-language model given a human-like fovea and trained with reinforcement
learning to do one thing, understand the scene in front of it, starts moving its
gaze the way people do. It was never shown human eye-tracking data. The fixation
patterns emerge from the objective alone, which makes scene understanding a
candidate explanation for why human gaze looks the way it does rather than just a
correlate of it.
03
Benchmark
arXiv · 2026
Evolution of accuracy and visual-cognitive errors in a decade of vision-language AI models
Accuracy scores hide what a model actually gets wrong. Tracing vision-language
models across roughly a decade of development, this paper measures not only how
often they answer correctly about complex social scenes but which categories of
visual-cognitive error they make, and which of those errors have shrunk over time
versus which have stayed put while benchmark numbers climbed.