Shravan Murlidaran

Vision Science · Machine Learning

Where people look, and why.

Shravan Murlidaran, Ph.D. Vision scientist bridging human and artificial intelligence

Open to research scientist and quantitative UX roles

Download CV (PDF)

  • DoctoratePh.D. Psychological & Brain Sciences, UC Santa Barbara, 2026
  • PublishedFirst author, Nature Communications 17, 940 (2026)
  • Under reviewFirst author, Nature Human Behaviour
  • PatentU.S. provisional application, filed May 2026 (pending)
  • Invited talkNVIDIA Research, 2026
Portrait of Shravan Murlidaran

Santa Barbara, California

I study where people look and why, then build models that have to reproduce it. My doctoral work showed that free viewing is active information seeking: people move their eyes toward regions that improve their comprehension of a scene. Those same characteristic fixation patterns then emerge on their own in a foveated vision-language model optimized for understanding scenes.

I come to this from engineering, with a master's in robotics, so I like to work at both ends: describing how human vision behaves, and building systems that reproduce it.

More about me

Highlights

What the work shows

Fixation heatmap · 50 participants Model fixations · 1 model

On the left, a fixation heatmap from 50 participants viewing the scene. On the right, a single model. The model was trained to do one thing, understand the scene, and was never shown a human eye movement. The similarity is the result.

Read the research

Research

Three results that build on each other

The through line is a single question posed three ways: first measuring where humans look, then building a model that reproduces it from scratch, then asking how far today's vision-language models still are from human visual competence.

01 Measurement
Nature Communications · 2026

Eye movements during free viewing to maximize scene understanding

When people look at a scene with no task at all, their eyes still go somewhere specific. This paper introduces the Winograd Images dataset and Scene Understanding Maps, a way of scoring how much each region of an image contributes to grasping what the scene is about. Human fixations track those scene-understanding regions more closely than they track classical visual saliency, which suggests free viewing is less free than it looks: the eyes are already working toward comprehension.

02 Model
Under review · Nature Human Behaviour

Why we look where we look: emergent human-like fixations of a foveated visual language model

A vision-language model given a human-like fovea and trained with reinforcement learning to do one thing, understand the scene in front of it, starts moving its gaze the way people do. It was never shown human eye-tracking data. The fixation patterns emerge from the objective alone, which makes scene understanding a candidate explanation for why human gaze looks the way it does rather than just a correlate of it.

Fixation heatmap, 50 participants Model fixations

S1 People
S2 Grasped objects
S3 Regions critical to scene understanding
S4 Text

Each clip is split: a fixation heatmap from 50 participants on the left, the model on the right.

03 Benchmark
arXiv · 2026

Evolution of accuracy and visual-cognitive errors in a decade of vision-language AI models

Accuracy scores hide what a model actually gets wrong. Tracing vision-language models across roughly a decade of development, this paper measures not only how often they answer correctly about complex social scenes but which categories of visual-cognitive error they make, and which of those errors have shrunk over time versus which have stayed put while benchmark numbers climbed.

About

From robot vision to human vision, and back

Understanding vision has directed my path since my undergraduate degree. As a mechanical engineering student I built a computer vision pipeline that recognised characters from a live camera feed. What I was really chasing was robustness: I wanted it to hold up across lighting conditions and character colours the way a person reads a sign without noticing the effort. Closing that gap turned out to be far harder than it looked, and it is what sent me to WPI for a master's in robotics.

At WPI I worked on computer vision for the simulated Valkyrie R5 humanoid in NASA's Space Robotics Challenge, alongside a separate strand building HoloLens tools for visualising medical imaging data. The robotics work made the point sharply. Deciding where to look, what matters, and what to ignore is something people do continuously and without noticing, and getting a robot to do the same for a single task required an enormous amount of computation. I came away wanting to understand what the brain is doing to make it look so easy.

That question became my Ph.D. at UC Santa Barbara, where I studied how people visually reason and measured their eye movements at scale. In the end I arrived back at the question I started with, approached from the other side. As an undergraduate I wanted machines to comprehend the visual world as reliably as people do. What I ended up showing is that people are already near optimal at it, and the evidence is a model that moves its eyes to comprehend a scene: optimize it for understanding, and it looks where humans look.

Outside the lab I play sitar and guitar, sing with a South Asian a cappella group, and play tennis and squash.

Experience

  • 2019 – 2026 Graduate Student Researcher and Teaching Assistant UC Santa Barbara. Eye tracking and psychophysics experiments, reinforcement learning systems for vision-language models, gaze-contingent prosthetic vision simulation. Taught statistics and experimental design.
  • 2016 – 2019 Graduate Research Assistant Worcester Polytechnic Institute. Mixed reality analytics for neurological datasets with AbbVie, terrain navigation perception for NASA's R5 Valkyrie robot in the Space Robotics Challenge, and an HCI study on mixed reality for surgical navigation.
  • 2013 – 2016 Robotics Club (RMI) National Institute of Technology, Tiruchirappalli. Real-time chess piece detection for a robotic arm and a CNN-based word recognition system.

Education

  • 2026Ph.D., Psychological and Brain SciencesUniversity of California, Santa Barbara
  • 2019M.S., RoboticsWorcester Polytechnic Institute, Massachusetts
  • 2016B.Tech., Mechanical EngineeringNational Institute of Technology, Tiruchirappalli, India

Publications

Full record

Everything published, under review, or presented, including co-authored work.

Journal articles and preprints

  • 2026Murlidaran, S., & Eckstein, M. P. Eye movements during free viewing to maximize scene understanding. Nature Communications, 17, 940.
  • 2026Murlidaran, S., Wen, Z., Shehabi, S., & Eckstein, M. P. Why we look where we look: emergent human-like fixations of a foveated visual language model maximizing scene understanding. arXiv:2605.17823. Under review at Nature Human Behaviour.
  • 2026Murlidaran, S., & Eckstein, M. P. Evolution of accuracy and visual-cognitive errors in a decade of vision-language AI models. arXiv:2607.09654.
  • 2025Murlidaran, S., Wen, Z., Skaza, J., Wang, W., & Eckstein, M. P. Semantic saliency from multi-modal large language model scene understanding maps. arXiv preprint.
  • 2025Wen, Z., Skaza, J., Murlidaran, S., Wang, W. Y., & Eckstein, M. P. Predicting reaction time to comprehend scenes with foveated scene understanding maps. arXiv:2505.12660.
  • 2025Skaza, J., Murlidaran, S., Varshney, A., Wen, Z., Wang, W., Eckstein, M. P., & Beyeler, M. A deep learning framework for predicting functional visual performance in bionic eye users. bioRxiv, 2025-06.
  • 2021Murlidaran, S., Wang, W. Y., & Eckstein, M. P. Comparing visual reasoning in humans and AI. ICLR Brain2AI workshop.

Patent

  • 2026Automated eye movement response systems and methods. U.S. provisional patent application, filed May 2026. Patent pending.

Talks and conference presentations

  • 2026Invited talk, NVIDIA Research. Why we look where we look: emergent human-like fixations of a foveated visual language model maximizing scene understanding.
  • 2025Scene understanding maps: predicting most frequently fixated object during free viewing with multi-modal large language models. Journal of Vision, 25 (VSS 2025).
  • 2024Eye movements during free viewing to maximize scene understanding. Journal of Vision, 24 (VSS 2024).
  • 2023Eye movements during free viewing and scene description are similarly directed to objects critical to scene understanding. Gordon Research Conference on Eye Movements, Mount Holyoke College, July 2023.
  • 2023Eye movements during free viewing and scene description are similarly directed to objects critical to scene understanding. Journal of Vision, 23 (VSS 2023), abstract 5905.

Methods

How the work gets done

Most of my projects run the same loop: design a study that can actually answer the question, collect behaviour from real participants, then build a model that has to reproduce it. These are the tools that loop runs on.

Human experimentation
Psychophysics, gaze-contingent displays, EyeLink 1000 Plus eye tracking, saccade and fixation parsing, Qualtrics survey instruments, IRB protocols, user studies.
Analysis
GLMs, mixed-effects models, ANOVA, Bayesian methods, time-series analysis, cross-validation, exploratory data analysis, scanpath similarity metrics.
Modeling
Reinforcement learning with policy gradients and REINFORCE, vision-language models, CNNs, transformers, embedding extraction, large-scale inference with vLLM and HuggingFace.
Languages
Python, C, C++, C#, R, MATLAB, SQL
Frameworks and tools
PyTorch, OpenCV, NumPy, SciPy, PsychoPy, Pulse2Percept, Unity, Git, Linux, TensorBoard, Weights & Biases
Infrastructure
GPU computing (A6000), distributed training, parallelization, data pipelines and ETL
Also
Mixed reality development on HoloLens, robotics perception, real-time systems

Contact

Open to research scientist roles

I am looking for work at the intersection of human vision and machine perception. Email is the fastest way to reach me.