Shravan Murlidaran

Vision Science · Machine Learning

Where people look, and why.

Shravan Murlidaran, Ph.D. Vision scientist bridging human and artificial intelligence

Open to research scientist and quantitative UX roles

Download CV (PDF)

  • DoctoratePh.D. Psychological & Brain Sciences, UC Santa Barbara, 2026
  • PublishedFirst author, Nature Communications 17, 940 (2026)
  • Under reviewFirst author, Nature Human Behaviour
  • PatentU.S. provisional application, filed May 2026 (pending)
  • Invited talkNVIDIA Research, 2026
Portrait of Shravan Murlidaran

Santa Barbara, California

I study where people look and why, then build models that have to reproduce it. My doctoral work showed that free viewing is active information seeking: people move their eyes toward regions that improve their comprehension of a scene. Those same characteristic fixation patterns then emerge on their own in a foveated vision-language model optimized for understanding scenes.

I come to this from engineering, with a master's in robotics, so I like to work at both ends: describing how human vision behaves, and building systems that reproduce it.

More about me

Highlights

What the work shows

Fixation heatmap · 50 participants Model fixations · 1 model

On the left, a fixation heatmap pooled from 50 people freely viewing the scene. On the right, a foveated vision-language model whose eye movements are optimized for its own scene comprehension. The signature fixation behaviours of human free viewing emerge on their own, with no human eye-tracking data used in training, as an optimal strategy for maximizing scene understanding under the biological constraints of foveated vision.

Read the research

Research

Three results that build on each other

The through line is a single question posed three ways: first measuring where humans look, then building a model that reproduces it from scratch, then asking how far today's vision-language models still are from human visual competence.

01 Measurement
Nature Communications · 2026

Eye movements during free viewing to maximize scene understanding

What humans look at when freely viewing a scene is not well understood. We measured eye movements under different instructions while observers viewed Winograd images: pairs differing by a small visual alteration that greatly changes how the scene is interpreted. Free-viewing fixations resemble those of observers describing scenes, and differ from those counting or searching for objects. They are directed toward the people and objects whose removal most alters scene interpretation, rather than toward the most salient or most meaningfully judged regions. Constraining observers to fixate objects irrelevant to understanding degrades their descriptions, showing that free-viewing eye movements are functionally important for scene comprehension.

02 Model
Under review · Nature Human Behaviour

Why we look where we look: emergent human-like fixations of a foveated visual language model

When humans view scenes without a specific task, they first direct their eye movements toward the scene center, then fixate on people, text, objects being looked at or grasped, and semantically meaningful regions. What these signature fixation patterns reflect, and whether they optimize an underlying perceptual task, has remained unknown. We show that a computational agent with simulated foveation, trained to optimize scene comprehension, exhibits emergent human fixation signature patterns. Versions of the agent trained instead to search or classify scenes, or equipped with peripheral vision better or worse than human vision, predicted human fixation patterns less accurately. Human free-viewing fixation patterns may therefore emerge as a functional byproduct of optimizing scene comprehension under the biological constraints of foveated vision.

Fixation heatmap, 50 participants Model fixations

Fixations to people in the scene
Fixations to objects being grasped
Fixations to regions critical to scene understanding
Fixations to text in the scene

Each clip is split. On the left, a fixation heatmap pooled from 50 participants, showing the signature patterns we observe in human free viewing: fixations drawn to people, to the objects they are handling, to text, and to whatever else the scene depends on for its meaning. On the right, the same properties emerging in the trained model, which arrived at them without ever seeing human eye-tracking data.

03 Benchmark
arXiv · 2026

Evolution of accuracy and visual-cognitive errors in a decade of vision-language AI models

Visual reasoning benchmarks mostly use simple scenes, few human descriptions, and rarely ask what models actually get wrong. We introduce the Complex Social Behavior dataset, 100 images depicting complex social interactions, and trace scene description accuracy across a decade of vision-language models against 20 human descriptions and a gold standard. Multimodal LLMs now reach accuracies similar to the top-ranked human descriptions, and have closed the gap between simple MS-COCO scenes and scenes depicting complex behaviors. Across five visual-cognitive error types, nearly all have been eliminated, except that models still occasionally rely on different image regions than humans do.

About

From robot vision to human vision, and back

Understanding vision has directed my path since my undergraduate degree. As a mechanical engineering student I built a computer vision pipeline that recognised characters from a live camera feed. What I was really chasing was robustness: I wanted it to hold up across lighting conditions and character colours the way a person reads a sign without noticing the effort. Closing that gap turned out to be far harder than it looked, and it is what sent me to WPI for a master's in robotics.

At WPI I worked on computer vision for the simulated Valkyrie R5 humanoid in NASA's Space Robotics Challenge, alongside a separate strand building HoloLens tools for visualising medical imaging data. The robotics work made the point sharply. Deciding where to look, what matters, and what to ignore is something people do continuously and without noticing, and getting a robot to do the same for a single task required an enormous amount of computation. I came away wanting to understand what the brain is doing to make it look so easy.

That question became my Ph.D. at UC Santa Barbara, where I studied how people visually reason and measured their eye movements at scale. In the end I arrived back at the question I started with, approached from the other side. As an undergraduate I wanted machines to comprehend the visual world as reliably as people do. What I ended up showing is that people are already near optimal at it, and the evidence is a model that moves its eyes to comprehend a scene: optimize it for understanding, and it looks where humans look.

Outside the lab I play sitar and guitar, sing with a South Asian a cappella group, and play tennis and squash.

Experience

  • 2019 – 2026 Graduate Student Researcher and Teaching Assistant UC Santa Barbara. Eye tracking and psychophysics experiments, reinforcement learning systems for vision-language models, gaze-contingent prosthetic vision simulation. Taught statistics and experimental design.
  • 2016 – 2019 Graduate Research Assistant Worcester Polytechnic Institute. Mixed reality analytics for neurological datasets with AbbVie, terrain navigation perception for NASA's R5 Valkyrie robot in the Space Robotics Challenge, and an HCI study on mixed reality for surgical navigation.
  • 2013 – 2016 Robotics Club (RMI) National Institute of Technology, Tiruchirappalli. Real-time chess piece detection for a robotic arm and a CNN-based word recognition system.

Education

  • 2026Ph.D., Psychological and Brain SciencesUniversity of California, Santa Barbara
  • 2019M.S., RoboticsWorcester Polytechnic Institute, Massachusetts
  • 2016B.Tech., Mechanical EngineeringNational Institute of Technology, Tiruchirappalli, India

Publications

Full record

Everything published, under review, or presented, including co-authored work.

Journal articles and preprints

  • 2026Murlidaran, S., & Eckstein, M. P. Eye movements during free viewing to maximize scene understanding. Nature Communications, 17, 940.
  • 2026Murlidaran, S., Wen, Z., Shehabi, S., & Eckstein, M. P. Why we look where we look: emergent human-like fixations of a foveated visual language model maximizing scene understanding. arXiv:2605.17823. Under review at Nature Human Behaviour.
  • 2026Murlidaran, S., & Eckstein, M. P. Evolution of accuracy and visual-cognitive errors in a decade of vision-language AI models. arXiv:2607.09654.
  • 2025Murlidaran, S., Wen, Z., Skaza, J., Wang, W., & Eckstein, M. P. Semantic saliency from multi-modal large language model scene understanding maps. arXiv preprint.
  • 2025Wen, Z., Skaza, J., Murlidaran, S., Wang, W. Y., & Eckstein, M. P. Predicting reaction time to comprehend scenes with foveated scene understanding maps. arXiv:2505.12660.
  • 2025Skaza, J., Murlidaran, S., Varshney, A., Wen, Z., Wang, W., Eckstein, M. P., & Beyeler, M. A deep learning framework for predicting functional visual performance in bionic eye users. bioRxiv, 2025-06.
  • 2021Murlidaran, S., Wang, W. Y., & Eckstein, M. P. Comparing visual reasoning in humans and AI. ICLR Brain2AI workshop.

Patent

  • 2026Automated eye movement response systems and methods. U.S. provisional patent application, filed May 2026. Patent pending.

Talks and conference presentations

  • 2026Invited talk, NVIDIA Research. Why we look where we look: emergent human-like fixations of a foveated visual language model maximizing scene understanding.
  • 2025Scene understanding maps: predicting most frequently fixated object during free viewing with multi-modal large language models. Journal of Vision, 25 (VSS 2025).
  • 2024Eye movements during free viewing to maximize scene understanding. Journal of Vision, 24 (VSS 2024).
  • 2023Eye movements during free viewing and scene description are similarly directed to objects critical to scene understanding. Gordon Research Conference on Eye Movements, Mount Holyoke College, July 2023.
  • 2023Eye movements during free viewing and scene description are similarly directed to objects critical to scene understanding. Journal of Vision, 23 (VSS 2023), abstract 5905.

Methods

How the work gets done

Most of my projects run the same loop: design a study that can actually answer the question, collect behaviour from real participants, then build a model that has to reproduce it. These are the tools that loop runs on.

Human experimentation
Psychophysics, gaze-contingent displays, EyeLink 1000 Plus eye tracking, saccade and fixation parsing, Qualtrics survey instruments, IRB protocols, user studies.
Analysis
GLMs, mixed-effects models, ANOVA, Bayesian methods, time-series analysis, cross-validation, exploratory data analysis, scanpath similarity metrics.
Modeling
Reinforcement learning with policy gradients and REINFORCE, vision-language models, CNNs, transformers, embedding extraction, large-scale inference with vLLM and HuggingFace.
Languages
Python, C, C++, C#, R, MATLAB, SQL
Frameworks and tools
PyTorch, OpenCV, NumPy, SciPy, PsychoPy, Pulse2Percept, Unity, Git, Linux, TensorBoard, Weights & Biases
Infrastructure
GPU computing (A6000), distributed training, parallelization, data pipelines and ETL
Also
Mixed reality development on HoloLens, robotics perception, real-time systems

Contact

Open to research scientist roles

I am looking for work at the intersection of human vision and machine perception. Email is the fastest way to reach me.