Ruohan Gao
Assistant Professor
Department of Computer Science, University of Maryland, College Park
Office: IRB-4248
Email: rhgao[AT]umd.edu
UMD Profile / Google Scholar / Short Bio
I am an assistant professor in the Department of Computer Science at University of Maryland, College Park, where I lead the UMD Multisensory Machine Intelligence Lab. I am also affiliated with the University of Maryland Institute for Advanced Computer Studies (UMIACS), Maryland Robotics Center (MRC), and Artificial Intelligence Interdisciplinary Institute at Maryland (AIM).
I received my Ph.D. in Computer Science from The University of Texas at Austin advised by Kristen Grauman, and then spent two years as a PostDoc at Stanford Vision and Learning Lab working with Fei-Fei Li, Jiajun Wu, and Silvio Savarese.
My research primarily focuses on computer vision and machine learning with a particular emphasis on multisensory machine intelligence involving sight, sound, and touch. The overarching goal of my research is to empower machines to emulate and enhance human capabilities in seeing, hearing, and feeling, ultimately enabling them to comprehensively perceive, understand, and interact with the multisensory world.
Prospective Students: I am always seeking self-motivated students to join my group. If you are interested, here is some more information.
News
| [Oct 2026] | Invited Talk at Speech and Audio in the Northeast (SANE 2026) Workshop at MIT. |
|---|---|
| [Sep 2026] | Invited Talk at ECCV 2026 Workshop on Audio-Visual Generation & Learning. |
| [Sep 2026] | Invited Talk at ECCV 2026 Workshop on 3D Modeling, Reconstruction, and Generation in the Wild. |
| [Sep 2026] | Invited Talk at ECCV 2026 Workshop on Generative AI for Audio-Visual Content Creation. |
| [Jun 2026] | Invited Talk at CVPR 2026 Workshop on Sight and Sound. |
| [Apr 2026] | Invited Talk at Symposium on Machine Learning across Modalities at Yale. |
| [Mar 2026] | Received the 2025 Sony Faculty Innovation Award. Thanks, Sony! |
| [May 2025] | Invited Talk at ICRA 2025 Workshop on Acoustic Sensing and Representations for Robotics. |
| [Dec 2024] | Selected for AAAI New Faculty Highlights 2025. |
| [Jun 2024] | We are organizing the Sight and Sound Workshop at CVPR 2024. |
| [Oct 2023] | We are organizing the AV4D Workshop at ICCV 2023. |
| [Feb 2023] | We are organizing the Creative AI Across Modalities Workshop at AAAI 2023. |
| [May 2021] | We are organizing the Embodied Multimodal Learning Workshop at ICLR 2021. |
| [May 2021] | I am very honored to have received the Michael H. Granof Award that recognizes UT Austin’s Top 1 Doctoral Dissertation of 2021. |
Multisensory Machine Intelligence Lab
PhD students:
Selected Publications [full list]
2026
-
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual PoliciesConference on Robot Learning (CoRL), 2026 -
Objects as Audio-Visual Modal Sound FieldsEuropean Conference on Computer Vision (ECCV), 2026 -
SonoWorld: From One Image to a 3D Audio-Visual SceneConference on Computer Vision and Pattern Recognition (CVPR), 2026
2025
-
Differentiable Room Acoustic Rendering with Multi-View Vision PriorsInternational Conference on Computer Vision (ICCV), 2025 -
Hearing Anywhere in Any EnvironmentConference on Computer Vision and Pattern Recognition (CVPR), 2025
2024
2023
-
The ObjectFolder Benchmark: Multisensory Object-Centric Learning with Neural and Real ObjectsConference on Computer Vision and Pattern Recognition (CVPR), 2023
2022
-
See, Hear, and Feel: Smart Sensory Fusion for Robotic ManipulationConference on Robot Learning (CoRL), 2022 -
ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real TransferConference on Computer Vision and Pattern Recognition (CVPR), 2022 -
Visual Acoustic MatchingConference on Computer Vision and Pattern Recognition (CVPR), 2022
2021
-
Geometry-Aware Multi-Task Learning for Binaural Audio Generation from VideoBritish Machine Vision Conference (BMVC), 2021 -
Look and Listen: From Semantic to Spatial Audio-Visual PerceptionPh.D. Dissertation, 2021
2019
-
2.5D Visual SoundConference on Computer Vision and Pattern Recognition (CVPR), 2019