Ruohan.jpg

Ruohan Gao

Assistant Professor
Department of Computer Science, University of Maryland, College Park

Office: IRB-4248
Email: rhgao[AT]umd.edu

UMD Profile / Google Scholar / Short Bio

I am an assistant professor in the Department of Computer Science at University of Maryland, College Park, where I lead the UMD Multisensory Machine Intelligence Lab. I am also affiliated with the University of Maryland Institute for Advanced Computer Studies (UMIACS), Maryland Robotics Center (MRC), and Artificial Intelligence Interdisciplinary Institute at Maryland (AIM).

I received my Ph.D. in Computer Science from The University of Texas at Austin advised by Kristen Grauman, and then spent two years as a PostDoc at Stanford Vision and Learning Lab working with Fei-Fei Li, Jiajun Wu, and Silvio Savarese.

My research primarily focuses on computer vision and machine learning with a particular emphasis on multisensory machine intelligence involving sight, sound, and touch. The overarching goal of my research is to empower machines to emulate and enhance human capabilities in seeing, hearing, and feeling, ultimately enabling them to comprehensively perceive, understand, and interact with the multisensory world.

Prospective Students: I am always seeking self-motivated students to join my group. If you are interested, here is some more information.

News

[Oct 2026] Invited Talk at Speech and Audio in the Northeast (SANE 2026) Workshop at MIT.
[Sep 2026] Invited Talk at ECCV 2026 Workshop on Audio-Visual Generation & Learning.
[Sep 2026] Invited Talk at ECCV 2026 Workshop on 3D Modeling, Reconstruction, and Generation in the Wild.
[Sep 2026] Invited Talk at ECCV 2026 Workshop on Generative AI for Audio-Visual Content Creation.
[Jun 2026] Invited Talk at CVPR 2026 Workshop on Sight and Sound.
[Apr 2026] Invited Talk at Symposium on Machine Learning across Modalities at Yale.
[Mar 2026] Received the 2025 Sony Faculty Innovation Award. Thanks, Sony!
[May 2025] Invited Talk at ICRA 2025 Workshop on Acoustic Sensing and Representations for Robotics.
[Dec 2024] Selected for AAAI New Faculty Highlights 2025.
[Jun 2024] We are organizing the Sight and Sound Workshop at CVPR 2024.
[Oct 2023] We are organizing the AV4D Workshop at ICCV 2023.
[Feb 2023] We are organizing the Creative AI Across Modalities Workshop at AAAI 2023.
[May 2021] We are organizing the Embodied Multimodal Learning Workshop at ICLR 2021.
[May 2021] I am very honored to have received the Michael H. Granof Award that recognizes UT Austin’s Top 1 Doctoral Dissertation of 2021.

Multisensory Machine Intelligence Lab

Selected Publications [full list]

2026

  1. omnitactune_corl2026.jpg
    OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
    Conference on Robot Learning (CoRL), 2026
  2. avmsf_eccv2026.png
    Objects as Audio-Visual Modal Sound Fields
    Zisen Shao*Zihao Wei*Derong Jin, and Ruohan Gao
    European Conference on Computer Vision (ECCV), 2026
  3. sonoworld_cvpr206.png
    SonoWorld: From One Image to a 3D Audio-Visual Scene
    Derong Jin*Xiyi Chen*Ming C. Lin, and Ruohan Gao
    Conference on Computer Vision and Pattern Recognition (CVPR), 2026

2025

  1. multisensory_objects_space.png
    Multisensory Machine Intelligence
    Ruohan Gao
    AI Magazine, 2025
  2. avdar_iccv2025.png
    Differentiable Room Acoustic Rendering with Multi-View Vision Priors
    Derong Jin, and Ruohan Gao
    International Conference on Computer Vision (ICCV), 2025
  3. waspaa2025_hrtf.png
    Towards Perception-Informed Latent HRTF Representations
    You Zhang, Andrew Francl, Ruohan GaoPaul CalamiaZhiyao Duan, and Ishwarya Ananthabhotla
    IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025
  4. HAAE_cvpr2025.jpeg
    Hearing Anywhere in Any Environment
    Xiulong LiuAnurag KumarPaul Calamia, Sebastià V. Amengual Garí, Calvin Murdock, Ishwarya Ananthabhotla, Philip Robinson, Eli ShlizermanVamsi Krishna Ithapu, and Ruohan Gao
    Conference on Computer Vision and Pattern Recognition (CVPR), 2025

2024

  1. hearing_anything_anywhere_cvpr2024.png
    Hearing Anything Anywhere
    Mason L. Wang*, Ryosuke Sawata*, Samuel ClarkeRuohan GaoShangzhe Wu, and Jiajun Wu
    Conference on Computer Vision and Pattern Recognition (CVPR), 2024

2023

  1. of_benchmark_cvpr2023.jpg
    The ObjectFolder Benchmark: Multisensory Object-Centric Learning with Neural and Real Objects
    Conference on Computer Vision and Pattern Recognition (CVPR), 2023
  2. realimpact_cvpr2023.jpg
    RealImpact: A Dataset of Impact Sound Fields for Real Objects
    Samuel ClarkeRuohan GaoMason WangMark Rau, Julia Xu, Mark RauJui-Hsien WangDoug James, and Jiajun Wu
    Conference on Computer Vision and Pattern Recognition (CVPR), 2023

2022

  1. see_hear_feel_corl2022.png
    See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
    Hao Li*, Yizhi Zhang*, Junzhe Zhu, Shaoxiong WangMichelle A. LeeHuazhe XuEdward AdelsonLi Fei-FeiRuohan Gao†, and Jiajun Wu†
    Conference on Robot Learning (CoRL), 2022
  2. objectfolderV2.png
    ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer
    Conference on Computer Vision and Pattern Recognition (CVPR), 2022
  3. visual_acoustic_matching_cvpr2022.png
    Visual Acoustic Matching
    Changan ChenRuohan GaoPaul Calamia, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2022

2021

  1. bmvc2021.png
    Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video
    Rishabh GargRuohan Gao, and Kristen Grauman
    British Machine Vision Conference (BMVC), 2021
  2. thesis_teaser.png
    Look and Listen: From Semantic to Spatial Audio-Visual Perception
    Ruohan Gao
    Ph.D. Dissertation, 2021

2019

  1. 2.5D_visual_sound_cvpr2019.png
    2.5D Visual Sound
    Ruohan Gao, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2019

2018

  1. audioobjects_eccv2018.png
    Learning to Separate Object Sounds by Watching Unlabeled Video
    Ruohan GaoRogerio Feris, and Kristen Grauman
    European Conference on Computer Vision (ECCV), 2018
  2. im2flow_cvpr2018.jpg
    Im2Flow: Motion Hallucination from Static Images for Action Recognition
    Ruohan GaoBo Xiong, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2018