Publications

*equal contribution, †equal advising

2026

  1. omnitactune_corl2026.jpg
    OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
    Conference on Robot Learning (CoRL), 2026
  2. controltac_corl2026.jpg
    ControlTac: Scaling Tactile Data with Physically Controlled Tactile Image Generation
    Conference on Robot Learning (CoRL), 2026
  3. odorspaces_corl2026.png
    OdorSpaces: Visual-Olfactory Embodied Navigation in 3D Environments
    Changyeon Lee, Kelin Yu, Zheng Wei, Shuna Ni, and Ruohan Gao
    Conference on Robot Learning (CoRL), 2026
  4. avmsf_eccv2026.png
    Objects as Audio-Visual Modal Sound Fields
    Zisen Shao*, Zihao Wei*, Derong Jin, and Ruohan Gao
    European Conference on Computer Vision (ECCV), 2026
  5. sonoworld_cvpr206.png
    SonoWorld: From One Image to a 3D Audio-Visual Scene
    Derong Jin*, Xiyi Chen*, Ming C. Lin, and Ruohan Gao
    Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  6. avllm_cvpr2026.png
    Do Audio-Visual Large Language Models Really See and Hear?
    Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi, Sreyan Ghosh, Ruohan Gao, and Dinesh Manocha
    Conference on Computer Vision and Pattern Recognition (CVPR Findings), 2026

2025

  1. multisensory_objects_space.png
    Multisensory Machine Intelligence
    Ruohan Gao
    AI Magazine, 2025
  2. avdar_iccv2025.png
    Differentiable Room Acoustic Rendering with Multi-View Vision Priors
    Derong Jin, and Ruohan Gao
    International Conference on Computer Vision (ICCV), 2025
  3. egoadapt_iccv2025.png
    EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
    Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan, Calvin Murdock, Ishwarya Ananthabhotla, Yijun Qian, Vamsi Krishna Ithapu, Dinesh Manocha, and Ruohan Gao
    International Conference on Computer Vision (ICCV), 2025
  4. genflowrl_iccv2025.png
    GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
    Kelin Yu*, Sheng Zhang*, Harshit Soora, Furong Huang, Heng Huang, Pratap Tokekar, and Ruohan Gao
    International Conference on Computer Vision (ICCV), 2025
  5. avtrustbench_iccv2025.jpg
    AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
    International Conference on Computer Vision (ICCV), 2025
  6. aurelia_iccv2025.png
    Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
    International Conference on Computer Vision (ICCV), 2025
  7. waspaa2025_hrtf.png
    Towards Perception-Informed Latent HRTF Representations
    You Zhang, Andrew Francl, Ruohan Gao, Paul Calamia, Zhiyao Duan, and Ishwarya Ananthabhotla
    IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025
  8. HAAE_cvpr2025.jpeg
    Hearing Anywhere in Any Environment
    Xiulong Liu, Anurag Kumar, Paul Calamia, Sebastià V. Amengual Garí, Calvin Murdock, Ishwarya Ananthabhotla, Philip Robinson, Eli Shlizerman, Vamsi Krishna Ithapu, and Ruohan Gao
    Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  9. visal_cvpr_2025.jpg
    Learning to Highlight Audio by Watching Movies
    Chao Huang, Ruohan Gao, J. M. F. Tsang, Jan Kurcius, Cagdas Bilen, Chenliang Xu, Anurag Kumar, and Sanjeel Parekh
    Conference on Computer Vision and Pattern Recognition (CVPR), 2025

2024

  1. swl_eccv2024.png
    Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
    Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar, Jacob Donley, Chao Li, Gunhee Kim, Vamsi Krishna Ithapu, and Calvin Murdock
    European Conference on Computer Vision (ECCV), 2024
  2. meercat_eccv2024.png
    Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
    European Conference on Computer Vision (ECCV), 2024
  3. diffsound_siggraph2024.png
    DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
    Xutong Jin*, Chenxi Xu*, Ruohan Gao, Jiajun Wu, Guoping Wang, and Sheng Li
    ACM Special Interest Group on Computer Graphics and Interactive Techniques Conference (SIGGRAPH), 2024
  4. hearing_anything_anywhere_cvpr2024.png
    Hearing Anything Anywhere
    Mason L. Wang*, Ryosuke Sawata*, Samuel Clarke, Ruohan Gao, Shangzhe Wu, and Jiajun Wu
    Conference on Computer Vision and Pattern Recognition (CVPR), 2024
  5. av-conv-cvpr2024.png
    The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
    Conference on Computer Vision and Pattern Recognition (CVPR), 2024

2023

  1. soundcam_neurips2023.png
    SoundCam: A Dataset for Tasks in Tracking and Identifying Humans from Real Room Acoustics
    Conference on Neural Information Processing Systems Datasets and Benchmarks Track (NeurIPS), 2023
  2. noir_corl2023.png
    NOIR: Neural Signal Operated Intelligent Robot for Everyday Activities
    Ruohan Zhang*, Sharon Lee*, Minjune Hwang*, Ayano Hiranaka*, Chen Wang, Wensi Ai, Jin Jie Ryan Tan, Shreya Gupta, Yilun Hao, Gabrael Levine, and 4 more authors
    Conference on Robot Learning (CoRL), 2023
  3. bmvc2021.png
    Visually-Guided Audio Spatialization in Video with Geometry-Aware Multi-task Learning
    Rishabh Garg, Ruohan Gao, and Kristen Grauman
    International Journal of Computer Vision (IJCV), 2023
  4. of_benchmark_cvpr2023.jpg
    The ObjectFolder Benchmark: Multisensory Object-Centric Learning with Neural and Real Objects
    Conference on Computer Vision and Pattern Recognition (CVPR), 2023
  5. realimpact_cvpr2023.jpg
    RealImpact: A Dataset of Impact Sound Fields for Real Objects
    Samuel Clarke, Ruohan Gao, Mason Wang, Mark Rau, Julia Xu, Mark Rau, Jui-Hsien Wang, Doug James, and Jiajun Wu
    Conference on Computer Vision and Pattern Recognition (CVPR), 2023
  6. osf_tmlr2023.jpg
    Learning Object-Centric Neural Scattering Functions for Free-Viewpoint Relighting and Scene Composition
    Transactions on Machine Learning Research (TMLR), 2023
  7. dano_ral2023.jpg
    Differentiable Physics Simulation of Dynamics-Augmented Neural Objects
    Simon Le Cleac’h, Hong-Xing Yu, Michelle Guo, Taylor A. Howell, Ruohan Gao, Jiajun Wu, Zachary Manchester, and Mac Schwager
    Robotics and Automation Letters (RA-L), 2023
  8. sonicverse_icra2023.jpg
    Sonicverse: A Multisensory Simulation Platform for Training Household Agents that See and Hear
    Ruohan Gao*, Hao Li*, Gokul Dharan, Zhuzhu Wang, Chengshu Li, Fei Xia, Silvio Savarese, Li Fei-Fei, and Jiajun Wu
    International Conference on Robotics and Automation (ICRA),, 2023
  9. emma_dataset_iclr2023.jpg
    An Extensible Multi-modal Multi-task Object Dataset with Materials
    Trevor Scott Standley, Ruohan Gao, Dawn Chen, Jiajun Wu, and Silvio Savarese
    International Conference on Learning Representations (ICLR), 2023

2022

  1. see_hear_feel_corl2022.png
    See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
    Hao Li*, Yizhi Zhang*, Junzhe Zhu, Shaoxiong Wang, Michelle A. Lee, Huazhe Xu, Edward Adelson, Li Fei-Fei, Ruohan Gao†, and Jiajun Wu†
    Conference on Robot Learning (CoRL), 2022
  2. objectfolderV2.png
    ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer
    Conference on Computer Vision and Pattern Recognition (CVPR), 2022
  3. visual_acoustic_matching_cvpr2022.png
    Visual Acoustic Matching
    Changan Chen, Ruohan Gao, Paul Calamia, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2022

2021

  1. objectfolder_corl2021.png
    ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations
    Ruohan Gao, Yen-Yu Chang, Shivani Mall, Li Fei-Fei, and Jiajun Wu
    Conference on Robot Learning (CoRL), 2021
  2. diffImpact_corl2021.png
    DiffImpact: Differentiable Rendering and Identification of Impact Sounds
    Conference on Robot Learning (CoRL), 2021
  3. bmvc2021.png
    Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video
    Rishabh Garg, Ruohan Gao, and Kristen Grauman
    British Machine Vision Conference (BMVC), 2021
  4. thesis_teaser.png
    Look and Listen: From Semantic to Spatial Audio-Visual Perception
    Ruohan Gao
    Ph.D. Dissertation, 2021
  5. VisualVoice_cvpr2021.jpg
    Visualvoice: Audio-visual speech separation with cross-modal consistency
    Ruohan Gao, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2021
  6. av_wan_iclr2021.jpg
    Learning to Set Waypoints for Audio-Visual Navigation
    International Conference on Learning Representations (ICLR), 2021

2020

  1. visualEchoes_eccv2020.png
    VisualEchoes: Spatial Visual Representation Learning through Echolocation
    European Conference on Computer Vision (ECCV), 2020
  2. listen_to_look_cvpr2020.png
    Listen to Look: Action Recognition by Previewing Audio
    Conference on Computer Vision and Pattern Recognition (CVPR), 2020

2019

  1. co-separation-iccv2019.png
    Co-Separating Sounds of Visual Objects
    Ruohan Gao, and Kristen Grauman
    International Conference on Computer Vision (ICCV), 2019
  2. 2.5D_visual_sound_cvpr2019.png
    2.5D Visual Sound
    Ruohan Gao, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2019

2018

  1. audioobjects_eccv2018.png
    Learning to Separate Object Sounds by Watching Unlabeled Video
    Ruohan Gao, Rogerio Feris, and Kristen Grauman
    European Conference on Computer Vision (ECCV), 2018
  2. shapecodes_eccv2018.jpg
    ShapeCodes: Self-Supervised Feature Learning by Lifting Views to Viewgrids
    Dinesh Jayaraman, Ruohan Gao, and Kristen Grauman
    European Conference on Computer Vision (ECCV), 2018
  3. im2flow_cvpr2018.jpg
    Im2Flow: Motion Hallucination from Static Images for Action Recognition
    Ruohan Gao, Bo Xiong, and Kristen Grauman
    Conference on Computer Vision and Pattern Recognition (CVPR), 2018

2017

  1. ondemand_iccv2017.jpg
    On-Demand Learning for Deep Image Restoration
    Ruohan Gao, and Kristen Grauman
    International Conference on Computer Vision (ICCV), 2017

2016

  1. objectcentric_accv2016.jpg
    Object-Centric Representation Learning from Unlabeled Videos
    Ruohan Gao, Dinesh Jayaraman, and Kristen Grauman
    Asian Conference on Computer Vision (ACCV), 2016
  2. IEEE ICC
    Accelerating Graph Mining Algorithms via Uniform Random Edge Sampling
    Ruohan Gao, Huanle Xu, Pili Hu, and Wing Cheong Lau
    IEEE International Conference on Communications (ICC), 2016

2015

  1. IEEE GLOBECOM
    Graph Property Preservation under Community-Based Sampling
    Ruohan Gao, Pili Hu, and Wing Cheong Lau
    IEEE Global Communications Conference (GLOBECOM), 2015