About Me

  • Computer Vision
  • Vision & Language
  • Embodied AI

My research lies at the intersection of computer vision, natural language, and robotics, with a recent focus on Embodied AI. I aim to build robots that can understand human instructions, perceive and reason about 3D environments, and act reliably in the physical world. My current work centers on:

  • Robot Manipulation: language-guided manipulation with robotic arms, and mobile manipulation with arm-equipped quadruped robots.
  • Human-Robot Interaction: interactive grounding that resolves ambiguous human instructions through dialogue.
  • Multimodal Perception: 3D visual grounding and 3D affordance understanding for actionable scene perception.

Previously, I worked on language-driven video understanding and open-vocabulary image/video recognition, as well as hand detection, hand pose estimation, face recognition, and person re-identification.

🎓 Welcome students who are interested in the research of Embodied AI and Vision & Language to join us!

News

Publications

Google Scholar citations * equal contribution  ·  † corresponding author
  1. AmbiRefer3D: 3D Visual Grounding with Referential Ambiguity

    Rongjiang Zhu*, Wei Kang*, Zeqi Liu, Junyu Chen, Shuo Yang†, Xinxiao Wu†

    International Conference on Machine Learning (ICML), 2026

    
      
  2. Image-free Multi-label Image Recognition via LLM-powered Hierarchical Prompt Tuning

    Shuo Yang†, Zirui Shang, Yongqi Wang, Derong Deng, Hongwei Chen, Xinxiao Wu, Qiyuan Cheng

    Pattern Recognition (PR), 2026

    
      
  3. LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching

    Mengxiao Tian, Xinxiao Wu, Shuo Yang†

    International Conference on Computer Vision (ICCV), 2025

    
      
  4. METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection

    Yongqi Wang, Xinxiao Wu, Shuo Yang†

    International Joint Conference on Artificial Intelligence (IJCAI), 2025

  5. End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting

    Yongqi Wang, Xinxiao Wu, Shuo Yang, Jiebo Luo

    IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

    
      

Education

  • 2018.09 - 2024.06
    Ph.D. in Computer Science, School of Computer Science & Technology, Beijing Institute of Technology
    Advisor: Shuliang Wang(2018.09 - 2021.06) and Xinxiao Wu from 2021.06.
  • 2014.09 - 2017.07
    M.S. in Computer Science, Institute of Software, Chinese Academic of Science
    Advisor: Xiaoming Deng.
  • 2010.09 - 2014.07
    B.S. in Computer Science, School of Information, Beijing Union University.

Experience