About Me
- Computer Vision
- Vision & Language
- Embodied AI
My research lies at the intersection of computer vision, natural language, and robotics, with a recent focus on Embodied AI. I aim to build robots that can understand human instructions, perceive and reason about 3D environments, and act reliably in the physical world. My current work centers on:
- Robot Manipulation: language-guided manipulation with robotic arms, and mobile manipulation with arm-equipped quadruped robots.
- Human-Robot Interaction: interactive grounding that resolves ambiguous human instructions through dialogue.
- Multimodal Perception: 3D visual grounding and 3D affordance understanding for actionable scene perception.
Previously, I worked on language-driven video understanding and open-vocabulary image/video recognition, as well as hand detection, hand pose estimation, face recognition, and person re-identification.
News
- 2026.10 A language-driven action localization paper is accepted by IJCV 2026 (CCF-A, 中科院一区, JCR Q1, IF=10.3)!
- 2026.07 An open-vocabulary multi-label action recognition paper is accepted by CVIU 2026 (CCF-B, JCR Q2, IF=3.6)!
- 2026.05 An interactive 3D grounding framework and dataset paper is accepted by ICML 2026 (CCF-A conference)!
- 2025.12 An image-free multi-label image recognition paper is accepted by Pattern Recognition 2026 (中科院一区, JCR Q1, IF=7.6)!
- 2025.06 An image-text matching paper is accepted by ICCV 2025 (CCF-A conference)!
- 2025.04 A video visual relationship detection paper is accepted by IJCAI 2025 (CCF-A conference)!
- 2025.04 A video visual relationship detection paper is accepted by IEEE TPAMI 2025 (CCF-A, 中科院一区, JCR Q1, IF=20.8)!
- 2025.01 An open-vocabulary multi-label action classification paper is published in 《计算机研究与发展》 2025 (CCF-A Chinese, IF=2.65)!
- 2024.12 A video Summarization paper is accepted by AAAI 2025 (CCF-A conference)!
- 2024.10 An image-text matching paper is accepted by IEEE Signal Processing Letter 2024 (JCR Q2, 中科院三区, IF=3.2)!
- 2024.10 A language-driven action localization paper is accepted by PRCV 2024 (CCF-C conference)!
- 2024.06 I graduated from Beijing Institute of Technology (北京理工大学) and got a position as an Associate Professor at Shenzhen MSU-BIT University (深圳北理莫斯科大学)!
- 2024.02 A language-driven action localization paper is accepted by IEEE TMM 2024 (中科院一区, JCR Q1, IF=7.3)!
- 2023.12 A video visual relationship detection paper is accepted by AAAI 2024 (CCF-A conference)!
- 2023.07 A frame-supervised language-driven action localization paper is accepted by ACM MM 2023 (CCF-A conference)!
- 2022.04 A language-driven action localization paper is accepted by IJCAI 2022 (CCF-A conference)!
- 2021.06 I attend a new research group under supervised by Prof.Xinxiao Wu.
- 2020.03 A person re-identification paper is accepted by CVPR 2020 (CCF-A conference)!
Publications
-
AmbiRefer3D: 3D Visual Grounding with Referential Ambiguity
International Conference on Machine Learning (ICML), 2026
-
Image-free Multi-label Image Recognition via LLM-powered Hierarchical Prompt Tuning
Pattern Recognition (PR), 2026
-
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
International Conference on Computer Vision (ICCV), 2025
-
International Joint Conference on Artificial Intelligence (IJCAI), 2025
-
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
-
Video Summarization using Denoising Diffusion Probabilistic Model
AAAI Conference on Artificial Intelligence (AAAI), 2025
-
Dynamic Pathway for Query-Aware Feature Learning in Language-Driven Action Localization
IEEE Transactions on Multimedia (TMM), 2024
-
Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship Detection
AAAI Conference on Artificial Intelligence (AAAI), 2024
-
Probability Distribution Based Frame-supervised Language-driven Action Localization
ACM International Conference on Multimedia (ACM MM), 2023
-
Entity-aware and Motion-aware Transformers for Language-driven Action Localization
International Joint Conference on Artificial Intelligence (IJCAI), 2022
-
High-Order Information Matters: Learning Relation and Topology for Occluded Person Re-Identification
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020
-
Joint Hand Detection and Rotation Estimation Using CNN
IEEE Transactions on Image Processing (TIP), 2018
- CVIU 2026 Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning. Rongjiang Zhu, Xinxiao Wu, Shuo Yang†, Yuheng Shi, Ziyi Wang
- 计算机研究与发展 2025 大语言模型知识引导的开放域多标签动作识别. 朱荣江, 石语珩, 杨硕, 王子奕, 吴心筱
- SPL 2024 Source-free Image-text Matching via Uncertainty-aware Learning. Mengxiao Tian, Shuo Yang†, Xinxiao Wu, Yunde Jia
- PRCV 2024 Efficient Language-Driven Action Localization by Feature Aggregation and Prediction Adjustment. Zirui Shang, Shuo Yang†, Xinxiao Wu
- arXiv 2017 Hand3D: Hand Pose Estimation using 3D Neural Network. Xiaoming Deng*, Shuo Yang*, Yinda Zhang*, Ping Tan, Liang Chang, Hongan Wang
- Acta Automatica Sinica 2016 Convolutional neural networks in image understanding. Liang Chang, Xiaoming Deng, Mingquan Zhou, Zhongke Wu, Ye Yuan, Shuo Yang, Hongan Wang
Education
- 2018.09 - 2024.06Ph.D. in Computer Science, School of Computer Science & Technology, Beijing Institute of TechnologyAdvisor: Shuliang Wang(2018.09 - 2021.06) and Xinxiao Wu from 2021.06.
- 2014.09 - 2017.07M.S. in Computer Science, Institute of Software, Chinese Academic of ScienceAdvisor: Xiaoming Deng.
- 2010.09 - 2014.07B.S. in Computer Science, School of Information, Beijing Union University.
Experience
- 2024.06 - nowAssociate Professor at Shenzhen MSU-BIT University, Shenzhen, China.
- 2019.05 - 2020.02Research intern at Megvii-inc, Beijing, China.
- 2017.07 - 2018.08Algorithm engineer at JD Finance, Beijing, China.