个人简介
- 计算机视觉
- 视觉与语言
- 具身智能
本人研究方向为计算机视觉、自然语言与机器人的交叉领域,近期主要聚焦于具身智能,致力于让机器人能够理解人类指令、感知并推理三维环境,并在真实物理世界中可靠地完成任务。目前的研究主要包括:
- 机器人操作:语言引导的机械臂操作,以及机器狗搭载机械臂的移动操作。
- 人机交互:交互式视觉定位,通过对话消解人类指令中的歧义。
- 多模态感知:三维视觉定位(3D Visual Grounding)与三维可供性(3D Affordance)理解,面向可执行的场景感知。
此前从事语言驱动的视频理解、开放词汇图像/视频识别,以及手部检测、手部姿态估计、人脸识别与行人重识别等研究。
学术动态
- 2026.10 语言驱动的动作定位 论文被 IJCV 2026 录用(CCF-A, 中科院一区, JCR Q1, IF=10.3)!
- 2026.07 开放词汇多标签动作识别 论文被 CVIU 2026 录用(CCF-B, JCR Q2, IF=3.6)!
- 2026.05 交互式三维场景定位(3D grounding)框架与数据集 论文被 ICML 2026 录用(CCF-A)!
- 2025.12 无图像多标签图像识别 论文被 Pattern Recognition 2026 录用(中科院一区, JCR Q1, IF=7.6)!
- 2025.06 图像—文本匹配 论文被 ICCV 2025 录用(CCF-A)!
- 2025.04 开放词汇视频视觉关系检测 论文被 IJCAI 2025 录用(CCF-A)!
- 2025.04 开放词汇视频视觉关系检测 论文发表于 IEEE TPAMI 2025(CCF-A, 中科院一区, JCR Q1, IF=20.8)!
- 2025.01 开放词汇多标签动作分类 论文发表于 《计算机研究与发展》 2025(CCF-A 中文期刊, IF=2.65)!
- 2024.12 视频摘要 论文被 AAAI 2025 录用(CCF-A)!
- 2024.10 图像—文本匹配 论文被 IEEE SPL 2024 录用(JCR Q2, 中科院三区, IF=3.2)!
- 2024.10 语言驱动的动作定位 论文被 PRCV 2024 录用(CCF-C)!
- 2024.06 毕业于北京理工大学,入职 深圳北理莫斯科大学 任副教授。
- 2024.02 语言驱动的动作定位 论文被 IEEE TMM 2024 录用(中科院一区, JCR Q1, IF=7.3)!
- 2023.12 开放词汇视频视觉关系检测 论文被 AAAI 2024 录用(CCF-A)!
- 2023.07 帧监督语言驱动的动作定位 论文被 ACM MM 2023 录用(CCF-A)!
- 2022.04 语言驱动的动作定位 论文被 IJCAI 2022 录用(CCF-A)!
- 2021.06 加入 吴心筱 教授课题组。
- 2020.03 行人重识别 论文被 CVPR 2020 录用(CCF-A)!
论文
-
AmbiRefer3D: 3D Visual Grounding with Referential Ambiguity
International Conference on Machine Learning (ICML), 2026
-
Image-free Multi-label Image Recognition via LLM-powered Hierarchical Prompt Tuning
Pattern Recognition (PR), 2026
-
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
International Conference on Computer Vision (ICCV), 2025
-
International Joint Conference on Artificial Intelligence (IJCAI), 2025
-
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
-
Video Summarization using Denoising Diffusion Probabilistic Model
AAAI Conference on Artificial Intelligence (AAAI), 2025
-
Dynamic Pathway for Query-Aware Feature Learning in Language-Driven Action Localization
IEEE Transactions on Multimedia (TMM), 2024
-
Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship Detection
AAAI Conference on Artificial Intelligence (AAAI), 2024
-
Probability Distribution Based Frame-supervised Language-driven Action Localization
ACM International Conference on Multimedia (ACM MM), 2023
-
Entity-aware and Motion-aware Transformers for Language-driven Action Localization
International Joint Conference on Artificial Intelligence (IJCAI), 2022
-
High-Order Information Matters: Learning Relation and Topology for Occluded Person Re-Identification
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020
-
Joint Hand Detection and Rotation Estimation Using CNN
IEEE Transactions on Image Processing (TIP), 2018
- CVIU 2026 Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning. Rongjiang Zhu, Xinxiao Wu, Shuo Yang†, Yuheng Shi, Ziyi Wang
- 计算机研究与发展 2025 大语言模型知识引导的开放域多标签动作识别. 朱荣江, 石语珩, 杨硕, 王子奕, 吴心筱
- SPL 2024 Source-free Image-text Matching via Uncertainty-aware Learning. Mengxiao Tian, Shuo Yang†, Xinxiao Wu, Yunde Jia
- PRCV 2024 Efficient Language-Driven Action Localization by Feature Aggregation and Prediction Adjustment. Zirui Shang, Shuo Yang†, Xinxiao Wu
- arXiv 2017 Hand3D: Hand Pose Estimation using 3D Neural Network. Xiaoming Deng*, Shuo Yang*, Yinda Zhang*, Ping Tan, Liang Chang, Hongan Wang
- Acta Automatica Sinica 2016 Convolutional neural networks in image understanding. Liang Chang, Xiaoming Deng, Mingquan Zhou, Zhongke Wu, Ye Yuan, Shuo Yang, Hongan Wang
教育背景
- 2018.09 - 2024.06
- 2014.09 - 2017.07计算机科学与技术硕士,中国科学院软件研究所导师:邓小明。
- 2010.09 - 2014.07计算机科学与技术学士,北京联合大学信息学院。