Biography
I am a Researcher at Shanghai AI Laboratory. My research focuses on large language model post-training, agentic RL and recursive self-improvement, with an emphasis on improving reasoning capabilities and enabling autonomous self-improvement.
I received my Ph.D. from The Hong Kong University of Science and Technology (HKUST) in 2023, advised by Prof. Chi-Keung Tang and Prof. Yu-Wing Tai. Before that, I completed my undergraduate studies in Computer Science at Shanghai Jiao Tong University in 2017 and worked at Tencent YouTu Lab from 2017 to 2019.
Selected Publications [Full list on Google Scholar]
* Equal contribution (co-first authors); † Corresponding author.
-
MegaStyle++: Scaling Image Style Space through Hierarchical Style Definition
arXiv preprint, 2026.
-
Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026.
[Paper]
-
Intern-S2-Preview: Scientific Agentic Foundation Model
Technical report, 2026.
[Paper] [GitHub] [Hugging Face]
-
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
Findings of the Association for Computational Linguistics (ACL), 2026.
[Paper]
-
Human-Agent Collaborative Paper-to-Page Crafting
Findings of the Association for Computational Linguistics (ACL), 2026.
-
MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing
arXiv preprint, 2026.
-
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
arXiv preprint, 2026.
-
MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
arXiv preprint, 2026.
[Paper] [Code] [Project] [Hugging Face] [Dataset]
-
Test-Time Correction: An Online 3D Detection System via Visual Prompting
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026.
-
StyleShot: A Snapshot on Any Style
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026.
-
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
arXiv preprint, 2025.
-
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
IEEE/CVF International Conference on Computer Vision (ICCV), 2025.
[Paper]
-
Detect Anything 3D in the Wild
IEEE/CVF International Conference on Computer Vision (ICCV), 2025.
-
FaceShot: Bring Any Character into Life
International Conference on Learning Representations (ICLR), 2025.
-
CharacterShot: Controllable and Consistent 4D Character Animation
arXiv preprint, 2025.
-
AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation
European Conference on Computer Vision (ECCV), 2024.
[Paper] [Code] [Project] [Hugging Face]
-
Visual Point Cloud Forecasting Enables Scalable Autonomous Driving
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
-
Semantic Image Matting: General and Specific Semantics
International Journal of Computer Vision (IJCV), 2024.
-
Ultrahigh Resolution Image/Video Matting with Spatio-Temporal Sparsity
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
-
Human Instance Matting via Mutual Guidance and Multi-Instance Refinement
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. Oral presentation.
-
A Unified Query-Based Paradigm for Point Cloud Understanding
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
-
Semantic Image Matting
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
-
Deep Video Matting via Spatio-Temporal Alignment and Aggregation
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
-
3DSSD: Point-Based 3D Single Stage Object Detector
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. Oral presentation.
-
GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-Aware Supervision
European Conference on Computer Vision (ECCV), 2020.
-
CN: Channel Normalization for Point Cloud Recognition
European Conference on Computer Vision (ECCV), 2020.
[Paper]
-
STD: Sparse-to-Dense 3D Object Detector for Point Cloud
IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
[Paper]
Experience
Shanghai AI Laboratory
Researcher
2023 – Present
Tencent YouTu Lab
Researcher
2017 – 2019
Education
The Hong Kong University of Science and Technology
Ph.D., Department of Computer Science and Engineering
Advisors: Prof. Chi-Keung Tang and Prof. Yu-Wing Tai
2019 – 2023
Shanghai Jiao Tong University
Undergraduate studies, Department of Computer Science and Technology
2013 – 2017
Academic Service
Conference Reviewer
- Annual Meeting of the Association for Computational Linguistics (ACL)
- Conference on Empirical Methods in Natural Language Processing (EMNLP)
- IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- IEEE/CVF International Conference on Computer Vision (ICCV)
- European Conference on Computer Vision (ECCV)
- Conference on Neural Information Processing Systems (NeurIPS)
- International Conference on Learning Representations (ICLR)
- ACM International Conference on Multimedia (ACM MM)
- British Machine Vision Conference (BMVC)
- ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
Journal Reviewer
- IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
- International Journal of Computer Vision (IJCV)
- IEEE Transactions on Image Processing (TIP)
Honors & Awards
- Postdoctoral Innovative Talent Support Program2025 – Present
- Shanghai Overseas High-Level Talent Recruitment Program2024 – 2027
- Hong Kong Postgraduate Scholarship2019 – 2023
- Outstanding Graduate of Shanghai Jiao Tong University2017
- National Scholarship2013 – 2017