I am an AI Researcher. I obtained my master's degree from Nanyang Technological University. My current research interests primarily focus on Vision-Language Models (VLM), post-training data flywheel, pretrained visual encoders, 3D scene understanding, and embodied intelligence.
Additionally, I have a strong passion for portrait photography and have been practicing it intensively lately. The application of large language models (LLMs) in quantitative finance is another area I actively follow in my spare time—I've recently been replicating and designing quantitative investment agents related to this field.
We are now recruiting project/research interns. If you are interested in, please directly send your CV to yzhou037@e.ntu.edu.sg.
") does not match the recommended repository name for your site ("").
", so that your site can be accessed directly at "http://".
However, if the current repository name is intended, you can ignore this message by removing "{% include widgets/debug_repo_name.html %}" in index.html.
",
which does not match the baseurl ("") configured in _config.yml.
baseurl in _config.yml to "".
MM Team, S Bai, L Bing, C Chen, G Chen, Y Chen, Z Chen, J Dai, et al.
arXiv preprint 2025
MiroThinker scales open-source research agents along three dimensions—model capacity, context length, and interactive reasoning—achieving state-of-the-art performance on research agent benchmarks.
MM Team, S Bai, L Bing, C Chen, G Chen, Y Chen, Z Chen, J Dai, et al.
arXiv preprint 2025
MiroThinker scales open-source research agents along three dimensions—model capacity, context length, and interactive reasoning—achieving state-of-the-art performance on research agent benchmarks.
J Zhang, Y Chen, Y Xu, Z Huang, Yanpeng Zhou, YJ Yuan, X Cai, G Huang, et al.
Neural Information Processing Systems (NeurIPS) 2025
4D-VLA introduces spatiotemporal pretraining for Vision-Language-Action models with cross-scene calibration, enabling embodied agents to better understand and act in dynamic 3D environments.
J Zhang, Y Chen, Y Xu, Z Huang, Yanpeng Zhou, YJ Yuan, X Cai, G Huang, et al.
Neural Information Processing Systems (NeurIPS) 2025
4D-VLA introduces spatiotemporal pretraining for Vision-Language-Action models with cross-scene calibration, enabling embodied agents to better understand and act in dynamic 3D environments.
J Zhang, Y Chen, Yanpeng Zhou, Y Xu, Z Huang, J Mei, J Chen, YJ Yuan, X Cai, et al.
Neural Information Processing Systems (NeurIPS) 2025
We propose a framework that lifts Vision-Language Models from 2D flat image understanding to full 3D spatial perception and reasoning, bridging the gap between visual recognition and geometric understanding.
J Zhang, Y Chen, Yanpeng Zhou, Y Xu, Z Huang, J Mei, J Chen, YJ Yuan, X Cai, et al.
Neural Information Processing Systems (NeurIPS) 2025
We propose a framework that lifts Vision-Language Models from 2D flat image understanding to full 3D spatial perception and reasoning, bridging the gap between visual recognition and geometric understanding.
Y Xu, J Zhang, Z Huang, Y Chen, Yanpeng Zhou, Z Chen, YJ Yuan, P Xia, et al.
International Conference on Learning Representations (ICLR) 2025
UniUGG presents a unified framework for 3D scene understanding and generation by jointly encoding geometric structure and semantic content, enabling coherent cross-task 3D reasoning.
Y Xu, J Zhang, Z Huang, Y Chen, Yanpeng Zhou, Z Chen, YJ Yuan, P Xia, et al.
International Conference on Learning Representations (ICLR) 2025
UniUGG presents a unified framework for 3D scene understanding and generation by jointly encoding geometric structure and semantic content, enabling coherent cross-task 3D reasoning.
H Li, Yanpeng Zhou, T Tang, J Song, Y Zeng, M Kampffmeyer, H Xu, X Liang
International Conference on Learning Representations (ICLR) 2025
UniGS unifies language, image, and 3D understanding through pretraining with Gaussian Splatting representations, enabling a single model to reason across all three modalities.
H Li, Yanpeng Zhou, T Tang, J Song, Y Zeng, M Kampffmeyer, H Xu, X Liang
International Conference on Learning Representations (ICLR) 2025
UniGS unifies language, image, and 3D understanding through pretraining with Gaussian Splatting representations, enabling a single model to reason across all three modalities.
S Zheng, Z Peng, Yanpeng Zhou, Y Zhu, H Xu, X Huang, Y Fu
arXiv preprint 2025
VidCraft3 enables fine-grained control over camera movement, object motion, and lighting conditions in image-to-video generation, advancing controllable video synthesis.
S Zheng, Z Peng, Yanpeng Zhou, Y Zhu, H Xu, X Huang, Y Fu
arXiv preprint 2025
VidCraft3 enables fine-grained control over camera movement, object motion, and lighting conditions in image-to-video generation, advancing controllable video synthesis.
Y Zhu, Yanpeng Zhou, C Wang, Y Cao, J Han, L Hou, H Xu
Neural Information Processing Systems (NeurIPS) 2024
UNIT is a unified vision encoder that jointly handles image and text recognition tasks within a single model, enabling versatile visual understanding without task-specific architectures.
Y Zhu, Yanpeng Zhou, C Wang, Y Cao, J Han, L Hou, H Xu
Neural Information Processing Systems (NeurIPS) 2024
UNIT is a unified vision encoder that jointly handles image and text recognition tasks within a single model, enabling versatile visual understanding without task-specific architectures.
D Yang, Yanpeng Zhou, Z Zhang, TJJ Li, R Lc
Joint Proceedings of the ACM IUI Workshops 2022
We explore interaction strategies between human writers and AI-generated text in collaborative fiction writing, examining how positioning AI as an active writer changes the creative process.
D Yang, Yanpeng Zhou, Z Zhang, TJJ Li, R Lc
Joint Proceedings of the ACM IUI Workshops 2022
We explore interaction strategies between human writers and AI-generated text in collaborative fiction writing, examining how positioning AI as an active writer changes the creative process.