Portrait
Yanpeng Zhou
Research Scientist
--
About Me

I am an AI Researcher. I obtained my master's degree from Nanyang Technological University. My current research interests primarily focus on Vision-Language Models (VLM), post-training data flywheel, pretrained visual encoders, 3D scene understanding, and embodied intelligence.

Additionally, I have a strong passion for portrait photography and have been practicing it intensively lately. The application of large language models (LLMs) in quantitative finance is another area I actively follow in my spare time—I've recently been replicating and designing quantitative investment agents related to this field.

We are now recruiting project/research interns. If you are interested in, please directly send your CV to yzhou037@e.ntu.edu.sg.

Education
  • Nanyang Technological University
    Nanyang Technological University
    Master's Degree
Experience
  • Huawei Noah's Ark Laboratory
    Huawei Noah's Ark Laboratory
    AI Researcher
    2022 - 2025
Selected Publications (view all )
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

MM Team, S Bai, L Bing, C Chen, G Chen, Y Chen, Z Chen, J Dai, et al.

arXiv preprint 2025

MiroThinker scales open-source research agents along three dimensions—model capacity, context length, and interactive reasoning—achieving state-of-the-art performance on research agent benchmarks.

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

MM Team, S Bai, L Bing, C Chen, G Chen, Y Chen, Z Chen, J Dai, et al.

arXiv preprint 2025

MiroThinker scales open-source research agents along three dimensions—model capacity, context length, and interactive reasoning—achieving state-of-the-art performance on research agent benchmarks.

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

J Zhang, Y Chen, Y Xu, Z Huang, Yanpeng Zhou, YJ Yuan, X Cai, G Huang, et al.

Neural Information Processing Systems (NeurIPS) 2025

4D-VLA introduces spatiotemporal pretraining for Vision-Language-Action models with cross-scene calibration, enabling embodied agents to better understand and act in dynamic 3D environments.

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

J Zhang, Y Chen, Y Xu, Z Huang, Yanpeng Zhou, YJ Yuan, X Cai, G Huang, et al.

Neural Information Processing Systems (NeurIPS) 2025

4D-VLA introduces spatiotemporal pretraining for Vision-Language-Action models with cross-scene calibration, enabling embodied agents to better understand and act in dynamic 3D environments.

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

J Zhang, Y Chen, Yanpeng Zhou, Y Xu, Z Huang, J Mei, J Chen, YJ Yuan, X Cai, et al.

Neural Information Processing Systems (NeurIPS) 2025

We propose a framework that lifts Vision-Language Models from 2D flat image understanding to full 3D spatial perception and reasoning, bridging the gap between visual recognition and geometric understanding.

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

J Zhang, Y Chen, Yanpeng Zhou, Y Xu, Z Huang, J Mei, J Chen, YJ Yuan, X Cai, et al.

Neural Information Processing Systems (NeurIPS) 2025

We propose a framework that lifts Vision-Language Models from 2D flat image understanding to full 3D spatial perception and reasoning, bridging the gap between visual recognition and geometric understanding.

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

Y Xu, J Zhang, Z Huang, Y Chen, Yanpeng Zhou, Z Chen, YJ Yuan, P Xia, et al.

International Conference on Learning Representations (ICLR) 2025

UniUGG presents a unified framework for 3D scene understanding and generation by jointly encoding geometric structure and semantic content, enabling coherent cross-task 3D reasoning.

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

Y Xu, J Zhang, Z Huang, Y Chen, Yanpeng Zhou, Z Chen, YJ Yuan, P Xia, et al.

International Conference on Learning Representations (ICLR) 2025

UniUGG presents a unified framework for 3D scene understanding and generation by jointly encoding geometric structure and semantic content, enabling coherent cross-task 3D reasoning.

UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting

H Li, Yanpeng Zhou, T Tang, J Song, Y Zeng, M Kampffmeyer, H Xu, X Liang

International Conference on Learning Representations (ICLR) 2025

UniGS unifies language, image, and 3D understanding through pretraining with Gaussian Splatting representations, enabling a single model to reason across all three modalities.

UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting

H Li, Yanpeng Zhou, T Tang, J Song, Y Zeng, M Kampffmeyer, H Xu, X Liang

International Conference on Learning Representations (ICLR) 2025

UniGS unifies language, image, and 3D understanding through pretraining with Gaussian Splatting representations, enabling a single model to reason across all three modalities.

VidCraft3: Camera, Object, and Lighting Control for Image-to-Video Generation

S Zheng, Z Peng, Yanpeng Zhou, Y Zhu, H Xu, X Huang, Y Fu

arXiv preprint 2025

VidCraft3 enables fine-grained control over camera movement, object motion, and lighting conditions in image-to-video generation, advancing controllable video synthesis.

VidCraft3: Camera, Object, and Lighting Control for Image-to-Video Generation

S Zheng, Z Peng, Yanpeng Zhou, Y Zhu, H Xu, X Huang, Y Fu

arXiv preprint 2025

VidCraft3 enables fine-grained control over camera movement, object motion, and lighting conditions in image-to-video generation, advancing controllable video synthesis.

UNIT: Unifying Image and Text Recognition in One Vision Encoder

Y Zhu, Yanpeng Zhou, C Wang, Y Cao, J Han, L Hou, H Xu

Neural Information Processing Systems (NeurIPS) 2024

UNIT is a unified vision encoder that jointly handles image and text recognition tasks within a single model, enabling versatile visual understanding without task-specific architectures.

UNIT: Unifying Image and Text Recognition in One Vision Encoder

Y Zhu, Yanpeng Zhou, C Wang, Y Cao, J Han, L Hou, H Xu

Neural Information Processing Systems (NeurIPS) 2024

UNIT is a unified vision encoder that jointly handles image and text recognition tasks within a single model, enabling versatile visual understanding without task-specific architectures.

AI as an Active Writer: Interaction Strategies with Generated Text in Human-AI Collaborative Fiction Writing

D Yang, Yanpeng Zhou, Z Zhang, TJJ Li, R Lc

Joint Proceedings of the ACM IUI Workshops 2022

We explore interaction strategies between human writers and AI-generated text in collaborative fiction writing, examining how positioning AI as an active writer changes the creative process.

AI as an Active Writer: Interaction Strategies with Generated Text in Human-AI Collaborative Fiction Writing

D Yang, Yanpeng Zhou, Z Zhang, TJJ Li, R Lc

Joint Proceedings of the ACM IUI Workshops 2022

We explore interaction strategies between human writers and AI-generated text in collaborative fiction writing, examining how positioning AI as an active writer changes the creative process.

All publications