Yuanze Lin
I am a DPhil student in the Computer Science Department at the University of Oxford , working with Prof. Ronald Clark and Prof. Philip Torr on diffusion and vision-language models.
I have also worked at Microsoft Research , MSR Asia , Alibaba , and CCVL @ Johns Hopkins University . I appreciate collaborating with distinguished professors and researchers from these institutions.
My research builds generative and multimodal foundation models that perceive, reason about, and generate the physical 3D world:
• Generative models of the physical world — geometry, illumination, avatars, and video
• Spatial, embodied, and 3D world models with physical grounding and spatial reasoning
• Efficient multimodal foundation models for vision–language reasoning and task routing
yuanze.lin [at] cs.ox.ac.uk
Google Scholar
GitHub
LinkedIn
News
Selected Publications
Papers are sorted by recency, * denotes equal contribution.
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
Yuanze Lin ,
Yi-Wen Chen ,
Yi-Hsuan Tsai ,
Ronald Clark ,
Ming-Hsuan Yang
NeurIPS , 2025
ArXiv
/
Project Page
/
Video
/
Code
/
BibTeX
We present IllumiCraft, a unified framework that unifies geometry and illumination diffusion for controllable video generation.
Olympus: A Universal Task Router for Computer Vision Tasks
Yuanze Lin ,
Yunsheng Li ,
Dongdong Chen ,
Weijian Xu ,
Ronald Clark ,
Philip Torr
CVPR , 2025 ★ Highlight
ArXiv
/
Project Page
/
Video
/
Poster
/
Code
/
BibTeX
Turns MLLMs into a universal task router that handles a wide array of computer vision tasks within a single unified framework.
Text-Driven Image Editing via Learnable Regions
Yuanze Lin ,
Yi-Wen Chen ,
Yi-Hsuan Tsai ,
Lu Jiang ,
Ming-Hsuan Yang
CVPR , 2024
ArXiv
/
Project Page
/
Video
/
Poster
/
Code
/
BibTeX
A region-based network trained with a CLIP-guided text-driven loss, editing images from freely provided language descriptions.
SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training
Yuanze Lin ,
Chen Wei ,
Huiyu Wang ,
Alan Yuille ,
Cihang Xie
ICCV , 2023
ArXiv
/
Poster
/
Slides
/
BibTeX
An efficient video-language pre-training framework that stays competitive on retrieval and video QA while cutting pre-training cost by 1.9X or more.
REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question Answering
Yuanze Lin , Yujia Xie , Dongdong Chen , Yichong Xu , Chenguang Zhu ,
Lu Yuan
NeurIPS , 2022
ArXiv /
Poster /
Supplementary Material /
OpenReview /
Code /
BibTeX
A knowledge-based VQA method exploiting explicit object-region information in both retrieval and answering, reaching state-of-the-art on OK-VQA.
Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding
Haojun Jiang* ,
Yuanze Lin* ,
Dongchen Han ,
Shiji Song ,
Gao Huang
CVPR , 2022
ArXiv /
Poster /
Code /
BibTeX
Automatically generates pseudo language queries for supervised training, matching or beating weakly-supervised visual grounding across five datasets.
AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video Recognition
Yulin Wang* ,
Yang Yue* ,
Yuanze Lin ,
Haojun Jiang ,
Zihang Lai ,
Victor Kulikov ,
Nikita Orlov ,
Humphrey Shi ,
Gao Huang
CVPR , 2022
ArXiv /
Code /
BibTeX
Reformulates AdaFocus as a one-stage algorithm via differentiable patch selection and an improved training scheme, validated on six benchmarks.
Self-supervised video representation learning with meta-contrastive network
Yuanze Lin ,
Xun Guo ,
Yan Lu
ICCV , 2021
ArXiv /
Poster /
BibTeX
A Meta-Contrastive Network combining contrastive and meta-learning for pre-training, surpassing prior methods on UCF101 and HMDB51.
EVA-GCN: Head Pose Estimation Based on Graph Convolutional Networks
Miao Xin ,
Shentong Mo ,
Yuanze Lin
CVPR AMFG Workshop , 2021   🏆 Best Paper Award
Paper /
Code /
BibTeX
Builds a landmark-connection graph and uses Graph Convolutional Networks to model nonlinear mappings from graph topology to head-pose angles.
Researcher Intern, Jul 2026 - Present
Researcher Intern, Feb 2024 - Nov 2024
hosted by Dr.
Dongdong Chen , on unified MLLMs for multi-task vision understanding and generation.
Senior Algorithm Engineer, Feb 2023 - Aug 2023
on large-scale vision-language pre-training, fine-tuning, and downstream applications of LLMs.
Researcher Intern, Feb 2022 - June 2022
Researcher Intern, Dec 2020 - Sep 2021
with Dr.
Xun Guo and Dr.
Yan Lu on self-supervised learning for video.
Professional Services
Program Comittee: AAAI 2025, AAAI 2026
Journal Reviewer: IJCV 2025
Conference Reviewer: ICLR 2026, AISTATS 2026, CVPR 2026, ECCV 2026
Conference Reviewer: ICLR 2025, AISTATS 2025, CVPR 2025, ICML 2025, ICCV 2025, NeurIPS 2025
Conference Reviewer: ICRA 2024, CVPR 2024, ECCV 2024, NeurIPS 2024
Conference Reviewer: ICLR 2023, CVPR 2023, ICCV 2023, NeurIPS 2023
Conference Reviewer: CVPR 2022
Photos
A few things I walked past. Swipe, or use the arrows.