Yuanze Lin

I am a DPhil student in the Computer Science Department at the University of Oxford, working with Prof. Ronald Clark and Prof. Philip Torr on diffusion and vision-language models.

Before going to Oxford, I had a great time at Microsoft Redmond, MSR Asia, CCVL @ Johns Hopkins University, Alibaba, etc. I appreciate collaborating with distinguished professors and researchers from these institutions.

My research interests span machine learning, especially:

 •   Multimodal large language models (MLLMs)

 •   Diffusion-based image/video generation and editing

 •   Applications of large language models (LLMs)

profile photo

  News


[07/2026] Started a research internship at Microsoft Research Cambridge.

[09/2025] IllumiCraft accepted to NeurIPS 2025.

[06/2025] We introduced IllumiCraft for high-fidelity video relighting.

[04/2025] Olympus was selected as a Highlight at CVPR 2025!

[03/2025] Released the code of Olympus (CVPR 2025) for vision tasks.

[02/2025] Olympus accepted to CVPR 2025.

[12/2024] We presented Olympus to solve over 20 different computer vision tasks.

[08/2024] Released the code of Learnable Regions (CVPR2024) for image editing.

[07/2024] Rethinking Visual Prompting for MLLMs with External Knowledge was presented.

[03/2024] Check out DreamPolisher for high-quality text-to-3D generation!

[02/2024] Started a research internship at GenAI @ Microsoft.

[02/2024] Text-Driven Image Editing via Learnable Regions accepted to CVPR 2024.

[10/2023] Started my DPhil journey at CS @ University of Oxford.

[07/2023] SMAUG accepted to ICCV 2023.

[09/2022] REVIVE accepted to NeurIPS 2022.

[03/2022] Pseudo-Q and AdaFocus V2 accepted to CVPR 2022.

[07/2021] MCN accepted to ICCV 2021.

[06/2021] EVA-GCN accepted to CVPR 2021 AMFG Workshop and won 🏆 Best Paper Award

 Selected Publications


Papers are sorted by recency, * denotes equal contribution.

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark, Ming-Hsuan Yang
NeurIPS, 2025 
ArXiv / Project Page / Video / Code / BibTeX

We present IllumiCraft, a unified framework that unifies geometry and illumination diffusion for controllable video generation.

Olympus: A Universal Task Router for Computer Vision Tasks
Yuanze Lin, Yunsheng Li, Dongdong Chen, Weijian Xu, Ronald Clark, Philip Torr
CVPR, 2025  ★ Highlight
ArXiv / Project Page / Video / Poster / Code / BibTeX

Turns MLLMs into a universal task router that handles a wide array of computer vision tasks within a single unified framework.

Text-Driven Image Editing via Learnable Regions
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Lu Jiang, Ming-Hsuan Yang
CVPR, 2024
ArXiv / Project Page / Video / Poster / Code / BibTeX

A region-based network trained with a CLIP-guided text-driven loss, editing images from freely provided language descriptions.

SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training
Yuanze Lin, Chen Wei, Huiyu Wang, Alan Yuille, Cihang Xie
ICCV, 2023
ArXiv / Poster / Slides / BibTeX

An efficient video-language pre-training framework that stays competitive on retrieval and video QA while cutting pre-training cost by 1.9X or more.

REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question Answering
Yuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu, Chenguang Zhu, Lu Yuan
NeurIPS, 2022
ArXiv / Poster / Supplementary Material / OpenReview / Code / BibTeX

A knowledge-based VQA method exploiting explicit object-region information in both retrieval and answering, reaching state-of-the-art on OK-VQA.

Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding
Haojun Jiang*, Yuanze Lin*, Dongchen Han, Shiji Song, Gao Huang
CVPR, 2022
ArXiv / Poster / Code / BibTeX

Automatically generates pseudo language queries for supervised training, matching or beating weakly-supervised visual grounding across five datasets.

AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video Recognition
Yulin Wang*, Yang Yue*, Yuanze Lin, Haojun Jiang, Zihang Lai, Victor Kulikov, Nikita Orlov, Humphrey Shi, Gao Huang
CVPR, 2022
ArXiv / Code / BibTeX

Reformulates AdaFocus as a one-stage algorithm via differentiable patch selection and an improved training scheme, validated on six benchmarks.

Self-supervised video representation learning with meta-contrastive network
Yuanze Lin, Xun Guo, Yan Lu
ICCV, 2021
ArXiv / Poster / BibTeX

A Meta-Contrastive Network combining contrastive and meta-learning for pre-training, surpassing prior methods on UCF101 and HMDB51.

EVA-GCN: Head Pose Estimation Based on Graph Convolutional Networks
Miao Xin, Shentong Mo, Yuanze Lin
CVPR AMFG Workshop, 2021   🏆 Best Paper Award
Paper / Code / BibTeX

Builds a landmark-connection graph and uses Graph Convolutional Networks to model nonlinear mappings from graph topology to head-pose angles.

 Experiences


Industrial
Researcher Intern, Jul 2026 - Present
hosted by Dr. Marek Kowalski, working on spatial reasoning in MLLMs.
Researcher Intern, Feb 2024 - Nov 2024
hosted by Dr. Dongdong Chen, working on MLLMs.
Senior Algorithm Engineer, Feb 2023 - Aug 2023
Working on vision-language pre-training and applications of LLMs.
Researcher Intern, Feb 2022 - June 2022
with Dr. Yujia Xie, Dr. Dongdong Chen and Dr. Yichong Xu on knowledge-based VQA.
Researcher Intern, Dec 2020 - Sep 2021
with Dr. Xun Guo and Dr. Yan Lu on self-supervised learning for video.
Academic
Visiting Student, May 2022 - Nov 2023
hosted by Prof. Ming-Hsuan Yang, working on text-driven image editing.
Research Assistant, May 2022 - Feb 2023
with Prof. Cihang Xie and Prof. Alan Yuille on MAE-based vision-language pre-training.

  Professional Services


Program Comittee: AAAI 2025, AAAI 2026

Journal Reviewer: IJCV 2025

Conference Reviewer: ICLR 2026, AISTATS 2026, CVPR 2026, ECCV 2026

Conference Reviewer: ICLR 2025, AISTATS 2025, CVPR 2025, ICML 2025, ICCV 2025, NeurIPS 2025

Conference Reviewer: ICRA 2024, CVPR 2024, ECCV 2024, NeurIPS 2024

Conference Reviewer: ICLR 2023, CVPR 2023, ICCV 2023, NeurIPS 2023

Conference Reviewer: CVPR 2022


Last Update: 07/2026        Template