Songyou Peng (彭崧猷)

I am a Research Scientist at Google DeepMind in San Francisco, USA, where I teach Gemini and Nano Banana multimodal understanding and generation.

I received my PhD from ETH Zurich and Max Planck Institute for Intelligent Systems under the supervision of Marc Pollefeys and Andreas Geiger. After the PhD, I worked as a Senior Researcher/Postdoc at ETH Zurich for 6 months.

I was a research intern at Google Research with Tom Funkhouser, Meta Reality Labs Research with Michael Zollhoefer, Technical University of Munich with Daniel Cremers, and INRIA with Peter Sturm. I completed an Erasmus Mundus Masters in Computer Vision and Robotics (VIBOT) with distinction, and a Bachelors in Automation at Xi'an Jiaotong University.

Email  |  CV  |  GitHub  |  Google Scholar  |  LinkedIn  |  Twitter

headshot
ETH Zurich Max Planck Institute for Intelligent Systems Google Research Meta Reality Labs INRIA Technical University of Munich VIBOT Xi'an Jiaotong University

News

Research ( | )

Vision Banana: Image Generators are Generalist Vision Learners preview
Vision Banana: Image Generators are Generalist Vision Learners
Project Co-Lead
Google DeepMind
tech reportwebsite
Gemini 4, Gemini 3 & Gemini 2.5 preview
Gemini 4, Gemini 3 & Gemini 2.5 preview
Gemini 4, Gemini 3 & Gemini 2.5
Core Contributor
Google DeepMind
Gemini 2.5 tech reportwebsite
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction preview
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu, Yepeng Liu, Lin Long, Yichen Guo, Nuo Chen, Zhaotian Weng, Elena Kochkina, Simerjot Kaur, Charese Smiley, Xiaomo Liu, James Zou, Sheng Liu, Yuheng Bu, Songyou Peng, Xin Eric Wang
Conference on Neural Information Processing Systems (NeurIPS), 2026
paperproject pagecodedata

A 400-task benchmark that evaluates how multimodal agents write, maintain, retrieve, and use memory through interaction.

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
Gene Chou, Charles Herrmann, Kyle Genova, Boyang Deng, Songyou Peng, Bharath Hariharan, Jason Y. Zhang, Noah Snavely, Philipp Henzler
Conference on Neural Information Processing Systems (NeurIPS), 2026 (Spotlight, top 0.95%)
paperproject page

Step into real cities with minutes-long, spatially grounded video generation.

GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification preview
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
Yijia Weng, Zhicheng Wang, Songyou Peng, Saining Xie, Howard Zhou, Leonidas Guibas
European Conference on Computer Vision (ECCV), 2026 (Spotlight, top 1.3%)
paperproject page

On-demand Gaussian densification brings high-resolution detail to just the regions you care about.

NeRFs in Robotics: A Survey preview
NeRFs in Robotics: A Survey
Guangming Wang, Lei Pan, Songyou Peng, Shaohui Liu, Chenfeng Xu, Yanzi Miao, Wei Zhan, Masayoshi Tomizuka, Marc Pollefeys, Hesheng Wang
The International Journal of Robotics Research (IJRR), 2026
journal | paper

A comprehensive survey of NeRF applications, advances, and open challenges in robotics.

Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment preview
Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
Youming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal, Tiancheng Sun, John Flynn, Steve Marschner, Lucy Chai
Conference on Computer Vision and Pattern Recognition (CVPR), 2026 (Oral, top 0.9%)
paperproject page

Teach a 3D foundation model to improve itself. No ground truth needed.

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving preview
Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
Jiahao Wang, Bo Sun, Yijing Bai, Vincent Casser, Songyou Peng, Zehao Zhu, Meng-Li Shih, Xander Masotto, Shih-Yang Su, Kanaad V Parvate, Tiancheng Ge, Linn Bieske, Dragomir Anguelov, Mingxing Tan, Chiyu "Max" Jiang
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
paper

A prototype for the Waymo World Model, translating in-the-wild monocular videos into high-fidelity multi-modal sensor logs.

Do 3D Large Language Models Really Understand 3D Spatial Relationships? preview
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
Xianzheng Ma*, Tao Sun*, Shuai Chen, Yash Bhalgat, Jindong Gu, Angel X Chang, Iro Armeni, Iro Laina, Songyou Peng†, Victor Adrian Prisacariu†
International Conference on Learning Representations (ICLR), 2026
Best Paper Runner-up Award at the 3D-LLM/VLA Workshop at CVPR 2026
(* equal contribution, † equal supervision)
paperproject pagecodedata

Your 3D-LLM isn't understanding 3D. It might be just guessing without seeing.

UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images preview
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images preview
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
Junhwa Hur, Charles Herrmann, Songyou Peng, Philipp Henzler, Zeyu Ma, Todd Zickler, Deqing Sun
International Conference on Learning Representations (ICLR), 2026
paperproject pagecode

Feedforward 4D reconstruction from just two unposed images.

LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering preview
LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering
Jonas Kulhanek, Marie-Julie Rakotosaona, Fabian Manhardt, Christina Tsalicoglou, Michael Niemeyer, Torsten Sattler, Songyou Peng, Federico Tombari
Conference on Neural Information Processing Systems (NeurIPS), 2025 (Spotlight, top 3%)
paperproject page

City-scale 3DGS in real-time with an iPhone.

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images preview
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images preview
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
Boyang Deng, Songyou Peng*, Kyle Genova*, Gordon Wetzstein, Noah Snavely, Leonidas Guibas, Thomas Funkhouser
International Conference on Computer Vision (ICCV), 2025 (Highlight, top 2.3%)
paperproject page

We help you find "unusual" things and trends in NYC and SF, like 200+ abstract sculptures, see left for an example.

CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization preview
CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization
Jan Ackermann, Jonas Kulhanek, Shengqu Cai, Haofei Xu, Marc Pollefeys, Gordon Wetzstein, Leonidas Guibas, Songyou Peng
International Conference on Computer Vision (ICCV), 2025
paperproject pagecode

We give you great 3DGS even after you add, delete, change stuff in your room.

SplatTalk: 3D VQA with Gaussian Splatting preview
SplatTalk: 3D VQA with Gaussian Splatting preview
SplatTalk: 3D VQA with Gaussian Splatting
Anh Thai, Songyou Peng, Kyle Genova, Leonidas Guibas, Thomas Funkhouser
International Conference on Computer Vision (ICCV), 2025
paperproject page

3D language Gaussian field benefits 3D VQA tasks.

DepthSplat: Connecting Gaussian Splatting and Depth preview
DepthSplat: Connecting Gaussian Splatting and Depth
Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, Marc Pollefeys
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
paperproject pagecode

Depths helps 3DGS, 3DGS helps depth prediction.

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation preview
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
Haotong Lin, Sida Peng, Jingxiao Chen, Songyou Peng Jiaming Sun, Minghuan Liu, Hujun Bao, Jiashi Feng, Xiaowei Zhou, Bingyi Kang
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
paperproject pagecode

4K accurate metric depth estimation from low-res LiDAR.

Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views preview
Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views
Chong Bao, Zehao Yu, Jiale Shi, Guofeng Zhang, Songyou Peng, Zhaopeng Cui
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
paperproject pagevideocode

Video models enable unbounded 360° scene reconstruction from 3-4 unposed views.

WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments preview
WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments
Jianhao Zheng*, Zihan Zhu, Valentin Bieri, Marc Pollefeys, Songyou Peng, Iro Armeni
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
paperproject pagecode

Robust SLAM for dynamic scenes in the wild.

No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images preview
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, Songyou Peng
International Conference on Learning Representations (ICLR), 2025 (Oral, top 1.8%)
paperproject pagecode

Unposed 3DGS made easy, also enables SoTA relative pose estimation performance!

WildGaussians: 3D Gaussian Splatting in the Wild preview
WildGaussians: 3D Gaussian Splatting in the Wild
Jonas Kulhanek, Songyou Peng, Zuzana Kukelova, Marc Pollefeys, Torsten Sattler
Conference on Neural Information Processing Systems (NeurIPS), 2024
paperproject pagecode

Boost 3DGS for in-the-wild scenes with appearance and dynamic changes.

Renovating Names in Open-Vocabulary Segmentation Benchmarks preview
Renovating Names in Open-Vocabulary Segmentation Benchmarks preview
Renovating Names in Open-Vocabulary Segmentation Benchmarks
Haiwen Huang, Songyou Peng, Dan Zhang, Andreas Geiger
Conference on Neural Information Processing Systems (NeurIPS), 2024
paperproject pagecode

Wanna enhance your segmentation model or benchmark? Renovate names now!

Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels preview
Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels preview
Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels
Rui Huang, Songyou Peng, Ayça Takmaz, Federico Tombari, Marc Pollefeys, Shiji Song, Gao Huang, Francis Engelmann
European Conference on Computer Vision (ECCV), 2024
paperproject pagecodedemo

A self-supervised segmentation approach that outperforms fully-supervised methods.

NeRF On-the-go : Exploiting Uncertainty for Distractor-free NeRFs in the Wild preview
NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild
Weining Ren*, Zihan Zhu*, Boyang Sun, Jiaqi Chen, Marc Pollefeys, Songyou Peng
Conference on Computer Vision and Pattern Recognition (CVPR), 2024
(* equal contribution)
paperproject pagevideocode

We enable robust novel view synthesis from casually captured in-the-wild images.
Master thesis project.

3D Neural Edge Reconstruction preview
3D Neural Edge Reconstruction
Lei Li, Songyou Peng, Zehao Yu, Shaohui Liu, Rémi Pautrat, Xiaochuan Yin, Marc Pollefeys
Conference on Computer Vision and Pattern Recognition (CVPR), 2024
paperproject pagevideocode

The straight line belongs to men, the curved one to God.-- Antonio Gaudi
Master thesis project.

Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization preview

Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
Weiyang Liu*, Zeju Qiu*, Yao Feng**, Yuliang Xiu**, Yuxuan Xue**, Longhui Yu**, Haiwen Feng, Zhen Liu, Juyeon Heo, Songyou Peng, Yandong Wen, Michael J. Black, Adrian Weller, Bernhard Schölkopf
(*/** equal contribution)
International Conference on Learning Representations (ICLR), 2024
paperproject pagecode

BOFT (Orthogonal Butterfly) is a general finetuning technique that adapts foundation models to different tasks such as Vision, NLP, Math QA, and Controllable Generation.

NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM preview
NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM
Zihan Zhu*, Songyou Peng*, Viktor Larsson, Zhaopeng Cui, Martin R. Oswald, Andreas Geiger, Marc Pollefeys
International Conference on 3D Vision (3DV), 2024 (Oral, Best Paper Honorable Mention)
(* equal contribution)
paperproject pagevideocode

RGB-only version of our NICE-SLAM, making it NICER.

FastHuman: Reconstructing High-Quality Clothed Human in Minutes preview

FastHuman: Reconstructing High-Quality Clothed Human in Minutes
Lixiang Lin, Songyou Peng, Qijun Gan, Jianke Zhu
International Conference on 3D Vision (3DV), 2024 (Spotlight, top 8.2%)
paperproject pagecode

Shape As Points (SAP) for fast human body reconstruction.

Neural Scene Representations for 3D Reconstruction and Scene Understanding preview
Neural Scene Representations for 3D Reconstruction and Scene Understanding preview
Neural Scene Representations for 3D Reconstruction and Scene Understanding
Songyou Peng
PhD Thesis, 2023
ECVA PhD Award, 2024
3DV Outstanding Dissertation Award Honorable Mention, 2026
thesisslides

PhD supervisors: Prof. Marc Pollefeys (ETH Zurich), Prof. Andreas Geiger (MPI-IS)
External committee: Prof. Leonidas J. Guibas (Stanford), Prof. Vincent Sitzmann (MIT)

DiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion Models preview

DiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion Models
Shengqu Cai, Eric R. Chan, Songyou Peng, Mohamad Shahbazi, Anton Obukhov, Luc Van Gool, Gordon Wetzstein
International Conference on Computer Vision (ICCV), 2023
paperproject page

A diffusion-model based unsupervised framework capable of synthesizing novel views depicting a long camera trajectory.

OpenScene: 3D Scene Understanding with Open Vocabularies preview
OpenScene: 3D Scene Understanding with Open Vocabularies
Songyou Peng, Kyle Genova, Chiyu "Max" Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser
Conference on Computer Vision and Pattern Recognition (CVPR), 2023
paperproject pagevideocode

Zero-shot approach for novel 3D scene understanding tasks with open-vocabulary queries.

: A Unified Framework for Surface Reconstruction preview
: A Unified Framework for Surface Reconstruction
Zehao Yu, Anpei Chen, Bozidar Antic, Songyou Peng, Apratim Bhattacharyya, Michael Niemeyer, Siyu Tang, Torsten Sattler, Andreas Geiger
Open Source Project, 2023
project pagecode

We provide a unified framework and benchmark for neural implicit surface reconstruction.

MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction preview
MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction
Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler, Andreas Geiger
Conference on Neural Information Processing Systems (NeurIPS), 2022
paperproject page

Monocular depth and normal cues significantly boost the performance of neural implicit surface reconstruction methods.

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM preview
NICE-SLAM: Neural Implicit Scalable Encoding for SLAM
Zihan Zhu*, Songyou Peng*, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R. Oswald, Marc Pollefeys
Conference on Computer Vision and Pattern Recognition (CVPR), 2022
(* equal contribution)
paperproject pagevideocode

A neural implicit-based RGB-D SLAM that can be applied to large-scale scenes.

Shape As Points: A Differentiable Poisson Solver preview
Shape As Points: A Differentiable Poisson Solver
Songyou Peng, Chiyu "Max" Jiang, Yiyi Liao, Michael Niemeyer, Marc Pollefeys, Andreas Geiger
Conference on Neural Information Processing Systems (NeurIPS), 2021 (Oral, top 0.6%)
paperproject pagevideo (6 min)video (12 min)podcastcode

An interpretable hybird shape representation that yields HQ watertight meshes at low inference times.

UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction preview
UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction preview
UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction
Michael Oechsle, Songyou Peng, Andreas Geiger
International Conference on Computer Vision (ICCV), 2021 (Oral, top 3%)
paperproject pagevideoteaser videocode

Our method enables to reconstruct accurate surfaces without input masks.

KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs preview
KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs
Christian Reiser, Songyou Peng, Yiyi Liao, Andreas Geiger
International Conference on Computer Vision (ICCV), 2021
paperproject pageblogvideoteaser videocode

Over 2000x speed-ups for NeRF are possible by utilizing thousands of tiny MLPs.

Dynamic Plane Convolutional Occupancy Networks preview
Dynamic Plane Convolutional Occupancy Networks preview
Dynamic Plane Convolutional Occupancy Networks
Stefan Lionar*, Daniil Emtsev*, Dusan Svilarkovic*, Songyou Peng
Winter Conference on Applications of Computer Vision (WACV), 2021
(* equal contribution)
papervideocode

A student project of 3D Vision course at ETH Zurich where I served as the advisor.

Convolutional Occupancy Networks preview
Convolutional Occupancy Networks preview
Convolutional Occupancy Networks
Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, Andreas Geiger
European Conference on Computer Vision (ECCV), 2020 (Spotlight, top 5%)
paperproject pageblogvideoteaser videocode
Most influential ECCV'20 papers #13

A flexible implicit representation for accurate large-scale 3D reconstruction.

DIST: Rendering Deep Implicit Signed Distance Function with Differentiable Sphere Tracing preview
DIST: Rendering Deep Implicit Signed Distance Function with Differentiable Sphere Tracing preview
DIST: Rendering Deep Implicit Signed Distance Function with Differentiable Sphere Tracing
Shaohui Liu, Yinda Zhang, Songyou Peng, Boxin Shi, Marc Pollefeys, Zhaopeng Cui
Conference on Computer Vision and Pattern Recognition (CVPR), 2020
paperproject pageteaser videopostercode

A differentiable renderer for deep implicit signed distance functions.

Calibration Wizard: A Guidance System for Camera Calibration Based on Modelling Geometric and Corner Uncertainty preview
Calibration Wizard: A Guidance System for Camera Calibration Based on Modelling Geometric and Corner Uncertainty preview
Calibration Wizard: A Guidance System for Camera Calibration Based on Modelling Geometric and Corner Uncertainty
Songyou Peng and Peter Sturm
International Conference on Computer Vision (ICCV), 2019 (Oral, top 4.6%)
papervideopostercode

A novel system that interactively guides a user to take optimal calibration images.

Photometric Depth Super-Resolution preview
Photometric Depth Super-Resolution preview
Photometric Depth Super-Resolution
Bjoern Haefner*, Songyou Peng*, Alok Verma*, Yvain Queau, Daniel Cremers
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2019
(* equal contribution)
paperproject page

Recover high-resolution depth maps with fine geometric details using photometric techniques.

PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship preview
PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship preview
PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship
Le Zhang, Songyou Peng, Stefan Winkler
IEEE Transactions on Affective Computing (TAFFC), 2019. In press.
papercode

A journal extension of our ACM MM 2018 paper.

Give Me One Portrait Image, I Will Tell You Your Emotion and Personality preview
Give Me One Portrait Image, I Will Tell You Your Emotion and Personality preview
Give Me One Portrait Image, I Will Tell You Your Emotion and Personality
Songyou Peng, Le Zhang, Stefan Winkler, Marianne Winslett
ACM International Conference on Multimedia (ACM MM), 2018
paperslidescode

Technical Demo. A deep Siamese-like network is introduced to predict one's Big-Five personality and arousal-valence emotion from one portrait photo.

Depth Super-Resolution Meets Uncalibrated Photometric Stereo preview
Depth Super-Resolution Meets Uncalibrated Photometric Stereo preview
Depth Super-Resolution Meets Uncalibrated Photometric Stereo
Songyou Peng, Bjoern Haefner, Yvain Queau, Daniel Cremers
International Conference on Computer Vision (ICCV) Workshops, 2017
paperslidescode & data

A novel depth super-resolution approach for RGB-D sensors is presented.

This paper a part of my master thesis, and subsumed by our TPAMI paper.

High Quality Shape from a RGB-D Camera using Photometric Stereo preview
High Quality Shape from a RGB-D Camera using Photometric Stereo preview

High Quality Shape from a RGB-D Camera using Photometric Stereo
Songyou Peng
M.Sc. Thesis, Techinical University of Munich
Supervisor: Yvain Queau and Daniel Cremers
thesisbibtexposter

Mentored Students and Interns

I am fortunate to (co-)mentor some talented and highly motivated students and interns. I have learnt from and gotten inspired by them:
  • Benlin Liu (2026): PhD student at University of Washington
    • Internship project at Google DeepMind: Improving Video Understanding via Self-Play

  • Gene Chou (2025-2026): PhD student at Cornell University
    • Internship project at Google: City-RAG (NeurIPS'26 spotlight)

  • Youming Deng (2025): PhD student at Cornell University
    • Internship project at Google: Selfi (CVPR'26 Oral)

  • Jiahao Wang (2025): PhD student at John Hopkins University

  • Zehao Yu (2025): PhD student at University of Tübingen
    • Internship project at Google DeepMind: 3D Reconstruction with Multimodal LLM

  • Jonas Kulhanek (2024-2025): PhD student at Czech Technical University
    • Internship project at Google: LODGE (NeurIPS'25 spotlight paper)
    • PhD project at ETH Zurich: WildGaussians (NeurIPS'24)

  • Boyang Deng (2024): PhD student at Stanford University

  • Anh Thai (2024): PhD student at Georgia Tech

  • Jan Ackermann (2024): MSc student at ETH Zurich
    • Semester thesis: Continual Learning of Gaussian Splatting with Local Optimization (ICCV'25)
    • → Master thesis and PhD student at Stanford University, advised by Gordon Wetzstein

  • Gonca Yilmaz (2024): MSc student at University of Zurich

  • Weining Ren (2023): MSc student at ETH Zurich
    • Master thesis: NeRF On-the-go (CVPR'24)
    • → PhD student at The University of Hong Kong (HKU), advised by Kai Han

  • Lei Li (2023): MSc student at ETH Zurich

  • Mirlan Karimov (2023): MSc student at ETH Zurich
    • Master thesis: Interactive Preprocessing via Multi-Modal Prompting for NeRFs
    • → PhD student at Mercedes-Benz AG

  • Junru Lin (2023): BSc student at Univeristy of Toronto

  • Shengqu Cai (2022): MSc student at ETH Zurich

  • Zihan Zhu (2021): BSc student at Zhejiang University
    • Bachelor internship project: NICE-SLAM (CVPR'22)
    • Semester project: NICER-SLAM (3DV'24, Best Paper Honorable Mention)
    • Semester project: WildGS-SLAM (CVPR'25)
    • Master thesis: Leverage geometric constraints for better depth estimation
    • → Direct doctorate student at ETH Zurich, advised by Marc Pollefeys

  • Pfister Severin (2021): MSc student at ETH Zurich
    • Master thesis: Online Implicit Reconstruction
    • → Consultant at McKinsey & Company

  • Weirong Chen (2020): MSc student at ETH Zurich
    • Semester thesis: Real-time 3D Reconstruction through Neural Implicit Representation
    • → PhD student at TU Munich, advised by Daniel Cremers and Andrea Vedaldi

Invited Talks

Vision Banana: Image Generators are Generalist Vision Learners preview
Vision Banana: Image Generators are Generalist Vision Learners preview
Vision Banana: Image Generators are Generalist Vision Learners
Waymo, hosted by Max Jiang, 2026
World Labs, hosted by Ben Mildenhall, 2026
3DSUN Workshop at CVPR, 2026
slides
My 10-Year Journey in 3D Vision preview
My 10-Year Journey in 3D Vision preview
My 10-Year Journey in 3D Vision and Finding Our Poses in the Current World
International Conference on 3D Vision (3DV), 2026
PhD Outstanding Dissertation Award Talk
recordingslides
Building Visual Intelligence preview
Building Visual Intelligence preview
Building Visual Intelligence
Meta, 2025
Amazon Frontier AI & Robotics (FAR), 2025
slides
A "Splatacular" Year of 3D Reconstruction preview
A "Splatacular" Year of 3D Reconstruction preview
A "Splatacular" Year of 3D Reconstruction
Stanford University, hosted by Iro Armeni, 2025
KAIST, hosted by Minhyuk Sung, 2025 (Guest Lecture)
recordingslides
2D Magic in a 3D World preview
2D Magic in a 3D World preview
2D Magic in a 3D World
Imperial College London, hosted by Andrew Davison, 2024
Czech Technical University (CTU), hosted by Torsten Sattler, 2024
The University of Hong Kong (HKU), hosted by Kai Han, 2024
slides
Dive into Neural Explicit-Implicit 3D Representations and Their Applications preview
Dive into Neural Explicit-Implicit 3D Representations and Their Applications preview
Dive into Neural Explicit-Implicit 3D Representations and Their Applications
Symposium of Geometry Processing (SGP) Graduate School, 2023 (Invited Lecture)
slides
Learning to Reconstruct and Understand the 3D World preview
Learning to Reconstruct and Understand the 3D World preview
Learning to Reconstruct and Understand the 3D World
Microsoft Mixed Reality & AI Labs - Zurich, 2023
slides
Learning Neural Scene Representations for 3D Reconstruction and Understanding preview
Learning Neural Scene Representations for 3D Reconstruction and Understanding preview
Learning Neural Scene Representations for 3D Reconstruction and Understanding
Shanghai AI Lab, 2023
slides
OpenScene: 3D Scene Understanding with Open Vocabularies preview
OpenScene: 3D Scene Understanding with Open Vocabularies preview
OpenScene: 3D Scene Understanding with Open Vocabularies
Peking University, hosted by Baoquan Chen, 2023
Apple, 2023
Stability.ai, 2023
slides
How do NeRF and CLIP advance 3D Scene Reconstruction and Understanding preview
How do NeRF and CLIP advance 3D Scene Reconstruction and Understanding preview
How do NeRF and CLIP advance 3D Scene Reconstruction and Understanding
Chinese University of Hong Kong (CUHK) Shenzhen, 2023
Bosch Center for Artificial Intelligence (BCAI), 2023
slides
Large-Scale 3D Scene Reconstruction with NeRF preview
Large-Scale 3D Scene Reconstruction with NeRF preview
Large-Scale 3D Scene Reconstruction with NeRF
Stanford University, hosted by Gordon Wetzstein, 2022
slides
Towards Practical Applications of NeRF preview
Towards Practical Applications of NeRF preview
Towards Practical Applications of NeRF
Adobe Research, hosted by Zexiang Xu, 2022
slides
Neural Scene Representations for 3D Reconstruction preview
Neural Scene Representations for 3D Reconstruction preview
Neural Scene Representations for 3D Reconstruction
University of Basel, 2022
slides
Shape As Points: A Differentiable Poisson Solver preview
Shape As Points: A Differentiable Poisson Solver preview
Shape As Points: A Differentiable Poisson Solver
Talking Papers Podcast, 2022
videopodcast
Shape As Points: A Differentiable Poisson Solver preview
Shape As Points: A Differentiable Poisson Solver preview
Shape As Points: A Differentiable Poisson Solver
Graphics And Mixed Environment Seminar (GAMES), 2021
slidestalk (in Chinese)
Towards Practical Applications of NeRF preview
Towards Practical Applications of NeRF preview
Towards Practical Applications of NeRF
Graphics And Mixed Environment Seminar (GAMES), 2021
slidestalk (in Chinese)

Teaching

GAMES003 course Lecturer, GAMES003: 图形视觉科研基本素养 (How To Do Research in CV/CG), Fall 2024

Together with Sida Peng, Jun Gao, and Qianqian Wang.
ETH Zurich Teaching Assistant (Lead), 3D Vision, Spring 2023
Teaching Assistant, Computer Vision, Fall 2022
Teaching Assistant (Lead), 3D Vision, Spring 2022
Teaching Assistant, Deep Learning for Computer Vision: Seminal Work, Spring 2022
Teaching Assistant, 3D Vision, Spring 2020
Teaching Assistant, Deep Learning for Computer Vision: Seminal Work, Spring 2020

University of Tübingen Teaching Assistant, Deep Learning, Winter 2020/2021

Academic Services

  • Publicity Chair: 3DV'25
  • Area Chair: NeurIPS'26, ECCV'26, CVPR'26, ICLR'26, ICCV'25, ICML'25, 3DV'24
  • Workshop Organizer:
  • Conference Reviewer: CVPR, ICCV, ECCV, ICLR, NeurIPS, SIGGRAPH, SIGGRAPH Asia
  • Journal Reviewer: TPAMI, IJCV, CVIU

template adapted from this awesome website