News
---- show more ----
09/2025 LODGE is accepted to NeurIPS 2025 as a spotlight !
08/2025 I will serve as an Area Chair at ICLR 2026 and CVPR 2026 .
07/2025 Three papers (Visual Chronicles , CL-Splats , and SplatTalk ) are accepted to ICCV 2025 !
05/2025 Invited to give talks on A "Splatacular" Year of 3D Reconstruction at Stanford University and KAIST (as a guest lecture).
02/2025 4 papers (DepthSplat , Prompt Depth Anything , Free360 , WildGS-SLAM ) are accepted to CVPR 2025 !
02/2025 NoPoSplat is accepted to ICLR 2025 as an Oral (top 1.8%) !
12/2024 I will serve as an Area Chair at ICCV 2025 and ICML 2025 .
11/2024 I give lectures for an 11-week course GAMES003: 图形视觉科研基本素养 (How To Do Research in CV/CG) . All lecture recordings and slides are online (in Chinese).
10/2024 My PhD thesis received the ECVA PhD Award !
09/2024 WildGaussians and RENOVATE are accepted to NeurIPS 2024 !
07/2024 Segment3D is accepted to ECCV 2024! Also, I will be co-organizing two workshops: Foundation Models Creators Meet Users (FOCUS) and Open-Vocabulary 3D Scene Understanding (OpenSUN3D) .
05/2024 Career Update : I start working as a research scientist at Google DeepMind !
03/2024 Our papers NeRF On-the-go and 3D Neural Edge Reconstruction are accepted to CVPR 2024 ! Congrats on the successful master theses of my amazing students at ETH Zurich, Lei Li and Weining Ren !
03/2024 NICER-SLAM received the Best Paper Honorable Mention Award at 3DV 2024 !
03/2024 Invited to give a talk on 2D Magic in a 3D World at Imperial College London and the University of Hong Kong.
11/2023 I successfully defended my PhD ! [Thesis ][Defense Slides ]
10/2023 Our papers NICER-SLAM and FastHuman are accepted to 3DV 2024.
07/2023 I served as an Area Chair at 3DV 2024 .
07/2023 I received the Best Presentation Award at ICVSS 2023 !
07/2023 Our paper DiffDreamer is accepted to ICCV 2023! Congrats Shengqu on a successful master thesis!
07/2023 Invited to give a 90-min lecture at SGP 2023 graduate school in neural explicit-implicit representations!
03/2023 I Co-organize OpenSUN3D : 1st Workshop on Open-Vocabulary 3D Scene Understanding in ICCV 2023 .
03/2023 Our paper OpenScene from my internship at Google Research is accepted to CVPR 2023 !
10/2022 Invited to give a talk on large-scale scene reconstruction with NeRF at Stanford University (Slides ).
09/2022 Our paper MonoSDF is accepted to NeurIPS 2022 .
09/2022 Invited to give a talk on neural rendering at Adobe Research (Slides ).
06/2022 This summer I will be a research intern at Google Research .
06/2022 1st place winner in partial object recovery and 2nd place overall of SHARP Challenge ! Congratulations on my students Lei, Zhizheng, Weining, and Liudi from 3D Vision course project at ETH Zurich!
05/2022 Selected as an outstanding reviewer at CVPR 2022.
03/2022 Our paper NICE-SLAM is accepted to CVPR 2022 !
02/2022 Invited to talk about Shape As Points at Talking Papers Podcast . Great chat with Yizhak Ben-Shabat !
12/2021 Gave a talk again this year at GAMES Seminar Series on Shape As Points .
09/2021 Our Shape As Points is accepted to NeurIPS 2021 as oral presentation (top 0.6%) !
08/2021 Join Facebook Reality Labs (FRL) as a research intern this fall.
07/2021 Two papers (UNISURF and KiloNeRF ) are accepted to ICCV 2021!
06/2021 Gave a talk at GAMES Seminar Series on Towards Practical Applications of NeRF .
11/2020 : A master course project that I advised on got accepted to WACV 2021.
08/2020 : Start my 1-year stay at Autonomous Vision Group (AVG) at MPI Tübingen.
07/2020 Our paper Convolutional Occupancy Networks is accepted to ECCV 2020 as spotlight (top 5%) !
09/2019 Start PhD journey at Max Planck ETH Center for Learning Systems !
07/2019 Our paper Calibration Wizard is accepted to ICCV 2019 as oral presentation (top 4.6%) .
06/2019 : The extension of my master thesis got accepted to TPAMI 2019!
Research
(Selected |
Full List )
Your browser does not support the video tag.
Vision Banana: Image Generators are Generalist Vision
Learners
Project Co-Lead
Google DeepMind
tech report |
website
Gemini 4, Gemini 3 & Gemini 2.5
Core Contributor
Google DeepMind
Gemini 2.5 tech report |
website
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Chengzhi Liu ,
Yuzhe Yang ,
Sophia Xiao Pu ,
Yepeng Liu ,
Lin Long ,
Yichen Guo ,
Nuo Chen ,
Zhaotian Weng ,
Elena Kochkina ,
Simerjot Kaur ,
Charese Smiley ,
Xiaomo Liu ,
James Zou ,
Sheng Liu ,
Yuheng Bu ,
Songyou Peng ,
Xin Eric Wang
Conference on Neural Information Processing Systems (NeurIPS ) , 2026
paper |
project page |
code |
data
A 400-task benchmark that evaluates how multimodal agents write, maintain, retrieve, and use memory through interaction.
Your browser does not support the video tag.
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
Gene Chou ,
Charles Herrmann ,
Kyle Genova ,
Boyang Deng ,
Songyou Peng ,
Bharath Hariharan ,
Jason Y. Zhang ,
Noah Snavely ,
Philipp Henzler
Conference on Neural Information Processing Systems (NeurIPS ) , 2026 (Spotlight , top 0.95%)
paper |
project page
Step into real cities with minutes-long, spatially grounded video generation.
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
Yijia Weng ,
Zhicheng Wang ,
Songyou Peng ,
Saining Xie ,
Howard Zhou ,
Leonidas Guibas
European Conference on Computer Vision (ECCV ) , 2026 (Spotlight , top 1.3%)
paper |
project page
On-demand Gaussian densification brings high-resolution detail to just the regions you care about.
NeRFs in Robotics: A Survey
Guangming Wang ,
Lei Pan ,
Songyou Peng ,
Shaohui Liu ,
Chenfeng Xu ,
Yanzi Miao ,
Wei Zhan ,
Masayoshi Tomizuka ,
Marc Pollefeys ,
Hesheng Wang
The International Journal of Robotics Research (IJRR ) , 2026
journal |
paper
A comprehensive survey of NeRF applications, advances, and open challenges in robotics.
Your browser does not support the video tag.
Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
Youming Deng ,
Songyou Peng ,
Junyi Zhang ,
Kathryn Heal ,
Tiancheng Sun ,
John Flynn ,
Steve Marschner ,
Lucy Chai
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2026 (Oral , top 0.9%)
paper |
project page
Teach a 3D foundation model to improve itself. No ground truth needed.
Your browser does not support the video tag.
Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
Jiahao Wang ,
Bo Sun ,
Yijing Bai ,
Vincent Casser ,
Songyou Peng ,
Zehao Zhu ,
Meng-Li Shih ,
Xander Masotto ,
Shih-Yang Su ,
Kanaad V Parvate ,
Tiancheng Ge ,
Linn Bieske ,
Dragomir Anguelov ,
Mingxing Tan ,
Chiyu "Max" Jiang
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2026
paper
A prototype for the Waymo World Model , translating in-the-wild monocular videos into high-fidelity multi-modal sensor logs.
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
Xianzheng Ma *,
Tao Sun *,
Shuai Chen ,
Yash Bhalgat ,
Jindong Gu ,
Angel X Chang ,
Iro Armeni ,
Iro Laina ,
Songyou Peng †,
Victor Adrian Prisacariu †
International Conference on Learning Representations (ICLR ) , 2026
Best Paper Runner-up Award at the 3D-LLM/VLA Workshop at CVPR 2026
(* equal contribution, † equal supervision)
paper |
project page |
code |
data
Your 3D-LLM isn't understanding 3D. It might be just guessing without seeing.
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
Junhwa Hur ,
Charles Herrmann ,
Songyou Peng ,
Philipp Henzler ,
Zeyu Ma ,
Todd Zickler ,
Deqing Sun
International Conference on Learning Representations (ICLR ) , 2026
paper |
project page |
code
Feedforward 4D reconstruction from just two unposed images.
Your browser does not support the video tag.
LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering
Jonas Kulhanek ,
Marie-Julie Rakotosaona ,
Fabian Manhardt ,
Christina Tsalicoglou ,
Michael Niemeyer ,
Torsten Sattler ,
Songyou Peng ,
Federico Tombari
Conference on Neural Information Processing Systems (NeurIPS ) , 2025 (Spotlight , top 3%)
paper |
project page
City-scale 3DGS in real-time with an iPhone.
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
Boyang Deng ,
Songyou Peng* ,
Kyle Genova *,
Gordon Wetzstein ,
Noah Snavely ,
Leonidas Guibas ,
Thomas Funkhouser
International Conference on Computer Vision (ICCV ) , 2025 (Highlight , top 2.3%)
paper |
project page
We help you find "unusual" things and trends in NYC and SF, like 200+ abstract sculptures, see left for an example.
Your browser does not support the video tag.
CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization
Jan Ackermann ,
Jonas Kulhanek ,
Shengqu Cai ,
Haofei Xu ,
Marc Pollefeys ,
Gordon Wetzstein ,
Leonidas Guibas ,
Songyou Peng
International Conference on Computer Vision (ICCV ) , 2025
paper |
project page |
code
We give you great 3DGS even after you add, delete, change stuff in your room.
SplatTalk: 3D VQA with Gaussian Splatting
Anh Thai ,
Songyou Peng ,
Kyle Genova ,
Leonidas Guibas ,
Thomas Funkhouser
International Conference on Computer Vision (ICCV ) , 2025
paper |
project page
3D language Gaussian field benefits 3D VQA tasks.
Your browser does not support the video tag.
DepthSplat: Connecting Gaussian Splatting and Depth
Haofei Xu ,
Songyou Peng ,
Fangjinhua Wang ,
Hermann Blum ,
Daniel Barath ,
Andreas Geiger ,
Marc Pollefeys
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2025
paper |
project page |
code
Depths helps 3DGS, 3DGS helps depth prediction.
Your browser does not support the video tag.
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
Haotong Lin ,
Sida Peng ,
Jingxiao Chen ,
Songyou Peng
Jiaming Sun ,
Minghuan Liu ,
Hujun Bao ,
Jiashi Feng ,
Xiaowei Zhou ,
Bingyi Kang
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2025
paper |
project page |
code
4K accurate metric depth estimation from low-res LiDAR.
Your browser does not support the video tag.
Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views
Chong Bao ,
Zehao Yu ,
Jiale Shi ,
Guofeng Zhang ,
Songyou Peng ,
Zhaopeng Cui
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2025
paper |
project page |
video |
code
Video models enable unbounded 360° scene reconstruction from 3-4 unposed views.
Your browser does not support the video tag.
WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments
Jianhao Zheng *,
Zihan Zhu ,
Valentin Bieri ,
Marc Pollefeys ,
Songyou Peng ,
Iro Armeni
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2025
paper |
project page |
code
Robust SLAM for dynamic scenes in the wild.
Your browser does not support the video tag.
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
Botao Ye ,
Sifei Liu ,
Haofei Xu ,
Xueting Li ,
Marc Pollefeys ,
Ming-Hsuan Yang ,
Songyou Peng
International Conference on Learning Representations (ICLR ) , 2025 (Oral , top 1.8%)
paper |
project page |
code
Unposed 3DGS made easy, also enables SoTA relative pose estimation performance!
Your browser does not support the video tag.
WildGaussians: 3D Gaussian Splatting in the Wild
Jonas Kulhanek ,
Songyou Peng ,
Zuzana Kukelova ,
Marc Pollefeys ,
Torsten Sattler
Conference on Neural Information Processing Systems (NeurIPS ) , 2024
paper |
project page |
code
Boost 3DGS for in-the-wild scenes with appearance and dynamic changes.
Renovating Names in Open-Vocabulary Segmentation Benchmarks
Haiwen Huang ,
Songyou Peng ,
Dan Zhang ,
Andreas Geiger
Conference on Neural Information Processing Systems (NeurIPS ) , 2024
paper |
project page |
code
Wanna enhance your segmentation model or benchmark? Renovate names now!
Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels
Rui Huang ,
Songyou Peng ,
Ayça Takmaz ,
Federico Tombari ,
Marc Pollefeys ,
Shiji Song ,
Gao Huang ,
Francis Engelmann
European Conference on Computer Vision (ECCV ) , 2024
paper |
project page |
code |
demo
A self-supervised segmentation approach that outperforms fully-supervised methods.
Your browser does not support the video tag.
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
Weiyang Liu *,
Zeju Qiu *,
Yao Feng **,
Yuliang Xiu **,
Yuxuan Xue **,
Longhui Yu **,
Haiwen Feng ,
Zhen Liu ,
Juyeon Heo ,
Songyou Peng ,
Yandong Wen ,
Michael J. Black ,
Adrian Weller ,
Bernhard Schölkopf
(*/** equal contribution)
International Conference on Learning Representations (ICLR ) , 2024
paper |
project page |
code |
BOFT (Orthogonal Butterfly) is a general finetuning technique that adapts foundation models to different tasks such as Vision, NLP, Math QA, and Controllable Generation.
Your browser does not support the video tag.
FastHuman: Reconstructing High-Quality Clothed Human in Minutes
Lixiang Lin ,
Songyou Peng ,
Qijun Gan ,
Jianke Zhu
International Conference on 3D Vision (3DV ) , 2024 (Spotlight , top 8.2%)
paper |
project page |
code
Shape As Points (SAP) for fast human body reconstruction.
Neural Scene Representations for 3D Reconstruction and Scene Understanding
Songyou Peng
PhD Thesis , 2023
ECVA PhD Award, 2024
3DV Outstanding Dissertation Award Honorable Mention, 2026
thesis |
slides
PhD supervisors: Prof. Marc Pollefeys (ETH Zurich), Prof. Andreas Geiger (MPI-IS)
External committee: Prof. Leonidas J. Guibas (Stanford), Prof. Vincent Sitzmann (MIT)
Your browser does not support the video tag.
OpenScene: 3D Scene Understanding with Open Vocabularies
Songyou Peng ,
Kyle Genova ,
Chiyu "Max" Jiang ,
Andrea Tagliasacchi ,
Marc Pollefeys ,
Thomas Funkhouser
Conference on Computer Vision and Pattern Recognition (CVPR ) , 2023
paper |
project page |
video |
code
Zero-shot approach for novel 3D scene understanding tasks with open-vocabulary queries.
Your browser does not support the video tag.
: A Unified Framework for Surface Reconstruction
Zehao Yu ,
Anpei Chen ,
Bozidar Antic ,
Songyou Peng ,
Apratim Bhattacharyya ,
Michael Niemeyer ,
Siyu Tang ,
Torsten Sattler ,
Andreas Geiger
Open Source Project, 2023
project page |
code
We provide a unified framework and benchmark for neural implicit surface reconstruction.
PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship
Le Zhang , Songyou Peng , Stefan Winkler
IEEE Transactions on Affective Computing (TAFFC ), 2019. In press.
paper |
code
A journal extension of our ACM MM 2018 paper.
Give Me One Portrait Image, I Will Tell You Your Emotion and Personality
Songyou Peng , Le Zhang , Stefan Winkler , Marianne Winslett
ACM International Conference on Multimedia (ACM MM ) , 2018
paper |
slides |
code
Technical Demo. A deep Siamese-like network is introduced to predict one's Big-Five personality and arousal-valence emotion from one portrait photo.
Depth Super-Resolution Meets Uncalibrated Photometric Stereo
Songyou Peng , Bjoern Haefner , Yvain Queau , Daniel Cremers
International Conference on Computer Vision (ICCV ) Workshops , 2017
paper |
slides |
code & data
A novel depth super-resolution approach for RGB-D sensors is presented.
This paper a part of my master thesis, and subsumed by our TPAMI paper .
High Quality Shape from a RGB-D Camera using Photometric Stereo
Songyou Peng
M.Sc. Thesis , Techinical University of Munich
Supervisor: Yvain Queau and Daniel Cremers
thesis |
bibtex |
poster
Mentored Students and Interns
I am fortunate to (co-)mentor some talented and highly motivated students and interns. I have learnt from and gotten inspired by them:
Benlin Liu (2026): PhD student at University of Washington
Internship project at Google DeepMind: Improving Video Understanding via Self-Play
Gene Chou (2025-2026): PhD student at Cornell University
Internship project at Google: City-RAG (NeurIPS'26 spotlight )
Youming Deng (2025): PhD student at Cornell University
Internship project at Google: Selfi (CVPR'26 Oral )
Jiahao Wang (2025): PhD student at John Hopkins University
Zehao Yu (2025): PhD student at University of Tübingen
Internship project at Google DeepMind: 3D Reconstruction with Multimodal LLM
Jonas Kulhanek (2024-2025): PhD student at Czech Technical University
Internship project at Google: LODGE (NeurIPS'25 spotlight paper)
PhD project at ETH Zurich: WildGaussians (NeurIPS'24)
Boyang Deng (2024): PhD student at Stanford University
Anh Thai (2024): PhD student at Georgia Tech
Jan Ackermann (2024): MSc student at ETH Zurich
Semester thesis: Continual Learning of Gaussian Splatting with Local Optimization (ICCV'25)
→ Master thesis and PhD student at Stanford University, advised by Gordon Wetzstein
Gonca Yilmaz (2024): MSc student at University of Zurich
Weining Ren (2023): MSc student at ETH Zurich
Master thesis: NeRF On-the-go (CVPR'24)
→ PhD student at The University of Hong Kong (HKU), advised by Kai Han
Lei Li (2023): MSc student at ETH Zurich
Mirlan Karimov (2023): MSc student at ETH Zurich
Master thesis: Interactive Preprocessing via Multi-Modal Prompting for NeRFs
→ PhD student at Mercedes-Benz AG
Junru Lin (2023): BSc student at Univeristy of Toronto
Shengqu Cai (2022): MSc student at ETH Zurich
Zihan Zhu (2021): BSc student at Zhejiang University
Bachelor internship project: NICE-SLAM (CVPR'22)
Semester project: NICER-SLAM (3DV'24, Best Paper Honorable Mention )
Semester project: WildGS-SLAM (CVPR'25)
Master thesis: Leverage geometric constraints for better depth estimation
→ Direct doctorate student at ETH Zurich, advised by Marc Pollefeys
Pfister Severin (2021): MSc student at ETH Zurich
Master thesis: Online Implicit Reconstruction
→ Consultant at McKinsey & Company
Weirong Chen (2020): MSc student at ETH Zurich
Semester thesis: Real-time 3D Reconstruction through Neural Implicit Representation
→ PhD student at TU Munich, advised by Daniel Cremers and Andrea Vedaldi
Invited Talks
Vision Banana: Image Generators are Generalist Vision Learners
Waymo , hosted by Max Jiang , 2026
World Labs , hosted by Ben Mildenhall , 2026
3DSUN Workshop at CVPR , 2026
slides
My 10-Year Journey in 3D Vision and Finding Our Poses in the Current World
International Conference on 3D Vision (3DV) , 2026
PhD Outstanding Dissertation Award Talk
recording |
slides
Building Visual Intelligence
Meta , 2025
Amazon Frontier AI & Robotics (FAR) , 2025
slides
A "Splatacular" Year of 3D Reconstruction
Stanford University , hosted by Iro Armeni , 2025
KAIST , hosted by Minhyuk Sung , 2025 (Guest Lecture )
recording | slides
2D Magic in a 3D World
Imperial College London , hosted by Andrew Davison , 2024
Czech Technical University (CTU) , hosted by Torsten Sattler , 2024
The University of Hong Kong (HKU) , hosted by Kai Han , 2024
slides
Dive into Neural Explicit-Implicit 3D Representations and Their Applications
Symposium of Geometry Processing (SGP ) Graduate School , 2023 (Invited Lecture )
slides
Learning to Reconstruct and Understand the 3D World
Microsoft Mixed Reality & AI Labs - Zurich , 2023
slides
Learning Neural Scene Representations for 3D Reconstruction and Understanding
Shanghai AI Lab , 2023
slides
OpenScene: 3D Scene Understanding with Open Vocabularies
Peking University , hosted by Baoquan Chen , 2023
Apple , 2023
Stability.ai , 2023
slides
How do NeRF and CLIP advance 3D Scene Reconstruction and Understanding
Chinese University of Hong Kong (CUHK ) Shenzhen , 2023
Bosch Center for Artificial Intelligence (BCAI ) , 2023
slides
Large-Scale 3D Scene Reconstruction with NeRF
Stanford University , hosted by Gordon Wetzstein , 2022
slides
Towards Practical Applications of NeRF
Adobe Research , hosted by Zexiang Xu , 2022
slides
Neural Scene Representations for 3D Reconstruction
University of Basel , 2022
slides
Shape As Points: A Differentiable Poisson Solver
Talking Papers Podcast , 2022
video |
podcast
Shape As Points: A Differentiable Poisson Solver
Graphics And Mixed Environment Seminar (GAMES ) , 2021
slides |
talk (in Chinese)
Towards Practical Applications of NeRF
Graphics And Mixed Environment Seminar (GAMES ) , 2021
slides |
talk (in Chinese)
Teaching
Teaching Assistant (Lead), 3D Vision , Spring 2023
Teaching Assistant, Computer Vision , Fall 2022
Teaching Assistant (Lead), 3D Vision , Spring 2022
Teaching Assistant, Deep Learning for Computer Vision: Seminal Work , Spring 2022
Teaching Assistant, 3D Vision , Spring 2020
Teaching Assistant, Deep Learning for Computer Vision: Seminal Work , Spring 2020
Academic Services
Publicity Chair : 3DV'25
Area Chair : NeurIPS'26, ECCV'26, CVPR'26, ICLR'26, ICCV'25, ICML'25, 3DV'24
Workshop Organizer :
Conference Reviewer : CVPR, ICCV, ECCV, ICLR, NeurIPS, SIGGRAPH, SIGGRAPH Asia
Journal Reviewer : TPAMI, IJCV, CVIU