Wild3D

3D Modeling, Reconstruction, and Generation in the Wild

in conjunction with ECCV 2026, Malmö, Sweden.

Time: September 9th, 2026, 13:30 - 17:15 CEST.


Speakers Schedule Related Workshop Call For Posters Organizers

Overview

The goal of this workshop is to bring together researchers and practitioners interested in modeling, reconstructing, or generating (dynamic) 3D objects/scenes in challenging, in-the-wild settings. With recent advances in 3D learning, the widespread availability of 2D and 3D visual data, and the predominance of image/video generative models, we believe now is a pivotal moment to tackle these challenges and make 3D vision more robust, accessible, and cost-effective. By fostering communication and highlighting important work in these areas, we hope to inspire new research topics and breakthroughs. Given recent advances in video generative models and dynamics modeling, we strongly encourage contributions not only from standard 3D topics, but also from broader 4D-related directions. We also welcome contributions that challenge the necessity of explicit 3D or 4D representations.

Invited Speakers (More to come...)

Katerina Fragkiadaki

Katerina Fragkiadaki

Carnegie Mellon University

Katerina Fragkiadaki is the JPMorgan Chase Associate Professor of Computer Science in Carnegie Mellon University's Machine Learning Department. Her work connects computer vision, machine learning, language, and robotics, with an emphasis on 3D-aware perception, embodied agents, world models, and generative simulation for robot learning.
Sagie Benaim

Sagie Benaim

Hebrew University of Jerusalem & Amazon

Sagie Benaim is an Assistant Professor (Senior Lecturer) at the Hebrew University of Jerusalem and an Amazon Scholar. His research brings together computer vision, machine learning, and graphics to develop efficient methods for reconstructing, generating, and understanding dynamic 3D worlds.
Vincent Sitzmann

Vincent Sitzmann

Massachusetts Institute of Technology

Vincent Sitzmann is an Associate Professor at MIT and leads the Scene Representation Group at CSAIL. He develops world models that learn from vision and interaction, enabling machines to represent their environments, anticipate what happens next, and act autonomously across vision, graphics, and robotics.
Georgios Pavlakos

Georgios Pavlakos

The University of Texas at Austin

Georgios Pavlakos is an Assistant Professor of Computer Science at UT Austin. His research spans computer vision, machine learning, and robotics, with a focus on recovering people and human-object interactions in 3D and 4D from images and videos, including challenging in-the-wild observations.

Schedule (Tentative)

13:30 - 13:35 Opening Remarks
13:35 - 14:05 Invited Talk 1 Vincent Sitzmann (MIT)
14:05 - 14:35 Invited Talk 2 Katerina Fragkiadaki (Carnegie Mellon University)
14:35 - 14:50 Spotlight Presentation TBA
14:50 - 15:30 Poster Session
15:30 - 15:45 Spotlight Presentation TBA
15:45 - 16:15 Invited Talk 3 Sagie Benaim (Hebrew University of Jerusalem)
16:15 - 16:45 Invited Talk 4 Georgios Pavlakos (UT Austin)
16:45 - 17:15 Invited Talk 5 TBA

Call For Posters

Wild3D 2026 will not run a separate paper-review track. Instead, we will be collecting interest for poster presentations at our workshop. If you are interested in presenting your work, please fill out this Google form. We especially welcome relevant work that has already been accepted at ECCV 2026 or another peer-reviewed venue. Posters are non-archival and will not appear in proceedings. We will follow up on the status of the poster presentations in the end of August.
  • Poster Interest Form: here.
  • Poster Interest Deadline: August 25, 2026, 23:59 Anywhere on Earth.
  • Notification to Presenters: August 28, 2026.



Topics of Interest

  • Data and Modality: What type of data provides the most useful information for (dynamic) 3D modeling? Do we need explicit 3D data or is video data sufficient? What is currently lacking in this area? What datasets and benchmarks are crucial to validate the effectiveness of 3D/4D algorithms in the wild?
  • Alignment: How can we align observations that exhibit significant variations in appearance, motion (articulation), lighting, contents, and viewpoints? How can we register images or videos with little or no overlap?
  • Modeling: How can we construct accurate 3D models from sparse, noisy, incomplete, or dynamic observations?
  • Representation: What are the most suitable representations for 3D modeling and reasoning? Do we truly need explicit 3D representations, or could view synthesis and video generative models be sufficient?
  • Knowledge and Reasoning: How can we represent, learn, and encode commonsense knowledge of 3D objects and scenes -- such as part structures, articulations, and affordances -- and leverage it for various 3D tasks, including reasoning, dynamic modeling, reconstruction, and generation?
  • 4D (Dynamic 3D): What is the best way to represent and model the dynamic 3D world? What priors are critical for its success? How can we improve 3D understanding via 4D modeling?
  • Risks and ethical considerations: How can we mitigate the risks of robust 3D modeling, reconstruction, and generation techniques? How do we address relevant ethical questions, such as invasion of privacy, spreading misinformation, and credit attribution?
  • Applications: What new applications can be unlocked by developing more robust 3D algorithms, and what modifications are needed? For example, how can we adapt existing (dynamic) 3D modeling techniques to better support robots operating in challenging environments? How can we leverage 3D priors learned from images to enable photorealistic content creation? How can we build on video foundation models to enhance our 3D understanding of the world? Are there other exciting applications for in-the-wild 3D modeling for domains such as construction, agriculture, and remote sensing?