Computer Graphics
TU Braunschweig

Seminar Computer Vision WS'26/27
Seminar

Prof. Dr.-Ing. Martin Eisemann

Hörerkreis: Bachelor & Master
Kontakt: sekretariat@cg.cs.tu-bs.de

Modul: INF-STD-66, INF-STD-68
Vst.Nr.: 4216031, 4216032

Topic: Recent research in Visual Computing

Latest News

Content

In this seminar we discuss current research results in computer vision, visual computing and image/video processing. The task of the participants is to understand and explain a certain research topic to the other participants. In a block seminar in the middle of the semester the background knowledge required for the final talks will be presented in oral presentations and at the end of the semester, the respective research topic is presented in an oral presentation. This must also be rehearsed beforehand in front of another student and his/her suggestions for improvement must be integrated.

Participants

The course is aimed at bachelor's and master's students from the fields of computer science (Informatik), IST, business informatics (Wirtschaftsinformatik), and data science.

Registration takes place centrally via StudIP. The number of participants is initially limited to 8 students, but can be extended in the kickoff if necessary.

Important Dates

All dates listed here must be adhered to. Attendance at all events is mandatory.

Events in person
Submission deadlines / action required

  • 09.07.2026 12:00 - 14.07.2026 12:00: Registration via Stud.IP
  • 27.10.2026, 10:30-12:00, (G30, ICG): Kickoff Meeting 
  • 02.11.2026: End of the deregistration period
  • tba, 10:30-12:00, G30 (ICG): Gather topics for fundamentals talk
  • tba: Submission of presentation slides for fundamentals talk (please use the following naming scheme: Lastname_FundamentalsPresentation_SeminarCV.pdf)
  • tba, 09:00 - 12:00, G30 (ICG): Fundamentals presentations, Block
  • Till 14.01.2026: Trial presentation for final presentation (between tandem partners from fundamentals talk)
  • 27.01.2027: Submission of presentation slides for final talk (ALL participants!) (please use the following naming scheme: Lastname_FinalPresentation_SeminarCV.pdf)
  • 28.01.2027, 09:00 - 15:00, G30 (ICG): Presentations - Block Event Part 1 (canceled)
  • 29.01.2027, 09:00 - 15:00, G30 (ICG): Presentations - Block Event Part 2 

Registered students have the option to deregister up to two weeks after the official start of lectures for this semester have started. For a successful deregistration it is necessary to deregister with the seminar supervisor.

The respective drop-offs are done by email to , and your advisor, and if necessary by email to the tandem partner. Unless otherwise communicated, submissions must be made by 11:59pm on the submission day.

If you would like to be provided with a presentation notebook for your talk, please let us know and send your presentation directly or via download link (TU-Cloud) in PPTX or PDF format at least 3 days in advance to .

If you have any questions about the event, please contact .

Format

  • The topics for the final talks will be distributed amongst the participants during the Kickoff event.
  • The topics for the fundamentals talks will be distributed amongst the participants during the second meeting.
  • The topics will be presented in approximately 20 minute presentations followed by a discussion, see important dates.
  • For the on-site lectures, a laptop of the institute or an own laptop can be used. If the institute laptop is to be used, it is necessary to contact seminarcv@cg.tu-bs.de in time, at least two weeks before the presentations. In this case, the presentation slides must be made available at least one week before the lecture.
  • The presentations will be given on site. If, for some reason, the presentations take place online, Big Blue Button will be used as a platform. In this case, students need their own PC with microphone. In addition, a video transmission during the own lecture would be desirable. If these requirements cannot be met, it is necessary to contact seminarcv@cg.cs.tu-bs.de in time.
  • The language for the presentations can be either German or English.
  • The presentations are mandatory requirements to pass the course successfully.

Files and Templates

    Topics - Bachelor Level

    1. Dual Photography
      Sen, P., Chen, B., Garg, G., Marschner, S. R., Horowitz, M., Levoy, M., & Lensch, H. P. (2005). In ACM SIGGRAPH 2005 Papers (pp. 745-755).
      [ paper ]

      We present a novel photographic technique called dual photography, which exploits Helmholtz reciprocity to interchange the lights and cameras in a scene. With a video projector providing structured illumination, reciprocity permits us to generate pictures from the viewpoint of the projector, even though no camera was present at that location. The technique is completely image-based, requiring no knowledge of scene geometry or surface properties, and by its nature automatically includes all transport paths, including shadows, interreflections and caustics. In its simplest form, the technique can be used to take photographs without a camera; we demonstrate this by capturing a photograph using a projector and a photo-resistor. If the photo-resistor is replaced by a camera, we can produce a 4D dataset that allows for relighting with 2D incident illumination. Using an array of cameras we can produce a 6D slice of the 8D reflectance field that allows for relighting with arbitrary light fields. Since an array of cameras can operate in parallel without interference, whereas an array of light sources cannot, dual photography is fundamentally a more efficient way to capture such a 6D dataset than a system based on multiple projectors and one camera. As an example, we show how dual photography can be used to capture and relight scenes.

      Advisor: Fabian Friederichs

    2. Acquiring the Reflectance Field of a Human Face
      Debevec, P., Hawkins, T., Tchou, C., Duiker, H. P., Sarokin, W., & Sagar, M. (2000, July). In Proceedings of the 27th annual conference on Computer graphics and interactive techniques (pp. 145-156).
      [ paper | project page ]

      We present a method to acquire the reflectance field of a human face and use these to render the face under arbitrary changes in lighting and viewpoint. We first  images of the face from a small set of viewpoints under a dense sampling of in cident illumination directions using a light stage. We then construct a reflectance  image for each observed image pixel from its values over the space of illumination directions. From the reflectance functions, we can directly generate images of the face from the original viewpoints in any form of sampled or computed illumination. To change the viewpoint, we use a model of skin reflectance to estimate the  of the reflectance functions for novel viewpoints. We demonstrate the technique with  renderings of a person’s face under novel illumination and viewpoints.

      Advisor: Fabian Friederichs

    3. Fast separation of direct and global components of a scene using high frequency illumination
      Nayar, S. K., Krishnan, G., Grossberg, M. D., & Raskar, R. (2006). In ACM SIGGRAPH 2006 Papers (pp. 935-944).
      [ paper ]

      We present fast methods for separating the direct and global illumination components of a scene measured by a camera and illuminated by a light source. In theory, the separation can be done with just two images taken with a high frequency binary illumination pattern and its complement. In practice, a larger number of images are used to overcome the optical and resolution limitations of the camera and the
      source. The approach does not require the material properties of objects and media in the scene to be known. However, we require that the illumination frequency is high enough to adequately sample the global components received by scene points. We present separation results for scenes that include complex interreflections, subsurface
      scattering and volumetric scattering. Several variants of the separation approach are also described. When a sinusoidal illumination pattern is used with different phase shifts, the separation can be done using just three images. When the computed images are of lower resolution than the source and the camera, smoothness constraints are used to perform the separation using a single image. Finally, in
      the case of a static scene that is lit by a simple point source, such as the sun, a moving occluder and a video camera can be used to do the separation. We also show several simple examples of how novel images of a scene can be computed from the separation results.

      Advisor: Fabian Friederichs

    4. Shape-from-shading: a survey
      Zhang, R., Tsai, P. S., Cryer, J. E., & Shah, M. (2002). IEEE transactions on pattern analysis and machine intelligence, 21 (8), 690-706.
      [ paper ]

      Since the first shape-from-shading (SFS) technique was developed by Horn in the early 1970s, many different approaches have emerged. In this paper, six well-known SFS  are implemented and compared. The performance of the algorithms was analyzed on synthetic images using mean and standard deviation of depth (Z) error, mean of gradient (p, q) error, and CPU timing. Each algorithm works well for certain images, but performs poorly for others. In general, minimization approaches are more robust, while the other approaches are faster. The implementation of these algorithms in C and images used in this paper are available by anonymous ftp under the /tech_paper/survey directory at eustis.cs.ucf.edu (132.170.108.42). These are also part of the electronic version of paper.

      Advisor: Fabian Friederichs

    5. Computational Parquetry: Fabricated Style Transfer with Wood Pixels
      Iseringhausen, J., Weinmann, M., Huang, W., & Hullin, M. B. (2020). ACM TOG
      [ paper | project page ]

      Parquetry is the art and craft of decorating a surface with a pattern of differently colored veneers of wood, stone, or other materials. Traditionally, the process of designing and making parquetry has been driven by color, using the texture found in real wood only for stylization or as a decorative effect. Here, we introduce a computational pipeline that draws from the rich natural structure of strongly textured real-world veneers as a source of detail to approximate a target image as faithfully as possible using a manageable number of parts. This challenge is closely related to the established problems of patch-based image synthesis and stylization in some ways, but fundamentally different in others. Most importantly, the limited availability of resources (any piece of wood can only be used once) turns the relatively simple problem of finding the right piece for the target location into the combinatorial problem of finding optimal parts while avoiding resource collisions. We introduce an algorithm that efficiently solves an approximation to the problem. It further addresses challenges like gamut mapping, feature characterization, and the search for fabricable cuts. We demonstrate the effectiveness of the system by fabricating a selection of pieces of parquetry from different kinds of unstained wood veneer.

      Advisor: Fabian Friederichs

    6. Stable fluids
      Stam, J. (1999). SIGGRAPH
      [ paper ]

      Building animation tools for fluid-like motions is an important and challenging problem with many applications in computer graphics. The use of physics-based models for fluid flow can greatly assist in creating such tools. Physical models, unlike key frame or procedural based techniques, permit an animator to almost effortlessly create interesting, swirling fluid-like behaviors. Also, the interaction of flows with objects and virtual forces is handled elegantly. Until recently, it was believed that physical fluid models were too expensive to allow real-time interaction. This was largely due to the fact that previous models used unstable schemes to solve the physical equations governing a fluid. In this paper, for the first time, we propose an unconditionally stable model which still produces complex fluid-like flows. As well, our method is very easy to implement. The stability of our model allows us to take larger time steps and therefore achieve faster simulations. We have used our model in conjuction with advecting solid textures to create many fluid-like animations interactively in two- and three-dimensions.

      Advisor: Jannis Möller

    7. Denoising Diffusion Probabilistic Models
      Ho, J., Jain, A., & Abbeel, P. (2020). NeurIPS
      [ paper | project page ]

      We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN.

      Advisor: Jannis Möller

    8. U-net: Convolutional networks for biomedical image segmentation
      Ronneberger, O., Fischer, P., & Brox, T. (2015). MICCAI
      [ paper | project page ]

      There is large consent that successful training of deep networks requires many thousand annotated training samples. In this paper, we present a network and training strategy that relies on the strong use of data augmentation to use the available annotated samples more efficiently. The architecture consists of a contracting path to capture context and a symmetric expanding path that enables precise localization. We show that such a network can be trained end-to-end from very few images and outperforms the prior best method (a sliding-window convolutional network) on the ISBI challenge for segmentation of neuronal structures in electron microscopic stacks. Using the same network trained on transmitted light microscopy images (phase contrast and DIC) we won the ISBI cell tracking challenge 2015 in these categories by a large margin. Moreover, the network is fast. Segmentation of a 512x512 image takes less than a second on a recent GPU.

      Advisor: Jannis Möller

    Topics - Master Level

    1. Neural Importance Sampling
      Müller, T., McWilliams, B., Rousselle, F., Gross, M., & Novák, J. (2019). ACM Transactions on Graphics (ToG), 38(5), 1-19.
      [ paper ]

      We propose to use deep neural networks for generating samples in Monte Carlo integration. Our work is based on non-linear independent components estimation (NICE), which we extend in numerous ways to improve performance and enable its application to integration problems. First, we introduce piecewise-polynomial coupling transforms that greatly increase the modeling power of individual coupling layers. Second, we propose to preprocess the inputs of neural networks using one-blob encoding, which stimulates localization of computation and improves inference. Third, we derive a gradient-descent-based optimization for the Kullback-Leibler and the χ 2 divergence for the specific application of Monte Carlo integration with unnormalized stochastic estimates of the target distribution. Our approach enables fast and accurate inference and efficient sample generation independently of the dimensionality of the integration domain. We show its benefits on generating natural images and in two applications to light-transport simulation: first, we demonstrate learning of joint path-sampling densities in the primary sample space and importance sampling of multi-dimensional path prefixes thereof. Second, we use our technique to extract conditional directional densities driven by the product of incident illumination and the BSDF in the rendering equation, and we leverage the densities for path guiding. In all applications, our approach yields on-par or higher performance than competing techniques at equal sample count.

      Advisor: Fabian Friederichs

    2. GI-GS: Global Illumination Decomposition on Gaussian Splatting for Inverse Rendering
      Chen, H., Lin, Z., Zhang, L. (2025). ICLR 2025.
      [ paper | project page ]

      We present GI-GS, a novel inverse rendering framework that leverages 3D Gaussian Splatting (3DGS) and deferred shading to achieve photo-realistic novel view synthesis and relighting. In inverse rendering, accurately modeling the shading processes of objects is essential for achieving high-fidelity results. Therefore, it is critical to incorporate global illumination to account for indirect lighting that reaches an object after multiple bounces across the scene. Previous 3DGS-based methods have attempted to model indirect lighting by characterizing indirect illumination as learnable lighting volumes or additional attributes of each Gaussian, while using baked occlusion to represent shadow effects. These methods, however, fail to accurately model the complex physical interactions between light and objects, making it impossible to construct realistic indirect illumination during relighting. To address this limitation, we propose to calculate indirect lighting using efficient path tracing with deferred shading. In our framework, we first render a G-buffer to capture the detailed geometry and material properties of the scene. Then, we perform physically-based rendering (PBR) only for direct lighting. With the G-buffer and previous rendering results, the indirect lighting can be calculated through a lightweight path tracing. Our method effectively models indirect lighting under any given lighting conditions, thereby achieving better novel view synthesis and relighting. Quantitative and qualitative results show that our GI-GS outperforms existing baselines in both rendering quality and efficiency.

      Advisor: Fabian Friederichs

    3. High-Gloss SVBRDF Capture Using Bounce Light
      Iser, T.,  Ardelean, A., Weyrich, T. (2026). Computer Graphics Forum. 2026.
      [ paper | project page ]

      Reflectance capture aims at the visual reproduction of an object under varying illumination. Past works differ substantially in their experimental overhead, from single- or few-image approaches, that employ significant (often learned) priors at the expense of biased reconstructions, to more accurate approaches that tend to be time-consuming, which to a good part is due to the need for carefully controlled illumination. Moreover, as we will show, the frequently employed point-light or directional lighting tends to clip highlights and under-sample the reflectance of glossy surfaces, leading to incorrect reconstructions under previously unseen illumination. Our work aims to strike a new balance, combining a low-overhead capture methodology with a fast (neural) model fit. A key feature of our approach is the use of handheld, indirect bounce light that enables a convenient capture methodology, limits the dynamic range of the reflectance (effectively avoiding highlight clipping) and ensures contiguous hemispherical incidence, even with few images, eliminating under-sampling of highly specular reflectance lobes. Moreover, our approach does not require training on pre-existing material datasets and thus is not restricted by the choice of dataset, and its inference scales linearly with the number of pixels, scaling exceptionally well to large image sizes. As a result, our method enables high-resolution capture of a spatially-varying reflectance distribution function (SVBRDF) from a small set of casually captured, indirectly lit photographs, making high-quality material acquisition practical even on consumer hardware. Overall, we believe that our method occupies a unique trade-off between acquisition effort, model assumptions and resulting quality, and it has the potential to transform areas that routinely use handheld point-light sources, such as the popular reflectance transformation imaging (RTI), leading to more faithful reproductions of artefacts and their surface characteristics.

      Advisor: Fabian Friederichs

    4. Fluid Simulation on Neural Flow Maps
      Deng, Y., Yu, H.-X., Zhang, D., Wu, J., & Zhu, B. (2023). SIGGRAPH Asia
      [ paper | project page ]

      This work introduces Neural Flow Maps, a novel method bridging implicit neural representations with flow map theory to achieve state-of-the-art inviscid fluid simulation. It utilizes a hybrid representation fusing small neural networks with multi-resolution sparse grids to compactly and accurately model long-term spatiotemporal velocity fields. This neural velocity buffer enables the symmetric computation of long-term, bidirectional flow maps and their Jacobians, drastically improving accuracy over existing solutions. These flow maps provide high advection accuracy with low dissipation, facilitating high-fidelity incompressible simulations of intricate vortical structures.

      Advisor: Jannis Möller

    5. V-jepa 2: Self-supervised video models enable understanding, prediction and planning.
      Assran, M., Bardes, A., Fan, D., Garrido, Q., Howes, R., Muckley, M., ... & Ballas, N. (2025). arXiv
      [ paper | project page ]

      A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data (robot trajectories), to develop models capable of understanding, predicting, and planning in the physical world. We first pre-train an action-free joint-embedding-predictive architecture, V-JEPA 2, on a video and image dataset comprising over 1 million hours of internet video. V-JEPA 2 achieves strong performance on motion understanding (77.3 top-1 accuracy on Something-Something v2) and state-of-the-art performance on human action anticipation (39.7 recall-at-5 on Epic-Kitchens-100) surpassing previous task-specific models. Additionally, after aligning V-JEPA 2 with a large language model, we demonstrate state-of-the-art performance on multiple video question-answering tasks at the 8 billion parameter scale (e.g., 84.0 on PerceptionTest, 76.9 on TempCompass). Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the Droid dataset. We deploy V-JEPA 2-AC zero-shot on Franka arms in two different labs and enable picking and placing of objects using planning with image goals. Notably, this is achieved without collecting any data from the robots in these environments, and without any task-specific training or reward. This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.

      Advisor: Jannis Möller

    6. PPISP: Physically-Plausible Compensation and Control of Photometric Variations in Radiance Field Reconstruction
      Deutsch, I., Moënne-Loccoz, N., & Gojcic, Z. (2026). CVPR
      [ paper | project page ]

      Multi-view 3D reconstruction methods remain highly sensitive to photometric inconsistencies arising from camera optical characteristics and variations in image signal processing (ISP). Existing mitigation strategies such as per-frame latent variables or affine color corrections lack physical grounding and generalize poorly to novel views. We propose the Physically-Plausible ISP (PPISP) correction module, which disentangles camera-intrinsic and capture-dependent effects through physically based and interpretable transformations. A dedicated PPISP controller, trained on the input views, predicts ISP parameters for novel viewpoints, analogous to auto exposure and auto white balance in real cameras. This design enables realistic and fair evaluation on novel views without access to ground-truth images. PPISP achieves SoTA performance on standard benchmarks, while providing intuitive control and supporting the integration of metadata when available.

      Advisor: Jannis Möller

    7. ReLive: Walking into Virtual Reality Spaces from Video Recordings of One's Past Can Increase the Experiential Detail and Affect of Autobiographical Memories
      Danry, V., Villa, E., Chan, S., Maes, P. (2025). IEEE VR/IEEE Transactions on Visualization and Computer Graphics
      [ paper ]

      With the rapid development of advanced machine learning methods for spatial reconstruction, it becomes important to understand the psychological and emotional impacts of such technologies on autobiographical memories. In a within-subjects study, we found that allowing users to walk through old spaces reconstructed from their videos significantly enhances their sense of traveling into past memories, increases the vividness of those memories, and boosts their emotional intensity compared to simply viewing videos of the same past events. These findings highlight that, regardless of the technological advancements, the immersive experience of VR can profoundly affect memory phenomenology and emotional engagement. As systems enabling immersive memory reconstruction become more ubiquitous, it is crucial to critically examine their effects on human cognition and perception of reality.

      Advisor: Anika Jewst

    8. RippleVision: Unobtrusive Gaze-Dependent Guidance via Directed Wave Motion in Virtual Reality 
      Kudnick, J., Mayer, D., Groth, C., Mohanto, B., Dörner, R. (2026). IEEE VR/IEEE Transactions on Visualization and Computer Graphics
      [ paper ]

      The inherent freedom of exploration in virtual reality poses a challenge for directing user attention toward relevant points or objects of interest, which are potentially located outside the user's field of view or are temporarily obscured. Accordingly, guidance must sustain users' perceptual focus while preserving the sense of presence. This paper presents RippleVision, a subtle gaze guidance technique using wave-like ripples that modulate brightness with inverted polarity between the eyes and appear only within a cone from the point of interest to the gaze location. A calibration study defined detectability and acceptability thresholds for the cues, followed by a comparative user study against four state-of-the-art techniques. Results show RippleVision achieves similar search times to clearly visible techniques while significantly reducing visual dominance. Moreover, its effectiveness scales with cue visibility, improving search performance with minimal impact on perceived visual obstruction and presence.

      Advisor: Anika Jewst

    9. From Novelty to Knowledge: A Longitudinal Investigation of the Novelty Effect on Learning Outcomes in Virtual Reality
      Lee, J., Chen, C., Basu, A. (2025). IEEE VR/IEEE Transactions on Visualization and Computer Graphics
      [ paper ]

      Virtual reality (VR) is increasingly recognized as a powerful educational platform, but the novelty effect–where users experience heightened engagement during initial interactions with new technology–can interfere with learning outcomes. This study investigates how the novelty effect influences learning using a three-wave longitudinal design, tracking changes in information recall and exploratory behavior over three weeks. Our findings reveal that while initial novelty impedes learning, learners' ability to encode educational content improves as they become more familiar with the virtual environment. Additionally, sustained exploratory behavior positively impacts learning over time, reinforcing the importance of active engagement in VR-based education. This study enhances the understanding of VR's long-term educational impact and provides guidance for improving learning effectiveness in immersive learning environments.

      Advisor: Anika Jewst

    Useful Resources

    Example of a good presentation (video on the website under the Presentation section, note how little text is needed, and how much has been visualized to create an intuitive understanding).

    General writing tips for scientific papers (mainly intended for writing scientific articles, but also good to use for summaries).

    Advisors

    Anika Jewst

    Fabian Friederichs

    Jannis Möller