Master thesis : Feed-Forward Novel View Synthesis For Soccer Scenes With Priors
Gardier, Simon
Promotor(s) :
Cioppa, Anthony
Date of defense : 29-Jun-2026/30-Jun-2026 • Permalink : http://hdl.handle.net/2268.2/26108
Details
| Title : | Master thesis : Feed-Forward Novel View Synthesis For Soccer Scenes With Priors |
| Translated title : | [fr] Synthèse de nouvelles vues par propagation avant pour les scènes de football avec a priori |
| Author : | Gardier, Simon
|
| Date of defense : | 29-Jun-2026/30-Jun-2026 |
| Advisor(s) : | Cioppa, Anthony
|
| Committee's member(s) : | Geuzaine, Christophe
Massoz, Quentin |
| Language : | English |
| Number of pages : | 71 |
| Keywords : | [en] 3DGS [en] Gaussian Splatting [en] Deep Learning [en] Computer Vision [en] Pose alignment [en] Pytorch [en] Python [en] Novel View Synthesis [en] NVS [en] 3D reconstruction [en] Radiance Fields [en] 3D [en] 3D Vision |
| Discipline(s) : | Engineering, computing & technology > Computer science |
| Research unit : | Vision and Information Understanding Laboratory (VIULab) EVS Broadcast Equipment |
| Target public : | Researchers Professionals of domain Student General public |
| Institution(s) : | Université de Liège, Liège, Belgique |
| Degree: | Master en sciences informatiques, à finalité spécialisée en "intelligent systems" |
| Faculty: | Master thesis of the Faculté des Sciences appliquées |
Abstract
[en] Novel View Synthesis ( NVS ) for soccer broadcasting would enable new production
effects such as free-viewpoint replays, 3D camera transitions, immersive 3D analyses,
etc. State-of-the-art NVS pipelines based on Neural Radiance Field (NeRF) and
3D Gaussian Splatting ( 3DGS ) produce high-quality novel views. However, their
per-scene optimization, dense input requirements, and dependence on Structure-From-
Motion ( SFM ) make them incompatible with the time and sparse input constraints of
soccer broadcasting. Recent feed-forward 3D foundation models such as π3, Visual
Geometry Grounded Transformer ( VGGT), and Depth-Anything-3 (DA3) recon-
struct an entire scene in a single forward pass, but were neither trained nor evaluated
on soccer scenes, where the scenes are large and the priors from the broadcast pipeline
(camera calibration, pitch geometry, segmentation masks) are not used.
Our work answers what a sparse-input feed-forward NVS pipeline made for soccer
scenes would look like. We propose a modular four-stage architecture. A frozen 3D
foundation model that outputs a coarse point cloud from a sparse input, a per-view
depth alignment that aligns this point cloud to the broadcast world frame using the
pitch as a reference, a learnable Dense Prediction Transformer (DPT) Gaussian head
that regresses 3D Gaussians from the features of the foundation model, and a 3DGS
rasterizer that renders novel views in real time. The pipeline is trained and evaluated
on a feed-forward training and validation split that we create from SoccerNet-NVS.
We benchmark three multi-view 3D foundation models and one monocular baseline
on a custom evaluation pipeline. π3 ranks first on five of six metrics and is selected as
the backbone. Our depth alignment beats global transform methods while preserving
the broadcast world frame. On top of this aligned geometry, the trained Gaussian
head produces plausible feed-forward novel views on a test scene in around one second
per scene on a single GPU. Our pipeline is several orders of magnitude faster than
the per-scene optimization methods. While the experiments do not currently produce
high-quality NVS results, our approach presents a credible path towards real-time,
photo-realistic 3D soccer scene reconstruction.
File(s)
Document(s)
thesis.pdf
Description: Main document
Size: 35.8 MB
Format: Adobe PDF
Annexe(s)
code.zip
Description:
Size: 10.65 MB
Format: Unknown
Cite this master thesis
The University of Liège does not guarantee the scientific quality of these students' works or the accuracy of all the information they contain.

Master Thesis Online

