Master thesis : Feed-Forward Novel View Synthesis For Soccer Scenes With Priors
Gardier, Simon
Promoteur(s) :
Cioppa, Anthony
Date de soutenance : 29-jui-2026/30-jui-2026 • URL permanente : http://hdl.handle.net/2268.2/26108
Détails
| Titre : | Master thesis : Feed-Forward Novel View Synthesis For Soccer Scenes With Priors |
| Titre traduit : | [fr] Synthèse de nouvelles vues par propagation avant pour les scènes de football avec a priori |
| Auteur : | Gardier, Simon
|
| Date de soutenance : | 29-jui-2026/30-jui-2026 |
| Promoteur(s) : | Cioppa, Anthony
|
| Membre(s) du jury : | Geuzaine, Christophe
Massoz, Quentin |
| Langue : | Anglais |
| Nombre de pages : | 71 |
| Mots-clés : | [en] 3DGS [en] Gaussian Splatting [en] Deep Learning [en] Computer Vision [en] Pose alignment [en] Pytorch [en] Python [en] Novel View Synthesis [en] NVS [en] 3D reconstruction [en] Radiance Fields [en] 3D [en] 3D Vision |
| Discipline(s) : | Ingénierie, informatique & technologie > Sciences informatiques |
| Centre(s) de recherche : | Vision and Information Understanding Laboratory (VIULab) EVS Broadcast Equipment |
| Public cible : | Chercheurs Professionnels du domaine Etudiants Grand public |
| Institution(s) : | Université de Liège, Liège, Belgique |
| Diplôme : | Master en sciences informatiques, à finalité spécialisée en "intelligent systems" |
| Faculté : | Mémoires de la Faculté des Sciences appliquées |
Résumé
[en] Novel View Synthesis ( NVS ) for soccer broadcasting would enable new production
effects such as free-viewpoint replays, 3D camera transitions, immersive 3D analyses,
etc. State-of-the-art NVS pipelines based on Neural Radiance Field (NeRF) and
3D Gaussian Splatting ( 3DGS ) produce high-quality novel views. However, their
per-scene optimization, dense input requirements, and dependence on Structure-From-
Motion ( SFM ) make them incompatible with the time and sparse input constraints of
soccer broadcasting. Recent feed-forward 3D foundation models such as π3, Visual
Geometry Grounded Transformer ( VGGT), and Depth-Anything-3 (DA3) recon-
struct an entire scene in a single forward pass, but were neither trained nor evaluated
on soccer scenes, where the scenes are large and the priors from the broadcast pipeline
(camera calibration, pitch geometry, segmentation masks) are not used.
Our work answers what a sparse-input feed-forward NVS pipeline made for soccer
scenes would look like. We propose a modular four-stage architecture. A frozen 3D
foundation model that outputs a coarse point cloud from a sparse input, a per-view
depth alignment that aligns this point cloud to the broadcast world frame using the
pitch as a reference, a learnable Dense Prediction Transformer (DPT) Gaussian head
that regresses 3D Gaussians from the features of the foundation model, and a 3DGS
rasterizer that renders novel views in real time. The pipeline is trained and evaluated
on a feed-forward training and validation split that we create from SoccerNet-NVS.
We benchmark three multi-view 3D foundation models and one monocular baseline
on a custom evaluation pipeline. π3 ranks first on five of six metrics and is selected as
the backbone. Our depth alignment beats global transform methods while preserving
the broadcast world frame. On top of this aligned geometry, the trained Gaussian
head produces plausible feed-forward novel views on a test scene in around one second
per scene on a single GPU. Our pipeline is several orders of magnitude faster than
the per-scene optimization methods. While the experiments do not currently produce
high-quality NVS results, our approach presents a credible path towards real-time,
photo-realistic 3D soccer scene reconstruction.
Fichier(s)
Document(s)
thesis.pdf
Description: Main document
Taille: 35.8 MB
Format: Adobe PDF
Annexe(s)
code.zip
Description:
Taille: 10.65 MB
Format: Unknown
Citer ce mémoire
L'Université de Liège ne garantit pas la qualité scientifique de ces travaux d'étudiants ni l'exactitude de l'ensemble des informations qu'ils contiennent.

Master Thesis Online

