Apparatus and method for covisibility map based image generation

The covisibility map and sparsity score enhance sparse view synthesis by addressing high-uncertainty regions, leading to improved reconstruction accuracy and photorealistic novel view generation.

WO2026104031A1PCT designated stage Publication Date: 2026-05-21HUAWEI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-11-13
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing methods for sparse view synthesis in computer vision struggle with incomplete scene reconstructions due to limited visual information, particularly in high-uncertainty mono-view regions, and lack effective evaluation metrics, leading to unrealistic renderings and error propagation.

Method used

A covisibility map is generated to identify pixel-wise correspondences across multiple input images, allowing for improved handling of occluded or sparsely visible areas, and a sparsity score is calculated to guide the learning process, enhancing reconstruction accuracy by distinguishing between high- and low-uncertainty regions.

Benefits of technology

The method achieves more accurate and complete scene reconstructions by selectively guiding the learning process based on uncertainty levels, resulting in photorealistic novel view synthesis from a minimal number of input images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082258_21052026_PF_FP_ABST
    Figure EP2024082258_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An image processing method (800) for use in novel view image synthesis comprising: receiving (801) multiple input images, each comprising multiple pixels; and for each input image: identifying (802) correspondences between pixels of the respective input image and pixels of each of the other input images; accumulating (803) the correspondences between the pixels of the respective input image and the other input images; and generating (804) a covisibility map for the respective input image in dependence on the accumulated correspondences indicating how many of the multiple input images each pixel of the respective input image appears in. The generation of the covisibility map may allow for improved handling of occluded or sparsely visible areas and may enhance the reliability of novel view synthesis in challenging regions. It may also be used to enhance reconstruction accuracy by guiding the learning process based on uncertainty levels, resulting in more accurate rendered scenes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] APPARATUS AND METHOD FOR COVISIBILITY MAP BASED IMAGE GENERATION

[0002] FIELD OF THE INVENTION

[0003] This invention relates to novel view image synthesis, which may be used in computer vision applications.

[0004] BACKGROUND

[0005] Recent advances in novel view synthesis have made significant progress in overcoming various challenges, such as generalized model building, view synthesis in un-bounded scenes and fast training.

[0006] Sparse view synthesis in computer vision and computer graphics aims to generate novel views of a scene using a minimal number of training images. This task is inherently challenging due to the limited visual information available from sparse input data, which can lead to incomplete or inaccurate scene reconstructions.

[0007] While the latest works using dense training views can achieve photo-realistic fidelity in generated novel view synthesis images, sparse view synthesis methods generally still exhibit unrealistic patterns or floating artifacts due to the fundamental challenge of having a limited number of training views. Despite the limitations of available training view images, recent studies have proposed various approaches, such as generalized novel view synthesis methods, diffusion-based techniques, densification updates, supervision with novel loss functions and the use of pseudo ground truth to improve performance in sparse view synthesis. Although these approaches have improved performance in sparse view synthesis by incorporating multiview geometry understanding and gradient-descent supervision, several challenges remain.

[0008] Firstly, benchmark tables using different sparse view synthesis datasets do not consistently show a single winning method across all datasets. The variability in the proportion of effective multiview regions across datasets can complicate fair supervision and evaluation in sparse view synthesis.

[0009] Current methods for sparse view synthesis typically leverage Structure from Motion (SfM) techniques, such as COLMAP (as described in Johannes Lutz Schonberger and Jan-Michael Frahm, “Structure-from-motion revisited”, Conference on Computer Vision and Pattern Recognition (CVPR), 2016), to produce initial point clouds from a small set of images. However, these methods struggle in sparse settings, resulting in insufficient geometric detail and incomplete scene representations. The initial point clouds generated are often too sparse to capture the necessary scene geometry, leading to artefacts and unrealistic renderings in the synthesized views. Conventional sparse point clouds generated by feature-based keypoint matching (such as SIFT (as described in David G. Lowe, “Distinctive image features from scale-invariant keypoints” International Journal of Computer Vision, 60(2):91— 110, 2004)) in COLMAP with limited training images fail to provide the geometric structure information that dense view settings can achieve.

[0010] Figure la shows an exemplary input image and Figures lb and 1c show examples of initial point clouds for the input image that represent dense view capturing scenario and sparse view capturing scenarios respectively. Methods relying heavily on multiview geometry struggle to reconstruct mono-view regions, where pixels do not appear in other views, which remains a significant challenge in sparse view synthesis. More specifically, errors from high-uncertainty (mono-view) regions can propagate into covisible regions during training, adversely affecting the model’s performance

[0011] Gaussian splatting methods, such as that described in Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis, “3D Gaussian Splatting for Real-time Radiance Field Rendering”, ACM Transactions on Graphics, 42 (4), 2023, have emerged as a promising technique, utilizing 3D points and Gaussian representations to model continuous scene geometry and improve novel view generation. Despite their advancements, these methods still face significant challenges, particularly in handling high-uncertainty regions: that is, areas visible from only a single viewpoint (mono- view regions). Existing approaches often fail to adequately reconstruct these mono-view regions, causing errors that can propagate into low-uncertainty (multiview) regions and degrade the overall quality of the synthesized scene.

[0012] Additionally, there is a lack of effective evaluation metrics for sparse view synthesis. Without a standardized sparsity score metric, it is difficult to benchmark and compare different methods fairly, hindering the identification of specific limitations and areas for improvement in existing technologies.

[0013] Recent advancements in sparse view synthesis have focused on improving scene reconstruction from a minimal number of input images. Methods such as Few-Shot Gaussian Splatting (FSGS) (as described in Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang, “Fsgs: Real-time few-shot view synthesis using gaussian splatting”, European Conference on Computer Vision, Milano, Italy, 2024) and Co-Regularized Gaussian Splatting (CoR-GS) (Jiawei Zhang, Jiahe Li, Xiaohan Yu, Lei Huang, Lin Gu, Jin Zheng, and Xiao Bai, “Cor-gs: Sparse-view 3d gaussian splatting via co-regularization”, European Conference on Computer Vision, Milano, Italy, 2024) have introduced enhancements such as pseudo-ground truth generation and strict multiview constraints.

[0014] Furthermore, approaches used in dynamic and dense view synthesis contexts may exhibit limitations when applied to sparse view synthesis, particularly in effectively handling high-uncertainty mono- view regions.

[0015] Many existing methods in sparse view synthesis exhibit significant limitations in handling high-uncertainty mono- view regions and generally rely heavily on multiview correspondences and lack explicit mechanisms to reconstruct areas visible from only a single viewpoint, resulting in incomplete scene reconstructions and artifacts in synthesized views.

[0016] It is desirable to develop an approach that may overcome at least some of the above issues.

[0017] SUMMARY OF THE INVENTION

[0018] According to a first aspect, there is provided an image processing method for use in novel view image synthesis, the method comprising: receiving multiple input images, each input image comprising multiple pixels; and for each input image: identifying correspondences between pixels of the respective input image and pixels of each of the other input images; accumulating the correspondences between the pixels of the respective input image and the other input images; and generating a covisibility map for the respective input image in dependence on the accumulated correspondences, the covisibility map indicating how many of the multiple input images each pixel of the respective input image appears in.

[0019] The generation of the covisibility map may allow for improved handling of occluded or sparsely visible areas of the input images. It may enhance the reliability of novel view synthesis in challenging regions. It may also be used to enhance reconstruction accuracy by guiding the learning process based on uncertainty levels, resulting in more accurate rendered scenes.

[0020] The step of identifying correspondences between pixels of the respective input image and pixels of each of the other input images may comprise obtaining dense correspondences between the respective input image and each of the other input images. The dense correspondences may be predicted using a model. This may allow state-of-the-art models to be used that can predict the dense correspondences, allowing for the production of a detailed covisibility map with pixel-wise correspondences. Accumulating the correspondences between the respective input image and the other input images may comprise accumulating pixel-wise counts to form the covisibility’ map. This may allow for the production of detailed covisibility maps with pixel-wise counts of covisible views.

[0021] The method may comprise, for each input image, computing a sparsity score by averaging the pixel-wise counts for all pixels of the input image. This may provide a metric to assess dataset sparsity and to apply different weights to supervise the training of Gaussian models.

[0022] The multiple images may represent multiple views of a scene. This may allow the covisibility maps to be formed for training images comprising multiple views of a scene in an image dataset.

[0023] The method may further comprise classifying pixels of the respective input image into one or more high-uncertainty regions and one or more low-uncertainty regions. This may allow for the enhancement of initial points clouds in high-uncertainty regions and for separate supervision of high- and low-certainty image regions during the training of Gaussian splatting models.

[0024] The high-uncertainty' regions may be regions visible from a single input image of the multiple input images. The low-uncertainty regions may be regions visible from more than one input image of the multiple input images. This may allow for the differentiation of different pixels in an image based on how many other input images they are visible in.

[0025] The method may further comprise using the generated covisibility map for the respective input image to selectively guide learning of a structure from motion point cloud and / or update a structure from motion point cloud for the respective input image. This may allow for more complete and accurate image synthesis.

[0026] The method may comprise, for each input image: forming an initial structure from motion point cloud for the respective input image, the initial structure from motion point cloud comprising multiple initial three-dimensional points; generating one or more additional three-dimensional points; and merging the initial three-dimensional points and the additional three-dimensional points to form an updated point cloud. This may allow for full densification and accurate reconstruction of scenes.

[0027] The method may comprise generating the one or more additional three-dimensional points for the low-uncertainty regions and the one or more additional three-dimensional points for the high-certainty regions of the respective input image. This may help to ensure that both low-certainty and high-certainty regions are cohesively reconstructed.

[0028] The method may comprise generating the one or more additional three-dimensional points for the low-certainty regions from dense correspondences between the respective input image and the other input images. The additional points may be generated through triangulation the of dense correspondences. This may preserve geometrically plausible structures and fill in areas where the initial point cloud is sparse.

[0029] The method may comprise rescaling and aligning the one or more additional three-dimensional points generated for the high-uncertainty regions using a linear regression model learned from the one or more additional three-dimensional points generated for the low-uncertainty regions. This may ensure that the additional points are aligned with the scene’s consistent geometry.

[0030] The method may comprise, generating the one or more additional three-dimensional points for the high-uncertainty regions using monocular depth estimation. This may allow further points to be added in high-uncertainty’ regions of the input image. The method may further comprise: splitting the updated point cloud into a first part and a second part; and training separate models for Gaussian primitives having respective positions corresponding to three-dimensional points in the first and second parts. This may allow different regions to be trained separately and for different constraints to be applied.

[0031] The method may comprise training Gaussian primitives having positions corresponding to three-dimensional points in the first part of the updated point cloud using standard supervision techniques and training Gaussian primitives having positions corresponding to three-dimensional points in the second part of the updated point cloud uncertainty using one or more specialized constraints. This may allow points in different regions to be treated differently according to their uncertainty, optimizing the reconstruction process for different regions.

[0032] The method may further comprise training a neural network to classify a respective three-dimensional point corresponding to a respective Gaussian primitive in dependence on the proximity of the respective three-dimensional point to an original scene geometry. This may be a convenient implementation for classifying the points.

[0033] According to another aspect, there is provided an image processing device for use in novel view image synthesis, the device comprising one or more processors configured to: receive multiple input images, each input image comprising multiple pixels; and for each input image: identify correspondences between pixels of the respective input image and pixels of each of the other input images; accumulate the correspondences between the pixels of the respective input image and the other input images; and generate a covisibility map for the respective input image in dependence on the accumulated correspondences, the covisibility map indicating how many of the multiple input images each pixel of the respective input image appears in. This may allow for improved handling of occluded or sparsely visible areas of the input images. It may enhance the reliability of novel view synthesis in challenging regions. It may also be used to enhance reconstruction accuracy by guiding the learning process based on uncertainty levels, resulting in more accurate rendered scenes.

[0034] According to another aspect, there is provided a method for reconstructing a model of a scene in dependence on multiple input images, the model defining a set of primitives shaped as three-dimensional Gaussians, the method comprising, for each input image: forming a structure from motion point cloud comprising multiple three-dimensional points; classifying each of the three-dimensional points in dependence on their respective proximity to an original scene geometry; and supervising the training of Gaussian primitives having positions corresponding to the three-dimensional points based on the respective classifications of the three-dimensional points. This may allow for separate supervision of different regions in dependence on their proximity to the original scene geometry,

[0035] The method may comprise training Gaussian primitives corresponding to three-dimensional points classified as being in proximity to the original scene geometry using standard supervision techniques and training Gaussian primitives corresponding to three-dimensional points classified as not being in proximity to the original scene geometry using one or more specialized constraints. This may allow the image reconstruction process to be optimized for different regions of the image.

[0036] The method may further comprise training a neural network to classify a respective three-dimensional point corresponding to a respective Gaussian primitive as being in proximity to the original scene geometry. This may be a convenient implementation for classifying the points.

[0037] The method may comprise using different loss functions for the training of the Gaussian primitives based on their classifications. This may help to optimize the training process. According to another aspect, there is provided a device for reconstructing a model of a scene in dependence on multiple input images, the model defining a set of primitives shaped as three-dimensional Gaussians, the device comprising one or more processors configured to, for each input image: form a structure from motion point cloud comprising multiple three-dimensional points; classify each of the three-dimensional points in dependence on their respective proximity to an original scene geometry; and supervise the training of Gaussian primitives having positions corresponding to the three-dimensional points based on the respective classifications of the three-dimensional points. This may allow for separate supervision of different regions in dependence on their proximity to the original scene geometry',

[0038] According to another aspect, there is provided a method for generating novel views of a scene from multiple input images by applying a model defining a set of primitives shaped as three-dimensional Gaussians, the method comprising: receiving multiple input images, each input image comprising multiple pixels and representing a view of the scene; generating a covisibility map for each input image, the covisibility’ map indicating how many of the multiple input images each pixel of the respective input image appears in; and in dependence on the generated covisibility maps, generating one or more novel views of the scene using the model. This may improve the quality and consistency of synthesized scenes.

[0039] The method may comprise adjusting a size, density and position of one or more Gaussian splats each comprising multiple Gaussian primitives based on the covisibility map and / or or uncertainty levels indicated by the covisibility map. This may help to ensure that each area of the image is reconstructed with appropriate consideration.

[0040] The model may be a Gaussian splatting model. This may result in the formation of high-quality' rendered images of novel views of a scene.

[0041] According to a further aspect, there is provided a device for generating novel views of a scene from multiple input images by applying a model defining a set of primitives shaped as three-dimensional Gaussians, the device comprising one or more processors configured to: receive multiple input images, each input image comprising multiple pixels and representing a view of the scene; generate a covisibility map for each input image, the co visibility map indicating how many of the multiple input images each pixel of the respective input image appears in; and in dependence on the generated covisibility maps, generate one or more novel view of the scene using the model. This may improve the quality and consistency of synthesized scenes.

[0042] According to a further aspect, there is provided one or more computer programs for instructing a computer comprising one or more processors to implement the methods above.

[0043] According to a further aspect there is provided a data carrier storing in non-transitory form the one or more computer programs above.

[0044] BRIEF DESCRIPTION OF THE FIGURES

[0045] Figure la shows an exemplary input image;

[0046] Figure lb shows an example of an initial point clouds for the input image that represents a dense view capturing scenario; Figure 1c shows an example of an initial point clouds for the input image that represents a sparse view capturing scenarios; Figure 2 schematically illustrates a pipeline of an exemplary’ co visibility’ map generation process;

[0047] Figures 3a, 3b and 3c show examples of covisibility maps corresponding to three training views;

[0048] Figure 4 schematically illustrates a pipeline for enhancing the initial point cloud process;

[0049] Figures 5a, 5b, 5c and 5d show examples of enhanced point clouds; Figure 6 schematically illustrates a pipeline of a separate supervision process;

[0050] Figure 7 schematically illustrates an overview of a covisibility map based Gaussian splatting system;

[0051] Figure 8 shows exemplary steps of a method for generating a co visibility map for use in novel view image synthesis;

[0052] Figure 9 shows exemplary steps of a method for reconstructing a model of a scene in dependence on multiple input images; Figure 10 shows exemplary steps of a method for generating novel views of a scene from multiple input images;

[0053] Figure 11 shows an example of a device for implementing the methods described herein and some of its associated components.

[0054] DETAILED DESCRIPTION

[0055] The goal of novel view synthesis using the system described herein is to reconstruct and render a photorealistic scene from a minimal number of input images, achieving consistent novel view synthesis by leveraging a covisibility map that addresses both high- and low-uncertainty regions. Embodiments of the present invention may achieve significant advancement over existing technologies by effectively addressing core challenges in novel view synthesis, particularly the handling of monoview (high-uncertainty) regions. By solving such technical problems, the present method enables the reconstruction and rendering of photorealistic scenes from a minimal number of input images, achieving consistent and high-quality novel view synthesis that is not possible with prior approaches.

[0056] In the present approach, a covisibility map is generated for improved supervision and fair evaluation. This may be performed by leveraging the results of techniques such as state-of-the-art 3D vision foundation models, for example MASt3R (as described in Vincent Leroy, Yohann Cabon, and Jerome Revaud, “Grounding Image Matching in 3D with MASt3R”, https: / / arxiv.org / abs / 2406.09756, 2024) or DuSt3R (as described in Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii and Jerome Revaud, “DUSt3R: Geometric 3D Vision Made Easy, The IEEE / C VF Conference on Computer Vision and Pattern Recognition (CVPR), 2024) to produce correspondences between input images, combined with a morphological method to finalize map construction.

[0057] Initial point clouds for the input images can be enhanced by incorporating dense correspondences and monocular depth estimation, compensating for sparse point clouds from conventional results, ultimately improving the quality of the final reconstruction. The method can minimize high-uncertainty regions by incorporating additional cues from the monocular depth estimation, while boosting low-uncertainty regions by allocating more resources through updated initial point clouds generated by traditional triangulation methods, resulting in denser reconstructions.

[0058] Leveraging the corrected initial points and covisibility map, weighted super-vision may be applied across covisible and non-covisible regions, adjusting supervision strength based on covisibility levels.

[0059] Additionally, a novel metric to benchmark the difficulty of sparse view datasets is introduced, enabling consistent evaluation and identification of effective methods across datasets.

[0060] The covisibility map reflects the correlation between the training views or between the training and test views. The co visibility map can be used to properly manage the trade-off between low-uncertainty (multiview regions) and high-uncertainty (monoview regions) and can efficiently identify dense correspondences between a given training view image and all other training view images.

[0061] The generation of a covisibility map will now be described. In these examples, the method uses the 3D foundational model, MASt3R, and morphological operations. The available resources, such as dense correspondences (used to create the covisibility maps) and a monocular depth estimation model, also allows to update the initial point clouds used in 3D Gaussian splatting methods across both covisible and mono-view regions.

[0062] Generally, the formation of the covisibility map comprises accumulating covisible counts on each pixel of the image from dense correspondences between input images. The covisibility map is created by identifying correspondences between a given training view image and all other training view images. For each input image training view, a covisibility map is generated by accumulating the correspondences between that view and the remaining views. In the examples described herein, the dense correspondences between two input images are predicted using the MASt3R model, which provides data to track the visibility of each pixel across the training images.

[0063] Dense correspondences are predicted between pairs of input images. For example, to generate a covisibility map for a particular training view, the dense correspondences between that view and the other remaining views are obtained. These correspondences directly indicate where pixels in one image match those in other images. Based on these correspondences, pixel-wise counts are accumulated to form the covisibility map. The resulting map contains values indicating how many views each pixel appears in. For example, where there are three input images depicting different views of a scene, the map represents whether each respective pixel of the input image is seen in one, two, or all available views.

[0064] This process is repeated for all training views, leading to co visibility maps for the entire dataset. Unlike known methods which only identify if a pixel is visible, such covisibility maps provide detailed pixel-wise visibility counts. The pixel-wise covisible counts are accumulated across all input training view's. This additional level of information is crucial for distinguishing between low- and high-uncertainty regions, thereby improving the effectiveness of both training and evaluation in sparse view synthesis tasks. The co visibility maps can thus serve as a key tool for optimizing the representation and handling of uncertainty in scene reconstruction.

[0065] The covisibility maps can be refined using morphological operations to handle occlusions and sparsely visible areas, explicitly distinguishing between high-uncertainty (mono-view) and low-uncertainty (multiview) regions. The morphological operations can be used to fill gaps in the covisibility map, refining it for accurate representation.

[0066] Furthermore, a sparsity score can be calculated by summing the pixel counts across the covisibility map, providing a metric to assess dataset sparsity. The sparsity (or covisibility) score can be determined using the average covisibility count across all pixels in an image. This score can serve as an indicator in different contexts. For instance, in training, where the entire set of training images represents a scene, the sparsity score can be further averaged across all training images for a scene. In some implementations, the sparsity scores of the individual scenes can be averages across a dataset comprising multiple scenes, each scene having multiple training views. This final average score can help to set a threshold for the supervision, allowing to adjust the weighting ratio in training for a scene depending on the score. In evaluation, the sparsity score, averaged across test images (compared to the given training images), indicates the difficulty level of the scene. This score reveals how challenging it may be to achieve high performance and prevents direct comparison of performance across scenes with different sparsity scores. By using this sparsity score as a benchmark, fairer comparisons can be promoted in the evaluation process for sparse view synthesis.

[0067] Figure 2 shows a pipeline of the covisibility’ map generation process. The input is a set of image frames from each camera, shown at 201. In this example, there are three input training view' images. However, the method may be use for a different number of input images. For example, there may be up to (or exactly) 3, 5, 6, 9, 12, 24 or fewer input images. There may be a larger number of input images. Dense correspondences for all pairs of input images are extracted, as shown at 202. The resulting dense correspondences for each pixel are shown at 203 for an exemplar}' input image. These examples show' covisibility count values of 0, 1, and 2 for each pixel. A value of 0 (dark grey region) represents pixels that appear only in the mono-view, 1 (grey region) represents pixels that appear in one other view, and 2 (white region) represents pixels that appear in all views. This is performed for all input image views, as illustrated at 204.

[0068] Then, as illustrated at 205, the pixel-wise covisible view counts are accumulated. At each pixel, accumulated co visible counts are stored in the covisibility map. The generated covisibility maps for each of the three input training view' images 201 are shown at 206.

[0069] As shown at 207, using morphological operations, such as dilation and erosion, the covisibility maps can be enhanced by filling holes generated from distortions.

[0070] The resulting map for one of the images is shown at 208. This map, indicating pixel-wise uncertainty information, can be used not only for better supervision but also for generating a sparsity score to understand the difficulties of a test view. The generation of a sparsity score is indicated at 210 for the co visibility map shown at 209.

[0071] The covisibility map generation process will now be described mathematically. Let T =

[0072]

[0073] represent the set of training view images, where n is the total number of training views. For each training view

[0074] a co visibility map Mtis generated by accumulating the correspondences

[0075]

[0076] between and the set of remaining view images / yj, where j + i. For every pair (Ii, I{j}), with j ≠ i, dense correspondences Ci^j can be predicted by a dense correspondence prediction function fcor, where Ci^j = fcor(Ii, Ij).

[0077] For example, where there are three training view images, to generate a covisibility map for the training view image the dense correspondences C1^[2:3] between I1 and the remaining images I[2:3] are obtained.

[0078] Since C{ directly provides the correspondences, pixel-wise counts in M1 are accumulated based on the locations where these correspondences exist.

[0079] The covisibility map Mt, initialized as a zero matrix of size W x H for a pixel (x, y) in view 1,- is defined as:

[0080] Mi(x,y) = Σ δ(x,y)∈풫(Ci^j), ∀(x,y) ∈

[0081]

[0082] where δ(x,y)∈풫(Ci^j) is an indicator function that returns 1 if pixel (x,y) in Ii has a correspondence match in view Ij and 0

[0083]

[0084] otherwise. J’(C ) represents the set of pixel coordinates in 1,- that have correspondences in Ij.

[0085] The resulting covisibility map Mtcaptures pixel-wise covisibility counts, ranging from 0 to n — 1, indicating the frequency a pixel appears across views.

[0086] For example, for the three view case, once the correspondences C1^[2:3] are obtained, the correspondence counts are accumulated for each pixel in I1. The resulting covisibility map has values 0, 1, or 2, indicating whether a pixel appears in one, two, or all three views, respectively. This process is repeated for all training images, resulting in covisibility maps M = {M1, M2, ..., Mn} for the entire set of training images.

[0087] Figures 3a, 3b and 3c each show an example of covisibility maps corresponding to three training views, for three different scenes. These examples show covisibility count values of 0, 1, and 2 for each pixel. A value of 0 (black region) represents pixels that appear only in the mono- view, 1 (grey region) represents pixels that appear in one other view, and 2 (white region) represents pixels that appear in all views.

[0088] This covisibility map provides detailed pixel-wise covisibility counts, providing precise covisibility details at the pixel level, offering crucial in-sights into low- and high-uncertainty regions, which is particularly advantageous in sparse data scenarios. This enhanced detail can improve both supervision and evaluation, particularly in sparse view synthesis tasks.

[0089] By accurately identifying regions with varying degrees of visibility and uncertainty, the covisibility map can be used to guide the reconstruction process to treat each area appropriately. High-uncertainty regions, which are challenging due to limited or no overlap between views, are effectively managed, preventing errors from propagating into low-uncertainty regions. This will be described in more detail later.

[0090] As mentioned previously, points clouds can be generated to form multiple three-dimensional (3D) points that can be the positions of 3D Gaussians in a Gaussian splatting model.

[0091] In sparse view synthesis, starting from overly sparse Structure from Motion (SfM) initial point clouds using approaches such as COLMAP often limits the full potential of Gaussian splatting methods. COLMAP tends to neglect regions with insufficient overlap, resulting in incomplete geometry recovery, which can restrict the densification strategy of 3D Gaussian splatting methods that depend on the current point clouds to fully densify and accurately reconstruct scenes. While recent approaches have shown the benefits of using dense initial point clouds in dense view scenarios, these methods do not easily extend to sparse view synthesis, where large mono- view regions may remain unhandled and cannot be effectively used for multiview triangulation.

[0092] The updating of initial point clouds will now be described. This method focusses on both low-uncertainty' (multiview-effective) and high-uncertainty (mono- view) regions.

[0093] In the examples described herein, the point cloud is enhanced via combined triangulation with mono-depth rescaling. Initial points clouds are enhanced by adding triangulated points derived from dense correspondences. Additional points can be added to the initial point cloud selectively, for example based on a distance threshold, to cover sparse regions without introducing redundancy.

[0094] In one example, the additional points may be generated through triangulation using the MASt3R model. The additional points can preserve geometrically plausible structures and fill in areas where the initial point cloud misses points, particularly where traditional keypoint matching and bundle adjustment techniques are inadequate. After generating these triangulated points, they can be combined with the initial point cloud points in regions where COLMAP’ s data is sparse, leading to a more complete and accurate scene reconstruction.

[0095] Additionally, high-uncertainty regions (which may be determined in dependence on the covisibility map) can be recovered by learning the correlation between mono-depth estimation and triangulated multiview point clouds, rescaling and positioning points in these regions. In high-uncertainty regions, use monocular depth estimation to generate 3D points and refine their scale through a linear regression model trained on co visible points.

[0096] For low-uncertainty regions where sparse SIM points initially exist, an initial point cloud Pc (for example, obtained using COLMAP) can be updated by incorporating a set of triangulated 3D points PTderived from dense correspondences. These additional points, which may be generated using a MASt3R model, as described previously, not only pre-serve geometrically plausible structures but also fill in areas that COLMAP misses, particularly in regions where keypoint-based matching and bundle adjustment are inadequate.

[0097] To generate PT, a triangulation process can be performed for each input image view Ii, utilizing dense correspondences 풫(Ci{j}) with other views I{j} (e.g., either 2 or 3 correspondences in a 3-view scenario). After each triangulation, the utilized correspondences may be removed from subsequent views to prevent redundant calculations. The complete triangulated point set, PT, is then combined with the initial COLMAP points. This may be done only in regions where COLMAP’s point cloud is sparse or missing.

[0098] To formalize this process, let D(Pc, PT) the distance between points in PTand their nearest neighbors in Pc. The updated point cloud, Pu, for low-uncertainty regions is then defined as:

[0099] Pu = Pc u {pTG PTI D(pT, PC) > e}

[0100] where e is a predefined threshold ensuring that only points from PTsufficiently distant from existing points in Pcare added to the updated point cloud. This strategy ensures that MASt3R’s triangulated points can complement COLMAP by filling in areas where COLMAP’s point cloud is sparse or missing, thereby improving coverage in covisible, low-uncertainty regions.

[0101] In high-uncertainty mono- view regions, monocular depth estimation can be used to generate a preliminary set of 3D points, Pd, by unprojecting depth estimates using provided camera intrinsics. This set can be defined as Pd= PdowU p^lgh^ where Pdowcomprises points unprojected from multiview regions and p^lghcomprises points unprojected from mono- view regions. A linear regression model can be applied to correct depth distortions. For example, the scale of the additional points can be refined with an anisotropic linear regression model learned from covisible regions, incorporating regularization to prevent overfitting. This scaling model, learned from the low-uncertainty' covisible regions, ensures that the depth-based points are aligned with the scene’s consistent geometry. By applying linear regression, the scale and position of the monocular depth points is adjusted to better integrate them into the overall point cloud.

[0102] This approach adapts the scale consistently across both low- and high-uncertainty regions, accounting for axis-specific distortions and improving the integration of depth-based points into the point cloud.

[0103] To align the randomly scaled unprojected points with triangulated points from multiview geometry, an anisotropic scaling transformation fscale is learned using covisible points from Pu^low ⊆ Pu as ground truth and Pd^low as input. This transformation applies scaling parameters learned through linear regression to match the multiview-derived geometry:

[0104]

[0105] For consistency, Pu = Pu^low ∪ Pu^high, where Pu^low is selected based on covisibility:

[0106] Pu^low = {p | Mi(π(Pu, Hi)) ≥ 1} where Tt(Pu, Hi) denotes the projection of Puusing the camera transformation matrix Htwhich returns (x, y) image coordinates for a camera viewpoint, ensuring only points visible in multiple views are included.

[0107] After training, fscaieis applied to the randomly scaled unprojected points from depth estimates in high-uncertainty regions:

[0108] phigh >f( phlgh

[0109] rs lscale\ d J

[0110] This transformation aligns the randomly selected scaled points p^lghto the triangulated metric scale P'llfl' enhancing spatial consistency across regions. The final point cloud, Pflnai, is then formed by combining the updated points from multiview and monocular depth sources, ensuring that both low- and high-uncertainty regions are cohesively reconstructed. Pflnaiis defined as:

[0111] — plow phigh

[0112] ‘ final

[0113] Therefore, the initial point cloud is updated by combining an initial point cloud, for example a traditional SIM point cloud generated by COLMAP, with additional 3D points derived from dense correspondences, for example provided by MASt3R. For high-uncertainty' regions where multiview triangulation is not feasible, monocular depth estimation can be used to generate 3D points. These points are then rescaled and aligned using a linear regression model learned from covisible regions.

[0114] By integrating monocular depth ranking and dense correspondences, initial point cloud quality can be significantly enhanced, resulting in better reconstructions even under sparse view conditions. Augmenting the initial, sparse point cloud with additional points compensates for limitations in traditional SIM techniques when few input images are available and can provide better geometric detail and completeness, especially in regions with limited visibility.

[0115] This can result in more dense and complete point clouds by filling gaps where the initial point cloud is sparse or incomplete, preserving overall scene geometry. This may result in enhanced depth accuracy. Monocular depth points are rescaled to align with multiview-derived geometry', improving depth consistency across the scene. This is also a robust initialization for Gaussian splatting, and serves as a strong foundation for subsequent reconstruction stages, leading to more accurate and stable 3D reconstructions.

[0116] Figure 4 shows an example of a pipeline for enhancing the initial point cloud process in the three-view case. The covisibility maps generated using the method described above are shown at 400.

[0117] The initial point cloud for an input image 401 is shown at 402. The dense correspondences for the input image, determined as described above, are shown at 403.

[0118] Given the correspondences and the initial point clouds (which may be generated using COLMAP), 3D points from the correspondences are triangulated at 404. At 405, these points are then merged from the correspondences with the initial point clouds in areas where data does not exist to produce an updated initial point cloud for the low'-uncertainty regions, as shown at 406.

[0119] For the 3D points located in low-uncertainty regions and the unprojected points (generated from depth images 408 by using mono-depth estimation of the input image 401 at 407), a linear regression model is learned, as shown at 409. Using the learned parameters (for scaling random unprojected points to metric scale), the high-uncertainty regions (which are unprojected monodepth points) are scaled to update the merged point clouds, as shown at 410. The final updated point cloud is shown at 411.

[0120] In general, the process comprises forming an initial structure from motion point cloud for a respective input image, the initial structure from motion point cloud comprising multiple initial three-dimensional points, generating one or more additional three-dimensional points, and merging the initial three-dimensional points and the additional three-dimensional points to form an updated point cloud.

[0121] This process enables more accurate and comprehensive 3D scene reconstructions in sparse view synthesis, overcoming the limitations of prior methods that struggle to handle uncertainty in mono-view regions.

[0122] Figures 5a, 5b, 5c and 5d show examples of enhanced point clouds for different input images of different scenes. This example shows the enhanced point clouds at each step of the enhancement process. Initial point clouds are shown at 501, 504, 507 and 510 in Figures 5a, 5b, 5c and 5d respectively. Additional points are shown at 502, 504, 506 and 511 in Figures 5a, 5b, 5c and 5d respectively. The final point clouds are shown at 503, 506, 509 and 512 in Figures 5a, 5b, 5c and 5d respectively

[0123] The following describes a covisibility map-based supervision method, which leverages the depth-adjusted and updated point cloud, Pfinai, along with the covisibility map Mtto adaptively supervise regions with varying levels of uncertainty as represented by covisible counts at each pixel. This approach differs from state-of-the-art Gaussian splatting methods by offering targeted supervision and region -specific weighting across varying covisibility regions, where Mt= {0: n — 1}. Additionally, supervising a subset of Gaussians located outside the training view frustum may be supervised by applying weights based on a scene sparsity or average covisibility score S = avg M^.nj, as will be described in more detail later.

[0124] This method can be integrated into any 3D Gaussian splatting-based framework by introducing an additional proximity loss term that applies an inverse proximity score from a classifier to weight the supervision of the Gaussians’ 3D coordinates. The proximity loss term may be weighted by an inverse proximity score learned through an MLP classifier trained on the final point clouds.

[0125] An exemplary implementation for within camera- view frustum will now be described. To implement this, a neural network, such as an MLP scene geometry proximity classifier fP, is trained on Pflnai, labeling these points with a binary indicator y = 1 positive points (representing the original scene geometry) and y = 0 for negative points (randomly distributed points, Prandomy constrained to remain distant from Pflnai). The classifier fP: R3-> [0, 1] outputs a proximity score s for each 3D point p as:

[0126] s = fP(P) G [0, 1]

[0127] where s approaches 1 for points near the original scene geometry and approaches 0 for distant points. Using batch processing, the proximity scores for the entire Gaussian set, ff, can be efficiently computed. The classifier’s results can be used as a loss value for supervision. A proximity loss term Lpcan be defined as:

[0128] + (1 - X(5))wout) • (1 - S)

[0129]

[0130] sec where g denotes the 3D position of a Gaussian in the set G, representing the complete Gaussian set. The indicator function X(g) may be defined as:

[0131] r(nx=[L.if iitg. Ht) e [0,w) x [0,h)

[0132]

[0133] I 0, otherwise

[0134] where n(g, is the projection of g under camera transformation

[0135]

[0136] onto the covisibility map Mt.

[0137] Within the camera-view frustum, the weight a)inis used for Gaussians projected onto covisibility mask regions, defined as:

[0138] 1

[0139]

[0140] M^g. H^ + l

[0141] where Ml(n(_g, Hl)') is the covisibility count of g in the covisibility map for the corresponding projection. This weighting ensures that high-covisibility regions, which are frequently supervised by other loss terms, receive lower proximity-related supervision, while mono- view regions benefit from a more direct application of £p.

[0142] An exemplary implementation for outside the camera-view frustum will now be described. In typical 3D Gaussian splatting-based methods, supervision occurs image-by-image, updating weights iteratively rather than across all views simultaneously. Only Gaussians within the view frustrum are supervised per iteration. Thus, Gaussians outside the current camera view frustum may be ignored in each iteration. However, particularly for scenes with high sparsity scores (S > 0.7), supervision can be extended to Gaussians located outside the view frustum, applying a scene-level weight a>outfor these points:

[0143] Mout= max(0,— )

[0144] This weight scales linearly from 1 to 0 as S ranges from 1 to 0.7. For higherS, more weight can be applied to supervise the Gaussians outside the view frustum, whereas lower values reduce this weight, ensuring efficient and targeted supervision. This approach can improve supervision outside the view frustum in high-covisibility score scenes, ensuring comprehensive guidance while making efficient use of the proximity classifier.

[0145] This approach allows supervision based on the classification results (indicating correctness of being close to the initial geometry), while still leveraging the covisibility map to supervise each uncertainty region(s) differently with different weights. Taking the example of three view cases, 1, 2, and 3 covisibility count regions can be weighted differently.

[0146] This classification approach provides similar benefits to a conventional nearest neighbor search by identifying how far any 3D point is from the original geometry. However, unlike nearest neighbor searches, which would require N computations per gaussian, this classifier-based method significantly reduces computational demands, making it much more practical.

[0147] While recent 3D Gaussian splatting methods have focused on improving the densification of point clouds through pseudoground truths or enforcing strict multiview constraints, these approaches do not effectively address the unique challenges of high-uncertainty mono-view regions, where no overlapping views are available. In some implementations, an additional supervision term can be used specifically for high-uncertainty regions, on top of the conventional objective function used for low-uncertainty regions, providing a more robust and balanced solution for sparse view synthesis. The depth-adjusted and updated initial point cloud, Pftnai, along with the covisibility map Mt, can be used to independently supervise both covisible (low-uncertainty) and mono-view (high-uncertainty) regions. Distinct supervision strategies can be applied in low- and high-uncertainty regions, effectively preventing error propagation between them and improving the overall performance in sparse view synthesis.

[0148] The supervision can advantageously be tailored for high-uncertainty regions, where Mt= 0. The approach introduces an additional high-uncertainty-targeted supervision term on top of the state-of-the-art objective function for low-uncertain regions, to improve model ro-bustness and accuracy in sparse data scenarios.

[0149] To implement this strategy, the updated point cloud Pflnaican be divided into two Gaussian models: GMi0Wfor low-uncertain (covisible) regions and GMhlghfor high-uncertain (mono-view) regions. The overall training objective is optimized through backpropagation with separate Gaussian models for each region, while en-suring a shared optimization framework. This enables to address the different characteristics and requirements of each region effectively.

[0150] For the low-uncertainty model GMi0W, a conventional objective function can be utilized that minimizes reconstruction error between the Gaussian model and the ground truth:

[0151]

[0152] =(1—

[0153] where / ;owrepresents the rendered image from GMi0W, supervised against the ground truth image / *, £1 is the LI reconstruction loss, and GD-SSIMdenotes the D-SSIM term, A is a balancing parameter between the LI and D-SSIM loss terms, consistent with conventional objectives in Gaussian splatting. This objective function reinforces the multiview-derived structure in covisible regions, ensuring accurate ge-ometry where points are observed in multiple views.

[0154] For the high-uncertainty model GMhlgh, two primary constraints may be used: 1) proximity to depth-adjusted points: Gaussians in GMhlghshould align closely with the updated point cloud P^gdh, derived from the monocular depth-based 3D points. 2) mono-visibility constraint: points in GMhlghshould be visible in only a single view to respect the mono-view nature of these regions. The objective function for high-uncertainty regions can therefore be formulated as follows:

[0155] Ghtgh= i\GMhlgh{p} - P^adh^\2+ 2 l{GMhiflh(p) n GMhigh(q) = 0 })

[0156]

[0157] qET’q*p

[0158] where GM / li(; / 1(p) denotes the Gaussian position for point p in the high-uncertainty region, P^pdh(p) represents the depth-adjusted point for p derived from monocular depth estimates, l{GM / li(; / 1(p) n GM / li(; / 1(q) = 0 } is an indicator function that enforces the mono- visibility constraint by ensuring that GM / li(; / 1(p) does not overlap with points from other views q, 2 is the weighting factor for the mono- visibility constraint.

[0159] To verify mono-visibility during training, the points from GMhlghcan be projected onto the image plane for each training view using associated camera intrinsics and extrinsics. A mask can then be applied where Mt= 0 to penalize points that erroneously appear in other views. The overall super- vision objective integrates both low- and high-uncertainty loss terms, ensuring separate and focused training for each region:

[0160] G

[0161]

[0162] In another implementation, the objective function can be expressed as:

[0163] L = (1 - 2)£i( / ,r) +D-sslM(i, )+ Lp

[0164] where I is the rendered image, supervised against ground truth / *. The proximity loss term Lpadds adaptive supervision to improve geometry alignment.

[0165] By applying distinct supervision terms, error propagation between covisible and mono-view regions can be mitigated. This adaptive and region-specific approach enhances training stability, leading to improved accuracy and completeness in scene reconstruction, particularly in sparse view synthesis scenarios.

[0166] This separate supervision framework effectively addresses the unique uncertainty characteristics of sparse view synthesis, providing more robust and consistent results over prior methods by directly targeting high-uncertainty regions while leveraging established multiview constraints in low-uncertainty areas.

[0167] Generally, the updated point cloud can be split into parts: one set containing points located in covisible (low-uncertainty) regions and the other set representing points in mono-view (high-uncertainty) regions. These sets are then supervised using separate Gaussian models - one for the low-certainty regions and one for the high-certainty regions. The covisible regions can use standard supervision techniques, ensuring the points adhere to the ground truth geometry across multiple views. In contrast, the mono- view regions can be handled with specialized constraints, where points are adjusted to align with depth-adjusted 3D points from monocular depth estimation and are constrained to appear in only one view.

[0168] Additionally, to improve this supervision process, a simple neural network, such as a MLP, to classify any 3D point based on its location in either a low-uncertainty or high-uncertainty region. The network can be trained using labeled data, where points within covisible areas are labeled as low-uncertainty and points in mono- view regions are labeled as high-uncertainty. The MLP may also provide a confidence score for each classification. This classification network is then applied during the densification process of Gaussian splatting, where newly added points are classified into high- or low-uncertainty regions, allowing for appropriate supervision based on their classification. This ensures that points in each region are treated differently according to their level of uncertainty, optimizing the reconstruction process for both regions.

[0169] By applying this region-specific supervision, errors may be prevented from propagating between high- and low-uncertainty regions. This separation can stabilize the training process, leading to more accurate and complete 3D reconstructions in sparse view synthesis, particularly in challenging areas where traditional methods struggle. This approach can offer a significant improvement over existing methods by directly addressing both high- and low-uncertainty regions through tailored supervision.

[0170] Figure 6 shows a pipeline of the separate supervision process. The set of input images are shown at 601. At 602, a co visibility map is generated for each of the input images, as described above. At 603, mono-depth estimation is performed for each input image to produce depth images 604. Taking the covisibility maps, initial point clouds (PCL) 605 and the depth images 604 as input, the initial PCLs are updated at 606, as described above.

[0171] For each input image, given the updated PCL and the co visibility map, each point of the PCL is annotated with an uncertainty label (such as 0 or 1 ). Then, an uncertainty-level classifier is learnt of xyz coordinates at 607. This is used to classify 3D points corresponding to the positions of newly added Gaussians at 608. This acts as a predictor that scores the uncertainty level of densified (cloned or copied) Gaussians in the point cloud. Based on the classifications of the points as low- or high-uncertainty, they can be supervised separately at 609 and the resulting images rendered at 610. Different loss functions can be applied for Gaussians based on their assigned uncertainty levels (higher low-uncertainty) to optimize training. Mono- visibility constraints can be enforced by implementing constraints that prevent Gaussians in high-uncertaintv areas from appearing in multiple views, isolating these regions to prevent cross-region errors.

[0172] At 611, uncertainty-aware update (regularization) can be performed, as described above.

[0173] The model adjusts supervision dynamically based on uncertainty levels. By applying different training strategies for high-uncertainty and low-uncertainty regions based on the co visibility map, conventional objective functions can be utilized for low-uncertainty regions that leverage multiview consistency, minimizing reconstruction error between the Gaussian model and ground truth. The approach introduces additional constraints for high-uncertainty regions using depth-adjusted monocular estimates and can enforce mono- visibility constraints to ensure that Gaussians align closely with monocular depth points and do not erroneously appear in other views. High-uncertainty (mono-view) regions can be more effectively handled by applying strict supervision while preventing errors from propagating into co visible regions, leading to stabilized training and improved overall performance.

[0174] Effective multiview regions can be fully leveraged through updated initial point clouds, while enforcing strict supervision over high-uncertainty regions using the latest mono- view depth estimation.

[0175] This approach tailors the supervision to the specific characteristics of each region, with differentiated supervision for low- and high-uncertainty regions, effectively handling the unique challenges posed by high-uncertainty areas, and prevents errors from high-uncertainty regions from adversely affecting the reconstruction of low-uncertainty regions. This can allow for mitigated error propagation and enhances training stability by preventing errors in mono- view regions from impacting multiview regions. This can also allow for improved reconstruction accuracy, allowing to achieve more accurate and complete scene reconstructions, even in high-uncertainty areas. The adaptive training framework allows for focused optimization in each region, leading to overall performance improvements.

[0176] As mentioned previously, a sparsity score can be determined for the covisibility map. Datasets often capture regions that only appear in a few training views, affecting the quality of initial point clouds based on the availability of effective multiviews for each region. The performance of state-of-the-art methods varies depending on the size of the effective multiview region, especially in sparse view synthesis where limited training views lead to fewer effective multiview regions in the test set. Consequently, a method that performs well on one dataset may not rank consistently across others due to these variations. To address this issue, using the sparsity score to consider the availability of effective multiviews in each dataset, performance of different methods can be more fairly evaluated for synthetic view cases. This approach allows to consistently assess methods across different datasets

[0177] Using the approach described herein, a covisibility map-based Gaussian splatting system that can advantageously be used for spare view synthesis can be realised, focusing on effectively handling both low-uncertainty (covisible) and high-uncertainty (mono- view) regions. Covisibility information can be incorporated directly into the Gaussian splatting process. The covisibility map to adaptively guide the Gaussian splatting process. Splat sizes, densities, and positions of Gaussians can be dynamically updated based on the visibility’ and uncertainty levels indicated by the covisibility map. The system can adapt to the specific needs of both covisible and non-co visible areas within a unified framework. In such implementations, the combination of covisibility map generation, enhanced point cloud creation, and separate supervision results in a comprehensive system that leverages the covisibility map for improved sparse view synthesis. Unlike conventional 3D Gaussian splatting methods, which do not address uncertainty-aware point cloud densification or supervision, the system can effectively handle both high- and low-uncertainty regions by adapting the reconstruction process based on the co visibility information. The covisibility map guides the Gaussian splatting process, allowing for adaptive adjustment of splat parameters such as position, size, and density according to the regional uncertainty levels, ensuring that each area of the scene is reconstructed with appropriate consideration.

[0178] Figure 7 schematically illustrates an overview of an exemplary complete covisibility map based Gaussian splatting system.

[0179] The system takes as input the input images 701 and the initial point clouds for the input images, which in this example is in the form of SfM points 702. Dense correspondences are extracted from each of the input images at 703 and the covisibility map is generated from the dense correspondences at 704. At 705, holes in the map are filled using morphological operations, refining the map for accurate representation, to give the final covisibility maps at 706.

[0180] At 707, additional 3D points for the PCLs are reconstructed from the dense correspondences, as described above. The points are merged at 708 to give the first updated PCLs at 709.

[0181] The first updated PCLs are then further updated by performing monocular depth estimation on the input images at 710 to produce depth images at 711 from which additional 3D points can be formed, as described earlier.

[0182] At 712a and 712b, the 3D points in the first updated PCL for a respective input image are classified into high-uncertainty and low-uncertainty points in dependence on the covisibility map for that input image. The processes at 712a and 712b are the collection process to identify the corresponding points from two different sets of points (712a from dense points, for example once recovered from SfM + MASt3R, and 712b from unprojected mono depth estimates).

[0183] From the low-uncertainty points from each of these sets, a scale mapping function is learned at 713 to give a scale converter at 714. This is then applied to the high-uncertainty points at 715 to scale the depth-based points in a metric scale to give the fully enhanced PCLs at 716.

[0184] At 717, the Gaussians of a Gaussian splatting model are initialized and supervision based on the co visibility map is performed at 718 to give the trained Gaussians at 719. The covisibility map is utilized to adaptively guide the Gaussian splatting process. Using the covisibility’ map, uncertainty’ labels can be assigned to newly added Gaussians and to supervise the high- and low-uncertainty' Gaussians in their respective regions.

[0185] The system can handle both low- (covisible) and high-uncertainty (mono-view) regions by optimizing splat positions using covisibility information.

[0186] Camera poses are input at 720 for the formation of novel views. Images are rendered at 721 to give the resulting novel view images at 722. The rendering module can perform uniform rendering quality across both covisible and non-covisible regions in novel views.

[0187] The system can be implemented as an integrated pipeline. The system can take a minimal number of input images and generate novel views of a 3D scene using Gaussian splatting, ensuring consistent reconstruction across both covisible and non-covisible areas in novel view synthesis. This approach ensures that regions with varying levels of visibility are reconstructed accurately by accommodating their unique requirements. The photorealism and consistency of the rendered scenes are enhanced by optimizing splat placement and density according to co visibility.

[0188] This approach effectively balances the representation of visible and occluded regions, leading to a cohesive 3D reconstruction. The visual quality’ of synthesized views is improved, which is particularly important in applications demanding high fidelity.

[0189] The sparsity score metric for the co visibility maps can be used for sparse view synthesis benchmark evaluation.

[0190] Although the described approach is particularly advantageous when used in extremely sparse view synthesis scenarios (for example, using less than 5 input images per scene), making it particularly suitable for practical applications with minimal input data, the concept of high- and low-uncertainty regions also applies to settings with a dense number of training views. The number of training views does not guarantee that every' region of a scene is equally visible across all views; occlusions, camera angles, and scene complexity can result in some areas being observed in fewer views than others. The present approach can be advantageously used in scenarios with high view counts but low region-specific covisibility'. Conventional methods often provide unbalanced supervision, leading to failures in reconstructing under-represented high-uncertainty' regions. The present system addresses this issue by utilizing the co visibility map to balance supervision and reconstruction efforts across different regions, improving the overall quality and consistency of the synthesized scenes. Region-wise adaptive supervision is applied based on covisbility counts, improving results for imbalanced datasets. In such cases, central objects are well-captured, while outer areas receive minimal coverage. This problem is common in real-world scenarios. Therefore, the approach provides a versatile solution for various applications where high-quality 3D scene reconstruction from minimal input data is desirable.

[0191] Figure 8 shows an example of an image processing method for use in novel view image synthesis. At step 801, the method comprises receiving multiple input images, each input image comprising multiple pixels. The steps 802, 803, 804 are then performed for each input image. At step 802, the method comprises identifying correspondences between pixels of the respective input image and pixels of each of the other input images. At step 803, the method comprises accumulating the correspondences between the pixels of the respective input image and the other input images. At step 804, the method comprises generating a covisibility map for the respective input image in dependence on the accumulated correspondences, the covisibility’ map indicating how many of the multiple input images each pixel of the respective input image appears in.

[0192] Figure 9 shows an example of a method for reconstructing a model of a scene in dependence on multiple input images, the model defining a set of primitives shaped as three-dimensional Gaussians. The method comprising performing steps 901-903 for each input image. At step 901, the method comprises forming a structure from motion point cloud comprising multiple three-dimensional points. At step 902, the method comprises classifying each of the three-dimensional points in dependence on their respective proximity to an original scene geometry. At step 903, the method comprises supervising the training of Gaussian primitives having positions corresponding to the three-dimensional points based on the respective classifications of the three-dimensional points.

[0193] Figure 10 shows an example of a method for generating novel views of a scene from multiple input images by applying a model defining a set of primitives shaped as three-dimensional Gaussians. At step 1001, the method comprises receiving multiple input images, each input image comprising multiple pixels and representing a view of the scene. At step 1002, the method comprises generating a co visibility map for each input image, the covisibility map indicating how many of the multiple input images each pixel of the respective input image appears in. At step 1003, the method comprises, in dependence on the generated covisibility maps, generating one or more novel views of the scene using the model. Figure 11 shows an example of a device 1100 configured to output and / or implement the methods described herein. The device 1100 comprises a processor 1101 and a memory 1102. The memory 1102 stores in a non-transient way code that is executable by the processor 1101 to implement the respective entity in the manner described herein. The device 1100 may be implemented by hardware or may be service-based computing device, for example it may be implemented as a cloud-based computing device. Therefore, the methods may be deployed in multiple ways, for example in the cloud, on the device, or in dedicated hardware.

[0194] The device 1100 may in some implementations also comprise a transceiver that is capable of communicating over a network with other entities. For example, the device may receive videos and / or images from other entities. Those entities may be physically remote from the device 1100. The network may be a publicly accessible network such as the internet. The entities may in some cases be based in the cloud. These entities may be logical entities. In practice they may each be provided by one or more physical devices such as servers and data stores, and the functions of two or more of the entities may be provided by a single physical device. Each physical device implementing an entity comprises a processor and a memory.

[0195] Embodiments of the present invention can overcome the disadvantages of the prior art by introducing a co visibility map-based Gaussian splatting system that generates detailed covisibility maps, enhances initial point clouds through dense correspondences and monocular depth estimation, and applies separate supervision strategies for high- and low-uncertainty regions, effectively reconstructing mono-view areas and preventing error propagation into covisible regions, thereby achieving consistent and high-quality novel view synthesis from minimal input images.

[0196] Using the covisibility map, which provides pixel-wise counts representing the number of covisible views among training image viewpoints, a Gaussian splatting model can be trained for sparse novel view synthesis. By fully leveraging both the covisibility map and the updated initial point clouds (benefiting from mono-view depth ranking estimation and dense correspondences) a separate supervision strategy can be employed that focuses differently on covisible regions and non-covisible (high-uncertainty) regions. A metric, the sparsity score, can be determined which reflects the degree to which the training images contain covisible regions compared to test views by analysing the covisible counts in the proposed covisibility maps. This metric provides a fair comparison across scenes and benchmarks, explaining the meaningful performance differences of existing methods across various datasets, which can be ranked according to the sparsity-score. The method may achieve state-of-the-art results in sparse view synthesis across both datasets and sparsity scores.

[0197] The present approach addresses the challenges of reconstructing photorealistic 3D scenes from a minimal number of input images. Unlike existing methods that primarily focus on multiview geometry and often neglect high-uncertainty (mono- view) regions, the system effectively manages both low- and high-uncertainty regions by leveraging covisibility maps. This ensures consistent and accurate novel view synthesis even when the input images have limited overlap.

[0198] The approach described herein offers a comprehensive solution that directly addresses the limitations of existing sparse view synthesis techniques. By effectively managing uncertainty through detailed co visibility maps, targeted handling of both high-and low-uncertainty regions can be enabled, leading to more accurate and reliable novel view synthesis even with minimal input images. Enhancing initial point clouds with dense correspondences and monocular depth estimates provides a denser and more accurate representation of scene geometry, improving the foundation for reconstruction processes. Adaptive supervision strategies prevent errors from propagating between regions of different uncertainty levels, resulting in more stable training and higher-quality outputs. The unified integration of covisibility maps into the Gaussian splatting framework allows for dynamic adaptation based on regional characteristics, ensuring consistent handling of varying visibility levels within the same system. By effectively reconstructing high-uncertainty'' mono-view regions and preventing error propagation into covisible regions (challenges inadequately addressed by prior art) embodiments of the present invention can achieve consistent and high-quality novel view synthesis from minimal input images, marking a significant advancement over existing methods.

[0199] The method introduces explicit supervision and depth scaling for high-uncertainty' regions using a covisibility’ map-based approach. Monocular depth estimates and covisibility' information can be leveraged to guide the reconstruction in mono-view regions, ensuring a more complete and accurate scene synthesis.

[0200] The detailed co visibility maps provide pixel-wise counts of covisible views and can be sued not only for evaluation but also to enhance the initialization of point clouds and to supervise both high- and low-uncertainty regions during training. This comprehensive use of covisibility information enables effective reconstruction in sparse view scenarios. Consistent scene reconstruction may be achieved even when multiview correspondences are unavailable.

[0201] The generation of the covisibility map also allows for improved handling of occluded or sparsely visible areas and enhances the reliability of novel view synthesis in challenging regions. It can guide the learning process based on uncertainty levels, resulting in more accurate rendered scenes. It can also provide a structured representation of the scene’s uncertainty', supporting consistent benchmarking across different datasets.

[0202] In summary, this co visibility’ map-based Gaussian splatting system represents a significant advancement in novel view synthesis by effectively leveraging co visibility information and addressing uncertainty at every' stage of the reconstruction process. By integrating the covisibility map into point cloud enhancement and applying separate supervision strategies, accurate and consistent reconstructions across both covisible and non-covisible areas can be achieved. This adaptive approach not only excels in sparse view scenarios but may also enhance dense view synthesis, making it a versatile solution for various applications requiring high-quality' 3D scene reconstruction from minimal input data.

[0203] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.

Claims

CLAIMS1. A image processing method (800) for use in novel view image synthesis, the method comprising:receiving (801) multiple input images (201), each input image comprising multiple pixels; andfor each input image:identifying (802) correspondences between pixels of the respective input image and pixels of each of the other input images;accumulating (803) the correspondences between the pixels of the respective input image and the other input images; andgenerating (804) a covisibility map (208) for the respective input image in dependence on the accumulated correspondences, the covisibility map indicating how many of the multiple input images each pixel of the respective input image appears in.

2. The method as claimed in claim 1, wherein identifying correspondences between pixels of the respective input image and pixels of each of the other input images comprises obtaining dense correspondences between the respective input image and each of the other input images, wherein the dense correspondences are predicted using a model.

3. The method as claimed in any preceding claim, wherein accumulating the correspondences between the respective input image and the other input images comprises accumulating pixel-wise counts to form the covisibility map.

4. The method as claimed in claim 3, wherein the method comprises, for each input image, computing a sparsity score by averaging the pixel-wise counts for all pixels of the input image.

5. The method as claimed in any preceding claim, wherein the multiple images represent multiple views of a scene.

6. The method as claimed in any preceding claim, wherein the method further comprises classifying pixels of the respective input image into one or more high-uncertainty regions and one or more low-uncertainty regions.

7. The method as claimed in claim 6, wherein the high-uncertainty regions are regions visible from a single input image of the multiple input images and wherein the low-uncertainty regions are regions visible from more than one input image of the multiple input images.

8. The method as claimed in any preceding claim, wherein the method further comprises using the generated co visibility map for the respective input image to selectively guide learning of a structure from motion point cloud and / or update a structure from motion point cloud for the respective input image.

9. The method as claimed in any preceding claim, wherein the method comprises, for each input image:forming an initial structure from motion point cloud for the respective input image, the initial structure from motion point cloud comprising multiple initial three-dimensional points;generating one or more additional three-dimensional points; andmerging the initial three-dimensional points and the additional three-dimensional points to form an updated point cloud.

10. The method as claimed in claim 9, as dependent on claim 6 or claim 7, wherein the method comprises generating the one or more additional three-dimensional points for the low-uncertainty regions and the one or more additional three-dimensional points for the high-certainty regions of the respective input image.

11. The method as claimed in claim 10, wherein the method comprises generating the one or more additional three-dimensional points for the low-certainty regions from dense correspondences between the respective input image and the other input images.

12. The method as claimed in claim 10 or claim 11, wherein the method comprises rescaling and aligning the one or more additional three-dimensional points generated for the high-uncertainty regions using a linear regression model learned from the one or more additional three-dimensional points generated for the low-uncertainty regions.

13. The method as claimed in any of claims 10 to 12, wherein the method comprises, generating the one or more additional three-dimensional points for the high-uncertainty regions using monocular depth estimation.

14. The method as claimed in any of claims 9 to 13, wherein the method further comprises:splitting the updated point cloud into a first part and a second part; andtraining separate models for Gaussian primitives having respective positions corresponding to three-dimensional points in the first and second parts.

15. The method as claimed in claim 14, wherein the method comprises training Gaussian primitives having positions corresponding to three-dimensional points in the first part of the updated point cloud using standard supervision techniques and training Gaussian primitives having positions corresponding to three-dimensional points in the second part of the updated point cloud uncertainty using one or more specialized constraints.

16. The method as claimed in claim 14 orclaim 15, wherein the method further comprises training a neural network to classify a respective three-dimensional point corresponding to a respective Gaussian primitive in dependence on the proximity of the respective three-dimensional point to an original scene geometry.

17. An image processing device (1100) for use in novel view image synthesis, the device comprising one or more processors (1101) configured to:receive (801) multiple input images (201), each input image comprising multiple pixels; andfor each input image:identify (802) correspondences between pixels of the respective input image and pixels of each of the other input images;accumulate (803) the correspondences between the pixels of the respective input image and the other input images; andgenerate (804) a covisibility map (208) for the respective input image in dependence on the accumulated correspondences, the covisibility map indicating how many of the multiple input images each pixel of the respective input image appears in.

18. A method (900) for reconstructing a model of a scene in dependence on multiple input images, the model defining a set of primitives shaped as three-dimensional Gaussians, the method comprising, for each input image:forming (901) a structure from motion point cloud comprising multiple three-dimensional points; classifying (902) each of the three-dimensional points in dependence on their respective proximity to an original scene geometry; andsupervising (903) the training of Gaussian primitives having positions corresponding to the three- dimensional points based on the respective classifications of the three-dimensional points.

19. The method as claimed in claim 18, wherein the method comprises training Gaussian primitives corresponding to three-dimensional points classified as being in proximity to the original scene geometry using standard supervision techniques and training Gaussian primitives corresponding to three-dimensional points classified as not being in proximity to the original scene geometry using one or more specialized constraints.

20. The method as claimed in claim 18 or claim 19, wherein the method further comprises training a neural network to classify a respective three-dimensional point corresponding to a respective Gaussian primitive as being in proximity to the original scene geometry.

21. The method as claimed in any of claims 18 to 20, wherein the method comprises using different loss functions for the training of the Gaussian primitives based on their classifications.

22. A device (1100) for reconstructing a model of a scene in dependence on multiple input images, the model defining a set of primitives shaped as three-dimensional Gaussians, the device comprising one or more processors (1101) configured to, for each input image:form (901) a structure from motion point cloud comprising multiple three-dimensional points; classify (902) each of the three-dimensional points in dependence on their respective proximity to an original scene geometry; andsupervise (903) the training of Gaussian primitives having positions corresponding to the three-dimensional points based on the respective classifications of the three-dimensional points.

23. A method (1000) for generating novel views of a scene from multiple input images by applying a model defining a set of primitives shaped as three-dimensional Gaussians, the method comprising:receiving (1001) multiple input images, each input image comprising multiple pixels and representing a view of the scene;generating (1002) a covisibility map for each input image, the co visibility’ map indicating how many of the multiple input images each pixel of the respective input image appears in; andin dependence on the generated covisibility maps, generating (1003) one or more novel views of the scene using the model.

24. The method as claimed in claim 23, wherein the method comprises adjusting a size, density and position of one or more Gaussian splats each comprising multiple Gaussian primitives based on the covisibility map and / or or uncertainty levels indicated by the covisibility map.

25. The method as claimed in claim 23 or claim 24, wherein the model is a Gaussian splatting model.

26. A device (1100) for generating novel views of a scene from multiple input images by applying a model defining a set of primitives shaped as three-dimensional Gaussians, the device comprising one or more processors (1101) configured to: receive (1001) multiple input images, each input image comprising multiple pixels and representing a view of the scene;generate (1002) a covisibility map for each input image, the covisibility map indicating how many of the multiple input images each pixel of the respective input image appears in; andin dependence on the generated covisibility maps, generate (1003) one or more novel view of the scene using the model.

27. One or more computer programs for instructing a computer comprising one or more processors to implement the method as claimed in any of claims 1 to 16, 18 to 21 or 23 to 25.

28. A data carrier (1102) storing in non-transitory form the one or more computer programs as claimed in claim 27.