Pseudo-view collaborative geometric perception sparse view 3DGS reconstruction method, apparatus and device, and storage medium

By employing a geometry-aware sparse view 3DGS reconstruction method based on pseudo-view collaboration, and utilizing Gaussian solution pooling and pseudo-view regularization techniques, the overfitting problem in 3D reconstruction under sparse views is solved, achieving high-quality 3D scene reconstruction and new view synthesis.

CN121582453APending Publication Date: 2026-02-27XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511613380.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing 3D reconstruction techniques suffer from overfitting under sparse view input conditions, leading to a decline in reconstruction quality. In particular, traditional neural radiation field methods have long training times and are difficult to edit scenes, while 3D Gaussian sputtering techniques are prone to degradation in reconstruction quality under sparse viewpoints.

Method used

A geometrically perceptive sparse view 3DGS reconstruction method based on pseudo-view collaboration is adopted. The 3D scene is initialized by Gaussian solution pooling, and the Gaussian representation is optimized by combining the depth stamping surface reconstruction module, the equalization shape constraint module, and the opacity constraint module. The pseudo-view regularization and neighborhood smoothing modules are used to generate high-quality 3D scene reconstruction results.

Benefits of technology

High-quality 3D scene reconstruction and new view synthesis were achieved under sparse view conditions, effectively overcoming the overfitting problem and improving reconstruction quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582453A_ABST
    Figure CN121582453A_ABST
Patent Text Reader

Abstract

The invention provides a pseudo-view collaborative geometric perception sparse view 3DGS reconstruction method, device and equipment and a storage medium, Gaussian representation of a three-dimensional scene is initialized through Gaussian solution pool densification, then geometric perception optimization is executed, outlier gauss deviating from the scene surface are forcibly embedded into a correct position by using a depth stamping surface reconstruction module, and the 3D scene is reconstructed. Excessive tensile deformation of the Gaussian ellipsoid is prevented through equilibrium form constraint, and redundant gauss are eliminated by using opacity constraint, so that an accurate scene geometric structure is established. Then, double-source collaborative pseudo-visual-angle regularization is implemented, a pseudo-training visual angle is intelligently generated through a visual angle interpolation dynamic scoring mechanism, and under the visual angle, a pseudo-true value generated by current Gaussian rendering and original view warping transformation is utilized at the same time for bidirectional constraint; and finally, a neighborhood smoothing module is applied, a neighborhood graph is constructed based on a Gaussian space neighborhood relation, and color jump and geometric discontinuity caused by discrete Gaussian distribution are eliminated through luminosity and depth smoothing constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sparse view reconstruction, and particularly to a pseudo-view cooperative geometric perception sparse view reconstruction method. Figure 3 DGS reconstruction methods, apparatus, equipment and storage media. Background Technology

[0002] With the rapid development of computer vision and 3D reconstruction technologies, novel view synthesis has become a core technology in fields such as virtual reality, autonomous driving, and digital twins. This technology aims to reconstruct 3D scenes from limited viewing angles and generate realistic images from arbitrary, unseen perspectives. However, existing 3D reconstruction techniques face numerous challenges under sparse view input conditions.

[0003] While traditional Neural Radiation Field (NeRF) methods can generate high-quality rendering results, they suffer from drawbacks such as long training times, difficulties in scene editing, and high computational resource consumption. 3D Gaussian Splatting (3DGS) technology, which combines the advantages of differentiable rendering and explicit editing, has gained widespread acceptance in the field of 3D reconstruction. However, its unstructured explicit representation is prone to severe overfitting under sparse viewpoint input, leading to a decline in reconstruction quality.

[0004] In view of the above, this application is hereby submitted. Summary of the Invention

[0005] This invention discloses a pseudo-view collaborative geometry-aware sparse view Figure 3 The DGS reconstruction method, apparatus, equipment, and storage medium aim to overcome the overfitting problem of 3D Gaussian splashing under conditions with only a small number of input views, and achieve high-quality 3D scene reconstruction and new view synthesis.

[0006] The first embodiment of the present invention provides a geometry-aware sparse view for pseudo-view collaboration. Figure 3 DGS reconstruction methods include: Obtain the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters. Initialize the 3D scene to be reconstructed using the Gaussian pooling densification method to generate an initial 3D Gaussian representation, wherein each Gaussian contains position, rotation, scale, opacity and color attributes. Based on the initial 3D Gaussian representation, geometrically aware Gaussian optimization is performed to generate a geometrically optimized 3D Gaussian representation. The optimization includes embedding outliers into the scene surface through a depth stamping surface reconstruction module, adjusting the shape of the Gaussian ellipsoid through an equalization shape constraint module, and controlling the visibility of the Gaussian through an opacity constraint module. A dual-source collaborative pseudo-viewpoint regularization is applied to the geometrically optimized 3D Gaussian representation to obtain a pseudo-viewpoint enhanced 3D Gaussian representation. The pseudo-viewpoint regularization includes generating a camera pose of a pseudo-training viewpoint from the existing training viewpoint through a viewpoint interpolation dynamic scoring mechanism, rendering a pseudo-viewpoint image and depth using the 3D Gaussian representation under the pseudo-training viewpoint, simultaneously projecting the original training viewpoint onto the pseudo-training viewpoint through a warp transformation to generate a pseudo-ground value, and updating the 3D Gaussian representation by minimizing the difference between the rendering result and the pseudo-ground value. The neighborhood smoothing module is applied to the pseudo-viewpoint enhanced 3D Gaussian representation. By analyzing the spatial relationships between Gaussians, a neighborhood graph is constructed, and photometric smoothing constraints and depth smoothing constraints are implemented to generate the final 3D Gaussian representation for new view synthesis.

[0007] Preferably, the embedding of outlier Gaussians into the scene surface via the deep stamping surface reconstruction module specifically involves: When rendering the initial 3D Gaussian representation, the opacity parameter is frozen, assuming that all Gaussians are close to opaque in order to identify outliers. Predict fine-grained depth maps from training views using the diffusion-based monocular depth estimation model Lotus. ; The scene depth is rendered using the 3D Gaussian representation through the 3DGS rendering pipeline. ; The scene surface is regularized by minimizing the Pearson correlation loss, forcing the 3D Gaussian representation to converge to the real scene surface. Its expression is: ,in, This indicates a deep loss.

[0008] Preferably, adjusting the Gaussian ellipsoid shape through the equilibrium shape constraint module specifically involves: The scale parameter of each Gaussian in the 3D Gaussian representation is constrained. The maximum and minimum principal axis scale ratios are limited by scale difference regularization, while small Gaussians are encouraged to improve texture reconstruction quality. The expression is as follows: ,in, To balance the loss due to morphological constraints, The weight of the first regularization term. The weight of the second regularization term, The number of Gaussian ellipsoids, Let be the scale vector of the i-th Gaussian ellipsoid; Preferably, controlling the visibility of Gaussians through the opacity constraint module specifically involves: The opacity of each Gaussian in the three-dimensional Gaussian representation Applying a suppression constraint, its expression is: ,in, For the loss due to opacity constraints, The weight of the opacity constraint term. The number of Gaussian ellipsoids; By employing a low-opacity pruning strategy, Gaussians with opacities below a threshold are removed, thereby optimizing the spatial distribution of the three-dimensional Gaussian representation.

[0009] Preferably, the dynamic scoring mechanism specifically includes: From the set of training viewpoints represented by the geometrically optimized 3D Gaussian representation, the score between the candidate viewpoint and the current training viewpoint is calculated, and its expression is: ,in, For dynamic scoring, To control the weights of the space-angle constraint balance, To control the weights of spatial distance constraints, To address the difference angle between viewing directions exceeding a threshold The penalty weights and Euclidean distances exceed the mean. The penalty weight, For indicator functions, The angle of difference in viewing direction; Select the K candidate viewpoints with the lowest scores, and interpolate between them and the current viewpoint. Linear interpolation of position components: ,in, The position components of the newly generated camera pose. For the position components of the current training viewpoint, To match the position components of the target viewpoint, These are the interpolation coefficients for the position components; Rotational component spherical linear interpolation: ,in, For the rotation component of the newly generated camera pose, The rotation component is the current training viewpoint. To match the rotation component of the target viewpoint, For the interpolation coefficients of the rotation component, The angle of difference in viewing direction; Generate new camera pose: ,in, This is the newly generated camera pose.

[0010] Preferably, the photometric smoothing constraint includes: For each Gaussian in the pseudo-viewpoint-enhanced 3D Gaussian representation A neighborhood graph is constructed by finding K nearest neighbor Gaussians centered on the given location, and the color transition between adjacent Gaussians is constrained by the following loss function:

[0011] in, For photometric smoothing loss based on nearest neighbor graph, for The set of Gaussian neighbors of the head node and its K-nearest Euclidean distance. Let represent the center positions of the i-th and j-th Gaussians, respectively. Parameters used to control sensitivity to spatial distance.

[0012] Preferably, the depth smoothing constraint includes: The depth gradient is estimated using a monocular depth pre-trained model, and an adaptive constraint is applied to the depth gradient rendered by the 3D Gaussian representation:

[0013] in, For adaptive depth smoothing loss, and These represent the gradients of the depth information estimated by the monocular depth pre-trained model in the x and y directions, respectively. and These represent the gradients of the rendering depth in the x and y directions, respectively.

[0014] The second embodiment of the present invention provides a geometry-aware sparse view for pseudo-view collaboration. Figure 3 DGS reconstruction apparatus, characterized in that it comprises: An initial Gaussian representation generation unit is used to obtain the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters. The 3D scene to be reconstructed is initialized by the Gaussian unpooling densification method to generate an initial 3D Gaussian representation, wherein each Gaussian contains position, rotation, scale, opacity and color attributes. The Gaussian representation optimization unit is used to perform geometry-aware Gaussian optimization based on the initial 3D Gaussian representation to generate a geometry-optimized 3D Gaussian representation. The optimization includes embedding outliers into the scene surface through a depth stamping surface reconstruction module, adjusting the shape of the Gaussian ellipsoid through an equalization shape constraint module, and controlling the visibility of the Gaussian through an opacity constraint module. A pseudo-viewpoint regularization unit is used to perform dual-source collaborative pseudo-viewpoint regularization on the geometrically optimized 3D Gaussian representation to obtain a pseudo-viewpoint enhanced 3D Gaussian representation. The pseudo-viewpoint regularization includes generating a camera pose of a pseudo-training viewpoint from the existing training viewpoint through a viewpoint interpolation dynamic scoring mechanism, rendering a pseudo-viewpoint image and depth using the 3D Gaussian representation under the pseudo-training viewpoint, simultaneously projecting the original training viewpoint onto the pseudo-training viewpoint through a warp transformation to generate a pseudo-ground value, and updating the 3D Gaussian representation by minimizing the difference between the rendering result and the pseudo-ground value. A smoothing unit is used to apply a neighborhood smoothing module to the pseudo-viewpoint-enhanced 3D Gaussian representation. By analyzing the spatial relationships between Gaussians, a neighborhood graph is constructed, and photometric and depth smoothing constraints are implemented to generate the final 3D Gaussian representation for new view synthesis.

[0015] The third embodiment of the present invention provides a geometry-aware sparse view for pseudo-view collaboration. Figure 3 The DGS reconstruction device is characterized by comprising a memory and a processor, wherein the memory stores a computer program that can be executed by the processor to achieve a pseudo-view cooperative geometry-aware sparse view as described in any of the above claims. Figure 3 DGS reconstruction method.

[0016] The fourth embodiment of the present invention provides a computer-readable storage medium, characterized in that it stores a computer program, the computer program being executable by a processor of the device in which the computer-readable storage medium is located, to implement a pseudo-view cooperative geometry-aware sparse view as described in any of the preceding claims. Figure 3 DGS reconstruction method.

[0017] Based on the present invention, a geometry-aware sparse view is used for pseudo-view collaboration. Figure 3 The DGS reconstruction method, apparatus, equipment, and storage medium initialize the Gaussian representation of the 3D scene through Gaussian solution pooling, followed by geometry-aware optimization. A depth-stamped surface reconstruction module forcibly embeds outliers from the scene surface into their correct positions. Equilibrium shape constraints prevent excessive stretching and deformation of the Gaussian ellipsoid, and opacity constraints eliminate redundant Gaussians, thus establishing an accurate scene geometry. Next, dual-source collaborative pseudo-viewpoint regularization is implemented. A pseudo-training viewpoint is intelligently generated through a viewpoint interpolation dynamic scoring mechanism. Under this viewpoint, bidirectional constraints are applied using both the current Gaussian rendering and pseudo-ground values ​​generated from the original view's warp transformation. Finally, a neighborhood smoothing module is applied to construct a neighborhood graph based on Gaussian spatial proximity relationships. Luminosity and depth smoothing constraints eliminate color abruptness and geometric discontinuities caused by discrete Gaussian distributions. Attached Figure Description

[0018] Figure 1This is a pseudo-view collaborative geometry-aware sparse view provided in the first embodiment of the present invention. Figure 3 A flowchart illustrating the DGS reconstruction method; Figure 2 This is a comparative visualization of the results on the LLFF dataset provided by this invention; Figure 3 This is a diagram showing a comparison of visualization results on the TNT dataset and the MipNeRF360 dataset provided by this invention; Figure 4 This is a pseudo-view cooperative geometry-aware sparse view provided in the second embodiment of the present invention. Figure 3 A schematic diagram of the DGS reconstruction device modules. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0021] This invention discloses a pseudo-view collaborative geometry-aware sparse view Figure 3 The DGS reconstruction method, apparatus, equipment, and storage medium aim to overcome the overfitting problem of 3D Gaussian splashing under conditions with only a small number of input views, and achieve high-quality 3D scene reconstruction and new view synthesis.

[0022] The first embodiment of the present invention provides a geometry-aware sparse view for pseudo-view collaboration. Figure 3 The DGS reconstruction method can be executed by a reconstruction device or system, specifically by one or more processors within the reconstruction device or system, to at least perform the following steps: S101, obtain the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters, initialize the 3D scene to be reconstructed by the Gaussian pooling densification method, and generate an initial 3D Gaussian representation, wherein each Gaussian contains position, rotation, scale, opacity and color attributes. In this embodiment, the reconstruction device can be a desktop computer, laptop computer, server or other terminal with data processing capabilities. The reconstruction device can be equipped with a corresponding operating system and application software, and the functions required in this embodiment can be realized through the combination of the operating system and application software. First, when acquiring the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters, this embodiment can use an open-source structured dataset containing scene images and corresponding precise camera pose information, or obtain it through a self-made static scene dataset. The self-made dataset can be created by taking photos from multiple angles or extracting frames from moving videos, thus ensuring the diversity and applicability of the sparse view input. Subsequently, this embodiment integrates a structured motion recovery module to accurately predict the camera pose parameters corresponding to the image set. This module utilizes a structured motion recovery algorithm to process the geometric relationships of the input view, achieving accurate estimation of the camera position and pose to provide a reliable pose foundation. Next, the 3D scene to be reconstructed is initialized using the Gaussian unpooling densification method. First, a preliminary Gaussian point cloud representation is generated based on the acquired sparse view and camera pose. Then, a Gaussian unpooling strategy is introduced to densify the Gaussian distribution of the scene, ensuring uniform coverage of Gaussian points in space. The original 3DGS pipeline is used for parameter optimization and rendering initialization to generate an initial 3D Gaussian representation. Each Gaussian contains position parameters to define its spatial coordinates, rotation parameters to describe its orientation, scale parameters to control its ellipsoidal shape, opacity parameters to adjust its visibility and contribution weight, and color attributes to represent its appearance features. This serves as the basic representation for subsequent geometry-aware optimization and pseudo-view collaboration.

[0023] S102, Perform geometry-aware Gaussian optimization based on the initial 3D Gaussian representation to generate a geometry-optimized 3D Gaussian representation. The optimization includes embedding outliers into the scene surface through a depth stamping surface reconstruction module, adjusting the shape of the Gaussian ellipsoid through an equalization shape constraint module, and controlling the visibility of the Gaussian through an opacity constraint module. In this embodiment, outlier Gaussians are first embedded into the scene surface using a depth-stamped surface reconstruction module. This module freezes the opacity parameter when rendering the initial 3D Gaussian representation and assumes all Gaussians are nearly opaque, which helps to highlight and identify outlier Gaussians floating at the front. Simultaneously, a fine-grained depth map is predicted from the training view using the diffusion-based monocular depth estimation model Lotus. Then, the scene depth is rendered using the 3D Gaussian representation through the 3DGS rendering pipeline. The scene surface is regularized by minimizing the Pearson correlation loss, forcing the 3D Gaussian representation to converge to the real scene surface. The Pearson loss function is defined as follows: , It is a very small positive number used to prevent division by zero errors in numerical calculations. This constraint effectively promotes the reconstruction quality of scene surface geometry.

[0024] Next, the Gaussian ellipsoid shape is adjusted through the equilibrium shape constraint module, which constrains the scale parameter of each Gaussian in the 3D Gaussian representation. Scale difference regularization is used to limit the maximum and minimum principal axis scale ratios, while encouraging the generation of smaller Gaussians to improve the quality of high-precision texture reconstruction. The specific loss function is... ,in To balance the loss due to morphological constraints, The weight of the first regularization term. The weight of the second regularization term, The number of Gaussian ellipsoids, Let be the scale vector of the i-th Gaussian ellipsoid. This mechanism suppresses excessive anisotropy of the Gaussian ellipsoid, thereby reducing the risk of geometric distortion and overfitting.

[0025] Finally, the visibility of Gaussians is controlled by an opacity constraint module, which controls the opacity of each Gaussian in the 3D Gaussian representation. Apply suppression constraints, and the loss function is: ,in, For the loss due to opacity constraints, The weight of the opacity constraint term. The number of Gaussian ellipsoids is determined, and a low-opacity pruning strategy is used to remove Gaussians with opacity below a threshold. This optimizes the spatial distribution of the 3D Gaussian representation, effectively controls the generation of useless Gaussians and floating artifacts, and reduces spatial costs. The entire geometry-aware optimization process, through the synergistic effect of these modules, ensures a smooth transition from the initial representation to the geometry-optimized representation, providing a reliable geometric basis for subsequent pseudo-viewpoint regularization.

[0026] S103, apply dual-source collaborative pseudo-viewpoint regularization to the geometrically optimized 3D Gaussian representation to obtain a pseudo-viewpoint enhanced 3D Gaussian representation. The pseudo-viewpoint regularization includes generating a camera pose of a pseudo-training viewpoint from the existing training viewpoint through a viewpoint interpolation dynamic scoring mechanism, rendering a pseudo-viewpoint image and depth using the 3D Gaussian representation under the pseudo-training viewpoint, simultaneously projecting the original training viewpoint onto the pseudo-training viewpoint through a warp transformation to generate a pseudo-ground value, and updating the 3D Gaussian representation by minimizing the difference between the rendering result and the pseudo-ground value. First, a pseudo-training viewpoint is generated from the existing training viewpoints using a viewpoint interpolation dynamic scoring mechanism. This mechanism calculates the score between the candidate viewpoints and the current training viewpoint from the set of training viewpoints represented by the geometrically optimized 3D Gaussian representation. The scoring formula is as follows: ,in For Euclidean distance, For dynamic scoring, To control the weights of the space-angle constraint balance, To control the weights of spatial distance constraints, To address the difference angle between viewing directions exceeding a threshold The penalty weights and Euclidean distances exceed the mean. The penalty weight, For indicator functions, The difference angle is the viewpoint direction difference angle, which is derived through the Frobinius inner product of the relative rotation matrix. It should be noted that the dynamic scoring integrates spatial distance and viewpoint direction similarity to ensure a high overlap of the visible area between the pseudo-viewpoint and the original training viewpoint, and to address angles exceeding a threshold. Candidates whose distance exceeds a threshold are subject to an exponential penalty to achieve uniform spatial coverage and efficient matching of pseudo-viewpoints. Then, a Top-K random sampling strategy is used to select matching targets from the K candidate viewpoints with the lowest scores, and interpolation is performed between the current viewpoint and the selected candidate. The position component is obtained through linear interpolation. ,in, The position components of the newly generated camera pose. For the position components of the current training viewpoint, To match the position components of the target viewpoint, These are the interpolation coefficients for the position components; Rotational components via spherical linear interpolation ,in, For the rotation component of the newly generated camera pose, The rotation component is the current training viewpoint. To match the rotation component of the target viewpoint, For the interpolation coefficients of the rotation component, The angle of difference in view direction is defined, where the interpolation coefficients are randomly sampled within a specified range to enhance sampling continuity, ultimately generating a new camera pose. ,in, For the newly generated camera pose; Next, using the geometrically optimized 3D Gaussian representation, this embodiment combines the training viewpoint scene rendering depth. and , train view pixels and depth information Warp transformation to The corresponding new view and In Chinese, it is expressed as: ,in and These represent projection from 3D point cloud to 2D and backprojection from 2D to 3D point cloud, respectively.

[0027] The same method can be used to obtain... corresponding and To eliminate photometric errors caused by warping transformation, a photometric confidence filter is proposed. Its expression is: ,in The confidence threshold; The simulated value is obtained by weighting the rendered RGB values ​​and depth values ​​from the two training views. and :

[0028]

[0029]

[0030] in Here are the camera center coordinates. To ensure geometric consistency across multiple viewpoints, a depth residual filter is introduced. :

[0031] in Using the depth residual threshold, the RGB pseudo-real values ​​are finally obtained from the pseudo-training perspective. and depth pseudo-true value :

[0032]

[0033] This embodiment regularizes the model by optimizing the L2 norm loss value between the pseudo-true value and the rendered value in the effective region M under the pseudo-training perspective:

[0034]

[0035] in , Indicates false truth value and In pixels Color and depth at the location; and Represents the pixels of the scene in a pseudo-view. The color and depth of the rendered area.

[0036] This embodiment introduces a monocular depth estimation model (DPT) to quickly predict depth information from rendered images under pseudo-training perspectives. Use Pearson loss to constrain the rendering of depth information. : .

[0037] S104, apply the neighborhood smoothing module to the pseudo-viewpoint enhanced 3D Gaussian representation, construct a neighborhood graph by analyzing the spatial relationship between Gaussians, implement photometric smoothing constraints and depth smoothing constraints, and generate the final 3D Gaussian representation for new view synthesis.

[0038] It should be noted that this embodiment introduces a photometric smoothing strategy based on a neighboring graph and an adaptive depth smoothing strategy, which aim to solve the color jump artifacts and non-physical fluctuations, such as floating outliers and surface fractures, caused by the anisotropic Gaussian distribution of 3DGS discretization.

[0039] A photometric smoothing strategy based on the nearest neighbor graph, with each Gaussian The set of Gaussian neighbors with the head node and the K nearest neighbors by Euclidean distance. Medium Gaussian Construct directed edges for the tail node Characterizing the local photometric propagation path, using Controlling sensitivity to spatial distance forces a smooth color transition between adjacent Gaussian pairs:

[0040] in, For photometric smoothing loss based on nearest neighbor graph, for The set of Gaussian neighbors of the head node and its K-nearest Euclidean distance. Let represent the center positions of the i-th and j-th Gaussians, respectively. Parameters used to control sensitivity to spatial distance.

[0041] The adaptive depth smoothing strategy utilizes the gradient of depth information estimated by a monocular depth pre-trained model to adaptively constrain the rendering depth smoothness:

[0042] in, For adaptive depth smoothing loss, and These represent the gradients of the depth information estimated by the monocular depth pre-trained model in the x and y directions, respectively. and These represent the gradients of the rendering depth in the x and y directions, respectively.

[0043] It should be noted that this embodiment provides a progressive scene reconstruction method, which progresses from coarse-grained pre-training to fine-grained reconstruction of high-frequency information, and from targeted optimization of scene geometry to pseudo-viewpoint consistency regularization. By balancing external prior knowledge with full utilization of training set information, it effectively achieves sparse viewpoint reconstruction. Figure 3 The high-precision reconstruction of the 3D scene alleviates the overfitting problem of the original 3DGS.

[0044] Table 1. Comparison metrics on the LLFF dataset

[0045] Table 2. Comparison metrics on the TNT dataset

[0046] Table 3. Comparison metrics on the MipNeRF360 dataset

[0047] Structural similarity (SSIM), peak signal-to-noise ratio (PSNR), and image similarity (LPIPS) were used to measure the quality of model reconstruction and novel attempt synthesis. The quantitative results in Tables 1, 2, and 3 are presented below. Figure 2 and Figure 3 The visualization results show that this invention performs excellently on various datasets, while the comparative methods all exhibit varying degrees of defects. 3DGS, due to its discrete nature, shows severe overfitting. DNGaussian introduces depth priors for global and local regularization, but due to scale errors in monocular depth estimation and optimization only under the original training viewpoint, it does not effectively address the overfitting problem, still exhibiting numerous artifacts and geometric distortions in the depth map. FSGS shows overfitting and texture loss in complex texture areas because the pseudo-viewpoint depth supervision information relies excessively on high-error rendering results and the accuracy of the monocular depth estimation model. CoR-GS, due to the mutual constraint between the two Gaussian models, gets stuck in local optima in complex texture areas, exhibiting blurred artifacts with smooth color; this constraint is fatal when both models have poor pseudo-viewpoint quality. This invention makes targeted optimizations to geometry and pseudo-viewpoints, significantly improving geometric reconstruction quality and achieving remarkable artifact optimization.

[0048] The second embodiment of the present invention provides a geometry-aware sparse view for pseudo-view collaboration. Figure 3 DGS reconstruction apparatus, characterized in that it comprises: An initial Gaussian representation generation unit is used to obtain the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters. The 3D scene to be reconstructed is initialized by the Gaussian unpooling densification method to generate an initial 3D Gaussian representation, wherein each Gaussian contains position, rotation, scale, opacity and color attributes. The Gaussian representation optimization unit is used to perform geometry-aware Gaussian optimization based on the initial 3D Gaussian representation to generate a geometry-optimized 3D Gaussian representation. The optimization includes embedding outliers into the scene surface through a depth stamping surface reconstruction module, adjusting the shape of the Gaussian ellipsoid through an equalization shape constraint module, and controlling the visibility of the Gaussian through an opacity constraint module. A pseudo-viewpoint regularization unit is used to perform dual-source collaborative pseudo-viewpoint regularization on the geometrically optimized 3D Gaussian representation to obtain a pseudo-viewpoint enhanced 3D Gaussian representation. The pseudo-viewpoint regularization includes generating a camera pose of a pseudo-training viewpoint from the existing training viewpoint through a viewpoint interpolation dynamic scoring mechanism, rendering a pseudo-viewpoint image and depth using the 3D Gaussian representation under the pseudo-training viewpoint, simultaneously projecting the original training viewpoint onto the pseudo-training viewpoint through a warp transformation to generate a pseudo-ground value, and updating the 3D Gaussian representation by minimizing the difference between the rendering result and the pseudo-ground value. A smoothing unit is used to apply a neighborhood smoothing module to the pseudo-viewpoint-enhanced 3D Gaussian representation. By analyzing the spatial relationships between Gaussians, a neighborhood graph is constructed, and photometric and depth smoothing constraints are implemented to generate the final 3D Gaussian representation for new view synthesis.

[0049] The third embodiment of the present invention provides a geometry-aware sparse view for pseudo-view collaboration. Figure 3 The DGS reconstruction device is characterized by comprising a memory and a processor, wherein the memory stores a computer program that can be executed by the processor to achieve a pseudo-view cooperative geometry-aware sparse view as described in any of the above claims. Figure 3 DGS reconstruction method.

[0050] The fourth embodiment of the present invention provides a computer-readable storage medium, characterized in that it stores a computer program, the computer program being executable by a processor of the device in which the computer-readable storage medium is located, to implement a pseudo-view cooperative geometry-aware sparse view as described in any of the preceding claims. Figure 3 DGS reconstruction method.

[0051] Based on the present invention, a geometry-aware sparse view is used for pseudo-view collaboration. Figure 3The DGS reconstruction method, apparatus, equipment, and storage medium initialize the Gaussian representation of the 3D scene through Gaussian solution pooling, followed by geometry-aware optimization. A depth-stamped surface reconstruction module forcibly embeds outliers from the scene surface into their correct positions. Equilibrium shape constraints prevent excessive stretching and deformation of the Gaussian ellipsoid, and opacity constraints eliminate redundant Gaussians, thus establishing an accurate scene geometry. Next, dual-source collaborative pseudo-viewpoint regularization is implemented. A pseudo-training viewpoint is intelligently generated through a viewpoint interpolation dynamic scoring mechanism. Under this viewpoint, bidirectional constraints are applied using both the current Gaussian rendering and pseudo-ground values ​​generated from the original view's warp transformation. Finally, a neighborhood smoothing module is applied to construct a neighborhood graph based on Gaussian spatial proximity relationships. Luminosity and depth smoothing constraints eliminate color abruptness and geometric discontinuities caused by discrete Gaussian distributions.

[0052] Exemplary, the computer program described in the third and fourth embodiments of the present invention can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, wherein the instruction segments describe the computer program's implementation of a pseudo-view cooperative geometrically perceptual sparse view. Figure 3 The execution process in the DGS reconstruction device. For example, the apparatus described in the second embodiment of the present invention.

[0053] The processor referred to can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and the processor is the aforementioned pseudo-view cooperative geometric perception sparse view. Figure 3 The control center of the DGS reconstruction method utilizes various interfaces and lines to connect the entire system, enabling geometric perception of sparse views in a pseudo-view collaboration. Figure 3 The various parts of the DGS reconstruction method.

[0054] The memory can be used to store the computer program and / or modules. The processor, by running or executing the computer program and / or modules stored in the memory, and by calling data stored in the memory, implements a pseudo-view cooperative geometrically perceptive sparse view. Figure 3 The DGS reconstruction method includes various functions. The memory primarily comprises a program storage area and a data storage area. The program storage area stores the operating system and at least one application program required for a given function (e.g., sound playback, text conversion, etc.). The data storage area stores data created based on the phone's usage (e.g., audio data, text message data, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0055] If the implemented module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0056] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0057] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A pseudo-view collaborative geometry-aware sparse view 3DGS reconstruction method, characterized in that, include: Obtain the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters. Initialize the 3D scene to be reconstructed using the Gaussian pooling densification method to generate an initial 3D Gaussian representation, wherein each Gaussian contains position, rotation, scale, opacity and color attributes. Based on the initial 3D Gaussian representation, geometrically aware Gaussian optimization is performed to generate a geometrically optimized 3D Gaussian representation. The optimization includes embedding outliers into the scene surface through a depth stamping surface reconstruction module, adjusting the shape of the Gaussian ellipsoid through an equalization shape constraint module, and controlling the visibility of the Gaussian through an opacity constraint module. A dual-source collaborative pseudo-viewpoint regularization is applied to the geometrically optimized 3D Gaussian representation to obtain a pseudo-viewpoint enhanced 3D Gaussian representation. The pseudo-viewpoint regularization includes generating a camera pose of a pseudo-training viewpoint from the existing training viewpoint through a viewpoint interpolation dynamic scoring mechanism, rendering a pseudo-viewpoint image and depth using the 3D Gaussian representation under the pseudo-training viewpoint, simultaneously projecting the original training viewpoint onto the pseudo-training viewpoint through a warp transformation to generate a pseudo-ground value, and updating the 3D Gaussian representation by minimizing the difference between the rendering result and the pseudo-ground value. The neighborhood smoothing module is applied to the pseudo-viewpoint enhanced 3D Gaussian representation. By analyzing the spatial relationships between Gaussians, a neighborhood graph is constructed, and photometric smoothing constraints and depth smoothing constraints are implemented to generate the final 3D Gaussian representation for new view synthesis.

2. The method for reconstructing geometry-aware sparse views using pseudo-view collaboration according to claim 1, characterized in that, The process of embedding outlier Gaussians into the scene surface using the deep stamping surface reconstruction module is as follows: When rendering the initial 3D Gaussian representation, the opacity parameter is frozen, assuming that all Gaussians are close to opaque in order to identify outliers. Predict fine-grained depth maps from training views using the diffusion-based monocular depth estimation model Lotus. ; The scene depth is rendered using the 3D Gaussian representation through the 3DGS rendering pipeline. ; The scene surface is regularized by minimizing the Pearson correlation loss, forcing the 3D Gaussian representation to converge to the real scene surface. Its expression is: ,in, This indicates a deep loss.

3. The method for reconstructing geometry-aware sparse views using pseudo-view collaboration according to claim 1, characterized in that, The adjustment of the Gaussian ellipsoid shape through the equilibrium shape constraint module is specifically as follows: The scale parameter of each Gaussian in the 3D Gaussian representation is constrained. The maximum and minimum principal axis scale ratios are limited by scale difference regularization, while small Gaussians are encouraged to improve texture reconstruction quality. The expression is as follows: ,in, To balance the loss due to morphological constraints, The weight of the first regularization term. The weight of the second regularization term, The number of Gaussian ellipsoids, Let be the scale vector of the i-th Gaussian ellipsoid.

4. The method for reconstructing geometry-aware sparse views using pseudo-view collaboration according to claim 1, characterized in that, The control of Gaussian visibility through the opacity constraint module is specifically as follows: The opacity of each Gaussian in the three-dimensional Gaussian representation Applying a suppression constraint, its expression is: ,in, For the loss constrained by opacity, The weight of the opacity constraint term. The number of Gaussian ellipsoids; By employing a low-opacity pruning strategy, Gaussians with opacities below a threshold are removed, thereby optimizing the spatial distribution of the three-dimensional Gaussian representation.

5. The method for geometrically aware sparse view 3DGS reconstruction based on pseudo-view collaboration according to claim 1, characterized in that, The dynamic scoring mechanism specifically includes: From the set of training viewpoints represented by the geometrically optimized 3D Gaussian representation, the score between the candidate viewpoint and the current training viewpoint is calculated, and its expression is: ,in, For dynamic scoring, To control the weights of the space-angle constraint balance, To control the weights of spatial distance constraints, To address the difference angle between viewing directions exceeding a threshold The penalty weights and Euclidean distances exceed the mean. The penalty weight, For indicator functions, The angle of difference in viewing direction; Select the K candidate viewpoints with the lowest scores, and interpolate between them and the current viewpoint. Linear interpolation of position components: ,in, The position components of the newly generated camera pose. For the position components of the current training viewpoint, To match the position components of the target viewpoint, These are the interpolation coefficients for the position components; Rotational component spherical linear interpolation: ,in, For the rotation component of the newly generated camera pose, The rotation component is the current training viewpoint. To match the rotation component of the target viewpoint, For the interpolation coefficients of the rotation component, The angle of difference in viewing direction; Generate new camera pose: ,in, This is the newly generated camera pose.

6. The method for reconstructing geometry-aware sparse views using pseudo-view collaboration according to claim 1, characterized in that, The photometric smoothing constraint includes: For each Gaussian in the pseudo-viewpoint-enhanced 3D Gaussian representation A neighborhood graph is constructed by finding K nearest neighbor Gaussians centered on the given location, and the color transition between adjacent Gaussians is constrained by the following loss function: in, For photometric smoothing loss based on nearest neighbor graph, for The set of Gaussian neighbors with head node and Euclidean distance K-nearest neighbors. Let represent the center positions of the i-th and j-th Gaussians, respectively. Parameters used to control sensitivity to spatial distance.

7. The method for geometrically aware sparse view 3DGS reconstruction based on pseudo-view collaboration according to claim 1, characterized in that, The depth smoothing constraint includes: The depth gradient is estimated using a monocular depth pre-trained model, and an adaptive constraint is applied to the depth gradient rendered by the 3D Gaussian representation: in, For adaptive depth smoothing loss, and These represent the gradients of the depth information estimated by the monocular depth pre-trained model in the x and y directions, respectively. and These represent the gradients of the rendering depth in the x and y directions, respectively.

8. A pseudo-view collaborative geometry-aware sparse view 3DGS reconstruction device, characterized in that, include: An initial Gaussian representation generation unit is used to obtain the sparse view dataset of the 3D scene to be reconstructed and the corresponding camera pose parameters. The 3D scene to be reconstructed is initialized by the Gaussian unpooling densification method to generate an initial 3D Gaussian representation, wherein each Gaussian contains position, rotation, scale, opacity and color attributes. The Gaussian representation optimization unit is used to perform geometry-aware Gaussian optimization based on the initial 3D Gaussian representation to generate a geometry-optimized 3D Gaussian representation. The optimization includes embedding outliers into the scene surface through a depth stamping surface reconstruction module, adjusting the shape of the Gaussian ellipsoid through an equalization shape constraint module, and controlling the visibility of the Gaussian through an opacity constraint module. A pseudo-viewpoint regularization unit is used to perform dual-source collaborative pseudo-viewpoint regularization on the geometrically optimized 3D Gaussian representation to obtain a pseudo-viewpoint enhanced 3D Gaussian representation. The pseudo-viewpoint regularization includes generating a camera pose of a pseudo-training viewpoint from the existing training viewpoint through a viewpoint interpolation dynamic scoring mechanism, rendering a pseudo-viewpoint image and depth using the 3D Gaussian representation under the pseudo-training viewpoint, simultaneously projecting the original training viewpoint onto the pseudo-training viewpoint through a warp transformation to generate a pseudo-ground value, and updating the 3D Gaussian representation by minimizing the difference between the rendering result and the pseudo-ground value. A smoothing unit is used to apply a neighborhood smoothing module to the pseudo-viewpoint-enhanced 3D Gaussian representation. By analyzing the spatial relationships between Gaussians, a neighborhood graph is constructed, and photometric and depth smoothing constraints are implemented to generate the final 3D Gaussian representation for new view synthesis.

9. A pseudo-view collaborative geometry-aware sparse view 3DGS reconstruction device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program that can be executed by the processor to implement a pseudo-view collaborative geometry-aware sparse view 3DGS reconstruction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The device contains a computer program that can be executed by a processor of the device in which the computer-readable storage medium is located, to implement the pseudo-view cooperative geometry-aware sparse view 3DGS reconstruction method as described in any one of claims 1 to 7.