Geometric consistency carving and gradient conflict driven densification of gaussian reconstruction method

By employing a Gaussian reconstruction method that combines geometric consistency sculpting with gradient conflict-driven densification, the problem of inaccurate reconstruction under sparse perspectives is solved, achieving efficient and robust 3D scene reconstruction and rendering, suitable for virtual reality, augmented reality, and digital twin applications.

CN121120951BActive Publication Date: 2026-02-27CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511651242.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-27
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Under sparse view conditions, existing technologies such as NeRF and 3DGS have inaccurate geometric reconstruction and lighting estimation, resulting in reconstruction degradation and decreased rendering quality. Furthermore, when relying on monocular depth estimation, depth prediction is distorted, making it difficult to achieve depth consistency across multiple viewpoints.

Method used

A Gaussian reconstruction method using geometric consistency sculpting and gradient conflict-driven densification is adopted. The scene image is analyzed by a multimodal depth estimation network, and the Gaussian sculpting is carried out using geometric consistency loss. Combined with gradient conflict-driven densification, the fine splitting and replication of Gaussian splatter model is achieved, generating a Gaussian splatter model.

Benefits of technology

It significantly improves the robustness and accuracy of 3D modeling under sparse perspective, reduces computational complexity, supports efficient view composition and rendering, and enhances its application value in virtual reality, augmented reality, and digital twins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120951B_ABST
    Figure CN121120951B_ABST
Patent Text Reader

Abstract

The present application provides a kind of geometric consistency engraving and gradient conflict driven densification gauss reconstruction method, it is related to image reconstruction technical field, the method is to obtain scene image training data;Multi-modal depth estimation network is used to analyze scene image training data, obtain multi-modal depth hypothesis data;Based on scene image training data, spatial gauss depth is sampled, and synthetic view multi-modal depth prediction is obtained;Based on synthetic view multi-modal depth prediction and multi-modal depth hypothesis data, using geometric consistency loss engraving, gauss is engraved in space, and gauss spatial engraving result is obtained;Gauss gradient conflict is used to drive gauss densification, and gauss densification result is obtained;Based on gauss spatial engraving result and gauss densification result, by calculation, obtain gauss splatter model result.The present application solves the problem that the object reconstruction of depth blur appears gauss depth is not accurate and geometric distortion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of image reconstruction, and in particular, to a Gaussian reconstruction method of geometric consistency carving and gradient conflict driven densification. BACKGROUND

[0002] In recent years, the representation and rendering technology of three-dimensional scene has been rapidly developed, especially the introduction of NeRF (Neural Radiance Fields), which effectively alleviates the problems of low efficiency and poor rendering quality in traditional methods. NeRF introduces deep learning mechanism, based on multi-view 2D pixel sampling and coding, optimizes the implicit scene representation function, and realizes high-quality scene reconstruction and new view synthesis. However, its MLP neural network has a high parameter quantity of millions, and in tens of thousands of iterations, high-density spatial sampling of 1024 rays and particle sampling of 68-198 on each ray are performed, resulting in high consumption of computing resources and time cost in rendering and training stages. Based on the theory of neural radiance field, 3DGS learns the geometric and appearance features of the scene through explicit three-dimensional Gaussian representation, and realizes efficient scene reconstruction and new view synthesis by combining CUDA multi-thread parallel rendering. However, under the condition of sparse view, NeRF and 3DGS cause inaccurate geometry reconstruction and lighting estimation due to insufficient view information, and the visual error and reprojection error are aggravated. It is difficult to fully learn the 3D details and structure of the scene, resulting in problems such as reconstruction degradation and rendering quality decline. To solve this problem, various improvement schemes have been proposed, but there are still the following shortcomings: 1) Based on NeRF, three-dimensional scene representation and reconstruction under sparse view are realized. Generally, monocular depth estimation is used as a geometric prior, which is prone to distortion in dealing with reflection, refraction or transparent medium, further amplifying the geometric regularization error and causing serious structural reconstruction bias. At the same time, the random virtual view sampling strategy used by some methods lacks modeling of real camera trajectories or occlusion distribution, and the guiding effect is poor at extreme elevation or edge view. In addition, the above methods often require large-scale pixel pair ordering, depth map reconstruction and high-dimensional prior modeling, increasing the computational cost of the training process, which seriously restricts the efficiency and actual deployment ability of the method. 2) Based on 3DGS, three-dimensional scene representation and reconstruction under sparse view are realized. First, the gradient update directions of pixels from different views to the same Gaussian position (μ) and opacity (α) are conflicting, resulting in optimization competition phenomenon in the gradient backpropagation process: some pixels push the Gaussian forward along the sight line, while others push it backward. The conflict between the gradient contributions of these pixels makes the Gaussian position oscillate in the wrong geometric area, making it difficult to converge to a stable solution. Second, in the light-transmitting or light-reflecting scene, the nonlinear shift of light path and the inconsistency of multi-view observation directly destroy the multi-view geometric and photometric consistency assumption, making the traditional depth-guided geometric optimization strategy ineffective; at the same time, the complex reflection and refraction effect makes it difficult for color information to provide reliable guidance for the true geometry, causing essential conflicts between the geometry and appearance optimization objectives, and further leading to problems such as geometric reconstruction error increasing, depth prior invalidation and misguidance, and color supervision noise increasing.The last part method still relies on monocular depth estimation, and cannot achieve depth consistency in multi-view in a depth ambiguous scene, thereby causing Gauss to be gathered on a wrong depth plane, and further aggravating geometric distortion of reconstruction. SUMMARY

[0003] In view of the above problems in the prior art, the geometric consistency carving and gradient conflict driven densification Gaussian reconstruction method provided by the present application solves the problems of inaccurate Gaussian depth and geometric distortion in the reconstruction of a depth ambiguous object.

[0004] In order to achieve the above-mentioned purposes, the technical scheme adopted by the present application is as follows: a geometric consistency carving and gradient conflict driven densification Gaussian reconstruction method, comprising:

[0005] S1: acquiring scene image training data; the scene point cloud training data comprises object image data of a depth ambiguous object;

[0006] S2: analyzing the scene image training data by using a multi-modal depth estimation network to obtain multi-modal depth hypothesis data;

[0007] S3: based on the scene image training data, initializing Gauss, sampling the spatial Gauss depth, and obtaining synthetic view multi-modal depth prediction;

[0008] S4: based on the synthetic view multi-modal depth prediction and the multi-modal depth hypothesis data, using geometric consistency loss carving to carve the Gauss in space to obtain a Gauss spatial carving result; using Gauss gradient conflict to drive the Gauss to be densified to obtain a Gauss densification result;

[0009] S5: based on the Gauss spatial carving result and the Gauss densification result, obtaining a Gauss splashing model result by calculation to complete the Gaussian reconstruction of the sparse view.

[0010] The present application has the beneficial effects that: the present application provides a geometric consistency carving and gradient conflict driven densification Gaussian reconstruction method, multi-modal depth estimation network is used to analyze scene image training data, and multi-modal depth hypothesis data is obtained; based on the scene image training data, the spatial Gaussian depth is sampled to obtain the synthetic view multi-modal depth prediction; based on the synthetic view multi-modal depth prediction; and the multi-modal depth hypothesis data, the geometric consistency loss carving is used to obtain the Gaussian space carving result; at the same time of obtaining the Gaussian space carving result, the gradient conflict driven densification is used for processing, and the Gaussian reconstruction of the sparse view is completed.(1) The gradient conflict driven densification is used for fine splitting and copying of local Gaussian particles, dynamic modeling of occlusion boundary, texture details and small scale structure is realized, and the integrity and detail restoration ability of the final reconstruction result are significantly enhanced.(2) The geometric consistency carving is used to realize the automatic 'carving' and removal of the error Gaussian in continuous iteration, without relying on the hard threshold judgment or post-processing screening.Finally, the distribution of Gaussian particles in the scene gradually tends to be reasonable and sparse, avoiding the pseudo structure caused by the error initial modeling or sparse input, and improving the structure clarity and representation simplicity of the model.(3) The multi-modal depth estimation network is used to make the Gaussian distribution gradually converge to at least one feasible geometric explanation, especially in the area with high depth uncertainty caused by reflection, texture missing or transparent object, the geometric structure continuity and consistency can still be maintained, and the robustness and accuracy of three-dimensional modeling under sparse view angle are significantly improved.(4) Without relying on traditional global dense matching or large-scale optimization framework, stable convergence can be realized under limited computing resources, and the computational complexity and training overhead are significantly reduced.(5) Since the Gaussian distribution can be directly combined with the rendering framework, the view synthesis and rendering can be efficiently supported while completing the geometric reconstruction, and the application value in virtual reality, augmented reality and digital twin is improved.

[0011] Further, the S1 comprises:

[0012] Based on the image-based RGB dataset, the camera pose training data is obtained by structure reconstruction of motion;

[0013] When the three-dimensional reconstruction toolkit fails to reconstruct the camera pose training data, the dense stereo matching function is called to continue depth estimation, and the depth maps of each view are fused into a dense point cloud by stereo vision fusion to obtain the scene image training data.

[0014] (1) By introducing the dynamic modulation mechanism of the style transfer layer, the network can maintain the consistency and stability of depth estimation when facing different scenes and different lighting conditions; (2) The generation mode of K depth hypothesis samples effectively increases the diversity of depth distribution, so that the subsequent carving link can fully utilize the redundant information and avoid being limited by single depth interpretation; (3) Compared with the scheme relying on single-mode depth prediction, the method has better generalization ability and good adaptability to cross-domain data (such as outdoor and indoor, artificial and natural scenes).

[0015] Further, the S2 comprises:

[0016] A pre-trained multi-modal depth estimation network is used as the core analysis model. The scene image training data RGB image is input into the multi-modal depth estimation network, and after the feature extraction by the LeReS backbone network in the multi-modal depth estimation network, dynamic modulation is performed by the multiple style transfer layers to generate K depth hypothesis samples.

[0017] The multi-modal depth hypothesis data corresponding to the pixels of N positions of each image in the depth hypothesis sample is sampled to obtain N*K multi-modal depth hypothesis data.

[0018] (1) By iteratively sampling the Gaussian depth at the same pixel position, the consistency between the synthesized view prediction and the multi-modal hypothesis is ensured; (2) This mechanism automatically filters out unreasonable Gaussian depths during the training process, so that the spatial distribution of the Gaussian gradually fits the real scene geometry; (3) It can maintain the global continuity of depth prediction under limited viewing angles, effectively avoiding the problems of "suspended Gaussian" and geometric discontinuity.

[0019] Further, the S3 comprises:

[0020] The Gaussian is initialized, and the same N pixel positions as the multi-modal depth hypothesis are selected. The ray from the pixel to the training view center of the scene image training data passes through the depth of all Gaussians to obtain the synthesized view multi-modal depth prediction, i.e., the p Gaussian prediction depth data of N pixels in each iteration.

[0021] (1) Through multi-view iterative optimization, the Gaussian is continuously adjusted in position and opacity, making the spatial distribution sparse and reasonable while preserving texture, boundaries, and small-scale structural details, and improving the clarity of geometric structures; (2) The Gaussian depth is regularized using geometric consistency loss, effectively reducing the influence of abnormal Gaussians on the overall reconstruction, and making the Gaussian distribution gradually converge to a reasonable global consistent depth; (3) In complex scenes (such as occlusion, reflection, texture loss, or transparent objects), the Gaussian distribution can still maintain coherence and stability, improving the accuracy and robustness of three-dimensional reconstruction under sparse viewing angles.

[0022] Further, the S4 comprises:

[0023] Based on the synthesized view multi-modal depth prediction and multi-modal depth hypothesis data, the depth distribution of the Gaussian ball in the three-dimensional space is geometrically consistent regularized by using a geometric consistency loss, to obtain an overall space carving loss;

[0024] The Gaussian ball depth error is soft constrained by using the overall space carving loss, and in multiple iterations, the Gaussian under different viewing angles is continuously optimized in position and opacity, to find the consistent depth of the Gaussian, and obtain a Gaussian space carving result.

[0025] (1) The overall space carving loss function proposed jointly constrains the Gaussian depth at the single-pixel scale and the global scale, so that each Gaussian gradually converges to a reasonable global consistent depth in the entire scene; (2) By optimizing the minimum distance between the predicted depth of each Gaussian and the multi-modal depth hypothesis, unreasonable depth explanations are automatically excluded, and the generation of pseudo-structures and local abnormal depths is reduced; (3) Compared with the method that only relies on Gaussian projection error, this design can handle occlusion, multi-view depth conflict and uncertainty under sparse view in complex scenes, thereby ensuring the consistency and stability of the Gaussian depth globally.

[0026] Further, the expression of the overall space carving loss is:

[0027] ;

[0028] ;

[0029] ;

[0030] wherein, the overall space carving loss is represented by Lsc, the number of pixels is represented by N, the geometric consistency error of a single pixel is represented by e, the number of Gaussians corresponding to a single pixel is represented by k, the minimum distance between the predicted depth of the Gaussian and all hypotheses is represented by d, the minimum value among all k depth hypotheses corresponding to the pixel is represented by min, the predicted depth of the i-th pixel given by the j-th visible Gaussian is represented by di,j, the m-th depth hypothesis of the i-th pixel is represented by di,m, m represents the multi-depth hypothesis index, and k represents the number of multi-depth hypotheses, the absolute value is represented by | |.

[0031] (1) A Gaussian densification mechanism driven by gradient conflict can gradually refine the geometry during training, significantly reducing the dependence on manually set hyperparameters; (2) By jointly constraining opacity and position gradient, the newly generated Gaussians can be adaptively distributed in the areas with the most detailed geometry, thereby improving the effectiveness of densification; (3) The number and distribution of Gaussians can be dynamically adjusted in multiple iterations, and the final model balances between density and simplicity, ensuring reconstruction accuracy while avoiding redundant calculations.

[0032] Further, the S5 comprises:

[0033] Based on the Gaussian space carving result and the Gaussian densification result, for each Gaussian under the current iteration training view, the pixel set covered by the Gaussian is counted, and the opacity gradient and position gradient of each pixel to the Gaussian are extracted, respectively, to obtain the sign feature and gradient amplitude corresponding to the opacity gradient and position gradient;

[0034] Based on the gradient conflict driven densification, the sign feature and gradient amplitude corresponding to the opacity gradient and position gradient are calculated to obtain the opacity conflict score and the position conflict score;

[0035] Based on the opacity conflict score and the position conflict score, the two-dimensional gradient features of the covered pixels are extracted, and a weighted clustering method is used for clustering to obtain a clustering result;

[0036] The clustering center of the clustering result is projected back to the three-dimensional space to obtain a conversion result;

[0037] Based on the conversion result, the new Gaussian is copied and split, and the Gaussian splashing model result is obtained through multiple iterations to complete the Gaussian reconstruction of the sparse view.

[0038] (1) By quantitatively defining the conflict score, the method can quantify the contribution conflict degree of Gaussians on their covered pixels, providing a reliable basis for Gaussian splitting and replication, thereby effectively improving the accuracy of Gaussian modeling under sparse view (2) Combined with the weighted clustering and back projection mechanism, the two-dimensional gradient conflict information can be accurately mapped to the three-dimensional space, improving the spatial positioning accuracy of Gaussian replication and splitting; (3) The final Gaussian splashing model is closer to the real surface structure, ensuring the integrity of the reconstruction while maintaining high rendering efficiency.

[0039] Further, the expression of the opacity conflict score is:

[0040] ;

[0041] wherein, the opacity conflict score is represented by, the i-th Gaussian is represented by, represents the jth pixel in the Gauss projection region, represents a pixel set, represents the sign feature of the Gauss opacity, represents the gradient amplitude of the Gauss opacity;

[0042] The expression of the position conflict score is:

[0043] ;

[0044] wherein, represents the position conflict score, represents the sign feature of the position gradient, represents the gradient amplitude of the position gradient.

[0045] Further, the expression of the clustering result is:

[0046] ;

[0047] wherein, represents the center feature of the kth clustering cluster, represents the clustering operation, represents the jth pixel feature, represents the clustering weight of each pixel feature, represents the number of clustering clusters, represents the absolute value of the ith Gauss opacity gradient corresponding to the jth pixel, represents the absolute value of the ith Gauss position gradient corresponding to the jth pixel.

[0048] Further, the expression of the conversion result is:

[0049] ;

[0050] wherein, represents the conversion result, represents the projection from the camera coordinate system to the world coordinate system, represents the horizontal coordinate of the clustering center, represents the vertical coordinate of the clustering center, represents the two-dimensional coordinate of the clustering center, represents the depth value. BRIEF DESCRIPTION OF DRAWINGS

[0051] The present specification will be further illustrated in the manner of exemplary embodiments, which will be described in detail through the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same numbers represent the same structures, wherein:

[0052] Figure 1 is an exemplary flowchart of a geometry consistency carving and gradient conflict driven densification Gaussian reconstruction method according to some embodiments of the present specification;

[0053] Figure 2 is an exemplary schematic diagram of obtaining multi-depth hypothesis data by a multi-modal depth estimation network according to some embodiments of the present specification;

[0054] Figure 3 is an exemplary schematic diagram of multi-modal depth loss of Gauss in each iteration according to some embodiments of the present specification;

[0055] Figure 4 is an exemplary schematic diagram of gradient conflict driven densification according to some embodiments of the present specification;

[0056] Figure 5 is an exemplary schematic diagram of a geometry consistency carving and gradient conflict driven densification Gaussian reconstruction method according to some embodiments of the present specification. DETAILED DESCRIPTION

[0057] The specific embodiments of the present application are described below to enable those skilled in the art to understand the present application, but it should be clear that the present application is not limited to the scope of the specific embodiments. It is obvious to those skilled in the art that various changes are within the spirit and scope of the present application as defined by the appended claims, and all applications utilizing the concept of the present application are within the scope of protection.

[0058] Embodiment One

[0059] Figure 1 is an exemplary flowchart of a geometry consistency carving and gradient conflict driven densification Gaussian reconstruction method according to some embodiments of the present specification. As shown in Figure 1 and Figure 5 , the flow includes the following steps. In some embodiments, the flow can be executed by a processor.

[0060] S1: Obtain scene image training data.

[0061] The scene image training data is sparse view image data used for inference of a multi-modal depth estimation network (Ambiguity-aware Depth Estimation Network). For example, the scene image training data can include sparse view table, chair, depth ambiguous object image data such as light transmission, etc.

[0062] In some embodiments, the processor can select sparse viewpoints from the multi-view of the real scene as the training data of the entire scene, and obtain the scene image training data through structure from motion reconstruction.

[0063] In some embodiments, as shown in Figure 2 The processor can obtain camera pose training data through structure from motion (SfM) based on the RGB dataset of the image; when the three-dimensional reconstruction toolkit (clomap) sparse reconstruction fails, call the dense stereo matching function patch_match_stereo to continue depth estimation, and use stereo fusion to fuse the depth maps of each viewpoint into a dense point cloud to obtain the scene image training data.

[0064] S2: Analyze the scene image training data using a multi-modal depth estimation network to obtain multi-modal depth hypothesis data.

[0065] The multi-modal depth estimation network (Ambiguity-aware Depth Estimation Network) is a neural network used to predict multiple depth values for each pixel. For example, the multi-modal depth estimation network (Ambiguity-aware Depth Estimation Network) can be an improved LeReS monocular depth estimation network.

[0066] In some embodiments, the architecture of the multi-modal depth estimation network can include: the input layer receives a 480x640x3 RGB image and a 32-dimensional latent code vector; the feature encoder contains 5 levels of down-sampling modules (based on an improved ResNet-50 architecture), each level processes the features through 3x3 convolution deformable convolution and instance normalization, and finally outputs the feature map; the AdaIn modulation system contains 4 adaptive normalization layers, and outputs the modulated multi-scale features; the multi-hypothesis decoder adopts a 4-level up-sampling structure, each level processes the features through nearest neighbor up-sampling, 3x3 convolution and skip connection, and finally generates a 480x640 depth map through 20 parallel output branches.

[0067] The multi-modal depth hypothesis data is data reflecting the assumed depth of the training picture.

[0068] In some embodiments, the processor can employ a pre-trained completed multi-modal depth estimation network as a core analysis model, input the scene image training data RGB image to the multi-modal depth estimation network, extract features through the LeReS backbone network in the multi-modal depth estimation network, and then perform dynamic modulation by multiple adaptive instance normalization layers to generate K depth hypothesis samples; sample the multi-modal depth hypothesis data corresponding to the pixels at N positions of each image in the depth hypothesis samples to obtain N*K multi-modal depth hypothesis data; wherein the multi-modal depth estimation network can include 4 adaptive instance normalization layers.

[0069] The LeReS backbone network is a depth estimation network based on a single image.

[0070] S3: Based on the scene image training data, initialize the Gaussian, sample the spatial Gaussian depth, and obtain the synthesized view multi-modal depth prediction.

[0071] The synthesized view multi-modal depth prediction is the depth prediction value of each pixel in the view synthesis in the iteration process.

[0072] In some embodiments, the processor can initialize the Gaussian, select the same N pixel positions as the multi-modal depth hypothesis, and obtain the synthesized view multi-modal depth prediction by passing through the depth of all Gaussians from the pixel to the ray of the training view center of the scene image training data, that is, the p Gaussian prediction depth data preparation of N pixels in each iteration.

[0073] In some embodiments, the processor can obtain K depth hypothesis images of each picture of the scene image training data. Extract K hypothesis depths of N pixels of each training picture to complete the multi-modal depth hypothesis data preparation of Gaussian carving. After the SFM initializes the Gaussian, the training view is sampled by Gaussian in each iteration of the representation and rendering of each scene. Specifically, select the same N pixel positions as the multi-depth hypothesis, and the depth of all Gaussians passing through the ray from the pixel to the training view center as the prediction depth. Complete the p Gaussian depth prediction data preparation of N pixels in each iteration.

[0074] S4: Based on the synthesized view multi-modal depth prediction and the multi-modal depth hypothesis data, use the geometric consistency loss carving to carve the Gaussian in space to obtain the Gaussian space carving result; use the Gaussian gradient conflict to drive the Gaussian densification to obtain the Gaussian densification result.

[0075] The Gaussian space carving result is the processing result of driving the Gaussian to move to the correct depth position and reducing the opacity, losing the small Gaussian position, and increasing the Gaussian opacity.

[0076] The Gaussian densification result is the Gaussian processing result of combining the gradient features at the pixel level for clustering and guiding more reasonable new Gaussian generation positions.

[0077] In some embodiments, the processor can perform geometric consistency regularization on the depth distribution of the Gaussian sphere in three-dimensional space based on the synthesized view multi-modal depth prediction and the multi-modal depth hypothesis data, using a geometric consistency loss, to obtain an overall space carving loss;

[0078] The overall space carving loss is used to perform soft constraint on the depth error of the Gaussian sphere, and in multiple iterations, the Gaussians under different viewing angles are continuously optimized in position and opacity to find the consistent depth of the Gaussians, and the Gaussian space carving result is obtained.

[0079] The overall space carving loss is a loss function used to perform soft constraint on the depth error of the Gaussian sphere and improve its geometric rationality in three-dimensional space.

[0080] In some embodiments, the processor can perform geometric consistency regularization on the depth distribution of the Gaussian sphere in three-dimensional space. This loss is derived from projection geometry and directly acts on the radial arrangement of the Gaussian sphere along the pixel ray direction, making the Gaussian distribution closer to the geometric representation of the real scene. Specifically, suppose N pixel points are sampled from each input image, and each pixel contains P Gaussian depth sampling points corresponding to the Gaussian predicted depth At the same time, N*K depth hypotheses are introduced to represent the various geometric interpretations that may exist for the sampling point.

[0081] In some embodiments, the expression of the overall space carving loss can be:

[0082] ;

[0083] ;

[0084] ;

[0085] wherein, the overall space carving loss is represented by Lsc, the number of pixel points is represented by N, the geometric consistency error of a single pixel is represented by e, the number of Gaussians corresponding to a single pixel is represented by P, the minimum distance between the Gaussian predicted depth and all hypotheses is represented by dmin, and denotes the minimum value of all k depth hypotheses corresponding to the pixel, denotes the predicted depth of the i-th pixel given by the j-th visible Gaussian, denotes the m-th depth hypothesis of the i-th pixel, m denotes the multi-depth hypothesis index, and k denotes the number of multi-depth hypotheses, denotes the absolute value.

[0086] In some embodiments, the overall space carving loss improves the geometric reasonableness of the Gaussian sphere depth error in three-dimensional space through a soft constraint, so that the prediction result can be aligned with at least one feasible hypothesis. Especially in sparse or highly uncertain scenes, the regularization term can effectively guide the Gaussian distribution in the three-dimensional reconstruction process to converge towards the real geometry, improving the structured quality of the reconstruction.

[0087] In some embodiments, the processor can obtain the multi-modal depth loss of each Gaussian based on the overall space carving loss. As shown in Figure 3 In each iteration, a large multi-modal depth loss of a Gaussian drives the Gaussian to move to the correct depth position and reduce the opacity, and a small loss keeps the Gaussian position unchanged and increases the Gaussian opacity. In multiple iterations, the Gaussians under different perspectives are continuously optimized in position and opacity, and finally find the consistent depth of the Gaussian. The Gaussians with large loss under multiple perspectives will gradually reduce their opacity, and finally be deleted by the densification strategy in the 3D Gaussian spattering algorithm (3DGS), obtaining the Gaussian space carving result.

[0088] Through the geometric consistency loss as a feedback signal, the opacity (a value) and the spatial position of each Gaussian particle are jointly regulated in each training iteration. When the predicted depth of a certain Gaussian is inconsistent with the multi-modal hypothesis for a long time, that is, its geometric consistency error is large, then by reducing its opacity and allowing position update, the weight of the Gaussian in the scene representation is gradually weakened. On the contrary, for the Gaussian with small consistency error, its opacity is enhanced and its spatial position is kept stable, so as to strengthen its dominant contribution to the final scene structure.

[0089] S5: based on the Gaussian space carving result and the Gaussian densification result, a Gaussian spattering model result is obtained by calculation, and the Gaussian reconstruction of the sparse view is completed.

[0090] In some embodiments, as shown in Figure 4As shown, after SFM initialization, in each iteration, each Gaussian under the camera view will cover certain pixels, and the covered pixels contribute to the gradient of multiple Gaussians. When the gradient contribution of multiple pixels to a certain Gaussian is close to 0, but the gradient contribution of the covered pixels can be seriously conflicting, these Gaussians will be ignored by the original 3D GS Gaussian densification strategy. The new Gaussian position is not accurate enough and other problems. In order to solve this problem, a clustering algorithm is introduced, which combines the gradient features at the pixel level to guide the generation of more reasonable new Gaussian positions.

[0091] In some embodiments, the processor can, based on the Gaussian space carving result and the Gaussian densification result, for each Gaussian under the current iteration training view, count the pixel set covered by the Gaussian, and extract the opacity gradient and position gradient of each pixel to the Gaussian, respectively, to obtain the sign feature and gradient amplitude corresponding to the opacity gradient and position gradient;

[0092] Based on the gradient conflict driven densification, the sign feature and gradient amplitude corresponding to the opacity gradient and position gradient are calculated to obtain the opacity conflict score and position conflict score;

[0093] Based on the opacity conflict score and the position conflict score, the two-dimensional gradient features of the covered pixels are extracted, and a weighted clustering method is used for clustering to obtain a clustering result;

[0094] The cluster center of the clustering result is back projected to the three-dimensional space to obtain a conversion result;

[0095] Based on the conversion result, the new Gaussian is copied and split, and the Gaussian splatter model result is obtained through multiple iterations to complete the Gaussian reconstruction of the sparse view.

[0096] In some embodiments, the processor first counts, for each Gaussian under the current iteration training view, the pixel set covered by the Gaussian extracts the opacity gradient and position gradient of each pixel to the Gaussian. Further, the sign feature and gradient amplitude of the Gaussian opacity and position gradient of the pixel are extracted, which are used to measure the gradient conflict of the Gaussian as a whole.

[0097] The opacity conflict score is a score quantifying the opacity conflict degree of the Gaussian under the current view.

[0098] In some embodiments, the expression of the opacity conflict score can be:

[0099] ;

[0100] wherein, a table opacity conflict score, denotes the i-th Gaussian, denotes the j-th pixel within the Gaussian projection region, denotes the pixel set, denotes the sign feature of the Gaussian opacity, denotes the gradient magnitude of the Gaussian opacity.

[0101] The position conflict score is a score quantifying the degree of position conflict of the Gaussian under the current view angle.

[0102] In some embodiments, the expression of the position conflict score can be:

[0103] ;

[0104] wherein, denotes the position conflict score, denotes the sign feature of the position gradient, denotes the gradient magnitude of the position gradient.

[0105] In some embodiments, when is less than a threshold , it is determined that the Gaussian has a significant pixel gradient conflict, indicating that the different covering pixels of the Gaussian cancel each other out in the direction of Gaussian update, hindering effective training update.

[0106] The clustering result is the result of clustering the pixel points of the Gaussian conflict.

[0107] In some embodiments, for the Gaussian determined to be in conflict, the processor can further extract the two-dimensional gradient features of the covering pixels of the Gaussian wherein is the coordinate of the pixel on the image plane, denotes the opacity and position gradient sign, respectively, is the absolute value of the opacity and position gradient, respectively, and a weighted clustering method is used for clustering to obtain the clustering result.

[0108] In some embodiments, the expression of the clustering result can be:

[0109] ;

[0110] wherein, denotes the center feature of the k-th clustering cluster, denotes the clustering operation, denotes the j-th pixel feature, denotes the clustering weight of each pixel feature, number of clusters, denotes the absolute value of the i-th Gaussian opacity gradient corresponding to the j-th pixel, denotes the absolute value of the i-th Gaussian position gradient corresponding to the j-th pixel.

[0111] The conversion result is the conversion of the two-dimensional coordinates of the cluster center to the three-dimensional space.

[0112] In some embodiments, the processor can take the two-dimensional coordinates of the cluster center and convert to the three-dimensional space by back projection operation.

[0113] In some embodiments, the expression of the conversion result can be:

[0114] ;

[0115] wherein, denotes the conversion result, denotes the projection from the camera coordinate system to the world coordinate system, denotes the horizontal coordinate of the cluster center, denotes the vertical coordinate of the cluster center, denotes the two-dimensional coordinates of the cluster center, denotes the depth value.

[0116] In some embodiments, the opacity and scale parameters of the newly generated Gaussian are initialized as: , directly inheriting the source Gaussian opacity and scale parameters. By introducing a gradient conflict awareness mechanism, new geometric details are dynamically supplemented during the training process, further improving the integrity and fineness of three-dimensional reconstruction.

[0117] In each training iteration, for each Gaussian particle under the current view, first, the set of image pixels covered by it is counted, and the gradient contributions of these pixels to the Gaussian position and opacity under the current view are recorded, including the direction (sign) and amplitude (size) of the gradient.

[0118] Then, the gradient conflict scoring function of opacity and position is defined to measure whether the Gaussian has gradient direction conflict under the supervision of multiple pixels, i.e., whether the optimization direction of the Gaussian by different pixels is contradictory. If the gradient conflict degree of a Gaussian exceeds the set threshold, it indicates that it cannot converge effectively in the current optimization process, and may be in a geometric edge, occlusion boundary or local detail area. At this time, the two-dimensional coordinates and gradient features of the covered pixels are extracted, a weighted feature vector is constructed, and a clustering algorithm is used to identify multiple conflict areas.

[0119] The Gaussian splashing model result is a reasonable distribution of Gaussians in space, and its parameters include position, opacity and shape.

[0120] In some embodiments, based on the Ambiguity-aware Depth Estimation Network, multiple (K) candidate depth maps are generated on each training image, K depth hypotheses of N pixels are sampled to form a set of multi-modal depth hypotheses to capture the inherent depth ambiguity in the scene. Then, during the training process, N pixels of the same position are sampled from the current image and ray tracing is performed along the camera principal axis direction, and the intersection depths of the ray with P Gaussians are recorded as a set of predicted depth views. The geometric consistency loss is calculated between the two, and the Gaussian parameters are updated through multiple iterations of gradient backpropagation, gradually approaching the consistent depth of the Gaussian representation under multi-view, and achieving depth carving in the Gaussian space. At the same time, during the iterative optimization process, the gradient conflict score of the opacity and position of each Gaussian also needs to be calculated. For Gaussians with significant conflicts, the pixel regions covered by them in the two-dimensional image projection are clustered to separate different geometric components. Then, new Gaussian instances are generated for the conflict region through the back projection operation, realizing the replication and splitting densification of Gaussians, and thus enhancing the geometric details of the scene representation. Finally, this process can effectively complete the three-dimensional scene reconstruction based on sparse views.

[0121] Based on the LeReS multi-modal depth estimation network, multiple (K) candidate depth maps are generated on each training image to form a set of multi-modal depth hypotheses to capture the inherent depth ambiguity in the scene. Then, during the training process, N pixels are randomly sampled from the current image, and ray tracing is performed along the camera principal axis direction, and the intersection depths of the ray with P Gaussians are recorded as a set of predicted depth views.

[0122] In some embodiments, as Figure 5As shown, a Gaussian reconstruction method based on geometric consistency carving and gradient conflict driven densification is proposed. Firstly, for the input sparse view RGB image, a multi-depth estimation network is used to generate multiple depth hypothesis maps, and the corresponding depth hypothesis set is obtained through pixel sampling as the supervision data of the geometric consistency loss. At the same time, the camera pose is obtained through motion structure reconstruction, and when the sparse reconstruction fails, the dense stereo matching and multi-view depth fusion are used to generate a dense point cloud, and the three-dimensional Gaussian initialization is carried out accordingly. In the training stage, pixels are sampled from the depth hypothesis every iteration and rays are emitted towards the scene center. After the interaction between the rays and the Gaussian, the synthesized depth prediction is formed, and the position and opacity of the Gaussian are optimized by the geometric consistency loss constraint. At the same time, in order to alleviate the local geometric deficiency, the gradient conflict score of the opacity and position of the Gaussian is calculated, the features of the conflict significant area are extracted and clustered to generate new Gaussian centers, which are mapped to the three-dimensional space through back projection, so as to gradually realize the densification carving of the scene. Finally, under the double driving of geometric consistency and gradient conflict, the Gaussian parameters are continuously optimized, and the fine and complete three-dimensional scene reconstruction result is obtained.

[0123] In some embodiments, the processor can start from sparse RGB images and gradually realize three-dimensional scene reconstruction. Firstly, a multi-depth estimation network is used to generate k depth hypothesis maps for each RGB image. Then N pixels are sampled on each depth map, and for each pixel, its corresponding K depth hypotheses are retained, thereby obtaining a depth hypothesis set with a size of N×K as the supervision data of the geometric consistency loss. Secondly, the camera pose information is obtained through motion structure reconstruction (SfM). When the sparse reconstruction based on the three-dimensional reconstruction toolkit COLMAP fails, the depth estimation is performed by calling the dense stereo matching function patch_match_stereo, and the multi-view depth maps are fused into a dense point cloud through stereo_fusion. The dense point cloud is further used for Gaussian initialization of the three-dimensional Gaussian splatting technology (3DGaussianSplatting, 3DGS). After entering the training stage, the same N pixels are selected from the multi-depth hypothesis data every iteration, and rays are emitted from these pixels towards the scene center. The P Gaussians intersected in the ray path are regarded as the multi-depth prediction of the pixel, thereby obtaining synthesized depth prediction data with a size of N×P. Subsequently, as Figure 3Based on the prediction result, the geometric consistency loss is calculated, and the position and opacity of the Gaussian are optimized by back propagation. At the same time, in the iteration process, the gradient conflict score of the opacity and position of each Gaussian is calculated. For the Gaussians with significant conflict, the pixel features of the coverage area are extracted in the two-dimensional image plane, and clustering is performed to generate new Gaussian centers and depth information. Subsequently, these new Gaussians are mapped back to the three-dimensional space through back projection, realizing the dense complement of the scene geometry. Finally, after multiple iterations, under the dual action of geometric consistency carving and gradient conflict driven densification, the parameters of the Gaussian are continuously optimized, and the three-dimensional scene reconstruction is gradually completed.

[0124] In some embodiments of the present specification, a Gaussian reconstruction method of geometric consistency carving and gradient conflict driven densification is provided, which uses a multi-modal depth estimation network to analyze scene image training data to obtain multi-modal depth hypothesis data; based on the scene image training data, the Gaussian is initialized, the spatial Gaussian depth is sampled to obtain synthetic view multi-modal depth prediction; based on the synthetic view multi-modal depth prediction; and the multi-modal depth hypothesis data, the geometric consistency loss is carved to obtain the Gaussian spatial carving result; at the same time of obtaining the Gaussian spatial carving result, the gradient conflict driven densification is processed to complete the Gaussian reconstruction of the sparse view. (1) The gradient conflict driven densification realizes the dynamic modeling of the occlusion boundary, texture details and small-scale structure through the fine splitting and replication of local Gaussian particles, which significantly enhances the integrity and detail restoration ability of the final reconstruction result. (2) The geometric consistency carving realizes the automatic "carving" and removal of error Gaussians in continuous iterations without relying on hard threshold judgment or post-processing screening. Finally, the distribution of Gaussian particles in the scene gradually tends to be reasonable and sparse, avoiding false structures caused by error initial modeling or sparse input, and improving the structure clarity and representation simplicity of the model. (3) The use of multi-modal depth estimation network enables the Gaussian distribution to gradually converge to at least one feasible geometric interpretation, especially in areas with high depth uncertainty caused by reflection, texture loss or transparent objects, which still maintains the coherence and consistency of the geometric structure, significantly improving the robustness and accuracy of three-dimensional modeling under sparse view. (4) Without relying on traditional global dense matching or large-scale optimization framework, stable convergence can be achieved under limited computing resources, significantly reducing the computational complexity and training overhead. (5) Since the Gaussian distribution can be directly combined with the rendering framework, the geometric reconstruction can be completed while efficiently supporting view synthesis and rendering, improving the application value in virtual reality, augmented reality and digital twin.

[0125] Embodiment two

[0126] The NeRF and 3DGS are inaccurate in geometry reconstruction and illumination estimation due to insufficient view information, and the visual error and reprojection error are aggravated under the condition of sparse view. There are problems such as difficulty in fully learning the 3D details and structure of the scene, resulting in reconstruction degradation and rendering quality decline.

[0127] A geometric consistency carving and gradient conflict driven densification Gaussian reconstruction method, comprising:

[0128] S1: acquiring scene image training data of virtual reality three-dimensional automated modeling;

[0129] S2: analyzing the scene image training data by using a multi-modal depth estimation network to obtain multi-modal depth hypothesis data;

[0130] S3: based on the scene image training data, initializing a Gaussian, sampling the spatial Gaussian depth, and obtaining a synthesized view multi-modal depth prediction;

[0131] S4: based on the synthesized view multi-modal depth prediction and the multi-modal depth hypothesis data, using a geometric consistency loss carving to carve the Gaussian in space, obtaining a Gaussian spatial carving result; using Gaussian gradient conflict to drive Gaussian densification, obtaining a Gaussian densification result;

[0132] S5: based on the Gaussian spatial carving result and the Gaussian densification result, calculating to obtain a Gaussian splatting model result of virtual reality three-dimensional automated modeling, completing sparse view Gaussian reconstruction of virtual reality three-dimensional automated modeling.

[0133] In some embodiments, the S1 comprises:

[0134] Based on the RGB dataset of the image in the virtual reality three-dimensional automated modeling, the camera pose training data is obtained by motion structure reconstruction;

[0135] When the three-dimensional reconstruction toolkit fails to reconstruct the camera pose training data sparsely, a dense stereo matching function is called to continue depth estimation, and each view depth map is fused into a dense point cloud by stereo vision fusion to obtain the scene image training data of virtual reality three-dimensional automated modeling.

[0136] In some embodiments, the S2 comprises:

[0137] A pre-trained multi-modal depth estimation network is used as a core analysis model, the scene image training data RGB image is input into the multi-modal depth estimation network, features are extracted by a LeReS backbone network in the multi-modal depth estimation network, and K depth hypothesis samples are generated by dynamic modulation of multiple style transfer layers;

[0138] The N*K multi-modal depth hypothesis data are obtained by sampling the multi-modal depth hypothesis data corresponding to N positions of pixels of each image in the depth hypothesis sample.

[0139] In some embodiments, the S3 comprises:

[0140] The Gaussians are initialized, the same N pixel positions as the multi-modal depth hypothesis are selected, the rays from the pixels to the training view center of the training data of the scene image are obtained, the depths of all Gaussians that the rays pass through are obtained, and the synthesized view multi-modal depth prediction is obtained, that is, the p Gaussian prediction depth data of N pixels in each iteration are prepared.

[0141] In some embodiments, the S4 comprises:

[0142] Based on the synthesized view multi-modal depth prediction and the multi-modal depth hypothesis data, the depth distribution of the Gaussian sphere in the three-dimensional space is geometrically consistent regularized by using a geometric consistency loss, and an overall space carving loss is obtained;

[0143] The Gaussian sphere depth error is soft constrained by using the overall space carving loss, and the Gaussians under different views are continuously optimized in position and opacity in multiple iterations, the consistent depth of the Gaussians is found, and a Gaussian space carving result is obtained.

[0144] In some embodiments, the expression of the overall space carving loss is:

[0145] ;

[0146] ;

[0147] ;

[0148] wherein, the overall space carving loss is represented by Lsc, the number of pixel points is represented by N, the geometric consistency error of a single pixel is represented by egi, the number of Gaussians corresponding to a single pixel is represented by ki, the minimum distance between the Gaussian prediction depth and all hypotheses is represented by di, the minimum value on all k depth hypotheses corresponding to the pixel is represented by min, the prediction depth of the i-th pixel given by the j-th visible Gaussian is represented by di,j, the m-th depth hypothesis of the i-th pixel is represented by di,m, m represents the multi-depth hypothesis index, and k represents the number of multi-depth hypotheses, the absolute value is represented by | |.

[0149] In some embodiments, the S5 comprises:

[0150] Based on the Gaussian space sculpting results and Gaussian density results, for each Gaussian in the current iterative training perspective, the set of pixels it covers is counted, and the opacity gradient and position gradient of each pixel to the Gaussian are extracted, and the sign features and gradient magnitudes corresponding to the opacity gradient and position gradient are obtained respectively.

[0151] Based on gradient conflict-driven densification, the sign features and gradient magnitudes corresponding to the opacity gradient and position gradient are calculated to obtain the opacity conflict score and position conflict score.

[0152] Based on the opacity conflict score and the position conflict score, the two-dimensional gradient features of the covered pixels are extracted, and the weighted clustering method is used to cluster them to obtain the clustering results.

[0153] The cluster centers of the clustering results are back-projected into three-dimensional space to obtain the transformation result;

[0154] Based on the transformation results, the new Gaussian is copied and split, and the Gaussian splash model result of virtual reality 3D automated modeling is obtained through multiple iterations, thus completing the sparse view Gaussian reconstruction of virtual reality 3D automated modeling.

[0155] In some embodiments, the expression for the opacity conflict score is:

[0156] ;

[0157] in, Table of Opacity Conflict Scores Let i represent the i-th Gaussian. express The j-th pixel within the Gaussian projection region, Represents a set of pixels. The symbolic feature representing Gaussian opacity. This represents the gradient magnitude of the Gaussian opacity.

[0158] The expression for the location conflict score is:

[0159] ;

[0160] in, Indicates the score for positional conflict. The sign characteristics representing the position gradient, This represents the magnitude of the gradient at the location.

[0161] In some embodiments, the expression for the clustering result is:

[0162] ;

[0163] in, denotes the center feature of the k-th cluster, denotes the clustering operation, denotes the j-th pixel feature, denotes the clustering weight of each pixel feature, denotes the number of cluster, denotes the absolute value of the i-th Gaussian opacity gradient corresponding to the j-th pixel, denotes the absolute value of the i-th Gaussian position gradient corresponding to the j-th pixel.

[0164] In some embodiments, the expression of the conversion result is:

[0165] ;

[0166] wherein, denotes the conversion result, denotes the projection from the camera coordinate system to the world coordinate system, denotes the horizontal coordinate of the cluster center, denotes the vertical coordinate of the cluster center, denotes the two-dimensional coordinate of the cluster center, denotes the depth value.

[0167] By utilizing the Gaussian reconstruction of sparse views based on the Gaussian space carving result and the Gaussian densification result for the sparse views of the virtual reality three-dimensional automated modeling scene, through the fine splitting and replication of local Gaussian particles, dynamic modeling of occlusion boundaries, texture details, and small-scale structures is achieved, significantly enhancing the integrity and detail restoration ability of the final virtual reality three-dimensional automated modeling scene reconstruction result. In continuous iterations, automatic "carving" and removal of error Gaussians are achieved without relying on hard threshold judgment or post-processing screening. Finally, the distribution of Gaussian particles in the scene gradually tends to be reasonable and sparse, avoiding pseudo-structures caused by error initial modeling or sparse input, and improving the structural clarity and simplicity of the model.

Claims

1. A Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification, characterized in that, include: S1: Acquire scene image training data; The scene image training data includes depth-blurred object image data; S2: Analyze the scene image training data using a multimodal depth estimation network to obtain multimodal depth hypothesis data; S3: Based on the scene image training data, initialize Gaussian, sample the spatial Gaussian depth, and obtain the synthetic view multimodal depth prediction; S4: Based on the synthetic view multimodal depth prediction and multimodal depth assumption data, Gaussian spatial sculpting is performed using geometric consistency loss sculpting to obtain Gaussian spatial sculpting results; Gaussian gradient collision is used to drive Gaussian denserization, and the Gaussian denserization result is obtained. This includes: based on synthetic view multimodal depth prediction and multimodal depth assumption data, using geometric consistency loss to perform geometric consistency regularization on the depth distribution of the Gaussian sphere in 3D space, and obtaining the overall spatial sculpting loss; The expression for the overall spatial sculpting loss is: ; ; ; in, This indicates the overall loss of sculpted space. Indicates the number of pixels. This represents the geometric consistency error of a single pixel. This represents the number of Gaussians corresponding to a single pixel. This represents the minimum distance between the Gaussian prediction depth and all hypotheses. This means taking the minimum value over all k depth assumptions corresponding to that pixel. This represents the predicted depth of the i-th pixel given by the j-th visible Gaussian. Let m represent the m-th depth hypothesis for the i-th pixel, where m is the multi-depth hypothesis index. Indicates the number of multi-depth hypotheses. Indicates taking the absolute value; By using the overall spatial sculpting loss to apply soft constraints to the depth error of the Gaussian sphere, the position and opacity of the Gaussian sphere are continuously optimized under different viewpoints in multiple iterations, and the consistent depth of the Gaussian sphere is found to obtain the Gaussian spatial sculpting result. S5: Based on the Gaussian space sculpting results and Gaussian densification results, for each Gaussian in the current iteration training perspective, the set of pixels it covers is counted, and the opacity gradient and position gradient of each pixel to the Gaussian are extracted, and the sign features and gradient magnitudes corresponding to the opacity gradient and position gradient are obtained respectively. Based on gradient conflict-driven densification, the sign features and gradient magnitudes corresponding to the opacity gradient and position gradient are calculated to obtain the opacity conflict score and position conflict score. Based on the opacity conflict score and the position conflict score, the two-dimensional gradient features of the covered pixels are extracted, and the weighted clustering method is used to cluster them to obtain the clustering results. The cluster centers of the clustering results are back-projected into three-dimensional space to obtain the transformation result; Based on the transformation results, the new Gaussian model is copied and split, and the Gaussian splash model results are obtained through multiple iterations, thus completing the Gaussian reconstruction of the sparse view.

2. The Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification according to claim 1, characterized in that, S1 includes: Camera pose training data is obtained by reconstructing the structure of motion based on the RGB dataset of the image. When the 3D reconstruction toolkit fails to sparsely reconstruct camera pose training data, it calls the dense stereo matching function to continue depth estimation and uses stereo vision fusion to fuse the depth maps from each viewpoint into a dense point cloud to obtain scene image training data.

3. The Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification according to claim 1, characterized in that, S2 includes: A pre-trained multimodal depth estimation network is used as the core analysis model. The RGB images of the scene image training data are input into the multimodal depth estimation network. After the LeReS backbone network in the multimodal depth estimation network extracts features, multiple style transfer layers dynamically modulate the data to generate K depth hypothesis samples. For each image in the depth hypothesis sample, multimodal depth hypothesis data corresponding to N pixel positions are sampled to obtain N*K multimodal depth hypothesis data.

4. The Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification according to claim 1, characterized in that, S3 includes: Initialize Gaussians, select N pixel positions that are the same as the multimodal depth hypothesis, and calculate the depth of all Gaussians through which the ray from the pixel to the training viewpoint center of the scene image training data passes. This yields the multimodal depth prediction of the synthetic view, i.e., the preparation of p Gaussian predicted depth data for N pixels in each iteration.

5. The Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification according to claim 1, characterized in that, The expression for the opacity conflict score is: ; in, Table of Opacity Conflict Scores Let i represent the i-th Gaussian. express The j-th pixel within the Gaussian projection region, Represents a set of pixels. The symbolic characteristics representing Gaussian opacity. This represents the gradient magnitude of the Gaussian opacity. The expression for the location conflict score is: ; in, Indicates the score for positional conflict. The sign characteristics representing the position gradient, This represents the magnitude of the gradient at the location.

6. The Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification according to claim 1, characterized in that, The expression for the clustering result is: ; in, This represents the central feature of the k-th cluster. This represents the clustering operation. Represents the feature of the j-th pixel. This represents the clustering weight for each pixel feature. Indicates the number of clusters. This represents the absolute value of the Gaussian opacity gradient corresponding to the j-th pixel. It represents the absolute value of the gradient at the i-th Gaussian position corresponding to the j-th pixel.

7. The Gaussian reconstruction method based on geometric consistency sculpting and gradient conflict-driven densification according to claim 1, characterized in that, The expression for the conversion result is: ; in, Indicates the conversion result. This represents the projection from the camera coordinate system to the world coordinate system. The x-coordinate represents the cluster center. The ordinate representing the cluster center, Two-dimensional coordinates representing the cluster centers This represents the depth value.

Citation Information

Patent Citations

  • 3DGS new view rendering quality improvement method based on Gaussian visibility

    CN120472067A

  • 5G signal tower three-dimensional rendering and high-altitude panoramic image synthesis method based on GSII

    CN120495079A