Dynamic Gaussian scene reconstruction method based on depth regularization

By introducing a new loss form based on depth regularization in the dynamic Gaussian reconstruction method, the problem of lack of constraints in the Gaussian body geometric structure in dynamic scenarios is solved, and more stable geometric consistency and detail fidelity are achieved.

CN119991974AActive Publication Date: 2025-05-13HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510480633.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The geometric structure of the Gaussian body in dynamic scenes lacks constraints, resulting in shape drift and instability. The existing dynamic Gaussian reconstruction methods are difficult to capture object motion and time-varying geometric information.

Method used

A new loss form based on depth regularization is adopted, a monocular depth map is generated through the pre-trained depth estimation network, and the correlation between the rendered depth and the reference depth is used as an additional loss term to constrain the geometry during the Gaussian body motion.

Benefits of technology

It significantly improves the geometric consistency and detail fidelity of the 3D Gaussian model in dynamic scenarios, improves the consistency and detail performance of reconstruction results from a new perspective, and solves the problems of shape drift and instability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991974A_ABST
    Figure CN119991974A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic Gaussian scene reconstruction method based on depth regularization. The method comprises the following steps: step (1), initializing a group of 3D Gaussian sets by using an SfM point cloud, wherein each Gaussian body is described by parameters such as a central position and a covariance matrix; (2) introducing a control network decoupling motion and a geometric structure, performing offset calculation by using a time variable and a Gaussian center position, and controlling dynamic transformation of a Gaussian body; (3) introducing a loss function of scale-independent depth alignment based on local depth normalization and Pearson correlation; and (4) a new loss function is obtained through calculation according to the rendered image, and the Gaussian field and the control network are optimized through reverse gradient return. Experiments on dynamic scene data sets such as D-NeRF and NeRF-DS show that the method is superior to an existing method in the aspects of reconstruction quality, geometric consistency and new view rendering effect, the stability and detail fidelity of the dynamic scene are improved, and the method has important research value and application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is directed to the field of dynamic three-dimensional scene reconstruction and mentions a new loss form based on deep regularization for geometric constraints in dynamic Gaussian reconstruction tasks. Background Art

[0002] The Gaussian reconstruction method achieves high-fidelity reconstruction and real-time rendering by explicitly modeling the scene with a set of learnable 3D Gaussian distributions. This method uses the center position, shape parameters (covariance matrix decomposition into rotation and scale) and opacity of the 3D Gaussian to form a sparse but efficient scene representation. When dealing with static scenes, 3D-GS has demonstrated the powerful capabilities of point-based rendering technology, but when the scene changes dynamically, traditional static methods find it difficult to capture object motion and time-varying geometric information.

[0003] However, in dynamic scenes, on the one hand, the shapes and positions of objects change over time, making it difficult for modeling methods based on static assumptions to accurately capture these changes. In addition, objects in dynamic scenes may involve complex non-rigid deformations, making it difficult for traditional methods based on rigid body assumptions to maintain high-precision reconstruction effects. How to establish a consistent point-based representation between different time frames to ensure the stability and temporal consistency of reconstruction remains a key challenge.

[0004] On the other hand, due to the lack of sufficient geometric constraints, especially when dealing with temporal interpolation tasks, the generated depth information may be jittery. This jitter may come from the rapid motion of objects, occlusion changes, or uneven data sampling, resulting in unstable geometric structures in the temporal dimension. In addition, due to the limited use of geometric information of the input image by the current dynamic Gaussian reconstruction method and the failure to fully combine multi-view depth priors, the reconstructed 3D structure still has certain limitations in scale, detail, and consistency. Summary of the invention

[0005] In view of this, the purpose of the present invention is to address the lack of constraints on the geometric structure of Gaussian bodies in dynamic scene reconstruction, which leads to shape drift and instability, and propose a dynamic Gaussian scene reconstruction method based on depth regularization. The strategy based on depth regularization can add a new depth loss to the reconstruction of the image, which is used to regulate the geometric shape of the Gaussian body during motion.

[0006] Specifically, by using a pre-trained depth estimation network (DPT) to generate a monocular depth map and taking the correlation between the rendered depth and the reference depth as an additional loss term, the model can obtain a depth correlation loss. This loss constrains the structural consistency of the rendered depth and the reference depth by calculating the Pearson correlation coefficient between the rendered depth and the reference depth, thereby providing a stronger geometric prior and making the reconstruction result have better consistency and detail performance under new perspectives.

[0007] A dynamic Gaussian scene reconstruction method based on deep regularization comprises the following steps: Step 1: Select the data set and initialize the 3D Gaussian; Firstly, the D-NeRF and NeRF-DS datasets are selected, and the SfM algorithm is used to recover sparse point clouds from multi-view images, and the recovered sparse point clouds are used as the global geometric prior of the scene.

[0008] In order to efficiently and differently represent and render the scene, the 3D Gaussian Splatting method is used to represent the scene as a set of 3D Gaussians, each of which is represented by a central position , and the 3D covariance matrix to describe.

[0009] Specifically, the probability distribution of a single 3D Gaussian Written in the form: (1); in, represents a random variable in 3D space, Represents a random variable A specific value of .

[0010] Pixel color calculation is based on point The specific formula for volume rendering is as follows: (2); (3); in, Indicate point The pixel color value after rendering, Is The transmittance is defined as represents the Gaussian distribution of colors along the light direction, and Represents the coordinates of the 3D Gaussian model projected onto the 2D image plane, Represents a 2D covariance matrix. Represents the weight of adjusting the opacity contribution of the Gaussian on the image plane during projection.

[0011] Step 2: Introduce the control network and process it; The present invention introduces a control network, inputs the time variable and the 3D Gaussian center position , the motion and geometric structure are decoupled through the multi-layer perceptron MLP, and the learning process is converted to the canonical space to obtain a time-dependent 3D Gaussian model and output the offset.

[0012] The control network introduced in the present invention is improved based on the multi-layer perceptron MLP. The control network adopted in the present invention includes eight groups of linear layers and batch normalization layers; wherein the input information is connected with the residual information output by the fourth group and then input into the fifth group, such as Figure 2 shown.

[0013] It is used to transform the 3D Gaussian in the specification space to the dynamic control space as follows: (4);

[0014] in, represents the dynamic control parameters of the output, represents the control network, Indicates stopping the gradient operation. Indicates time code, Indicates positional encoding: (5);

[0015] in, Indicates the position encoding order. In the synthetic scene, the 3D Gaussian center position The position encoding order is L=10, 3D Gaussian center position The corresponding position encoding order of time frame t is L=6; in the real scene, L=10 is used for both and t.

[0016] Step 3: Calculate the loss function based on the global depth constraint; In order to optimize the rendering quality and geometric stability of the 3D Gaussian model and the structural control of the control network, the present invention combines image loss, image structure similarity loss and depth constraint loss for optimization; Step 3-1. In order to optimize the reconstruction accuracy of the image, L1 loss is used to measure the absolute error between the image rendered by the model and the target image at the pixel level. The calculation formula is as follows (6);

[0017] in, represents the target image, Represents the image rendered by the model.

[0018] Step 3-2. In order to improve the details and structural consistency of the image, the image structural similarity loss SSIM is introduced to measure the image rendered by the model. With the target image The similarity in terms of brightness, contrast and structure is calculated as follows: (7);

[0019] in, It is a function used to measure the consistency of brightness, contrast and structure between two images. Compared with simple pixel differences.

[0020] Step 3-3. In terms of 3D structure optimization, a depth constraint loss is introduced to ensure that the rendered depth is consistent with the real depth at global and local scales. This method includes local depth normalization (LDN) and scale-independent depth alignment based on Pearson correlation to improve the robustness of the model in different depth ranges.

[0021] Step 3-3-1. Introduce local depth regularization to the image and target image rendered by the model under the same viewing angle to supplement the constraints of the global depth regularization on the local area. In particular, the present invention performs local normalization on the depth map, encourages the acquisition of the depth map corresponding to the image rendered by the model, and obtains the local block of the depth map ; At the same time, obtain the local block of the depth map of the target image generated by the pre-trained depth prediction transformer (DPT) ; Calculate the similarity between two local blocks. The normalization process is expressed as follows: (8);

[0022] in, represents the depth of the normalized local block, represents the local blocks sampled from the rendering depth map and the target depth map, respectively, Respectively represent the mean and standard deviation within the local block,

[0023] Step 3-3-2. Define two scale-independent depth transforms: negative depth transform and reciprocal depth transform; negative depth transform is used to alleviate the problem that the distant depth value is greater than the set threshold, and reciprocal depth transform performs a reciprocal transform on the reference depth to solve the problem that the close-view depth is less than the set threshold but the importance is enhanced.

[0024] Final depth loss Pearson correlation is used to measure the correlation between the rendered depth map and the target depth map. The specific calculation is as follows: (9);

[0025] Among them, Corr(⋅,⋅) represents the Pearson correlation coefficient, represents the depth of the target depth map after the negative depth transformation, Indicates the depth of the rendered depth map, Represents the depth of the target depth map after the reciprocal depth transformation. The specific formula of Corr(⋅,⋅) is as follows: (10);

[0026] in, represents the depth of the normalized local block corresponding to the target depth map, represents the normalized local block depth corresponding to the rendered depth map, represents the covariance of two depths, Represents the variance of a single depth.

[0027] Step 3-4. Add the above losses according to the weight distribution to get the final loss function It is expressed as: (11);

[0028] in, , , Controls the relative weight of each loss term.

[0029] Step 4: Use an end-to-end training method to reversely optimize the 3D Gaussian field and control network, so that it can accurately model the geometry and motion information of dynamic scenes and achieve high-quality dynamic scene reconstruction.

[0030] The beneficial effects of the present invention are as follows: The present invention proposes a dynamic Gaussian scene reconstruction method based on depth regularization to solve the problems of insufficient geometric constraints, shape drift and unstable rendering quality in existing dynamic Gaussian reconstruction methods. By introducing global and local deep geometric regularization, the present invention significantly improves the geometric consistency and detail fidelity of the 3D Gaussian model in dynamic scenes. By utilizing the geometric prior provided by the monocular depth estimation network and using negative depth contrast and reciprocal depth contrast for scale-independent depth alignment, the present invention effectively improves the stability of the 3D structure in dynamic scenes.

[0031] Local depth normalization (LDN) further enhances the ability to reconstruct details and ensures the geometric consistency of objects of different scales. At the same time, the optimized L1+SSIM+depth constraint loss design improves the generalization ability and training stability of the model. Experiments show that this method is superior to existing methods in rendering quality, geometric stability and new perspective synthesis. The image rendering indicators PSNR, SSIM, and LPIPS are improved on the dynamic scene synthesis dataset D-NeRF and the real dataset NeRF-DS, which has important research value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a complete flow chart of the dynamic Gaussian scene reconstruction method based on depth regularization of the present invention.

[0033] Figure 2 This is a control network structure diagram of the present invention. DETAILED DESCRIPTION

[0034] The method of the present invention and its detailed parameters are further described in detail below in conjunction with the accompanying drawings.

[0035] Step (1), Initialize 3D Gaussian Firstly, we select the dynamic scene synthetic dataset D-NeRF and the real dataset NeRF-DS, and use the SfM (Structure-from-Motion) technology to recover the sparse point cloud from multi-view images as the global geometric prior of the scene.

[0036] In order to represent and render the scene efficiently and differently, we use the 3D Gaussian Splatting method to represent the scene as a set of 3D Gaussians.

[0037] Each 3D Gaussian can be found at its center position , and the covariance matrix to describe.

[0038] Specifically, the probability distribution of a single Gaussian can be written as: (12);

[0039] In order to make the learning process of the 3D Gaussian model easier, it is split into two learnable parts: the quaternion r represents the rotation and the three-dimensional vector s represents the scaling. Then, these parts are converted into the corresponding rotation matrix R and scaling matrix S. The resulting 3D covariance matrix can be expressed as: (13);

[0040] In order to optimize the parameters of the 3D Gaussian models in the canonical space, it is important to differentiate and render 2D images from these 3D Gaussian models. In this invention, the differential Gaussian rasterization process is adopted. The 3D Gaussian model can be projected onto a 2D plane and the following 2D covariance matrix is ​​used Rendering for each pixel: (14);

[0041] where J is the Jacobian matrix of the affine approximation of the projective transformation, The pixel color represented by on the image plane is obtained by sequentially rendering using the point-based volume rendering technique: (15); (16);

[0042] in, Indicate point The pixel color value after rendering, Depend on The transmittance is defined as represents the Gaussian distribution of colors along the light direction, and Represents the coordinates of the 3D Gaussian model projected onto the 2D image plane, Represents a 2D covariance matrix. Represents the weight of adjusting the opacity contribution of the Gaussian on the image plane during projection.

[0043] 3DGS uses an adaptive density control strategy to manage the number of Gaussians. For areas with insufficient reconstruction, the Gaussians are copied and moved along the position gradient direction; large Gaussians in high variance areas are split into smaller Gaussians. In addition, the algorithm removes Gaussians whose opacity is below a threshold within a certain iteration cycle.

[0044] Sparse point clouds are recovered from multi-view images using SfM as the global geometric prior of the scene, and 3D Gaussian projection is initialized to represent the scene as a set of 3D Gaussians.

[0045] Step (2): Introduce the control network and process An intuitive way to use 3D Gaussians to model dynamic scenes is to train a 3D Gaussian sputtering set separately for each time-dependent set of views, but it is insufficient for continuous monocular capture in a time series.

[0046] We use a control network that works in conjunction with a 3D Gaussian model to decouple motion and geometry and transform the learning process into a canonical space to obtain a time-independent 3D Gaussian model. This decoupling approach introduces geometric prior information about the scene and links the changes in the position of the 3D Gaussian model to time and coordinates. The core of the control network is a multi-layer perceptron (MLP).

[0047] The control network introduced in the present invention is improved based on the multi-layer perceptron MLP. The control network adopted in the present invention includes 8 groups of linear layers and batch normalization layers; the input information is connected with the residual information of the fourth group output and then input into the fifth group, such as Figure 2 shown.

[0048] Taking as input the time and center position of the 3D Gaussian, the control network produces offsets that subsequently transform the 3D Gaussian in the canonical space to the control space: (17);

[0049] in Indicates stopping the gradient operation. Indicates positional encoding: (18);

[0050] In the synthetic scene, the 3D Gaussian center position The position encoding order is L=10, 3D Gaussian center position The corresponding position encoding order of time frame t is L=6; in the real scene, and t are both L=10. We set the depth D of the control network to 8 and the hidden layer width W to 256. As shown in Tables 1 and 2, related experiments show that using position encoding in the input of the control network can enhance the details of the rendering results. Step (3), calculate the loss function based on the global depth constraint 3-1 Image Loss In order to ensure the closeness between the model rendered image and the target image at the pixel level, we use L1 loss as the reconstruction loss, which is defined as follows: (19);

[0051] in, represents the target image, Represents the image rendered by the model. This loss directly compares the absolute difference of each pixel to ensure accurate restoration of information such as overall brightness and color. We use the weight To regulate.

[0052] 3-2 Image Structural Similarity Loss (SSIM Loss) In order to further improve the details and structural consistency of the generated images, we introduce SSIM loss, which is defined as follows: (20);

[0053] in It is used to measure the consistency of two images in terms of brightness, contrast and structure. Compared with the simple pixel difference, The loss is more robust in terms of the fidelity of local structures such as edges and textures, and can effectively improve the overall visual effect of the image.

[0054] 3-3 Depth Constraint Loss In the task of dynamic scene reconstruction, depth information provides a key geometric prior for the 3D Gaussian volume, which helps to ensure the stability of the 3D structure. However, due to the ever-changing camera perspective in dynamic scenes, the scene may contain objects of different scales, and the depth span between the near and far scenes is large. This change in depth scale makes it impossible to fully restore complex geometric structures by simply relying on global depth regularization. Global depth often focuses on the overall shape constraints, but tends to ignore local details, resulting in geometric drift or incomplete reconstruction in some areas.

[0055] To solve the above problems, we introduce local depth regularization (LDR) on the basis of global depth regularization, so that the rendered depth is consistent with the real depth at the local scale. Specifically, this method includes local depth normalization (LDN) and scale-independent depth alignment based on Pearson correlation to improve the robustness of the model in different depth ranges.

[0056] ① Local depth normalization; Previous studies have used global depth information to assist 3D geometry reconstruction. However, this method performs poorly in complex scenes containing objects of multiple scales. In particular, global depth information usually focuses on global features and ignores fine details in depth information, which may lead to the loss of local geometry.

[0057] To solve this problem, the present invention introduces local depth regularization to the image and target image rendered by the model under the same viewing angle to supplement the constraints of the global depth regularization on the local area. In particular, the present invention performs local normalization on the depth map, encourages the acquisition of the depth map corresponding to the image rendered by the model, and obtains the local block of the depth map. ; At the same time, obtain the local block of the depth map of the target image generated by the pre-trained depth prediction transformer (DPT) ; Calculate the similarity between two local blocks. The normalization process is expressed as follows: (twenty one);

[0058] in, represents the depth of the normalized local block, represents the local blocks sampled from the rendering depth map and the target depth map, respectively, Respectively represent the mean and standard deviation within the local block,

[0059] ②Scale-independent depth alignment based on Pearson correlation; In dynamic scenes, the depth range of each frame varies due to the change in camera perspective, and the depth distribution of near and far view areas can be significantly different. Therefore, matching based only on the original depth value may lead to errors. In order to improve the stability of depth alignment, we use negative depth comparison and reciprocal depth comparison to simultaneously consider the similarity at different depth scales.

[0060] Considering that the depth value estimated by the DPT deep model is negative, we define two scale-independent depth transforms: Negative depth transform: to the reference depth map Take the opposite number, that is - , used to alleviate the problem of large perspective depth values.

[0061] Reciprocal depth transformation: Perform a reciprocal transformation on the reference depth, i.e. , which is used to solve the problem of enhancing the importance of close-up depth.

[0062] Final depth loss Pearson correlation is used to measure the correlation between the rendered depth map and the target depth map. The specific calculation is as follows: (twenty two);

[0063] Among them, Corr(⋅,⋅) represents the Pearson correlation coefficient, represents the depth of the target depth map after the negative depth transformation, Indicates the depth of the rendered depth map, Represents the depth of the target depth map after the reciprocal depth transformation. The specific formula of Corr(⋅,⋅) is as follows: (twenty three);

[0064] in, represents the depth of the normalized local block corresponding to the target depth map, Represents the depth of the normalized local block corresponding to the rendered depth map, represents the covariance of two depths, Represents the variance of a single depth.

[0065] ③Loss calculation; Add the above losses by weight to get the final loss function It is expressed as: (twenty four);

[0066] in, , , Control the relative weights of each loss term. By introducing local depth regularization, the model can not only maintain global geometric consistency, but also capture local structural details more accurately, thereby improving the final reconstruction quality.

[0067] Step (4): Model training and rendering The method of the present invention adopts an end-to-end training method to optimize the 3D Gaussian field and the control network, so that it can perform high-quality image rendering and geometric reconstruction in dynamic scenes. During the training process, we first load the D-NeRF and NeRF-DS datasets, and use the SfM technology to restore the sparse point cloud as the global geometric prior. Subsequently, a time-independent scene representation is established by initializing the 3D Gaussian set and the control network. We use a combination of L1 loss, SSIM loss, and geometric loss based on global-local depth regularization to continuously optimize the 3D Gaussian field and the control network to ensure the quality of the rendered image and the stability of the 3D structure.

[0068] As shown in Tables 1 and 2, when compared with various dynamic scene reconstruction methods on the D-NeRF and NeRF-DS datasets, the present invention achieved better results in indicators such as PSNR, SSIM and LPIPS, proving the superiority of the present invention in dynamic scene reconstruction.

[0069] Table 1 shows the performance comparison of various methods on the D-NeRF synthetic dataset

[0070] Table 2 shows the performance comparison of various methods on the NeRF-DS real dataset

[0071] In summary, experiments on dynamic scene datasets such as D-NeRF and NeRF-DS show that the present invention is superior to existing methods in terms of reconstruction quality, geometric consistency and new perspective rendering effect, improves the stability and detail fidelity of dynamic scenes, and has important research value and application prospects.

Claims

1. A dynamic Gaussian scene reconstruction method based on deep regularization, characterized in that: The strategy based on depth regularization adds a new depth loss to the image reconstruction to regulate the geometric shape of the Gaussian body during motion, including the following steps: Step 1: Select the data set and initialize the 3D Gaussian; Step 2: Introduce the control network and process it; Step 3: Calculate the loss function based on the global depth constraint, combining the image loss, image structure similarity loss and depth constraint loss; Step 4: Use an end-to-end training method to reverse gradient optimize the 3D Gaussian field and control network.

2. The dynamic Gaussian scene reconstruction method based on depth regularization according to claim 1, characterized in that: Step 1 is as follows: The SfM algorithm is used to recover sparse point clouds from multi-view images, and the recovered sparse point clouds are used as the global geometric prior of the scene; The 3DGS method is used to represent the scene as a set of 3D Gaussians, each of which is represented by a central position , and the 3D covariance matrix to describe; Probability distribution of a single 3D Gaussian Written in the form: (1); in, represents a random variable in 3D space, Represents a random variable A specific value of ; Pixel color calculation is based on point The specific formula for volume rendering is as follows: (2); (3); in, Indicate point The pixel color value after rendering, Is The transmittance is defined as represents the Gaussian distribution of colors along the light direction, and Represents the coordinates of the 3D Gaussian model projected onto the 2D image plane, represents the 2D covariance matrix; Represents the weight of adjusting the opacity contribution of the Gaussian on the image plane during projection.

3. The dynamic Gaussian scene reconstruction method based on depth regularization according to claim 1 or 2, characterized in that: Step 2 is as follows: Introduce the control network, input the time variable t and the 3D Gaussian center position , decouple motion and geometry through multi-layer perceptron MLP, transform the learning process into canonical space, thereby obtaining a time-dependent 3D Gaussian model and outputting the offset; The control network includes eight groups of linear layers and batch normalization layers; the input information is connected with the residual information of the fourth group output and then input into the fifth group; It is used to transform the 3D Gaussian in the specification space to the dynamic control space as follows: (4); in, represents the dynamic control parameters of the output, represents the control network, Indicates stopping the gradient operation. Indicates time code, Indicates positional encoding: (5); in, Indicates the position encoding order. In the synthetic scene, the 3D Gaussian center position The position encoding order is L=10, 3D Gaussian center position The corresponding position encoding order of time frame t is L=6; in the real scene, L=10 is used for both and t.

4. The dynamic Gaussian scene reconstruction method based on depth regularization according to claim 3, characterized in that: Step 3 is as follows: Step 3-1. Use L1 loss to measure the absolute error between the image rendered by the model and the target image at the pixel level. The calculation formula is as follows: (6); in, represents the target image, Represents the image rendered by the model; Step 3-2. Introduce the image structure similarity loss SSIM to measure the image rendered by the model With the target image The similarity in terms of brightness, contrast and structure is calculated as follows: (7); in, It is a function used to measure the consistency of two images in terms of brightness, contrast and structure; Step 3-3. Introduce depth constraint loss to ensure that the rendered depth is consistent with the real depth at global and local scales, including local depth normalization and scale-independent depth alignment based on Pearson correlation; Step 3-4. Add the above losses according to the weight distribution to get the final loss function It is expressed as: (8); in, , , Controls the relative weight of each loss term.

5. The method for dynamic Gaussian scene reconstruction based on depth regularization according to claim 4, characterized in that: Step 3-3 is implemented as follows: Step 3-3-1. Introduce local depth regularization to the image and target image rendered by the model under the same viewing angle to supplement the constraints of the global depth regularization on the local area; perform local normalization on the depth map to encourage the acquisition of the depth map corresponding to the image rendered by the model, and obtain the local block of the depth map ; At the same time, obtain the local block of the depth map of the target image generated by the pre-trained depth prediction transformer ; Calculate the similarity between two local blocks. The normalization process is expressed as follows: (9); in, represents the depth of the normalized local block, represents the local blocks sampled from the rendering depth map and the target depth map, respectively, Respectively represent the mean and standard deviation within the local block, represents a numerical stability term; Step 3-3-2. Define two scale-independent depth transforms: negative depth transform and reciprocal depth transform; negative depth transform is used to alleviate the problem that the distant depth value is greater than the set threshold, and reciprocal depth transform performs a reciprocal transform on the reference depth to solve the problem that the near depth is less than the set threshold but the importance is enhanced; Final depth loss Pearson correlation is used to measure the correlation between the rendered depth map and the target depth map. The specific calculation is as follows: (10); Among them, Corr(⋅,⋅) represents the Pearson correlation coefficient, represents the depth of the target depth map after the negative depth transformation, Indicates the depth of the rendered depth map, Represents the depth of the target depth map after the reciprocal depth transformation; the specific formula of Corr(⋅,⋅) is as follows: (11); in, represents the depth of the normalized local block corresponding to the target depth map, represents the normalized local block depth corresponding to the rendered depth map, represents the covariance of two depths, Represents the variance of a single depth.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method and device based on geometric prior network

    CN116597112A

  • Denoising diffusion model and local linear embedding regularization three-dimensional Gaussian sputtering method

    CN118941457A

  • Three-dimensional Gaussian scene stylization method based on text driving

    CN119006760A

  • New view angle synthesis method and system based on generative adversarial strategy and Gaussian sputtering

    CN119379548A

  • Method and apparatus for generating from a quantised image having a first bit depth a corresponding image having a second bit depth

    EP2869260A1

Cited By

  • 3D scene generation method and device guided by structure control information, equipment and medium

    CN120833439A

  • Monocular dynamic scene reconstruction method and system based on self-supervised flow matching

    CN120833442A

  • Monocular dynamic scene reconstruction method and system based on self-supervised flow matching

    CN120833442B

  • Dynamic scene reconstruction method using Kalman filter to guide Gaussian rendering

    CN121074215A

  • Bridge support dynamic digital twinborn model construction method based on 4D Gaussian sputtering

    CN121095433A