Dynamic Gaussian Scene Reconstruction Method Based on Deep Regularization
By introducing deep regularization strategies and local depth alignment technology, the problems of shape drift and instability in dynamic scene reconstruction are solved, and high-quality dynamic scene reconstruction is achieved, which improves rendering quality and geometric stability.
Patent Information
- Application Number
- CN202510480633.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-17
AI Technical Summary
When dealing with dynamic scene reconstruction methods, it is difficult to accurately capture object motion and time-varying geometric information. The lack of sufficient geometric constraints leads to shape drift and reconstruction instability, especially in the depth information generated in the time interpolation task, and fails to fully utilize the multi-view depth prior.
A strategy based on depth regularization is introduced, a monocular depth map is generated through the pre-trained depth estimation network, and the correlation between rendering depth and reference depth is used as an additional loss term, combining local depth normalization and Pearson correlation for depth alignment, optimize the rendering quality and geometric stability of the 3D Gaussian model.
It significantly improves the geometric consistency and detail fidelity of dynamic scene reconstruction, improves the rendering quality and stability of the model from a new perspective, enhances the geometric consistency of objects of different scales, and improves the generalization ability and training stability of the model.
Smart Images

Figure CN119991974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dynamic three-dimensional scene reconstruction, and mentions a new loss form based on depth regularization for geometric constraints in dynamic Gaussian reconstruction tasks. Background Art
[0002] The Gaussian reconstruction method realizes high-fidelity reconstruction and real-time rendering by explicitly modeling the scene with a set of learnable 3D Gaussian distributions. This method uses the central position, shape parameters (covariance matrix decomposed into rotation and scale), and opacity of the 3D Gaussian to form a sparse but efficient scene representation. When dealing with static scenes, such as 3D-GS has demonstrated the powerful capabilities of point-based rendering techniques, but when the scene changes dynamically, traditional static methods are difficult to capture object motion and time-varying geometric information.
[0003] However, in dynamic scenes, on the one hand, due to the continuous change of the shape and position of objects over time, it is difficult for modeling methods based on static assumptions to accurately capture these changes. In addition, objects in dynamic scenes may involve complex non-rigid deformations, making it difficult for traditional methods based on rigid body assumptions to maintain high-precision reconstruction effects. How to establish a consistent point-based representation between different time frames to ensure the stability and temporal consistency of the reconstruction remains a key challenge.
[0004] On the other hand, due to the lack of sufficient geometric constraints, especially when dealing with time interpolation tasks, the generated depth information may exhibit jitter phenomena. This jitter may stem from the rapid movement of objects, occlusion changes, or uneven data sampling, resulting in unstable geometric structures in the time dimension. In addition, due to the limited utilization of the geometric information of the input image by current dynamic Gaussian reconstruction methods and the failure to fully combine multi-view depth priors, there are still certain limitations in the scale, details, and consistency of the reconstructed 3D structures. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to propose a dynamic Gaussian scene reconstruction method based on depth regularization for the lack of constraints on the geometric structure of Gaussian bodies in current dynamic scene reconstruction, resulting in shape drift and instability. The strategy based on depth regularization can add a new depth loss to the reconstruction of images to regularize the geometric shape during the movement of Gaussian bodies.
[0006] Specifically, by using a pre-trained depth estimation network (DPT) to generate a monocular depth map and taking the correlation between the rendered depth and the reference depth as an additional loss term, the model can obtain a depth-related loss. This loss calculates the Pearson correlation coefficient between the rendered depth and the reference depth to constrain the structural consistency between the two, thereby providing a stronger geometric prior and making the reconstruction results have better consistency and detail performance from new perspectives.
[0007] A dynamic Gaussian scene reconstruction method based on depth regularization, comprising the following steps:
[0008] Step 1, select a dataset and initialize the 3D Gaussian;
[0009] First, select the D-NeRF and NeRF-DS datasets, and use the SfM algorithm to recover the sparse point cloud from multi-view images, and use the recovered sparse point cloud as the global geometric prior of the scene.
[0010] To efficiently and differentiably represent and render the scene, the 3D Gaussian Splatting method is used to represent the scene as a set of 3D Gaussians, and each 3D Gaussian is represented by the center position , and the 3D covariance matrix to describe.
[0011] Specifically, the probability distribution of a single 3D Gaussian is written in the form of:
[0012] (1);
[0013] where, represents a random variable in a 3D space, represents a specific value of the random variable .
[0014] Pixel color calculation uses volume rendering based on points , and the specific formula is as follows:
[0015] (2);
[0016] (3);
[0017] where, represents the pixel color value after rendering the point , is the transmittance defined by , represents the color of the Gaussian distribution along the ray direction, while represents the coordinates of the 3D Gaussian model projected onto the 2D image plane, Represents a 2D covariance matrix. Represents the weight for adjusting the opacity of the Gaussian volume contribution on the image plane during the projection process.
[0018] Step 2: Introduce the control network and process it;
[0019] The present invention introduces a control network, which inputs the time variable and the 3D Gaussian center position , decouples the motion and geometric structure through a multi-layer perceptron MLP, transforms the learning process into the canonical space, thereby obtaining a time-related 3D Gaussian model, and outputs the offset.
[0020] The control network introduced in the present invention is improved based on the multi-layer perceptron MLP. The control network adopted in the present invention includes eight groups of linear layers and batch normalization layers; among them, the input information is input into the fifth group after residual connection with the information output by the fourth group, as Figure 2 shown.
[0021] For transforming the 3D Gaussian in the canonical space to the dynamic control space, specifically as follows:
[0022] (4);
[0023] Among them, represents the output dynamic control parameter, represents the control network, represents the stop gradient operation, represents the time encoding, represents the position encoding:
[0024] (5);
[0025] Among them, represents the order of the position encoding. In the synthetic scene, in the synthetic scene, the position encoding order L of the 3D Gaussian center position is L = 10, and the position encoding order of the time frame t corresponding to the 3D Gaussian center position takes L = 6; while in the real scene, and t both adopt L = 10.
[0026] Step 3: Calculate the loss function based on the global depth constraint;
[0027] In order to optimize the rendering quality and geometric stability of the 3D Gaussian model and the structural control of the control network, the present invention combines the image loss, the image structure similarity loss, and the depth constraint loss for optimization;
[0028] Step 3-1. To optimize the reconstruction accuracy of the image, the L1 loss is used to measure the absolute error between the image rendered by the model and the target image at the pixel level. Its calculation formula is as follows
[0029] (6);
[0030] where represents the target image, represents the image rendered by the model.
[0031] Step 3-2. To improve the details and structural consistency of the image, the Structural Similarity Index Measure (SSIM) loss is introduced to measure the similarity between the image rendered by the model and the target image in terms of brightness, contrast, and structure. Its calculation formula is as follows:
[0032] (7);
[0033] where is a function used to measure the consistency between two images in terms of brightness, contrast, and structure. Compared with the simple pixel difference.
[0034] Step 3-3. In terms of 3D structure optimization, a depth constraint loss is introduced to ensure that the rendered depth is consistent with the true depth at both the global and local scales. This method includes Local Depth Normalization (LDN) and scale-invariant depth alignment based on Pearson correlation to improve the robustness of the model in different depth-of-field ranges.
[0035] Step 3-3-1. Local depth regularization is introduced for the images rendered by the model and the target image from the same perspective to supplement the global depth regularization's constraint on the local area. In particular, the present invention performs local normalization on the depth map, encourages obtaining the depth map corresponding to the image rendered by the model, and obtaining the local patch ; at the same time, obtaining the local patch of the depth map of the target image generated by the pre-trained Depth Prediction Transformer (DPT) ; calculating the similarity between the two local patches. This normalization process is expressed as follows:
[0036] (8);
[0037] where represents the depth of the normalized local patch, represents the local patches sampled from the rendered depth map and the target depth map respectively, represent the mean and standard deviation within the local patch respectively,
[0038] Step 3-3-2. Define two scale-invariant depth transformations: negative depth transformation and reciprocal depth transformation; the negative depth transformation is used to alleviate the problem that the depth value of the far view is greater than the set threshold, and the reciprocal depth transformation performs a reciprocal transformation on the reference depth to solve the problem of enhancing the importance when the near view depth is less than the set threshold.
[0039] The final depth loss is measured using Pearson correlation to measure the correlation between the rendered depth map and the target depth map, and the specific calculation is as follows:
[0040] (9);
[0041] where Corr(⋅,⋅) represents the Pearson correlation coefficient, represents the depth of the target depth map after negative depth transformation, represents the depth of the rendered depth map, represents the depth of the target depth map after reciprocal depth transformation. The specific formula of Corr(⋅,⋅) is as follows:
[0042] (10);
[0043] where, represents the depth of the normalized local block corresponding to the target depth map, represents the depth of the normalized local block corresponding to the rendered depth map, represents the covariance of the two depths, represents the variance of a single depth.
[0044] Step 3-4. Add the above-mentioned losses according to the weight distribution to obtain the final loss function which is expressed as:
[0045] (11);
[0046] where, , , controls the relative weights of each loss term.
[0047] Step 4. Adopt an end-to-end training method to optimize the 3D Gaussian field and the control network in the reverse gradient direction, so that they can accurately model the geometry and motion information of the dynamic scene and achieve high-quality dynamic scene reconstruction.
[0048] The beneficial effects of the present invention are as follows:
[0049] The present invention proposes a dynamic Gaussian scene reconstruction method based on depth regularization to solve the problems of insufficient geometric constraints, shape drift, and unstable rendering quality in existing dynamic Gaussian reconstruction methods. By introducing global and local depth geometric regularization, the present invention significantly improves the geometric consistency and detail fidelity of 3D Gaussian models in dynamic scenes. Utilizing the geometric prior provided by a monocular depth estimation network and adopting negative depth contrast and reciprocal depth contrast for scale-invariant depth alignment, the present invention effectively enhances the stability of 3D structures in dynamic scenes.
[0050] Local depth normalization (LDN) further enhances the detail reconstruction ability, ensuring geometric consistency for objects at different scales. Meanwhile, the optimized L1+SSIM+depth constraint loss design improves the generalization ability and training stability of the model. Experiments show that this method is superior to existing methods in terms of rendering quality, geometric stability, and novel view synthesis, with improvements in image rendering metrics such as PSNR, SSIM, and LPIPS on the dynamic scene synthesis dataset D-NeRF and the real dataset NeRF-DS, having important research value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a complete flowchart of the dynamic Gaussian scene reconstruction method based on depth regularization of the present invention.
[0052] Figure 2 It is a control network structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0053] The method of the present invention and its detailed parameters are further specifically described below in conjunction with the accompanying drawings.
[0054] Step (1), Initialize 3D Gaussian
[0055] First, select the dynamic scene synthesis dataset D-NeRF and the real dataset NeRF-DS, and use the SfM (Structure-from-Motion) technology to recover the sparse point cloud from multi-view images as the global geometric prior of the scene.
[0056] To efficiently and differentiably represent and render the scene, we adopt the 3D Gaussian Splatting method to represent the scene as a set of 3D Gaussians.
[0057] Each 3D Gaussian can be represented by its center position , and the covariance matrix .
[0058] Specifically, the probability distribution form of a single Gaussian volume can be written as:
[0059] (12);
[0060] To make the learning process of the 3D Gaussian model easier, it is split into two learnable parts: the quaternion r represents rotation, and the three-dimensional vector s represents scaling. Then, these parts are converted into the corresponding rotation matrix R and scaling matrix S. The resulting 3D covariance matrix can be expressed as:
[0061] (13);
[0062] To optimize the parameters of the 3D Gaussian model in the canonical space, it is crucial to differentially render 2D images from these 3D Gaussian models. In this invention, the present invention adopts a differential Gaussian rasterization process. The 3D Gaussian model can be projected onto a 2D plane and rendered for each pixel using the following 2D covariance matrix for rendering:
[0063] (14);
[0064] where J is the Jacobian matrix of the affine approximation of the projection transformation, represents the pixel color represented on the image plane of the view matrix for converting from world coordinates to camera coordinates, and is obtained by sequentially rendering through point-based volume rendering technology:
[0065] (15);
[0066] (16);
[0067] where, represents the point rendered pixel color value, is the transmittance defined by , represents the color of the Gaussian distribution along the ray direction, and represents the coordinates of the 3D Gaussian model projected onto the 2D image plane, represents the 2D covariance matrix. represents the weight for adjusting the opacity of the contribution of the Gaussian volume on the image plane during the projection process.
[0068] 3DGS uses an adaptive density control strategy to manage the number of Gaussian volumes. For under-reconstructed regions, the Gaussian volumes are copied and moved along the position gradient direction; while large Gaussian distributions in high-variance regions are split into smaller Gaussian volumes. In addition, the algorithm also deletes Gaussian volumes with opacity lower than the threshold within a certain number of iterations.
[0069] Sparse point clouds are recovered from multi-view images using SfM as the global geometric prior of the scene, and 3D Gaussian projection is initialized to represent the scene as a set of 3D Gaussians.
[0070] Step (2), introduce the control network and process
[0071] An intuitive way to model a dynamic scene using 3D Gaussians is to train a set of 3D Gaussian splashes separately in each time-related view set, but it has limitations for continuous monocular captures in a time series.
[0072] We use a control network that works in tandem with the 3D Gaussian model to decouple motion and geometric structure, transform the learning process into the canonical space, and obtain a time-independent 3D Gaussian model. This decoupling method introduces the geometric prior information of the scene, relating the changes in the positions of the 3D Gaussian model to time and coordinates. The core of the control network is a multi-layer perceptron (MLP).
[0073] The control network introduced in the present invention is improved based on the multi-layer perceptron MLP. The control network adopted in the present invention includes 8 groups of linear layers and batch normalization layers; among them, the input information is connected with the information residual output by the fourth group and then input into the fifth group, as Figure 2 shown.
[0074] Taking time and the central position of the 3D Gaussian as inputs, the control network generates offsets that will then transform the 3D Gaussian in the canonical space to the control space:
[0075] (17);
[0076] where represents the stop gradient operation, represents the position encoding:
[0077] (18);
[0078] In the synthetic scene, the position encoding order L of the 3D Gaussian central position is 10, and the position encoding order of the time frame t corresponding to the 3D Gaussian central position takes L = 6; while in the real scene, and t both adopt L = 10. We set the depth D of the control network to 8 and the hidden layer width W to 256. As shown in Table 1 and Table 2, relevant experiments show that using position encoding in the input of the control network can enhance the details of the rendering results. Step (3), calculate the loss function based on the global depth constraint
[0079] 3-1 Image Loss
[0080] To ensure the pixel-level proximity between the model-rendered image and the target image, we adopt the L1 loss as the reconstruction loss, which is defined as follows:
[0081] (19);
[0082] where denotes the target image, denotes the image rendered by the model. This loss directly compares the absolute differences of each pixel, ensuring the accurate restoration of information such as overall brightness and color. We use the weight for regulation.
[0083] 3-2 Structural Similarity Index Measure Loss (SSIM Loss)
[0084] To further improve the details and structural consistency of the generated image, we introduce the SSIM loss, which is defined as follows:
[0085] (20);
[0086] where is used to measure the consistency between two images in terms of brightness, contrast, and structure. Compared with the simple pixel difference, the loss is more robust in the fidelity of local structures such as edges and textures, and can effectively improve the overall visual effect of the image.
[0087] 3-3 Depth Constraint Loss
[0088] In the dynamic scene reconstruction task, depth information provides crucial geometric priors for 3D Gaussian volumes, helping to ensure the stability of the 3D structure. However, due to the constantly changing camera viewpoints in dynamic scenes, the scene may contain objects of different scales, and there is a large depth span between the foreground and the background. This variation in depth scales makes it possible that simply relying on global depth regularization may not be sufficient to fully recover complex geometric structures. Global depth often focuses on the overall shape constraints but tends to overlook local details, resulting in geometric structure drift or incomplete reconstruction in some regions.
[0089] To solve the above problems, we introduce Local Depth Regularization (LDR) on the basis of global depth regularization, so that the rendered depth is consistent with the true depth at the local scale. Specifically, this method includes Local Depth Normalization (LDN) and scale-invariant depth alignment based on Pearson correlation to improve the robustness of the model in different depth-of-field ranges.
[0090] ① Local Depth Normalization;
[0091] Previous studies have used global depth information to assist 3D geometry reconstruction. However, this method performs poorly in complex scenes containing objects of multiple scales. In particular, global depth information usually focuses on global features while ignoring the fine details in the depth information, which may lead to the loss of local geometric shapes.
[0092] To solve this problem, the present invention introduces local depth regularization to the images rendered from the model and the target image under the same perspective to supplement the constraint of global depth regularization on the local region. In particular, the present invention performs local normalization on the depth map, encourages obtaining the depth map corresponding to the image rendered from the model, and obtains the local block of this depth map ; at the same time, obtain the local block of the depth map of the target image generated by the pre-trained Depth Prediction Transformer (DPT) ; calculate the similarity between the two local blocks, and this normalization process is expressed as follows:
[0093] (21);
[0094] Where represents the depth of the normalized local block, represents the local blocks sampled from the rendered depth map and the target depth map respectively, represent the mean and standard deviation within this local block respectively,
[0095] ② Scale-invariant depth alignment based on Pearson correlation;
[0096] In a dynamic scene, due to the change of the camera perspective, the depth-of-field range of each frame of picture is different, and there will be significant differences in the depth distribution of the near and far regions. Therefore, matching based only on the original depth values may lead to errors. To improve the stability of depth alignment, we use negative depth contrast and reciprocal depth contrast to consider the similarity at different depth scales simultaneously.
[0097] Considering that the depth values estimated by the DPT deep model are negative, we define two scale-invariant depth transformations:
[0098] Negative depth transformation: For the reference depth map take the opposite number, i.e., - , which is used to alleviate the problem of large depth values in the far scene.
[0099] Reciprocal depth transformation: Perform a reciprocal transformation on the reference depth, i.e., , which is used to solve the problem of enhancing the importance of small but important near-scene depths.
[0100] The final depth loss is measured using Pearson correlation to measure the correlation between the rendered depth map and the target depth map. The specific calculation is as follows:
[0101] (22);
[0102] where Corr(⋅,⋅) represents the Pearson correlation coefficient, represents the depth of the target depth map after negative depth transformation, represents the depth of the rendered depth map, represents the depth of the target depth map after reciprocal depth transformation. The specific formula for Corr(⋅,⋅) is as follows:
[0103] (23);
[0104] where, represents the depth of the normalized local block corresponding to the target depth map, represents the depth of the normalized local block corresponding to the rendered depth map, represents the covariance of the two depths, represents the variance of a single depth.
[0105] ③ Loss calculation;
[0106] Add the above losses according to the weights to obtain the final loss function which is expressed as:
[0107] (24);
[0108] where, , , controls the relative weights of each loss term. By introducing local depth regularization, the model can not only maintain global geometric consistency but also capture local structural details more precisely, thereby improving the final reconstruction quality.
[0109] Step (4), model training and rendering
[0110] The method of the present invention adopts an end-to-end training method to optimize the 3D Gaussian field and the control network, enabling it to perform high-quality image rendering and geometric reconstruction in dynamic scenes. During the training process, we first load the D-NeRF and NeRF-DS datasets, and use the SfM technique to recover the sparse point cloud as the global geometric prior. Subsequently, by initializing the 3D Gaussian set and the control network, a time-independent scene representation is established. We comprehensively use the L1 loss, SSIM loss, and geometric loss based on global-local depth regularization to continuously optimize the 3D Gaussian field and the control network to ensure the quality of the rendered image and the stability of the 3D structure.
[0111] As shown in Table 1 and Table 2, when compared with various dynamic scene reconstruction methods on the D-NeRF and NeRF-DS datasets, the present invention achieves better results in metrics such as PSNR, SSIM, and LPIPS, demonstrating the superiority of the present invention in dynamic scene reconstruction.
[0112] Table 1 shows the performance comparison of various methods on the D-NeRF synthetic dataset
[0113]
[0114] Table 2 shows the performance comparison of various methods on the NeRF-DS real dataset
[0115]
[0116] In summary, experiments on dynamic scene datasets such as D-NeRF and NeRF-DS show that the present invention is superior to existing methods in terms of reconstruction quality, geometric consistency, and new view rendering effect, improving the stability and detail fidelity of dynamic scenes, and having important research value and application prospects.
Claims
1. A dynamic Gaussian scene reconstruction method based on depth regularization, characterized in that The strategy based on depth regularization adds a new depth loss to the reconstruction of images, which is used to regularize the geometry during the Gaussian volume movement process, including the following steps: Step 1: Select a dataset and initialize the 3D Gaussian; Step 2: Introduce and process the control network; Step 3: Calculate the loss function based on global depth constraints, combining image loss, image structural similarity loss, and depth constraint loss, specifically including: Step 3-3. Introduce the depth constraint loss L depth , to ensure that the rendered depth is consistent with the true depth at both the global and local scales, specifically including local depth normalization and scale-invariant depth alignment based on Pearson correlation; Step 3-3-1. Introduce local depth regularization to the images rendered from the model under the same perspective and the target images to supplement the constraints of global depth regularization on local regions; perform local normalization on the depth map to encourage obtaining the depth map corresponding to the image rendered from the model, and obtain the local blocks of this depth map Meanwhile, obtain the local blocks of the depth map of the target image generated by the pre-trained depth prediction transformer Calculate the similarity between the two local blocks, and the normalization process is expressed as follows: where d LN (x) represents the depth of the normalized local patch, represent local patches sampled from the rendered depth map and the target depth map respectively, μ” and σ represent the mean and standard deviation within the local patch respectively, and ∈ represents a numerical stability term; Step 3-3-2. Define two scale-invariant depth transformations: negative depth transformation and reciprocal depth transformation; the negative depth transformation is used to alleviate the problem that the far-view depth value is greater than the set threshold, and the reciprocal depth transformation performs a reciprocal transformation on the reference depth to solve the problem that the near-view depth is less than the set threshold but the importance is enhanced; The final depth constraint loss L depth It is measured using Pearson correlation to measure the correlation between the rendered depth map and the target depth map, and the specific calculation is as follows: Among them, Corr(·,·) represents the Pearson correlation coefficient, represents the depth of the target depth map after negative depth transformation, represents the depth of the rendered depth map, represents the depth of the target depth map after reciprocal depth transformation; the specific formula of Corr(·,·) is shown as follows: Among them, d LN represents the depth of the normalized local block corresponding to the target depth map, represents the depth of the normalized local block corresponding to the rendered depth map, Cov(·,·) represents the covariance of two depths, and Var(·) represents the variance of a single depth; Step 4: Adopt an end-to-end training method to optimize the 3D Gaussian field and the control network by backpropagation gradient.
2. The method for dynamically reconstructing a Gaussian scene based on depth regularization according to claim 1, wherein Step 1 is specifically as follows: Use the SfM algorithm to recover the sparse point cloud from multi-view images, and use the recovered sparse point cloud as the global geometric prior of the scene; The 3DGS method is used to represent a scene as a set of 3D Gaussians, each described by a center position and a 3D covariance matrix ; The probability distribution G(X) form of a single 3D Gaussian is written as: where X represents a random variable in 3D space, and x represents a specific value of the random variable X; Pixel color calculation adopts volume rendering based on point p, and the specific formula is as follows: C(p) = ∑T i α i c i Among them, C(p) represents the pixel color value after rendering of point p, and T i is the transmittance defined by , c i represents the color of the Gaussian distribution along the ray direction, and μ' i represents the coordinates of the 3D Gaussian model projected onto the 2D image plane, and Σ' represents the 2D covariance matrix; σ i represents the weight for adjusting the opacity contribution of the Gaussian volume on the image plane during the projection process.
3. The method for reconstructing a dynamic Gaussian scene based on depth regularization according to claim 1 or 2, characterized in that Step 2 is specifically as follows: Introduce the control network, input the time variable t and the 3D Gaussian center position μ, decouple the motion and geometric structure through the multi-layer perceptron MLP, transform the learning process into the canonical space, so as to obtain a time-related 3D Gaussian model and output the offset; The control network includes eight groups of linear layers and batch normalization layers; among them, the input information is connected in residual with the information output by the fourth group and then input into the fifth group; It is used to transform the 3D Gaussian in the canonical space into the dynamic control space, specifically as follows: (δx, δr, δs) = F θ (γ(sg(x)), γ(t)) where δx, δr, δs represent the output dynamic control parameters, F θ represents the control network, sg(·) represents the stop gradient operation, γ(t)) represents the time encoding, and γ(sg(x)) represents the position encoding: where L represents the position encoding order. In the synthetic scene, the position encoding order L of the 3D Gaussian center position μ is 10, and the position encoding order of the time frame t corresponding to the 3D Gaussian center position μ is taken as L = 6; while in the real scene, both μ and t adopt L = 10.
4. The method for dynamically reconstructing a Gaussian scene based on depth regularization according to claim 1, wherein Step 1 is specifically as follows: Step 3-1. Use the L1 loss to measure the absolute error between the image rendered by the model and the target image at the pixel level, and its calculation formula is as follows Among them, represents the target image, and I represents the image obtained by model rendering; Step 3-2. Introduce the structural similarity loss of the image SSIM, which is used to measure the similarity between the image I rendered by the model and the target image in terms of brightness, contrast, and structure. Its calculation formula is as follows: where SSIM is a function used to measure the consistency of two images in terms of brightness, contrast, and structure; Step 3-3. Introduce the depth constraint loss L depth , to ensure that the rendered depth is consistent with the true depth at both the global and local scales, specifically including local depth normalization and scale-invariant depth alignment based on Pearson correlation; Step 3-4. Add the above losses according to the weight distribution to obtain the final loss function L total Expressed as: L total = λ1L1 + λ2L SSIM + λ3L depth where λ1 = 0.8, λ2 = 0.2, λ3 = 0.01 control the relative weights of each loss term.