A 3D Scene Reconstruction and Real-Time SLAM System and Method Based on 2D Gaussian Sputtering

By combining monocular geometric priors and consistent views, the reconstruction distortion problem of SLAM method in complex environments is solved, efficient and accurate three-dimensional scene reconstruction and depth estimation are achieved, and the performance of the SLAM system is improved.

CN120070794BActive Publication Date: 2025-07-08YANTAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510541203.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-08
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Existing SLAM methods have poor performance in strong light, radiation variation, and dynamic or low texture environments, and there is distortion during sparse 3D modeling and reconstruction, making it difficult to achieve high-fidelity view synthesis and accurate geometric reconstruction.

Method used

Using a three-dimensional scene reconstruction and real-time SLAM system based on 2D Gaussian sputtering, combining monocular geometric priors and learning-based intensive SLAM, the map is constructed through 2D Gaussian and guided camera pose optimization to achieve efficient and accurate scene reconstruction.

Benefits of technology

High-quality indoor scene reconstruction is achieved, the accuracy of depth estimation and the accuracy of camera pose estimation are enhanced, and exposure compensation and Gaussian pruning are supported, ensuring rendering quality and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070794B_ABST
    Figure CN120070794B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of three-dimensional reconstruction, and specifically relates to a three-dimensional scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering. The method includes sub-map construction and sub-map optimization after sub-map initialization based on the collected video stream; performing frame-to-model tracking based on the constructed map, and jointly optimizing the key frame poses and sub-maps through local bundle adjustment; a view-consistent 2D Gaussian scene representation, calculating the intersection points of the camera rays and the 2D Gaussian and performing depth estimation; merging all sub-maps into a global map, optimizing the global map, performing Gaussian pruning and exposure compensation to obtain a color map, a depth map, and a normal map, and completing three-dimensional scene reconstruction according to the color map, the depth map, and the normal map; the system includes a data acquisition module, a map construction module, a tracking module, a scene representation module, a data optimization module, and a data output module; through the technical solution of the present invention, high-quality indoor scene reconstruction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and in particular to a three-dimensional scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering. Background Art

[0002] SLAM was originally designed for robots and automation systems. Today, its application requirements have expanded to multiple fields, including augmented reality (AR), visual surveillance, medical applications, etc. To meet these requirements, researchers are developing methods that enable machines to autonomously build increasingly accurate scene representations, leveraging advancements in robotics, computer vision, sensor technology, and artificial intelligence (AI).

[0003] Over the years, SLAM methods have evolved significantly to adapt to these specific requirements. Initially, hand-designed algorithms demonstrated excellent real-time performance and scalability. However, they had problems in strong lighting, radiation changes, and dynamic or low-texture environments, resulting in less-than-satisfactory performance. Advanced techniques that incorporate deep learning methods are crucial for improving the accuracy and reliability of localization and mapping. By leveraging the powerful feature extraction capabilities of deep neural networks, SLAM is particularly effective in challenging conditions. However, deep learning methods rely on large amounts of training data sets and precise ground truth annotations, which limit their generalization ability in unseen scenes. In addition, both hand-designed methods and deep learning-based methods are limited to using discrete surface representations, such as point / surface clouds, voxel hashing, voxel grids, and octrees. These representations lead to sparse 3D modeling, limited spatial resolution, and distortion during the reconstruction process. Accurately estimating the geometry of unobserved regions remains an unsolved challenge. Recent research has focused on improving map quality and density. With the emergence of neural scene representations (such as Neural Radiance Fields, NeRF), which enable high-fidelity view synthesis, dense neural SLAM methods have developed rapidly. Although these methods have made impressive progress in scene representation quality, they are still limited to small-scale synthetic scenes, and the re-rendered results are far from realistic. Recently, 3DGS has been shown to be comparable to NeRF in rendering quality, while achieving an order-of-magnitude improvement in rendering and training speed. This representation is explicit and actionable, making it a strong candidate for downstream tasks. However, due to the lack of geometric priors and Figure 1 consistency, their geometric reconstruction performance has not yet surpassed dense neural SLAM methods. The present invention aims to solve this problem.

[0004] This paper presents a Gaussian-based SLAM system that uses RGB-D input to achieve accurate and fast monocular scene reconstruction. The method of the present invention models the scene in an efficient and accurate manner by combining monocular geometric priors and learning-based dense SLAM, while using 2DGS as a compact map representation, improving the accuracy of geometric estimation. The present invention uses 2DGS to construct a map and uses the existing map to guide camera pose optimization. This design ensures efficiency and accuracy, breaks the trade-off between speed and quality, thus achieving fast and accurate surface reconstruction while maintaining high-quality results, and successfully overcomes the limitations of previous methods. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention develops a three-dimensional scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering, which can achieve high-quality indoor scene reconstruction, and provides an enhanced depth estimation method and a camera pose estimation method.

[0006] The technical solution for the present invention to solve the technical problems is: A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, comprising the following steps:

[0007] Collect the RGB-D video stream of the target scene;

[0008] Based on the first frame of the RGB-D video stream, use the initialization strategy of camera motion to create a sub-map, initially estimate and optimize the camera pose for the key frames of the created sub-map, and optimize the sub-map through loss calculation;

[0009] Perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment;

[0010] Based on the optimized sub-map, use Figure 1 consistent 2D Gaussian for scene representation, calculate the intersection of the camera ray and the 2D Gaussian and perform unbiased depth estimation;

[0011] Merge all optimized sub-maps into a global map, and optimize the color of the global map and the geometric properties of the Gaussian by minimizing the loss function;

[0012] Perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the three-dimensional scene reconstruction based on the color map, the depth map, and the normal map.

[0013] The initialization strategy of the camera motion includes that the translation amount of the current frame relative to the first frame of the active sub-map exceeds a predefined threshold , or when the estimated Euler angle exceeds a predetermined threshold When a new sub-map is created, only the current active sub-map is processed at any given time point.

[0014] After the camera pose is initially estimated for the key frames of the created sub-map using visual odometry, a dense point cloud is calculated using the RGB-D measurements of the key frames. At the beginning of each sub-map, a number of points are uniformly sampled from the point cloud of the key frame, and regions with high color gradients are preferentially selected to add new Gaussian models. For subsequent key frames in the sub-map, a number of points are uniformly sampled from regions where the transparency value is lower than a certain threshold. A new Gaussian model is added based on the sampled points within the search radius of the current sub-map when there are no neighbors in the neighborhood of the current sub-map. a number of points, and regions with high color gradients are preferentially selected to add new Gaussian models; for subsequent key frames in the sub-map, a number of points are uniformly sampled from regions where the transparency value is lower than a certain threshold a number of points, and the new Gaussian model is added when there are no neighbors in the neighborhood of the current sub-map based on the search radius of the current sub-map of the sampled points.

[0015] The optimization of the sub-map by loss calculation includes calculating the minimum loss to optimize the sub-map for rendering the color and depth of the key frames;

[0016] The minimum loss includes color loss and depth loss, and the minimum loss function is:

[0017] ,

[0018] where is the color loss, is the real color image, is the predicted color image, reflects the structural similarity between the predicted color map and the real color map; is the depth loss, is the real depth image, is the predicted depth image, = 0.8, is used to control the weights of color difference and structural difference, represents the 1-norm, that is and .

[0019] The accurate estimation of the camera pose is achieved by calculating the tracking loss, and the camera pose is optimized by local bundle adjustment, including introducing a soft mask and an error mask , the soft mask is directly from the transparency map of the active sub-map, and the error boolean mask eliminates all pixels where the color or depth error exceeds the threshold to obtain the tracking loss function; the local bundle adjustment loss function that keeps the Gaussian parameters unchanged is used to optimize the camera pose;

[0020] The tracking loss function is:

[0021] ,

[0022] Among them: Used to control the weights of color difference and depth difference;

[0023] The local bundle adjustment loss function is:

[0024] ,

[0025] Wherein, Represents the tracking loss, the surface-to-surface projection error Represents the distance between the true surface observation points from two different viewpoints, wherein, Represents from the th frame camera to the th frame camera's rigid body transformation, Represents the 3D coordinates of the object surface observation under the current relative transformation mapping, Represents the th frame camera coordinate system's 3D coordinates of the object surface; the Gaussian-to-Gaussian projection error Represents the distance between the Gaussian points from two different viewpoints, Represents the 3D coordinates of the Gaussian observation under the current relative transformation mapping, Represents the th frame camera's 3D coordinates of the Gaussian observation.

[0026] The method of using a visually Figure 1 consistent 2D Gaussian for scene representation, calculating the intersection of the camera ray and the 2D Gaussian and performing unbiased depth estimation, including compressing the 3D Gaussian into a 2D Gaussian by removing one dimension of the scaling matrix, using the 2D Gaussian for scene representation, solving the intersection of the camera ray and the 2D Gaussian and calculating the unbiased depth to determine the Gaussian depth;

[0027] The expression of the 2D Gaussian is:

[0028] ,

[0029] Wherein, and are the coordinates in the local coordinate system, and the parameters of the 2D Gaussian include: is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, serves as the scaling matrix, serves as the rotation matrix, wherein represents the normal vector of the 2D Gaussian, and the spatial transformation matrix of the 2D Gaussian The expression is:

[0030] ,

[0031] where, is the projection matrix of the local coordinate system of the 2D Gaussian, is the center position of the 2D Gaussian, is the two principal tangent vectors of the 2D Gaussian, is the scaling factor of the two principal tangent vectors of the 2D Gaussian; is the transformation matrix from world coordinates to pixel coordinates, and the formula for the 3D spatial point

[0032] is: The formula is:

[0033] ,

[0034] where, represents the uniform light emitted from the camera, passing through the pixel and intersecting with the 2D Gaussian at the depth , is the homogeneous coordinate of the 3D spatial point, is the transformation matrix from the world coordinate system to the pixel coordinate system, is the local projection matrix of the 2D Gaussian, is the homogeneous coordinate of a point in the local coordinate system of the 2D Gaussian, is the transpose of the matrix;

[0035] The intersection point of the camera ray and the 2D Gaussian is:

[0036] ,

[0037] where: and are the i-th parameters of the plane vector;

[0038] The unbiased depth estimation function is:

[0039] ,

[0040] where, is the opacity of the j-th Gaussian, is the opacity of the j-th pixel, is the intersection point of the ray and the Gaussian, and the vector is calculated from the extrinsic parameters of the camera, the scaling matrix and the rotation matrix of the Gaussian, is the Gaussian mean.

[0041] To improve the accuracy of unbiased depth calculation, Gaussian optimization must be performed. The Gaussian optimization includes regularization loss and normal loss calculation. The regularization loss includes scale regularization that restricts the Gaussian scale and scale ratio regularization that prevents excessive stretching of the Gaussian. The normal loss enables the 2D Gaussian to be tangent to the actual surface;

[0042] The regularization loss function is:

[0043] ,

[0044] where: Scale regularization in controls the regularization weight, = 0.05, is the scaling matrix of the i-th 2D Gaussian, is the maximum scale threshold, is the number of Gaussians exceeding the threshold; Scale ratio regularization in and are the major axis vector and minor axis vector of the i-th 2D Gaussian respectively, controls the regularization weight, = 0.05, is the maximum scale ratio threshold, is the number of all Gaussians exceeding the threshold;

[0045] The normal loss function is:

[0046] ,

[0047] where, is the mixing weight of the current intersection point, where , is the Gaussian opacity at the j-th pixel, is the opacity at the j-th pixel. is the normal of the current 2D Gaussian in the camera direction, is the object surface normal estimated from the real depth map, and are the gradients of the depth map in and directions respectively.

[0048] Optimizing the color of the global map and the geometric properties of the Gaussian by minimizing the loss function includes calculating color loss, depth loss, normal loss, and regularization loss for the global map to optimize the global map. The minimizing loss function is:

[0049] ,

[0050] Among them, is the color loss, is the depth loss, is the normal loss, is the regularization loss.

[0051] Performing Gaussian pruning and exposure compensation on the optimized global map includes a pruning method based on color contribution to judge the importance by evaluating the color contribution of each Gaussian, pruning the Gaussians that have no color contribution to the final rendering, and using a method of applying a transformation matrix to the key-frame images for correction to calculate the exposure compensation. The formula for the multi-view contribution of Gaussian pruning is:

[0052] ,

[0053] Among them, is the multi-view set with the highest contribution,

[0054] is the formula for the single-view contribution of Gaussian pruning. Among them, is the color contribution of the Gaussian at the j-th view, is the 2D projection area of the Gaussian at the j-th view, is the pixel value at the j-th view, is the opacity of the current 2D Gaussian, is the color value of the current 2D Gaussian, is the accumulated opacity of the current pixel, is a factor used to control the balance between opacity and color contribution. When , the color-based pruning degrades to transparency-based pruning;

[0055] The exposure compensation formula is:

[0056] ,

[0057] Among them, is the color map after exposure compensation, is the color map predicted by the current frame, is a transformation matrix used to adjust the intensity and proportional relationship of the image color channels. By independently adjusting different color channels, it can effectively correct the color deviation caused by lighting changes, is a bias vector.

[0058] A three-dimensional scene reconstruction and real-time SLAM system based on 2D Gaussian sputtering includes:

[0059] A data acquisition module configured to acquire the RGB-D video stream of the target scene;

[0060] A map construction module, configured to create a sub-map based on the first frame of an RGB-D video stream using an initialization strategy for camera motion, initially estimate the camera pose for the key frames of the created sub-map using visual odometry, and optimize the sub-map through loss calculation;

[0061] A tracking module, configured to perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment;

[0062] A scene representation module, configured to perform scene representation based on the optimized sub-map using Figure 1 consistent 2D Gaussians, calculate the intersections of camera rays with the 2D Gaussians and perform unbiased depth estimation;

[0063] A data optimization module, configured to merge all optimized sub-maps into a global map, and optimize the color of the global map and the geometric properties of the Gaussians by minimizing a loss function;

[0064] A data output module, configured to perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the reconstruction of a three-dimensional scene based on the color map, the depth map, and the normal map.

[0065] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects:

[0066] 1. By combining monocular geometric priors and Figure 1 consistent scene representations, high-quality indoor scene reconstruction is achieved;

[0067] 2. Through an enhanced depth estimation method, the depth estimation deviation is corrected by the camera ray and Gaussian intersection method, achieving more accurate depth estimation;

[0068] 3. To improve the accuracy of pose estimation, the present invention introduces an additional bundle adjustment (BA) optimization strategy during the frame-to-model tracking process. This optimization is achieved by minimizing the surface-to-surface and Gaussian-to-Gaussian projection errors, thereby obtaining a more accurate camera pose estimation;

[0069] 4. The system supports an offline processing method for exposure compensation and Gaussian pruning, which improves the rendering quality while reducing the storage overhead;

[0070] 5. The system can render high-fidelity RGB images, accurate depth maps, and normal maps in real time, and still maintain stability and noise resistance in complex scenarios. Description of the Drawings

[0071] Figure 1 This is the overall flowchart of a 3D scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to Embodiment 1 of the present invention;

[0072] Figure 2 This is the rendering effect diagram of Embodiment 1 of the present invention;

[0073] Figure 3 This is the comparison diagram between the reconstruction result and the real scene of Embodiment 1 of the present invention. Detailed Embodiments

[0074] In order to clearly illustrate the technical features of this solution, the present invention will be elaborated in detail below through specific embodiments and in conjunction with its drawings.

[0075] Embodiment 1

[0076] See Figures 1 to 3 , a 3D scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, includes the following steps:

[0077] Step 1: Collect the RGB-D video stream of the target scene. The construction of the sub-map starts from the first frame of the RGB-D video stream.

[0078] The collection of the RGB-D video stream depends on the RGB-D camera. During collection, the RGB-D camera needs to synchronously output the color map and the depth map to form the RGB-D video stream.

[0079] The RGB-D camera should be fixed to the robot body to ensure an unobstructed field of view. During the startup phase, the camera internal parameters need to be fixed. During the collection phase, the robot needs to move at a constant speed and avoid excessive rotation; the robot's movement speed cannot be too fast to prevent motion blur.

[0080] With the input of the video stream, a frame-to-model tracking strategy is implemented at the backend, and the sub-map is constructed frame by frame in real time. Using the sub-map to represent the scene helps to alleviate catastrophic forgetting and avoid the need to store all Gaussian models in the GPU memory. Each sub-map is constructed from a part of the key frames and contains a set of Gaussian point clouds.

[0081] Step 2: Based on the first frame of the RGB-D video stream, use the initialization strategy of the camera movement to create a sub-map, initially estimate and optimize the camera pose for the key frames of the created sub-map, and optimize the sub-map through loss calculation;

[0082] The specific method of the initialization strategy for camera movement is as follows: The submap starts from the first frame of the video stream and expands as new key frames are input. In terms of obtaining RGB-D video stream data, by sequentially reading the color image and the corresponding depth image data of each frame in chronological order, it provides the original data support for the construction of the submap. As the exploration area increases, new submaps are needed to cover previously unobserved areas, while avoiding storing all Gaussian models in the GPU memory. Different from the method of creating new submaps at fixed intervals, the present invention adopts an initialization strategy based on camera movement. Specifically, when the translation amount of the current frame relative to the first frame of the active submap exceeds a predefined threshold , or the Euler angle exceeds the threshold , a new submap will be created. Specifically: Given the external camera parameters of the current frame and the external camera parameters of the first frame of the current submap , the relative rotation and translation can be calculated by . When the norm of is greater than = 0.5m or the norm of the Euler angle of is greater than = 50°, a new submap will be created. Here, is the external camera parameter matrix, is the rotation matrix, is the translation vector, is the relative rotation matrix, is the relative translation vector. At any point in time, only the current active submap is processed. This method limits the computational cost and ensures that the optimization process remains efficient when exploring larger scenes.

[0083] After initially estimating and optimizing the camera pose of the key frames of the created submap using visual odometry, a dense point cloud is calculated using the RGB-D measurements of the key frames. At the beginning of each submap, points are uniformly sampled from the point cloud of the key frame, and regions with high color gradients are preferentially selected to add new Gaussian models; for subsequent key frames in the submap, points are uniformly sampled from regions with transparency values lower than a certain threshold. This can ensure the expansion of the map in sparse regions covered by Gaussian models. The new Gaussian models will be added based on the sampled points within the search radius of the current submap, provided that these points are within the There are no neighbors in the neighborhood; the key frames are selected using a fixed interval strategy, that is, a key frame is determined every 5 frames. The first frame is used as the initial key frame, and the frames every 5 frames thereafter are the new key frames. After initialization, the sub-map continuously expands with the input of new key frames, integrating the scene information carried by the new key frames into the existing sub-map to achieve continuous update and refinement of the explored area.

[0084] The method for optimizing the sub-map is as follows: Whenever a new Gaussian model is added to the active sub-map, all Gaussian models in the sub-map are jointly optimized for a fixed number of iterations; to accelerate the optimization, the RGB color is directly optimized, only focusing on the key frames in the active sub-map, and at least 40% of the iteration times are spent on the optimization of the new key frames;

[0085] The minimum loss function includes color loss and depth loss, and the expression is:

[0086] ,

[0087] where, is the color loss, is the real color image, is the predicted color image, reflects the structural similarity between the predicted color map and the real color map; is the depth loss, is the real depth image, is the predicted depth image, = 0.8, is used to control the weights of color difference and structural difference, represents the 1-norm, that is and .

[0088] Step 3: Perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment;

[0089] When performing frame-to-model tracking based on the constructed sub-map, assume that the current camera pose is the uniform motion assumption: , and predict the camera pose through the loss function. The loss function is:

[0090] ,

[0091] where, and are the predicted color map and depth map of the current frame respectively, while and are the real color map and depth map of the current frame respectively, It is a rendered transparency map used to calculate the mask of sparse regions , which reflects the differences in color and depth under a single view; to ensure the stability of tracking, a soft mask and an error mask are introduced to prevent pixels from unobserved or poorly reconstructed regions from affecting the tracking loss. The soft mask is directly from the active sub-map image, and the error boolean mask discards all pixels with depth errors exceeding the error threshold. Let \(w_c\) and \(w_d\) be the weights for controlling color differences and depth differences, and the loss function is obtained as:

[0092] ;

[0093] The camera pose is predicted by minimizing the color loss and depth loss of the image under relative camera pose transformation. However, this method only relies on single-view information, and when the camera observes sparse regions, it may cause artifacts in the color map and distortion in the depth map. These problems will invalidate the loss calculation of a large number of pixels, thereby reducing the tracking accuracy. To enhance the robustness of tracking, the present invention adopts a method similar to bundle adjustment (BA), introducing surface-to-surface projection error and Gaussian-to-Gaussian projection error. Due to the inherent difficulty in directly constraining the properties of Gaussians, the present invention adopts a method different from the traditional BA method. The traditional BA method jointly optimizes the camera pose and the point cloud, while the present invention chooses to only optimize the camera pose and keep the Gaussian parameters unchanged. The pose optimization loss function of the local bundle adjustment is:

[0094] ,

[0095] where \(L_t\) represents the tracking loss, and the surface-to-surface projection error \(d_{s2s}\) represents the distance between true surface observation points from two different views, where \(T_{i,j}\) represents the rigid body transformation from the \(i\)-th frame camera to the \(j\)-th frame camera, \(p_{s,j}\) represents the 3D coordinates of the object surface observation under the current relative transformation mapping, \(p_{s,i}\) represents the 3D coordinates of the object surface in the \(i\)-th frame camera coordinate system; the Gaussian-to-Gaussian projection error \(d_{g2g}\) represents the distance between Gaussian points from two different views, \(p_{g,j}\) represents the 3D coordinates of the Gaussian observation under the current relative transformation mapping, \(p_{g,i}\) represents the 3D coordinates of the Gaussian in the \(i\)-th Gaussian observation 3D markers under a frame camera

[0096] Step Four: Based on the optimized sub - map, use consistent 2D Gaussians for scene representation, calculate the intersections of the camera rays with the 2D Gaussians, and perform unbiased depth estimation; Figure 1 The 2D Gaussian is obtained by compressing a 3D Gaussian by removing one dimension of the scaling matrix. The 3D Gaussian represents the scene through a set of translucent Gaussian ellipsoids and renders the image using a specific Gaussian rasterizer. Each Gaussian has a mean, covariance matrix, color, and opacity. The 3D Gaussian can be expressed as:

[0097]

[0098] ,

[0099] To optimize the covariance matrix, the present invention decomposes the covariance matrix into a scaling matrix and a rotation matrix :

[0100] ,

[0101] where, is the covariance matrix of the 3D Gaussian, which is decomposed into the scaling matrix and the rotation matrix , is the transpose of the matrix.

[0102] To project the 3D Gaussian onto the rendering plane, the present invention uses the world - to - camera projection matrix to transform the coordinates of the Gaussian and uses the Jacobian matrix for local affine transformation approximation:

[0103] ,

[0104] is the Jacobian matrix, which prevents distortion of the Gaussian in non - linear projection, is the extrinsic camera parameters, which project the Gaussian from the world coordinate system to the pixel coordinate system, is the covariance matrix of the Gaussian in the world coordinate system, is the covariance matrix of the Gaussian in the pixel coordinate system.

[0105] After depth - sorting the projected Gaussians, the pixel color is calculated by blending along the camera ray direction:

[0106] ,

[0107] where, Represents the color of the current pixel, Represents the Gaussian set intersecting with the ray, Is the color of the th Gaussian, Is the opacity of the th Gaussian.

[0108] Due to the lack of Figure 1 consistency, the surface reconstruction effect of 3D Gaussians is poor. To solve this problem, the present invention adopts a Figure 1 consistent 2D Gaussian representation. By removing one dimension of the scaling matrix, the 3D Gaussian is compressed into a 2D disk. The 2D Gaussian can be expressed as:

[0109] ,

[0110] where, and are coordinates in the local coordinate system; The parameters of the 2D Gaussian include: is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, serves as the scaling matrix, serves as the rotation matrix, where represents the normal vector of the 2D Gaussian. The spatial transformation matrix of the 2D Gaussian has the following expression:

[0111] ,

[0112] where, is the projection matrix of the local coordinate system of the 2D Gaussian, is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, and are the scaling coefficients of the two principal tangent vectors of the 2D Gaussian;

[0113] Assume is the transformation matrix from world coordinates to pixel coordinates. Then the 3D space point is calculated by the following formula:

[0114] ;

[0115] where, represents a uniform light ray emitted from the camera, passing through the pixel and at a depth intersects with the 2D Gaussian at a point, is the homogeneous coordinate of a 3D space point, is the transformation matrix from the world coordinate system to the pixel coordinate system, is the local projection matrix of the 2D Gaussian, is the homogeneous coordinate of a point in the local coordinate system of the 2D Gaussian.

[0116] Calculating the intersection point of the camera ray and the 2D Gaussian specifically involves obtaining the intersection point through two planes. Given a pixel point , define the intersection point of this point in the plane and the plane: The normal vector of the plane is (-1, 0, 0), and the offset is The definition of the plane is as follows: The normal vector of the plane is (0, -1, 0), and the offset is The definition of the plane is as follows:

[0117] ,

[0118] The local coordinate point of the 2D Gaussian is The intersection point of the two planes is obtained using the following equation:

[0119] ,

[0120] Thus, the solution is:

[0121] ,

[0122] where: and are the i-th parameters of the plane vector.

[0123] Previous methods used the component of the Gaussian mean as the depth. However, when the intersection point of the ray and the Gaussian distribution is far from the mean, significant errors will be introduced in depth estimation. The present invention calculates an unbiased depth by solving the intersection point of the ray and the Gaussian surface. This depth can be obtained through the following equation to calculate the unbiased depth:

[0124] ,

[0125] where, is the opacity of the j-th Gaussian, is the opacity of the j-th pixel, is the intersection point of the ray and the Gaussian, and the vector It is calculated from the extrinsic parameters of the camera, the scaling matrix and the rotation matrix of Gauss. By converting homogeneous coordinates to inhomogeneous coordinates, we can obtain , where is the extrinsic parameters of the camera, is the spatial transformation matrix of the 2D Gauss. is the Gauss mean. Since all rays come from the same viewpoint and are coplanar with the given Gauss distribution, the intersection equation of each Gauss distribution only needs to be solved once.

[0126] Step Five: Merge all optimized submaps into a global map, and optimize the color of the global map and the geometric properties of Gauss by minimizing the loss function;

[0127] To obtain an accurate global map for improving the accuracy of unbiased depth, Gauss optimization must be performed. The Gauss optimization includes the calculation of regularization loss and normal loss. The regularization loss includes scale regularization and scale ratio regularization. The regularization loss function is: :

[0128] Scale regularization: An overly large Gauss will cause artifacts and reduce the accuracy of depth estimation. 3DGS prevents this by splitting an overly large Gauss into multiple smaller Gausses. However, this will increase the memory overhead and training time, making it inapplicable in SLAM and causing the Gauss scale to grow uncontrollably. To solve this problem, the present invention proposes a scale regularization method to limit the scale of Gauss. The scale regularization loss is as follows:

[0129] ,

[0130] where: controls the regularization weight, =0.05, is the scaling matrix of the i-th 2D Gauss, is the maximum scale threshold, is the number of Gausses exceeding the threshold; Scale regularization dynamically calculates the scale of each Gauss during training, compares it with the threshold, and imposes a penalty on the Gausses exceeding the threshold;

[0131] Scale Ratio Regularization: To prevent the Gaussian from being overly stretched, previous methods introduced isotropic regularization to ensure that the Gaussian expands uniformly in all directions and suppress anisotropy. However, this method affects the rendering quality. Isotropic regularization limits the diversity of the Gaussian distribution and reduces the robustness of the model. During the rendering process, appropriate anisotropy helps the Gaussian better adapt to the geometric structure, thereby improving the rendering effect and making the image more physically realistic. To prevent the Gaussian from being overly stretched while maintaining high rendering quality, the present invention proposes a scale ratio regularization method, and its loss function is as follows:

[0132] ,

[0133] where, controls the regularization weight, = 0.05, and are the major axis vector and minor axis vector of the i-th 2D Gaussian respectively, is the maximum scale ratio threshold, is the number of Gaussians that exceed the threshold. The scale ratio regularization only restricts the extreme cases that exceed the threshold to prevent the excessive stretching of the Gaussian shape. For Gaussians with a scale ratio within a reasonable range, this regularization does not intervene, thus maintaining moderate anisotropy;

[0134] When optimizing 2D Gaussians using the depth loss, some 2D Gaussians may have unreasonable directions, such as being embedded inside the object surface. This increases the depth prediction error and may lead to the degradation of a large number of 2D Gaussians. To solve this problem and ensure that the normal vector of the 2D Gaussian is aligned with the surface normal of the object, thereby improving the accuracy of depth prediction and enhancing the stability of the Gaussian distribution. The present invention regards the real depth map as the object surface and defines a normal loss to constrain the normal vector of the 2D Gaussian to be aligned with the surface normal of the object. The normal loss function is:

[0135] ,

[0136] where, is the mixing weight of the current intersection point, where , is the opacity of the Gaussian at the j-th pixel, is the opacity at the j-th pixel. is the normal of the current 2D Gaussian in the camera direction, is the surface normal of the object estimated from the real depth map, and are the gradients of the depth map in the and directions respectively.

[0137] After calculating the regularization loss and the normal loss, the color of the global map and the geometric properties of the Gaussian are optimized by minimizing the loss function, which includes color loss, depth loss, normal loss, and regularization loss;

[0138] The loss function to be minimized is: , where is the color loss, is the depth loss, is the normal loss, is the regularization loss.

[0139] Step 6: Perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the reconstruction of the 3D scene based on the color map, the depth map, and the normal map.

[0140] Perform Gaussian pruning and exposure compensation on the optimized global map. The most critical factor in Gaussian pruning is how to correctly evaluate the contribution of each Gaussian. 3DGS adopts a pruning method based on opacity, where the contribution of a Gaussian is determined by its opacity parameter, and Gaussians with lower opacity are pruned. However, due to the instability of the opacity parameter, some Gaussians that are crucial for the rendering quality may be wrongly assigned a lower opacity and thus be pruned by mistake. In contrast, the pruning method based on color contribution judges the importance of each Gaussian by evaluating its color contribution and only prunes Gaussians that have almost no color contribution to the final rendering. This pruning strategy is more stable and can effectively avoid having a negative impact on the rendering quality. The Gaussian pruning calculation formula is:

[0141] ,

[0142] where is the color contribution of the Gaussian at the j-th view, is the 2D projection area of the Gaussian at the j-th view, is the pixel value at the j-th view, is the opacity of the current 2D Gaussian, is the color value of the current 2D Gaussian, is the accumulated opacity of the current pixel, is a factor used to control the balance between transparency and color contribution. When , the color-based pruning degrades to transparency-based pruning;

[0143] Since the contribution of a Gaussian is high in a few views, its overall contribution is determined by the average contribution of these high-contribution views. The multi-view contribution calculation formula of the Gaussian is as follows:

[0144] ,

[0145] Among them, is the multi-view set with the highest contribution. Here, the 5 single views with the highest contribution are selected for calculation;

[0146] During the image acquisition process in real scenarios, light changes and perspective-related reflection phenomena are very common. These factors often lead to different exposure effects in images from different perspectives. This exposure difference will introduce serious color inconsistencies, thus greatly affecting the quality of scene reconstruction. To solve this problem, the present invention uses an exposure compensation algorithm to optimize the image color and ensure the consistency and accuracy of image information during the reconstruction process. The present invention applies an affine transformation matrix to each key-frame image for correction, and the exposure compensation formula is:

[0147] ,

[0148] Among them, is the color map predicted for the current frame, is a transformation matrix, which is used to adjust the intensity and proportional relationship of the image color channels. By independently adjusting different color channels, it can effectively correct the color deviation caused by light changes, is a bias vector, which is used to adjust the overall brightness of the image to compensate for the light intensity difference between different perspectives. The matrix and the vector are obtained through training, so as to achieve adaptive compensation for exposure differences.

[0149] To comprehensively evaluate the performance, we adopted a variety of evaluation metrics. The tracking accuracy is measured by the root mean square of the absolute trajectory error (ATE RMSE). A lower ATE RMSE value indicates higher accuracy. The rendering quality is evaluated using the peak signal-to-noise ratio (PSNR), the structural similarity index (SSIM), and the perceptual image patch similarity (LPIPS). These metrics are calculated after full-resolution rendering along the estimated trajectory, and the mapping interval is set to 5 frames. The reconstruction performance is evaluated on the mesh generated by the Truncated Signed Distance Function algorithm and quantified by the F1 score (the harmonic mean of precision and recall). In addition, the depth L1 error of the unobserved perspective is also reported as a supplementary evaluation metric.

[0150] Multiple representative datasets were selected to evaluate the performance of our method. For the synthetic dataset, the Replica dataset was used, which contains high-quality 3D reconstructions of indoor scenes acquired through RGB-D sensor trajectories. The real-world datasets include TUM-RGBD and ScanNet, where the poses of TUM-RGBD were obtained by an external motion capture system, while the poses of ScanNet were estimated by BundleFusion. Such a dataset selection enables us to comprehensively and objectively evaluate the effectiveness and robustness of the algorithm under different scenarios.

[0151] Example 2

[0152] A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering in Example 1, the system includes:

[0153] A data acquisition module, configured to acquire an RGB-D video stream of the target scene;

[0154] A map construction module, configured to create a sub-map using the initialization strategy of camera motion based on the first frame of the RGB-D video stream, initially estimate the camera pose for the key frames of the created sub-map using visual odometry, and optimize the sub-map through loss calculation;

[0155] A tracking module, configured to perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment;

[0156] A scene representation module, configured to use consistent 2D Gaussians for scene representation based on the optimized sub-map, calculate the intersections of camera rays with the 2D Gaussians and perform unbiased depth estimation; Figure 1 Calculate the intersections of camera rays with the 2D Gaussians and perform unbiased depth estimation;

[0157] A data optimization module, configured to merge all optimized sub-maps into a global map and optimize the color and geometric properties of the Gaussians by minimizing the loss function;

[0158] A data output module, configured to perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the three-dimensional scene reconstruction based on the color map, the depth map, and the normal map.

[0159] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.

Claims

1. A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, characterized in that: It includes the following steps: Collect the RGB-D video stream of the target scene; Based on the first frame of the RGB-D video stream, use the initialization strategy of camera motion to create a sub-map, initially estimate and optimize the camera pose for the key frames of the created sub-map, and optimize the sub-map through loss calculation; Perform frame-to-model tracking based on the constructed submaps, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; the accurate estimation of the camera pose by calculating the tracking loss and the optimization of the camera pose through local bundle adjustment include introducing soft mask and error mask , soft mask directly from the transparency map of the active submap, the error boolean mask rejects all pixel points with color or depth errors exceeding the threshold to obtain the tracking loss function; optimize the camera pose using a local bundle adjustment loss function that keeps the Gaussian parameters unchanged; The tracking loss function is: , Wherein: is the weight for controlling color difference and depth difference; The local bundle adjustment loss function is: , Among them, represents the tracking loss, the surface-to-surface projection error represents the distance between the true surface observation points from two different perspectives, where represents from the th frame camera to the th frame camera's rigid body transformation, represents the 3D coordinates of the object surface observation under the current relative transformation mapping, represents the th frame camera coordinate system's object surface 3D coordinates; Gaussian-to-Gaussian projection error represents the distance between Gaussian points from two different perspectives, represents the 3D coordinates of the Gaussian observation under the current relative transformation mapping, represents the th frame camera's Gaussian observation 3D coordinates; Based on the optimized sub-map, use view-consistent 2D Gaussian for scene representation, calculate the intersection of the camera ray and the 2D Gaussian and perform unbiased depth estimation; using view-consistent 2D Gaussian for scene representation, calculating the intersection of the camera ray and the 2D Gaussian and performing unbiased depth estimation includes compressing the 3D Gaussian into a 2D Gaussian by removing one dimension of the scaling matrix, using the 2D Gaussian for scene representation, solving the intersection of the camera ray and the 2D Gaussian and calculating the unbiased depth to determine the Gaussian depth; The expression of the 2D Gaussian is: , wherein, and are coordinates in a local coordinate system. The parameters of the 2D Gaussian include: is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, serves as a scaling matrix, serves as a rotation matrix, where represents the normal vector of the 2D Gaussian. The spatial transformation matrix of the 2D Gaussian has the following expression: , Wherein: is the projection matrix of the 2D Gaussian in the local coordinate system, is the center position of the 2D Gaussian, are the two principal tangent vectors of the 2D Gaussian, are the scaling factors of the two principal tangent vectors of the 2D Gaussian; is the transformation matrix from world coordinates to pixel coordinates, and the 3D spatial point has the following formula: , Wherein: represents uniform light emitted from the camera, passing through the pixel and intersecting with the 2D Gaussian at depth wherein, is the homogeneous coordinate of the 3D spatial point, is the transformation matrix from the world coordinate system to the pixel coordinate system, is the local projection matrix of the 2D Gaussian, is the homogeneous coordinate of a point within the local coordinate system of the 2D Gaussian, is the transpose of the matrix; The intersection of the camera ray and the 2D Gaussian is: , Wherein: and are the i-th parameters of the planar vector; The unbiased depth estimation function is: , Among them, is the opacity of the j-th Gaussian, is the opacity of the j-th pixel, is the intersection point of the ray and the Gaussian, and the vector is calculated from the extrinsic parameters of the camera, the scaling matrix, and the rotation matrix of the Gaussian, is the Gaussian mean; Merge all optimized sub-maps into a global map, and optimize the color of the global map and the geometric properties of the Gaussian by minimizing the loss function; Perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map and a normal map, and complete the reconstruction of the three-dimensional scene based on the color map, the depth map and the normal map.

2. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 1, wherein, The initialization strategy for the camera movement includes that when the translation amount of the current frame relative to the first frame of the active sub-map exceeds a predefined threshold , or when the estimated Euler angles exceed a predefined threshold , a new sub-map will be created, and only the current active sub-map is processed at any point in time.

3. A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 2, characterized in that, After the camera pose is initially estimated for the key frames of the created submaps using visual odometry, a dense point cloud is calculated using the RGB-D measurements of the key frames. At the beginning of each submap, points are uniformly sampled from the point cloud of the key frame, and regions with high color gradients are preferentially selected to add new Gaussian models. For subsequent key frames in the submap, points are uniformly sampled from regions with transparency values lower than a certain threshold. A new Gaussian model is added based on the sampled points within the neighborhood of the current submap when there are no neighbors within the search radius of the current submap .

4. A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 3, characterized in that, Optimizing the sub-map through loss calculation includes calculating the minimum loss to optimize the sub-map to render the color and depth of the key frames; The minimum loss includes color loss and depth loss, and the minimum loss function is: , Among them, is the color loss, is the real color image, is the predicted color image, reflecting the structural similarity between the predicted color map and the real color map; is the depth loss, is the real depth image, is the predicted depth image, used to control the weights of the color difference and the structural difference, represents the 1-norm, that is and .

5. A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 4, characterized in that To improve the accuracy of unbiased depth calculation, Gaussian optimization must be performed. The Gaussian optimization includes calculating regularization loss and normal loss. The regularization loss includes scale regularization that restricts the Gaussian scale and scale ratio regularization that prevents the Gaussian from being over-stretched. The normal loss enables the 2D Gaussian to be tangent to the actual surface; The regularization loss function is: , Where: scale regularization China Control the regularization weight, is the scaling matrix of the i-th 2D Gaussian, is the maximum scale threshold, is the number of Gaussians exceeding the threshold; Scale ratio regularization medium and are the major axis vector and minor axis vector of the i-th 2D Gaussian, respectively, controls the weight of regularization, is the maximum scale ratio threshold, is the number of Gaussians that exceed the threshold; The normal loss function is: , Among them, is the mixed weight of the current intersection point, where , is the Gaussian opacity at the j-th pixel, is the opacity of the j-th pixel, is the normal of the current 2D Gaussian in the camera direction, is the object surface normal estimated from the true depth map, and are the gradients of the depth map in the and directions respectively.

6. A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 5, characterized in that, Optimizing the color of the global map and the geometric properties of the Gaussian by minimizing the loss function includes calculating color loss, depth loss, normal loss and regularization loss for the global map to optimize the global map. The minimum loss function is: , Among them, is the color loss, is the depth loss, is the normal loss, is the regularization loss.

7. A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 6, characterized in that Performing Gaussian pruning and exposure compensation on the optimized global map includes pruning the Gaussian based on the color contribution method by evaluating the color contribution of each Gaussian to determine its importance, pruning the Gaussian that has no color contribution to the final rendering, and calculating the exposure compensation by applying a transformation matrix to the key frame image for correction. The multi-view contribution calculation formula for Gaussian pruning is: , Among them, is the multi-view set with the highest contribution, is the calculation formula for the Gaussian pruning single-view contribution, where, is the color contribution of the Gaussian at the j-th view, is the 2D projection area of the Gaussian at the j-th view, is the pixel value at the j-th view, is the opacity of the current 2D Gaussian, is the color value of the current 2D Gaussian, is the accumulated opacity for the current pixel, is a factor used to control the balance between opacity and color contribution. When the color-based pruning degrades to transparency-based pruning; The exposure compensation formula is: , Among them, is the color map after exposure compensation, is the color map predicted by the current frame, is a transformation matrix for adjusting the intensity and proportional relationship of the image color channels. By independently adjusting different color channels, the color deviation caused by light changes can be effectively corrected, is a bias vector.

8. A three-dimensional scene reconstruction and real-time SLAM system based on 2D Gaussian sputtering, which executes the method for three-dimensional scene reconstruction and real-time SLAM based on 2D Gaussian sputtering described in claim 1, characterized in that, It includes: A data acquisition module, configured to collect the RGB-D video stream of the target scene; A map construction module, configured to create a sub-map based on the first frame of the RGB-D video stream using the initialization strategy of camera motion, initially estimate the camera pose for the key frames of the created sub-map, and optimize the sub-map through loss calculation; The tracking module is configured to perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; The scene representation module is configured to perform scene representation using view-consistent 2D Gaussians based on the optimized sub-map, calculate the intersections of camera rays with the 2D Gaussians, and perform unbiased depth estimation; The data optimization module is configured to merge all the optimized sub-maps into a global map and optimize the color of the global map and the geometric properties of the Gaussians by minimizing the loss function; The data output module is configured to perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the reconstruction of the three-dimensional scene based on the color map, the depth map, and the normal map.

Citation Information

Patent Citations

  • Dense RGB-D SLAM method based on multistage 3D Gaussian

    CN118781189A

  • Method for 3D scene dense reconstruction based on monocular visual slam

    US20200273190A1