3D scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering

By combining monocular geometric priors and learning-based intensive SLAM, using 2D Gaussian sputtering to represent scenes, the problem of poor performance of existing SLAM methods in challenging environments is solved, achieving high-quality and efficient three-dimensional scene reconstruction.

CN120070794AActive Publication Date: 2025-05-30YANTAI UNIV

Patent Information

Application Number
CN202510541203.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Existing SLAM methods perform poorly in strong light, radiation variation, and dynamic or low texture environments, and sparse 3D modeling based on discrete surface representations leads to limited spatial resolution and reconstruction distortion, accurately estimating unobserved area geometry remains a challenge.

Method used

Using a three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, combining monocular geometric priors and learning-based intensive SLAM, 2DGS is used as a compact map representation, and the color and Gaussian geometric properties of the global map are optimized through unbiased depth estimation of the intersection of camera rays and 2D Gaussians.

Benefits of technology

High-quality indoor scene reconstruction is achieved, the accuracy of geometric estimation is improved, the trade-off between speed and quality is broken, and the ability to quickly and accurately surface reconstruction is overcome, overcoming the limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070794A_ABST
    Figure CN120070794A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of three-dimensional reconstruction, and particularly relates to a three-dimensional scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering, and the method comprises the steps: carrying out the construction and optimization of a sub-map after the initialization of the sub-map based on a collected video stream; frame-to-model tracking is executed based on the constructed map, and joint optimization of key frame poses and sub-maps is carried out through local beam adjustment; the 2D Gaussian scene with consistent views is represented, an intersection point of a camera ray and 2D Gaussian is calculated, and depth estimation is carried out; combining all the sub-maps into a global map, optimizing the global map, performing Gaussian pruning and exposure compensation to obtain a color map, a depth map and a normal map, and completing three-dimensional scene reconstruction according to the color map, the depth map and the normal map; the system comprises a data acquisition module, a map construction module, a tracking module, a scene representation module, a data optimization module and a data output module. Through the technical scheme of the invention, high-quality indoor scene reconstruction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a three-dimensional scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering. Background Art

[0002] SLAM was originally designed for robots and automation systems. Today, its application requirements have expanded to multiple fields, including augmented reality (AR), visual surveillance, medical applications, etc. To meet these requirements, researchers are developing methods that enable machines to autonomously build increasingly accurate scene representations, leveraging advancements in robotics, computer vision, sensor technology, and artificial intelligence (AI).

[0003] Over the years, SLAM methods have evolved significantly to adapt to these specific requirements. Initially, hand-designed algorithms demonstrated excellent real-time performance and scalability. However, they had problems in strong lighting, radiation changes, and dynamic or low-texture environments, resulting in less-than-satisfactory performance. Advanced techniques that incorporate deep learning methods are crucial for improving the accuracy and reliability of localization and mapping. By leveraging the powerful feature extraction capabilities of deep neural networks, SLAM is particularly effective in challenging conditions. However, deep learning methods rely on large amounts of training data sets and precise ground truth annotations, which limit their generalization ability in unseen scenes. In addition, both hand-designed methods and deep learning-based methods are limited by the use of discrete surface representations, such as point / surface clouds, voxel hashing, voxel grids, and octrees. These representations lead to sparse 3D modeling, limited spatial resolution, and distortion during the reconstruction process. Accurately estimating the geometry of unobserved regions remains an unsolved challenge. Recent research has focused on improving map quality and density. With the emergence of neural scene representations (such as Neural Radiance Fields, NeRF), which enable high-fidelity view synthesis, dense neural SLAM methods have rapidly developed. Although these methods have made impressive progress in scene representation quality, they are still limited to small-scale synthetic scenes, and the re-rendered results are far from realistic. Recently, 3DGS has been shown to be comparable to NeRF in rendering quality, while achieving an order-of-magnitude improvement in rendering and training speed. This representation is explicit and manipulable, making it a strong candidate for downstream tasks. However, due to the lack of geometric priors and Figure 1 consistency, their geometric reconstruction performance has not yet surpassed that of dense neural SLAM methods. The present invention aims to solve this problem.

[0004] This paper presents a Gaussian-based SLAM system that uses RGB-D input to achieve accurate and fast monocular scene reconstruction. The method of the present invention models the scene in an efficient and accurate manner by combining monocular geometric priors and learning-based dense SLAM, while using 2DGS as a compact map representation, improving the accuracy of geometric estimation. The present invention uses 2DGS to construct a map and uses the existing map to guide camera pose optimization. This design ensures efficiency and accuracy, breaks the trade-off between speed and quality, thus achieving fast and accurate surface reconstruction while maintaining high-quality results, and successfully overcomes the limitations of previous methods. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention develops a three-dimensional scene reconstruction and real-time SLAM system and method based on 2D Gaussian sputtering, which can achieve high-quality indoor scene reconstruction, and provides an enhanced depth estimation method and a camera pose estimation method.

[0006] The technical solution for the present invention to solve the technical problem is: A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, comprising the following steps: Collect the RGB-D video stream of the target scene; Based on the first frame of the RGB-D video stream, use the initialization strategy of camera motion to create a sub-map, initially estimate and optimize the camera pose for the key frames of the created sub-map, and optimize the sub-map through loss calculation; Perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; Based on the optimized sub-map, use the Figure 1 Consistent 2D Gaussian for scene representation, calculate the intersection of the camera ray and the 2D Gaussian and perform unbiased depth estimation; Merge all optimized sub-maps into a global map, and optimize the color and geometric properties of the Gaussian of the global map by minimizing the loss function; Perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the three-dimensional scene reconstruction based on the color map, the depth map, and the normal map.

[0007] The initialization strategy of the camera motion includes that when the translation amount of the current frame relative to the first frame of the active sub-map exceeds a predefined threshold , or when the estimated Euler angle exceeds a predetermined threshold , a new sub-map will be created, and only the current active sub-map is processed at any time point.

[0008] After the camera pose is initially estimated for the key frames of the created sub-map using visual odometry, a dense point cloud is calculated using the RGB-D measurements of the key frames. At the beginning of each sub-map, points are uniformly sampled from the point cloud of the key frame, and regions with high color gradients are preferentially selected to add new Gaussian models. For subsequent key frames in the sub-map, points are uniformly sampled from regions where the transparency value is lower than a certain threshold. New Gaussian models are added based on the sampled points within the neighborhood of the current sub-map when there are no neighbors within the search radius of the current sub-map.

[0009] The optimization of the sub-map through loss calculation includes calculating the minimum loss to optimize the sub-map for rendering the color and depth of the key frames; The minimum loss includes color loss and depth loss, and the minimum loss function is: , where is the color loss, is the real color image, is the predicted color image, reflects the structural similarity between the predicted color map and the real color map; is the depth loss, is the real depth image, is the predicted depth image, = 0.8, is used to control the weights of color differences and structural differences, represents the 1-norm, that is and .

[0010] The accurate estimation of the camera pose is achieved by calculating the tracking loss, and the camera pose is optimized through local bundle adjustment, including introducing a soft mask and an error mask . The soft mask is directly from the transparency map of the active sub-map, and the error boolean mask removes all pixels with color or depth errors exceeding the threshold to obtain the tracking loss function; the local bundle adjustment loss function that keeps the Gaussian parameters unchanged is used to optimize the camera pose; The tracking loss function is: , where: is used to control the weights of color differences and depth differences; The local bundle adjustment loss function is: , wherein, represents the tracking loss, the surface-to-surface projection error represents the distance between the true surface observation points from two different perspectives, wherein, represents the th frame camera to the th frame camera rigid body transformation, represents the 3D coordinates of the object surface observation under the current relative transformation mapping, represents the distance between the Gaussian points from two different perspectives, represents the 3D coordinates of the Gaussian observation under the current relative transformation mapping, represents the th

[0011] The method of using Figure 1 consistent 2D Gaussians for scene representation, calculating the intersection of the camera ray and the 2D Gaussian and performing unbiased depth estimation, including compressing the 3D Gaussian into a 2D Gaussian by removing one dimension of the scaling matrix, using the 2D Gaussian for scene representation, solving the intersection of the camera ray and the 2D Gaussian and calculating the unbiased depth to determine the Gaussian depth; The expression of the 2D Gaussian is: , wherein, and are the coordinates in the local coordinate system. The parameters of the 2D Gaussian include: is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, serves as the scaling matrix, serves as the rotation matrix, wherein represents the normal vector of the 2D Gaussian. The spatial transformation matrix of the 2D Gaussian is expressed as: , wherein, is the projection matrix of the local coordinate system of the 2D Gaussian, is the center position of the 2D Gaussian, are the two principal tangent vectors of the 2D Gaussian, are the scaling coefficients of the two principal tangent vectors of the 2D Gaussian; is the transformation matrix from world coordinates to pixel coordinates, and the 3D spatial point has the formula: , where represents the uniform light emitted from the camera, passing through the pixel and intersecting with the 2D Gaussian at the depth ; is the homogeneous coordinate of the 3D spatial point, is the transformation matrix from the world coordinate system to the pixel coordinate system, is the local projection matrix of the 2D Gaussian, is the homogeneous coordinate of a point in the local coordinate system of the 2D Gaussian, is the transpose of the matrix; The intersection point of the camera ray and the 2D Gaussian is: , where: and are the i-th parameters of the plane vector; The unbiased depth estimation function is: , where is the opacity of the j-th Gaussian, is the opacity of the j-th pixel, is the intersection point of the ray and the Gaussian, and the vector is calculated from the extrinsic parameters of the camera, the scaling matrix and the rotation matrix of the Gaussian, is the Gaussian mean.

[0012] To improve the accuracy of unbiased depth calculation, Gaussian optimization is required. The Gaussian optimization includes the calculation of regularization loss and normal loss. The regularization loss includes scale regularization that restricts the Gaussian scale and scale ratio regularization that prevents the Gaussian from being over-stretched. The normal loss enables the 2D Gaussian to be tangent to the actual surface; The regularization loss function is: , where: Scale regularization in which controls the regularization weight, = 0.05, is the scaling matrix of the i-th 2D Gaussian, is the maximum scale threshold, is the number of Gaussians exceeding the threshold; Scale ratio regularization in which and are the major axis vector and minor axis vector of the i-th 2D Gaussian respectively, Control the weight of regularization, = 0.05, is the maximum scale ratio threshold, is the number of Gaussians that exceed the threshold; The normal loss function is: , where, is the mixing weight of the current intersection point, where , is the Gaussian opacity at the j-th pixel, is the opacity of the j-th pixel. is the normal of the current 2D Gaussian in the camera direction, is the object surface normal estimated from the true depth map, and are respectively the gradients of the depth map in and directions.

[0013] The color and geometric properties of the global map are optimized by minimizing the loss function, including calculating color loss, depth loss, normal loss, and regularization loss for the global map to optimize the global map. The minimized loss function is: , where, is the color loss, is the depth loss, is the normal loss, is the regularization loss.

[0014] Gaussian pruning and exposure compensation are performed on the optimized global map, including a pruning method based on color contribution to judge the importance of each Gaussian by evaluating its color contribution, pruning Gaussians that have no color contribution to the final rendering, and calculating exposure compensation using a method of applying a transformation matrix to the key frame image for correction. The multi-view contribution calculation formula for Gaussian pruning is: , where, is the multi-view set with the highest contribution, is the single-view contribution calculation formula for Gaussian pruning, where, is the color contribution of the Gaussian at the j-th view, is the 2D projection area of the Gaussian at the j-th view, is the pixel value at the j-th view, is the opacity of the current 2D Gaussian, is the color value of the current 2D Gaussian, is the opacity accumulated for the current pixel, is a factor used to control the balance between transparency and color contribution. When , color-based pruning degrades to transparency-based pruning; The exposure compensation formula is: , where, is the color map after exposure compensation, is the color map predicted for the current frame, is a transformation matrix, which is used to adjust the intensity and proportional relationship of the image color channels. By independently adjusting different color channels, it can effectively correct the color deviation caused by illumination changes, is a bias vector.

[0015] A three-dimensional scene reconstruction and real-time SLAM system based on 2D Gaussian sputtering, comprising: A data acquisition module, configured to acquire an RGB-D video stream of a target scene; A map construction module, configured to, based on the first frame of the RGB-D video stream, create a sub-map using the initialization strategy of camera movement, initially estimate the camera pose for the key frames of the created sub-map using visual odometry, and optimize the sub-map through loss calculation; A tracking module, configured to perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; A scene representation module, configured to, based on the optimized sub-map, use visually Figure 1 consistent 2D Gaussians for scene representation, calculate the intersection points of the camera rays and the 2D Gaussians, and perform unbiased depth estimation; A data optimization module, configured to merge all the optimized sub-maps into a global map, and optimize the color and geometric properties of the Gaussians in the global map by minimizing the loss function; A data output module, configured to perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the reconstruction of the three-dimensional scene based on the color map, the depth map, and the normal map.

[0016] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: 1. By combining monocular geometric priors and visually Figure 1 consistent scene representations, high-quality indoor scene reconstruction is achieved; 2. Through the enhanced depth estimation method, the depth estimation deviation is corrected by the camera ray and Gaussian intersection method, achieving more accurate depth estimation; 3. To improve the accuracy of pose estimation, the present invention introduces an additional Bundle Adjustment (BA) optimization strategy during the frame-to-model tracking process. This optimization is achieved by minimizing the surface-to-surface and Gaussian-to-Gaussian projection errors, thereby obtaining a more accurate camera pose estimation; 4. The system supports an offline processing method for exposure compensation and Gaussian pruning, improving the rendering quality while reducing the storage overhead; 5. The system can render high-fidelity RGB images, accurate depth maps, and normal maps in real time, and still maintain stability and noise resistance in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the general flow block diagram of a three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to Embodiment 1 of the present invention; Figure 2 It is the rendering effect diagram of Embodiment 1 of the present invention; Figure 3 It is the comparison diagram of the reconstruction result and the real scene of Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to clearly illustrate the technical features of the present solution, the present invention will be elaborated in detail below through specific embodiments and in combination with its accompanying drawings.

[0019] Embodiment 1 Refer to Figures 1 to 3 , a three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, comprising the following steps: Step 1: Collect the RGB-D video stream of the target scene, and the construction of the sub-map starts from the first frame of the RGB-D video stream.

[0020] The collection of the RGB-D video stream depends on the RGB-D camera. During collection, the RGB-D camera needs to synchronously output the color map and the depth map to form the RGB-D video stream.

[0021] The RGB-D camera should be fixed to the robot body to ensure unobstructed vision. During the startup phase, the camera internal parameters need to be fixed. During the collection phase, the robot needs to move at a constant speed and avoid excessive rotation; the movement speed of the robot cannot be too fast to prevent motion blur.

[0022] With the input of the video stream, a frame-to-model tracking strategy is implemented at the backend, and a sub-map is constructed frame by frame in real time. Using the sub-map to represent the scene helps to alleviate catastrophic forgetting and avoid the need to store all Gaussian models in the GPU memory. Each sub-map is constructed from a part of the key frames and contains a set of Gaussian point clouds.

[0023] Step 2: Based on the first frame of the RGB-D video stream, use the initialization strategy of camera motion to create a sub-map. For the key frames of the created sub-map, use visual odometry to initially estimate and optimize the camera pose, and optimize the sub-map through loss calculation. The specific method of the initialization strategy of camera motion is as follows: The sub-map starts from the first frame of the video stream and expands with the input of new key frames. In terms of obtaining RGB-D video stream data, by reading the color image and the corresponding depth image data of each frame in sequence according to the time series, it provides the original data support for the construction of the sub-map. As the exploration area increases, new sub-maps are needed to cover the previously unobserved areas, while avoiding storing all Gaussian models in the GPU memory. Different from the method of creating new sub-maps at fixed intervals, the present invention adopts an initialization strategy based on camera motion. Specifically, when the translation amount of the current frame relative to the first frame of the active sub-map exceeds a predefined threshold , or the Euler angle exceeds the threshold , a new sub-map will be created. Specifically: Given the external camera parameters of the current frame and the external camera parameters of the first frame of the current sub-map , the relative rotation and translation can be calculated by . When the norm of is greater than = 0.5m or the norm of the Euler angle of is greater than = 50°, a new sub-map will be created. Here, is the external camera parameter matrix, is the rotation matrix, is the translation vector, is the relative rotation matrix, is the relative translation vector. At any point in time, only the current active sub-map is processed. This method limits the computational cost and ensures that the optimization process remains efficient when exploring a larger scene.

[0024] After initially estimating and optimizing the camera pose of the key frames of the created sub-map using visual odometry, a dense point cloud is calculated using the RGB-D measurements of the key frames. At the beginning of each sub-map, uniformly sample from the point cloud of the key frames Points are selected, and regions with high color gradients are preferentially chosen to add new Gaussian models; for subsequent key frames in the sub-map, points are uniformly sampled from regions with transparency values lower than a certain threshold, which can ensure the expansion of the map in sparse regions covered by Gaussian models. The new Gaussian models will be added based on the sampled points within the search radius of the current sub-map, provided that these points have no neighbors in the neighborhood of the current sub-map; Key frames are selected using a fixed interval strategy, that is, a key frame is determined every 5 frames. The first frame is used as the initial key frame, and the frames every 5 subsequent frames are the new key frames. After initialization, the sub-map continuously expands with the input of new key frames, integrating the scene information carried by the new key frames into the existing sub-map to achieve continuous update and refinement of the explored area.

[0025] The method for optimizing the sub-map is as follows: Whenever a new Gaussian model is added to the active sub-map, all Gaussian models in the sub-map are jointly optimized for a fixed number of iterations; To accelerate the optimization, the RGB color is directly optimized, only focusing on the key frames in the active sub-map, and at least 40% of the iteration times are spent on the optimization of new key frames; The minimum loss function includes color loss and depth loss, and the expression is: , where is the color loss, is the real color image, is the predicted color image, reflects the structural similarity between the predicted color map and the real color map; is the depth loss, is the real depth image, is the predicted depth image, = 0.8, is used to control the weights of color differences and structural differences, represents the 1-norm, that is and .

[0026] Step 3: Perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; When performing frame-to-model tracking based on the constructed sub-map, assume that the current camera pose is the uniform motion assumption: , and predict the camera pose through the loss function. The loss function is: , where ​​and are the color map and depth map for current frame prediction respectively, while and are the true color map and depth map of the current frame respectively, is the rendered transparency map, which is used to calculate the mask for sparse regions , reflects the difference between color and depth in a single view; to ensure the stability of tracking, a soft mask and an error mask are introduced to prevent pixels from unobserved or poorly reconstructed regions from affecting the tracking loss. The soft mask comes directly from the map of the active sub-map, and the error boolean mask discards all pixels with depth errors exceeding the error threshold, is the weight used to control the color difference and depth difference, and the loss function is obtained as: ; The camera pose is predicted by minimizing the color loss and depth loss of the image under relative camera pose transformation. However, this method only relies on single-view information, and when the camera observes sparse regions, it may cause artifacts in the color map and distortion in the depth map. These problems will invalidate the loss calculation of a large number of pixels, thus reducing the tracking accuracy. To enhance the robustness of tracking, the present invention adopts a method similar to bundle adjustment (BA), introducing surface-to-surface projection error and Gaussian-to-Gaussian projection error. Due to the inherent difficulty in directly constraining the properties of Gaussians, the present invention adopts a method different from the traditional BA method. The traditional BA method jointly optimizes the camera pose and the point cloud, while the present invention chooses to only optimize the camera pose and keep the Gaussian parameters unchanged. The pose optimization loss function of the local bundle adjustment is: , where represents the tracking loss, and the surface-to-surface projection error represents the distance between the true surface observation points from two different views. Among them, represents the rigid body transformation from the camera of the th frame to the camera of the th frame, represents the 3D coordinates of the object surface observation under the current relative transformation mapping, represents the th frame; the Gaussian-to-Gaussian projection error represents the distance between Gaussian points from two different views, Represents the Gaussian observation 3D coordinates under the current relative transformation mapping, Represents the Gaussian observation 3D coordinates under the camera of the

[0027] Step Four: Based on the optimized submap, use visually Figure 1 Consistent 2D Gaussian for scene representation, calculate the intersection of the camera ray and the 2D Gaussian, and perform unbiased depth estimation; The 2D Gaussian is obtained by removing one dimension of the scaling matrix to compress the 3D Gaussian into a 2D Gaussian. The 3D Gaussian represents the scene through a set of translucent Gaussian ellipsoids and renders the image using a specific Gaussian rasterizer. Each Gaussian has a mean, covariance matrix, color, and opacity. The 3D Gaussian can be expressed as: , To optimize the covariance matrix, the present invention decomposes the covariance matrix into a scaling matrix and a rotation matrix : , where, is the covariance matrix of the 3D Gaussian, which is decomposed into the scaling matrix and the rotation matrix , is the transpose of the matrix.

[0028] To project the 3D Gaussian onto the rendering plane, the present invention uses the world-to-camera projection matrix to transform the coordinates of the Gaussian and uses the Jacobian matrix for local affine transformation approximation: , is the Jacobian matrix, which prevents the Gaussian from being distorted in the non-linear projection, is the external camera parameters, which project the Gaussian from the world coordinate system to the pixel coordinate system, is the covariance matrix of the Gaussian in the world coordinate system, is the covariance matrix of the Gaussian in the pixel coordinate system.

[0029] After depth sorting the projected Gaussian, the pixel color is calculated by blending along the camera ray direction: , where, represents the color of the current pixel, represents the set of Gaussians intersecting with the ray, is the th color of the Gaussian, is the opacity of the th Gaussian, representing the transparency of the previous

[0030] Due to the lack of Figure 1 consistency, the surface reconstruction effect of the 3D Gaussian is poor. To solve this problem, the present invention adopts a Figure 1 consistent 2D Gaussian representation. By removing one dimension of the scaling matrix, the 3D Gaussian is compressed into a 2D disk. The 2D Gaussian can be expressed as: , where and are the coordinates in the local coordinate system; the parameters of the 2D Gaussian include: is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, serves as the scaling matrix, serves as the rotation matrix, where represents the normal vector of the 2D Gaussian. The spatial transformation matrix of the 2D Gaussian has the following expression: , where is the projection matrix of the local coordinate system of the 2D Gaussian, is the center position of the 2D Gaussian, and are the two principal tangent vectors of the 2D Gaussian, and are the scaling coefficients of the two principal tangent vectors of the 2D Gaussian; Assume is the transformation matrix from the world coordinates to the pixel coordinates. Then the 3D spatial point is calculated by the following formula: ; where represents the uniform light emitted from the camera, passing through the pixel and intersecting with the 2D Gaussian at the depth , is the homogeneous coordinate of the 3D spatial point, is the transformation matrix from the world coordinate system to the pixel coordinate system, is the local projection matrix of the 2D Gaussian, is the homogeneous coordinate of a point within the local coordinate system of the 2D Gaussian.

[0031] The calculation of the intersection of the camera ray and the 2D Gaussian specifically includes obtaining the intersection point through two planes, given a pixel point , defining the intersection point of this point on the plane and the plane: The normal vector of the plane is (-1, 0, 0), and the offset is , and the definition of the plane is as follows: , the normal vector of the plane is (0, -1, 0), and the offset is , and the definition of the plane is as follows: , The local coordinate point of the 2D Gaussian is , and the intersection point of the two planes is obtained using the following equation: , Thus, the solution is: , where: and are the i-th parameters of the plane vector.

[0032] Previous methods used the component of the Gaussian mean as the depth. However, when the intersection point of the ray and the Gaussian distribution is far from the mean, significant errors will be introduced in depth estimation. The present invention calculates the unbiased depth by solving the intersection point of the ray and the Gaussian surface, and this depth can be obtained through the following equation for calculating the unbiased depth: , where, is the opacity of the j-th Gaussian, is the opacity of the j-th pixel, is the intersection point of the ray and the Gaussian, and the vector is calculated from the extrinsic camera parameters, the scaling matrix, and the rotation matrix of the Gaussian. By converting the homogeneous coordinate to the inhomogeneous coordinate, can be obtained, where, is the extrinsic camera parameter, is the spatial transformation matrix of the 2D Gaussian. is the Gaussian mean. Since all rays come from the same viewpoint and are coplanar with the given Gaussian distribution, the intersection point equation of each Gaussian distribution only needs to be solved once.

[0033] Step 5: Merge all optimized sub - maps into a global map, and optimize the color of the global map and the geometric properties of the Gaussian by minimizing the loss function; To obtain an accurate global map by improving the accuracy of unbiased depth, Gaussian optimization is required. The Gaussian optimization includes the calculation of regularization loss and normal loss. The regularization loss includes scale regularization and scale - ratio regularization. The regularization loss function is: : Scale regularization: An overly large Gaussian will cause artifacts and reduce the accuracy of depth estimation. 3DGS prevents this by splitting an overly large Gaussian into multiple smaller Gaussians. However, this increases memory overhead and training time, making it inapplicable in SLAM and causing the Gaussian scale to grow uncontrollably. To solve this problem, the present invention proposes a scale - regularization method to limit the scale of the Gaussian. The scale - regularization loss is as follows: , where: controls the regularization weight, = 0.05, is the scaling matrix of the i - th 2D Gaussian, is the maximum scale threshold, is the number of Gaussians exceeding the threshold; Scale regularization dynamically calculates the scale of each Gaussian during training, compares it with the threshold, and imposes a penalty on Gaussians exceeding the threshold; Scale - ratio regularization: To prevent the Gaussian from being overly stretched, previous methods introduced isotropic regularization to ensure that the Gaussian expands uniformly in all directions and suppress anisotropy. However, this method affects the rendering quality. Isotropic regularization limits the diversity of the Gaussian distribution and reduces the robustness of the model. During the rendering process, appropriate anisotropy helps the Gaussian better adapt to the geometric structure, thereby improving the rendering effect and making the image more physically realistic. To prevent the Gaussian from being overly stretched while maintaining high rendering quality, the present invention proposes a scale - ratio regularization method, and its loss function is as follows: , where, controls the regularization weight, = 0.05, and are the major - axis vector and minor - axis vector of the i - th 2D Gaussian respectively, is the maximum scale - ratio threshold, is the number of all Gaussians exceeding the threshold. Scale - ratio regularization only restricts extreme cases exceeding the threshold to prevent the Gaussian shape from being overly stretched. For Gaussians with a reasonable scale ratio, this regularization does not intervene, thus maintaining moderate anisotropy; When using depth loss to optimize 2D Gaussians, some 2D Gaussians may have unreasonable directions, such as being embedded inside the object surface. This increases the depth prediction error and may lead to the degradation of a large number of 2D Gaussians. To solve this problem, ensure that the normal vector of the 2D Gaussian is aligned with the surface normal of the object, thereby improving the accuracy of depth prediction and enhancing the stability of the Gaussian distribution. The present invention regards the real depth map as the object surface and defines a normal loss to constrain the normal vector of the 2D Gaussian to be aligned with the object surface normal. The normal loss function is: , where is the mixing weight of the current intersection point, where , is the Gaussian opacity at the j-th pixel, is the opacity at the j-th pixel. is the normal of the current 2D Gaussian in the camera direction, is the object surface normal estimated from the real depth map, and are the gradients of the depth map in the and directions respectively.

[0034] After calculating the regularization loss and the normal loss, the color of the global map and the geometric properties of the Gaussian are optimized by minimizing the loss function. Minimizing the loss function includes color loss, depth loss, normal loss, and regularization loss; The minimizing loss function is: where is the color loss, is the depth loss, is the normal loss, is the regularization loss.

[0035] Step Six: Perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map, and a normal map, and complete the reconstruction of the three-dimensional scene based on the color map, the depth map, and the normal map.

[0036] Perform Gaussian pruning and exposure compensation on the optimized global map. The most critical factor in Gaussian pruning is how to correctly evaluate the contribution of each Gaussian. 3DGS adopts a pruning method based on opacity, where the contribution of a Gaussian is determined by its opacity parameter, and Gaussians with lower opacity are pruned. However, due to the instability of the opacity parameter, some Gaussians that are crucial for rendering quality may be wrongly assigned a low opacity and thus be mispruned. In contrast, the pruning method based on color contribution judges the importance of each Gaussian by evaluating its color contribution, and only prunes Gaussians that have almost no color contribution to the final rendering. This pruning strategy is more stable and can effectively avoid negative impacts on rendering quality. The Gaussian pruning calculation formula is as follows: , where, is the color contribution of the Gaussian at the j-th view, is the 2D projection area of the Gaussian at the j-th view, is the pixel value at the j-th view, is the opacity of the current 2D Gaussian, is the color value of the current 2D Gaussian, is the accumulated opacity for the current pixel, is a factor used to control the balance between transparency and color contribution. When the color-based pruning degrades to transparency-based pruning; Since the contribution of a Gaussian is high in a few views, its overall contribution is determined by the average contribution of these high-contribution views. The multi-view contribution calculation formula of the Gaussian is as follows: , where, is the multi-view set with the highest contribution. Here, the 5 single views with the highest contribution are selected for calculation; During the image acquisition process in a real scene, light changes and view-related reflection phenomena are very common. These factors often lead to different exposure effects for images from different views. This exposure difference will introduce serious color inconsistencies, thus greatly affecting the quality of scene reconstruction. To solve this problem, the present invention uses an exposure compensation algorithm to optimize the image color and ensure the consistency and accuracy of image information during the reconstruction process. The present invention applies an affine transformation matrix to each key-frame image for correction. The exposure compensation formula is: , where, is the predicted color map of the current frame, is a The transformation matrix is used to adjust the intensity and proportional relationship of the image color channels. By independently adjusting different color channels, it can effectively correct color deviations caused by lighting changes. is a bias vector, which is used to adjust the overall brightness of the image to compensate for the lighting intensity difference between different viewpoints. The matrix and the vector are obtained through training, so as to achieve adaptive compensation for exposure differences.

[0037] To comprehensively evaluate the performance, we adopted a variety of evaluation metrics. The tracking accuracy is measured by the root mean square of the absolute trajectory error (ATE RMSE). A lower ATE RMSE value indicates higher accuracy. The rendering quality is evaluated using the peak signal-to-noise ratio (PSNR), the structural similarity index (SSIM), and the perceptual image patch similarity (LPIPS). These metrics are calculated after full-resolution rendering along the estimated trajectory, and the mapping interval is set to 5 frames. The reconstruction performance is evaluated on the mesh generated by the Truncated Signed Distance Function algorithm and quantified by the F1 score (the harmonic mean of precision and recall). In addition, the depth L1 error of unobserved viewpoints is reported as a supplementary evaluation metric.

[0038] Meanwhile, multiple representative datasets were selected to evaluate the performance of our method. For the synthetic dataset, the Replica dataset was used, which contains 3D reconstructions of high-quality indoor scenes collected through RGB-D sensor trajectories. The real-world datasets include TUM-RGBD and ScanNet. The poses of TUM-RGBD are obtained by an external motion capture system, while the poses of ScanNet are estimated by BundleFusion. Such a dataset selection enables us to comprehensively and objectively evaluate the effectiveness and robustness of the algorithm in different scenarios.

[0039] Example 2 A three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering in Example 1. Its system includes: A data acquisition module, configured to acquire the RGB-D video stream of the target scene; A map construction module, configured to create a sub-map based on the first frame of the RGB-D video stream using the initialization strategy of camera motion, initially estimate the camera pose for the key frames of the created sub-map using visual odometry, and optimize the sub-map through loss calculation; A tracking module, configured to perform frame-to-model tracking based on the constructed sub-map, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; The scene representation module is configured to use visual Figure 1 The consistent 2D Gaussian is used to represent the scene, the intersection of the camera ray and the 2D Gaussian is calculated, and the unbiased depth estimation is performed; The data optimization module is configured to merge all optimized sub-maps into a global map and optimize the color and Gaussian geometric properties of the global map by minimizing a loss function; The data output module is configured to perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map and a normal map, and complete the reconstruction of the three-dimensional scene based on the color map, the depth map and the normal map.

[0040] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.

Claims

1. A 3D scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering, characterized in that: The following steps are involved: Collect RGB-D video stream of the target scene; Based on the first frame of the RGB-D video stream, a submap is created using the camera motion initialization strategy. The visual odometer is used to preliminarily estimate and optimize the camera pose for the key frame of the created submap, and the submap is optimized through loss calculation. Perform frame-to-model tracking based on the constructed submap, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose through local bundle adjustment; Based on the optimized submap, a view-consistent 2D Gaussian is used to represent the scene, the intersection of the camera ray and the 2D Gaussian is calculated, and unbiased depth estimation is performed; Merge all optimized sub-maps into a global map, and optimize the color and Gaussian geometric properties of the global map by minimizing the loss function; Gaussian pruning and exposure compensation are performed on the optimized global map to obtain a color map, a depth map, and a normal map, and the three-dimensional scene is reconstructed based on the color map, the depth map, and the normal map.

2. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 1, characterized in that: The initialization strategy of the camera motion includes the translation of the current frame relative to the first frame of the active submap exceeding a predefined threshold. , or when the estimated Euler angle exceeds a predetermined threshold , a new submap is created, and at any point in time, only the currently active submap is processed.

3. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 2, characterized in that: After the keyframes of the created submaps are preliminarily estimated using visual odometers, dense point clouds are calculated using the RGB-D measurements of the keyframes. At the beginning of each submap, the point clouds of the keyframes are uniformly sampled. points, and give priority to adding new Gaussian models to areas with high color gradients; for subsequent keyframes in the submap, uniform sampling is performed from areas where the transparency value is below a certain threshold points, the new Gaussian model is in the current sub-map If there are no neighbors in the neighborhood, the search radius will be based on the current submap. The sampling points are added.

4. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 3 is characterized in that: The optimizing the submap by loss calculation includes calculating the minimum loss optimized submap to render the color and depth of the key frame; The minimum loss includes color loss and depth loss, and the minimum loss function is: , in, is color loss, is the true color image, To predict the color image, Reflects the structural similarity between the predicted color map and the true color map; is the depth loss, is the real depth image, To predict the depth image, The weights used to control color differences and structural differences, represents the 1 norm, that is and .

5. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 4 is characterized in that: The camera pose is accurately estimated by calculating the tracking loss and the camera pose is optimized by local bundle adjustment, including the introduction of soft Mask and error mask ,soft Mask Transparency map directly from the active submap, error boolean mask Eliminate all pixels whose color or depth error exceeds the threshold and derive the tracking loss function; use the local bundle adjustment loss function that keeps the Gaussian parameters unchanged to optimize the camera pose; The tracking loss function is: , in: is the weight used to control the color difference and depth difference; The local bundle adjustment loss function is: , in, represents the tracking loss, the surface-to-surface projection error represents the distance between the real surface observation points from two different perspectives, where Indicates that from Frame camera to The rigid body transformation of the frame camera, Represents the 3D coordinates observed on the object surface under the current relative transformation mapping. Indicates 3D coordinates of the object surface in the frame camera coordinate system; Gaussian to Gaussian projection error represents the distance between Gaussian points from two different perspectives, Represents the Gaussian observed 3D coordinates under the current relative transformation mapping, Indicates Gaussian observation of 3D landmarks under frame camera.

6. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 5, characterized in that: The method comprises: using a view-consistent 2D Gaussian to represent the scene, calculating the intersection of the camera ray and the 2D Gaussian and performing unbiased depth estimation, including compressing the 3D Gaussian into a 2D Gaussian by removing one dimension of the scaling matrix, using the 2D Gaussian to represent the scene, solving the intersection of the camera ray and the 2D Gaussian and calculating the unbiased depth to determine the Gaussian depth; The expression of the 2D Gaussian is: , in, and For local Coordinates in the coordinate system, the parameters of the 2D Gaussian include: is the center position of the 2D Gaussian, and are the two main tangent vectors of the 2D Gaussian, As the scaling matrix, As a rotation matrix, Represents the normal vector of 2D Gaussian, the spatial transformation matrix of 2D Gaussian The expression is: , in: For a 2D Gaussian The projection matrix of the local coordinate system, is the center position of the 2D Gaussian, are the two main tangent vectors of the 2D Gaussian, is the scaling factor of the two main tangent vectors of the 2D Gaussian; is the transformation matrix from world coordinates to pixel coordinates, 3D space point The formula is: , in: Represents a uniform ray of light emitted from the camera, passing through the pixel and at depth It intersects with the 2D Gaussian at is the homogeneous coordinate of a point in 3D space, is the transformation matrix from the world coordinate system to the pixel coordinate system, is the local projection matrix of the 2D Gaussian, is the homogeneous coordinate of a point in the 2D Gaussian local coordinate system, is the transpose of the matrix; The intersection of the camera ray and the 2D Gaussian is: , in: and is the i-th parameter of the plane vector; The unbiased depth estimation function is: , in, is the opacity of the jth Gaussian, is the opacity of the jth pixel, is the intersection of the ray and the Gaussian, the vector It is calculated by the camera external parameters, Gaussian scaling matrix and rotation matrix. is the Gaussian mean.

7. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 6, characterized in that: In order to improve the accuracy of unbiased depth calculation, Gaussian optimization must be performed, and the Gaussian optimization includes regularization loss and normal loss calculation. The regularization loss includes scale regularization to limit the Gaussian scale and scale ratio regularization to prevent Gaussian from over-stretching. The normal loss can make the 2D Gaussian tangent to the actual surface; The regularization loss function is: , Where: scale regularization middle Controls the regularization weight, is the scaling matrix of the i-th 2D Gaussian, is the maximum scale threshold, is the number of Gaussians exceeding the threshold; scale ratio regularization middle and are the major axis vector and minor axis vector of the i-th 2D Gaussian, Controls the regularization weight, is the maximum scale ratio threshold, is the number of Gaussians exceeding the threshold; The normal loss function is: , in, is the mixing weight of the current intersection, where , is the Gaussian opacity at the jth pixel, is the opacity of the jth pixel, is the normal of the current 2D Gaussian in the direction of the camera, is the object surface normal estimated from the true depth map, and The depth map is and Directional gradient.

8. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 7, characterized in that: The color and Gaussian geometric properties of the global map are optimized by minimizing the loss function, including calculating the color loss, depth loss, normal loss and regularization loss of the global map to optimize the global map. The minimization loss function is: , in, is color loss, is the depth loss, is the normal loss, is the regularization loss.

9. The three-dimensional scene reconstruction and real-time SLAM method based on 2D Gaussian sputtering according to claim 8, characterized in that: The optimized global map is subjected to Gaussian pruning and exposure compensation, including a pruning method based on color contribution, which determines the importance of each Gaussian by evaluating its color contribution, prunes the Gaussian that has no color contribution to the final rendering, and uses a method of applying a transformation matrix to the key frame image for correction to perform exposure compensation calculation. The Gaussian pruning multi-view contribution calculation formula is: , in, is the multi-view set with the highest contribution, is the Gaussian pruning single view contribution calculation formula, where is the color contribution of Gaussian at the jth viewing angle, is the 2D projection area of ​​Gauss at the jth viewing angle, is the pixel value at the jth viewing angle, is the opacity of the current 2D Gaussian, is the color value of the current 2D Gaussian, The accumulated opacity of the current pixel. is a factor used to control the balance between transparency and color contribution. When , color-based pruning degenerates into transparency-based pruning; The exposure compensation formula is: , in, is the color map after exposure compensation, is the color map predicted for the current frame, is a The transformation matrix is ​​used to adjust the intensity and proportional relationship of the image color channels. By independently adjusting different color channels, the color deviation caused by lighting changes can be effectively corrected. is a The bias vector of .

10. A 3D scene reconstruction and real-time SLAM system based on 2D Gaussian sputtering, characterized in that: include: The data acquisition module is configured to acquire an RGB-D video stream of a target scene; The map construction module is configured to create a submap based on the first frame of the RGB-D video stream using a camera motion initialization strategy, perform preliminary camera pose estimation on the key frames of the created submap using a visual odometry, and optimize the submap through loss calculation; The tracking module is configured to perform frame-to-model tracking based on the constructed submap, accurately estimate the camera pose by calculating the tracking loss, and optimize the camera pose by local bundle adjustment; The scene representation module is configured to use a view-consistent 2D Gaussian for scene representation based on the optimized submap, calculate the intersection of the camera ray and the 2D Gaussian and perform unbiased depth estimation; The data optimization module is configured to merge all optimized sub-maps into a global map and optimize the color and Gaussian geometric properties of the global map by minimizing a loss function; The data output module is configured to perform Gaussian pruning and exposure compensation on the optimized global map to obtain a color map, a depth map and a normal map, and complete the reconstruction of the three-dimensional scene based on the color map, the depth map and the normal map.

Citation Information

Patent Citations

  • Large-scale three-dimensional scene real-time reconstruction method based on Gaussian expression

    CN118314280A

  • Dense RGB-D SLAM method based on multistage 3D Gaussian

    CN118781189A

  • Underground long-distance pipeline three-dimensional mapping method and system based on three-dimensional Gaussian sputtering

    CN119516125A

  • Method for 3D scene dense reconstruction based on monocular visual slam

    US20200273190A1

Cited By

  • Point projection type three-dimensional reconstruction and segmentation method and system based on semi-Gaussian pruning

    CN120635367A

  • A point projection type three-dimensional reconstruction and segmentation method and system based on semi-gaussian pruning

    CN120635367B

  • Improved AD-GS three-dimensional reconstruction method based on 2DGS

    CN120807798A

  • Improved AD-GS three-dimensional reconstruction method based on 2DGS

    CN120807798B

  • Three-dimensional scene simulation system and method based on VR technology

    CN120997460A