Dynamic environment three-dimensional Gaussian spatter SLAM mapping method and device based on uncertainty prediction

Through an improved three-dimensional Gaussian splash SLAM mapping method, uncertainty prediction and multimodal consistency estimation technology are used to eliminate dynamic feature points, which solves the positioning error and map quality problems of the visual SLAM system in dynamic environments and realizes high-precision, real-time dynamic environment mapping.

CN120672978APending Publication Date: 2025-09-19ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510692904.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing visual SLAM systems have difficulty accurately distinguishing dynamic and static elements in dynamic environments, resulting in the accumulation of pose estimation errors and the degradation of map quality, and are unable to meet the dual requirements of real-time performance and accuracy in complex dynamic scenes.

Method used

A three-dimensional Gaussian splash SLAM mapping method for dynamic environments based on uncertainty prediction is adopted. By improving structured 3DGS, sinusoidal position encoding, shallow multi-layer perceptron and DINO visual feature distillation technology, combined with multimodal consistency estimation method, feature points in dynamic masks are eliminated, dynamic areas are optimized, and robust construction of static maps and effective suppression of dynamic interference are achieved.

Benefits of technology

It significantly improves the robustness and positioning accuracy of mapping in dynamic environments, reduces dependence on prior information, supports parallel execution of tracking and map optimization, and improves the real-time performance of the system and the compactness of map representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672978A_ABST
    Figure CN120672978A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic environment three-dimensional Gaussian spatter SLAM mapping method and device based on uncertainty prediction, and the method comprises the steps: capturing an image and a depth map in a dynamic environment, and achieving a self-adaptive structured 3DGS through the improvement of the structured 3DGS; time sequence information is embedded into a 3DGS attribute decoder through a sine position encoder; the visual features of the DINO are distilled to 3DGS through a superficial multilayer perceptron (MLP); aiming at the key frame, utilizing a dynamic region identification method based on multi-modal consistency to remove feature points in the identified dynamic mask; the invention discloses an uncertainty prediction method based on DINO visual features, and the method is used for carrying out the continuous optimization of a dynamic region, and achieves the steady construction of a static map and the effective inhibition of dynamic interference. According to the method, map representation compactness and dynamic adaptability are improved, dependence on prior is eliminated, parallel execution of tracking and map optimization is supported, and mapping robustness and positioning precision in a dynamic environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent robots, and in particular to a dynamic environment three-dimensional Gaussian splash SLAM mapping method and device based on uncertainty prediction. Background Art

[0002] With the rapid development of robotics, autonomous driving, augmented reality, and other fields, Visual Simultaneous Localization and Mapping (Visual SLAM) has garnered widespread attention as a key enabling technology for environmental perception and autonomous navigation. This technology processes image data captured by cameras to enable autonomous positioning and mapping of devices in unknown environments. With its low cost and rich information, it has become a crucial component of many intelligent systems.

[0003] Existing visual SLAM systems have achieved relatively mature research results in static environments, capable of achieving high-precision positioning and map reconstruction. However, in practical applications, the environment often contains moving objects such as pedestrians, vehicles, and animals, constituting typical dynamic scenes. In such environments, traditional visual SLAM systems, unable to accurately distinguish between dynamic and static elements, easily mistakenly incorporate dynamic targets into the mapping and positioning process, resulting in accumulated pose estimation errors, degraded map quality, and even system failure. Therefore, how to improve the robustness and accuracy of visual SLAM systems in dynamic environments has become one of the key and difficult issues in current research.

[0004] To address these challenges, researchers have proposed various improved methods. Patent No. CN119540942A proposes a dynamic SLAM method that combines a lightweight target detection network. This method introduces a balanced convolutional structure (GSConv) and a feature fusion module (VoVGSCS) to optimize the YOLOv11 network structure, making it more lightweight and efficient. The improved YOLOv11 is then integrated into the ORB-SLAM3 framework, enabling localization and mapping in dynamic environments. However, this method still relies on prior knowledge of target categories during dynamic target detection, making it difficult to handle dynamic interference from unknown categories or those without clear semantic features. Another technology, Patent No. CN116299525A, utilizes a spatial partitioning strategy to divide the scene into several independent regions and identifies dynamic regions based on the local motion characteristics of each region. This method not only effectively separates dynamic and static regions, but also eliminates the reliance on prior information about dynamic objects, offering certain advantages in dynamic region identification. However, this patent still suffers from a lack of map information, limiting its application in representing environmental semantics. Overall, although the above patents have improved the system's adaptability to dynamic environments to a certain extent, they still face problems such as reliance on prior information of dynamic objects and lack of map semantic information. It is still difficult to meet the dual requirements of real-time and accuracy in complex dynamic scenarios. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a dynamic environment three-dimensional Gaussian splash SLAM mapping method and device based on uncertainty prediction.

[0006] To achieve the above objectives, the first aspect of the present invention provides the following technical solutions:

[0007] A method and device for constructing three-dimensional Gaussian splash SLAM mapping in a dynamic environment based on uncertainty prediction includes the following steps:

[0008] S1: Capture images and depth maps in dynamic environments and implement adaptive structured 3DGS by improving structured 3DGS;

[0009] S2: embeds the timing information into the 3DGS attribute decoder through the sinusoidal position encoder;

[0010] S3: Distilling DINO visual features into 3DGS through a shallow multi-layer perceptron (MLP);

[0011] S4: Using a dynamic region recognition method based on multimodal consistency for key frames, the feature points within the identified dynamic mask are removed;

[0012] S5: An uncertainty prediction method based on DINO visual features is used to continuously optimize dynamic areas, achieve robust construction of static maps and effectively suppress dynamic interference.

[0013] Furthermore, the step S1 includes the following steps:

[0014] S11: Capture images and depth maps in dynamic environments, and estimate camera pose to obtain camera pose T cw ;

[0015] S12: Combine the depth map and camera pose T cw Generate point cloud, perform voxel filtering on the point cloud, and generate sparse anchor points;

[0016] S13: Based on the ray pointing from the camera's optical center to the anchor point, the probability update formula is combined;

[0017]

[0018] The update equation is based on Bayes’ theorem and requires the prior probability P(n), the current observation z t And update the likelihood model P(n|z) of each anchor point 1:t-1 ). Among them, P(n|z t ) represents the probability that anchor point n is occupied given an observation.

[0019] Furthermore, step S2 includes the following steps:

[0020] S21: normalize the frame number t by the maximum and minimum values;

[0021] S22: Map the sequence number described in S21 to the implicit space through sinusoidal position coding, where the sinusoidal position coding is:

[0022]

[0023] S23: Embed the implicit spatial information proposed in S22 into the attribute decoder of the 3D Gaussian ellipsoid, taking the decoding of color attributes as an example:

[0024]

[0025] The input of the decoder is the distance between the camera view and the anchor point Relative direction δ vc , Anchor point characteristics And the position code 1 described in S22 t Through multiple independent shallow perceptrons, the transparency {o}, scale {s}, rotation {q}, color {c}, and low-dimensional visual features {f} attributes can be decoded.

[0026] S24: Render the Gaussian ellipsoid attributes decoded in S23 using 3DGS technology to obtain color depth Low-dimensional visual features And calculate the cumulative refractive index

[0027] Furthermore, step S3 includes the following steps:

[0028] S31: The low-dimensional visual feature {f} attribute obtained by decoding S23 is passed through a shallow multi-layer perceptron Mapping to high-dimensional space and converting to DINO visual features

[0029]

[0030] S32: Designing a visual feature supervision loss function for optimization

[0031]

[0032] Among them, F i Represents the i-th vector of visual features extracted by the DINOv2 model, N d The visual feature dimension extracted by the DINOv2 model.

[0033] Furthermore, step S4 includes the following steps:

[0034] S41: When ORBSLAM3 is about to generate a keyframe, the color and depth images rendered in S2 and the high-dimensional DINO visual features decoded in S31 are compared with the corresponding true values ​​to calculate the residual:

[0035]

[0036] B represents a 3×3 box filter, which is represented by the convolution symbol Acts on the depth map; Indicates that the upper limit of the output result is truncated to 1. The DINO feature F is bilinearly interpolated and upsampled to the image size.

[0037] S42: Input the residual into the objective function and minimize the objective function by optimizing the uncertainty graph σ:

[0038]

[0039] H and W are the image height and width respectively.

[0040] S43: Binarize the uncertainty map σ described in S42 to obtain a dynamic mask M:

[0041] M=δ(2σ 2 >1) (8)

[0042] S44: Fill the segmentation result generated by the YOLOv8-Seg network with the dynamic mask described in S43 to further refine the dynamic mask.

[0043] S45: Eliminate all feature points of M in the dynamic mask to prevent them from being converted into map points.

[0044] Furthermore, step S5 includes the following steps:

[0045] S51: Obtain the DINO visual feature F of the image through the DINOv2 visual feature large model;

[0046] S52: Take the DINO visual features as input and input them into the shallow multi-layer perceptron middle:

[0047]

[0048] S53: Use the residual of S41 as supervision to optimize the parameters of the shallow multilayer perceptron in S52. The loss function is:

[0049]

[0050] S54: Shallow Multilayer Perceptron The output uncertainty map is binarized to generate a dynamic mask M;

[0051] S55: Integrate the dynamic mask constructed in S54 into the total loss function to achieve supervised optimization of the 3D Gaussian ellipsoid parameters:

[0052]

[0053]

[0054] Among them, SSIM is the structural similarity function, C and D are the true image and depth respectively, and {λ} is a hyperparameter. is the mean of the Gaussian parameters.

[0055] A second aspect of the present invention relates to a dynamic environment three-dimensional Gaussian splash SLAM mapping system based on uncertainty prediction, which includes the following modules:

[0056] Adaptive Structured Anchor Module: Captures images and depth maps in dynamic environments and implements adaptive structured 3DGS by improving structured 3DGS;

[0057] Position encoding module: embeds timing information into the 3DGS attribute decoder through a sinusoidal position encoder;

[0058] Visual feature distillation module: distills DINO visual features into 3DGS through a shallow multi-layer perceptron;

[0059] Dynamic region detection and removal module: This module uses a dynamic region recognition method based on multimodal consistency for key frames and removes feature points within the identified dynamic mask.

[0060] Dynamic mask continuous optimization module: Based on the uncertainty prediction method of DINO visual features, it is used to continuously optimize dynamic areas, achieve robust construction of static maps and effectively suppress dynamic interference.

[0061] The third aspect of the present invention relates to a dynamic environment three-dimensional Gaussian splash SLAM mapping device based on uncertainty prediction, which includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the dynamic environment three-dimensional Gaussian splash SLAM method based on uncertainty prediction of the present invention.

[0062] A fourth aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the three-dimensional Gaussian splash SLAM mapping method for dynamic environments based on uncertainty prediction of the present invention.

[0063] This paper compresses three-dimensional Gaussians into structured anchors encoded by a shallow multi-layer perceptron (MLP), and introduces a probabilistic octree structure to achieve adaptive management of anchors, improving the compactness and dynamic adaptability of map representation. An optimization-based multimodal consistency estimation method is proposed, which integrates geometric information and visual features, eliminates reliance on priors, and supports parallel execution of tracking and map optimization. A temporal encoder combined with sinusoidal position coding is constructed to enhance the model's expressiveness. At the same time, an uncertainty estimation mechanism based on visual features is introduced. A shallow MLP is used for pixel-level prediction to continuously optimize dynamic masks, improving mapping robustness and positioning accuracy in dynamic environments.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] By compressing three-dimensional Gaussians (3DGS) into structured anchors encoded by a shallow multi-layer perceptron (MLP), and introducing a probabilistic octree structure to achieve adaptive management of the anchors, the compactness of the map representation and the adaptability in dynamic environments can be effectively improved. The design of decoupling motion mask generation from map optimization supports parallel execution of tracking and mapping, significantly improving the real-time performance of the system. An optimization-based multimodal consistency estimation method is proposed, which integrates geometric information and DINO features to achieve dynamic target recognition without training and reduce dependence on prior information. A temporal encoder combined with sinusoidal position encoding is introduced in the mapping process to embed inter-frame information into the MLP to enhance the expressive power. At the same time, an uncertainty estimation mechanism based on DINO features is constructed, and pixel-by-pixel prediction is performed through a shallow MLP to continuously optimize the motion mask, thereby improving the robustness and accuracy of mapping in dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a flow chart of the method of the present invention;

[0067] Figure 2 This demonstrates the ability of the present invention to eliminate the need for prior information about dynamic objects. The red dashed boxes in the figure represent dynamic objects that the model has never seen before, and the bright areas represent detected dynamic areas.

[0068] Figure 3 This is the comparison result of the present invention and other methods in terms of rendering effect. Figure 3The following graph shows the results of experiments on multiple sequences of the BonnDynamic RGB-D dataset, including Balloon, Person_track2 (Ps_track2), and Pserson_track (Ps_track). The rows in the graph correspond to different sequences, and the columns show the results of different methods. Wherein Groundtruth is the true value of the image, UP-SLAM (Our) is the method of the present invention, the SplaTAM algorithm is available in “SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM”, published in the Proceedings of the 2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), pp. 21357–21366; the Photo-SLAM algorithm is available in “Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular Stereo and RGB-D Cameras”, published in the Proceedings of the 2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), pp. 21584–21593; the DG-SLAM algorithm is available in “DG-SLAM: Robust Dynamic Gaussian Splatting SLAM with Hybrid Pose Optimization" was published in the proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS2024); the WildGS-SLAM algorithm can be found in "WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments", which was accepted by the 2025 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025). DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings. It should be understood that the examples described here are only used to explain the present invention and are not intended to limit the present invention.

[0070] Example 1

[0071] The present invention is a dynamic environment three-dimensional Gaussian splash SLAM mapping method based on uncertainty prediction, comprising the following steps:

[0072] S1: Capture images and depth maps in dynamic environments and implement adaptive structured 3DGS by improving structured 3DGS;

[0073] S2: embeds the timing information into the 3DGS attribute decoder through the sinusoidal position encoder;

[0074] S3: Distilling DINO visual features into 3DGS through a shallow multi-layer perceptron (MLP);

[0075] S4: Using a dynamic region recognition method based on multimodal consistency for key frames, the feature points within the identified dynamic mask are removed;

[0076] S5: An uncertainty prediction method based on DINO visual features is used to continuously optimize dynamic areas, achieve robust construction of static maps and effectively suppress dynamic interference.

[0077] The step S1 comprises the following steps:

[0078] S11: Capture images and depth maps in dynamic environments, and estimate camera pose to obtain camera pose T cw ;

[0079] S12: Combine the depth map and camera pose T cw Generate point cloud, perform voxel filtering on the point cloud, and generate sparse anchor points;

[0080] S13: Based on the ray pointing from the camera's optical center to the anchor point, the probability update formula is combined;

[0081]

[0082] The update equation is based on Bayes’ theorem and requires the prior probability P(n), the current observation z t And update the likelihood model P(n|z) of each anchor point 1:t-1 ). Among them, P(n|z t ) represents the probability that anchor point n is occupied given an observation.

[0083] The step S2 comprises the following steps:

[0084] S21: normalize the frame number t by the maximum and minimum values;

[0085] S22: Map the sequence number described in S21 to the implicit space through sinusoidal position coding, where the sinusoidal position coding is:

[0086]

[0087] S23: Embed the implicit spatial information proposed in S22 into the attribute decoder of the 3D Gaussian ellipsoid, taking the decoding of color attributes as an example:

[0088]

[0089] The input of the decoder is the distance between the camera view and the anchor point Relative direction δ vc , Anchor point characteristics And the position code 1 described in S22 t Through multiple independent shallow perceptrons The corresponding attributes of transparency {o}, scale {s}, rotation {q}, color {c}, and low-dimensional visual features {f} can be decoded.

[0090] S231: Further, the model structure of multiple independent shallow independent perceptrons (transparency MLP Scale MLP Rotating MLP Color MLP Visual Feature MLP );

[0091] S232: Transparency MLP The model structure is: linear layer, ReLU layer, linear layer, Tanh layer, and the hidden layer is 32 dimensions;

[0092] S233: Scale MLP Rotating MLP The same model structure is used, which is: linear layer, ReLU layer, linear layer, and the hidden layer is 32 dimensions;

[0093] S232: Color MLP The model structure is: linear layer, ReLU layer, linear layer, Sigmoid layer, and the hidden layer is 32 dimensions;

[0094] S232: Visual Feature MLP The model structure is: linear layer, Softplus layer, linear layer, and the hidden layer is 32 dimensions;

[0095] S24: Render the Gaussian ellipsoid attributes decoded in S23 using 3DGS technology to obtain color depth Low-dimensional visual features And calculate the cumulative refractive index

[0096] S25: Further details on the 3DGS rendering technology described in S24:

[0097] S251: The environment is constructed as a series of anisotropic 3D Gaussian models with opacity properties and spherical harmonics, represented as:

[0098] G={G i:(X i ,Σ i ,o i ,Η i )|i=1,...,N} (13)

[0099] Each Gaussian is represented by its position in the world coordinate system 3D covariance matrix Transparency o∈[0,1], and the first-order spherical harmonics representing the color of each Gaussian sphere Among them, each Gaussian first-order spherical harmonic function has a total of 12 coefficients, Σ i By the rotation matrix and the scale matrix Composition:

[0100] Σ=RSS T R T (14)

[0101] Note that it is difficult to ensure that the covariance matrix is ​​positive by directly optimizing it, so the rotation matrix R is replaced by a quaternion, and S is represented by a 3D scale vector.

[0102] S252: Obtain the camera's T through the front-end visual odometry wc ∈SE(3), the 3D Gaussian μ in the world coordinate system w Transform to the corresponding two-dimensional plane through the projection matrix μ I

[0103] μ I =π(T cw μ w ) (15)

[0104] The covariance matrix Σ w Projected onto the plane through the affine approximation Jacobian matrix J:

[0105] Σ I =JW -1 Σ w W -T J T (16)

[0106] where π(·) is the projection function and W is T wc The rotation matrix in .

[0107] S253: After obtaining the projected 3D Gaussian points, sort them according to the distance between the Gaussian points in front of and behind the canvas, and then efficiently obtain the pixel color value through the alpha blending method:

[0108]

[0109] S254: Replace the color value c with the depth value d to obtain the rendered depth value:

[0110]

[0111] c i is the color of the Gaussian function obtained by the spherical harmonic coefficient H, α i The density value is the product of the two-dimensional Gaussian equation and the transparency o:

[0112]

[0113] 3D Gaussian center μ 3D Splashing to flat pixels μ 2D , K is the known camera intrinsic parameter matrix, is the transformation matrix from the world coordinates of the kth frame to the camera coordinates, and d is the z-axis distance of the corresponding three-dimensional point, that is, the depth.

[0114] S255: Rendering low-dimensional visual features

[0115]

[0116] S256: Calculate the cumulative refractive index

[0117]

[0118] The step S3 comprises the following steps:

[0119] S31: The low-dimensional visual feature {f} attribute obtained by decoding S23 is passed through a shallow multi-layer perceptron Mapping to high-dimensional space and converting to DINO visual features

[0120]

[0121] S311: Among them The network model results are: convolutional layer, ReLU layer, convolutional layer, and the hidden layer is 128 dimensions.

[0122] S32: Designing a visual feature supervision loss function for optimization

[0123]

[0124] Among them, F i Represents the i-th vector of visual features extracted by the DINOv2 model, N d The visual feature dimension extracted by the DINOv2 model.

[0125] The step S4 comprises the following steps:

[0126] S41: When ORBSLAM3 is about to generate a keyframe, the color and depth images rendered in S2 and the high-dimensional DINO visual features decoded in S31 are compared with the corresponding true values ​​to calculate the residual:

[0127]

[0128] B represents a 3×3 box filter, which is represented by the convolution symbol Acts on the depth map; Indicates that the upper limit of the output result is truncated to 1. The DINO feature F is bilinearly interpolated and upsampled to the image size.

[0129] S42: Input the residual into the objective function and minimize the objective function by optimizing the uncertainty graph σ:

[0130]

[0131] H and W are the image height and width respectively.

[0132] S43: Binarize the uncertainty map σ described in S42 to obtain a dynamic mask M:

[0133] M=δ(2σ 2 >1) (8)

[0134] S44: Fill the segmentation result generated by the YOLOv8-Seg network with the dynamic mask described in S43 to further refine the dynamic mask.

[0135] S45: Eliminate all feature points of M in the dynamic mask to prevent them from being converted into map points.

[0136] The step S5 comprises the following steps:

[0137] S51: Obtain the DINO visual feature F of the image through the DINOv2 visual feature large model;

[0138] S52: Take the DINO visual features as input and input them into the shallow multi-layer perceptron middle:

[0139]

[0140] S521: Among them The network model results are: convolution layer, ReLU layer, convolution layer, Softplus layer, and the hidden layer is 128 dimensions.

[0141] S53: Use the residual of S41 as supervision to optimize the parameters of the shallow multilayer perceptron in S52. The loss function is:

[0142]

[0143] S54: Shallow Multilayer Perceptron The output uncertainty map is binarized to generate a dynamic mask M;

[0144] S55: Integrate the dynamic mask constructed in S54 into the total loss function to achieve supervised optimization of the 3D Gaussian ellipsoid parameters:

[0145]

[0146] Among them, SSIM is the structural similarity function, C and D are the true image and depth respectively, and {λ} is a hyperparameter. is the mean of the Gaussian parameters.

[0147] As attached Figure 2 As shown in Figure 1, the method proposed in this invention can detect dynamic objects that have not been trained. The red dotted box in the figure indicates a dynamic object that the model has never seen, and the bright area indicates the detected dynamic area. Figure 3The results show a comparison of the rendering effects of the present invention with other methods, intuitively demonstrating that the present method can construct high-quality static maps. Table 1 compares the positioning accuracy of the present invention with that of various existing methods on multiple sequences of the Bonn Dynamic RGB-D dataset, including Balloon, Balloon, Balloon_track (Ball_track), Person_track2 (Ps_track2), Pserson_track (Ps_track), and Moving_nonobstructing_box2 (Mv_box2). The results demonstrate that the present method achieves higher positioning accuracy. The result is the absolute pose error (unit: centimeter), Avg. is the mean of all sequence results, the ORB-SLAM3 algorithm can be found in the journal "IEEE Transactions on Robotics", 2021, Vol. 37, No. 6, pp. 1874–1890; the DynaSLAM algorithm can be found in the journal "IEEE Robotics and Automation Letters", 2018, Vol. 3, No. 4, pp. 4076–4083; the ESLAM algorithm can be found in the Proceedings of the CVPR 2023 Conference, pp. 17408–17419; the RoDyn-SLAM algorithm can be found in the journal "IEEE Robotics and Automation Letters", 2024; the Photo-SLAM algorithm can be found in "Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular Stereo and RGB-DCameras", published in the 2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) The GS-SLAM algorithm can be found in the Proceedings of the CVPR 2024 Conference, pp. 21584–21593; the DG-SLAM algorithm can be found in "DG-SLAM: Robust Dynamic Gaussian Splatting SLAM with Hybrid Pose Optimization," published in the Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024). Table 2 shows a comparative analysis of the runtime of our method with other similar methods. The results demonstrate that our method is capable of providing real-time pose estimation and has high runtime efficiency.Where Avg. / frame is the time required to process one frame, TotalTime(+refine) is the time required for the algorithm to complete, refine is the time required for map refinement, and ModelSize is the size of the output map model. UP-SLAM(Our) is the method of our invention. "-" indicates that this method does not provide this function. The SplaTAM algorithm can be found in "SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM", published in the proceedings of the 2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), pages 21357–21366; the DG-SLAM algorithm can be found in "DG-SLAM: Robust Dynamic Gaussian Splatting SLAM with Hybrid PoseOptimization", published in the proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024); the WildGS-SLAM algorithm can be found in "WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments", which was accepted by the 2025 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025).

[0148] Table 1. Absolute pose error evaluation on the Bonn Dynamic RGB-D dataset. The bolded ones are the best results, and the underlined ones are the suboptimal results.

[0149]

[0150] Table 2 evaluates the runtime on the Balloon sequence of the Bonn Dynamic RGB-D dataset. The bolded results are the best.

[0151] The down arrow means smaller is better.

[0152]

[0153] This method effectively addresses several current practical challenges, such as the vulnerability of robot positioning in dynamic environments, the inability to construct high-quality static dense maps, and insufficient real-time pose estimation. This method provides intelligent robots with high-precision pose estimation and information-rich dense maps in real time, significantly improving their environmental perception and autonomous navigation capabilities in complex dynamic scenes.

[0154] Example 2

[0155] This embodiment relates to a dynamic environment three-dimensional Gaussian splash SLAM mapping device based on uncertainty prediction, including a memory and one or more processors, wherein the memory stores executable code. When the one or more processors execute the executable code, they are used to implement the dynamic environment three-dimensional Gaussian splash SLAM mapping method based on uncertainty prediction of Example 1.

[0156] Example 3

[0157] This embodiment relates to a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for constructing three-dimensional Gaussian splash SLAM mapping in a dynamic environment based on uncertainty prediction of embodiment 1 is implemented.

[0158] It should be emphasized that the implementation cases described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific implementation plans. Any other similar implementation plans derived by those skilled in the art based on the technical solution of the present invention also fall within the scope of protection of the present invention.

Claims

1. A dynamic environment three-dimensional Gaussian splash SLAM mapping method based on uncertainty prediction includes the following steps: S1: Capture images and depth maps in dynamic environments and implement adaptive structured 3DGS by improving structured 3DGS; S2: embeds the timing information into the 3DGS attribute decoder through the sinusoidal position encoder; S3: Distill DINO visual features into 3DGS through shallow multi-layer perceptron; S4: Using a dynamic region recognition method based on multimodal consistency for key frames, the feature points within the identified dynamic mask are removed; S5: An uncertainty prediction method based on DINO visual features is used to continuously optimize dynamic areas, achieve robust construction of static maps and effectively suppress dynamic interference.

2. The method according to claim 1, wherein The step S1 comprises the following steps: S11: Capture images and depth maps in dynamic environments, and estimate camera pose to obtain camera pose T cw ; S12: Combine the depth map and camera pose T cw Generate point cloud, perform voxel filtering on the point cloud, and generate sparse anchor points; S13: Based on the ray pointing from the camera's optical center to the anchor point, the probability update formula is combined; The update equation is based on Bayes’ theorem and requires the prior probability P(n), the current observation z t And update the likelihood model P(n|z) of each anchor point 1:t-1 ). Among them, P(n|z t ) represents the probability that anchor point n is occupied given an observation.

3. The method according to claim 1, wherein The step S2 comprises the following steps: S21: normalize the frame number t by the maximum and minimum values; S22: Map the sequence number described in S21 to the implicit space through sinusoidal position coding, where the sinusoidal position coding is: S23: Embed the implicit spatial information proposed in S22 into the attribute decoder of the 3D Gaussian ellipsoid, taking the decoding of color attributes as an example: The input of the decoder is the distance between the camera view and the anchor point Relative direction δ vc , Anchor point characteristics And the position code 1 described in S22 t Through multiple independent shallow perceptrons, the transparency {o}, scale {s}, rotation {q}, color {c}, and low-dimensional visual features {f} attributes can be decoded. S24: Render the Gaussian ellipsoid attributes decoded in S23 using 3DGS technology to obtain color depth Low-dimensional visual features And calculate the cumulative refractive index 4. The method according to claim 1, wherein The step S3 comprises the following steps: S31: The low-dimensional visual feature {f} attribute obtained by decoding S23 is passed through a shallow multi-layer perceptron Mapping to high-dimensional space and converting to DINO visual features S32: Designing a visual feature supervision loss function for optimization Among them, F i Represents the i-th vector of visual features extracted by the DINOv2 model, N d The visual feature dimension extracted by the DINOv2 model.

5. The method according to claim 1, wherein The step S4 comprises the following steps: S41: When ORBSLAM3 is about to generate a keyframe, the color and depth images rendered in S2 and the high-dimensional DINO visual features decoded in S31 are compared with the corresponding true values ​​to calculate the residual: B represents a 3×3 box filter, which is represented by the convolution symbol Acts on the depth map; Indicates that the upper limit of the output result is truncated to 1. The DINO feature F is bilinearly interpolated and upsampled to the image size. S42: Input the residual into the objective function and minimize the objective function by optimizing the uncertainty graph σ: H and W are the image height and width respectively. S43: Binarize the uncertainty map σ described in S42 to obtain a dynamic mask M: M=δ(2σ 2 >1) (8) S44: Fill the segmentation result generated by the YOLOv8-Seg network with the dynamic mask described in S43 to further refine the dynamic mask. S45: Eliminate all feature points of M in the dynamic mask to prevent them from being converted into map points.

6. The method according to claim 1, wherein The step S5 comprises the following steps: S51: Obtain the DINO visual feature F of the image through the DINOv2 visual feature large model; S52: Take the DINO visual features as input and input them into the shallow multi-layer perceptron middle: S53: Use the residual of S41 as supervision to optimize the parameters of the shallow multilayer perceptron in S52. The loss function is: S54: Shallow Multilayer Perceptron The output uncertainty map is binarized to generate a dynamic mask M; S55: Integrate the dynamic mask constructed in S54 into the total loss function to achieve supervised optimization of the 3D Gaussian ellipsoid parameters: Among them, SSIM is the structural similarity function, C and D are the true image and depth respectively, and {λ} is a hyperparameter. is the mean of the Gaussian parameters.

7. A dynamic environment three-dimensional Gaussian splash SLAM mapping device based on uncertainty prediction, characterized in that: The invention comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the dynamic environment three-dimensional Gaussian splash SLAM mapping method based on uncertainty prediction according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, the method for constructing a three-dimensional Gaussian splash SLAM map of a dynamic environment based on uncertainty prediction according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Dynamic environment dense point cloud SLAM method and system based on YOLOv11 and ORB-SLAM3

    CN119540942A

Cited By

  • Positioning and mapping method, system and equipment based on dynamic vision and medium

    CN121026152A