A dynamic scene-oriented 3D Gaussian sputtering SLAM method and system

The 3D Gaussian sputtering SLAM method using bi-branch modeling and adaptive density control solves the problems of insufficient pose estimation robustness and low reconstruction accuracy of SLAM systems in dynamic scenes. It achieves a balance between static scene mapping accuracy and dynamic target modeling quality, meeting the real-time requirements of robot autonomous navigation.

CN122384772APending Publication Date: 2026-07-14SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610353508.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-23
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing 3D Gaussian sputtering SLAM algorithms suffer from the problem of misclassifying dynamic targets as static scenes in dynamic environments, leading to map artifacts and decreased camera pose estimation accuracy. Furthermore, current technologies struggle to balance the accuracy and real-time performance of dynamic target modeling, failing to meet the needs of robot autonomous navigation and 3D scene reconstruction.

Method used

We employ a bi-branch modeling, adaptive density control, and confidence-aware optimization approach. By segmenting human instances and aligning parameterized models, we construct static scene and dynamic human Gaussian tuples. Combining multi-constraint loss functions and confidence-aware loss functions, we achieve camera pose estimation and bi-branch 3D Gaussian map update in dynamic scenes.

Benefits of technology

It improves the robustness and reconstruction accuracy of the SLAM system in dynamic scenes, reduces the error of camera pose estimation, achieves a balance between the accuracy of static scene mapping and the quality of dynamic target modeling, meets real-time requirements, and achieves a frame rate of over 1.2 FPS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122384772A_ABST
    Figure CN122384772A_ABST
Patent Text Reader

Abstract

The application discloses a kind of 3D Gaussian sputtering SLAM methods and systems for dynamic scene, which comprises the following steps: collecting RGB image, depth image and IMU data;RGB image is divided into human dynamic area and static scene area;The posture, shape and facial expression estimation of human dynamic area are carried out, and the alignment of model and RGB frame is realized by multi-constraint optimization;Construct a double-branch Gaussian primitive set, and model the static scene area and human dynamic area with 3D Gaussian;Construct a confidence-aware sub-region loss function to jointly optimize the double-branch Gaussian primitive and camera pose;Based on the tracking and mapping alternately executed process, the original static SLAM process is adaptively modified, and the camera pose estimation and double-branch 3D Gaussian map updating in dynamic scene are realized.The application can balance the static scene mapping precision and dynamic target modeling quality, and improve the robustness of SLAM in dynamic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of synchronous positioning and mapping technology, specifically to a 3D Gaussian sputtering SLAM method and system for dynamic scenes. Background Technology

[0002] Simultaneous Localization and Mapping (SLAM) technology is the core support for autonomous robot movement. It estimates the robot's pose and builds an environmental map in real time using data from sensors such as cameras and IMUs. SLAM algorithms based on 3D Gaussian sputtering (3DGS) have become a research hotspot in the field of dense SLAM due to their advantages of differentiable rendering and real-time mapping. However, existing SLAM algorithms based on 3D Gaussian sputtering (3DGS) are mostly designed for static scenes. In actual indoor environments containing dynamic targets such as humans, dynamic targets are easily misjudged as static scene structures, resulting in geometric artifacts on the map, decreased camera pose estimation accuracy, and even tracking loss. Furthermore, existing dynamic scene adaptation technologies suffer from problems such as low accuracy in dynamic target modeling, poor real-time performance, and inability to simultaneously meet the needs of localization and mapping.

[0003] To address the dynamic scene adaptation problem, existing technologies mainly fall into three categories. The first category is dynamic target culling methods based on feature detection. These methods detect dynamic feature points and remove their contribution to optimization, achieving camera localization in dynamic scenes. However, they struggle with large dynamic regions. When dynamic targets occupy a significant image area, a large number of effective feature points are removed, leading to insufficient optimization constraints. Furthermore, they lack a dynamic target modeling mechanism, meaning that Gaussian primitives in dynamic regions are still included in static map optimization, resulting in artifacts in static scene mapping. The second category is separation and reconstruction methods based on mask segmentation. This method obtains dynamic target masks through semantic segmentation, achieving separation and reconstruction of static scenes and dynamic targets. However, it uses voxel modeling, whose computational complexity and storage space increase exponentially with scene size, resulting in poor real-time performance (frame rate only 0.5). The first category is dynamic human body reconstruction technology based on 3DGS. This method combines the SMPL-X model with 3DGS to achieve high-precision reconstruction of dynamic human bodies, but it is not deeply integrated with the SLAM system. It cannot simultaneously meet the requirements of real-time camera positioning and robust mapping of static scenes. In addition, it adopts a fixed-density Gaussian primitive initialization and optimization strategy, and lacks an adaptive modeling strategy for the fine structure of human hands, faces and other parts of the human body. It cannot adapt to the granularity differences of different parts of the human body, and it is difficult to balance reconstruction accuracy and real-time performance.

[0004] Furthermore, existing lightweight 3DGS SLAM algorithms such as SplaTAM-s use a single set of Gaussian primitives to model the entire scene without distinguishing between dynamic and static regions. The movement of dynamic targets leads to frequent updates of the corresponding Gaussian primitive parameters, which disrupts the consistency of the static scene map and interferes with the optimization of photometric and geometric residuals, thus affecting the accuracy of pose estimation. Summary of the Invention

[0005] To overcome the defects and shortcomings of existing technologies, this invention provides a 3D Gaussian sputtering SLAM method and system for dynamic scenes. Through bi-branch modeling, adaptive density control and confidence-aware optimization, it solves the core technical problems of insufficient pose estimation robustness, low accuracy of dynamic target reconstruction and poor real-time performance of SLAM systems in dynamic scenes, and adapts to the practical application needs of robot autonomous navigation, 3D scene reconstruction and other applications.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] This invention provides a 3D Gaussian sputtering SLAM method for dynamic scenes, comprising the following steps:

[0008] Acquire RGB images, depth images, and IMU data;

[0009] Segment the RGB image into dynamic human body regions and static scene regions;

[0010] Based on the human body parameterization model, the pose, shape and facial expression of the human body dynamic region are estimated, and the alignment of the model with the RGB frame is achieved through multi-constraint optimization.

[0011] Construct a static scene Gaussian set and a dynamic human body Gaussian set, and perform 3D Gaussian modeling on the static scene region and the dynamic human body region.

[0012] A confidence-aware regional loss function is constructed, and the optimization loss is calculated separately for the dynamic human body region and the static scene region, realizing the joint optimization of the bi-branch Gaussian unit and the camera pose.

[0013] Based on the alternating execution of tracking and mapping, the original static SLAM process is adapted to achieve camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes.

[0014] As a preferred technical solution, the RGB image is segmented into a dynamic human body region and a static scene region, specifically including:

[0015] Human instance segmentation is performed based on the SGHM lightweight human instance segmentation model, and a binary human mask is output. ,in, Represents pixels This is the dynamic area of ​​the human body. Represents pixels This is a static scene area.

[0016] As a preferred technical solution, the pose, shape, and facial expression of the human body dynamic region are estimated based on a human body parametric model. Alignment between the model and the RGB frame is achieved through multi-constraint optimization, specifically including:

[0017] Construct initial shape parameters, pose parameters, and facial expression parameters, and output 2D human body key points, 3D body joints, and 3D hand joints;

[0018] Construct a multi-constraint loss function, expressed as:

[0019] ;

[0020] in, Indicates attitude parameters, Indicates shape parameters, Indicates facial expression parameters. This represents the 2D keypoint constraint loss. This indicates the 3D prior loss of the body. For the 3D prior loss of the hand. , These represent the loss weights;

[0021] Iterative optimization based on multi-constraint loss function is used to align the model with the RGB frame.

[0022] As a preferred technical solution, 2D keypoint constraint loss Represented as:

[0023] ;

[0024] in, Let K be the projection function under the camera's intrinsic parameters. This is a joint rotation function driven by attitude parameters. , These are predefined weights and detection confidence weights, respectively. Representing key points of the human body in 2D;

[0025] 3D body prior loss Represented as:

[0026] ;

[0027] in, This is the low-dimensional embedding vector of the pose. For regularization terms;

[0028] 3D prior loss of hand Represented as:

[0029] ;

[0030] in, This represents 3D hand joints. Let be the robust error function acting on the z-axis.

[0031] As a preferred technical solution, a static scene Gaussian set and a dynamic human body Gaussian set are constructed, specifically including:

[0032] Based on the estimated camera pose and depth image, initialize the parameters of the static scene Gaussian unit in the world coordinate system, including the spatial mean, covariance matrix, opacity, and pixel color in the depth image.

[0033] The dynamic human body Gaussian tuple includes the human body mesh vertex coordinates, covariance matrix, opacity, and pixel color of the human body dynamic region, while assigning body part attributes to each Gaussian tuple.

[0034] As a preferred technical solution, 3D Gaussian modeling is performed on the static scene area and the dynamic human body area, specifically including:

[0035] Construct a context-aware adaptive density control strategy to adjust the density of human high-sigma units, including:

[0036] Based on the body part attributes and historical gradient information of Gaussian elements, the compaction threshold of each Gaussian element is calculated.

[0037] When the position gradient of a Gaussian unit exceeds the adaptive threshold, it is split or cloned.

[0038] Calculate the shortest distance between the center of the Gaussian cell and the surface of the human body mesh. If the shortest distance is greater than a set threshold, the Gaussian cell is determined to be invalid and is pruned.

[0039] As a preferred technical solution, the densification threshold of each Gaussian element is calculated and expressed as:

[0040] ;

[0041] in, Densification threshold, where R represents the densification time window. For hyperparameters related to the part attribute U, It represents the position gradient of the center of the i-th Gaussian element in the k-th frame.

[0042] As a preferred technical solution, a confidence-aware regional loss function is constructed, expressed as:

[0043] ;

[0044] ;

[0045] ;

[0046] in, This represents a static scene area. Indicates the dynamic area of ​​the human body. Indicates the loss in a static scene. Indicates loss in a specific area of ​​the human body. Pixel color and depth rendered using Gaussian primitives for static scenes. For the true colors and depth captured by the camera, For visibility weights in static scenes, , These represent the scaling parameter and average scaling value of the Gaussian element, respectively. For regularization weights, , , These represent the corresponding weights. This represents the perceived loss based on confidence level. Indicates mask loss. Represents structural similarity loss. This indicates perceived loss.

[0047] As a preferred technical solution, based on the alternating execution of tracking and mapping, the original static SLAM process is adapted to achieve camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes, specifically including:

[0048] During the tracking phase, a constant velocity motion model is used to initialize the camera pose of the current frame. The static scene Gaussian set and the dynamic human body Gaussian set are rendered separately to obtain the static region color map and depth map and the human body region color map and depth map. The camera pose and human body pose parameters are iteratively optimized based on the confidence-aware regional loss function.

[0049] During the mapping phase, the camera pose is fixed, and 2D scale adaptive filtering, anti-aliasing rendering and isotropic regularization optimization are performed on the static scene Gaussian primitives to update the static scene Gaussian primitives. Adaptive density control is performed on the dynamic human body Gaussian primitives, and the dynamic human body Gaussian primitives are optimized through the confidence-perceived loss function.

[0050] The optimized bi-branch Gaussian set is added to the global map to complete the map update.

[0051] The present invention also provides a 3D Gaussian sputtering SLAM system for dynamic scenes, used to implement the above-mentioned 3D Gaussian sputtering SLAM method for dynamic scenes, including: a data acquisition module, a dynamic region segmentation module, an alignment module, a bi-branch modeling module, a confidence-aware optimization module, and a tracking mapping module;

[0052] The data acquisition module is used to acquire RGB images, depth images, and IMU data;

[0053] The dynamic region segmentation module is used to segment the RGB image into a dynamic human body region and a static scene region.

[0054] The alignment module is used to estimate the posture, shape and facial expression of the human body dynamic region based on the human body parameterization model, and to achieve the alignment of the model with the RGB frame through multi-constraint optimization.

[0055] The dual-branch modeling module is used to construct a static scene Gaussian set and a dynamic human body Gaussian set, and to perform 3D Gaussian modeling on the static scene region and the dynamic human body region.

[0056] The confidence-aware optimization module is used to construct a region-specific loss function for confidence-awareness, calculate the optimization loss for the dynamic human body region and the static scene region respectively, and realize the joint optimization of the dual-branch Gaussian unit and the camera pose.

[0057] The tracking and mapping module is used to adapt the original static SLAM process based on the alternating execution of tracking and mapping, so as to realize camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes.

[0058] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0059] (1) For the dual-branch modeling framework, the existing static 3DGS SLAM (such as SplaTAM-s) uses a single set of Gaussian primitives for modeling, which cannot distinguish between dynamic and static regions. This invention divides dynamic and static regions by human instance segmentation, constructs independent static scene and dynamic human Gaussian primitive sets, realizes separate modeling and joint optimization, fundamentally avoids the interference of dynamic targets on camera pose estimation and static scene mapping, improves the robustness of SLAM system in dynamic scenes, and reduces the ATE RMSE of camera pose estimation by more than 20% compared with the static 3DGS SLAM algorithm. The dual-branch modeling framework constructed by this invention can achieve a balance between the accuracy of static scene mapping and the quality of dynamic target modeling, and provides support for the engineering implementation of dynamic scene SLAM technology.

[0060] (2) For SMPL-X alignment optimization, existing dynamic human reconstruction technologies (such as EVA) do not design an alignment mechanism for SLAM scenarios, which makes the model and image easily misaligned. This invention achieves accurate alignment between the model and RGB frames by using a multi-constraint loss function with 2D key point constraints and 3D body / hand prior constraints, combined with a three-stage optimization strategy, and provides reliable geometric priors. Existing technologies mostly use single constraints or no targeted optimization.

[0061] (3) For adaptive density control, existing 3DGS dynamic reconstruction technology (such as EVA) uses fixed density threshold control, which cannot adapt to the granularity differences of different parts of the human body. This invention combines body part attributes and historical gradient information to dynamically calculate the densification threshold and achieve a denser Gaussian unit distribution for fine parts. This is different from the fixed threshold strategy of existing technologies. It can control the amount of computation while ensuring reconstruction accuracy, and can improve the restoration ability of fine structures such as hands and faces, so that the PSNR of human body reconstruction reaches more than 24.8 dB.

[0062] (3) The present invention can guarantee the real-time performance of the algorithm and achieve a running frame rate of more than 1.2 FPS on mainstream GPU hardware (such as NVIDIA GTX 4070), which meets the real-time application requirements such as robot autonomous navigation. Attached Figure Description

[0063] Figure 1 This is a flowchart illustrating the 3D Gaussian sputtering SLAM method for dynamic scenes according to the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0065] Example 1

[0066] like Figure 1 As shown, this embodiment provides a 3D Gaussian sputtering SLAM method for dynamic scenes, which deeply integrates dynamic region segmentation, human parametric modeling, and 3DGS SLAM to form a complete technology chain of separate modeling, precise alignment, and adaptive optimization. This solves the problem of difficulty in balancing dynamic interference suppression, fine structure reconstruction, and real-time performance in existing technologies. Specifically, it includes the following steps:

[0067] S1: Data Acquisition;

[0068] This embodiment uses an RGB-D camera (such as RealSense D455) and a six-axis IMU (Inertial Measurement Unit) to form a data acquisition unit, which synchronously acquires RGB images, depth images and raw IMU data of a dynamic indoor scene. The resolution of the RGB-D camera is set to 640×480, the frame rate is 30FPS, the depth measurement error is less than or equal to 1%, and the sampling frequency of the IMU is 500 Hz to assist in camera pose initialization. The acquired data is synchronized and aligned through timestamps to ensure the time consistency of the RGB images, depth images and IMU data, providing reliable input for subsequent processing.

[0069] S2: Dynamic region segmentation;

[0070] In this embodiment, each synchronized RGB image is processed by a dynamic region segmentation module. This module uses the SGHM lightweight human instance segmentation model to achieve real-time detection and segmentation of human regions. Specific steps include:

[0071] The RGB image is input into the SGHM lightweight human instance segmentation model, which then extracts features and outputs a binary mask of the human body through an instance segmentation head. ,in, Represents pixels For the dynamic area of ​​the human body , Represents pixels Static scene area The model achieves an inference frame rate of ≥30 FPS and an inference latency of <33 ms on an NVIDIA GTX 4070 graphics card, which is consistent with the camera acquisition frame rate and meets the real-time requirements. This step clearly divides the input frame into dynamic and static parts, laying the foundation for subsequent separation modeling.

[0072] S3: SMPL-X model alignment optimization;

[0073] For the human dynamic region segmented in step S2 This paper utilizes the SMPL-X human parametric model for pose, shape, and facial expression estimation, and achieves precise alignment between the model and RGB frames through multi-constraint optimization, thus solving the problem of model-image misalignment in existing technologies. The specific implementation is as follows:

[0074] S31: Parameter initialization: Using VIBE to initialize the human body dynamic region Preliminary processing was performed to estimate the initial shape parameters of the SMPL-X model. Attitude parameters (Including full-body joint rotation) Facial expression parameters Output 2D key points of the human body. 3D body joints With 3D hand joints ;

[0075] S32: To achieve precise alignment, a multi-constraint loss function is constructed, expressed as:

[0076] ;

[0077] In this embodiment, This represents the 2D keypoint constraint loss, utilizing the detected 2D keypoints of the human body. To ensure supervision, the keypoints projected onto the SMPL-X model are consistent with the detected keypoints. The Geman-McClure robust error function is used to suppress the interference of noisy keypoints, expressed as:

[0078] ;

[0079] in, Let K be the projection function under the camera's intrinsic parameters. This is a joint rotation function driven by attitude parameters. , These are predefined weights and detection confidence weights, respectively.

[0080] In this embodiment, To represent the 3D body prior loss, a variational human pose prior is introduced to filter out infeasible body poses, while the estimated 3D body joints are used as the basis for further analysis. To guide and optimize the low-dimensional embedding of the pose, the expression is:

[0081] ;

[0082] in, This is the low-dimensional embedding vector of the pose. Regularization terms are used to avoid overfitting;

[0083] In this embodiment, For the 3D prior loss of the hand, and considering the fine pose of the hand, the estimated 3D hand joints are used. For supervision, only the z-axis coordinate is constrained to adapt to differences in camera viewpoint; the expression is:

[0084] ;

[0085] in, This is a robust error function that acts only on the z-axis;

[0086] In this embodiment, the loss weight is set as follows: , Extensive experiments have verified that this weighting ratio can balance the alignment accuracy between the body and the hands.

[0087] S33: Three-stage optimization: The L-BFGS optimizer is used to iteratively optimize the above loss function, which is divided into three stages: The first stage (50 iterations) initializes the human body embedding to the initial estimate and optimizes the overall pose and shape; the second stage (30 iterations) focuses on optimizing the spatial relationship of the hand to improve the hand alignment accuracy; the third stage (20 iterations) fine-tunes the fine areas of the hand and face, and finally obtains the SMPL-X parameters that are accurately aligned with the RGB frame. The human body mesh provides a reliable geometric prior for subsequent 3D Gaussian modeling of the human body.

[0088] S4: Dual-branch 3D Gaussian modeling;

[0089] In this embodiment, a static scene Gaussian metaset is constructed. With dynamic human body Gaussian set The dual-branch modeling framework performs 3D Gaussian modeling on static scenes and dynamic human bodies separately to achieve separate representations. The specific implementation is as follows:

[0090] S41: Static Scene Gaussian Metaset Initialization: Following the initialization strategy of static 3DGS SLAM, based on the estimated camera pose and depth images, the core parameters of the static scene Gaussian units are initialized in the world coordinate system: spatial mean. (Corresponding spatial point coordinates in the depth image), covariance matrix (Initialized to a scaled form of the identity matrix), Opacity (Initialized to 0.5), Color (Extract the corresponding pixel color from the RGB image) to ensure compatibility with the original static SLAM algorithm;

[0091] S42: Dynamic Human Body Gaussian Unit Set Initialization: Centered on the vertices of the SMPL-X human body mesh aligned in step S3, initialize the core parameters of the dynamic human body Gaussian elements: spatial mean. (Coordinates of human body mesh vertices in normalized space), covariance matrix (Initialized to a scaled form of the identity matrix, with a scaling factor of 0.01), Opacity (Initialized to 0.8), Color (Extract the corresponding pixel color from the human body region of the RGB image); At the same time, assign body part attribute U (body, hand, face) to each Gaussian pixel to provide a basis for subsequent adaptive density control;

[0092] S43: Adaptive Density Control of Gaussian Units: To address the granularity differences in different parts of the human body (hands and face require finer modeling, while body parts can use coarser modeling), a context-aware adaptive density control strategy is constructed to dynamically adjust the density of Gaussian units in the human body. The specific implementation is as follows:

[0093] Compacting threshold calculation: Combining the body part attribute U of Gaussian elements with historical gradient information, the compaction threshold of each Gaussian element is dynamically calculated. The expression is:

[0094] ;

[0095] in, Hyperparameters related to the location attribute U;

[0096] Specifically, the body: , Hands: , Face: , Smaller hand and face settings With larger To achieve more refined densification; R=50 is the densification time window; The gradient represents the position of the center of the i-th Gaussian unit at frame k. A continuously increasing gradient indicates that the region needs more Gaussian units to fit the fine structure.

[0097] Compactification operation: When the position gradient of the Gaussian unit exceeds the adaptive threshold When performing a splitting or cloning operation, the Gaussian element is split into two sub-Gaussian elements along the direction of maximum gradient, and the covariance matrix of the sub-Gaussian elements is half of the original matrix; the cloning operation replicates the Gaussian element and fine-tunes its position to increase the local Gaussian element density.

[0098] Trimming operation: Based on the SMPL-X human body prior, calculate the shortest distance between the center of the Gaussian element and the surface of the human body mesh. If the distance is greater than... If a value is found to be invalid, it is considered an invalid Gaussian element and is pruned to avoid redundant calculations.

[0099] This strategy ensures the accuracy of reconstruction of fine details such as hands and face while controlling the overall number of Gaussian elements, thus balancing reconstruction quality and real-time performance.

[0100] S5: Regional confidence perception optimization;

[0101] To address the distortion of supervision signals caused by human occlusion and motion blur in dynamic scenes, a confidence-aware region-specific loss function is constructed for static scene regions. With human dynamic area The optimization loss is calculated separately to achieve joint optimization of the bi-branch Gaussian elements and camera pose, as detailed below:

[0102] S51: Definition of Total Loss Function: The total optimization loss is the weighted sum of the static scene loss and the human body region loss, expressed as: ,in The loss weight is used to balance the optimization contributions of the two.

[0103] S52: Static Scene Loss : Adopting the loss function form of the static 3DGS SLAM algorithm, only in static scene areas Internal calculations, including photometric residuals, geometric residuals, and isotropic regularization terms, while retaining visibility filtering, are expressed as follows:

[0104] ;

[0105] in: Pixel color and depth rendered in Gaussian style for static scenes; For the true colors and depth captured by the camera; For visibility weights in static scenes; , These are the scaling parameters and average scaling value of the Gaussian element, respectively; These are regularization weights used to suppress excessive stretching of Gaussian elements.

[0106] S53: Loss of human body area This approach introduces confidence-perceived loss as its core, combining masking loss, structural similarity loss (SSIM), and perceptual loss (LPIPS) in the human body region. The optimization of the human body's Gaussian elements is expressed as follows:

[0107] ;

[0108] in, , , For loss weights;

[0109] Confidence-perceived loss Colors rendered using human body Gaussian primitives With depth As input, a lightweight CNN network is used to predict the confidence level C(p) of each pixel (the lower the confidence level, the greater the interference from occlusion and motion blur). The expression is:

[0110] ;

[0111] ;

[0112] in, It is a constant. This is a confidence prediction network (containing 3 convolutional layers and 2 fully connected layers). It is a pixel-level dot product;

[0113] Mask loss The cumulative opacity of the constrained human body Gaussian primitive rendering is consistent with the human body segmentation mask, and the expression is:

[0114] ;

[0115] in The cumulative opacity of the human body's Gaussian elements;

[0116] Structural similarity loss To constrain the consistency between the rendered image and the real image at the structural similarity level, the expression is:

[0117] ;

[0118] in The currently rendered image. For real images In the region The corresponding part above;

[0119] Perceived loss To constrain the quality of rendered images from the perspective of human visual perception, a pre-trained LPIPS network is used for computation.

[0120] S6: Tracking and Mapping: The original static SLAM process is adapted by alternating between tracking and mapping to achieve camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes. The specific process is as follows:

[0121] S61: Tracking phase (only pose optimization, no Homo sapiens updates), specifically including:

[0122] S611: Pose Initialization: Initialize the camera pose of the current frame using a constant velocity motion model. The expression is: ,in, These are the camera poses at frame t and frame (t-1), respectively.

[0123] S612: Regional Rendering: Rendering Gaussian sets of static scenes With dynamic human body Gaussian set Render each area separately to obtain a static area color map. Depth map Color map of human body regions Depth map ;

[0124] S613: Pose Optimization: Calculating the Total Loss Function The camera pose was optimized through 20 iterations using the Adam optimizer. With human posture parameters until the losses are mitigated;

[0125] S614: Result saving: The optimized camera pose is retained as a priori for the mapping stage, and the human pose parameters are used for the SMPL-X alignment initialization of the next frame.

[0126] S62: Mapping phase (fixed pose, updated Heteroskeleton elements):

[0127] S621: Static Scene Gaussian Element Optimization: With fixed camera pose, optimize the Gaussian elements for static scenes. Perform 2D scale-adaptive filtering, anti-aliasing rendering, and isotropic regularization optimization to update the mean of the Gaussian elements. Covariance ,color With Opacity ;

[0128] S622: Dynamic Human Gossky Element Optimization: Optimization of Human Gossky Elements Perform adaptive density control (densification and pruning) while using a confidence-aware loss function. Optimize the mean of Gaussian elements Covariance ,color With Opacity The optimization was completed after 15 iterations.

[0129] S623: Keyframe Selection and Map Update: If the current frame meets the keyframe selection criteria (IOU < 0.5 and OC < 0.6, or the ratio of camera relative displacement to median depth > 0.1), the optimized bi-branch Gaussian set is updated. Add to the global map and complete the map update.

[0130] In this embodiment, the hardware environment and parameter settings are as follows:

[0131] Running in the following hardware environment: Ubuntu 22.04 operating system, Intel i5 12400 CPU, NVIDIA GeForce GTX 4070 GPU. Memory; core parameter settings are as follows:

[0132] Tracking phase: Rendering resolution 640×480, optimization iterations 20 times per frame;

[0133] Map building phase: Static scene rendering resolution is 320×240, human body area rendering resolution is 640×480, and each map building optimization iteration is 15 times.

[0134] Gaussian optimization: 2000 total iterations, 400-1000 densification iterations;

[0135] Keyframe thresholds: IOU=0.5, OC=0.6, and the ratio of camera relative displacement to median depth threshold is 0.1.

[0136] This embodiment adopts a dual-branch 3D Gaussian modeling framework and regional confidence perception optimization, separating the optimization objectives of dynamic and static regions, effectively eliminating the interference of dynamic human bodies on camera pose estimation and static mapping, and has strong robustness to dynamic scenes and high accuracy of camera pose estimation.

[0137] The implementation method of this embodiment achieves high accuracy in dynamic human reconstruction, especially in the restoration of fine structures. Based on the SMPL-X model, multi-constraint alignment optimization and context-aware adaptive density control strategy provide accurate human geometric priors and realize adaptive adjustment of Gaussian pixel density for different parts of the human body. This results in a PSNR of 24.8-26.3dB for human region reconstruction, clearly restoring fine structures such as hands and face. The method employs a lightweight human instance segmentation model (SGHM), a region-based rendering strategy (low resolution for static regions and high resolution for human regions), and CUDA-accelerated Gaussian pixel density control. On an NVIDIA GTX 4070 graphics card, the frame rate is stable at over 1.2 FPS, significantly better than traditional dynamic 3DGS SLAM algorithms. It also has good real-time performance and meets the needs of engineering applications.

[0138] In this embodiment, other lightweight human instance segmentation models can be used to replace the lightweight human instance segmentation model SGHM, such as YOLO-Pose combined with a semantic segmentation head, MobileSeg, etc. As long as real-time segmentation of the human region can be achieved (frame rate ≥ 30 FPS and output of binary mask), the same region segmentation purpose can be achieved. In addition, multimodal segmentation methods (fusion of RGB and depth images) can also be used to improve segmentation robustness, which is suitable for low-light and severely occluded scenes.

[0139] In this embodiment, other human parametric models can be used to replace SMPL-X, such as a combination of SMPL, MANO (hands only) and FLAME (face only). As long as the human body shape, posture and expression can be represented by low-dimensional parameters and the human body mesh vertex information can be provided, it can be used as the geometric prior for dynamic human body modeling. The corresponding multi-constraint loss function can be adapted and adjusted according to the key point type of the model output (such as only body joints, no facial key points), for example, removing the constraints related to facial expression parameters.

[0140] In this embodiment, the densification threshold ϵi can be calculated using other features instead of historical gradient information, such as the color gradient of Gaussian elements and the rate of change of opacity. As long as it can reflect the fine structure requirements of the region, adaptive density adjustment can be achieved. The hyperparameters eU, λt, and U can also be dynamically calibrated through methods such as grid search and Bayesian optimization, rather than fixed values, to adapt to different scenarios.

[0141] In this embodiment, the photometric residual in the static scene loss Ls can be replaced by the L2 norm instead of the L1 norm, and the isotropic regularization term can be replaced by the L2 norm instead of the L1 norm; the confidence prediction network in the human body region loss can be replaced by a lightweight Transformer model instead of a CNN. As long as pixel-level confidence estimation can be achieved and interference signals can be suppressed, the same optimization effect can be achieved; the loss weights (such as λh, λm) can be replaced by fixed values ​​through adaptive adjustment strategies (such as real-time updates based on the scene's dynamics) to improve the algorithm's scene adaptability.

[0142] In this embodiment, pose initialization during the tracking phase can use IMU pre-integration results instead of a constant velocity motion model to improve initialization accuracy in fast motion scenarios; key frame selection conditions during the mapping phase can introduce more indicators (such as Gaussian sigma update amount and photometric loss value) to further optimize key frame quality; in addition, a parallel tracking-mapping architecture can be used to replace the alternating execution process, and the overall frame rate of the system can be improved through multi-threaded parallel processing.

[0143] Example 2

[0144] This embodiment provides a 3D Gaussian sputtering SLAM system for dynamic scenes, used to implement the 3D Gaussian sputtering SLAM method for dynamic scenes in Embodiment 1 above, including: a data acquisition module, a dynamic region segmentation module, an alignment module, a bi-branch modeling module, a confidence-aware optimization module, and a tracking mapping module;

[0145] In this embodiment, the data acquisition module is used to acquire RGB images, depth images, and IMU data;

[0146] In this embodiment, the dynamic region segmentation module is used to segment the RGB image into a dynamic human body region and a static scene region;

[0147] In this embodiment, the alignment module is used to estimate the pose, shape and facial expression of the human body dynamic region based on the human body parameterization model, and to achieve the alignment of the model with the RGB frame through multi-constraint optimization.

[0148] In this embodiment, the dual-branch modeling module is used to construct a static scene Gaussian set and a dynamic human body Gaussian set, and to perform 3D Gaussian modeling on the static scene region and the dynamic human body region.

[0149] In this embodiment, the confidence perception optimization module is used to construct a region-specific loss function for confidence perception, calculate the optimization loss for the dynamic human body region and the static scene region respectively, and realize the joint optimization of the dual-branch Gaussian unit and the camera pose.

[0150] In this embodiment, the tracking and mapping module is used to adapt the original static SLAM process based on the process of alternating tracking and mapping, so as to realize camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes.

[0151] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A 3D Gaussian sputtering SLAM method for dynamic scenes, characterized in that, Includes the following steps: Acquire RGB images, depth images, and IMU data; Segment the RGB image into dynamic human body regions and static scene regions; Based on the human body parameterization model, the pose, shape and facial expression of the human body dynamic region are estimated, and the alignment of the model with the RGB frame is achieved through multi-constraint optimization. Construct a static scene Gaussian set and a dynamic human body Gaussian set, and perform 3D Gaussian modeling on the static scene region and the dynamic human body region. A confidence-aware regional loss function is constructed, and the optimization loss is calculated separately for the dynamic human body region and the static scene region, realizing the joint optimization of the bi-branch Gaussian unit and the camera pose. Based on the alternating execution of tracking and mapping, the original static SLAM process is adapted to achieve camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes.

2. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 1, characterized in that, The RGB image is segmented into dynamic human body regions and static scene regions, specifically including: Human instance segmentation is performed based on the SGHM lightweight human instance segmentation model, and a binary human mask is output. ,in, Represents pixels This is the dynamic area of ​​the human body. Represents pixels This is a static scene area.

3. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 1, characterized in that, Based on a human parametric model, pose, shape, and facial expression estimation are performed on dynamic regions of the human body. Alignment between the model and RGB frames is achieved through multi-constraint optimization, specifically including: Construct initial shape parameters, pose parameters, and facial expression parameters, and output 2D human body key points, 3D body joints, and 3D hand joints; Construct a multi-constraint loss function, expressed as: ; in, Indicates attitude parameters, Indicates shape parameters, Indicates facial expression parameters. This represents the 2D keypoint constraint loss. This indicates a 3D prior loss of the body. For the 3D prior loss of the hand. , These represent the loss weights; Iterative optimization based on multi-constraint loss function is used to align the model with the RGB frame.

4. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 3, characterized in that, 2D Keypoint Constraint Loss Represented as: ; in, Let K be the projection function under the camera's intrinsic parameters. This is a joint rotation function driven by attitude parameters. , These are predefined weights and detection confidence weights, respectively. Representing key points of the human body in 2D; 3D body prior loss Represented as: ; in, This is the low-dimensional embedding vector of the pose. For regularization terms; 3D prior loss of hand Represented as: ; in, This represents 3D hand joints. Let be the robust error function acting on the z-axis.

5. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 1, characterized in that, Constructing static scene Gaussian set and dynamic human body Gaussian set, specifically including: Based on the estimated camera pose and depth image, initialize the parameters of the static scene Gaussian unit in the world coordinate system, including the spatial mean, covariance matrix, opacity, and pixel color in the depth image. The dynamic human body Gaussian tuple includes the human body mesh vertex coordinates, covariance matrix, opacity, and pixel color of the human body dynamic region, while assigning body part attributes to each Gaussian tuple.

6. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 5, characterized in that, Perform 3D Gaussian modeling on static scene areas and dynamic human body areas, specifically including: Construct a context-aware adaptive density control strategy to adjust the density of human high-sigma units, including: Based on the body part attributes and historical gradient information of Gaussian elements, the compaction threshold of each Gaussian element is calculated. When the position gradient of a Gaussian unit exceeds the adaptive threshold, it is split or cloned. Calculate the shortest distance between the center of the Gaussian cell and the surface of the human body mesh. If the shortest distance is greater than a set threshold, the Gaussian cell is determined to be invalid and is pruned.

7. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 6, characterized in that, The compaction threshold for each Gaussian cell is calculated and expressed as: ; in, Densification threshold, where R represents the densification time window. For hyperparameters related to the part attribute U, It represents the position gradient of the center of the i-th Gaussian element in the k-th frame.

8. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 1, characterized in that, Construct a confidence-aware regional loss function, expressed as: ; ; ; in, This represents a static scene area. Indicates the dynamic area of ​​the human body. Indicates the loss in a static scene. Indicates loss in a specific area of ​​the human body. Pixel color and depth rendered using Gaussian primitives for static scenes. For the true colors and depth captured by the camera, For visibility weights in static scenes, , These represent the scaling parameter and average scaling value of the Gaussian element, respectively. For regularization weights, , , These represent the corresponding weights. This represents the perceived loss based on confidence level. Indicates mask loss. Represents structural similarity loss. This indicates perceived loss.

9. The 3D Gaussian sputtering SLAM method for dynamic scenes according to claim 1, characterized in that, Based on the alternating execution of tracking and mapping, the original static SLAM process is adapted to achieve camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes, specifically including: During the tracking phase, a constant velocity motion model is used to initialize the camera pose of the current frame. The static scene Gaussian set and the dynamic human body Gaussian set are rendered separately to obtain the static region color map and depth map and the human body region color map and depth map. The camera pose and human body pose parameters are iteratively optimized based on the confidence-aware regional loss function. During the mapping phase, the camera pose is fixed, and 2D scale adaptive filtering, anti-aliasing rendering and isotropic regularization optimization are performed on the static scene Gaussian primitives to update the static scene Gaussian primitives. Adaptive density control is performed on the dynamic human body Gaussian primitives, and the dynamic human body Gaussian primitives are optimized through the confidence-perceived loss function. The optimized bi-branch Gaussian set is added to the global map to complete the map update.

10. A 3D Gaussian sputtering SLAM system for dynamic scenes, characterized in that, The method for implementing the 3D Gaussian sputtering SLAM method for dynamic scenes as described in any one of claims 1-9 includes: a data acquisition module, a dynamic region segmentation module, an alignment module, a bi-branch modeling module, a confidence-aware optimization module, and a tracking mapping module. The data acquisition module is used to acquire RGB images, depth images, and IMU data; The dynamic region segmentation module is used to segment the RGB image into a dynamic human body region and a static scene region. The alignment module is used to estimate the posture, shape and facial expression of the human body dynamic region based on the human body parameterization model, and to achieve the alignment of the model with the RGB frame through multi-constraint optimization. The dual-branch modeling module is used to construct a static scene Gaussian set and a dynamic human body Gaussian set, and to perform 3D Gaussian modeling on the static scene region and the dynamic human body region. The confidence-aware optimization module is used to construct a region-specific loss function for confidence-awareness, calculate the optimization loss for the dynamic human body region and the static scene region respectively, and realize the joint optimization of the dual-branch Gaussian unit and the camera pose. The tracking and mapping module is used to adapt the original static SLAM process based on the alternating execution of tracking and mapping, so as to realize camera pose estimation and dual-branch 3D Gaussian map update in dynamic scenes.