Dynamic scene four-dimensional gauss simultaneous localization and mapping method based on space-time weight field
By constructing a sliding window and factor graph in SLAM, and combining reliability weights and appearance gating factors, observation constraints are adaptively adjusted, solving the problem of unstable optimization in dynamic scenes for Gaussian SLAM methods, and achieving stable pose and depth solutions and high-quality map reconstruction.
Patent Information
- Application Number
- CN202610758555.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-25
AI Technical Summary
Existing Gaussian SLAM methods cannot adaptively modulate the observation constraint weights in scenarios involving dynamic occlusion, non-rigid motion, and coupled illumination changes, leading to instability in the optimization process and difficulty in achieving convergence.
The dynamic scene four-dimensional Gaussian synchronous localization and mapping method based on spatiotemporal weighted field constructs a factor map and performs dense bundle adjustment by setting a sliding window on the keyframe set in the continuous RGB video stream, and adaptively adjusting the observation constraints by combining reliability weight and appearance gating factor to achieve stable optimization convergence.
It effectively suppresses interference caused by dynamic occlusion, non-rigid body motion and lighting changes, improves the convergence stability and accuracy of pose and depth solutions in dynamic scenes, and achieves high-quality reconstruction of the global scene map.
Smart Images

Figure CN122636722A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of synchronous positioning and mapping technology, and in particular to a dynamic scene four-dimensional Gaussian synchronous positioning and mapping method based on spatiotemporal weighted fields. Background Technology
[0002] In recent years, 3DGS (3D Gaussian Splatting), as an explicit scene representation method, has demonstrated high efficiency in novel perspective synthesis and real-time rendering, and has therefore been gradually introduced into the field of SLAM (Simultaneous Localization and Mapping). Compared with traditional SLAM methods that only focus on camera trajectory solving and sparse geometry reconstruction, the 3DGS-based SLAM framework can construct dense, renderable scene models with realistic textures and fine structures while completing camera localization, providing richer and more complete environmental representation capabilities for downstream tasks such as autonomous driving, robot perception, and embodied intelligence.
[0003] Traditional visual SLAM is largely based on the static world assumption, using cross-frame geometric consistency to estimate camera pose and optimize the map. These methods, centered on sparse features and keyframe map optimization, are stable in static scenes. However, in dynamic environments with dynamic occlusion, non-rigid motion, and varying illumination, cross-frame consistency is disrupted, causing pixel correspondences and residual statistics to become unstable, leading to slower convergence and even local degradation in the backend optimization. For dynamic environments, early methods typically introduced semantic or geometric consistency priors to identify dynamic regions and filter unreliable observations, improving mapping robustness by removing dynamic components and preserving static structures. While such filtering strategies can achieve some results, they are prone to over-reducing effective constraints and causing gradient discontinuities and structural breaks in occluded boundary regions. With the development of learning-based methods, dense differentiable SLAM and neural implicit SLAM have achieved high-precision dense reconstruction with stronger scene modeling capabilities. However, in real-world scenes with dynamic interference and fluctuating illumination, they still cannot effectively overcome the optimization instability caused by observation inconsistency.
[0004] Building upon this foundation, 3DGS, with its advantages of explicit structure, efficient incremental updates, and high-quality appearance modeling, is widely used in dynamic SLAM tasks. Static 3DGS-SLAM methods use rendering consistency as the core supervisory signal, incorporating camera pose and Gaussian parameters into a unified optimization loop, enabling the construction of high-precision, renderable, dense maps while achieving stable localization. However, real-world scenes commonly feature dynamic target movement, frequent occlusion switching, and complex lighting perturbations. Static 3DGS-SLAM lacks an adaptive constraint mechanism for unreliable dynamic observations. Dynamic errors propagate, accumulate, and amplify along the highly coupled optimization path of pose and map parameters, ultimately leading to camera pose drift, scene structure blurring, and map degradation.
[0005] To suppress dynamic interference, existing dynamic 3DGS-SLAM methods often introduce external dynamic cues to mitigate the impact of unreliable observations. They identify dynamic regions through motion detection, temporal error discrimination, semantic masking, and high-level uncertainty representation, and deweight or remove dynamic pixels during the tracking phase. While these methods can alleviate the interference of dynamic targets on localization and mapping to some extent, the dynamic cues are all external, fixed priors and cannot be consistently updated with scene geometric changes, occlusion evolution, and lighting fluctuations. In scenes with strong occlusion, rapid visibility switching, and complex lighting, issues of constraint instability and continuous error accumulation still exist. In contrast, 4DGS and its derivative 4DGS-SLAM methods explicitly characterize scene dynamic changes by constructing spatiotemporal Gaussian representations, separating static and dynamic Gaussian elements, and modeling temporal deformations, effectively improving the consistency of dynamic temporal reconstruction and rendering. However, this type of method still has obvious defects in the dynamic target boundary region. The problems of occlusion and mixed pixels, inter-frame boundary misalignment and structural mismatch are difficult to be effectively constrained, which can easily cause target outline blurring and boundary artifact accumulation, severely limiting the detail accuracy and spatiotemporal stability of dynamic scene reconstruction.
[0006] In summary, existing Gaussian SLAM methods primarily improve the robustness of the system in dynamic scenes by identifying and removing or weakening unreliable observations, and suppress dynamic interference through methods such as [list of methods]. They exhibit a certain degree of robustness in simple dynamic scenes. However, in complex real-world scenes with multiple interferences coupled by dynamic occlusion, non-rigid motion, and lighting changes, existing technologies generally have certain shortcomings. Dynamic observation discrimination and weights cannot adaptively evolve with the scene state, making it difficult to achieve continuous and smooth constraint modulation. The optimization process is also susceptible to outlier observations, leading to gradient oscillations and pose drift. Summary of the Invention
[0007] Therefore, the technical problem to be solved by the present invention is to overcome the shortcomings of existing Gaussian SLAM methods in scenarios with coupled dynamic occlusion, non-rigid motion and illumination changes, which are unable to adaptively modulate the observation constraint weights and are difficult to stably optimize convergence.
[0008] To address the aforementioned technical problems, this invention provides a method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on a spatiotemporal weighted field, comprising: Set a sliding window for a set of keyframes in a continuous RGB video stream; Each sliding window is processed sequentially, and each keyframe image within the current sliding window is converted into a low-resolution optimized mesh, and a factor map of the current sliding window is constructed. In dense bundle adjustment, based on the reliability weights and appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window, the network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are corrected to construct a weighted geometric consistency objective function. The camera pose and depth of each keyframe within the current sliding window are obtained by minimizing the weighted geometric consistency objective function. Based on the camera pose and depth of each keyframe within all sliding windows, the final camera pose and depth of each keyframe are determined, and a global scene map is generated.
[0009] Preferably, the method for constructing the factor graph corresponding to each sliding window includes: For each sliding window, each keyframe within the sliding window is treated as a node, and the node state includes camera pose and depth information; bidirectional adjacency edges are constructed between temporally adjacent keyframes; based on the bidirectional inter-frame distance of any keyframe pair, target keyframe pairs are selected, cross-frame constraint edges are constructed, and the factor graph corresponding to each sliding window is obtained.
[0010] Preferably, the method for obtaining the reliability weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window includes: Based on the real-valued scores, preset maximum weight values, and preset minimum weight values of each low-resolution optimized grid point on the keyframes corresponding to each edge of the factor graph in the current sliding window, the initial reliability weights of each low-resolution optimized grid point on the keyframes corresponding to each edge of the factor graph in the current sliding window are constructed. Based on the initial reliability weights of each low-resolution optimized grid point in the factor graph of the current sliding window corresponding to two key frames, the minimum value is selected as the edge-level reliability gate of the low-resolution optimized grid point in the factor graph of the current sliding window. Based on the product of the edge-level reliability gate and the network prediction confidence weight, the reliability weights of each low-resolution optimized grid point in the factor graph of the current sliding window are obtained.
[0011] Preferably, the method for correcting the network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window based on the reliability weights and appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window includes: The product of the reliability weight of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window and the appearance gating factor is used as the final weight of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window. The network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are corrected by using the final weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window.
[0012] Preferably, the method for solving the problem by minimizing the weighted geometric consistency objective function includes: Step S41: Initialize the camera pose and depth of all keyframes in the current sliding window, and the real-valued scores of each low-resolution optimized mesh point on the two keyframes of each edge of the factor graph in the current sliding window. Set the number of iterations. ; Step S42: Based on the first In the next iteration, the final weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are used to perform a second-order update on the objective function of minimizing the weighted geometric consistency, resulting in the... The next iteration measures the camera pose and depth of each keyframe within the current sliding window; S43: Based on the first In the next iteration, the camera pose and depth of each keyframe within the current sliding window are updated, along with the edge-level reliability gating, normalized appearance residuals, and geometric residuals of each low-resolution optimized mesh point in the factor graph of the current sliding window. S44: Fusion In the next iteration, the normalized appearance residuals and geometric residuals of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are obtained, yielding the... The fusion residuals of each low-resolution optimized grid point in the next iteration on each edge of the factor graph in the current sliding window; S45: Based on the first In the next iteration, the edge-level reliability gating and square norm of the fused residuals of each low-resolution optimized grid point in the factor graph of the current sliding window are used to obtain the th iteration. The next iteration updates the data items; S46: Based on the first The spatiotemporal regularization term and the updated data term in the next iteration, the th The real-valued scores of each low-resolution optimized grid point in the next iteration are obtained. In the next iteration, the real-valued scores of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window on two keyframes; S47: Determine if the iterative convergence condition is met; if not, update. Return to step S42. If the condition is met, output the first... The next iteration calculates the camera pose and depth for each keyframe within the current sliding window.
[0013] Preferably, the method for obtaining the appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window includes: Based on the normalized appearance residuals of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window, and the gray-scale mean brightness of each edge of the factor graph of the current sliding window corresponding to two keyframes, the appearance gating factor of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window is constructed, and the formula is as follows: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. On the appearance gating factor, Optimize grid point coordinates for low resolution. This is the factor graph edge index of the current sliding window. The factor graph edges of the current sliding window The reference frame in the text is the edge. The starting frame, The factor graph edges of the current sliding window The target frame in the image is the edge. The terminating frame, Optimize grid point positions for low resolution. This is a truncation function. It is an exponential function with the natural constant as its base. The factor graph edges of the current sliding window Exposure adaptive coefficient, , For keyframes grayscale mean brightness, For keyframes grayscale mean brightness, For numerically stable terms, To control the intensity of exposure attenuation, Optimize grid points for low resolution On the factor graph edge of the current sliding window Normalized appearance residuals , This is the lower bound of the appearance gating factor.
[0014] Preferably, the process of obtaining the appearance residuals of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window includes: Use the starting frame of each edge of the factor graph of the current sliding window as the reference frame and the ending frame as the target frame. For each low-resolution optimized grid point on each edge, the difference between the gray intensity of the low-resolution optimized grid point in the reference frame of each edge and the gray intensity of the projection point of the low-resolution optimized grid point in the reference frame in the target frame is used as the appearance residual of the low-resolution optimized grid point on each edge.
[0015] Preferably, the method for determining the final camera pose and depth of each keyframe based on the camera pose and depth of each keyframe within all sliding windows includes: Based on the camera pose and depth of each keyframe within each sliding window, the final camera pose and depth of each keyframe are obtained through local BA, global BA, and loop closure detection.
[0016] Preferably, the method for generating a global scene map includes: Based on the motion mask of each keyframe, each keyframe pixel is divided into static and dynamic regions; Based on the final camera pose and depth of each keyframe, the effective pixels in the static and dynamic regions are back-projected into three-dimensional space to obtain the static Gaussian set and the dynamic Gaussian set. Based on the scaling factor and the distance from each pixel in the motion mask boundary neighborhood of each keyframe to the motion mask boundary, the soft weight of each pixel in the motion mask boundary neighborhood of each keyframe is determined, and boundary time consistency regularization constraints are set for it. The soft weight of each pixel in the neighborhood of the motion mask boundary of each keyframe is used as the weight of the photometric consistency loss to obtain the target photometric consistency loss. Keyframes within a local window are selected as observation data. A total loss function is constructed, consisting of target photometric consistency loss, geometric consistency loss, scale regularization loss, and boundary structure consistency regularization. The parameters of the static Gaussian and dynamic Gaussian are iteratively updated by minimizing the total loss function to obtain the global scene map.
[0017] Preferably, the formula for determining the soft weight of each pixel in the motion mask boundary neighborhood of each keyframe, based on the scaling factor and the distance from each pixel in the motion mask boundary neighborhood to the motion mask boundary, is as follows: , , in, For the first Pixels in the motion mask boundary neighborhood of each keyframe Soft weights For pixels within the neighborhood of the motion mask boundary, This represents the scaling factor for the soft weights. Indicates the maximum expansion radius of the boundary. For the boundary zone radius, Represents pixels To the dynamic mask boundary The minimum radius of expansion, For the boundary neighborhood of the motion mask, For expansion operation, For dynamic mask boundaries.
[0018] Compared with the prior art, the above-described technical solution of the present invention has the following advantages:
[0019] This invention presents a dynamic scene four-dimensional Gaussian synchronous localization and mapping method based on a spatiotemporal weighted field. It constructs a sliding window for continuous video keyframes, builds a factor graph based on keyframe nodes, temporal sequence, and cross-frame constraint edges within the window, and completes dense constraint modeling using a low-resolution optimized mesh, providing a stable optimization framework for subsequent adaptive weight modulation. Based on this, initial reliability weights are generated by combining real-valued scores of mesh points and weight boundary constraints. A boundary-level reliability gating is constructed through bidirectional keyframe weight filtering to correct the network prediction confidence weights and adaptively suppress invalid observation constraints caused by dynamic occlusion and non-rigid motion. Simultaneously, an exposure-adaptive appearance gating factor is introduced, dynamically adjusting residual weights based on inter-frame brightness differences and normalized appearance residuals, effectively offsetting observation biases caused by illumination changes. Through iterative optimization, the camera state, residual information, and reliability weights within the window are continuously updated, enabling observation constraint weights to dynamically change with the scene and self-consistently update with illumination fluctuations. This overcomes the shortcomings of traditional fixed weight constraints, which cannot adapt to complex dynamically coupled scenes, effectively avoids interference from dynamic outlier observations on the optimization process, and significantly improves the convergence stability and estimation accuracy of pose and depth solutions in dynamic scenes.
[0020] Furthermore, to address the lack of refined constraint mechanisms for dynamic target boundaries, which leads to structural misalignment, blurring, and artifact accumulation in boundary regions, making it difficult to simultaneously achieve stable positioning and high-quality boundary reconstruction, a refined global Gaussian map optimization and reconstruction is conducted based on accurately acquiring keyframe camera pose and depth. Abandoning the traditional binary mask hard segmentation constraint method, a pixel-by-pixel soft weight is calculated by combining mask boundary distance and attenuation coefficient. This provides continuous and smooth constraint modulation capabilities for mask boundary transition regions, effectively solving the problems of abrupt boundary gradient changes and constraint discontinuities caused by hard threshold segmentation. Pixel-level soft weights are adaptively weighted into photometric consistency loss, precisely weakening the invalid residual interference of mixed pixels in dynamic boundaries while preserving effective scene texture constraints. Simultaneously, a multi-dimensional joint loss function is constructed by fusing target photometric loss, geometric loss, scale regularization, and boundary structure consistency regularization, and the global static and dynamic Gaussian parameters are iteratively optimized based on local window keyframe observation information. This approach can constrain the Gaussian update process from multiple dimensions, including pixel luminosity, 3D geometry, Gaussian scale, boundary structure, and temporal stability. It can accurately correct the structural misalignment problem of dynamic target boundaries, suppress boundary blurring and artifact accumulation, and effectively make up for the shortcomings of existing methods in constraining dynamic boundary regions. Ultimately, it can achieve global scene map reconstruction with complete structure, clear boundaries, and temporal stability in dynamic scenes. Attached Figure Description
[0021] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein:
[0022] Figure 1 This is a flowchart illustrating a dynamic scene four-dimensional Gaussian synchronous localization and mapping method based on a spatiotemporal weighted field according to the present invention.
[0023] Figure 2 This is a structural diagram of a dynamic scene four-dimensional Gaussian synchronous localization and mapping method based on a spatiotemporal weighted field according to the present invention.
[0024] Figure 3 These are comparison images of the rendering effects of different methods. Figure 3 (a) in the figure is a comparison chart of the rendering effects of different methods on the TUM / fr3_wk_xyz data. Figure 3 (b) in the figure is a comparison chart of the rendering effects of different methods on the TUM / fr3_wk_st data. Figure 3 (c) in the figure is a comparison of the rendering effects of different methods on the Bonn / balloon dataset.
[0025] Figure 4 This is a diagram showing the results of a BARC ablation experiment. Figure 4 (a) in the figure shows the experimental results using BARC. Figure 4(b) in the figure shows the experimental results without BARC. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0027] In recent years, significant progress has been made in 4DGS-SLAM (4-Dimension Gaussian Splatting Simultaneous Localization and Mapping), which is based on 3DGS (3-Dimension Gaussian Splatting). However, current methods still struggle to handle outliers induced by dynamic occlusion, non-rigid motion, and illumination variations over time, leading to instability in the optimization process and pose estimation drift.
[0028] Therefore, referring to Figure 1 As shown, this embodiment proposes a dynamic scene four-dimensional Gaussian simultaneous localization and mapping method ST³GS-SLAM (Spatio-Temporal Consistency Soft-weighted Field for 4-Dimension Gaussian Splatting SLAM) based on a spatio-temporal weighted field. This method is a 4DGS-SLAM that can effectively adapt to the coexistence of the three dynamic scene changes mentioned above (time-varying dynamic occlusion, non-rigid motion, and illumination changes), including: like Figure 2 As shown, Figure 2 This is a structural diagram of a dynamic scene four-dimensional Gaussian synchronous localization and mapping method based on a spatiotemporal weighted field according to the present invention.
[0029] Step S1: Set a sliding window for the set of keyframes in a continuous RGB video stream; In RGB stream preprocessing, for each input RGB frame... The dense depth prior of each RGB frame is obtained through a depth estimation network. The motion mask for each RGB frame is obtained through the instance segmentation model. , Indicates the RGB frame index.
[0030] During the tracking phase, a set of keyframes is selected. and maintain a fixed-length sliding window. For keyframe selection, this invention selects keyframes based on inter-frame motion metrics and frame interval metrics. The inter-frame motion metrics are derived from displacement field statistics constructed from feature correlations and obtained through network inference; they are used to measure relative motion between frames. If the inter-frame motion metrics exceed a threshold... or the frame interval exceeds If so, then it is selected as a keyframe. In the experiment, Consistent with Splat-SLAM, it is set to 4.0 to control the keyframe density and baseline size, avoiding local graph redundancy or sparse keyframes caused by excessively large or small values. Refer to WildGS-SLAM setting to 9.
[0031] Step S2: Process each sliding window sequentially, convert each keyframe image within the current sliding window into a low-resolution optimized mesh, and construct the factor map of the current sliding window; In this embodiment, specifically, the method for constructing the factor graph corresponding to each sliding window includes: For each sliding window, each keyframe within the sliding window is treated as a node, and the node state includes camera pose and depth information; bidirectional adjacency edges are constructed between temporally adjacent keyframes; based on the bidirectional inter-frame distance of any keyframe pair, target keyframe pairs are selected, cross-frame constraint edges are constructed, and the factor graph corresponding to each sliding window is obtained.
[0032] After keyframe selection, this embodiment constructs a sliding window by selecting the latest keyframe and its 25 neighboring historical keyframes, and uses an incremental strategy to dynamically update the factor graph. Whenever a new candidate keyframe is added, constraints are updated only around the latest keyframe and the historical keyframes within the window, avoiding global repetition and effectively improving optimization efficiency. At the factor graph topology construction level, this invention introduces two types of relationships: temporal constraints and spatial co-visibility constraints. On the one hand, bidirectional adjacency edges are established for all temporally adjacent keyframes to ensure complete temporal associations and continuous local trajectories within the window. On the other hand, the bidirectional inter-frame distance between any keyframe pair is calculated based on the current camera pose and depth information. Co-visibility nearest neighbor keyframe pairs that meet the distance threshold are selected as target constraint pairs, and redundant connections are eliminated using a non-maximum suppression strategy, ultimately generating sparse and highly robust cross-frame constraint edges.
[0033] If a keyframe pair satisfies the temporal adjacency condition or the spatial distance threshold condition, a bidirectional constraint edge is constructed for it; if a keyframe pair satisfies both conditions, only a single bidirectional connection is retained to avoid repeated superposition of constraints.
[0034] In the constructed factor graph, each edge is used to characterize the multi-view geometric consistency between two keyframes: that is, the pixels in the source keyframe are reprojected onto the target keyframe under the current pose and depth estimation, and compared with the corresponding positions predicted by the network. At the same time, confidence weights are used to weight the importance of different pixel constraints. Therefore, the goal of factor graph optimization is to minimize the weighted projection residuals on all edges within the entire sliding window, and to combine photometric consistency information for auxiliary constraints when necessary, thereby jointly optimizing the pose and depth estimation of each keyframe within the window, so that the local trajectory and local map satisfy multi-view geometric consistency as much as possible, ultimately achieving the goal of reducing cumulative drift, improving tracking stability, and improving local mapping accuracy.
[0035] Within the sliding window, cross-frame constraints are constructed between keyframes to form a factor graph. This provides an optimized structure for backend beam adjustment. Based on this factor graph structure, the backend further employs the ST³-BA proposed in this invention, performing continuous soft weighting and adaptive reweighting of cross-frame constraints within a sliding window. Specifically, ST³-BA adjusts the observation reliability online based on the current residual statistics and adaptively adjusts the appearance gating intensity in conjunction with exposure differences to reduce the interference of brightness fluctuations on the appearance residuals. Through the above processing, effective observations can be screened more stably under dynamic occlusion, non-rigid interference, and illumination changes, thereby improving the stability of optimization convergence and the consistency of estimation results.
[0036] Step S3: In Dense Bundle Adjustment (DBA), based on the reliability weights and appearance residuals of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window, the network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are corrected, and a weighted geometric consistency objective function is constructed. ST³-BA addresses the unreliability of observations caused by factors such as occlusion, non-rigid motion, and appearance changes in dynamic scenes. Its goal is to improve convergence stability in dense geometric consistency optimization and maintain consistency of constraints across frames.
[0037] In this embodiment, the basic form and weighted modeling process of dense bundle adjustment are first introduced to clarify the expression of cross-frame constraints in the objective function. Then, a continuous reliability weight field is introduced to robustly weight the geometric constraints, and an auxiliary gating is constructed in conjunction with appearance consistency cues to suppress unreliable observations caused by occlusion, non-rigid perturbations, and illumination changes. Furthermore, residual feedback drives online weight updates, enabling observation reliability to adaptively adjust with the evolution of the geometric state. This invention also incorporates spatiotemporal regularization to maintain the smoothness and stability of the weight field. Finally, geometric variables and reliability variables are updated alternately, unifying the above mechanisms into a single optimization process, thereby obtaining more stable optimization convergence and more consistent estimation results under dynamic disturbance conditions.
[0038] In SLAM, the core task of backend optimization is to organize the constraints between consecutive frames into an iteratively solvable minimization problem, and then jointly optimize camera pose and scene geometric variables to ensure cross-frame geometric consistency. DROID-like methods employ dense bundle adjustment to model the consistency of dense correspondences across frames, and jointly optimize camera pose and geometric variables within a unified weighted least squares framework. Specifically, this method establishes dense cross-frame constraints in the low-resolution domain and introduces spatially adaptive observation weights to adjust the constraint strength at different locations and for different residual components, maintaining a stable second-order solution process under dense constraints. This modeling approach, which uniformly incorporates weights into the objective function, also provides a direct foundation for subsequent explicit modeling and online updates of observation reliability.
[0039] In the keyframe set Construct a set of directed edges above Each edge Corresponding to a set of reference frames Point to target frame Dense cross-frame constraints, whose domain is the set of low-resolution optimized mesh locations. ,in, This represents the coordinates of grid points in the low-resolution optimized grid.
[0040] Record No. The camera pose of each keyframe is Its depth map is defined as opposite side Reference frame To target frame relative pose It is obtained by combining the camera poses of two frames, and the formula is: , in, For the edge Middle target frame Camera pose, For the edge Middle reference frame The camera pose.
[0041] Define reprojection function It uses edges to correspond to the relative pose between two frames and the depth of the reference frame. Low-resolution optimized grid point coordinates The input is the reference frame's optimized grid points, which, after 3D backprojection and cross-frame transformation, are projected onto the target frame's image plane as input. Optimize grid point positions for low resolution. Optimize the set of grid point locations for low resolution.
[0042] To mitigate the linearization error and local mismatch caused by relying solely on the current geometric prediction, an update network is used to predict the two-dimensional correction amount on each edge of the sliding window. Using the linearization point of the current iteration Using the geometric projection at that location as a reference, the corresponding position of the reference projection can be obtained. and its corresponding position after network correction The formula is as follows: , , in, This represents the linearization point of the current iteration. The pose of the linearized point in the current iteration. The depth of the linearization point in the current iteration. This indicates the corrected position given by the network, while Then follow the variable to be optimized The process is updated and entered into a minimization process, thereby maintaining an iterative coupling between network correction and geometric optimization.
[0043] Furthermore, the low-resolution optimized grid points are located on the factor graph edges of the current sliding window. Network prediction confidence weights And satisfying that each component is positive, i.e. Based on this, construct the corresponding information matrix. As shown below: , Therefore, the weighted geometric consistency objective of DBA It can be written as: , in, This represents the set of edges in the sliding window factor graph. Reference frame depth, The weighted norm represents the position corresponding to the projection induced by the current geometric state, and is defined as follows: , in, The input represents the norm to be calculated. Represents a matrix.
[0044] Low-resolution optimized grid points on the factor graph edges of the current sliding window Network prediction confidence weights Through information matrix The objective function is explicitly entered to achieve adaptive weighting of different positions and two-dimensional residual components, which can be continuously adjusted according to changes in observation reliability. The objective, after local linearization, takes the form of a weighted quadratic shape and can be solved by Gauss-Newton iteration. Combined with Schur complement, it achieves efficient elimination of pose variables and dense geometric variables, maintaining controllable computational overhead and stable iterative convergence even under large-scale dense constraints.
[0045] In dynamic scenes, the error distribution of cross-frame geometric constraints is often not stable. Occlusion often introduces significant mismatches near the target boundary, while non-rigid body motion can continuously disrupt the geometric consistency of local areas in a short period. Exposure and illumination changes can also cause brightness shifts, further amplifying the time-varying characteristics of appearance residuals. When these factors overlap, structural anomalies are more likely to occur in dense constraints. If hard thresholding or discrete elimination methods are directly used to handle abnormal observations, the constraint set and effective information may suddenly change during iteration, thereby inducing unstable convergence behavior. Therefore, this invention introduces a continuous reliability weight field within the DBA framework and uses flexible reweighting to adjust the cross-frame constraint strength, so that the impact of unreliable observations can be smoothly decayed during the optimization process. After consistency is gradually restored, the corresponding weights can also recover, thus mitigating the discontinuity problem caused by direct elimination.
[0046] In this embodiment, preferably, the method for obtaining the reliability weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window includes: Based on the real-valued scores, preset maximum weight values, and preset minimum weight values of each low-resolution optimized grid point on the keyframes corresponding to each edge of the factor graph in the current sliding window, the initial reliability weights of each low-resolution optimized grid point on the keyframes corresponding to each edge of the factor graph in the current sliding window are constructed. The formula for calculating the initial reliability weight of each low-resolution optimized grid point in each keyframe is as follows: , in, Coordinates are Low-resolution optimized grid points in keyframes Initial reliability weights on Optimize grid point coordinates for low resolution. It is a truncation function. For the Sigmoid function, , This is the lower limit of the initial reliability weight. This represents the upper limit of the initial reliability weight.
[0047] Through the parameterization described above, reliability no longer appears as a discrete label, but varies within a continuous range. This preserves the ability to suppress anomalous observations while maintaining controllable constraint strength and stable gradient propagation during iteration.
[0048] The reliability of cross-frame constraints should reflect the consistency of observations at both ends. If weighting is based solely on the reliability of one end, constraints may still participate in optimization with a large weight even when occlusion, deformation, or appearance anomalies have occurred at the other end. To mitigate this risk, a conservative fusion strategy is adopted, defining edge-level reliability gating. This definition ensures that the effective reliability of an edge does not exceed any endpoint, so that when occlusion or deformation causes a significant mismatch at one end, the cross-frame constraints of the corresponding region can be consistently suppressed, thereby reducing the interference of outlier observations on the second-order update direction.
[0049] Based on the initial reliability weights of the two keyframes corresponding to each edge of the factor graph of each low-resolution optimized grid point in the current sliding window, the minimum value is selected as the edge-level reliability gating for each edge of the factor graph of that low-resolution optimized grid point in the current sliding window. The formula is as follows: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. Edge-level reliability gating, , Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. Reference Frame Initial reliability weights on Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. target frame The initial reliability weights.
[0050] Based on the product of edge-level reliability gating and network prediction confidence weights, the reliability weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are obtained, as shown in the formula: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. Reliability weight, As a scalar, for The two components act in equal proportions.
[0051] Thus, reliability weights are uniformly incorporated into the weighted least squares objective. When a position is determined to be of low reliability, the contribution of its corresponding residual term is continuously weakened. When subsequent iterations show that geometric consistency has been restored, the weights can also smoothly recover. This avoids the constraint discontinuity and information drop caused by hard removal, and reduces the risk of convergence oscillations caused by it.
[0052] In addition, dynamic occlusion, local deformation, and illumination perturbations typically exhibit a certain degree of spatial continuity in images. If only point-by-point independent weighting is used, the weight field is easily affected by local noise and becomes fragmented, which is not conducive to forming a stable suppression region. Based on this situation, this invention adopts a point-by-point maintenance method for reliability and uses spatial neighborhood regularization constraints to achieve coupling, so that the weight changes at adjacent positions remain moderately consistent. This constraint can also suppress weight fragmentation caused by local noise, making it easier to identify and suppress inconsistent regions caused by occlusion, non-rigid motion, or appearance anomalies as continuous segments, thereby reducing the impact of scattered noise on the optimization process. Furthermore, appearance consistency cues are subsequently used to assist in adjusting reliability, thereby enhancing the method's ability to distinguish between complex dynamic disturbances and non-stationary appearance changes.
[0053] In dynamic scenes, relying solely on geometric residuals makes it difficult to distinguish whether errors originate from modeling biases that can be absorbed through pose and geometric variable adjustments, or from more systemic disturbances such as occlusion, non-rigid deformation, and non-stationary changes in appearance. Occlusion and deformation directly alter cross-frame visibility and correspondences, causing some dense constraints to exhibit spatially patchy anomalies. Exposure and illumination variations also lead to global or local appearance shifts, making it difficult to maintain a consistent scale for luminance-based discrimination across different time periods. If judgments are based solely on geometric consistency, these factors can easily be mixed into the same residual source, thus weakening the ability to discriminate unreliable observations. Based on this, this invention introduces appearance consistency into the backend and uses gating to adjust the effective strength of cross-frame constraints, making occluded and deformed regions more easily suppressed. Under strong appearance perturbations, it can also adaptively reduce the impact of appearance criteria. This design does not treat appearance as an equivalent primary optimization objective to geometric terms, but rather uses it to form a gating factor that evolves synchronously with the geometric state.
[0054] For any edge in the factor graph, this invention constructs an appearance residual to characterize brightness consistency across frames. Let... Indicates the first Frames optimized grid point coordinates at low resolution The grayscale intensity at each edge determines the target frame. According to the current projection Backsample to reference frame coordinate system, to obtain When the geometry is consistent across frames and the appearance is approximately stable, The residual should exhibit small fluctuations locally, but it typically increases significantly in areas of occlusion, deformation, or significant appearance drift.
[0055] In this embodiment, preferably, the process of obtaining the appearance residuals of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window includes: Use the starting frame of each edge of the factor graph of the current sliding window as the reference frame and the ending frame as the target frame. For each low-resolution optimized grid point on each edge, the difference between the gray intensity of the low-resolution optimized grid point in the reference frame of each edge and the gray intensity of the projection point of the low-resolution optimized grid point in the reference frame in the target frame is used as the appearance residual of the low-resolution optimized grid point on each edge. For each low-resolution optimized grid point, the difference between the grayscale intensity of that low-resolution optimized grid point on each edge of the reference frame and the grayscale intensity of its projection point in the reference frame in the target frame is taken as the appearance residual of that low-resolution optimized grid point on each edge. The formula is as follows: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. Appearance defects Coordinates are The low-resolution optimized grid point grayscale intensity on each edge reference frame. The coordinates in the target frame are The grayscale intensity of the low-resolution optimized grid points projected onto the reference frame.
[0056] In this embodiment, preferably, the method for constructing the final weights of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window based on the reliability weights and appearance residuals of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window includes: Based on the normalized appearance residuals of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window, and the gray-scale mean brightness of each edge of the factor graph of the current sliding window corresponding to two key frames, the appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window are constructed. To enhance scale comparability between different edges and time periods, this invention robustly normalizes the normalized appearance residuals of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window. The normalized appearance residual is defined as follows: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. The normalized appearance residuals This represents median operations. This is a numerical stability term used to avoid numerical explosion caused by excessively small scale estimates.
[0057] This normalization statistically weakens the impact of extreme outliers on scale estimation, making the gating mechanism more stable in areas with varying noise levels and in locally strong texture regions.
[0058] Furthermore, changes in exposure can cause a systematic shift in appearance residuals; if directly used... To suppress unreliable observations, gating may misinterpret exposure-induced errors as dynamic mismatches. Therefore, this invention introduces an exposure adaptive coefficient. The effective intensity used to adjust the photometric data, the exposure adaptive coefficient. It monotonically decreases as the exposure difference increases, thereby automatically reducing the weight of appearance criteria under strong exposure difference conditions.
[0059] The formula for constructing the appearance gating factor of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window is: , in, Optimize grid points for low resolution On the factor graph edge of the current sliding window On the appearance gating factor, This is the factor graph edge index of the current sliding window. The factor graph edges of the current sliding window Reference frame in The factor graph edges of the current sliding window The target frame in the middle, Optimize grid point positions for low resolution. This is a truncation function. It is an exponential function with the natural constant as its base. The factor graph edges of the current sliding window Exposure adaptive coefficient, , For keyframes grayscale mean brightness, For keyframes grayscale mean brightness, For numerically stable terms, To control the intensity of exposure attenuation, Optimize grid points for low resolution On the factor graph edge of the current sliding window Normalized appearance residuals , This serves as a lower bound for the appearance gating factor to prevent information collapse caused by the gating setting constraints to zero locally. In the experiment, it was set to 1.5.
[0060] Appearance gating factors exhibit a clear monotonicity: the stronger the appearance inconsistency, The smaller the size, the stronger the suppression applied to the corresponding area; gating occurs when the appearance is consistent. The value approaches 1, so that the constraint strength is mainly determined by the weighted average of geometric confidence weight and continuous reliability weight.
[0061] The product of the reliability weight and the appearance gating factor of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window is used as the final weight of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window. The formula is as follows: , in, Optimize grid points for low resolution On the factor graph edge of the current sliding window The final weight on, Optimize grid points for low resolution On the factor graph edge of the current sliding window Reliability weights on Optimize grid points for low resolution On the factor graph edge of the current sliding window Appearance gating factor.
[0062] Using scalar gating, the two-dimensional components are treated proportionally. This fusion method incorporates appearance consistency into the backend optimization through gating without altering the second-order solution structure. The current appearance residual is built upon existing correspondences, and the gating changes synchronously with the geometric state update. Occlusion, deformation, and illumination change areas can then obtain additional discrimination information consistent with the geometric state. This information can also serve as a supplementary basis for subsequent residual feedback updates.
[0063] Step S4: Solve by minimizing the weighted geometric consistency objective function to obtain the camera pose and depth of each keyframe within the current sliding window; While continuous reweighting and appearance gating can mitigate the impact of unreliable observations on single linearization updates to some extent, the mismatch patterns in dynamic scenarios still exhibit significant non-stationarity. As the geometric state iterates continuously, projection correspondences, visibility relationships, and residual statistics constantly change. If the observation weights are always fixed according to external priors, their spatial distribution and intensity often fail to keep pace with the current residual structure, potentially leading to overly strong or weak local constraints, impacting convergence stability and solution efficiency. To address this issue, this invention incorporates reliability weights into the backend as updatable internal variables, using residual feedback for online updates. Simultaneously, spatial and temporal regularization is combined to smooth weight changes, preventing drastic oscillations during iteration.
[0064] Step S41: Initialize the camera pose and depth of all keyframes in the current sliding window, and the real-valued scores of each low-resolution optimized mesh point on the two keyframes of each edge of the factor graph in the current sliding window. Set the number of iterations. ; Step S42: Based on the first In the next iteration, the final weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are used to perform a second-order update on the objective function of minimizing the weighted geometric consistency, resulting in the... The next iteration measures the camera pose and depth of each keyframe within the current sliding window; Geometric estimation and reliability estimation are statistically interdependent. Reliability weights determine which cross-frame constraints need to be strengthened or suppressed in the current iteration, thus affecting the update direction of pose and geometric variables. Conversely, after the geometric state is updated, the projection correspondence and residual distribution also change. To preserve the second-order solution structure, adaptive weights are introduced. In the In the next iteration, the reliability scoring field is first fixed. and its induced block weights Edge-level reliability is obtained through endpoint conservative fusion. By combining grid point confidence weights and appearance gating, the final weights for this iteration can be constructed, as shown below: , , Among them, superscript Indicates the first iteration For the first The next iteration optimizes the low-resolution mesh points. On the factor graph edge of the current sliding window The final weight on, For the first The next iteration optimizes the low-resolution mesh points. On the factor graph edge of the current sliding window Edge-level reliability gating, For the first The next iteration optimizes the low-resolution mesh points. On the factor graph edge of the current sliding window The appearance gating factor on the surface, two scalar gating pairs The two-dimensional components act proportionally.
[0065] The formula for performing a second-order update on the objective function that minimizes the weighted geometric consistency is as follows: , in, For the first Camera pose in the next iteration. No. Depth of the next iteration For the first The next iteration optimizes the low-resolution mesh points. On the factor graph edge of the current sliding window The final weights, and the constructed information matrix. The weighted quadratic norm is used. Since the weights in the current iteration are determined by the reliability scoring field and appearance gating, performing a second-order update on the objective function of minimizing the weighted geometric consistency is actually equivalent to performing a standard second-order least squares update on the objective function of minimizing the weighted geometric consistency under fixed weights.
[0066] S43: Based on the first In the next iteration, the camera pose and depth of each keyframe within the current sliding window are updated, along with the edge-level reliability gating, normalized appearance residuals, and geometric residuals of each low-resolution optimized mesh point in the factor graph of the current sliding window. For any edge With grid points The geometric residual is defined by the network correction corresponding to the difference with the current projection, as shown below: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. geometric residuals, This is the network-corrected correspondence. This represents the projection correspondence obtained from the current camera pose and depth prediction.
[0067] Because the residual magnitude is affected by scene depth, parallax amplitude, and linearization points, the residual scales between different edges can vary significantly, making direct comparison impossible. To ensure that reliability updates reflect relative inconsistencies rather than absolute scale differences, this invention performs robust scale normalization on the geometric residuals, using the following formula: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. The normalized geometric residuals, This represents median operations. Indicates in the grid domain Top The set formed by taking the norm point-by-point or the absolute value of each component. It is a numerically stable term.
[0068] The normalized two-dimensional geometric residual vector is used. To facilitate integration with the two-dimensional geometric residual, the normalized appearance residual is used. Scalar expansion to two-dimensional vector The two components take the same value.
[0069] S44: Fusion In the next iteration, the normalized appearance residuals and geometric residuals of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are obtained, yielding the... The fusion residuals of each low-resolution optimized grid point in the next iteration on each edge of the factor graph in the current sliding window; To unify geometric and appearance cues in reliability updates, this invention optimizes the coordinates of each low-resolution mesh point. Define the fusion error intensity above The reason for this approach is that geometric residuals directly reflect the consistency across frames in the current optimized state and are the primary basis for reliability updates, while appearance residuals provide supplementary discrimination for occlusion, non-rigid deformation, and appearance anomalies. Only by organizing these two types of cues into a unified feedback quantity can subsequent updates simultaneously respond to the current geometric mismatch and retain the ability to distinguish appearance anomalies. Specifically, the two types of normalized residuals are first linearly combined according to their weights to obtain the fused residuals. Then, its square norm is used to characterize the error intensity, and the fused residual is defined as follows: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. The fusion residual, The geometric residual fusion coefficient, For the edge Appearance weights are used to adjust the appearance of items on edges. The intensity of the effect.
[0070] Since the reliability unit and the aforementioned residual are defined at the same sampling resolution, the error intensity... It can be directly used for subsequent reliability updates without introducing additional index mappings. Appearance weights Adjusted by the exposure difference between two frames, and stabilized within a bounded range by truncation constraints, it is defined as follows: ,in, As the upper bound of the baseline appearance weight, The lower realm Control the intensity of exposure difference attenuation. For numerically stable terms, For the edge The average grayscale brightness of the reference frame. For the edge The average grayscale brightness of the target frame.
[0071] S45: Based on the first In the next iteration, the edge-level reliability gating and square norm of the fused residuals of each low-resolution optimized grid point in the factor graph of the current sliding window are used to obtain the th iteration. Update data items in the next iteration ; The intensity of the fusion error is defined by the square norm of the fusion residuals, as follows: , in, For channel indexing.
[0072] To ensure that reliability evolves synchronously with residual statistics, this invention samples at each keyframe location. Upgrade Real Value Scoring And the reliability weights are obtained through bounded mapping. The edge-level reliability of cross-frame constraints adopts endpoint conservative fusion, and its definition follows the established rules. Under this setting, online updates directly affect the rating variables. ,and and This is used to modulate the effective strength of cross-frame constraints. To ensure that the score update is consistent with the current error strength, an update data item aligned with the error strength is introduced. Its form is as follows: , in, Let be the set of edges. Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. Edge-level reliability gating, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. The square norm of the fusion residual.
[0073] because The gradient obtained from minimum fusion is piecewise with respect to the endpoint weights, thus prioritizing updates to endpoints with lower reliability, forming a conservative and stable suppression strategy. The derivative of the score to the weight is given by the Sigmoid function, which takes the following form:
[0074] Among them, the truncation operator guarantees Always in Inside, in the experiment and The values are set to 0.05 and 1.0 respectively. This restricts the weights of unreliable observations to a bounded interval, thus helping to maintain numerical stability during the optimization process.
[0075] In obtaining Then, the geometric variables are fixed and the reliability score field is updated based on the current residual statistics. Specifically, the normalized geometric residuals and normalized appearance residuals are calculated using the new geometric state, and the error intensity is constructed according to the residual fusion method described in the previous section. Due to the reliability element and the mesh domain Alignment at the same resolution, error intensity It can be directly used for reliability updates. Based on this, reliability updates are achieved by minimizing the sum of data items and regularization terms, with the following objective:
[0076] S46: Based on the first The spatiotemporal regularization term and the updated data term in the next iteration, the th The real-valued scores of each low-resolution optimized grid point in the next iteration are obtained. In the next iteration, the real-valued scores of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window on two keyframes; No. The update rule for the real-value score of each low-resolution optimized grid point on each keyframe in the next iteration is as follows: , in, For the first In the next iteration, the coordinates are... The low-resolution optimized grid points in the first Real-valued scores on each keyframe For the first In the next iteration, the coordinates are... The low-resolution optimized grid points in the first Real-valued scores on each keyframe Step size, For spacetime regularization, For the first The next iteration updates the data items.
[0077] in, The step size is used to obtain the new value after the update through Sigmoid mapping and truncation. And then proceed to the next iteration.
[0078] Thus, the score update is driven by the current residual feedback, enabling it to respond promptly to error redistribution caused by occlusion, non-rigid motion, and appearance changes. On the other hand, it is also regulated by regularization constraints to prevent the update process from being overly dominated by local anomalies.
[0079] Only depend In practice, the scoring field is often overly sensitive to local noise, commonly manifested as spatial fragmentation and temporal jitter. Furthermore, overall drift may occur in the early stages of iteration. This is because while residual feedback can quickly reflect inconsistencies in the current region, the lack of constraints on neighborhood structure and cross-time continuity means that directly updating the scoring field based on this result can easily amplify local anomalies. To improve the smoothness and stability of the reliability field, this invention adds three types of regularization constraints—spatial smoothing, temporal smoothing, and prior anchoring—in addition to the data terms. The spatial smoothing term suppresses drastic changes in scores within the neighborhood, ensuring moderate consistency in the reliability evolution of adjacent locations within the same region. This invention defines the spatial regularization term using a discrete Laplace form: ;in, For discrete Laplace operators on block adjacency graphs.
[0080] Spatial smoothing alone is insufficient to guarantee the stable evolution of the scoring field across consecutive frames. The residual structure in dynamic scenes is continuously reconstructed with iteration and inter-frame state changes. Without cross-time constraints, the scoring is prone to high-frequency fluctuations between adjacent time points. Therefore, this invention further incorporates a temporal smoothing term to limit the variation amplitude of the scoring field between adjacent frames. The temporal smoothing term is: ; In addition to spatial and temporal continuity, the overall shift of the scoring field also needs to be constrained. Otherwise, when the residuals are not yet sufficiently stable in the early stages of iteration, the scoring field may shift upward or downward globally, weakening its ability to distinguish local anomalies. To address this issue, this invention introduces a priori anchoring terms to maintain a certain consistency between the current scoring field and the initial scoring or external priors. The priori anchoring terms are: ,in, To initialize the scoring field or external prior anchor point.
[0081] Spacetime regularization term , These are weighting coefficients, which were set to 0.01, 0.001, and 0.02 respectively in the experiment.
[0082] Through the above combination, the spatial smoothing term is mainly responsible for suppressing local fragmentation, the temporal smoothing term is mainly responsible for limiting cross-time jitter, and the prior anchoring term is responsible for constraining the overall drift. These three parts, together with the previous residual feedback, constitute the online adaptive mechanism of the reliability field, so that the reliability field no longer manifests as a pre-fixed static weight, but becomes an internal variable that can change together with the geometric state and observation consistency.
[0083] S47: Determine if the iterative convergence condition is met; if not, update. Return to step S42. If the condition is met, output the first... The next iteration calculates the camera pose and depth for each keyframe within the current sliding window.
[0084] From a mechanistic perspective, the aforementioned alternating process forms a closed-loop update framework. Geometric updates complete second-order least squares solutions under given weights, while reliability updates adjust weight configurations based on residual statistics under given geometric states. By employing endpoint minimum fusion for edge-level reliability, the update process becomes more sensitive to low-reliability endpoints, and outlier regions are often prioritized for suppression. Spatiotemporal regularization also constrains drastic spatial and temporal fluctuations in the scoring field, thus avoiding fragmented updates and temporal jitter dominated by local noise. Compared to a one-time fixed-weight approach, this alternating optimization allows the weight distribution to change synchronously with the reconstruction of the residual structure. When dynamic occlusion, non-rigid motion, and non-stationary changes in appearance coexist, it is more conducive to stabilizing and suppressing unreliable observations and maintaining consistency across frame geometric constraints. Geometric estimation and reliability estimation are no longer isolated but interact and converge together within the same optimization loop.
[0085] This invention innovatively proposes ST³-BA (Spatio-Temporal Self-adaptive BundleAdjustment by Soft Weight), which models observation reliability as a continuous spatio-temporal weight field to robustly reweight bundle adjustment geometric constraints. This smoothly suppresses unreliable observations caused by occlusion, non-rigid perturbations, and illumination variations, while also avoiding constraint discontinuities caused by hard rejection. Unlike approaches that introduce dynamic cues or uncertainties as external constraints into the optimization process, ST³-BA directly treats observation reliability as an internal variable in the optimization process and adaptively updates it based on residual feedback during the BA (BundleAdjustment) iterations. Simultaneously, robust normalization, adaptive exposure gating, and spatio-temporal regularization are combined to improve the continuity and stability of the weight field re-evolution, thereby improving convergence and estimation consistency under dynamic conditions.
[0086] Step S5: Based on the camera pose and depth of each keyframe within all sliding windows, determine the final camera pose and depth of each keyframe and generate a global scene map.
[0087] In this embodiment, preferably, the method for determining the final camera pose and depth of each keyframe based on the camera pose and depth of each keyframe within all sliding windows includes: Based on the camera pose and depth of each keyframe within each sliding window, the final camera pose and depth of each keyframe are obtained through local BA, global BA, and loop closure detection.
[0088] ST³-BA improves convergence stability under dynamic conditions through robust weighting within a sliding window. However, relying solely on in-window optimization is insufficient to eliminate long-term drift and error accumulation, so layered optimization on the keyframe map is still indispensable. Local BA, global BA, and loop closure detection maintain the consistency of the trajectory and map from three levels: short-term refinement, long-term consistency redistribution, and loop closure correction, respectively. The connection between ST³-BA and this process lies in the fact that it provides updatable reliability weights and appearance gating for cross-frame constraints in the keyframe map, allowing cross-frame constraints in both local and global optimization to receive consistent and robust weighting in dynamic scenarios.
[0089] Local Basis Analysis (BA) faces the current sliding window. The system iterates and updates the pose and geometric variables within the window at a high frequency to suppress short-time errors and stabilize tracking. This process maintains the weighted least squares solution structure unchanged. The effective information matrix in the system is jointly adjusted by the reliability weights and appearance gating in ST³-BA. Abnormal residuals in occluded regions and non-rigid deformation regions are also continuously reduced in weight during iteration. This reduces the interference of structured outliers on the update direction and improves the consistency of estimation results within the window.
[0090] Global BA corresponds to a larger range of keyframe maps. The system coordinates the accumulated error at a lower frequency to alleviate long-term drift that is difficult to cover by local windows. Global optimization also reuses the robust weighted interface provided by ST³-BA. After the optimization range is expanded, dynamic anomaly constraints can still be suppressed. This can reduce the risk of the global solution being pulled by a small number of error factors and undergoing unreasonable deformation, and maintain geometric consistency at the global scale.
[0091] Loop closure detection is used to explicitly add loop closure constraints to long sequences to correct drift. The system first obtains candidate loop closure pairs through retrieval, then performs geometric consistency verification to complete the screening. After screening, the candidate loop closure edges are added to the keyframe graph, and then global correction is triggered, allowing the loop closure information to propagate throughout the entire global trajectory. After the loop closure constraints enter the optimization process, the system still uses the ST³-BA reliability modulation mechanism, making the loop closure correction rely more on the contribution of high-reliability regions. In this way, even if dynamic foreground and appearance changes exist simultaneously, the effectiveness of the loop closure constraints and the stability of the correction can be improved.
[0092] In summary, local BA, global BA, and loop closure detection have a clear division of labor in terms of their scale of action and triggering frequency. They also obtain a unified robust weighted interface by leveraging the reliability weights and gating mechanisms provided by ST³-BA, ultimately improving both local convergence stability and global consistency in dynamic scenarios.
[0093] In dynamic scenes, motion masks often carry uncertainties at target boundaries that are difficult to completely eliminate. Resolution limitations, interpolation, and post-processing can all cause jagged edges or slight drift at the boundaries. Occlusion variations and thin structures can also cause inconsistencies in labels near the boundaries between adjacent frames. If the mask is directly used as a hard decision to determine whether to discard or retain observations, the constraints in the boundary neighborhood will suddenly change with small changes in labels, thus introducing discontinuous gradient responses during optimization and weakening convergence stability. To alleviate this problem, this invention no longer treats the boundary as a hard decision region. Instead, it forms continuous soft weights in the boundary neighborhood to smoothly adjust the contribution of observations and adds boundary consistency regularization to suppress temporal drift and local noise, allowing the mask to function in a smoother and more stable manner during optimization.
[0094] In this embodiment, preferably, the method for generating a global scene map includes: Based on the motion mask of each keyframe, each keyframe pixel is divided into static and dynamic regions; Based on the final camera pose and depth of each keyframe, the effective pixels in the static and dynamic regions are back-projected into three-dimensional space to obtain the static Gaussian set and the dynamic Gaussian set. Based on the scaling factor and the distance from each pixel in the motion mask boundary neighborhood of each keyframe to the motion mask boundary, the soft weight of each pixel in the motion mask boundary neighborhood of each keyframe is determined, using the following formula: , , in, For the first Pixels in the motion mask boundary neighborhood of each keyframe Soft weights For pixels within the neighborhood of the motion mask boundary, This represents the scaling factor for the soft weights. Indicates the maximum expansion radius of the boundary. Represents pixels To the dynamic mask boundary The minimum radius of expansion, The smaller the value, the closer it is to the boundary, and the higher the weight. The larger, The larger the value, the further away from the boundary, and the higher the weight. The smaller, For the boundary neighborhood of the motion mask, For expansion operation, For dynamic mask boundaries.
[0095] set up For the first For each keyframe's motion mask, to explicitly characterize the uncertain region near the boundary, and to characterize the mixed pixels and mask uncertain region near the mask boundary, extract its inner boundary set: , in, For the boundary zone radius, Indicates pixel-by-pixel XOR, Erosion is performed on a structuring element with a radius of 1. Subsequently, to obtain the boundary neighborhood... This invention performs multi-scale expansion of the boundary with radius (R) and takes the union of the results:
[0096] The soft weight of each pixel in the neighborhood of the motion mask boundary of each keyframe is used as the weight of the photometric consistency loss to obtain the target photometric consistency loss. In the loss function, This is applied as a multiplicative weight to the photometric consistency loss. This process continuously modulates the contribution of the boundary neighborhood, thereby avoiding gradient spikes and constraint discontinuities introduced by boundary jitter.
[0097] Keyframes within a local window are selected as observation data. A total loss function is constructed, consisting of target photometric consistency loss, geometric consistency loss, scale regularization loss, and boundary structure consistency regularization. The parameters of the global static Gaussian and dynamic Gaussian are iteratively updated by minimizing the total loss function to obtain the global scene map.
[0098] Relying solely on single-frame soft weights is insufficient to suppress temporal instability caused by cross-frame boundary drift. Therefore, a temporal consistency constraint is imposed on the soft weight field within the boundary band, allowing it to evolve smoothly over time. The boundary temporal consistency regularity can be defined as: ,in, For boundary time consistency regularization, For the first Pixels in the motion mask boundary neighborhood of each keyframe Soft weights For the first Pixels in the motion mask boundary neighborhood of each keyframe Soft weights for Norm, For the first Motion mask boundary neighborhoods for each keyframe are used to reduce the dominant effect of local anomalies on the regularization term. It should be noted that this boundary temporal consistency regularization is implemented through online updates of the tracking soft weight field, rather than being solved separately as an additional pixel-by-pixel mapping loss; during the mapping stage, a time-stabilized dynamic mask and boundary soft weights are used to constrain the reconstruction error and structural gradient of the boundary region.
[0099] During rendering, this invention introduces boundary structure consistency constraints within the boundary band to reduce the risk of boundary noise being written into the model parameters. The boundary structure consistency term is: ,in, For boundary structure consistency terms, Render an image for the current model. To observe the image, The image gradient operator is represented by the Scharr approximation used in this invention. This term aligns the gradient responses of rendering and observation within the boundary band, suppressing artifact accumulation at the boundary and improving boundary alignment stability.
[0100] The task of the mapping phase is to jointly optimize the Gaussian map parameters based on the camera pose and initial geometry provided by the tracking module, so that the rendered result is consistent with the observed image in both appearance and structure. Dynamic regions often have masking errors and pixel mixing near their boundaries. If only the conventional pixel-level rendering reprojection loss is considered, the optimization process can easily result in color leakage, blurred outlines, or local misalignment at the boundary locations. Therefore, this invention formulates map updating as a unified differentiable rendering optimization problem and explicitly incorporates boundary-related constraints into the objective function, allowing the optimization to obtain more stable gradients and more reliable alignment signals in the boundary neighborhood.
[0101] Record No. Frame observation image is The image rendered from the current Gaussian map is , depth rendering Map optimization is achieved by minimizing the weighted sum of multiple losses, and its total loss function is... As shown in the formula: , in, For the total loss function, The weight is 0.90 in the experiment. It is a loss of photometric uniformity. It is geometric consistency loss. For scale regularization loss, This is a boundary structure consistency term.
[0102] Loss of photometric uniformity L1 loss ensures pixel-level consistency between the rendered image and the actual observation, while geometric consistency loss... The accuracy of 3D geometry is constrained by depth loss, and scale regularization loss is used. Overfitting of Gaussian parameters is prevented by isotropic constraints and scale distribution control, ensuring that the scale parameters of Gaussians are within a reasonable range and that the boundary structure is consistent with regularization. These loss terms, used to enhance structural alignment and detail constraints in the boundary neighborhood, work together to improve the accuracy and robustness of the map, laying the foundation for the robot's high-quality perception of the scene.
[0103] , Geometric consistency loss Constraining 3D geometry accuracy using depth and gradient loss: , Scale regularization loss Overfitting of Gaussian parameters is prevented by isotropic constraints and scale distribution control, ensuring that the scale parameters of Gaussian parameters are within a reasonable range: , in, It contains a set of keyframes within a local window. It is the number of pixels in the keyframe. and These are the exposure compensation parameters for each keyframe, optimizing image brightness differences. and Representing color and depth respectively. and Let Gaussian be the scaling parameter and its mean. It is an L1 norm. For the first in the local window The true color of each keyframe This indicates the keyframe index within a local window. and These are the balance weights for photometric consistency loss and geometric consistency loss, and the weight for scale regularization loss, respectively.
[0104] , , in, Gaussian index, For Gaussian sets, For the first The center of the Gaussian is at a depth along the z-axis. For the first An opacity of one Gaussian unit. For the first An opacity of one Gaussian unit. , Indicates Gaussian index, , For the first An opacity parameter of Gaussians. For the first The two-dimensional mean of a Gaussian projection onto the image plane. For the first The two-dimensional covariance matrix corresponding to the Gaussian projection onto the image plane.
[0105] To address the issues of occlusion blending and structural mismatch that easily occur at dynamic boundaries, this invention proposes BARC (Boundary-Aware Reweighting and Consistency Regularization). This method is based on motion masks and constructs soft weights in the boundary neighborhood that decay according to distance, thereby strengthening the supervision of boundary region reconstruction. Furthermore, it introduces boundary consistency regularization at the boundary position to constrain the rendering results to be consistent with the observations near the contour edges, thereby reducing boundary artifacts and improving the clarity and stability of Gaussian maps in dynamic scenes.
[0106] In pose tracking, this invention first selects keyframes based on RGB stream preprocessing results and forms cross-frame constraints within a sliding window. Subsequently, ST³-BA iteratively updates the pose and depth within the current window in this factor graph. It treats observation reliability as an internal variable in the optimization process and integrates it with observation weights for robust reweighting of BA constraints. Specifically, ST³-BA continuously models observation reliability and combines it with appearance consistency gating, residual feedback, and spatiotemporal constraints to achieve online adaptive updates, suppressing unreliable observations caused by dynamic occlusion, non-rigid motion, and appearance perturbations during continuous adjustment. Based on this, local BA, global BA, and loop closure detection respectively undertake the tasks of local constraint refinement, global drift correction, and loop closure consistency maintenance, ultimately improving the stability and global consistency of the tracking results.
[0107] In scene reconstruction, this invention renders based on the current pose and a Gaussian map, and uses the error between the rendered result and the observation as the driving basis for Gaussian parameter optimization. Specifically, to address issues such as occlusion blending and structural mismatch that easily occur at the boundaries of dynamic objects, BARC forms a soft weighted band with distance decay in the neighborhood of the motion mask boundary, and then applies boundary-aware reweighting to the reconstruction error to strengthen boundary region supervision. This invention also introduces boundary consistency regularization to constrain the alignment of rendering and observation on the contour boundary, ultimately reducing boundary artifacts and improving the clarity and stability of the Gaussian map in dynamic scenes.
[0108] From the perspective of the core components of ST³GS-SLAM, this invention proposes ST³-BA, which explicitly models observation reliability as a continuous bounded spatiotemporal soft weight field to reweight the BA (Bundle Adjustment) constraints. This supports smooth weight reduction for dynamic occlusion and non-rigid perturbation regions, thereby enhancing convergence and estimation consistency in dynamic scenes. Simultaneously, to further mitigate the misleading effects of illumination changes and weak textures, ST³-BA uses appearance consistency and geometric residuals as complementary discriminative information. Based on this, a gating mechanism is constructed to modulate the effective weights of cross-frame constraints, and exposure adaptation and robust normalization are combined to enhance cross-frame comparability. Finally, a closed-loop update is formed between geometric variables and reliability weights through alternating optimization, and regularization constraints are applied in the spatial and temporal dimensions to ensure the effectiveness and stability of the weight field evolution. Furthermore, this invention proposes a boundary uncertainty handling mechanism during the mapping stage. For the drift and jaggedness of motion masks at target boundaries, observation errors are continuously buffered and weighted in the boundary neighborhood, and boundary consistency regularization is introduced to suppress temporal jitter, thereby improving boundary alignment accuracy and detail fidelity. Experimental results on the widely recognized TUM RGB-D and Bonn RGB-D datasets in the robotics field demonstrate that ST³GS-SLAM achieves state-of-the-art (SOTA) performance in both tracking accuracy and reconstruction rendering quality.
[0109] To verify the robustness and generalization ability of the method in dynamic scenes, the experiments selected the mainstream TUM RGB-D dataset and Bonn RGB-D dataset in the field of robotics as evaluation benchmarks. The two datasets have significant differences in terms of dynamic range, motion form and scene complexity, and can examine the system's tracking and mapping performance under dynamic occlusion, non-rigid motion and interference from lighting / appearance changes from different levels.
[0110] The TUM RGB-D dataset provides synchronized RGB images, depth information, and high-precision ground truth camera poses. Some sequences include phenomena such as human movement, local occlusion, and human-scene interaction, making it suitable for evaluating the camera tracking accuracy and optimization stability of methods in moderately dynamic indoor environments. The main challenges of this dataset lie in the frequent interference of foreground objects on the line of sight and the unreliable observations introduced by local dynamic regions, thus it can be used to analyze the robustness of methods under dynamic occlusion conditions.
[0111] The Bonn RGB-D dataset is designed for real-world scenes with higher dynamic levels, featuring a higher proportion of dynamic targets and more complex motion patterns, including significant non-rigid motion, prolonged occlusion, and larger dynamic regions. Compared to TUM RGB-D, this dataset places higher demands on the system's ability to handle time-varying visibility, non-rigid perturbations, and boundary mixing errors, making it more suitable for evaluating the mapping quality and robust optimization capabilities of methods under strong dynamic conditions. Furthermore, some sequences in real-world acquisition processes are accompanied by certain degrees of brightness fluctuations, exposure changes, or local appearance perturbations. These factors further weaken the appearance consistency assumption and interfere with cross-frame correlation and residual optimization. Based on these considerations, this invention further selects representative test sequences from both datasets and categorizes them according to their main dynamic characteristics. As shown in Table 1, Table 1 illustrates the coverage of dynamic occlusion, non-rigid motion, and illumination changes in each test sequence.
[0112] Table 1
[0113] The wk / st type sequences in TUM RGB-D mainly reflect local occlusion and tracking interference under moderate dynamic conditions, while the balloon, person, and crowd type sequences in Bonn RGB-D highlight the impact of non-rigid motion, continuous occlusion, and large-scale dynamic regions on mapping and optimization.
[0114] Therefore, subsequent experiments in this invention will focus on three categories of factors: dynamic occlusion, non-rigid motion, and changes in illumination / appearance, to conduct quantitative and qualitative analyses, thereby ensuring that the experimental results correspond to the method analysis and conclusions presented in this invention.
[0115] To comprehensively evaluate the camera tracking accuracy and renderable map quality of the proposed method in dynamic scenes, this invention employs an evaluation scheme combining geometric accuracy, reconstruction quality, and operational efficiency metrics. On datasets with true pose values, RMSE (Root Mean Square Error) is used as the primary trajectory error metric to measure the overall deviation between the estimated camera trajectory and the true value. This invention uses Sim(3) similarity transformation alignment when calculating ATE, and then calculates the RMSE and SD of ATE and converts them uniformly to centimeters (cm) for reporting, in order to quantitatively analyze the positioning stability and consistency of each method in dynamic environments. Furthermore, this invention uses the evo tool to calculate ATE.
[0116] RMSE comprehensively reflects the cumulative error level and is one of the most commonly used trajectory accuracy evaluation standards in the SLAM field. In addition, STD (Standard Deviation) measures the dispersion of trajectory errors, used to evaluate the stability and consistency of the system under dynamic disturbance conditions. A lower STD indicates that the system has more stable tracking performance across different time periods or under different experimental conditions.
[0117] To evaluate the quality of the constructed map in new perspective synthesis and rendering tasks, three complementary image quality evaluation metrics were adopted: PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index Measure), and LPIPS (Learned Perceptual Image Patch Similarity). PSNR measures the pixel-level error between the rendered result and the reference image, reflecting the overall reconstruction accuracy. SSIM evaluates the consistency of images in terms of brightness, contrast, and structural information from a structural similarity perspective. LPIPS measures perceptual similarity based on depth features, which is more consistent with human visual perception and is particularly sensitive to local artifacts and the reconstruction quality of dynamic regions.
[0118] In addition to tracking accuracy and reconstruction quality, this invention also uses frame rate (FPS) as a performance metric to measure the number of image frames the algorithm can process per unit of time. This metric can intuitively reflect the differences in computational overhead between different methods and provides a basis for a comprehensive comparison of accuracy, reconstruction quality, and performance. By examining these metrics together, a more comprehensive evaluation of the performance of the method in dynamic scenes can be made from multiple perspectives, including geometric accuracy, system stability, perceptual rendering quality, and performance.
[0119] Based on the above indicators, the performance of the method of the present invention in dynamic scenes can be comprehensively evaluated from multiple dimensions such as geometric accuracy, system stability, perceptual rendering quality, and operating efficiency.
[0120] All experiments in this invention were performed on the AutoDL cloud server. The operating system was Ubuntu 22.04, using PyTorch 2.1.0, Python 3.10, and CUDA 12.1. The hardware configuration was a single vGPU with 32 GB of video memory, a 12-vCPU Intel Xeon Platinum 8352V with a clock speed of 2.10 GHz, and 90 GB of RAM. Unless otherwise specified, all experiments were run under the above-mentioned unified hardware and software environment and consistent data preprocessing procedures to ensure the fairness and reproducibility of the comparison results.
[0121] To verify the effectiveness of the method in dynamic scenes, this invention was compared with several representative SLAM methods on the TUM RGB-D and Bonn RGB-D datasets, and a comprehensive analysis was conducted from three aspects: camera trajectory accuracy, reconstruction quality, and running time.
[0122] To systematically evaluate the camera tracking capability of the method in dynamic environments, comparative experiments were conducted on two publicly available datasets, TUM RGB-D and Bonn RGB-D, with multiple sequences exhibiting significant dynamic disturbances selected as test scenarios. The comparisons covered several mainstream and representative SLAM methods, including traditional SLAM frameworks based on features or geometric constraints (such as ORB-SLAM3, DROID-SLAM, and DynaSLAM), NeRF-SLAM methods based on implicit radiation field modeling (such as NICE-SLAM and ESLAM), and state-of-the-art SLAM methods employing 3D Gaussian representation (such as MonoGS, SplaTAM, and Splat-SLAM). Furthermore, to further demonstrate the differences in dynamic modeling capabilities, this invention also introduced several recently proposed dynamic SLAM methods based on 3D Gaussian Splatting for comparison, including DG-SLAM and 4DGS-SLAM, as well as Dy3DGS-SLAM and WildGS-SLAM based on RGB-only input.
[0123] The experimental results are shown in Tables 2 and 3. Table 2 shows the tracking results of the TUM_RGBD dataset, and Table 3 shows the tracking results of the Bonn RGB-D dataset.
[0124] Table 2
[0125] Table 3
[0126] Where “X” indicates that the sequence tracking failed, “-” indicates that the method did not report the result or is not applicable, and “Avg.” is the arithmetic mean of the successful sequences. If a method fails in a certain sequence, that sequence is not included in its “Avg.”.
[0127] Overall, the method of this invention achieved superior or tied-optimal results on both datasets, demonstrating stronger dynamic robustness in trajectory accuracy and stability, with a more pronounced advantage in highly dynamic sequences. This indicates that the method has a stronger adaptability to dynamic occlusion, non-rigid motion, and appearance perturbations introduced by real-world acquisition. On TUM_RGBD, the method of this invention achieved optimal tracking performance, demonstrating stronger consistency and stability under dynamic occlusion, observational perturbations caused by rapid motion, and local non-rigid changes introduced by the human body itself. Compared to WildGS-SLAM, the method of this invention performed slightly worse on the fr3_wk_hf sequence. This is mainly because this sequence involves more drastic rotation and rapid viewpoint changes, resulting in fewer effective static observations available for constructing geometric constraints. In addition, large rotation scenes are more sensitive to keyframe coverage and geometric constraint distribution, which can also cause insufficient constraints in local stages, ultimately leading to a slight increase in RMSE. Nevertheless, this difference is very limited, indicating overall tracking stability. On Bonn RGB-D, the proposed method also performs excellently overall. Under more dynamic conditions, the advantages of the proposed method are further amplified, indicating that the proposed method can still maintain relatively stable tracking accuracy even under conditions of higher dynamic levels, more complex motion patterns, and longer periods of occlusion. In contrast, balloon2 is a sequence for which the proposed method did not achieve optimal or suboptimal results. This may be because the boundary deformation and occlusion changes caused by the non-rigid motion of balloons and other objects in this sequence make it more difficult to fully explain the geometric consistency residuals. In this case, the reliability weighting strategy of the proposed method may further expand the low-weight region and reduce the effective constraint density, thus resulting in a slightly higher RMSE than the comparative method that uses more aggressive modeling on highly dynamic foregrounds and makes fuller use of foreground cues.
[0128] The TUM RGB-D results demonstrate the high accuracy advantage of the method in moderately dynamic scenarios, while the Bonn RGB-D results further verify that it can maintain stable tracking even under strong dynamic disturbances, indicating that the method can more effectively suppress unreliable constraints and improve the robustness and consistency of trajectory estimation when dynamic observations exist.
[0129] To verify the reconstruction quality of the method of the present invention, the present invention was compared with 3DGS-based SLAM methods (SplaTAM and Splat-SLAM) and 4DGS-based 4DGS-SLAM.
[0130] The quantitative results are shown in Tables 4 and 5. Table 4 compares the reconstruction quality of the TUM RGB-D dataset, and Table 5 compares the reconstruction quality of the BonnRGB-D dataset.
[0131] Table 4
[0132] Table 5
[0133] Visual comparison chart as follows Figure 3 As shown, Figure 3 Comparison charts showing the rendering effects of different methods. Figure 3 (a) in the figure represents a comparison of the rendering effects of different methods on the TUM / fr3_wk_xyz data. Figure 3 (b) in the figure represents a comparison of the rendering effects of different methods on the TUM / fr3_wk_st data. Figure 3 (c) in the figure represents a comparison of the rendering effects of different methods on the Bonn / balloon dataset. Figure 3 From left to right, the images are SplaTAM, Splat-SLAM, 4DGS-SLAM, the method of this invention, and a real reference image.
[0134] Overall, ST³GS-SLAM achieves more stable and consistent rendering and reconstruction results in most dynamic sequences. This advantage is reflected not only in clearer single-frame images but also in less drift and contamination of structures and textures across frames. Compared to the statically hypothesized 3DGS-SLAM method, when dynamic objects repeatedly pass through the viewpoint, static mapping often continuously writes inconsistent observations into the Gaussian map, ultimately manifesting as ghosting, trailing shadows, or local texture contamination in the rendering, accompanied by jitter at the edges of background structures. ST³GS-SLAM results exhibit fewer such dynamic remnants, indicating that it more consistently suppresses unreliable observations, and map updates are more dominated by long-term consistent data. Compared to dynamic 3DGS methods that rely solely on time-varying representations or simple static / dynamic separation, the lack of effective robust constraints can lead to frequent local failures in correspondences due to occlusion and non-rigid deformation. Model updates are easily driven by anomalous gradients, resulting in local geometric collapse, boundary flickering, or temporal inconsistencies. ST³GS-SLAM is more stable in such segments, indicating that it not only possesses dynamic representation capabilities but also the ability to suppress the long-term interference of structured outliers on the optimization process. Overall, these phenomena demonstrate that in dynamic scenes, the key to reconstruction quality lies not only in the ability to represent dynamics but also in the ability to avoid embedding short-term inconsistent observations into the map during long-term iterations, thereby maintaining rendering consistency and interpretability.
[0135] The aforementioned advantages are directly related to two key design features of this invention. First, ST³-BA explicitly models observation reliability as an updatable continuous weight field and adaptively adjusts it online through residual feedback. This continuously suppresses structured outliers caused by dynamic occlusion and deformation during the iteration process, rather than relying on a one-time fixed weight or hard removal. Since the projection correspondence and residual statistics are continuously reconstructed with the update of the geometric state, this closed-loop update allows the weight distribution to remain synchronized with the current error structure, thereby reducing the interference of outlier residuals on the second-order update direction and reducing long-term rendering pollution caused by erroneous correspondences being written into the map. Second, the boundary-aware soft weights introduced in the mapping stage address the issues of jagged edges, drift, and inconsistencies across frames at the boundaries of instances or motion masks. Hard decisions are transformed into continuous gating and combined with consistency constraints to suppress boundary jitter, making the gradient distribution in the boundary neighborhood smoother and the update more stable. This reduces the risk of dynamic edge flickering, thin structure fragmentation, and local noise being solidified into the Gaussian map. The two methods are functionally complementary. The former improves the robustness of tracking and geometric estimation, while the latter enhances the detail fidelity of boundary regions, enabling the system to achieve more consistent rendering and reconstruction results under complex dynamic conditions. Failure to achieve single-item optimality on a few sequences is usually related to the type of perturbation and the trade-offs in stability. When appearance changes dominate or boundaries are extremely thin and change rapidly, the method's conservative suppression strategy may result in slight detail smoothing, while some comparative methods may more aggressively fit the appearance in these segments, potentially gaining a slight advantage on a particular metric.
[0136] Overall, ST³GS-SLAM achieves more consistent and stable reconstruction results through continuous suppression and adaptive updates, with its advantages being particularly evident under complex dynamic conditions where occlusion, non-rigid deformation, and non-stationary appearance coexist.
[0137] To evaluate the computational overhead of the proposed method, this invention statistically analyzed the running efficiency of different methods under the same experimental environment. The results are shown in Table 6, which is a running time evaluation.
[0138] Table 6
[0139] Overall, Gaussian splash SLAM methods for static scenes generally have higher operating efficiency, while dynamic scene methods need to additionally handle target motion, temporal variations, and fluctuations in observation reliability, often resulting in higher computational costs. As shown in Table 6, the ST³GS-SLAM proposed in this invention exhibits better operating efficiency in dynamic Gaussian splash SLAM methods, with an overall speed superior to the compared dynamic methods. This indicates that after adding boundary consistency constraints, the system does not incur excessive additional computational burden and can still maintain relatively stable operating efficiency. Compared to static scene methods, ST³GS-SLAM does have a certain speed gap; however, considering that the system needs to handle more complex dynamic disturbances and additionally model the spatiotemporal variations in observation reliability, this efficiency level is sufficient to support the experimental verification and result analysis of this invention.
[0140] To verify the effectiveness of the two key designs proposed by ST³GS-SLAM in dynamic scenes, this invention conducted component ablation experiments on the BonnRGB-D dataset. In the w / o ST³-BA setting, the ST³-BA method was removed, while the DBA method of DROID-SLAM was used. In the w / o BARC setting, the BARC method was removed, and in the w / o both setting, both were removed simultaneously. The results are shown in Table 7, which presents the ablation experiment results.
[0141] Table 7
[0142] The purpose of this experiment is not only to compare the impact of different modules on the metrics, but also to examine whether the two designs are effective in addressing two core challenges in dynamic scenes: dynamic occlusion, non-rigid motion, and backend geometric instability caused by changes in lighting, as well as rendering detail degradation caused by uncertain dynamic boundaries.
[0143] Table 7 shows that removing ST³-BA significantly worsens the system's tracking stability and geometric consistency. In contrast, removing BARC results in a smaller change in trajectory metrics, but a more pronounced decrease in rendering quality. This indicates that the main source of performance improvement is ST³-BA's ability to continuously model observation reliability and its online updating function during iteration. Dynamic occlusion, non-rigid motion, and lighting changes cause continuous fluctuations in the residual distribution across frame geometric constraints. If only fixed weights or one-time weight configurations are used, the system often struggles to suppress the continuous interference from unreliable observations. ST³-BA, however, writes the observation reliability as a continuously updatable score field, and then combines residual feedback, appearance gating, and spatiotemporal regularization to achieve adaptive adjustment, allowing reliability to change along with the current geometric state. Consequently, removing this module makes it difficult to continuously suppress the structured mismatch caused by occlusion and non-rigid motion regions, and residual non-stationarity is more easily directly transmitted to the optimization update, ultimately resulting in decreased convergence stability and worse pose estimation consistency. Therefore, the main function of ST³-BA is to alleviate the time-varying reliability problem in backend optimization under dynamic scenarios, which is the key to the method of this invention to obtain stable tracking performance.
[0144] Compared to removing ST³-BA, removing BARC results in a smaller change in pose parameters, such as... Figure 4 As shown, Figure 4 This is a diagram showing the results of a BARC ablation experiment. Figure 4 (a) in the figure shows the experimental results using BARC. Figure 4 (b) in the diagram shows the experimental results without using BARC. However, Figure 4 The rendering results showed more pronounced blurring and structural mismatch near the boundaries of dynamic targets, indicating that BARC has a limited direct impact on pose estimation, but plays a more crucial role in detail recovery and structural alignment in boundary regions. BARC adds soft weights with distance decay characteristics to the boundary neighborhood, combined with boundary consistency constraints, to suppress boundary drift and local noise intrusion. Therefore, its gains are more reflected in rendering quality than in trajectory metrics. Removing both modules simultaneously resulted in a more significant overall degradation, indicating a strong complementary relationship between the two modules. ST³-BA is primarily responsible for stable convergence and consistent estimation under dynamic conditions, while BARC mainly handles detail constraints and rendering fidelity in boundary regions. The complete approach thus achieves superior overall performance in both tracking and rendering.
[0145] This invention proposes a renderable localization and mapping framework, ST³GS-SLAM, that relies solely on RGB input for 4DGS-SLAM in complex dynamic scenes. Addressing dynamic occlusion, non-rigid motion, and accompanying appearance perturbations, this invention primarily tackles two types of problems: cross-frame geometric mismatch and residual non-stationarity. It proposes ST³-BA, which expresses observation reliability as a continuously updated spatiotemporal weight field within a sliding window. Combined with robust normalization, exposure adaptive gating, and spatiotemporal regularization, it continuously suppresses unreliable observations and performs smooth backfilling after consistency restoration, thereby improving the convergence stability and estimation consistency of the backend optimization. Regarding artifacts and detail degradation caused by pixel mixing at mask boundaries and cross-frame inconsistencies, this invention further incorporates boundary-aware residual reweighting and boundary structure consistency constraints during the mapping stage. This adjusts the error contribution of the boundary neighborhood in a smoother manner, reducing color leakage and contour drift, and improving detail fidelity.
[0146] Experiments on TUM RGB-D and Bonn RGB-D dynamic sequences demonstrate that ST³GS-SLAM achieves leading performance in pose stability and rendering quality, reaching state-of-the-art levels in multiple metrics. Ablation results validate the complementarity of the two key designs: ST³-BA primarily enhances robust convergence and stable tracking under dynamic conditions, while the boundary mechanism mainly improves rendering details and boundary alignment quality.
[0147] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0148] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0149] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0150] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0151] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields, characterized in that, include: Set a sliding window for a set of keyframes in a continuous RGB video stream; Each sliding window is processed sequentially, and each keyframe image within the current sliding window is converted into a low-resolution optimized mesh, and a factor map of the current sliding window is constructed. In dense bundle adjustment, based on the reliability weights and appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window, the network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are corrected to construct a weighted geometric consistency objective function. The camera pose and depth of each keyframe within the current sliding window are obtained by minimizing the weighted geometric consistency objective function. Based on the camera pose and depth of each keyframe within all sliding windows, the final camera pose and depth of each keyframe are determined, and a global scene map is generated.
2. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 1, characterized in that, Methods for constructing the factor graph corresponding to each sliding window include: For each sliding window, each keyframe within the sliding window is treated as a node, and the node state includes camera pose and depth information; bidirectional adjacency edges are constructed between temporally adjacent keyframes; based on the bidirectional inter-frame distance of any keyframe pair, target keyframe pairs are selected, cross-frame constraint edges are constructed, and the factor graph corresponding to each sliding window is obtained.
3. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 1, characterized in that, Methods for obtaining the reliability weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window include: Based on the real-valued scores, preset maximum weight values, and preset minimum weight values of each low-resolution optimized grid point on the keyframes corresponding to each edge of the factor graph in the current sliding window, the initial reliability weights of each low-resolution optimized grid point on the keyframes corresponding to each edge of the factor graph in the current sliding window are constructed. Based on the initial reliability weights of each low-resolution optimized grid point in the factor graph of the current sliding window corresponding to two key frames, the minimum value is selected as the edge-level reliability gate of the low-resolution optimized grid point in the factor graph of the current sliding window. Based on the product of the edge-level reliability gate and the network prediction confidence weight, the reliability weights of each low-resolution optimized grid point in the factor graph of the current sliding window are obtained.
4. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 3, characterized in that, The method for correcting the network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window, based on the reliability weights and appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window, includes: The product of the reliability weight of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window and the appearance gating factor is used as the final weight of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window. The network prediction confidence weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are corrected by using the final weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window.
5. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 4, characterized in that, Methods for solving the problem by minimizing the weighted geometric consistency objective function include: Step S41: Initialize the camera pose and depth of all keyframes in the current sliding window, and the real-valued scores of each low-resolution optimized mesh point on the two keyframes of each edge of the factor graph in the current sliding window. Set the number of iterations. ; Step S42: Based on the first In the next iteration, the final weights of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are used to perform a second-order update on the objective function of minimizing the weighted geometric consistency, resulting in the... The next iteration measures the camera pose and depth of each keyframe within the current sliding window; S43: Based on the first In the next iteration, the camera pose and depth of each keyframe within the current sliding window are updated, along with the edge-level reliability gating, normalized appearance residuals, and geometric residuals of each low-resolution optimized mesh point in the factor graph of the current sliding window. S44: Fusion In the next iteration, the normalized appearance residuals and geometric residuals of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window are obtained, yielding the... The fusion residuals of each low-resolution optimized grid point in the next iteration on each edge of the factor graph in the current sliding window; S45: Based on the first In the next iteration, the edge-level reliability gating and square norm of the fused residuals of each low-resolution optimized grid point in the factor graph of the current sliding window are used to obtain the th iteration. The next iteration updates the data items; S46: Based on the first The spatiotemporal regularization term and the updated data term in the next iteration, the th The real-valued scores of each low-resolution optimized grid point in the next iteration are obtained. In the next iteration, the real-valued scores of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window on two keyframes; S47: Determine if the iterative convergence condition is met; if not, update. Return to step S42. If the condition is met, output the first... The next iteration calculates the camera pose and depth for each keyframe within the current sliding window.
6. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 1, characterized in that, Methods for obtaining the appearance gating factors of each low-resolution optimized grid point on each edge of the factor graph in the current sliding window include: Based on the normalized appearance residuals of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window, and the gray-scale mean brightness of each edge of the factor graph of the current sliding window corresponding to two keyframes, the appearance gating factor of each low-resolution optimized grid point on each edge of the factor graph of the current sliding window is constructed, and the formula is as follows: , in, Coordinates are The low-resolution optimized grid points are located on the factor graph edges of the current sliding window. On the appearance gating factor, Optimize grid point coordinates for low resolution. This is the factor graph edge index of the current sliding window. The factor graph edges of the current sliding window The reference frame in the text is the edge. The starting frame, The factor graph edges of the current sliding window The target frame in the image is the edge. The terminating frame, Optimize grid point positions for low resolution. This is a truncation function. It is an exponential function with the natural constant as its base. The factor graph edges of the current sliding window Exposure adaptive coefficient, , For keyframes grayscale mean brightness, For keyframes grayscale mean brightness, For numerically stable terms, To control the intensity of exposure attenuation, Optimize grid points for low resolution On the factor graph edge of the current sliding window Normalized appearance residuals , This is the lower bound of the appearance gating factor.
7. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 6, characterized in that, The process of obtaining the appearance residuals of each low-resolution optimized grid point on each edge of the factor map in the current sliding window includes: Use the starting frame of each edge of the factor graph of the current sliding window as the reference frame and the ending frame as the target frame. For each low-resolution optimized grid point on each edge, the difference between the gray intensity of the low-resolution optimized grid point in the reference frame of each edge and the gray intensity of the projection point of the low-resolution optimized grid point in the reference frame in the target frame is used as the appearance residual of the low-resolution optimized grid point on each edge.
8. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 1, characterized in that, The methods for determining the final camera pose and depth of each keyframe based on the camera pose and depth within all sliding windows include: Based on the camera pose and depth of each keyframe within each sliding window, the final camera pose and depth of each keyframe are obtained through local BA, global BA, and loop closure detection.
9. The method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on spatiotemporal weighted fields according to claim 1, characterized in that, Methods for generating a global scene map include: Based on the motion mask of each keyframe, each keyframe pixel is divided into static and dynamic regions; Based on the final camera pose and depth of each keyframe, the effective pixels in the static and dynamic regions are back-projected into three-dimensional space to obtain the static Gaussian set and the dynamic Gaussian set. Based on the scaling factor and the distance from each pixel in the motion mask boundary neighborhood of each keyframe to the motion mask boundary, the soft weight of each pixel in the motion mask boundary neighborhood of each keyframe is determined, and boundary time consistency regularization constraints are set for it. The soft weight of each pixel in the neighborhood of the motion mask boundary of each keyframe is used as the weight of the photometric consistency loss to obtain the target photometric consistency loss. Keyframes within a local window are selected as observation data. A total loss function is constructed, consisting of target photometric consistency loss, geometric consistency loss, scale regularization loss, and boundary structure consistency regularization. The parameters of the static Gaussian and dynamic Gaussian are iteratively updated by minimizing the total loss function to obtain the global scene map.
10. A method for dynamic scene four-dimensional Gaussian synchronous localization and mapping based on a spatiotemporal weighted field according to claim 9, characterized in that, The formula for determining the soft weight of each pixel in the motion mask boundary neighborhood of each keyframe, based on the scaling factor and the distance from each pixel in the motion mask boundary neighborhood to the motion mask boundary, is as follows: , , in, For the first Pixels in the motion mask boundary neighborhood of each keyframe Soft weights For pixels within the neighborhood of the motion mask boundary, This represents the scaling factor for the soft weights. Indicates the maximum expansion radius of the boundary. For the boundary zone radius, Represents pixels To the dynamic mask boundary The minimum radius of expansion, For the boundary neighborhood of the motion mask, For expansion operation, For dynamic mask boundaries.