Reinforcement learning adaptive multi-mode SLAM method based on 4D Gaussian splashing

By introducing a reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splash in the SLAM system, the problems of tracking loss and high computing cost in the existing technology in dynamic and complex environments are solved, and more efficient and robust positioning and map construction effects are achieved.

CN120141447AActive Publication Date: 2025-06-13SHANDONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510607098.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-13
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing multimodal SLAM technology has problems such as tracking loss, relocation failure, large trajectory drift, closed-loop detection failure, map construction ghosting in dynamic, complex and changeable environments, and the calculation cost is high, making it difficult to effectively apply in resource-constrained environments.

Method used

A reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splash is proposed. Through reinforcement learning, sensor mode is automatically selected based on real-time environmental information, and map updates and optimizations are used to integrate it into the SLAM system to improve work efficiency and real-timeness and save resources.

Benefits of technology

This method significantly improves the adaptability and resource optimization capabilities of SLAM systems in dynamic environments, improves positioning accuracy and map construction quality, reduces power consumption and computing complexity, and enhances the robustness and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120141447A_ABST
    Figure CN120141447A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning adaptive multi-mode SLAM (Simultaneous Localization and Mapping) method based on 4D Gaussian splashing, and belongs to the field of simultaneous localization and map construction. An overall solution of further combining 4D GS and reinforcement learning on the basis of an existing multi-sensor fusion SLAM system is provided, a sensor mode is autonomously selected according to real-time environment information through reinforcement learning, and meanwhile, map updating and optimization are performed by utilizing the 4D GS and are integrated into the SLAM system. Therefore, the purposes of improving the working efficiency and the real-time performance of the whole real-time multi-mode SLAM system and saving resources are achieved. Comprising the following steps: step (1), multi-source sensor input and data preprocessing; (2) selecting a sensor combination type; (3) carrying out multi-source sensor data fusion processing; and step (4), updating and optimizing the map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Simultaneous Localization and Mapping (SLAM), and specifically proposes an adaptive multi-modal online SLAM method based on reinforcement learning for 4D Gaussian Splatting (4D GS) applicable to dynamic environments. Background Art

[0002] The Simultaneous Localization and Mapping (SLAM) method is a technology that uses multi-sensor (such as cameras, lidar, IMU, etc.) data to estimate the sensor pose and simultaneously reconstruct a 3D map of the surrounding environment. The SLAM technology can solve the problems of autonomous localization and navigation of robots or devices in unknown environments and is one of the key technologies for realizing fully autonomous mobile robots.

[0003] Existing 3D Gaussian Splatting is a technology for 3D scene reconstruction and rendering. It represents points or surfaces in the scene using Gaussian functions and processes them in 3D space, but 3D GS focuses on static scenes. It can be extended to dynamic scenes for 4D representation, and the key technical issue is how to model complex point motions from sparse inputs. 4D Gaussian Splatting can be regarded as an extension of 3D Gaussian Splatting, which adds a fourth dimension on the basis of 3D space and is usually used to process dynamic scenes or time-varying data. 4D Gaussian Splatting has advantages in processing dynamic content, such as modeling and rendering of dynamic objects, scene reconstruction under time-varying lighting conditions, etc.

[0004] So far, the SLAM technology has developed SLAM systems that integrate lasers, vision, and IMU multi-sensors. As in the following prior publicly disclosed Chinese patent application, with the application number CN202410154510.8 and the title "A Dense Visual SLAM Method and System Using a Three-Dimensional Gaussian Backend Representation", which proposes a dense visual SLAM method using a three-dimensional Gaussian backend representation. By constructing a three-dimensional Gaussian scene representation, performing adaptive three-dimensional Gaussian expansion, reconstructing the scene geometric structure, and performing coarse-to-fine camera differentiable pose estimation based on the reconstructed scene geometric structure. It faces challenges in dealing with dynamic objects or rapidly changing scenes, and the computational cost is relatively high, restricting its application in resource-constrained environments.

[0005] For another example, the application number is CN202411103239.1, and the name is "A Laser-Enhanced Visual 3D Reconstruction Method and System Based on Gaussian Splashing". It is divided into two parts: pose calculation and scene reconstruction. During pose calculation, relative rotation is solved using image feature matching, the point cloud ICP pose is initialized, and the relative pose is optimized. During scene reconstruction, the scene is represented by a three-dimensional Gaussian sphere, and the depth map generated by projecting the image and point cloud with high-precision pose is used to supervise the synthetic image of differentiable rendering, and the parameters of the three-dimensional Gaussian sphere are optimized to achieve scene reconstruction. The dependence on lidar data may lead to the degradation of the reconstruction effect in case of inaccurate or occluded lidar data, and problems such as data inconsistency and calibration error are faced during the data fusion process. The overall complexity of the algorithm is relatively high, and the requirement for computing resources is also high.

[0006] For another example, the application number is CN202410277176.5, and the name is "A Dense Gaussian Map Reconstruction Method and System in a Dynamic Environment". It uses instance segmentation to divide potential moving objects to obtain a dynamic prior region, uses a dynamic point detection algorithm to detect and remove dynamic feature points in the segmented region, obtains camera pose estimation information based on static feature points, and combines a visual SLAM system and a 3D Gaussian Splatting framework to perform dense map reconstruction on the static region. The definition and detection of dynamic objects rely on instance segmentation and a preset dynamic prior region, and it cannot adapt to all types of dynamic scenes, especially the adaptability to fast or complex dynamic change scenes is insufficient, and the dependence on optical flow method and epipolar geometry method leads to inaccurate dynamic point detection in some cases.

[0007] The above-mentioned multi-modal SLAM technologies of the prior art mainly have the following deficiencies and defects: First, due to dynamic objects in the scene, problems such as tracking loss and relocalization failure occur in the SLAM system; that is, when working in a dynamically complex and changeable environment, it faces very serious degradation problems. Especially in low-texture areas, limited-view constraints, and dynamic environments, the pose estimation and map construction effects are not good. Dynamic objects cause problems such as large trajectory drift, closed-loop detection failure, and ghosting in map construction in the SLAM system. Second, due to the generally adopted fixed rules such as weighted fusion and static priority, it is difficult to cope with sudden environmental changes such as sensor occlusion or sudden light changes, and the simultaneous operation of multiple sensors will lead to redundant calculations, increasing the power consumption and latency of mobile devices. Third, the existing NeRF methods bear a large training and rendering cost, and there are non-negligible latency problems in the rendering process. Fourth, representing the scene as 3D Gaussian cannot meet the requirements of dynamic scenes for real-time rendering speed.

[0008] In view of this, the present application is specifically proposed. Summary of the Invention

[0009] The proposed Reinforcement Learning Adaptive Multi-Modal SLAM Method Based on 4D Gaussian Splashing aims to address the problems in the existing technologies by presenting an overall solution that further integrates 4D GS and reinforcement learning on the basis of an existing multi-sensor fusion SLAM system. Through reinforcement learning, it autonomously selects sensor modes according to real-time environmental information, while using 4D GS for map updating, optimization, and integration into the SLAM system, thereby achieving the goals of improving the working efficiency, real-time performance, and resource savings of the entire real-time multi-modal SLAM system.

[0010] To achieve the above design objectives, the Reinforcement Learning Adaptive Multi-Modal SLAM Method Based on 4D Gaussian Splashing includes the following steps: Step 1: Multi-source sensor input and data preprocessing; Receive point cloud data from lidar, RGB image data from depth cameras, and IMU data, and perform data preprocessing. Step 2: Select sensor combination types; Based on the reinforcement learning sensor selection module, the reinforcement learning policy network learns the environmental state in real time, dynamically selects the most reliable sensor combination type, and activates sensors as needed. Step 3: Multi-source sensor data fusion processing; Based on the data fusion processing module, multi-source sensor data fusion processing is performed by the sensor combination. Step 4: Map updating and optimization; Based on the map updating and backend optimization module, using the map data provided by the front-end data fusion processing module, combined with 4D GS, Gaussian map updating and optimization of the environmental map are carried out, and error correction and map optimization are performed through loop closure detection.

[0011] Furthermore, Step 1 includes the following steps: Step 1.1 Data input; The SLAM system synchronizes the point cloud data from lidar, RGB image data from depth cameras, and IMU data in time. Step 1.2 Incorporation of time information; Timestamps are introduced in the data preprocessing stage , and each lidar point cloud and image frame carries a timestamp. Step 1.3 Extrinsic parameter calibration of multi-sensors; Use the Kalibr framework to calibrate the extrinsic parameters between sensors. Step 1.4 Initialization of the reinforcement learning module; The sensor selection module based on reinforcement learning is initialized during the data preprocessing stage. By analyzing historical data and environmental characteristics, it learns the strategy of selecting the optimal sensor combination under different scenario conditions. In the initial state, a random exploration strategy is adopted.

[0012] Furthermore, step 2 includes the following steps: Step 2.1 State space design; Define the environmental information that the reinforcement learning system can observe at each decision-making moment. Step 2.2 Define the action space formula; It includes defining the action of the system to switch sensor modes to optimize data acquisition and system performance. The SLAM system adopts a hybrid discrete-continuous space to achieve refined control. The mode selection part is discrete and includes three sensor-dominated modes: depth camera-IMU mode, lidar-depth camera mode, and lidar-IMU-depth camera mode. The weight allocation part is continuous and includes sensor gain coefficients and resource limit parameters. The sensor gain coefficients are the weights of the lidar, depth camera, and IMU, and the range is between 0 and 1. The expression is as follows: (9) The resource limit parameter is the lidar sampling rate, and the range is between 5Hz and 40Hz. The expression is as follows: (10) Step 2.3 Reward function module; It includes a main reward, an auxiliary reward, and a penalty term. Step 2.4 Network architecture design; It includes a state encoder, which serves as the basis of the entire network. It is responsible for converting multi-modal sensor data into a unified feature representation so that the subsequent policy network and value network can make decisions and evaluations based on these features. The policy network is responsible for outputting specific action selections based on the encoded state features, including discrete action modes and continuous parameter adjustments. The value network is used to evaluate the quality of the current state and action combination, provide a learning signal for the policy network, and help it optimize the decision-making process. Step 2.5 Execute the training strategy; Through an efficient data collection mechanism, an intelligent exploration strategy, and strict safe learning constraints, a comprehensive training framework is provided for the reinforcement learning sensor adaptive selection module. It includes: The data collection mechanism provides the system with empirical data for interacting with the environment and is the basis for learning and optimizing strategies. Exploration strategy, which determines how the system effectively explores in an unknown or partially known environment to obtain more information and experience; Safety learning constraints, ensuring that the system not only pursues high performance but also meets a series of safety and practicality requirements during the learning and decision-making processes.

[0013] Furthermore, step 2.1 includes the following: Geometric dynamic feature design, including lidar point cloud density gradient and dynamic object coverage rate; The lidar point cloud density gradient reflects the change in the density of the point cloud in space, helps the system perceive the geometric structure of the environment and potential obstacles, and quantifies the observation confidence of the lidar in the local area. The expression is as follows: (1) Where, represents the number of lidar hits in a unit voxel, represents the voxel volume, represents the range of reflection intensity normalization values; Calculating the dynamic object coverage rate represents the proportion and distribution of dynamic objects in the scene, enabling the system to timely understand the dynamic characteristics of the environment. The expression is as follows: (2) Where, represents the estimated velocity vector of the dynamic object, represents the projected area of the dynamic object detection box, represents the field of view area of the sensor; Sensory reliability, including IMU confidence and radar-vision depth consistency. IMU confidence reflects the credibility of the inertial measurement unit data and is calculated based on its noise level and bias stability. The higher the resulting value, the more reliable the IMU data. The expression is as follows: (3) Where, represents the gyroscope zero-bias estimation error, represents the accelerometer noise standard deviation (calculated using a sliding window); Radar-vision depth consistency measures the degree of agreement between the visual sensor and the radar depth data. It is calculated by comparing the depth measurement values of the two. The higher the consistency, the more coordinated the depth data of the two sensors. The expression is as follows: (4) Where, represents the observed depth distribution histogram, represents the depth distribution histogram of the depth camera, Represents the intersection - union ratio of the effective regions of the two. System resource status monitoring, including real - time computing power load and remaining energy budget. The real - time computing power load reflects the current computing pressure of the system, helping the module decide whether to reduce the amount of sensor data processing or simplify the computing tasks; the expression is as follows: (5) Among them, Represents the used video memory capacity of the current GPU, Represents the total video memory capacity of the GPU, Represents the rendering time per frame, Represents the expected frame period; The remaining energy budget indicates the energy reserve situation of the system, enabling the module to select a sensor combination or adjust the working mode based on the current remaining energy budget; the expression is as follows: (6) Among them, Represents the remaining battery level, Represents the initial total power, Represents the running time of the system, Represents the power decay factor; Spatio - temporal context, including historical decision - making memory and scene category probability. The historical decision - making memory records the success rate and effect of past sensor selection decisions, helping the module learn from experience, avoid repeating mistakes, and improve decision - making efficiency; the expression is as follows: (7) Among them, Represents the th step action in the past, Represents the long - short - term memory network; The scene category probability calculates the probabilities of belonging to a low - texture scene, a limited - view - constraint scene, and a dynamic scene according to the current environmental characteristics, enabling the module to adjust the sensor selection strategy according to the scene characteristics; the expression is as follows: (8) Among them, Represents CLIP The image - text joint embedding vector extracted by the model, Represents the learnable weight matrix, To normalize this distribution, the sum is 1.

[0014] Furthermore, step 2.3 includes, The main rewards include the positioning accuracy reward and the energy - efficiency economy reward, which are used to improve the accuracy of system positioning and the energy utilization efficiency respectively; the expression is as follows: (11) (12) Among them, ATE represents the absolute trajectory error, represents the total instantaneous power consumption, represents the reward compensation item when the remaining power exceeds the threshold; Auxiliary rewards, including rendering quality rewards and policy stability rewards, are designed to improve the visual effect of scene reconstruction and ensure the continuity and reliability of the policy; the expression is as follows: (13) (14) Among them, SSIM represents the structural similarity index, and LPIPS represents the learned perceptual image patch similarity, represents the action mutation penalty, represents the smoothness constraint of the policy network parameters; Penalty terms, including fatal penalty errors and resource overload penalties, are used to avoid unreasonable sensor selection and excessive consumption of resources, thereby guiding the system to learn the optimal sensor selection strategy and achieving comprehensive optimization of system performance; the expression is as follows: (15) (16) The comprehensive reward calculation formula is as follows: (17) Furthermore, step 3 described above includes the following steps, The reinforcement learning sensor selection module adaptively selects the most suitable sensor mode according to the current environmental conditions and system state, considering the influence of each parameter on the sensor mode selection; among them, the weight function expression is as follows: + + + + + (21) Among them, , , , , and respectively represent the lidar confidence , the dynamic object coverage rate , the IMU confidence , the radar-vision consistency , computing power load and remaining energy budget with high weights; Step 3.1 Depth camera - IMU mode; When the reinforcement learning sensor selection module selects according to the current environmental characteristics and based on the weight function S≥ , as the judgment threshold, it is judged that the SLAM system selects the sensor combination of the depth camera and the IMU, with the depth camera as the main and the IMU as the auxiliary; Step 3.2 LiDAR - depth camera mode; When the reinforcement learning sensor selection module selects according to the current environmental characteristics and based on the weight function , as the judgment threshold, it is judged that the SLAM system selects the sensor combination of the LiDAR and the depth camera, with the depth camera as the main and the LiDAR as the auxiliary; Step 3.3 LiDAR - IMU - depth camera multi - sensor balanced mode; When the reinforcement learning sensor selection module selects according to the current environmental characteristics and through the weight function based on the weight function , it is judged that the SLAM system selects the sensor mode of the combination of the LiDAR, the depth camera and the IMU; After the reinforcement learning sensor selection module executes actions and interacts with the environment in Step 3.4, it records the historical decision memory through the reward function module and optimizes the update of the entire policy network through value network evaluation.

[0015] Furthermore, the said Step 3.1 includes, Step 3.1.1 Data input and key - frame selection; Frames containing new scene structures or significant feature increases are preferentially selected as key frames to improve the integrity and efficiency of the map; Step 3.1.2 Point - cloud registration using Generalized ICP tracking; Select the RGB image of the th frame and the depth image to generate a point cloud , where each point , and calculate the covariance matrix of each point . Estimate the relative pose transformation between the current frame source point cloud and the target map point cloud through the GICP algorithm; Step 3.1.3. IMU pre - integration; Fuse the high-frequency measurement values of the IMU to generate the inter-frame motion pre-integration quantity as the initial guess of GICP, improve the tracking efficiency, and combine with the GICP algorithm to enhance the robustness and accuracy of the system; Step 3.1.4 Update the pre-integration; Use the optimized result of GICP to correct the drift of the IMU integration caused by noise and bias over time.

[0016] Furthermore, the step 3.2 includes the following steps, Step 3.2.1 Input the images from the depth camera and the point cloud from the lidar; Use the calibrated external signal for integration to convert the time-aligned LiDAR point cloud into a depth image; the expression is as follows: (38) Where, is the LiDAR point cloud, and are the rotation matrix and translation vector from the lidar to the camera coordinate system respectively, is the intrinsic matrix of the camera; Step 3.2.2 Use the incremental error minimization function; Ensure the exact correspondence between the plane and the points, the formula is as follows: (39) Where, represents a point in the LiDAR point cloud in, is the result after iterations based on the current pose estimate from the previous moment to the world coordinate system, is the Gaussian center closest to is the normal vector of is is the weight of the point is the regularization term, which is used to enhance the stability and accuracy of the error function and considers the direction error between the normal vectors; Introduce the regularization term to enhance the stability and precision of the error function and consider the error in the normal direction; the expression is as follows: (40) Where, is the normal of the current Gaussian distribution; Step 3.2.3 Weight function calculation; The calculation steps of the weight function are as follows: a. Determine the Gaussian center within the local spherical region: Find all the nearest Gaussian distribution centers therein, where is the center of the sphere, is the radius; b. Calculate the density function of the Gaussian points. The density function is calculated by the following formula: (41) where, is the reconstructed covariance matrix, which is constructed by selecting the minimum variance along the normal direction and the larger variance in the perpendicular direction ; c. Simplify the density function calculation. To speed up the calculation, simplify the density function calculation during the tracking process: (42) d. Consistency calculation. For each point , calculate the consistency between the normal of the current Gaussian distribution and the local average normal ; e. Complex texture calculation. For the image region corresponding to each radar point, calculate the local texture complexity. Project the radar point onto the pixel coordinate system of the camera to obtain its position in the image. Take this as the center to intercept ( = 16) image patches. Convert the image patches to grayscale images and calculate the variance of the pixel intensities: (43) where, is the pixel intensity, is the mean of the image patch; Convert the variance to a texture weight between 0 and 1 through the sigmoid function: (44) where, is the scaling factor used to adjust the variance sensitivity; f. Final weight function. Define the final weight function as the product of the normal consistency, the density function, and the texture complexity, that is .

[0017] Furthermore, step 3.3 includes the following steps, Step 3.3.1 Data input and hardware synchronization; The lidar provides sparse but highly accurate 3D point clouds, the depth camera captures the RGB texture of the scene, and the IMU outputs angular velocity and acceleration at high frequency for motion prediction. Ensure that the data timestamps of the three are strictly aligned to avoid timing drift; Step 3.3.2 selects key frames through the input of the depth camera; Select representative frames from the continuous data stream as key frames to reduce redundant calculations; Step 3.3.3. IMU data is used for state propagation; According to the input IMU data and the state of the previous key frame, predict the current state through IMU pre-integration, which is used for state estimation for prediction and forward prediction for motion de-distortion; Step 3.3.4 de-distorts the lidar input; Using the continuous poses predicted by the IMU, each lidar point is transformed from the local coordinate system at the scanning moment to the global coordinate system; the expression is as follows: (45) where, represents the acquisition timestamp of point , represents the pose at the moment obtained by IMU interpolation.

[0018] Furthermore, the step (4) includes the following steps, Step 4.1 Sliding window maintenance; Step 4.2 4D Gaussian distribution; Step 4.3 Introduce optical flow to solve the overfitting problem of 4D GS; Step 4.4 Map update and optimization; Step 4.5 Loop detection; By extracting lidar and visual features and generating feature descriptors, the module can detect potential loop candidates; Use geometric verification and consistency checks to confirm the loop hypothesis; Feed the verified loop constraints back into the map optimization process to globally optimize the map structure.

[0019] In summary, the reinforcement learning adaptive multi-modal SLAM method based on 4D Gaussian splash described in this application has the following advantages and beneficial effects: 1. Based on the existing multi-sensor fusion SLAM system, this application adaptively selects an appropriate sensor combination according to specific environmental characteristics through reinforcement learning, has significant environmental adaptability and resource optimization capabilities, improves the positioning accuracy and map construction quality, while reducing power consumption and computational complexity, and enhances the robustness, efficiency, and resource utilization rate of the system.

[0020] 2. This application integrates the 4D Gaussian splash technology into the multi-modal SLAM system, expands the scene adaptability of the SLAM system, and has significant advantages over traditional SLAM technologies in terms of rendering speed, memory efficiency, dynamic adaptability, robustness, resource efficiency, and global consistency.

[0021] 3. The 4D GS adopted in this application can efficiently represent and process complex changes in dynamic scenes by adding a time dimension on the basis of 3D GS. At the same time, optical flow is introduced in 4D GS to solve the overfitting problem, enabling it not only to capture the motion and deformation of dynamic objects, but also to significantly reduce memory occupancy and computational complexity through sparse Gaussian distribution representation, and improve the stability and generalization ability of the model.

[0022] 4. This application provides more prior information through optical flow, restricts the overfitting of the deformation field network of 4D GS to the noise in the training data, enables the deformation field network to learn a more reasonable and physically consistent Gaussian point deformation method, and improves the stability and generalization ability of the model.

[0023] 5. This application is particularly suitable for resource-constrained mobile devices and complex and changing dynamic scenes, and provides a more efficient and accurate solution for the application of the SLAM system in dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present application will be further described in conjunction with the following drawings; Figure 1 is a system framework diagram of the adaptive multi-modal SLAM method described in this application; Figure 2 is a flowchart of the reinforcement learning sensor adaptive selection module; Figure 3 is a flowchart of the depth camera-IMU mode; Figure 4 is a flowchart of the lidar-depth camera mode; Figure 5 is a flowchart of the lidar-depth camera-IMU mode; Figure 6 is a flowchart of the 4D Gaussian splash; Figure 7 is a schematic diagram of the Gaussian splash; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To further elaborate on the technical means adopted by this application to achieve the predetermined design objective, the following relatively preferred implementation solutions are proposed in combination with the accompanying drawings.

[0026] Specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementation manners disclosed below.

[0027] As Figure 1 shown, the system applying the adaptive multi-modal SLAM method described in this application includes a reinforcement learning sensor selection module, a data fusion processing module, a map update and backend optimization module, and a loop detection and global map optimization module. This SLAM system adopts three types of sensor combinations, namely depth camera - IMU, lidar - depth camera, and lidar - IMU - depth camera.

[0028] As Figures 2 to 7 shown, the reinforcement learning adaptive multi-modal SLAM method based on 4D Gaussian splash includes the following steps: Step 1: Multi-source sensor input and data preprocessing; Receive point cloud data from the lidar, RGB image data from the depth camera, and IMU data, and perform data preprocessing; Step 1.1 Data input; The SLAM system synchronizes the point cloud data received from the lidar, the RGB image data received from the depth camera, and the IMU data in time to ensure that all data at the same timestamp can accurately correspond; Step 1.2 Incorporation of time information; To handle dynamic scenes, timestamps are introduced in the data preprocessing stage , and each lidar point cloud and image frame is accompanied by a timestamp, so as to accurately track and compensate for the dynamic changes of the scene in subsequent processing, thereby providing the necessary time background for the deformation field network of 4D GS; Step 1.3 Calibration of multi-sensor extrinsic parameters; Use the Kalibr framework to calibrate the extrinsic parameters between sensors; Step 1.4 Initialization of the reinforcement learning module; Based on the reinforcement learning sensor selection module, it is initialized in the data preprocessing stage. This module analyzes historical data and environmental characteristics to learn the strategy of selecting the optimal sensor combination under different scene conditions; in the initial state, the module has no prior knowledge of the reliability of the sensors, so a random exploration strategy is adopted; Step 2: Select the type of sensor combination; Based on the reinforcement learning sensor selection module, the environmental state is learned in real time through the reinforcement learning policy network, including but not limited to data such as light, dynamic object density, and sensor noise, and the most reliable sensor combination type is dynamically selected and the sensors are activated as needed to adapt to the highly dynamic and complex environment; Step 2.1 State space design; Define the environmental information that the reinforcement learning system can observe at each decision-making moment to guide the sensor selection strategy; including the following aspects, Geometric dynamic feature design, including lidar point cloud density gradient and dynamic object coverage rate; The lidar point cloud density gradient reflects the change in the density of the point cloud in space, helps the system perceive the geometric structure of the environment and potential obstacles, and quantifies the observation confidence of the lidar in the local area. The expression is as follows, (1) Among them, represents the number of lidar hits in a unit voxel, represents the voxel volume, represents the range of the reflection intensity normalization value; Calculating the dynamic object coverage rate represents the proportion and distribution of dynamic objects in the scene, enabling the system to timely understand the dynamic characteristics of the environment. The expression is as follows, (2) Among them, represents the estimated velocity vector of the dynamic object, represents the projected area of the dynamic object detection frame, represents the field of view area of the sensor; The above two features jointly provide key information for the reinforcement learning module, enabling it to optimize the sensor selection strategy according to the static and dynamic characteristics of the environment, and improving the adaptability and accuracy of the system in complex environments; Sensory reliability, including IMU confidence and radar-vision depth consistency. IMU confidence reflects the credibility of the inertial measurement unit data, which is calculated based on its noise level and bias stability. The higher the result value, the more reliable the IMU data. The expression is as follows: (3) Among them, represents the gyroscope zero bias estimation error, represents the standard deviation of the accelerometer noise (calculated by a sliding window); The radar-vision depth consistency measures the degree of agreement between the visual sensor and the radar depth data. It is calculated by comparing the depth measurement values of the two. The higher the consistency, the more coordinated the depth data of the two sensors are. The expression is as follows: (4) Among them, represents the observed depth distribution histogram, represents the depth distribution histogram of the depth camera, represents the intersection over union of the effective regions of the two; Both of them jointly provide key information on the reliability of sensor data for the reinforcement learning sensor module, assist in optimizing sensor selection and weight allocation, and improve the perception ability and decision-making accuracy of the system in complex environments; System resource status monitoring includes real-time computing power load and remaining energy budget. The real-time computing power load reflects the current computing pressure of the system and helps the module decide whether to reduce the amount of sensor data processing or simplify the computing task. The expression is as follows: (5) Among them, represents the used video memory capacity of the current GPU, represents the total video memory capacity of the GPU, represents the rendering time per frame, represents the expected frame period; The remaining energy budget indicates the energy reserve situation of the system, enabling the module to preferentially select a sensor combination with lower energy consumption or adjust the working mode when the energy is limited. The expression is as follows: (6) Among them, represents the remaining battery level, represents the initial total power, represents the running time of the system, represents the power decay factor; Through the combined action of the two, it can ensure that the system realizes the efficient use of resources and the reasonable management of energy consumption while meeting the performance requirements; Spatio-temporal context includes historical decision memory and scene category probability. The historical decision memory records the success rate and effect of past sensor selection decisions, helps the module learn from experience, avoid repeating mistakes, and improve decision-making efficiency. The expression is as follows: (7) Among them, represents the th step action in the past, represents the long short-term memory network; The scene category probability calculates the likelihood of the scene belonging to different categories, such as long corridors, weak texture areas, etc., according to the current environmental characteristics, enabling the module to adjust the sensor selection strategy based on the scene characteristics; the expression is as follows: (8) Wherein, represents CLIP the image-text joint embedding vector extracted by the model, represents the learnable weight matrix, is to normalize this distribution, and the sum is 1; The combination of the two can make the sensor selection more intelligent and adaptable; Step 2.2 Define the action space formula; It includes defining all possible actions that the system can execute, that is, adjusting the sensor parameters to optimize data acquisition and system performance; The SLAM system adopts a hybrid discrete-continuous space to achieve refined control. The mode selection part is discrete, including three sensor-dominated modes: Mode 1 is the depth camera-IMU mode, which is suitable for situations with good lighting conditions and relatively stable environmental structures; Mode 2 is the lidar-depth camera mode, which is suitable for environments with many dynamic objects and low-light or no-light scenarios, etc.; Mode 3 is the lidar-IMU-depth camera mode, which is suitable for complex and changeable environments, systems running for a long time, and scenarios with high-precision positioning requirements, etc.; The weight allocation part is continuous, including the sensor gain coefficient and the resource limit parameter. The sensor gain coefficient is the weight of the lidar, depth camera, and IMU, and the range is between 0 and 1; the expression is as follows (9) The resource limit parameter is the lidar sampling rate, and the range is between 5Hz and 40Hz; the expression is as follows: (10) The action space formula provides a mathematical framework for the system to adjust the sensor configuration, enabling it to make flexible decisions in a dynamic environment, balance data quality and resource consumption, and thus improve the overall performance of the system; Step 2.3 Reward function module; It guides the SLAM system to learn the optimal sensor selection strategy to achieve the overall optimization of the system performance. It can improve the positioning accuracy and mapping quality, take into account the resource utilization efficiency, avoid excessive consumption, and constrain the system behavior through penalty terms to prevent unreasonable sensor selection, so as to ensure the stable and efficient operation of the entire SLAM system in a complex and changeable environment.

[0029] It includes a main reward, an auxiliary reward, and a penalty term; Among them, the main reward includes a positioning accuracy reward and an energy efficiency economy reward, which are respectively used to improve the accuracy of system positioning and the energy utilization efficiency; the expressions are as follows: (11) (12) Among them, ATE represents the absolute trajectory error, represents the total instantaneous power consumption, represents the reward compensation term when the remaining power exceeds the threshold; Step 2.3.2 Auxiliary reward; It includes a rendering quality reward and a policy stability reward, aiming to improve the visual effect of scene reconstruction and ensure the continuity and reliability of the policy; the expressions are as follows: (13) (14) Among them, SSIM represents the structural similarity index, and LPIPS represents the learned perceptual image patch similarity, represents the action mutation penalty, represents the smoothness constraint of the policy network parameters; The penalty term, including a fatal penalty error and a resource overload penalty, is used to avoid unreasonable sensor selection and excessive consumption of resources, thereby guiding the system to learn the optimal sensor selection strategy and achieving comprehensive optimization of system performance; the expressions are as follows: (15) (16) The above structure can enable the reward function to comprehensively guide the system to optimize sensor selection to improve the overall performance; The comprehensive reward calculation formula is as follows: (17) Step 2.4 Network architecture design; As the core of the reinforcement learning sensor selection module, it can determine how the system processes multi-modal sensor data and makes intelligent decisions; including, The state encoder, as the basis of the entire network, is responsible for converting multi-modal sensor data into a unified feature representation so that the subsequent policy network and value network can make decisions and evaluations based on these features; including, a. Backbone Network: The ResNet-18 architecture is adopted to process visual features. ResNet-18 is a classic deep residual network that can effectively extract high-level semantic information in images. At the same time, the degradation problem of deep networks is alleviated through residual connections, ensuring the training effect and feature extraction ability of the network.

[0030] b. Branch Network: Point cloud Transformer is used to extract geometric features. The Transformer architecture has advantages in processing sequence data, can capture global dependencies in point cloud data, and extract more representative geometric features, which helps the system understand the three-dimensional structure of the environment.

[0031] c. Temporal Network: Bi-LSTM is used to process IMU sequences. Bi-LSTM can consider both past and future information of IMU data, capture bidirectional dependencies in time series, and thus more accurately model the dynamic characteristics of IMU data, providing stable pose estimation for the system.

[0032] The policy network is responsible for outputting specific action selections according to the encoded state features, including discrete action modes and continuous parameter adjustments. Among them, a. Mode Selection Branch: Gumbel-Softmax is adopted to output discrete actions. Gumbel-Softmax is a technique that can balance between continuous relaxation and discrete sampling, enabling the network to effectively learn the selection of discrete actions during training while maintaining the propagability of gradients, and is suitable for discrete decision-making tasks such as selecting sensor-dominated modes.

[0033] b. Parameter Adjustment Branch: The Tanh activation function is used to output continuous weights. The Tanh activation function can limit the output within the range of [-1, 1]. Through appropriate scaling and offset, it can be mapped to the required continuous parameter space, such as sensor gain coefficients and sampling rate parameters, to achieve fine adjustment of sensor parameters.

[0034] The value network is used to evaluate the pros and cons of the current state and action combination, provide learning signals for the policy network, and help it optimize the decision-making process; including a. Dueling DQN Structure: Separate the state value and the advantage function. The Dueling DQN structure divides the value network into two parts, estimating the state value function (V) and the advantage function (A) respectively, and then obtaining the final Q value through a specific combination method. This separation structure enables the network to more accurately evaluate the relative pros and cons of different actions in the current state, improving learning efficiency and stability.

[0035] b. Multi-Head Attention: Dynamically weights multi-modal features. The multi-head attention mechanism can simultaneously focus on different aspects of different modal features and dynamically adjust the weights of each modal feature according to the current task requirements, achieving effective fusion of multi-modal information and enhancing the network's adaptability to complex environments.

[0036] Step 2.5 Execute the training strategy; Through an efficient data collection mechanism, intelligent exploration strategies, and strict safety learning constraints, a comprehensive training framework is provided for the reinforcement learning sensor adaptive selection module; including, Data collection mechanism, which provides the system with empirical data for interacting with the environment and is the basis for learning and optimizing strategies. Specifically, a. Deployment environment: Select the Gazebo simulation environment, which provides high-fidelity scene simulation and flexible sensor configuration options for testing and training robot algorithms.

[0037] b. Sampling frequency: Set to 120FPS real-time sampling. High frame rate captures more delicate environmental dynamics and robot motion states, providing richer information for subsequent data processing and policy learning.

[0038] Exploration strategies, which determine how the system effectively explores in unknown or partially known environments to obtain more information and experience. Specifically, a. Adaptive -greddy: As the number of training rounds increases, the value gradually decreases, and the system transitions from extensive exploration to exploiting the learned knowledge, balancing exploration and exploitation to improve policy performance; the expression is as follows: (18) b. Intrinsic curiosity: Add a state prediction error reward, that is, estimate the next state through a prediction model and use the prediction error as an intrinsic reward; the expression is as follows: (19) Where, is the intrinsic reward, indicating the intrinsic reward obtained by the system when taking action in state , is a tuning coefficient used to control the intensity of the intrinsic reward; is the system's prediction of the next state, estimated through a forward model, is the actual next state feature; Safety learning constraints, ensuring that the system not only pursues high performance but also meets a series of safety and practicality requirements during the learning and decision-making processes. Specifically, a. Action Masking: Prohibit the selection of high-power actions that exceed the battery capacity, that is, mask those actions in the action space that will cause the battery to deplete quickly. Ensure that the system takes into account energy limitations during the decision-making process, avoid unreasonable high-power consumption behaviors, and improve the practicality and sustainability of the system.

[0039] b. Policy gradient correction, the expression is as follows: (20) Among them, among represents the gradient of the probability distribution of the actions output by the policy network with respect to the network parameters, is the action value function, is the baseline function.

[0040] Through the deep combination of the above hierarchical state representation, hybrid action space, and multi-objective reward function in this application, the sensor decision-making achieves dynamic balance in the three dimensions of time-space-energy. Therefore, the SLAM system can achieve the advantages of adapting to dynamic environments, optimizing resource management, improving the accuracy of positioning and mapping, enhancing the robustness of the system, and supporting long-term stable operation.

[0041] Step 3: Multi-source sensor data fusion processing; Based on the data fusion processing module, multi-source sensor data fusion processing is performed by the sensor combination, so as to provide rich and accurate data for the construction of the Gaussian map; The reinforcement learning sensor selection module adaptively selects the most suitable sensor mode according to the current environmental conditions and system state, aiming at the influence of each parameter on the sensor mode selection; among them, the weight function expression is as follows: + + + + + (21) Among them, , , , , and respectively represent the lidar confidence , the coverage rate of dynamic objects , the IMU confidence , the radar-vision consistency , the computing power load and the remaining energy budget with high weights; It includes the following steps: Step 3.1 Depth camera-IMU mode; When the reinforcement learning sensor selection module, according to the current environmental characteristics, the lidar confidence is low, the coverage rate of dynamic objects is low, the IMU confidence is high, the radar-vision consistency is low, and the computing power load is moderate and the remaining energy budget is high, for example, in an environment with good lighting conditions, relatively stable environmental structure, many specular objects, or few dynamic objects with rich textures, etc., according to the weight function S≥ , as the judgment threshold, when it is judged that the SLAM system selects Mode 1, that is, the sensor combination of the depth camera and the IMU is adopted, with the depth camera as the main and the IMU as the auxiliary; Step 3.1.1 Data input and key frame selection; Frames containing new scene structures or significant feature increases are preferentially selected as key frames to improve the integrity and efficiency of the map; Step 3.1.2 Use Generalized ICP Tracking (GICP) for point cloud registration; Select the RGB image of the th frame and the depth image to generate a point cloud , where each point . Calculate the covariance matrix of each point of the current frame source point cloud and the relative pose transformation with the target map point cloud ; including, a. Distribution distance calculation: Model each point as a Gaussian distribution . After the source point cloud is transformed by , the distance (22) Its distribution is: (23) where, , represents the coordinates of the corresponding points in the target map and the source point cloud, , represents the covariance matrices of the target point cloud and the source point cloud; b. Maximum Likelihood Estimation: Solve for the optimal transformation by maximizing the log-likelihood of the probability density function : (24) The optimization objective is simplified to minimizing the Mahalanobis distance, and the expression is as follows: (25) Step 3.1.3 IMU Pre-integration; Fuse the high-frequency IMU measurements to generate the inter-frame motion pre-integration quantity as the initial guess for GICP, improve the tracking efficiency, and combine with the GICP algorithm to enhance the robustness and accuracy of the system; including, a. Construct the original measurement value of the IMU measurement model: (26) (27) Where, , represents the original acceleration and angular velocity measurements, represents the rotation from the IMU to the world coordinate system, represents the gravity vector, / represents the sensor bias, / represents the sensor noise; b. Perform motion mode recursion: The recursion formulas for position , velocity , rotation are as follows: (28) (29) (30) c. Calculate the IMU pre-integration quantity: To avoid repeated integration, define the relative inter-frame motion quantity from the th frame to the th frame, and the expression is as follows: (31) (32) (33) Where, , , respectively represent the integral of position, velocity, and the product of rotation ; represents the skew-symmetric matrix of angular velocity .

[0042] d. Calculate the relative transformation between consecutive frames : Use the external parameters between the camera and the IMU sensor Convert the relative transformation to camera coordinates. Obtain a good initial guess for GICP tracking from IMU pre-integration ; (34) (35) Step 3.1.4 Update the pre-integration; Adopt the GICP optimization result to correct the drift of the IMU integration caused by noise and bias over time; including a. Convert the camera pose optimized by GICP to the IMU coordinate system; the expression is as follows: (36) where represents the external parameters from the camera to the IMU, and respectively represent the results of GICP tracking for position and rotation.

[0043] b. Update the IMU state variables: , , (37) Step 3.2 LiDAR-depth camera mode; When the reinforcement learning sensor selection module, according to the current environmental characteristics, the LiDAR confidence is high, the coverage rate of dynamic objects is high, the IMU confidence is low, the radar-vision consistency is medium, the computing power load is large and the remaining energy budget is moderate, the reinforcement learning sensor adaptive selection module, according to the current environment, such as low-light and low-texture environments, according to the weight function , as the judgment threshold, when it is judged that the SLAM system selects mode 2, adopt the sensor combination of LiDAR and depth camera, with the depth camera as the dominant and the LiDAR as the auxiliary; including Step 3.2.1 Data input: Images from the depth camera and point clouds from the lidar are input; Integrate using a calibrated external signal to convert the time-aligned LiDAR point cloud into a depth image; The expression is as follows: (38) Where, is the LiDAR point cloud, and are the rotation matrix and translation vector from the lidar to the camera coordinate system respectively, is the intrinsic matrix of the camera; Step 3.2.2 Use the incremental error minimization function; Ensure the exact correspondence between the plane and the points, and the formula is as follows: (39) Where, represents a point in the LiDAR point cloud in, is the result after iterations based on the current pose estimate from the previous moment to the world coordinate system, is the Gaussian center closest to , is 's normal vector. is the weight of point , is the regularization term, which is used to enhance the stability and accuracy of the error function and takes into account the directional error between normal vectors; Introduce the regularization term to enhance the stability and accuracy of the error function and consider the error in the normal direction; The expression is as follows: (40) Where, is the normal of the current Gaussian distribution; Step 3.2.3 Weight function calculation; To distinguish between Gaussian points generated only by color supervision and Gaussian points generated by lidar depth simultaneously, the system introduces a weight function. This weight function combines the consistency of the normal vector, the density factor, and the texture complexity to evaluate the reliability of different Gaussian points. The calculation steps of the weight function are as follows: a. Determine the Gaussian center within the local spherical region: Find all the closest Gaussian distribution centers within, where is the center of the sphere, is the radius; b. Calculate the density function of the Gaussian points, and the density function is calculated by the following formula: (41) where is the reconstructed covariance matrix, which is constructed by selecting the minimum variance along the normal direction and the larger variance in the perpendicular direction for construction.

[0044] c. Simplify the density function calculation. To speed up the calculation, simplify the density function calculation during the tracking process: (42) d. Consistency calculation. For each point , calculate the consistency between the normal of the current Gaussian distribution and the local average normal ; ; e. Complex texture degree calculation. For the image area corresponding to each radar point, calculate the local texture complexity. Project the radar point onto the pixel coordinate system of the camera to obtain its position in the image. Take this as the center to intercept ( = 16) image patches, convert the image patches to grayscale images, and calculate the variance of the pixel intensity: (43) where is the pixel intensity, is the mean of the image patch; Convert the variance to a texture weight between 0 and 1 through the sigmoid function: (44) where is the scaling factor used to adjust the variance sensitivity; f. Final weight function. Define the final weight function as the product of the normal consistency, the density function, and the texture complexity, that is ; Step 3.3 LiDAR-IMU-depth camera multi-sensor equalization mode; When the reinforcement learning sensor selection module is based on the current environmental characteristics, the LiDAR confidence is medium, the dynamic object coverage is high, the IMU confidence is high, the radar-vision consistency is low, and the computing power load Larger and energy surplus budget When it is high, such as in a complex and changing dynamic environment, or an environment with high-precision positioning requirements for long-term operation, according to the weight function , when it is judged that the SLAM system selects Mode 3, a sensor mode combining lidar, depth camera and IMU is adopted; including Step 3.3.1 Data input and hardware synchronization; The lidar provides sparse but high-precision 3D point clouds, the depth camera captures the RGB texture of the scene, and the IMU outputs angular velocity and acceleration at high frequency for motion prediction, ensuring that the data timestamps of the three are strictly aligned to avoid timing drift; Step 3.3.2 Key frame selection through depth camera input; Select representative frames from the continuous data stream as key frames to reduce redundant calculations; Step 3.3.3. IMU data is used for state propagation; According to the input IMU data and the state of the previous key frame, predict the current state through IMU pre-integration for state estimation prediction and forward prediction for motion de-distortion; Step 3.3.4 De-distortion of lidar input; Using the continuous poses predicted by the IMU, transform each lidar point from the local coordinate system at the scanning moment to the global coordinate system; the expression is as follows: (45) Among them, represents the point acquisition timestamp, represents the pose at the moment obtained by IMU interpolation.

[0045] After the reinforcement learning sensor selection module in Step 3.4 executes actions and interacts with the environment, the reward function module records the historical decision memory, and through the value network evaluation, optimizes the update of the entire policy network.

[0046] Step 4: Map update and optimization; Based on the map update and backend optimization module, use the rich map data provided by the front-end multi-mode sensor data fusion processing module, combine with 4D GS to perform Gaussian map update and optimization on the environmental map, and perform error correction and map optimization through loop detection; Including the following steps, Step 4.1 Sliding window maintenance; In the data fusion processing module, the system maintains a sliding window that filters and selects point clouds from the nearest 10 time frames in the Gaussian map to construct Gaussian points while masking the remaining Gaussian points. This selection process ensures that the Gaussian points are relevant in the sub-map of current interest; Step 4.2 4D Gaussian distribution; The Gaussian deformation field network is used to model the motion and shape changes of the Gaussian distribution of dynamic objects. The network consists of an efficient spatio-temporal structure encoder and a multi-head Gaussian deformation decoder. Its goal is to transform the standard 3D Gaussian distribution to a new position and shape by learning the Gaussian deformation field, so as to achieve efficient representation and real-time rendering of dynamic scenes, through the Gaussian deformation field network of 4D GS Predict the deformation amount of the current frame point cloud ; The expression is as follows: (46) Among them, represents the deformed three-dimensional Gaussian function.

[0047] Specifically, the spatio-temporal structure encoder aims to efficiently encode the spatial and temporal features of the 3D Gaussian distribution. It consists of a multi-resolution HexPlane module and a small multi-layer perceptron MLP.

[0048] The multi-resolution HexPlane module is used to efficiently encode the spatial and temporal features of the 3D Gaussian distribution. It is achieved by decomposing the 4D neuro-voxel into multiple 2D planes, which can be sampled and encoded at different resolutions. According to the input Gaussian map and time stamp t , extract the center coordinates of the 3D Gaussian distribution G t and the time stamp , obtain the voxel features by querying the multi-resolution plane module , and use bilinear interpolation for querying. It contains 6 multi-resolution plane modules which , represent the resolution levels.

[0049] Each plane module is defined as where is the hidden dimension of the feature, is the basic resolution of the voxel grid.

[0050] Query the voxel features by bilinear interpolation, and the formula is: (47) Among them, is the feature of the neural voxel.

[0051] A small MLP , used to merge all features: (48) Among them, is the final feature representation.

[0052] Specifically, the multi-head Gaussian deformation decoder D is used to decode the deformation of each 3D Gaussian distribution from the features obtained by the encoder. It consists of three independent MLPs that calculate the deformations of position, rotation, and scale respectively.

[0053] The position deformation head, used to calculate the position deformation , and the formula is: (49) The rotation deformation head, used to calculate the position deformation r , and the formula is: (50) The scale deformation head, used to calculate the scale deformation , and the formula is: (51) Apply these deformation amounts to the original 3D Gaussian distribution to obtain the deformed Gaussian distribution , and the formula is: (52) Among them, is the new position, is the new rotation, is the new scale.

[0054] Step 4.3 introduces optical flow to solve the 4D GS overfitting problem; During the real-time operation of the SLAM system, due to a large amount of newly emerging data and the influence of noise in the data, the 4D GS will have an overfitting problem in the construction of the environmental map.

[0055] Specifically, the optical flow calculation method is adopted to capture the motion information of pixels in the time dimension, providing additional temporal consistency constraints for the model. The RAFT optical flow algorithm is used to calculate the pixel motion between adjacent timestamps. It includes: Step 4.3.1 Input images; For each pair of images at adjacent timestamps and , use the pre-trained RAFT to calculate the optical flow , and the pixel motion predicted by 4D GS is ; Step 4.3.2 Constrain the deformation of 3D Gaussian points through optical flow; Ensure that the dynamic part predicted by the model is consistent with the optical flow, and construct a loss function for optical flow constraint; The expression is as follows: (53) Where, represents the observed optical flow calculated from the RGB image through RAFT, represents the pixel-level motion prediction rendered based on the 4D GS deformation field; Step 4.3.3 Predict the change of Gaussian parameters based on the deformation field network of 4D GS; Adopt the deformation field network of 4D GS to predict the change of Gaussian parameters, and the parameters include position change , rotation change , scale change ; For the k th Gaussian, its deformed position is expressed by the following formula: (54) Where, represents the position of the initial Gaussian, represents the position offset predicted by the deformation field; Step 4.3.4 Project the Gaussian onto the image plane through a differentiable Splatting process to calculate the pixel-level motion; The expression is as follows: (55) Step 4.3.5 Introduce optical flow confidence to filter unreliable optical flow predictions; To reduce the noise of the optical flow algorithm, use the confidence map provided by the optical flow algorithm , and only apply the optical flow constraint loss in the regions with high confidence; The expression is as follows: (56) Where, Indicates the value of the optical flow confidence at position ; Step 4.4 Map update and optimization; Initialize the static 3D Gaussian distribution using the Structure-from-Motion (SfM) method. Only optimize the static 3D Gaussian for the first 3000 iterations, and then use the 3D Gaussian to replace the 4D Gaussian for image rendering, learn a reasonable initial 3D Gaussian distribution, separate the dynamic and static parts, reduce the pressure of large deformation learning, and avoid numerical instability problems when directly optimizing the deformation field network; For the construction of the loss function, use color loss to supervise the training process during reconstruction, and the optical flow constraint loss and the grid-based total variation loss are also applied; the expressions are as follows: (57) where and represent the weight parameters of the optical flow constraint loss and the grid-based total variation loss, used to balance the influence of different losses; Step 4.5 Loop detection; By extracting lidar and visual features and generating feature descriptors, the module can detect potential loop candidates; Use geometric verification and consistency checks to confirm the loop hypothesis; Feed the verified loop constraints back into the map optimization process to globally optimize the map structure.

[0056] Through the above loop detection steps, the robustness and efficiency of the SLAM system can be significantly improved, ensuring the map quality and positioning accuracy during long-term operation.

[0057] As described above, similar technical solutions can be derived based on the content of the solution given in combination with the drawings and descriptions. Any solution content that does not depart from the structure of the present invention still falls within the scope of the technical solutions of this application.

Claims

1. A reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing, characterized by: Including the following step, Step 1: Multi-source sensor input and data preprocessing; Receive point cloud data from the LiDAR, RGB image data from the depth camera, and IMU data for data preprocessing; Step 2: Select the sensor combination type; Based on the reinforcement learning sensor selection module, the reinforcement learning strategy network learns the environment status in real time, dynamically selects the most reliable sensor combination type and activates the sensor on demand; Step 3: Multi-source sensor data fusion processing; Based on the data fusion processing module, the sensor combination performs multi-source sensor data fusion processing; Step 4: Map update and optimization; Based on the map update and back-end optimization module, the environment map is updated and optimized by Gaussian map using the map data provided by the front-end data fusion processing module in combination with 4DGS, and error correction and map optimization are performed through loop detection.

2. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 1, characterized in that: The step 1 comprises the following steps, Step 1.1 Data input; The SLAM system synchronizes the point cloud data received from the LiDAR, the RGB image data from the depth camera, and the IMU data in time; Step 1.2: Incorporate time information; Introducing timestamps in data preprocessing ,Each lidar point cloud and image frame is timestamped; Step 1.3 Multi-sensor external parameter calibration; Use the Kalibr framework to calibrate the external parameters between sensors; Step 1.4: Initialize the reinforcement learning module; The reinforcement learning-based sensor selection module is initialized in the data preprocessing stage, and learns the strategy of selecting the optimal sensor combination under different scene conditions by analyzing historical data and environmental characteristics; In the initial state, a random exploration strategy is adopted.

3. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 1, characterized in that: The step 2 comprises the following steps, Step 2.1 state space design; Define the environment information that the reinforcement learning system can observe at each decision moment; Step 2.2 defines the action space formula; This includes defining actions for the system to switch sensor modes to optimize data acquisition and system performance; The SLAM system uses a hybrid discrete-continuous space to achieve refined control. The mode selection part is discrete, including three sensor-dominated modes: depth camera-IMU mode, lidar-depth camera mode, lidar-IMU-depth camera mode; The weight allocation part is continuous, including sensor gain coefficient and resource limitation parameters. The sensor gain coefficient is the weight of the lidar, depth camera and IMU, ranging from 0 to 1; the expression is as follows: (9) The resource limiting parameter is the lidar sampling rate, which ranges from 5 Hz to 40 Hz; the expression is as follows: (10) Step 2.3 Reward function module; Contains main rewards, auxiliary rewards and penalty items; Step 2.4 Network architecture design; Including the state encoder, as the basis of the entire network, which is responsible for converting multimodal sensor data into a unified feature representation so that the subsequent policy network and value network can make decisions and evaluations based on these features; The policy network is responsible for outputting specific action selections based on the encoded state features, including discrete action modes and continuous parameter adjustments; The value network is used to evaluate the pros and cons of the current state and action combination, providing learning signals for the policy network to help it optimize the decision-making process; Step 2.5 executes the training strategy; Through efficient data collection mechanism, intelligent exploration strategy and strict safety learning constraints, a comprehensive training framework is provided for the reinforcement learning sensor adaptive selection module; include, The data collection mechanism provides the system with experience data on interactions with the environment, which is the basis for learning and optimizing strategies; Exploration strategy, which determines how the system can effectively explore unknown or partially known environments to gain more information and experience; Safe learning constraints ensure that the system not only pursues high performance but also meets a series of safety and practicality requirements during the learning and decision-making process.

4. The method of reinforcement learning adaptive multimodal SLAM based on 4D Gaussian splashing according to claim 3, characterized in that: The step 2.1 includes: Geometric dynamic feature design, including LiDAR point cloud density gradient and dynamic object coverage; The density gradient of the lidar point cloud reflects the density change of the point cloud in space, helps the system perceive the geometric structure and potential obstacles of the environment, and quantifies the observation confidence of the lidar in the local area. The expression is as follows: (1) in, Indicates the number of lidar hit points within a unit voxel, represents the voxel volume, Indicates the normalized value range of reflection intensity; Calculating the dynamic object coverage rate represents the proportion and distribution of dynamic objects in the scene, so that the system can understand the dynamic characteristics of the environment in a timely manner. The expression is as follows: (2) in, represents the estimated velocity vector of the dynamic object, Represents the projection area of ​​the dynamic object detection box, Indicates the field of view area of ​​the sensor; Sensory reliability, including IMU confidence and radar-vision depth consistency. IMU confidence reflects the credibility of inertial measurement unit data and is calculated based on its noise level and bias stability. The higher the result value, the more reliable the IMU data. The expression is as follows: (3) in, represents the gyroscope bias estimation error, represents the standard deviation of accelerometer noise; Radar-vision depth consistency measures the degree of agreement between the depth data from the vision sensor and the radar. It is calculated by comparing the depth measurements of the two. The higher the consistency, the more coordinated the depth data from the two sensors. The expression is as follows: (4) in, represents the observed depth distribution histogram, represents the depth distribution histogram of the depth camera, It represents the intersection ratio of the effective areas of the two; System resource status monitoring, including real-time computing load and remaining energy budget. Real-time computing load reflects the current computing pressure of the system and helps the module decide whether to reduce the amount of sensor data processing or simplify computing tasks. The expression is as follows: (5) in, Indicates the video memory capacity currently used by the GPU. Indicates the total GPU memory capacity. Indicates the time taken to render a single frame. represents the expected frame period; The remaining energy budget indicates the energy reserve of the system, which enables the module to select the sensor combination or adjust the working mode based on the current remaining energy budget; the expression is as follows: (6) in, Indicates the current remaining battery level. Indicates the initial total power, Indicates the system running time. Indicates the power attenuation factor; Spatiotemporal context, including historical decision memory and scene category probability. Historical decision memory records the success rate and effect of past sensor selection decisions, helping the module learn from experience, avoid repeated mistakes, and improve decision efficiency. The expression is as follows: (7) in, Indicates the past Step action, represents the long short-term memory network; The scene category probability calculates the probability of belonging to low-texture scenes, limited-viewing-constrained scenes, and dynamic scenes based on the current environmental characteristics, so that the module can adjust the sensor selection strategy according to the scene characteristics; the expression is as follows: (8) in, express CLIP The joint image-text embedding vector extracted by the model, represents the learnable weight matrix, To normalize the distribution, the sum is 1.

5. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 3, characterized in that: The step 2.3 includes, The main rewards include positioning accuracy rewards and energy efficiency economic rewards, which are used to improve the accuracy of system positioning and energy efficiency respectively; the expressions are as follows: (11) (12) Where ATE represents the absolute trajectory error, represents the total instantaneous power consumption, Indicates the reward compensation item when the remaining power exceeds the threshold; Auxiliary rewards, including rendering quality rewards and policy stability rewards, are designed to improve the visual effect of scene reconstruction and ensure the continuity and reliability of the strategy; the expression is as follows: (13) (14) Among them, SSIM represents the structural similarity index, LPIPS represents the learning-perceptual image block similarity, represents the penalty for action mutation, Represents the smoothness constraint of policy network parameters; Penalty items, including fatal penalty errors and resource overload penalties, are used to avoid unreasonable sensor selection and excessive resource consumption, thereby guiding the system to learn the optimal sensor selection strategy and achieve comprehensive optimization of system performance; the expression is as follows: (15) (16) The comprehensive reward calculation formula is as follows: (17)。 6. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 1, characterized in that: The step 3 comprises the following steps, The reinforcement learning sensor selection module adaptively selects the most suitable sensor mode according to the current environmental conditions and system status and the influence of various parameters on the sensor mode selection; the weight function expression is as follows: + + + + + (21) in, , , , , and Respectively represent the laser radar confidence , dynamic object coverage , IMU confidence , radar vision consistency ,Computing load and energy remaining budget High weight; Step 3.1 Depth Camera-IMU Mode; When the reinforcement learning sensor selection module selects the weight function S≥ , To determine the threshold, the SLAM system selects the sensor combination of depth camera and IMU, with the depth camera as the main and IMU as the auxiliary; Step 3.2 LiDAR-Depth Camera Mode; When the reinforcement learning sensor selection module is based on the current environment characteristics, according to the weight function , To determine the threshold, the SLAM system selects a sensor combination of lidar and depth camera, with the depth camera as the main sensor and the lidar as the auxiliary sensor. Step 3.3 LiDAR-IMU-Depth Camera Multi-Sensor Balance Mode; When the reinforcement learning sensor selection module selects the current environment characteristics through the weight function according to the weight function , determine the sensor mode that the SLAM system chooses to use, which is a combination of lidar, depth camera, and IMU; Step 3.4 After the reinforcement learning sensor selection module executes the action and interacts with the environment, the historical decision memory is recorded through the reward function module, and the update of the entire policy network is optimized through value network evaluation.

7. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 6, characterized in that: The step 3.1 includes: Step 3.1.1 Data input and key frame selection; Prioritize frames that contain new scene structures or significant features as key frames to improve the integrity and efficiency of the map; Step 3.1.2 uses generalized ICP tracking for point cloud registration; Select RGB image of the frame and depth image , generate point cloud , where each point , and calculate the covariance matrix for each point ; Estimate the source point cloud of the current frame through the GICP algorithm With target map point cloud Relative pose transformation ; Step 3.1.3.IMU pre-integration; Fusion of IMU high-frequency measurements to generate inter-frame motion pre-integration as the initial guess of GICP, improves tracking efficiency, and combines with GICP algorithm to enhance system robustness and accuracy; Step 3.1.4: Update pre-integration; The GICP optimization results are used to correct the drift of IMU integral due to noise and bias over time.

8. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 6, characterized in that: The step 3.2 comprises the following steps, Step 3.2.1 Data input: images from depth camera and point cloud from LiDAR; The time-aligned LiDAR point cloud is converted into a depth image using the calibrated external signal for integration; the expression is as follows: (38) in, is the lidar point cloud, and are the rotation matrix and translation vector from the laser radar to the camera coordinate system, is the intrinsic matrix of the camera; Step 3.2.2 uses the incremental error minimization function; To ensure the exact correspondence between the plane and the point, the formula is as follows: (39) in, Represents a LiDAR point cloud A point in It is based on the current posture estimation from the previous moment to the world coordinate system. The result after iterations is is closest The Gaussian center of yes The normal vector of Yes The weight of is a regularization term used to enhance the stability and accuracy of the error function, taking into account the directional error between normal vectors; Introducing regularization terms To enhance the stability and accuracy of the error function and take into account the error in the normal direction; the expression is as follows: (40) in, is the normal of the current Gaussian distribution; Step 3.2.3 weight function calculation; The calculation steps of the weight function are as follows: a. Determine the Gaussian center within the local spherical region: Find all the nearest Gaussian distribution centers within The center of the ball. is the radius; b. Calculate the density function of Gaussian points, density function Calculated by the following formula: (41) in, is the reconstructed covariance matrix, by selecting the minimum variance along the normal direction and a larger variance in the vertical direction to build; c. Simplify the calculation of density function. In order to speed up the calculation, simplify the calculation of density function during the tracking process: (42) d. Consistency calculation, for each point , calculate the normal of the current Gaussian distribution With the local mean normal Consistency ; e. Complex texture calculation: For each image area corresponding to a radar point, calculate the local texture complexity; Project to the camera's pixel coordinate system to get its position in the image , intercepted with this as the center image blocks, where =16, convert the image block to grayscale and calculate the variance of pixel intensity: (43) in, is the pixel intensity, Image block mean; The variance is converted to a 0-1 texture weight through the sigmoid function: (44) in, is the scaling factor used to adjust the variance sensitivity; f. Final weight function, define the final weight function is the product of normal consistency, density function and texture complexity, that is, .

9. The method of reinforcement learning adaptive multimodal SLAM based on 4D Gaussian splashing according to claim 6, characterized in that: The step 3.3 comprises the following steps, Step 3.3.1 Data input and hardware synchronization; The LiDAR provides sparse but high-precision 3D point clouds, the depth camera captures the RGB texture of the scene, and the IMU outputs angular velocity and acceleration at high frequency for motion prediction, ensuring that the data timestamps of the three are strictly aligned to avoid timing drift; Step 3.3.2: select key frames through depth camera input; Select representative frames from continuous data streams as key frames to reduce redundant calculations; Step 3.3.3.IMU data is used for state propagation; According to the input IMU data and the state of the previous key frame, the current state is predicted by IMU pre-integration, which is used for state estimation and forward prediction of motion distortion removal; Step 3.3.4 dedistorts the lidar input; Using the continuous pose predicted by IMU, each lidar point Convert from the local coordinate system at the time of scanning to the global coordinate system; the expression is as follows: (45) in, Indicate point The acquisition timestamp, Indicates the value obtained through IMU interpolation The position of the moment.

10. The reinforcement learning adaptive multimodal SLAM method based on 4D Gaussian splashing according to claim 1, characterized in that: The step 4 comprises the following steps, Step 4.1 Sliding window maintenance; Step 4.2 4D Gaussian distribution; Step 4.3 introduces optical flow to solve the 4D GS overfitting problem; Step 4.4 Map update and optimization; Step 4.5 loop detection; By extracting lidar and visual features and generating feature descriptors, the module is able to detect potential loop closure candidates; Confirm loop closure assumptions using geometric verification and consistency checks; The verified loop constraints are fed back into the map optimization process to globally optimize the map structure.

Citation Information

Patent Citations

  • Indoor navigation method based on vision and radar information fusion and reinforcement learning

    CN116263335A

  • Dense vision SLAM (Simultaneous Localization and Mapping) method and system using three-dimensional Gaussian back-end representation

    CN117990088A

  • Method and system for reconstructing dense Gaussian map in dynamic environment

    CN118071873A

  • Three-dimensional reconstruction method and device of scene, electronic equipment and storage medium

    CN118840486A

  • Complex scene three-dimensional reconstruction method based on multi-source fusion

    CN119006700A

Cited By

  • Self-adaptive 4D Gaussian splashing high-precision three-dimensional reconstruction system and method

    CN121033280A

  • Adaptive 4D gaussian splatting high-precision three-dimensional reconstruction system and method

    CN121033280B

  • Resonant sensor reading adaptive optimization method and system based on reinforcement learning

    CN121388395A

  • Event camera image reconstruction method of multi-frame fusion network based on optical flow guidance

    CN121903863A

  • 4DGS and generative prior three-machine position state accurate sensing method and system

    CN122312636A