SLAM parameter adaptive method and system based on deep reinforcement learning

Through the SLAM parameter adaptive method of deep reinforcement learning, the SLAM system parameters are adjusted in real time, which solves the problem of incomplete positioning accuracy and map construction in complex environments in traditional SLAM systems, and improves the system's robustness and positioning accuracy.

CN120339571APending Publication Date: 2025-07-18HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455551.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The fixed parameters of traditional SLAM systems in complex unknown environments lead to reduced positioning accuracy and incomplete map construction, affecting path planning and security.

Method used

Adaptive SLAM parameter method based on deep reinforcement learning is adopted, and SLAM system parameters are adjusted in real time through DRL agents, and the SLAM system status and motion observation indicators are used to optimize keyframe windows and parameters, so as to improve the matching degree of scene information and decisions.

Benefits of technology

It significantly improves the robustness and positioning accuracy of the SLAM system in complex unknown environments, reduces the average trajectory error, and improves the system's adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339571A_ABST
    Figure CN120339571A_ABST
Patent Text Reader

Abstract

The invention discloses an SLAM parameter adaptive method and system based on deep reinforcement learning. A deep reinforcement learning technology is introduced into SLAM system parameter optimization. An intelligent agent based on a deep reinforcement learning algorithm is designed and trained, on the premise that normal operation of a bottom-layer SLAM method is not interfered, feature information of an SLAM system and an environment scene is collected in real time, and key parameters in the SLAM system are dynamically adjusted. The matching degree of the scene information and the system decision is obviously improved, and the robustness and the positioning precision of the SLAM system in a complex unknown environment are enhanced. Test results on different data sets show that compared with an SLAM system with fixed parameters, the method can effectively reduce the average trajectory error and greatly improve the performance, and the superiority and practicability of the method are proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of robotics, autonomous driving, and computer vision, and particularly relates to a method and system for SLAM parameter adaptation based on deep reinforcement learning. Background Art

[0002] SLAM (Simultaneous Localization and Mapping) technology is one of the core technologies in the fields of robotics, autonomous driving, and drones. It aims to construct an environmental map in real time through sensor data and determine its own position. It has been widely applied in indoor service robots such as hotels and shopping malls, outdoor unmanned delivery vehicles, autonomous driving systems, and AR / VR, etc., and provides a solid foundation for many intelligent applications.

[0003] Traditional SLAM systems usually rely on predefined parameter configurations, which perform well in expected scenarios. However, when facing complex and unknown scenarios, there is a problem of mismatch between the decisions obtained from fixed system parameters and the scenarios, which may lead to problems such as decreased positioning accuracy and incomplete map construction. This not only affects path planning but may also result in navigation failure, getting lost, and even safety accidents. To solve this problem, various solutions have been proposed in the prior art. For example, the multi-sensor fusion method improves the system's adaptability to the environment by integrating feature information from different sources. Semantic SLAM extracts high-level features through scene semantic segmentation to enhance the stability of localization and mapping. However, the former increases the cost of the SLAM system while not substantially solving the problem of decision mismatch between the SLAM system and the scenario. The latter heavily relies on initial semantic segmentation, and when the semantic segmentation is incorrect or inaccurate, the capabilities of all subsequent links will be affected.

[0004] Therefore, it is necessary to propose a method that can adaptively adjust the parameters of the SLAM system to adapt to the current environment. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention proposes a method and system for SLAM parameter adaptation based on deep reinforcement learning. Through the interaction between the agent and the environment, the parameters in the SLAM system are dynamically adjusted, improving the matching degree between the scene information and the decision, enhancing the adaptability of the SLAM system to complex and unknown environments, and improving the robustness and positioning accuracy of the system.

[0006] A SLAM parameter adaptation system based on deep reinforcement learning includes a SLAM system for real-time map construction and pose estimation, and a DRL agent for dynamically outputting parameters.

[0007] The DRL agent collects the SLAM system state and motion observation metrics at each time step t as inputs, and selects the most suitable system parameters for the SLAM system according to the inputs, using the Markov decision process and the tuple M = (s t , a t , r t , γ). Among them, s t represents the SLAM system state corresponding to time step t, including the image frame sequence collected by the camera, the current pose estimate, and other environment-related observation information. a t represents the action that the DRL agent can execute at time step t. γ represents the discount factor of the reward r, which is used to incorporate historical rewards into the current reward during training. The reward r t is the evaluation of the result obtained after the DRL agent executes the action a t in the state s t . The following reward function is designed around the pose estimation accuracy:

[0008] r t = r(s t , a t ) = (1 - λ imitation ) × (max(-1, λ threshold - e tran )) + (λ imitation )I

[0009] Among them, e tran represents the error between the aligned estimated pose and the true pose, λ threshld represents the expectation of e tran , and λ threshold - e tran ≥ -1. I represents the consistency reward between the agent's decision and the system's heuristic decision, and λ imitation represents the control coefficient.

[0010] The control policy π: s × a → R of the DRL agent uses the PPO algorithm. Its action input includes motion observation metrics and the state of the SLAM system, and the policy input is the true pose in the future time step. The action a to be taken in each state s is specified through the control policy π, and the probability trajectory τ ∼ π = {s0, a0, s1, a1, …, s T , a T-1} of the output state and action, and the expected sum V π (s i ) is expressed as:

[0011]

[0012] By maximizing the expected sum V π(s i ) value to train the DRL agent to understand the most suitable actions to execute in different states. Among them, the value of the discount factor γ gradually decreases as the time step t increases.

[0013] Preferably, the motion observation indicators include point cloud overlap, scene complexity, and pose change.

[0014] Preferably, the SLAM system state includes local map information, tracking state, positioning confidence, and all map information collected during operation.

[0015] A method for SLAM parameter self - adaptation based on deep reinforcement learning specifically includes the following steps:

[0016] Step 1: Construct the SLAM system to be used, construct the map in real - time and estimate the pose.

[0017] Step 2: Establish a DRL agent and construct a policy network of the PPO algorithm.

[0018] Step 3: Extract the optical flow change as frame information, the optical flow change trend, and the point cloud coincidence degree as map information from the key - frame window and the map in the SLAM system state s0, and input them into the DRL agent.

[0019] Step 4: The DRL agent outputs an action a0 through the policy network and returns it to the SLAM system. The actions include whether to retain the key frame and the size of the key - frame window. At t = 0, since the parameters of the DRL agent have not been determined, the output action a0 is random.

[0020] Step 5: The SLAM system operates on the key frame according to the output action a0 of the DRL agent, adjusts the size of the key - frame window, and then extracts information from the new state s1 and inputs it into the DRL agent.

[0021] Step 6: After the SLAM system executes the action a0, calculate the value of the reward r0 and evaluate the performance of the SLAM system before and after the action execution.

[0022] Step 7: Repeat steps 3 to 6. After T time steps, generate the probability trajectory τ~π={s0,a0,s1,a1,…,s T ,a T-1} of the states and actions output by the DRL agent. By maximizing the expected sum V π (s i ) value to train the DRL agent and modify the parameters of the DRL agent to complete one round of training.

[0023] Step 8: Use the trained DRL agent for the next round of training. Repeat Steps 3 to 7 for multiple rounds of training. Use the discount factor γ to control the proportion of historical rewards participating in the current training, so that the DRL agent maintains a balance between the learned patterns and the patterns being learned.

[0024] Step 9: Fix the parameters of the DRL agent after multiple rounds of training, embed it into the SLAM system, execute Steps 3 to 5, and output the actions corresponding to the current environment according to the system state.

[0025] The present invention has the following beneficial effects:

[0026] This method reduces the dependence on manually designed parameters, can dynamically generate appropriate parameters according to the changes in the scenario, thereby improving the matching degree of scenario information and decision-making, and enhancing its robustness and positioning accuracy in complex unknown environments. Description of the Drawings

[0027] Figure 1 It is a block diagram of a SLAM parameter adaptive system based on deep reinforcement learning;

[0028] Figure 2 It is a flowchart of SLAM parameter adaptation based on deep reinforcement learning. Detailed Embodiment

[0029] The following further explains and illustrates the present invention with reference to the accompanying drawings;

[0030] As Figure 1 shown, a SLAM parameter adaptive system based on deep reinforcement learning includes a SLAM system for real-time map construction and pose estimation and a DRL agent for dynamically outputting parameters.

[0031] The DRL agent collects the SLAM system state and motion observation metrics at each time step t as inputs, and selects the most suitable system parameters for the SLAM system according to the inputs. It directly utilizes various information provided inside the SLAM system and automatically outputs the parameters most suitable for the current scenario after training. Therefore, there is no need for additional special design for the traditional SLAM system, nor will it affect the original process of the SLAM system.

[0032] In this embodiment, the DPVO method proposed in the prior art is selected to construct the SLAM system to illustrate the specific process of the SLAM parameter adaptation method based on deep reinforcement learning:

[0033] Step 1: Use the DPVO method to construct the SLAM system to be used, and construct the map and estimate the pose in real time.

[0034] Step 2: Establish a DRL agent and construct a policy network for the PPO algorithm.

[0035] Step 3: As shown in Figure 2 , extract the optical flow change as frame information, the optical flow change trend, and the point cloud overlap as map information from the key frame window and the map in the SLAM system state s0, and input them into the DRL agent.

[0036] Step 4: The DRL agent uses the input information to observe the environment, thereby obtaining the characteristics of the current scene and the state of the current SLAM system. Then, it outputs an action a0 through the policy network and returns it to the SLAM system. The actions include whether to retain the key frame and the size of the key frame window. At t = 0, since the parameters of the DRL agent have not been determined, the output action a0 is random.

[0037] The policy network is a funnel-shaped network including 3 layers of MLP, which is beneficial to increasing the dimension of the input state information, thereby improving the expression ability of the model and the fitting ability for non-linear relationships. It can also determine the key dimension information and filter out redundant dimension information by reducing the dimension layer by layer, thereby improving the generalization ability of the model.

[0038] In the output actions, whether to retain the key frame is a comprehensive parameter decision, and the size of the key frame window is a specific parameter decision. Among them, the comprehensive parameter decision can integrate multiple related parameters into a specific behavior. On the one hand, it avoids the problem that the original parameters do not cover the scene completely, and on the other hand, it also reduces the dimension of the policy network output, which can reduce the training difficulty of the agent.

[0039] Step 5: The SLAM system operates on the key frame according to the output action a0 of the DRL agent, adjusts the size of the key frame window, and then extracts information from the new state s1 and inputs it into the DRL agent.

[0040] Step 6: After the SLAM system executes the action a0, calculate the value of the reward function r0 to evaluate the performance of the SLAM system before and after the action execution:

[0041] r t = r(s t , a t ) = (1 - λ imitation ) × (max(-1, λ threshold - e tran )) + (λ imitation )I

[0042] where e tran represents the error between the estimated pose and the true pose after alignment, λ threshold represents the expectation of e tran , λthreshold -e tran ≥ -1. I represents the consistency reward between the agent's decision and the system's heuristic decision, and λ imitation represents the control coefficient. In this embodiment, λ imitation is set to 0.01, λ threshold is set to 0.12, and I is set to 0.1.

[0043] Step 7: Repeat Steps 3 to 6. After T time steps, generate the probability trajectories τ~π = {s0, a0, s1, a1, …, s T , a T-1} of the state and action output by the DRL agent. Train the DRL agent by maximizing the value of the expected sum V π (s i ), modify the parameters of the DRL agent, and complete one round of training:

[0044]

[0045] Train the DRL agent by maximizing the value of the expected sum V π (s i ) so that it can learn the most suitable actions to perform in different states.

[0046] Step 8: Use the trained DRL agent in the previous step for the next round of training. Repeat Steps 3 to 7 for multiple rounds of training. And use the discount factor γ to control the proportion of historical rewards participating in the current training.

[0047] Step 9: Fix the parameters of the DRL agent after multiple rounds of training, embed it into the SLAM system, execute Steps 3 to 5, and output the actions corresponding to the current environment according to the system state.

[0048] In this embodiment, the synthetic dataset TartanAir with random data augmentation is used to train the DRL agent. The significance of using random data augmentation is to reduce the domain transfer with the real scenario and enhance the generalization ability of the policy network. To accelerate the training speed, different sequences are run on 10 concurrent DPVO instances to collect rewards at the same time. The initial value of the discount factor γ is set to 0.6 and gradually decreases as the time step t increases to balance historical rewards and current rewards.

[0049] To verify the improvement of this method on the performance of the SLAM system, tests are carried out with ORB - SLAM, DSO, and SVO on the real datasets EuRoC and TUM - RGBD, and the average trajectory error (ATE) is selected as the metric. Table 1 shows the test results on the EuRoC dataset, and Table 2 shows the test results on the fr1 dataset of TUM - RGBD:

[0050] Table 1

[0051]

[0052] Table 2

[0053] Method 360 fdesk desk2 floor plant room rpy teddy xyz Average ORB-SLAM - 0.017 0.210 - 0.034 - - - 0.009 - DSO 0.173 0.567 0.916 0.808 0.121 0.379 0.058 - 0.036 - DROID-VO 0.161 0.028 0.099 0.033 0.028 0.327 0.028 0.169 0.013 0.098 DPVO 0.135 0.038 0.048 0.040 0.036 0.394 0.034 0.064 0.012 0.089 Our method 0.138 0.032 0.052 0.040 0.030 0.347 0.034 0.052 0.012 0.082

[0054] According to the table data, the proposed method achieves the best average results on the EuRoC dataset. Moreover, compared with the fixed-parameter method of DPVO, the proposed method has achieved an improvement of more than 6%. On the fr1 dataset of TUM-RGBD, the proposed method has achieved an improvement of more than 7% compared with the fixed-parameter method of DPVO. The experimental results prove that the proposed method effectively improves the matching degree of scene information and decision-making, and improves the robustness and positioning accuracy of the SLAM system in complex unknown environments.

Claims

1. A SLAM parameter adaptive system based on deep reinforcement learning, characterized in that: It includes a SLAM system for real-time map construction and pose estimation and a DRL agent for dynamically outputting parameters. Using a Markov decision process and a tuple M = (s t , a t , r t , γ) to represent the DRL agent process, where s t represents the SLAM system state corresponding to time step t, including the sequence of image frames collected by the camera, the current pose estimate, and other environment-related observation information; a t represents the action executed by the DRL agent at time step t; γ represents the discount factor of the reward r, and r t represents the reward value at time step t: r t = r(s t , a t ) = (1 - λ imitation ) × (max(-1, λ threshold - e tran )) + (λ imitation )I Among them, e tran represents the error between the estimated pose after alignment and the true pose, and λ threshold represents the expectation of e tran , and λ threshold -e tran ≥ -1; I represents the consistency reward between the agent's decision and the system's heuristic decision, and λ imitation represents the control coefficient; The control policy π: s×a→R of the DRL agent uses the PPO algorithm. Its action input includes motion observation metrics and the state of the SLAM system, and the policy input is the true pose at future time steps. The action to be taken in each state is specified by the control policy π, and the probability trajectory τ~π = {s0, a0, s1, a1, …, s T , a T-1} of the output state and action, and the expected total sum V π (s i ) of the rewards under the discount factor γ is expressed as: By maximizing the expected sum V π (s i ) value, the DRL agent is trained to understand the most suitable actions to perform in different states; the value of the discount factor γ is set to gradually decrease as the time step t increases.

2. The SLAM parameter adaptive system based on deep reinforcement learning according to claim 1, characterized in that: The motion observation metrics include point cloud overlap, scene complexity, and pose change.

3. The SLAM parameter adaptive system based on deep reinforcement learning according to claim 1, wherein: The SLAM system state includes local map information, tracking state, localization confidence, and all map information collected during operation.

4. The SLAM parameter adaptive system based on deep reinforcement learning according to claim 1, characterized in that: The SLAM system is implemented using any one of DPVO, ORB-SLAM, DSO, and SVO.

5. The SLAM parameter adaptive system based on deep reinforcement learning according to claim 1, wherein: Set λ imitation = 0.01, λ threshold = 0.12, I = 0.

1.

6. The SLAM parameter adaptive system based on deep reinforcement learning according to claim 1, characterized in that: Set the initial value of the discount factor γ to 0.

6.

7. A method for SLAM parameter self - adaptation based on deep reinforcement learning, characterized in that: Specifically, it includes the following steps: Step 1: Construct a SLAM system to construct a map in real time and estimate the pose. Step 2: Establish a DRL agent and construct a policy network of the PPO algorithm. Step 3: Extract the optical flow change as frame information, the optical flow change trend, and the point cloud overlap as map information from the key frame window and the map in the SLAM system state s0 and input them into the DRL agent. Step 4: The DRL agent outputs an action a0 through the policy network and returns it to the SLAM system; the actions include whether to retain the key frame and the size of the key frame window. Step 5: The SLAM system operates on the key frame according to the output action a0 of the DRL agent, adjusts the size of the key frame window, then extracts information from the new state s1 and inputs it into the DRL agent. Step 6: After the SLAM system executes the action a0, calculate the value of the reward r0 and evaluate the performance of the SLAM system before and after the action execution; the reward function is designed as: r(s t ,a t )=(1-λ imitation )×(max(-1,λ threshold -e tran ))+(λ imitation )I where t represents the time step; e tran represents the pose error, and λ threshold represents the expectation of e tran and λ threshold -e tran ≥ -1; I represents the consistency reward, and λ imitation represents the control coefficient; Step 7. Repeat Step 3 to Step 6. After T time steps, generate the probability trajectory τ ∼ π = {s0, a0, s1, a1, …, s T , a r-1} of the states and actions output by the DRL agent. Train the DRL agent by maximizing the value of the expected sum V π (s i ). Modify the parameters of the DRL agent to complete one round of training: where γ represents the discount factor that decreases with the increase of time steps. Step 8: Use the trained DRL agent in the previous round for the next round of training, repeat steps 3 to 7 for multiple rounds of training. Step 9: Fix the parameters of the DRL agent after multiple rounds of training, embed it into the SLAM system, execute steps 3 to 5, and output the action corresponding to the current environment according to the system state.