Deep reinforcement learning type path planning method and device for indoor complex environment
Through deep reinforcement learning combined with sensor data and multi-objective reward function, the adaptability problem of traditional path planning algorithms in dynamic obstacle environments is solved, and the robot's autonomous and efficient navigation in complex environments is realized.
Patent Information
- Application Number
- CN202510746585.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Traditional path planning algorithms have poor adaptability when facing dynamic obstacles in complex indoor environments, and it is difficult to deal with rapid changes in obstacle location and state in real time, and lack adaptive learning capabilities, resulting in low robot navigation efficiency.
Deep reinforcement learning method is adopted, combined with radar and camera sensors to obtain environmental data, extract feature information through deep convolutional neural networks, and use near-end strategy optimization algorithm to build multi-objective reward functions to train the robot's path planning strategy in complex environments to achieve dynamic obstacle avoidance and autonomous navigation.
The robot can achieve autonomous and efficient navigation to the target point in a complex indoor environment, reducing the frequency of path re-planning, reducing the computing burden, and ensuring safety and smoothness.
Smart Images

Figure CN120333460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, deep learning, and path planning, and in particular, to a deep reinforcement learning-based path planning method and device for complex indoor environments. The entire path planning method and device are designed to achieve the autonomous navigation function of a robot in a complex indoor environment. Background Art
[0002] With the rapid development of technology, robots have been widely used in many fields, especially the service demand in indoor scenarios such as homes, offices, shopping malls, etc. is increasing. In these complex indoor environments, path planning has become one of the core problems for a robot to achieve efficient and safe autonomous navigation.
[0003] Traditional path planning algorithms are mainly graph search-based algorithms, such as the A* algorithm, Dijkstra algorithm, etc. They perform well in simple environments with static obstacles. However, when facing the actual dynamic obstacles in the indoor environment, such as moving pedestrians, suddenly appearing handling equipment, etc., the limitations of these traditional algorithms become prominent. On the one hand, it is difficult for them to respond to the rapid changes in the position and state of obstacles in real time; on the other hand, most traditional algorithms are based on pre-set rules and models, lacking the ability of adaptive learning for complex environments and being difficult to dynamically adjust strategies according to different scenarios.
[0004] Meanwhile, in recent years, sensor technology has made great progress. Radar and camera sensors can provide rich environmental perception data for robots. However, how to make full and efficient use of this massive perception data to assist path planning has become an urgent problem to be solved. The rise of deep reinforcement learning technology has brought new opportunities for robot path planning. It enables a robot to repeatedly explore and learn in a simulated environment, and continuously optimize its own behavior strategy according to the obtained reward feedback, thus having the ability to make autonomous decisions in complex dynamic environments. Compared with traditional methods, deep reinforcement learning can better adapt to the uncertainty of the environment, dynamically adjust path planning to avoid obstacles and efficiently approach the target. Developing a method and device that can make full use of advanced sensor technology, combine the advantages of deep reinforcement learning, and effectively solve the robot path planning problem in complex indoor environments has extremely high practical significance and value. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a deep reinforcement learning-based path planning method and device for indoor complex environments. In view of the poor adaptability of traditional algorithms to dynamic obstacles, deep reinforcement learning is introduced to endow the robot with real-time perception and obstacle avoidance capabilities, avoid frequent re-planning of paths, reduce the computational burden, and ensure its smooth and safe operation in dynamic scenarios. This method enables the robot to autonomously and efficiently navigate to the target point in a complex indoor environment.
[0006] To achieve the above object, the present invention provides the following technical solutions: A deep reinforcement learning-based path planning method for indoor complex environments, comprising the following steps:
[0007] S1. Obtain the perception data of the environment through radar and camera sensors, estimate and identify the pose state in the robot's body coordinate system, the spatial distribution of static obstacles, and the kinematic parameters of surrounding pedestrians, and generate, through multi-modal processing: a lidar information data map and a pedestrian pose data map;
[0008] S2. Use a deep convolutional neural network to perform multi-modal feature extraction on the lidar information data map and the pedestrian pose data map to generate feature information including environmental structure features, dynamic pedestrian motion features, and feasible region constraints;
[0009] S3. Construct a deep reinforcement learning training environment model, and perform network initialization on the deep reinforcement learning training environment model based on the proximal policy optimization (PPO) algorithm;
[0010] S4. Use the extracted feature information as the input of the partially observed state of the reinforcement learning, and train the optimal control strategy based on the multi-objective reward function and the deep reinforcement learning proximal policy optimization algorithm;
[0011] S5. Through the learned optimal control strategy, obtain the optimal action of the current observed state of the robot, and complete the dynamic obstacle avoidance and navigation tasks without stopping to achieve path planning and autonomous navigation in a complex indoor environment containing static and multiple dynamic obstacles.
[0012] As a preferred technical solution of the present invention, the specific steps of sensor data extraction and processing in step S1 are as follows:
[0013] S1.1. The lidar data is processed by multi-resolution rasterization, and the original point cloud data is converted into a two-dimensional occupancy grid map, that is, the lidar information data map;
[0014] S1.2. Convert the dynamic pedestrian motion parameters extracted by the depth camera into a two-dimensional position and velocity feature map in the local Cartesian coordinate system with the robot itself as the origin, that is, the pedestrian pose data map.
[0015] As a preferred technical solution of the present invention, the specific steps of the feature extraction of the local observation information in step S2 are as follows:
[0016] S2.1. Input the lidar information data map and the pose data map of the pedestrian into the convolutional neural network to further extract the environmental features and obtain the current partial observation result .
[0017] As a preferred technical solution of the present invention, the specific steps of the network initialization of the depth reinforcement learning training environment model based on the proximal policy optimization algorithm in step S3 are as follows:
[0018] S3.1. Regard the process of robot path planning and navigation as a partially observable Markov decision process. The POMDP model can be represented as a seven-tuple , where represents the state space, which contains all possible environmental states . represents the observation space, which contains all possible observations
[0019] . represents the action space, which contains all possible actions . is the state transition probability, which is the probability of transitioning from the observation state by executing the action to the observation state . represents the reward function, that is, the immediate reward obtained when executing the action in the observation state and transitioning to the observation state . represents the observation probability distribution. represents the discount factor, which is used to determine the current value of future rewards. The goal of the deep reinforcement learning network of the present invention is to maximize the expected discounted return. The objective function is defined as , and the specific formula is:
[0020]
[0021] where represents the discount coefficient, represents the expected value under the policy .
[0022] S3.2. Define the observation space as the relative position and speed of the obstacles in the local coordinate system of the robot;
[0023]
[0024] Among them, represents the relative position map of sub-goals during the training step . represents the data map of 10 scans of the lidar during the training step . represents the pose data map of the dynamic obstacle crowd during the training step .
[0025] S3.3. Define the action space The speed control signal for the robot includes the forward linear velocity and the turning angular velocity in the robot coordinate system;
[0026]
[0027] Among them, represents the forward linear velocity during the training step , with the unit of . represents the turning angular velocity during the training step , with the unit of .
[0028] S3.4. Build the Actor network of the PPO algorithm, which is a two-layer fully connected neural network. The activation function is the linear activation function. The Actor policy network takes the current observation state as the input, outputs a probability distribution with a dimension of 2, and outputs the action . The policy , where represents the parameters of the Actor network and can select the best action according to the current observation state. The random sampling process of the action satisfies the diagonal Gaussian probability distribution and is completed through the following formula:
[0029]
[0030] Among them, represents the probability density function, which gives the probability density at the position . represents the dimension of the random variable , that is, the dimension of the action space. represents a sampling result of an action . represents the mean vector of the normal distribution, is the mean of the th dimension, which determines the expected position of the action in this dimension. is the The standard deviation of a dimension, which represents the uncertainty or dispersion degree of the action in this dimension. , which is the square of the Mahalanobis distance. The smaller this term is, it indicates the closer it is to the mean. According to the properties of the exponential function, the probability density is greater. , which is a normalization term to ensure that the integral of the probability density function over the entire space is equal to 1.
[0031] Parameter is updated by maximizing the expected return, and the formula is:
[0032]
[0033] where, is the learning rate. is expressed as the error, which is the same as in the Critic network, that is, the difference between the actual reward plus the value estimate of the next observation state after discounting and the value estimate of the current observation state.
[0034] S3.5. Establish the Critic network of the PPO algorithm, which is also a two-layer fully connected layer neural network. The overall structure of the Critic network is basically the same as that of the Actor network, and the output dimension is 1, that is, the value function
[0035] , are the parameters of the Critic network. The activation function is the linear activation function. The input of the Critic network is the observation state information extracted from the experience pool. During each navigation process, the experience pool of PPO will store at most the step-by-step experience information, and then take samples from the experience pool. Input the extracted sample observation state information into the Critic network to evaluate the value function of the current policy, calculate the advantage function through the value function, and then update the Actor network backward. The of the present invention is set to 576, is set to 64.
[0036] Advantage function
[0037] The calculation formula is:
[0038]
[0039]
[0040]
[0041]
[0042]
[0043] Among them, is expressed as the advantage function. represents the discount factor, represents the decay coefficient of GAE, and the value is within . is the temporal difference error, The calculation formula of is:
[0044]
[0045] Among them, is the immediate reward. is the discount factor, which is set here to. is the value estimate of the next observation state. is the value function estimate of the current observation state.
[0046] Observation state value function Weight The weight of is updated by the gradient ascent method, and the policy improvement formula is:
[0047]
[0048] Among them, is the weight of the observation state value function. is the learning rate, which is used to control the update amplitude. is expressed as error. is the gradient of the observation state value function with respect to the weight .
[0049] S4. Take the extracted feature information as the input of the partially observable state of the reinforcement learning, and train the optimal control policy based on the multi-objective reward function and the proximal policy optimization algorithm of deep reinforcement learning.
[0050] As a preferred technical solution of the present invention, the multi-objective reward function of the proximal policy optimization PPO algorithm defined in step 4 is realized through the following sub-steps:
[0051] S4.1. The defined multi-objective reward function consists of four parts, namely: the reward for reaching the target, the reward for avoiding collisions, the reward for preventing the robot from oscillating, and the active avoidance reward based on speed plus steering control; that is:
[0052]
[0053] Among them, Represents the reward for reaching the target, Represents the reward for avoiding collisions, Represents the reward for preventing the robot from oscillating, Represents the active obstacle avoidance reward based on improved velocity obstacle theory for velocity control and lateral steering control.
[0054] S4.2. Define the reward for reaching the target It is:
[0055]
[0056] Among them, Represents the reward coefficient for reaching the target point, which is set to 18 here; Represents at the training step The distance from the current position of the robot to the target point; Represents the path point coefficient, which is set to 3 here; Represents the range tolerance for reaching the target point, which is set to a circle with a radius of 0.25 m; Represents the current number of training steps; Represents the maximum number of steps for a complete navigation.
[0057] S4.3. Define the reward for avoiding collisions It is:
[0058]
[0059] Among them, Represents the collision penalty coefficient, which is set to -15 here;
[0060] Is the minimum distance of non-zero values in the lidar scan, representing the distance between the robot and the nearest obstacle. Is the radius of the robot. Is the penalty coefficient for the scan distance, which is set to -0.2 here.
[0061] S4.4. Define the reward for avoiding collisions It is:
[0062] r o = r smooth × |θ t-1 - θ t |
[0063] Among them Represents the azimuth angle of the robot at step ; Represents the oscillation penalty coefficient, which is set to -0.02 here.
[0064] S4.5. Define a heuristic active avoidance reward function based on speed plus steering control according to the speed obstacle theory 。
[0065]
[0066] Where and represent weight coefficients, which are set here to , being 0.7. represents the optimal heading angle desired by the robot at training step . represents the maximum allowable deviation coefficient of the heading direction angle, which is set here to . represents the angle reward coefficient, which is set here to 0.6. is the optimal speed of the robot at training step , is the maximum deviation speed, which is set here to 0.1. represents the speed reward coefficient, which is set here to 0.8.
[0067] As a preferred technical solution of the present invention, the optimal control strategy is trained by the deep reinforcement learning algorithm in the step S5 , which further includes:
[0068] The deep reinforcement learning algorithm trains the control strategy model in the simulation environment GAZEBO under the Ubuntu 20.04 system . Starting from the start of training, this training method first determines whether the maximum number of training steps has been reached. If so, it evaluates and saves the best model and then ends; if not, it initializes the training environment, obtains the initial observation result and then executes the policy action , and then obtains the current observation result , and calculates the reward. Then it checks in turn whether the maximum number of iterations allowed in the current episode, the maximum number of collisions in the episode, and the target point have been reached. If the corresponding conditions are met, it resets the environment and updates the counter and enters the next training episode. If the target point has not been reached, it repeats the action steps.
[0069] After training, the performance and effectiveness of the trained strategy are tested in the environment. The test metrics include success rate, collision rate, average duration, and average path length. Success rate: The success rate refers to the proportion of samples in all test samples where the robot can reach the target point from the starting point without collision. Collision rate: The collision rate refers to the proportion of samples in all test samples where the robot collides with pedestrians. Average duration: The average duration refers to the average time taken by the robot to travel from the starting point to the end point in all test samples that successfully reach the target point. Average path length: The average path length refers to the average value of the actual path lengths traveled by the robot in all test samples that successfully reach the target point.
[0070] The technical solution of the present invention also relates to a planning and execution device for a robot in a complex indoor environment. The device includes a computer system, and the computer device includes the above-mentioned computer-readable storage medium and a deep reinforcement learning module.
[0071] Compared with the prior art, the present invention provides a deep reinforcement learning-based path planning method and device for a complex indoor environment, having the following beneficial effects:
[0072] In view of the poor adaptability of traditional algorithms to dynamic obstacles, the present invention introduces deep reinforcement learning to enable the robot to have real-time obstacle avoidance ability, avoid frequent re-planning of paths, reduce the computational burden, and ensure its smooth and safe operation in dynamic scenarios. This method mainly focuses on speed control and supplemented by steering control, aiming to guide the robot to find a path that can avoid obstacles and move forward quickly towards the target while maintaining the target direction. It can enable the robot to autonomously and efficiently navigate to the target point in a complex indoor environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 is a flowchart of a deep reinforcement learning-based path planning method for a complex indoor environment according to the present invention;
[0074] Figure 2 is the overall model diagram of a deep reinforcement learning-based path planning method for a complex indoor environment according to the present invention;
[0075] Figure 3 is a flowchart of the training steps of a deep reinforcement learning-based path planning method for a complex indoor environment according to the present invention;
[0076] Figure 4 is the model diagram of a deep reinforcement learning-based path planning device for a complex indoor environment according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0078] Refer to Figure 1 、 One A depth reinforcement learning-based path planning method for complex indoor environments, comprising the following steps:
[0079] S1. Obtain the perception data of the environment through radar and camera sensors, estimate and identify the pose state in the robot's body coordinate system, the spatial distribution of static obstacles, and the kinematic parameters of surrounding pedestrians, and generate, through modal processing: a lidar information data map and a pedestrian pose data map;
[0080] S2. Use a deep convolutional neural network to perform multi-modal feature extraction on the lidar information data map and the pedestrian pose data map to generate feature information including environmental structure features, dynamic pedestrian movement features, and feasible region constraints;
[0081] S3. Construct a deep reinforcement learning training environment model, and initialize the network of the deep reinforcement learning training environment model based on the proximal policy optimization (PPO) algorithm;
[0082] S4. Use the extracted feature information as the input of the partial observation state of the reinforcement learning, and train the optimal control strategy based on the multi-objective reward function and the deep reinforcement learning proximal policy optimization algorithm;
[0083] S5. Through the learned optimal control strategy, obtain the optimal action of the current observation state of the robot, and complete the dynamic obstacle avoidance and navigation tasks without stopping the vehicle, so as to realize path planning and autonomous navigation in a complex indoor environment containing static and multiple dynamic obstacles.
[0084] Refer to Figure 2 , construct a deep reinforcement learning training environment model, and initialize the network of the deep reinforcement learning training environment model based on the proximal policy optimization algorithm, which is realized through the following sub-steps:
[0085] Obtain the perception data of the environment through radar and camera sensors, estimate and identify the pose information of the robot itself, static obstacles, and surrounding pedestrians, and convert the depth camera and radar perception data into a network grid map with the same shape to represent the surrounding environment and the positions and motion states of pedestrians and obstacles in the local Cartesian coordinate system with the robot itself as the origin; then input the lidar information data map and the pose data map of pedestrians into a convolutional neural network to further extract environmental features and obtain the current partial observation results 。
[0086] Furthermore, regard the process of robot planning and navigation as a partially observable Markov decision process. The POMDP model can be represented as a seven-tuple where represents the state space, which contains all possible environmental states 。 represents the observation space, which contains all possible observations 。 represents the action space, which contains all possible actions 。 State transition probability, the probability of transitioning from the observation state executing the action and then transitioning to the observation state 。 represents the reward function, that is, the immediate reward obtained when executing the action in the observation state and transitioning to the observation state 。 。 represents the observation probability distribution represents the discount factor, which is used to determine the current value of future rewards. The goal of the deep reinforcement learning network of the present invention is to maximize the expected discounted return. The objective function is defined as , and the specific formula is
[0087]
[0088] where represents the discount coefficient, which is set to 0.98 here represents the expected value under the policy 。
[0089] S3.1. Define the observation space as the relative position and speed of obstacles in the local coordinate system of the robot
[0090]
[0091] where Indicates the training step Relative position map of sub-goals Indicates the training step Data map of 10 lidar scans Indicates the training step Pose data map of dynamic obstacle crowd
[0092] S3.2. Define the action space The speed control signal for the robot includes the forward linear velocity and the turning angular velocity in the robot coordinate system;
[0093]
[0094] Among them, Indicates the forward linear velocity at the training step , and the unit is . . Indicates the turning angular velocity at the training step , and the unit is . .
[0095] S3.3. Build the Actor network of the PPO algorithm, which is a two-layer fully connected neural network. The activation function is the linear activation function. The Actor policy network takes the current observation state as the input, and outputs a probability distribution with a dimension of 2 for the action , and the policy , where represents the parameters of the Actor network and can select the best action according to the current observation state. It satisfies the diagonal Gaussian probability distribution, and the specific formula is as follows:
[0096]
[0097] Among them, represents the probability density function, which gives the probability density at the position . represents the dimension of the random variable , that is, the dimension of the action space, which is set to 2 here. represents a sampling result of an action . represents the mean vector of the normal distribution, is the mean of the th dimension, which determines the expected position of the action in this dimension. is the standard deviation of the th dimension, which represents the uncertainty or dispersion degree of the action in this dimension. , which is the square of the Mahalanobis distance. The smaller this term is, the closer it is to the mean. According to the properties of the exponential function, the probability density is greater.
[0098] , which is a normalization term to ensure that the integral of the probability density function over the entire space is equal to 1.
[0099] Parameter is updated by maximizing the expected return, and the formula is:
[0100]
[0101] where is the learning rate. is expressed as the error, which is the same as in the Critic network, that is, the difference between the actual reward plus the value estimate of the next observation state after discounting and the value estimate of the current observation state.
[0102] S3.4. Establish the Critic network of the PPO algorithm, which is also a two-layer fully connected neural network. The overall structure of the Critic network is basically the same as that of the Actor network, and the output dimension is 1, that is, the value function
[0103] , are the parameters of the Critic network. The activation function is the linear activation function. The input of the Critic network is the observation state information extracted from the experience pool. During each navigation process, the experience pool of PPO will store at most the step-length experience information, and then take samples from the experience pool. The sampled observation state information is input into the Critic network to evaluate the value function of the current policy. The advantage function is calculated through the value function, and then the Actor network is updated backward. The of the present invention is set to 576, is set to 64.
[0104] The advantage function is calculated by the formula:
[0105]
[0106]
[0107]
[0108]
[0109]
[0110] Among them, is expressed as the advantage function. represents the discount factor, represents the attenuation coefficient of GAE, and its value is in . is the temporal difference error, and its calculation formula is:
[0111]
[0112] Among them, is the immediate reward. is the discount factor, which is set here to is the value estimation of the next observation state. is the value function estimation of the current observation state.
[0113] Observation state value function weight is updated through the gradient ascent method, and the policy improvement formula is:
[0114]
[0115] Among them, is the weight of the observation state value function. is the learning rate, which is used to control the update amplitude. is expressed as error. is the gradient of the observation state value function with respect to the weight .
[0116] For the multi-objective reward function of the Proximal Policy Optimization (PPO) algorithm, it is achieved through the following sub-steps:
[0117] S4.1. The defined multi-objective reward function consists of four parts, namely: the reward for reaching the target, the reward for avoiding collisions, the reward for preventing the robot from oscillating, and the active avoidance reward based on speed plus steering control; that is:
[0118]
[0119] Among them, represents the reward for reaching the target, represents the reward for avoiding collisions, represents the reward for preventing the robot from oscillating, represents the active obstacle avoidance reward based on improved speed obstacle theory speed control and lateral steering control.
[0120] S4.2. Define the reward for reaching the target is:
[0121]
[0122] wherein, represents the reward coefficient for reaching the target point, which is set to 18 here; represents at the training step the distance from the current position of the robot to the target point; represents the path point coefficient, which is set to 3 here; represents the range tolerance for reaching the target point, set as a circle with a radius of 0.25 m; represents the current number of training steps; represents the maximum number of steps for a complete navigation.
[0123] S4.3. Define the collision avoidance reward is:
[0124]
[0125] wherein, represents the collision penalty coefficient, which is set to -15 here; is the minimum distance of non-zero values in the lidar scan, representing the distance between the robot and the nearest obstacle. is the radius of the robot. is the penalty coefficient for the scanning distance, which is set to -0.2 here.
[0126] S4.4. Define the anti-oscillation reward is:
[0127] r o = r smooth × |θ t-1 - θ t |
[0128] where represents the azimuth angle of the robot at step ; represents the oscillation penalty coefficient, which is set to -0.02 here.
[0129] S4.5. Define an active avoidance reward function based on velocity plus steering control according to the velocity obstacle theory .
[0130]
[0131] where and represent the weight coefficients, which are set to , is 0.7. is represented as the training step the optimal heading angle expected by the robot. is represented as the maximum allowable deviation coefficient of the heading direction angle, which is set here to . is represented as the angle reward coefficient, which is set here to 0.6. is the robot training step the optimal speed at is the maximum deviation speed, which is set here to 0.1. is represented as the speed reward coefficient, which is set here to 0.8.
[0132] Referring to Figure 3 , the present invention proposes a training flowchart of a deep reinforcement learning-based path planning method for indoor complex environments. The simulation is to train the control policy model in the simulation environment GAZEBO under the Ubuntu 20.04 system , starting from the start of training, first judge whether the maximum number of training steps is reached. If it is reached, evaluate and save the best model and then end; if not, initialize the training environment and obtain the initial observation result and then execute the policy action , and then obtain the current observation result , and calculate the reward. Then check in turn whether the maximum number of iterations allowed in the current episode, the maximum number of collisions in the episode, and the target point are reached. If the corresponding conditions are reached, reset the environment and update the counter and enter the next training episode. If the target point is not reached, repeat the action steps.
[0133] When the reward value tends to be flat, the training ends, extract the model with the highest reward, and test the performance and effect of the trained policy model in the environment. In the configured simulation scenario, simulation experiments are carried out on a rule-based method and three learning-based navigation methods DRL, DRL_VO, and the navigation method used in the present invention, each with 500 random tests. The specific hardware configuration is 13th Gen Intel(R) Core(TM) i5-13490F 2.50 GHz with 16GBRAM. The simulation experiment is carried out in the simulated indoor area, and the comparison metrics include success rate, collision rate, average duration, and average path length. Success rate: The success rate refers to the proportion of samples in all test samples where the robot can reach the target point from the starting point without collision. Collision rate: The collision rate refers to the proportion of samples in all test samples where the robot collides with pedestrians. Average duration: The average duration refers to the average duration that the robot travels from the starting point to the end point in all test samples that successfully reach the target point, and the unit is 。Average path length: The average path length refers to the average value of the path lengths actually traveled by the robot among all the test samples that successfully reach the target point, with the unit of 。
[0134] The specific simulation experiment results are shown in the following table. OURS represents the use of the path planning method proposed in the present invention. The downward arrow indicates that the lower the value, the higher the model planning and navigation efficiency, and the upward arrow indicates that the higher the value, the higher the model planning and navigation efficiency. For a rule-based method and three learning-based navigation methods, the following conclusions can be drawn: In complex indoor environments with different numbers of pedestrians, the method proposed in the present invention demonstrates excellent performance, with relatively higher success rates, the lowest collision rates, and the shortest average path lengths in more complex environments, while maintaining a relatively low average travel time. This indicates that the method proposed in the present invention can not only effectively avoid collisions with pedestrians in path planning, but also optimize the navigation path and improve the navigation efficiency. Especially when the environment is more complex, the method proposed in the present invention has more significant advantages and can better cope with the challenges of complex environments to achieve efficient and safe navigation.
[0135] Table 1 Simulation Experiment Results
[0136]
[0137] Referring to Figure 4 , the present invention proposes a deep reinforcement learning-based path planning device for complex indoor environments. The device obtains the perception data of the environment through at least one sensor device, and then uses the deep neural network in the computer system device to process these data to generate an observation state space vector, providing the robot with an understanding and representation of the environment. Then, the trained deep reinforcement learning module outputs a control strategy based on these vectors, enabling the robot to make optimal decisions in complex environments. Finally, the chassis control actuator controls the steering angular velocity and linear velocity of the robot according to the output strategy, ensuring that the robot can move safely and efficiently along the expected planned path, thereby realizing autonomous navigation in complex indoor environments.
[0138] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A deep reinforcement learning-based path planning method for complex indoor environments, characterized in that: It includes the following steps: S1. Obtain the perception data of the environment through radar and camera sensors, estimate and identify the pose state in the robot's body coordinate system, the spatial distribution of static obstacles, and the kinematic parameters of surrounding pedestrians, and generate through modal processing: a lidar information data map and a pedestrian pose data map; S2. Use a deep convolutional neural network to perform multi-modal feature extraction on the lidar information data map and the pedestrian pose data map to generate feature information including environmental structure features, dynamic pedestrian movement features, and feasible region constraints; S3. Construct a deep reinforcement learning training environment model, and initialize the network of the deep reinforcement learning training environment model based on the Proximal Policy Optimization (PPO) algorithm; S4. Use the extracted feature information as the input of the partial observation state of the reinforcement learning, and train the optimal control strategy based on the multi-objective reward function and the deep reinforcement learning proximal policy optimization algorithm; S5. Through the learned optimal control strategy, obtain the optimal action of the current observation state of the robot, and complete the dynamic obstacle avoidance and navigation tasks without stopping, so as to achieve path planning and autonomous navigation in a complex indoor environment containing static and multiple dynamic obstacles.
2. The path planning method according to claim 1, characterized in that, The camera sensor is a depth camera, and the radar is a single-line lidar. Specifically, S1 includes: S1.
1. The lidar data is processed by multi-resolution rasterization to convert the original point cloud data into a two-dimensional occupancy grid map, that is, the lidar information data map; S1.
2. Convert the dynamic pedestrian movement parameters extracted by the depth camera into a two-dimensional position and velocity feature map in the local Cartesian coordinate system with the robot itself as the origin, that is, the pedestrian pose data map.
3. The path planning method according to claim 1, wherein Specifically, S2 includes: S2.
1. Input the lidar information data map and the pose data map of the pedestrian into the convolutional neural network to extract environmental features and obtain the current partial observation results .
4. The path planning method according to claim 1, wherein Specifically, S3 includes: S3.
1. Regard the process of robot planning and navigation as a partially observable Markov decision process, and use the POMDP model as the deep reinforcement learning training environment model. The model can be represented as a seven-tuple , where represents the state space, which contains all possible environmental states , represents the observation space, which contains all possible observations , represents the action space, which contains all possible actions , is the state transition probability, which is the probability of transitioning from the observation state executing the action to the observation state . represents the reward function, that is, the immediate reward obtained when executing the action in the observation state and transitioning to the observation state . , represents the observation probability distribution, represents the discount factor, which is used to determine the current value of future rewards; The goal of the deep reinforcement learning network is to maximize the expected discounted return, and its objective function is defined as , and the specific formula is: ; Among them, is expressed as a discount factor, represents the expected value under the policy ; S3.
2. Define the observation space which is the relative position and velocity of obstacles in the local coordinate system of the robot; ; Among them, represents the relative position coordinates of the sub-goal at the training step, represents the data graph of 10 scans of the lidar at the training step, represents the pose data graph of the dynamic obstacle crowd at the training step, including the position of the pedestrians in the Cartesian coordinate system with the robot as the origin and the speed of the pedestrians; S3.
3. Define the action space is the speed control signal of the robot, including the forward linear velocity and the turning angular velocity in the robot coordinate system; ; Among them, represents the forward linear velocity during the training step, represents the turning angular velocity during the training step; S3.
4. Establish the Actor network of the PPO algorithm, which is a two-layer fully connected neural network with a linear activation function. The Actor policy network takes the current observation state as input and outputs actions with a dimension of 2. The probability distribution of the forward speed and the angular velocity is output. The policy is , where represents the parameters of the Actor network and can select the best action according to the current observation state. The random sampling process of the action satisfies the diagonal Gaussian probability distribution and is completed through the following formula: ; denotes the probability density function, which gives the probability density at the position ; denotes the dimension of the random variable , that is, the dimension of the action space; denotes a sampled result of an action ; denotes the mean vector of the normal distribution, is the mean of the -th dimension, which determines the expected position of the action in this dimension; is the standard deviation of the -th dimension, which represents the uncertainty or dispersion degree of the action in this dimension; , which is the square of the Mahalanobis distance; is a normalization term used to ensure that the integral of the probability density function over the entire space is equal to 1; Parameter Updated by maximizing the expected return, with the formula: ; where, is the learning rate, is denoted as the error, which is the same as that in the Critic network; S3.
5. Establish the Critic network of the PPO algorithm, which is also a two-layer fully connected neural network. The overall structure of the Critic network is the same as that of the Actor network, and the output dimension is 1, that is, the value function , is the parameter of the Critic network, and the activation function is the linear activation function. The input of the Critic network is the observed state information extracted from the experience pool. During each navigation process, the experience pool of PPO will store at most the experience information of the previous step lengths, and then take samples from the experience pool, and input the observed state information of these samples extracted into the Critic network to evaluate the value function of the current policy . Calculate the advantage function through the value function , and then update the Actor network in reverse; Advantage function The calculation formula is as follows: ; Among them, is expressed as the advantage function, is expressed as the discount factor, is expressed as the decay coefficient of GAE, is the temporal difference error, and its calculation formula is: ; Among them, is the immediate reward, is the discount factor, is the value estimate of the next observation state, is the value function estimate of the current observation state; Observation state value function Weight is updated by the gradient ascent method, and the policy improvement formula is as follows: ; Among them, is the weight of the observation state value function, is the learning rate, which is used to control the amplitude of the update, is expressed as the error, and is the gradient of the observation state value function with respect to the weight 5. The path planning method according to claim 1, characterized in that In step S4, the multi-objective reward function in the Proximal Policy Optimization (PPO) algorithm is realized through the following sub-steps: S4.
1. The defined multi-objective reward function consists of four parts, namely: the reward for reaching the target, the reward for avoiding collisions, the reward for preventing the robot from oscillating, and the active avoidance reward based on speed plus steering control; that is: r t =r a +r c +r o +r v_θ Among them, r a represents the reward for reaching the target, r c represents the reward for avoiding collisions, r o represents the reward for preventing the robot from oscillating, r θ,v represents the active obstacle avoidance reward based on the improved velocity obstacle theory for velocity control and lateral steering control; S4.
2. Define the reward r for reaching the target a as follows: Among them, r arrival represents the reward coefficient for reaching the target point; represents the distance from the current position of the robot to the target point at training step t; r waypoint represents the waypoint coefficient; g r represents the range tolerance for reaching the target point; i t represents the current number of training steps; i m represents the maximum number of steps for a complete navigation; S4.
3. Define the collision avoidance reward r c as follows: Among them, r collision represents the collision penalty coefficient; d min is the minimum distance of non-zero values in lidar scanning, representing the distance between the robot and the nearest obstacle, r robot is the radius of the robot, r s is the penalty coefficient of the scanning distance; S4.
4. Define the anti-shock collision reward r o as follows: r o = r smooth × |θ t-1 - θ t | where θ t represents the azimuth angle of the robot at step t, and r smooth represents the oscillation penalty coefficient; S4.
5. Define an active avoidance reward function r based on speed and steering control according to the speed obstacle theory θ,v ; where ω1 and ω2 represent weight coefficients, denotes the optimal heading angle desired by the robot at training step t, θ m denotes the maximum allowable deviation coefficient for the heading direction angle, r θ denotes the angle reward coefficient, is the optimal speed of the robot at training step t, V m is the maximum deviation speed, r v denotes the speed reward coefficient.
6. The path planning method according to claim 1, wherein Training the optimal control strategy with the deep reinforcement learning algorithm in S5 Specifically including: The deep reinforcement learning algorithm trains a control policy model in the simulation environment GAZEBO under the Ubuntu 20.04 system Starting from the start of training, this training method first determines whether the maximum number of training steps has been reached. If so, it evaluates and saves the best model and then ends; if not, it initializes the training environment and obtains the initial observation results and then executes the policy action After that, it obtains the current observation results and calculates the reward; then it checks in turn whether the maximum number of iterations allowed in the current round, the maximum number of collisions in the round, and the target point have been reached. If the corresponding conditions are met, it resets the environment and updates the counter and enters the next training round. If the target point is not reached, it repeats the action steps 7. A deep reinforcement learning-based path planning device for complex indoor environments, characterized in that, It includes at least one sensor device for obtaining the perception data of the environment; a deep neural network for processing the perception data and generating an observation state space; a deep reinforcement learning module for training the control strategy; and a chassis control actuator for controlling the steering angle and speed of the robot according to the trained strategy. The path planning device is used to implement the path planning method described in any one of claims 1-6.
Citation Information
Patent Citations
Mobile robot obstacle avoidance method based on deep reinforcement learning
CN114237235A
Mobile robot end-to-end navigation method, system and equipment
CN114564004A
Robot navigation obstacle avoidance method and system based on reinforcement learning
CN115933675A
Indoor navigation method based on vision and radar information fusion and reinforcement learning
CN116263335A
Robot path collision avoidance planning method based on deep reinforcement learning in pedestrian environment
CN116360454A
Cited By
Tunnel construction robot path planning method and system based on deep reinforcement learning
CN121740035A
Deep reinforcement learning path planning method based on local environment driving
CN121977582A
Markov decision-based space debris laser ranging method
CN121978698A