Robot Path Collision Avoidance Planning Method Based on Deep Reinforcement Learning in Pedestrian Environment
By adopting a deep reinforcement learning method in intelligent mobile robot navigation, combining codec network and pedestrian prediction deep neural network, predicting dynamic pedestrians and planning collision avoidance paths, the problems of slow training convergence and insufficient pedestrian recognition in navigation are solved, and the safety and efficiency of navigation are improved.
Patent Information
- Application Number
- CN202310437715.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-04-18
AI Technical Summary
The prior art has problems in the navigation of intelligent mobile robots that training convergence is too slow and dynamic pedestrians are insufficient, resulting in possible collisions when navigating in pedestrian environments.
A robot path collision avoidance planning method based on deep reinforcement learning is proposed, using codecs and pedestrian prediction deep neural networks to predict dynamic pedestrian dynamics and quickly plan collision avoidance paths. This method generates pedestrian trajectory data sets through multi-step agent stage reward function, codec network and deep learning network, combining three collision avoidance algorithms: RVO, ORCA, and SFM, extracts pedestrian motion characteristics, estimates the robot's state-action value pair, and outputs the optimal action.
It effectively improves the robot's obstacle avoidance success rate and navigation efficiency in pedestrian environments, solves the problems of short-sightedness and occlusion in navigation in dynamic environments, and improves the safety and timeliness of obstacle avoidance behaviors.
Smart Images

Figure CN116360454B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and mobile robot navigation, and particularly relates to a robot path collision avoidance planning method based on deep reinforcement learning in a pedestrian environment. Background Art
[0002] With the relatively rapid development of the robot industry in recent years, intelligent mobile robots are increasingly used in the fields of transportation, logistics, social services, and emergency rescue. An intelligent mobile robot can rely on sensors carried by itself to perceive and understand the external environment, make real-time decisions according to task requirements, perform closed-loop control, and operate in an autonomous or semi-autonomous manner, with certain self-learning and adaptation capabilities in known or unknown environments. For an intelligent mobile robot, its most fundamental technologies are navigation and obstacle avoidance. Navigation is an important issue that needs to be solved for an intelligent mobile robot to achieve path planning, which refers to the process by which a mobile robot perceives the environment and its own state through sensors and learning, and realizes autonomous movement towards a target in an environment with obstacles.
[0003] Robot path planning refers to a robot perceiving the surrounding environment through various sensors and autonomously searching for a collision-free path from a starting point to an ending point. Traditional navigation and obstacle avoidance methods mainly include search-based methods, sampling-based methods, artificial potential field methods, etc. Search-based methods mainly include the Dijkstra algorithm and the A* algorithm. The Dijkstra algorithm adopts a greedy mode, which solves the shortest path problem from a single node to another node in a directed graph. Its main feature is that the next node selected in each iteration is the nearest child node of the current node. In order to ensure that the finally searched path is the shortest, in each iteration process, the shortest path between the starting node and all traversed points needs to be updated, and the algorithm ends when the target point is covered within the search range. The A* algorithm is a heuristic search algorithm, that is, a heuristic search rule is established during the search process to measure the distance relationship between the real-time search position and the target position, so that the search direction preferentially faces the direction where the target point is located, ultimately achieving the effect of improving the search efficiency. Sampling-based methods include the probabilistic roadmap method, the rapidly-exploring random tree method, etc. The probabilistic roadmap method means randomly sampling in the pathfinding space to form a roadmap and search for a path. The rapidly-exploring random tree method is to grow a path tree in the pathfinding space through sampled points to form a path. The basic idea of the artificial potential field method originates from the concept of "field" in physics. The algorithm simulates the process of a robot moving in the environment as moving in an abstract artificial "field".
[0004] Reinforcement learning is a type of learning that maps from environmental states to actions, with the goal of enabling an agent to obtain the maximum cumulative reward during the interaction with the environment. The Markov decision process can be used to model RL problems. In robot navigation technology, the factors that have a significant impact on the navigation effect are mainly divided into two parts: the environmental perception part and the decision-making and control part. Deep learning and reinforcement learning can respectively solve the environmental perception reasoning and decision-making and control problems in navigation. In recent years, deep reinforcement learning has achieved remarkable results in navigation decision-making problems.
[0005] However, in reinforcement learning, the setting of the reward function is the key to driving the agent to learn the optimal policy. An inappropriate reward function may cause the agent to learn an inefficient or even incorrect policy. And the existing algorithms use a single-step reward function and the inherent trial-and-error nature of reinforcement learning, which easily leads to too high variance of the Q value. The high-variance Q value will affect the convergence speed of the training return function.
[0006] In the application scenarios of intelligent mobile robots, it is necessary to consider the situation of surrounding pedestrians. In highly dynamic and partially observable situations, occlusion phenomena are very common. Existing crowd navigation methods often cannot predict human movement trajectories and assume that complete environmental knowledge is provided. When deployed in the real world, these algorithms only consider avoiding detected or observed humans. Therefore, when an occluded human suddenly appears on the robot's path, a collision may occur.
[0007] All of the above pose higher requirements for the navigation and obstacle avoidance of intelligent mobile robots. The required intelligent mobile robot needs to have a certain ability to adapt to unknown environments and be able to navigate in a pedestrian environment. Summary of the Invention
[0008] To solve the problems of slow convergence in existing technology during training and insufficient recognition of dynamic pedestrians, the present invention proposes a robot path collision avoidance planning method based on deep reinforcement learning in a pedestrian environment. This method is based on an encoder-decoder and a pedestrian prediction deep neural network, which can predict the dynamics of occluded pedestrians in a dynamic pedestrian environment and quickly plan an effective collision avoidance path, solving the short-sighted and unsafe navigation problems in robot navigation.
[0009] The technical solutions adopted by the present invention to solve its technical problems are as follows:
[0010] A robot path collision avoidance planning method based on deep reinforcement learning in a pedestrian environment, comprising the following steps:
[0011] S1: Define the state space with the states of the robot and humans at the current moment; according to the state space, model the mutual observations between the robot and pedestrians as a partially observable Markov tuple;
[0012] S2: Set the multi-step proxy stage reward function, including penalties for the robot's collision with dynamic obstacles, constraints on the distance between the robot and dynamic obstacles, penalties for the distance from the target point, and time costs;
[0013] S3: According to the action space of pedestrians, randomly generate a pedestrian trajectory dataset composed of pedestrian states based on the collision algorithm;
[0014] S4: Input the pedestrian trajectory dataset D into the encoder-decoder network to extract pedestrian motion features;
[0015] S5: According to the pedestrian motion features, the current state and observation of the robot obtained from the partially observable Markov tuple, use the deep learning network to estimate the state-action value pair of the robot;
[0016] S6: Use the action neural network to output the optimal action selected by the robot in the current state according to the estimated value of the state-action value pair of the robot, use the evaluation neural network to score the state-action of the robot, and use reinforcement learning for iterative training. Combine the multi-step proxy stage reward function to update and optimize the parameters of the encoder-decoder network, deep learning network, action neural network, and evaluation neural network;
[0017] S7: In the process of path collision avoidance planning, first use the trained encoder-decoder network to encode the pedestrian state to obtain pedestrian motion features; then according to the pedestrian motion features, the current state and observation of the robot obtained from the partially observable Markov tuple, use the trained deep learning network to estimate the state-action value pair of the robot; finally, output the optimal action in the current state through the trained action neural network to achieve safe navigation in a dynamic pedestrian environment.
[0018] Furthermore, the state space is represented as:
[0019] s = [d g , v pref , v x , v y , r]
[0020] h i = [p x , p y , v x , v y , r i , d i , r i + r]
[0021] Among them, s represents the robot state, d g represents the distance between the robot and the target point, v pref represents the preferred speed of the robot, vx , v y represents the speed of the robot in the x-axis and y-axis directions, and r represents the radius of the space occupied by the robot; h i represents the state of the i-th pedestrian, p x , p y represents the position of the pedestrian, v x , v y represents the speed of the pedestrian in the x-axis and y-axis directions, r i represents the radius of the space occupied by the i-th pedestrian, d i represents the distance between the robot and the i-th pedestrian, r i +r represents the minimum safety distance between the robot and the person.
[0022] Furthermore, the Markov tuple is represented as (S, A, P, R, Ω, O, γ), where S represents the state space, A represents the action space, P represents the transition state, R represents the reward, Ω represents the observation probability distribution, O represents the relationship mapping from the state space to the observation space, and γ represents the discount rate.
[0023] Furthermore, the multi-step agent stage reward function is as follows:
[0024]
[0025] where N represents the total number of steps of the multi-step agent stage reward function, r t represents the reward at the t-th step, T represents the total number of decisions in the navigation process, γ represents the discount rate, and k represents the current step number.
[0026] Furthermore, the r t is as follows:
[0027]
[0028] where t represents the current time, represents the joint state of the robot at time t in the robot navigation, a t represents the action at time t, d t is the distance between the robot and the pedestrian at time t. When the distance between the robot and the pedestrian is less than 0, a collision occurs and a penalty is imposed; when the distance between the robot and the pedestrian is less than 0.15, a penalty is imposed according to the degree of approach; when the robot reaches the target position within the specified time, a reward is given, and the time taken to reach the destination is negatively correlated with the reward.
[0029] Furthermore, the pedestrian trajectory dataset D in step S3 is generated by three collision avoidance algorithms: RVO, ORCA, and SFM.
[0030] Further, the codec network consists of an encoder and a decoder. The input of the encoder is the pedestrian state in the pedestrian trajectory dataset, and the output of the decoder is the pedestrian motion feature F = {f t , f t-1 , f t-2}, where f t represents the motion feature of the pedestrian at time t.
[0031] Further, the deep neural network is successively composed of a multi-layer perceptron network, a softmax activation function, a multi-layer perceptron network, and a fully connected layer; the input of the deep neural network is the current state s t of the robot and the observation o t , as well as the pedestrian motion feature F, and the output of the deep neural network is the estimated value of the state-action value pair of the robot.
[0032] The beneficial effects of the present invention are mainly reflected in:
[0033] (1) The present invention combines and uses three collision avoidance algorithms, namely RVO, ORCA, and SFM, randomly generates pedestrian data, strengthens the robustness of the data, and solves the problem that a single model is difficult to generalize to a complex dynamic pedestrian environment.
[0034] (2) The present invention uses a multi-step proxy stage reward function, reduces the variance of the value function, and speeds up the training convergence speed.
[0035] (3) Through the encoder network structure, the present invention uses the observed pedestrian behavior as an additional sensor measurement to estimate the position of the occluded pedestrian, and integrates this mechanism into the deep reinforcement learning, effectively improving the obstacle avoidance success rate and navigation efficiency of the robot in the pedestrian environment, solving the short-sightedness and occlusion problems of the robot navigation in the dynamic environment, and improving the safety and timeliness of the obstacle avoidance behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the robot path collision avoidance planning method based on deep reinforcement learning in the pedestrian environment of the present invention.
[0037] Figure 2 is a structural diagram of the model training network of the present invention.
[0038] Figure 3 is the simulation result when the robot starts for 2.5 s.
[0039] Figure 4 is the simulation result when the robot starts for 7.75 s. DETAILED DESCRIPTION OF THE INVENTION
[0040] The present invention will be further described below in conjunction with the drawings and some cases.
[0041] The robot path collision avoidance planning method based on deep reinforcement learning in a pedestrian environment of the present invention has a flow block diagram as Figure 1 shown, including defining a state space, a partially observable Markov tuple, setting a multi-step reward function considering the influence of future multi-step robot operations on the current operation, randomly generating a pedestrian trajectory data set D, establishing a local perception map of the robot according to this data set and initializing the state space of the robot, using a multi-layer encoder-decoder network to extract pedestrian motion features, training an action-evaluation network through reinforcement learning, and predicting pedestrian trajectories in a simulation scenario to prove that the obstacle avoidance success rate of the present invention is high, the navigation efficiency is high, and the obstacle avoidance effect is good.
[0042] As Figure 1-2 shown, the specific implementation steps are as follows:
[0043] Step 1, define the state space based on the states of the robot and the person at the current moment:
[0044] s = [d g , v pref , v x , v y , r] (1)
[0045] h i = [p x , p y , v x , v y , r i , d i , r i + r] (2)
[0046] Among them, d g represents the distance between the robot and the target point, v pref represents the preferred speed of the robot, v x , v y represents the speed of the machine. Among them, s represents the robot state, d g represents the distance between the robot and the target point, v pref represents the preferred speed of the robot, v x , v y represents the speed of the robot in the x-axis and y-axis directions, r represents the radius of the space occupied by the robot; h i represents the state of the i-th pedestrian, p x , p y represents the position of the pedestrian, v x , v y represents the speed of the pedestrian in the x-axis and y-axis directions, r i represents the radius of the space occupied by the i-th pedestrian, d i represents the distance between the robot and the i-th pedestrian, ri +r represents the minimum safe distance between the robot and the human.
[0047] Model the mutual observation between the robot and the pedestrian as a partially observable Markov tuple (S, A, P, R, Ω, O, γ), where S represents the state space, A represents the action space, P represents the transition state, R represents the reward, Ω represents the observation probability distribution, O represents the relationship mapping from the state space to the observation space, and γ represents the discount rate. At each time step, given the observation o ∈ O, the robot selects an action a ∈ A. According to the Markov assumption, the next state S′ is only determined by the current state S; the observation probability distribution Ω(o, s′, a) and the state transition P(s, a, s′) are determined by the conditional probability.
[0048] Step 2: Set the multi-step proxy stage reward function, including the penalty for the robot colliding with dynamic obstacles, the constraint on the distance between the robot and dynamic obstacles, the penalty for the distance from the target point, and the time cost;
[0049] In this embodiment, the multi-step proxy stage reward function is defined as follows:
[0050]
[0051] where N represents the total number of steps of the multi-step proxy stage reward function, r i represents the reward at the t-th step, T represents the total number of decisions during the navigation process, γ represents the discount rate, and k represents the current number of steps. Take the multi-step value as N, and this N-step trajectory comes from the experience stored in the temporary replay buffer. For the training segment of T steps, the buffer is a moving window of size T from the initial state s0 to the termination state s T It considers the influence of the robot's action selection in the future multi-steps on the current reward value, makes the robot's actions converge within the safe action interval, reduces the variance of the value function, and speeds up the training convergence speed.
[0052] where r t is expressed as follows:
[0053]
[0054] where t represents the current time, represents the joint state of the robot at time t during navigation, a t represents the action at time t, d tis the distance between the robot and the pedestrian at time t. In the reward function, it includes the penalty for collision with dynamic obstacles, the constraint on the distance between the robot and dynamic obstacles, the penalty for the distance from the target point, and the time cost. When the distance between the robot and the pedestrian is less than 0, a collision occurs and a penalty is imposed, that is, the penalty for collision with dynamic obstacles; when the distance between the robot and the pedestrian is less than 0.15, a penalty is imposed according to the degree of approach, that is, the constraint on the distance between the robot and dynamic obstacles. When the robot reaches the target position within the specified time, it is rewarded, and the time taken to reach the destination is negatively correlated with the reward, that is, the penalty for the distance from the target point and the time cost.
[0055] Step 3: According to the action space of the pedestrian, randomly generate a pedestrian trajectory dataset D based on the collision algorithm; in a specific implementation of the present invention, in the action space of the pedestrian, three collision avoidance algorithms, namely RVO, ORCA, and SFM, are used to randomly generate a pedestrian motion trajectory dataset D in a ratio of 1:1:1. The pedestrian motion trajectory dataset consists of several pedestrian states h i =[p x , p y , v x , v y , r i , d i , r i +r].
[0056] Step 4: Generate a surrounding grid map of the robot to capture the positions of surrounding dynamic pedestrians.
[0057] Step 5: Build an encoder-decoder structure, where the encoder and decoder. In this embodiment, the encoder and decoder adopt a three-layer structure. Among them, each encoder layer includes a feed-forward network and a self-attention module, and each decoder includes a self-attention module, an encoder-decoder attention module, and a feed-forward network. The self-attention module is used to learn the relationship between the current behavior of the pedestrian and the previous behavior, and the encoder-decoder attention module is used to learn the relationship between the current behavior of the pedestrian and the pedestrian motion feature F = {f t , f t-1 , f t-2} to be encoded. As Figure 2 shown, the input of the encoder is the pedestrian state in the pedestrian trajectory dataset, and the output of the decoder is the pedestrian motion feature F = {f t , f t-1 , f t-2}, where f t represents the motion feature of the pedestrian at time t.
[0058] Step 6: Use a deep learning network to estimate the state-action value pair of the robot.
[0059] The current state s of the robott With the observation o t , and the pedestrian motion feature F are input into the deep neural network to estimate the state-action value pair of the robot. In this embodiment, the deep neural network structure is successively composed of 3 multi-layer perceptron networks, a softmax activation function, 2 multi-layer perceptron networks, and 1 fully connected layer.
[0060] Step 7: Use the action neural network in the PPO algorithm to output the optimal action selected by the robot in the current state, use the evaluation neural network in the PPO algorithm to score the state-action of the robot, and select the action with the largest reward for update. Using the method of curriculum learning, gradually increase the number of pedestrians in the environment, change the pedestrian speed, etc. Iteratively train, update and optimize the network parameters, and select the optimal motion trajectory.
[0061] In a specific implementation of the present invention, a simulation environment is built as GYM, the scene is a square space with a side length of 12 meters, the initial positions of pedestrians are randomly distributed in this space, and the number of pedestrians is 8 dynamic pedestrians. In this environment, use the trained reinforcement learning network to select the optimal collision avoidance path to complete the safe navigation of the robot.
[0062] The simulation result of the robot path planning process is as Figure 3 , Figure 4 shown, where the square unit is the robot, and the circular units numbered 0-7 represent pedestrians. The legend in the upper left corner represents the instantaneous speed magnitudes of pedestrians numbered 0-7 respectively, the arrow represents the speed direction, and the pentagram is the destination. The black area is the blind area scanned by the lidar during the movement of the robot.
[0063] As Figure 3 can be seen, at 2.5 s after the robot starts, pedestrian No. 2 and pedestrian No. 5 are completely blocked by pedestrian No. 3 and pedestrian No. 7 and are in the blind area of the robot's vision. At the same time, the position of pedestrian No. 6 causes partial occlusion of pedestrian No. 0. Figure 4 In, when moving to 7.75 s, there is still an occlusion relationship between the two. The present invention can output the pedestrian feature vector F = {f i , f i-1 , f i-2} through the encoder-decoder, and the multi-step reward function, quickly capture the dynamics of pedestrians, select the optimal behavior action of the current robot, and avoid collisions. Experimental simulations prove that the present invention can solve the obstacle avoidance and path planning problems of robots in the actual pedestrian environment. And it proves the safety of navigation and the effectiveness of obstacle avoidance in a multi-pedestrian environment.
[0064] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept and is for illustrative purposes only. The protection scope of the present invention should not be regarded as limited to the specific forms stated in these embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept of the present invention.
Claims
1. A method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment, characterized in that It includes the following steps: S1: Define the state space based on the states of the robot and the human at the current moment; According to the state space, model the mutual observations between the robot and the pedestrian as a partially observable Markov tuple; S2: Set the multi-step proxy stage reward function, including penalties for the robot's collision with dynamic obstacles, constraints on the distance between the robot and dynamic obstacles, penalties for the distance to the target point, and time costs; The multi-step proxy stage reward function is as follows: where N represents the total number of steps of the multi-step proxy stage reward function, r t represents the reward at the t-th step, T represents the total number of decisions in the navigation process, γ represents the discount rate, and k represents the current number of steps; S3: Based on the action space of the pedestrian, randomly generate a pedestrian trajectory dataset composed of pedestrian states using a collision algorithm; S4: Input the pedestrian trajectory dataset into the encoder-decoder network to extract pedestrian motion features; S5: According to the pedestrian motion features, the current state and observations of the robot obtained from the partially observable Markov tuple, use a deep learning network to estimate the state-action value pair of the robot; S6: Use the action neural network to output the optimal action selected by the robot in the current state based on the estimated value of the state-action value pair of the robot, use the evaluation neural network to score the state-action of the robot, and use reinforcement learning for iterative training. Combine the multi-step proxy stage reward function to update and optimize the parameters of the encoder-decoder network, deep learning network, action neural network, and evaluation neural network; S7: During the robot path collision avoidance planning process, obtain the surrounding environment data. First, use the trained encoder-decoder network to encode the pedestrian state to obtain pedestrian motion features; then, according to the pedestrian motion features, the current state and observations of the robot obtained from the partially observable Markov tuple, use the trained deep learning network to estimate the state-action value pair of the robot; finally, output the optimal action in the current state through the trained action neural network to achieve safe navigation in a dynamic pedestrian environment.
2. The method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment according to claim 1, characterized in that: The state space is represented as: s = [d g , v pref , v x , v y , r] h i = [p x , p y , v x , v y , r i , d i , r i + r] Among them, s represents the robot state, and d g represents the distance between the robot and the target point, v pref represents the preferred speed of the robot, v x , v y represents the speed of the robot in the x-axis and y-axis directions, and r represents the radius of the space occupied by the robot; h i represents the state of the i-th pedestrian, p x , p y represents the position of the pedestrian, v x , v y represents the speed of the pedestrian in the x-axis and y-axis directions, r i represents the radius of the space occupied by the i-th pedestrian, d i represents the distance between the robot and the i-th pedestrian, r i + r represents the minimum safety distance between the robot and the person.
3. The method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment according to claim 1, characterized in that: The Markov tuple is represented as (S, A, P, R, Ω, O, γ), where S represents the state space, A represents the action space, P represents the transition state, R represents the return reward, Ω represents the observation probability distribution, O represents the relationship mapping from the state space to the observation space, and γ represents the discount rate.
4. The method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment according to claim 1, characterized in that: The described r t is as follows: where t represents the current time, represents the joint state at time t in robot navigation, a t represents the action at time t, d t is the distance between the robot and the pedestrian at time t. When the distance between the two is less than 0, a collision occurs and a penalty is imposed; when the distance between the robot and the pedestrian is less than 0.15, a penalty is imposed according to the degree of approach; when the robot reaches the target position within the specified time, a reward is given, and the time taken to reach the destination is negatively correlated with the reward.
5. The method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment according to claim 1, characterized in that: The pedestrian trajectory dataset D mentioned in step S3 is generated by three collision avoidance algorithms: RVO, ORCA, and SFM.
6. The method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment according to claim 1, characterized in that: The described codec network consists of an encoder and a decoder. The input of the encoder is the pedestrian state in the pedestrian trajectory dataset, and the output of the decoder is the pedestrian motion feature F = {f t , f t-1 , f t-2}, where f t represents the motion feature of the pedestrian at time t.
7. The method for robot path collision avoidance planning based on deep reinforcement learning in a pedestrian environment according to claim 1, characterized in that: The deep neural network is successively composed of a multi-layer perceptron network, a softmax activation function, a multi-layer perceptron network, and a fully connected layer; the input of the deep neural network is the current state s of the robot t and the observation o t , as well as the pedestrian motion feature F, and the output of the deep neural network is the estimated value of the state-action value pair of the robot