A distributed multi-robot path planning method based on secure reinforcement learning
By applying a distributed method based on security reinforcement learning in multi-robot path planning, a global and local neural network is built, and combined with the optimal reciprocity collision avoidance mechanism, the problems of high computational complexity and local optimal solutions in a dynamic environment are solved, and efficient and secure multi-robot path planning is achieved.
Patent Information
- Application Number
- CN202510002849.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-02
AI Technical Summary
Traditional multi-robot path planning algorithms are difficult to deal with environmental changes and dynamic interactions between robots in real time in dynamic environments, resulting in high computational complexity and local optimal solutions.
A distributed multi-robot path planning method based on security reinforcement learning is adopted, and by building global and local neural networks, using distance cost maps and observation map information, combined with the optimal reciprocity collision avoidance mechanism, a conflict-free motion strategy is planned in real time.
It effectively reduces the risk of conflict and collision between robots, improves the efficiency and safety of path planning, and adapts to complex and dynamic environmental changes.
Smart Images

Figure CN119394314B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-robot path planning, and in particular relates to a distributed multi-robot path planning method based on secure reinforcement learning. Background Art
[0002] Multi-robot path planning is one of the core research topics in the field of robotics. It aims to plan efficient and safe collision-free paths for multiple robots in a shared environment, so as to ensure that they can reach their respective target locations smoothly in a short time. Traditional graph search algorithms, such as A* and Dijkstra algorithms, are widely used in single-robot path planning due to their high efficiency in deterministic environments. However, when these algorithms are directly applied to multi-robot systems, they often face many challenges. For example, as the number of robots increases, the dimension of the path planning problem increases significantly, resulting in a sharp increase in computational complexity, which makes it easy to fall into local optimal solutions. In addition, traditional algorithms have poor adaptability in dynamic environments and find it difficult to respond to environmental changes and dynamic interactions between robots in real time.
[0003] To overcome these difficulties, researchers have proposed a variety of heuristic path search methods. For example, the conflict search-based method significantly improves the effect of multi-robot path planning by pre-detecting and resolving potential path conflicts. However, centralized offline planning methods have poor robustness and face challenges in scalability and real-time performance when dealing with large-scale complex environments.
[0004] In practical applications, such as hotel service robots, they need to operate efficiently in dynamic and complex environments. Traditional path planning algorithms often rely on predefined maps and static environment assumptions, which are difficult to adapt to the frequently changing environment and robot flow within the hotel. Improving the efficiency of path planning while ensuring safe interaction between robots and the environment has become a hot topic and difficulty in the current research of multi-robot path planning methods. Summary of the invention
[0005] In view of this, the present invention provides a distributed multi-robot path planning method based on secure reinforcement learning, which can provide a distributed motion strategy for each robot in a dynamic environment of multiple robots, effectively reduce the risk of conflicts and collisions between robots, and improve operational efficiency.
[0006] In order to achieve the above object, the technical solution of the present invention is:
[0007] A distributed multi-robot path planning method based on secure reinforcement learning includes the following steps:
[0008] Robot distance cost map acquisition: Create a two-dimensional environment containing multiple robots, and obtain the corresponding distance cost map for each robot;
[0009] Neural network construction: Build a global neural network and A local neural network, which is the same as the global neural network, is input as local observation map information , the output includes the discrete action probability distribution , Status Value and blocking prediction ; The global neural network is used for strategy training, and the local neural network is used for distributed interaction between the robot and the environment;
[0010] Neural network training: For each robot, use its distance cost map to obtain local observation map information , input into the corresponding local neural network, and the input and output of the local neural network form training samples for training the global neural network;
[0011] Path planning: Use the trained global neural network to obtain distributed multi-robot path planning.
[0012] Furthermore, the process of obtaining the distance cost map of the robot of the present invention is as follows:
[0013] First, a two-dimensional environment containing multiple robots is created, and the distribution of obstacles is represented by a binary occupancy map;
[0014] Secondly, for each robot, a breadth-first search algorithm is used to obtain the corresponding distance cost map based on the occupancy map and the target position.
[0015] Furthermore, the local observation map information of the present invention The composition is:
[0016]
[0017] in, Indicates The measurement matrix of channels, is the local distance cost map, is the neighbor relative position map, A map of relative positions of neighbor targets.
[0018] Furthermore, the present invention at each time step Next, local observation map information As a state, the calculation is in the state Next action , the reward obtained based on the set reward function And get the state of the next moment , and record the end mark And the network outputs discrete action probability distribution Status Value and a label indicating whether blocking actually occurred , record seven-tuple Store to experience replay pool B middle.
[0019] Furthermore, the reward function set in the present invention is:
[0020]
[0021]
[0022]
[0023] in, represents the dense penalty for each move, which is the corresponding value of the center position of the local cost map; Indicates sparse penalties for special situations, including collision or out-of-bounds and blocking other robots;
[0024] Represents a function of state s and action a.
[0025] Furthermore, the present invention uses the experience replay pool B The data is used as sample data to perform global neural network training. The loss function of the global neural network is for:
[0026]
[0027]
[0028]
[0029]
[0030]
[0031] in, are the parameters of the current network, is the strategy clipping loss, For loss of value, is the entropy of the strategy, is the blocking prediction loss, , , , is the weight hyperparameter, is the mean function of all sampled time steps 𝑡, For the new strategy and the old strategy in the state Next select action The probability ratio of To measure the current action The advantage estimate over the average strategy advantage, is the clipping function that limits the upper and lower bounds of the variable, is a hyperparameter that limits the magnitude of policy updates, Status The future cumulative return, A label indicating whether blocking actually occurs. is the predicted blocking probability of the network, Status The probability distribution of discrete actions under .
[0032] Furthermore, the cumulative reward of the present invention The calculation method is:
[0033]
[0034] in, is the discount factor, is the terminal time step of the sequence.
[0035] Furthermore, the advantage estimation of the present invention The calculation method is:
[0036]
[0037]
[0038] in, is the discount factor, is the terminal time step of the sequence, is the time difference residual, is a control coefficient used to gradually decay the weight during multi-step estimation.
[0039] Furthermore, the specific process of path planning of the present invention is as follows:
[0040] Before each action, the robot will load the global neural network as a path planning model and call it to obtain the probability distribution of discrete actions in the current state, and select the action with the highest probability as the preferred discrete action. ;
[0041] Based on the optimal reciprocal collision avoidance mechanism, the collision avoidance speed sets of multiple robots are calculated, and the continuous action planning results with safety constraints are solved by linear programming. .
[0042] Furthermore, the continuous action planning result of the present invention The acquisition process is:
[0043] S111: Calculate the time of robot A The relative velocity set that will collide with robot B ;
[0044]
[0045] in, , are the positions of robots A and B respectively, , are the safety radius of robots A and B respectively, So Centered on is a circle with radius Indicates the time range for considering whether there will be a collision in the future;
[0046] S112: Based on the relative speed set , define a set of allowed velocities for each robot ,This set ensures that collisions can be avoided when considering other robots adopting the same algorithm;
[0047]
[0048] in, is the current speed of robots A and B, yes The external normal vector of the nearest point on the boundary, is from The vector to the closest point on the speed barrier boundary;
[0049] S113: Based on the allowed speed set , for each robot Choose a safe speed ;
[0050]
[0051] in, Yes all The intersection of is the preferred action of robot A.
[0052] Beneficial effects:
[0053] First, the present invention provides a distributed multi-robot path planning method based on secure reinforcement learning, innovatively designs the state space of the reinforcement learning decision model, which is a small-scale three-channel local observation map, including a local distance cost map, a neighbor relative position map, and a neighbor target relative position map, which represents the information of the robot and the environment in the form of a local map, effectively improving the generalization of scale and scene. Among them, based on the breadth-first search algorithm, a distance cost map is constructed, thereby introducing global information into local observations, providing an effective solution to the problem of being easily trapped in local optimality.
[0054] Second, the present invention provides a distributed multi-robot path planning method based on safe reinforcement learning, which combines the rule-based optimal reciprocal collision avoidance mechanism with the learning-based distributed decision model. In order to reduce the occurrence of collisions as much as possible, the soft constraints set by the reinforcement learning reward function are insufficient. Therefore, by calculating the reciprocal collision avoidance speed domain of the robot and solving the linear programming problem, the optimal speed closest to the output result of the reinforcement learning model is found. The introduction of this constraint ensures the safety and credibility of the path planning results. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0056] Figure 1 A flow chart of multi-robot path planning based on secure reinforcement learning provided by the present invention;
[0057] Figure 2 A structural diagram of the security reinforcement learning model provided by the present invention; DETAILED DESCRIPTION
[0058] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0059] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.
[0060] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0061] The present invention provides a distributed multi-robot path planning method based on secure reinforcement learning, which can provide a distributed motion strategy for each robot in a dynamic environment of multiple robots, effectively reduce the risk of conflicts and collisions between robots, and improve operation efficiency.
[0062] The present application embodiment provides a distributed multi-robot path planning method based on secure reinforcement learning, comprising the following steps:
[0063] Robot distance cost map acquisition: Create a two-dimensional environment containing multiple robots, and obtain the corresponding distance cost map for each robot;
[0064] Neural network construction: Build a global neural network and A local neural network, which is the same as the global neural network, is input as local observation map information , the output includes the discrete action probability distribution , Status Value and blocking prediction ; The global neural network is used for strategy training, and the local neural network is used for distributed interaction between the robot and the environment;
[0065] Neural network training: For each robot, use its distance cost map to obtain local observation map information , input into the corresponding local neural network, and the input and output of the local neural network form training samples for training the global neural network;
[0066] Path planning: Use the trained global neural network to obtain distributed multi-robot path planning.
[0067] The multi-service robot path planning problem in the hotel lobby map is used as an implementation example. , where obstacles are distributed, can be converted into a binary occupancy map The robot set is defined as ,in Indicates A robot, For each robot, its local observation range is , the available information includes: occupancy map , Current location , target location , own speed , the position set of other robots in the field of view , the target position set of neighbor robots , the speed set of neighbor robots .
[0068] like Figure 1-2 As shown, the embodiment of the present application is a distributed multi-robot path planning method based on secure reinforcement learning, comprising the following steps:
[0069] S1: Create a N The two-dimensional environment of the robot is converted into a binary occupancy map ,The distribution of obstacles is represented by a binary occupancy map. In this map, the position coordinates of each robot and its target position coordinates are initialized.
[0070] S2: For each robot, a breadth-first search is used to obtain a corresponding distance cost map based on the occupancy map and its target position. In the distance cost map of this embodiment, a point with a value of 1 indicates that there is an obstacle at that position, and a point with a value less than 1 indicates the relative distance from that position to the target position.
[0071] Specifically, the method of obtaining the distance cost map using breadth-first search is:
[0072] S21: Initialize the distance cost map. Set the maximum value of the grid where the obstacle is located , order mark ;
[0073] S22: Perform breadth-first search from the target position. Set the grid where the target position is located to 0, and traverse the surrounding four grids each time, executing: , non-obstruction grid . Until the complete graph is traversed, there is a maximum distance from the target position M ;
[0074] S23: Normalization processing. The grid where the obstacle is located is set to 1, and the other grids are set to the original value and M commerce;
[0075] Finally, a distance cost map is obtained.
[0076] S3: Initialize a global neural network and A local neural network, each of which has the same structure as the global neural network, and the input of the neural network is set to the local observation map , the output includes: discrete action probability distribution , Status Value and blocking prediction Among them, the global neural network is used for strategy training, and the local neural network is used for distributed interaction between the robot and the environment.
[0077] S4: assigning the parameters of the global neural network to each local neural network;
[0078] S5: In a distributed computing framework, each robot constructs a three-channel local observation map based on the distance cost map and the robot's position and target information within the field of view. , and input it into the corresponding local neural network, and select the action with the highest probability as the optimal action according to the discrete action probability distribution output by the network .
[0079] In this embodiment, the specific composition of the local observation map is:
[0080]
[0081]
[0082]
[0083]
[0084] in, Indicates The observation matrix of each channel has the size . is the local distance cost map, Indicates that the position ( ) relative distance to the target location; is the neighbor relative position map, Indicates that the position ( ) Whether there are other robots; is the relative position map of neighbor targets, Indicates that the position ( ) is the target position of other robots. is the current position of the robot, is the position index of the center point of the local map in the map, is the distance cost map obtained in step S2, is the position set of other robots, is the target position set of neighbor robots. is the discriminant item. When the positions of other robots Located in the observation matrix ( ) is 1, otherwise it is 0; As a discriminant item, when the target position of the neighbor robot Located in the observation matrix ( ) is 1, otherwise it is 0.
[0085] In addition, the robot can perform actions The space is: moving in eight directions: up, down, left, right, upper left, lower left, upper right, lower right, and staying still, a total of 9 actions.
[0086] S6: Calculated by the set reward function in state Next action Corresponding rewards And get the next moment state , and record the end mark And the network outputs discrete action probability distribution Status Value and a label indicating whether blocking actually occurred Then in the state The next execution action is trained in this way, and the seven-tuple Store to experience replay pool B middle.
[0087] The reward function designed in this embodiment is specifically:
[0088]
[0089]
[0090]
[0091] in, represents the dense penalty for each move, which is the corresponding value of the center position of the local cost map; Indicates sparse penalties for special situations, including collision or out-of-bounds and blocking other robots;
[0092] Represents a function about state s and action a, that is, giving corresponding reward feedback r according to the result of executing action a in state s (such as whether the target point is reached or whether there is a collision). In the specific interaction process, the state s of each step t Next, perform action a t , will reach a state (whether it reaches the target point, whether it collides, etc.), and according to the reward function R, rt The value of .
[0093] S7: Replay from the experience pool in batches B The data are sampled as training samples, and the global neural network is trained by combining the PPO algorithm and supervised learning, and the loss function of the global neural network is calculated.
[0094] The current network parameters of the global neural network are , the input is , the output is ,
[0095] , the loss function of the global neural network is calculated as:
[0096]
[0097]
[0098]
[0099]
[0100]
[0101] in, are the parameters of the current network, is the strategy clipping loss, For loss of value, is the entropy of the strategy, is the blocking prediction loss, , , , is the weight hyperparameter, A label indicating whether blocking actually occurs. is the predicted blocking probability of the network, Status The future cumulative return, To measure the current action The advantage estimate over the average strategy advantage, For the new strategy and the old strategy in the state Next select action The probability ratio of is a hyperparameter that limits the magnitude of policy updates, is the mean function of all sampled time steps 𝑡.
[0102] Cumulative Return and advantage estimate The calculation method is:
[0103]
[0104]
[0105]
[0106] in, is the discount factor, is the terminal time step of the sequence, is the time difference residual, is a control coefficient used to gradually decay the weight during multi-step estimation.
[0107] The role of is to allow the strategy to better choose the best action in the current state, while avoiding drastic changes in the strategy; The role of is to accurately estimate the value of each state; The role of is to prevent the strategy from converging to the local optimal solution too early, thereby improving the exploration ability; The role of is to more accurately predict the blocking risk through supervised learning and improve the feature extraction ability of the model.
[0108] S8: Determine whether the current number of iterations has reached the upper limit T If yes, go to step S9; if no, return to step S4;
[0109] S9: After training, the global neural network is output as a distributed multi-robot path planning model;
[0110] S10: In the execution phase, before each action, the robot will call the path planning model to obtain the discrete action probability distribution under the current state and select the action with the highest probability as the preferred discrete action. .
[0111] S11: Based on the optimal reciprocal collision avoidance mechanism, the collision avoidance speed set is calculated among multiple robots, and the continuous action planning result with safety constraints is solved by linear programming. .
[0112] The specific steps to obtain the planning results with safety constraints are:
[0113] S111: Calculate the time of robot A The relative velocity set that will collide with robot B :
[0114]
[0115] in, , are the positions of robots A and B respectively, , are the safety radius of robots A and B respectively, So Centered on A circle with a radius of .
[0116] S112: Define a set of allowed speeds for each robot , which ensures that collisions are avoided when considering other robots using the same algorithm:
[0117]
[0118]
[0119] in, Using the current speed of robots A and B, yes The external normal vector of the nearest point on the boundary, is from The vector to the closest point on the speed barrier boundary.
[0120] S113: Each robot The safe speed is selected by solving the following optimization problem :
[0121]
[0122]
[0123] in, Yes all The intersection of is the maximum speed allowed by the robot, is the preferred action obtained by the robot in step S10.
[0124] In order to further illustrate the effectiveness of the provided method, the path planning method provided by the present invention is tested on 10 robots running simultaneously. The simulation experiment was conducted on a two-dimensional map. In 100 rounds and a total of 8,000 decision steps, there was no collision between the agents, and they all reached the end successfully.
[0125] Based on the above experiments, the multi-robot path planning method based on secure reinforcement learning provided by the present invention can provide each robot with a distributed conflict-free motion strategy in a multi-robot dynamic environment, thereby ensuring the safety and reliability of path planning.
[0126] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A distributed multi-robot path planning method based on secure reinforcement learning, characterized in that: The following steps are involved: Robot distance cost map acquisition: Create a two-dimensional environment containing multiple robots, and obtain the corresponding distance cost map for each robot; Neural network construction: Build a global neural network and A local neural network, which is the same as the global neural network, has local observation map information as input and outputs discrete action probability distribution, state value and blocking prediction; the global neural network is used for strategy training, and the local neural network is used for distributed interaction between the robot and the environment; Neural network training: For each robot, use its distance cost map to obtain local observation map information , input into the corresponding local neural network, and the input and output of the local neural network form training samples for training the global neural network; Path planning: Use the trained global neural network to obtain distributed multi-robot path planning; At each time step Next, local observation map information As a state, the calculation is in the state Next action , the reward obtained based on the set reward function And get the state of the next moment , and record the end mark And the network outputs discrete action probability distribution Status Value and a label indicating whether blocking actually occurred , record seven-tuple Store to experience replay pool B middle; The reward function set is: in, represents the dense penalty for each move, which is the corresponding value of the center position of the local cost map; Indicates sparse penalties for special situations, including collision or out-of-bounds and blocking other robots; Represents a function of state s and action a; Experience Replay Pool B The data is used as sample data to perform global neural network training. The loss function of the global neural network is for: in, are the parameters of the current network, is the strategy cutting loss, For loss of value, is the entropy of the strategy, is the blocking prediction loss, , , , is the weight hyperparameter, is the mean function of all sampled time steps 𝑡, For the new strategy and the old strategy in the state Next select action The probability ratio of To measure the current action The advantage estimate over the average strategy advantage, is the clipping function that limits the upper and lower bounds of the variable, is a hyperparameter that limits the magnitude of policy updates, Status The future cumulative return, A label indicating whether blocking actually occurs. is the predicted blocking probability of the network, Status The probability distribution of discrete actions under .
2. The distributed multi-robot path planning method based on secure reinforcement learning according to claim 1 is characterized in that: The process of obtaining the robot distance cost map is as follows: First, a two-dimensional environment containing multiple robots is created, and the distribution of obstacles is represented by a binary occupancy map; Secondly, for each robot, a breadth-first search algorithm is used to obtain the corresponding distance cost map based on the occupancy map and the target position.
3. The distributed multi-robot path planning method based on secure reinforcement learning according to claim 1 is characterized in that: The local observation map information The composition is: in, is the local distance cost map, is the neighbor relative position map, A map of relative positions of neighbor targets.
4. The distributed multi-robot path planning method based on secure reinforcement learning according to claim 1 is characterized in that: The cumulative return The calculation method is: in, is the discount factor, is the terminal time step of the sequence.
5. The distributed multi-robot path planning method based on secure reinforcement learning according to claim 1 is characterized in that: The advantage estimate The calculation method is: in, is the discount factor, is the terminal time step of the sequence, is the time difference residual, is a control coefficient used to gradually decay the weight during multi-step estimation.
6. The distributed multi-robot path planning method based on secure reinforcement learning according to claim 1 is characterized in that: The specific process of path planning is as follows: Before each action, the robot will load the global neural network as a path planning model and call it to obtain the probability distribution of discrete actions in the current state, and select the action with the highest probability as the preferred discrete action. ; Based on the optimal reciprocal collision avoidance mechanism, the collision avoidance speed sets of multiple robots are calculated, and the continuous action planning results with safety constraints are solved by linear programming. .
7. The distributed multi-robot path planning method based on secure reinforcement learning according to claim 6 is characterized in that: The continuous action planning result The acquisition process is: S111: Calculate the time of robot A The relative velocity set that will collide with robot B ; in, , are the positions of robots A and B respectively, , are the safety radius of robots A and B respectively, So Centered on is a circle with radius Indicates the time range for considering whether there will be a collision in the future; S112: Based on the relative speed set , define a set of allowed velocities for each robot ,This set ensures that collisions can be avoided when considering other robots adopting the same algorithm; in, is the current speed of robots A and B, yes The external normal vector of the nearest point on the boundary, is from The vector to the closest point on the speed barrier boundary; S113: Based on the allowed speed set , for each robot Choose a safe speed ; in, Yes all The intersection of is the preferred action of robot A.
Citation Information
Patent Citations
Motion planning method based on machine learning in complex environment
CN116551703A