Multi-robot path planning method and system under attention mechanism
Through the multi-robot path planning method under the attention mechanism, CBS algorithm, convolutional neural network and graph neural network are used to solve the problem of inaccurate information transmission in multi-robot path planning, efficient collaboration and collision-free path planning between robots are achieved, and the stability and efficiency of the system are improved.
Patent Information
- Application Number
- CN202510174630.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-08
AI Technical Summary
In multi-robot path planning, the accuracy and timeliness of information transmission are difficult to ensure. Especially under the decentralized method, the limited bandwidth and memory limitations between robots lead to instability in communication, environmental interference affects communication reliability, and it is difficult to achieve efficient and collision-free path planning.
The multi-robot path planning method under the attention mechanism is adopted, data is collected through the CBS algorithm, map perception features are extracted using the convolutional neural network, and the graph convolutional neural network transmits features, and action scores are decoded through multi-layer perceptrons, combining the multi-head attention mechanism and graph neural network dynamic weighted aggregation of neighbor information, and setting up anti-collision and deadlock cancellation mechanisms.
It improves the timeliness of collaboration among robots and the accuracy of path planning, avoids collisions and deadlocks, ensures stable operation of the system, and improves overall efficiency and safety.
Smart Images

Figure CN120274775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a multi-robot path planning method and system under an attention mechanism. Background Art
[0002] Compared with the path planning of a single robot, multi-robot path planning is to lead each robot from the starting position to the target position, and during this period, the paths of each robot cannot collide at the same moment. In recent years, great progress has been made in multi-robot path planning, but finding a method with both good performance and efficiency is still a great challenge. Multi-robot mobile path planning methods are usually divided into centralized and decentralized. The centralized method mainly relies on a central controller to centrally schedule each robot. As the scale of the system becomes larger and larger, the consumption of computing resources increases and the requirements for the central controller become higher and higher. As the scale of the system expands, the decentralized method becomes more suitable than the centralized method. Different from the centralized method that relies on a central controller, the decentralized method mainly relies on communication to coordinate the paths of robots. But this does not mean that this method has no disadvantages. Due to the limited bandwidth between robots, large data cannot be transmitted, and the relatively small memory cannot store a large amount of data. In addition, it is easily interfered by the environment and cannot ensure reliable and continuous communication. Summary of the Invention
[0003] The purpose of the present invention is to overcome the problem that at a certain specific moment, a robot only has partial information of the entire system, such as only the information of adjacent robots, etc. The model extracts features through a convolutional neural network, and the graph convolutional neural network transmits these features among multiple robots, and finally a multi-layer perceptron is used to make action decisions. However, under this model, the accuracy and timeliness of information transmission cannot be determined, that is, it cannot be determined whether the information is transmitted to a specific robot at a specific moment. The present invention provides a multi-robot path planning method and system under an attention mechanism, adding an attention mechanism to solve the problem of when the information is transmitted to whom.
[0004] The purpose of the present invention is achieved by the following technical solutions:
[0005] A multi-robot path planning method under an attention mechanism, comprising:
[0006] S1. Through the CBS algorithm and based on the map perception features around the current position of the robot obtained, input the obtained map perception features into the graph attention network for extraction;
[0007] S2. Through the attention network, perform weighted summation according to the features of neighbor nodes to obtain an updated feature representation;
[0008] S3. Compute different attention relationships in parallel through the multi-head attention mechanism to generate multiple attention outputs, and finally splice them to obtain the final features;
[0009] S4. Use the final node features to decode through a multi-layer perceptron to generate the action scores for each robot;
[0010] S5. Use the softmax function to convert the action scores into a probability distribution, and select the next action according to the probability.
[0011] Further, in the step S1, the dataset is collected through the CBS algorithm. The CBS algorithm consists of two search processes: a high-level search process and a low-level search process. The low-level search process is responsible for searching an effective path for each robot. The high-level search process is responsible for checking path conflicts and selecting the branch with the smallest cost value to re-perform the low-level path search until the high-level search process finds an effective path; the path is optimal for each robot.
[0012] Further, in the step S1, the map perception features are extracted through a convolutional neural network. The map perception features include the obstacles within a specific range and the position information of other robots.
[0013] Further, in the step S2, the graph neural network is used to summarize the information of its own node and adjacent nodes to update the characteristics expressed in the adjacency matrix, realizing the communication between robot nodes.
[0014] Further, in the step S2, the network under the attention mechanism can dynamically weight and aggregate the neighbor information according to the relationship between each robot and its neighbors, and can better handle complex and irregular map structures, thereby improving the accuracy of information transmission.
[0015] Further, through the mutual relationship between robots and environmental information, and by changing the values of the adjacency matrix of the graph neural network, the communication between robots is realized.
[0016] Further, in the step S3, the multi-head attention mechanism can compute multiple different attention patterns in parallel, so as to capture different interaction relationships between multiple robots.
[0017] Further, in the step S5, after obtaining the final output features, a multi-layer perceptron is set to decode the final output features. After obtaining the decoding scores, they are converted into probabilities by using the classification probability function softmax. Finally, the action with the highest probability is selected, and an anti-collision strategy and a deadlock release mechanism are set.
[0018] A multi-robot path planning system under the attention mechanism is provided, and this system is used to execute a multi-robot path planning method under the attention mechanism.
[0019] There is provided a readable storage medium storing instructions, characterized in that when the instructions are executed by one or more processors of a machine, the processors are caused to execute a multi-robot path planning method under an attention mechanism.
[0020] The beneficial effects of the present invention are as follows:
[0021] (1) The robots cooperate well, can adjust the path in time when encountering environmental changes or approaching each other, avoid collisions, and improve the overall efficiency;
[0022] (2) The anti-collision and deadlock release mechanisms ensure the safety of the robots, can solve problems in time when problems occur, and ensure the continuous operation of the system;
[0023] (3) The model training has a high accuracy rate, the robots make reasonable decisions, reduce wrong actions, and can operate stably in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A flowchart of a multi-robot path planning system under an attention mechanism provided for the embodiment;
[0025] Figure 2 A structural diagram of a multi-robot path planning system under an attention mechanism provided for the embodiment;
[0026] Figure 3 A graph of the model training accuracy rate of a multi-robot path planning method system under an attention mechanism provided for the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0028] Embodiment 1
[0029] Refer to Figure 1 , and there is provided a multi-robot path planning method under an attention mechanism, the steps including:
[0030] S1. Through the CBS algorithm and based on the map perception features around the current position of the robot obtained, the obtained map perception features are input into the graph attention network for extraction;
[0031] S2. Through the attention network, the updated feature representation is obtained by weighted summation of the features of the neighbor nodes;
[0032] S3. Compute different attention relationships in parallel through the multi-head attention mechanism to generate multiple attention outputs, and finally splice them to obtain the final features;
[0033] S4. Use the final node features to decode through a multi-layer perceptron to generate the action scores for each robot;
[0034] S5. Use the softmax function to convert the action scores into a probability distribution, and select the next action according to the probability.
[0035] Embodiment 2
[0036] In a warehouse environment, there are 5 robots (R1, R2, R3, R4, R5) that need to perform path planning during the process of carrying goods, while avoiding shelves (obstacles) and not colliding with each other. The warehouse is divided into grids, and each grid can be free, occupied by an obstacle, or occupied by a robot.
[0037] Dataset collection (based on the CBS algorithm) Use the CBS algorithm to plan paths for 5 robots from the starting position to the target position. During the low-level search process, use algorithms such as A* to find a preliminary valid path for each robot. For example, R1 starts from the starting point A1, and the preliminary planned path is A1 - B1 - C1 - D1 - target point T1; R2 starts from the starting point A2, and the preliminary path is A2 - B2 - C2 - target point T2. The high-level search process checks whether these paths conflict. If it is found that R1 and R2 conflict at the grid positions C1 and C2 (both need to enter this grid at the same time), then calculate the cost value of the conflicting paths (the cost value can be determined according to factors such as path length and the degree of danger near obstacles), and select the branch with the smallest cost value (assuming it is the path branch of R2) to re-perform the low-level path search. Repeat this process until the high-level search process finds a set of non-conflicting valid paths, and these paths are used as the dataset for subsequent model training.
[0038] Refer to Figure 2 , Environment perception and feature extraction At a certain moment t, each robot obtains the surrounding map information. Taking R3 as an example, it perceives the information within a specific range (such as a 3×3 grid area centered on itself), including the positions of obstacles (such as there are shelves in the grid) and the positions of other robots (such as R4 in the adjacent grid). Input these map perception features into a convolutional neural network (corresponding to Figure 2 the perception feature extraction module in
[0039] Refer to Figure 2, by using a graph neural network and an attention network to update the feature representation, each robot is regarded as a node in the graph neural network. The graph neural network of R3 aggregates the information of its own node and adjacent nodes (such as R4). By changing the values of the adjacency matrix of the graph neural network (assuming that R3 is adjacent to R4, and the corresponding element value in the adjacency matrix is changed to represent the connection relationship between them), communication between robots is achieved (corresponding to Figure 2 the state representation extraction module in
[0040] See Figure 2 , the multi-head attention mechanism is processed, and the multi-head attention mechanism calculates multiple different attention patterns in parallel. For R3, one head focuses on the distance relationship with nearby robots, and another head focuses on the relationship between the moving directions of other robots and its own target direction, etc. Through these different attention relationships, multiple attention outputs are generated, and finally these outputs are concatenated to obtain the final features, capturing a more comprehensive interaction relationship with other robots and enhancing the model's ability to process multi-dimensional information (corresponding to Figure 2 the weighted feature output module in
[0041] See Figure 2 , action decision-making, using the final node features, decoded through a multi-layer perceptron (corresponding to Figure 2 the action generation module in
[0042] ;
[0043] ;
[0044] ;
[0045] ;
[0046] ;
[0047] After calculation, it is found that the probability of "left" is the largest, so R3 chooses to move left.
[0048] Collision avoidance strategy and deadlock resolution mechanism During the execution of actions, if R3 moves left according to the calculated action, it will collide with R2 (judged by the collision detection algorithm). According to the collision avoidance strategy, the action of R3 will be replaced with an idle action. If multiple robots (such as R1, R2, R3) are in an idle state for a long time due to similar collision situations, the system enters a deadlock state (such as exceeding the set time threshold). At this time, a part of the cases are randomly selected from the training set to judge which cases are in a deadlock state, and the expert algorithm (such as re-planning the paths of some robots) is run to resolve the deadlock. The successful trajectories obtained after resolving the deadlock are added to the training set so that the model can learn to avoid similar deadlock situations from occurring again.
[0049] See Figure 3 , which is the training accuracy graph of the system model. Based on the above embodiments, it plays a key role in the entire multi-robot path planning implementation process: When training the model for multi-robot path planning, developers will continuously observe the accuracy. Taking the training of the path planning model of 5 robots in the warehouse as an example, as the number of training rounds gradually increases, the changing trend of the accuracy can be intuitively seen. If in the early stage of training, the accuracy curve shows a steady upward trend, it means that the model is effectively learning and continuously optimizing its ability to process and make decisions on information related to robot path planning. For example, after 50 rounds of training, the accuracy increases from the initial 40% to 60%, indicating that the model has gradually mastered how to better handle key information such as the relationships between robots and avoiding obstacles during this period, and makes more reasonable path planning decisions.
[0050] After the training is over, it can clearly show the final accuracy level reached by the model. Assuming that when the final model training is stable, the accuracy reaches 85%, this data can be directly used to evaluate the model performance. Compared with other similar multi-robot path planning models, if the highest accuracy of other models is 75%, then this model is more excellent in dealing with the multi-robot path planning problem in a complex environment such as a warehouse, and can more accurately plan collision-free and efficient robot paths. If it shows that the accuracy of the model fluctuates during the training process, such as a sudden drop in the accuracy at a certain stage, this prompts the developers that there may be problems with the model. For example, it may be that the learning rate is set unreasonably during the training process, resulting in unstable model training. The developers can adjust model parameters such as the learning rate accordingly, then retrain the model, and observe the change of the accuracy again. By continuously adjusting the parameters, the accuracy curve of the model tends to rise steadily, achieving a better training effect, and thus improving the accuracy and efficiency of multi-robot path planning.
[0051] Embodiment 3
[0052] Use the conflict-based multi-robot path planning CBS algorithm to obtain the multi-robot movement trajectories and use them as the data set during model training;
[0053] The required model is built by simulating communication between multiple robots by changing node features using a graph neural network. The model includes a convolutional neural network for extracting map features, which is used to transfer features between robots, and a graph attention network for dynamically adjusting information propagation according to each robot and its neighbors to improve the accuracy of information transfer.
[0054] By adding a multi-head attention mechanism, complex interaction relationships between robots can be learned in multiple subspaces, thus improving the model efficiency.
[0055] Finally, a multi-layer perceptron is used to decode the finally obtained enhanced feature output to obtain the score of the action taken, and the softmax classification probability function is used to probabilize the score to set the action prediction strategy of the robot.
[0056] A collision prevention mechanism and a deadlock release mechanism are set up.
[0057] According to the above embodiments, a multi-robot path planning model under an attention mechanism is provided. The model is mainly divided into two parts: feature aggregation of neighbor nodes and action mapping.
[0058] During the decision-making process of the robot, it is first necessary to perform necessary perception of the nearby environment: at moment, the robot obtains the surrounding map information, including the position information of obstacles and other robots.
[0059] The relevant map information obtained by each robot is input into the convolutional neural network for feature extraction, and finally each robot will obtain a feature vector.
[0060] According to the definition of the graph at moment represents a node (robot), represents the edge between nodes, represents the weight value. Changing the value of means that each robot node assigns different weights according to the different importance of its neighbor nodes. Adding the graph attention mechanism can also dynamically adjust the weights according to the relationship between the position of the robot at moment and its neighbor nodes, and change the graph network in real time. The feature of each robot node in the graph is , and the new representation of the robot node is obtained through the graph attention mechanism: . Then, the relationship between robots is processed through the multi-head attention mechanism to enhance the representation learning ability of these relationships: .
[0061] Regarding the robot's action decision, a multi-layer perceptron (MLP) network is used to decode the aggregated features output by the multi-head attention mechanism. Five actions (including up, down, left, right, and idle) are set. The aggregated features are decoded by MLP to obtain the scores of the actions. Then, the scores are converted into probability distributions according to the softmax function of multi-class classification. The mathematical expression of the softmax function is as follows: , represents the output of the MLP network, It is a normalization operation for all outputs to ensure that the sum of all output probabilities is 1. Represents different directions, paths or action choices. Softmax converts these selected values into probability distributions. The robot selects the action with the highest probability for output based on the probability distribution output by softmax.
[0062] Set up an anti-collision mechanism for action decisions: The anti-collision mechanism is in the action phase. It is necessary to ensure that each robot cannot collide with each other and avoid obstacles. The anti-collision mechanism needs to be set up and implemented as follows: If the inferred action will cause a collision with another robot or an obstacle, the action is replaced by an idle action; if the inferred actions of the two robots will cause an edge collision (let them exchange positions), these actions are replaced by idle actions. Another possibility is that the robot is always in an idle state until it times out due to the above two situations.
[0063] The existence of a timeout means that the system has entered a deadlock state. In this invention, a deadlock state is defined as a robot being idle for a long time until the timeout. The deadlock state is handled by a data aggregation method. Specifically, in each round, a portion of cases are randomly selected from the training set, and it is determined which cases are in a deadlock state, and the deadlock is released by running an expert algorithm. The obtained successful trajectory is added to the training set.
[0064] According to the above embodiments, the multi-robot path planning method under the attention mechanism can achieve various effects in the multi-robot collaborative operation scenario: by means of the CBS algorithm, a high-quality data set is obtained, and the model can accurately plan the path after training. In the warehouse scenario, the robot can quickly find the route from the starting point to the target point, avoid obstacles, and will not waste time due to circuitous routes. Taking the R1 robot as an example, it can quickly move from the goods storage area to the shipping area according to the plan, improving the overall handling efficiency and reducing the waiting time for goods backlog. By extracting map features through a convolutional neural network and dynamically aggregating neighbor information by the graph attention network, it can accurately perceive and adapt to complex environments. When the warehouse layout changes or new obstacles are added, the robot can quickly adjust the path. For example, when new shelves are temporarily set up in the warehouse, the robot can detect it in time and re-plan the route to keep the operation smooth. The multi-head attention mechanism enables the robot to capture complex interaction relationships and achieve intelligent cooperation. When multiple robots meet in a narrow passage, they can dynamically adjust their actions according to each other's positions, speeds, and goals to avoid congestion. For example, at the intersection, R3 and R4 can automatically coordinate the order to improve space utilization and operation efficiency. Under the anti-collision and deadlock release mechanism, the safe operation of the robot is guaranteed. In actual operation, it can effectively avoid collisions between robots or with obstacles, reducing the risk of equipment damage. Once a deadlock occurs, it can be released in time, such as obtaining experience from the training set and re-planning the paths of some robots to ensure the continuous and stable operation of the system. From the model training accuracy graph, it can be seen that the model training effect is good and the accuracy is relatively high. A high accuracy means that the robot path planning decision is more reasonable, which can reduce wrong actions in practical applications, lower the probability of path planning failure, enhance the reliability and stability of the system, and provide strong support for large-scale multi-robot collaborative operation.
[0065] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-robot path planning method under an attention mechanism, characterized in that Including: S1. Input the obtained map perception features into the graph attention network for extraction through the CBS algorithm and based on the map perception features around the current position of the robot. S2. Obtain the updated feature representation through the attention network and weighted summation according to the features of neighbor nodes. S3. Generate multiple attention outputs through parallel calculation of different attention relationships by the multi-head attention mechanism, and finally splice them to obtain the final features. S4. Use the final node features to generate the action scores of each robot through multi-layer perceptron decoding. S5. Use the softmax function to convert the action scores into a probability distribution, and select the next action according to the probability.
2. The multi-robot path planning method under the attention mechanism according to claim 1, wherein, In step S1, dataset collection is performed through the CBS algorithm. The CBS algorithm consists of two search processes, a high-level search process and a low-level search process. The low-level search process is responsible for searching an effective path for each robot, and the high-level search process is responsible for checking path conflicts and selecting the branch with the smallest cost value to re-perform the low-level path search until the high-level search process finds an effective path.
3. A multi-robot path planning method under an attention mechanism according to claim 1, characterized in that In step S1, the map perception features are extracted through a convolutional neural network. The map perception features include obstacles within a specific range and the position information of other robots.
4. A multi-robot path planning method under an attention mechanism according to claim 1, characterized in that, In step S2, the graph neural network is used to summarize the information of its own node and adjacent nodes to update the characteristics expressed in the adjacency matrix, realizing communication between robot nodes.
5. A multi-robot path planning method under an attention mechanism according to claim 1, characterized in that, In step S2, through the network dynamics under the attention mechanism and dynamically weighted aggregation of neighbor information according to the relationship between each robot and its neighbors.
6. A multi-robot path planning method under an attention mechanism according to claim 4, characterized in that, Through the mutual relationship between robots and environmental information, and by changing the values of the adjacency matrix of the graph neural network, communication between robots is realized.
7. A multi-robot path planning method under an attention mechanism according to claim 1, characterized in that In step S3, the multi-head attention mechanism parallelly calculates multiple different attention patterns to capture different interaction relationships among multiple robots.
8. A multi-robot path planning method under an attention mechanism according to claim 1, characterized in that, In step S5, after obtaining the final output features, a multi-layer perceptron is set to decode the final output features. After obtaining the decoding scores, they are converted into probabilities using the classification probability function softmax. Finally, the action with the highest probability is selected, and an anti-collision strategy and a deadlock release mechanism are set.
9. A multi-robot path planning system under an attention mechanism, characterized in that, This system is used to execute the method according to any one of claims 1-8.
10. A readable storage medium stores instructions, characterized in that, The instructions, when executed by one or more processors of a machine, cause the processors to execute the method according to any one of claims 1-8.
Citation Information
Cited By
Dynamic optimization control system for driving path of rail trolley based on AI vision
CN120704344A