Multi-agent path planning method based on attention mechanism

By adopting an encoder-decoder structure based on attention mechanism and a reinforcement learning framework in multi-agent path planning, the problems of low efficiency and relying on manual design in the prior art are solved, and more efficient and reliable path planning is achieved.

CN120027797APending Publication Date: 2025-05-23XI AN JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510162377.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing multi-agent path planning algorithm is inefficient when solving large-scale tasks, and the heuristic algorithm relies too much on manual design features, resulting in insufficient reliability and efficiency of path planning.

Method used

A multi-agent path planning method based on attention mechanism is adopted, and a path planning model of the encoder-decoder structure is constructed, combined with multi-head attention operation and feedforward operation, a high-dimensional feature representation is generated, and a reinforcement learning framework based on rollback benchmarks is designed to optimize the path planning results.

Benefits of technology

The quality and efficiency of multi-agent path planning are improved, the cost of path planning is reduced, and the dynamic perception ability of the model and the reliability of path planning are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120027797A_ABST
    Figure CN120027797A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of path planning, and particularly provides a multi-agent path planning method based on an attention mechanism, and the method comprises the following steps: S1, selecting path planning training data, and constructing a training set; s2, respectively designing an encoder and a decoder by using an attention mechanism, and constructing a path planning model based on an encoder-decoder structure; s3, designing a reinforcement learning framework based on a rollback reference, and constructing a multi-agent task-oriented reward function; and S4, inputting the data set obtained in the S1 into the path planning model constructed in the S2, training the path planning model by using the reinforcement learning framework designed in the S3, and applying the trained model to multi-agent autonomous path planning. According to the multi-agent path planning method, node high-dimensional feature representation is obtained through an encoder, probability distribution is calculated through a decoder, the model is trained through a reinforcement learning framework combined with a rollback reference, the multi-agent path planning cost is reduced, and the path planning quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of path planning, and specifically relates to a multi-agent path planning method based on an attention mechanism. Background Art

[0002] With the rapid development of intelligent technology, multi-agent systems are increasingly being used in logistics distribution, drone formations, autonomous driving, disaster relief and other tasks. How to plan paths for multi-agents so that they can traverse all task points at the lowest cost is the key to these tasks. However, in existing path planning algorithms, precise algorithms cannot cope with large-scale task solutions, and heuristic algorithms rely too much on artificially designed features. More advanced path planning algorithms are urgently needed to improve the reliability and efficiency of multi-agent path planning. Summary of the invention

[0003] In order to solve the current problems in multi-agent path planning, the purpose of the present invention is to provide a multi-agent path planning method based on an attention mechanism to improve the quality and efficiency of path planning.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] A multi-agent path planning method based on an attention mechanism comprises the following steps:

[0006] S1: Select path planning training data and construct a training set;

[0007] S2: Use the attention mechanism to design the encoder and decoder respectively, and build a path planning model based on the encoder-decoder structure;

[0008] S3: Design a reinforcement learning framework based on rollback benchmarks and construct reward functions for multi-agent tasks;

[0009] S4: Input the data set obtained by S1 into the path planning model constructed by S2, use the reinforcement learning framework designed by S3 to train the path planning model, and apply the trained model to multi-agent autonomous path planning.

[0010] Furthermore, in S1, the constructed training set includes a number of training samples, each of which is composed of a specified number of node coordinate information in the [0, 1]×[0, 1] two-dimensional region of the sample, including a starting node and remaining nodes.

[0011] Furthermore, each training sample is used as the input of the path planning model in the following matrix format:

[0012] X=[x 0 ;x 1 ;…x i;...;x N ]

[0013] Among them, x i represents the i-th node in the sample, which is composed of the coordinate data of the node; x 0 represents the starting node, and N is the number of nodes in the sample excluding the starting node.

[0014] Furthermore, in S2, the path planning model is composed of an encoder and a decoder. For each training sample input, the encoder first processes the two-dimensional node information and outputs the node high-dimensional feature representation. The decoder combines the node high-dimensional feature representation and context embedding information to generate a path set planned by the path planning model for the sample:

[0015] r[1,M]=(r[1],r[2],...,r[M])

[0016] Where M is the number of agents, r[1, M] is the set of M independent paths of the M agents, all agents start from the starting node and eventually return to the starting node, and at the end of the task, all nodes except the starting node have been visited once by one of the M agents.

[0017] Furthermore, the encoder processes the two-dimensional node information as follows:

[0018] First, for the input node information matrix X, an embedding layer is used to calculate its high-dimensional node embedding:

[0019] E (0) =XW X +B X

[0020] Among them, E (0) is the high-dimensional node embedding of the input node information matrix X, W X is the learnable weight parameter matrix of the embedding layer, B X is the learnable bias parameter matrix of the embedding layer;

[0021] Secondly, E (0) Input L end-to-end coding units and get the final output E at the last coding unit. (l) , as the high-dimensional feature representation of the node; these encoding units have the same structure and do not share parameters. For the kth encoding unit, its input is E (k-1) The encoding unit first performs multi-head attention and normalization operations on E (k-1) Processing to obtain primary feature representation:

[0022]

[0023] in, For E (k-1) The primary feature representation is, BN() is the normalization operation, Sigmoid() is the nonlinear activation operation, MHA() is the multi-head attention operation, is the learnable attention weight parameter matrix of the kth encoding unit, is the learnable attention bias parameter matrix of the kth encoding unit, is the corresponding element multiplication operation;

[0024] For the primary feature representation, the encoding unit processes it based on the feedforward operation and normalization operation to obtain the output feature representation of the kth encoding unit:

[0025]

[0026] Among them, E (k) is the output feature representation of the kth encoding unit, FF() is the feedforward operation, is the learnable feedforward weight parameter matrix of the kth encoding unit, is the learnable feedforward bias parameter matrix of the kth encoding unit;

[0027] After L end-to-end coding units, the final output of the encoder is:

[0028] E (l) =[e 0 ;e 1 ;...;e N ]

[0029] That is, the high-dimensional feature representation of each node.

[0030] Furthermore, the decoder constructs a context embedding vector based on the node high-dimensional feature representation and context embedding information:

[0031] c t =Concat(mean{e t:N},e 0 , e t-1 )W C

[0032] Among them, t is the current time step, c t is the context embedding vector, e 0 With e t-1 are high-dimensional feature representations of the starting node and the currently visited node, representing contextual embedding information. Concat() is a horizontal concatenation operation. C is the learnable context embedding weight parameter matrix, mean{e t : N} is the mean of the high-dimensional feature representation of the unvisited nodes, that is:

[0033]

[0034] Where N is the number of nodes in the sample excluding the starting node.

[0035] After that, the compatibility value is calculated based on the attention mechanism:

[0036]

[0037] in, is the compatibility value of the decoder for matching the i-th node to the m-th agent, c t is the context embedding vector, e i is the high-dimensional feature representation of the i-th node, d k for e i The number of vector dimensions of is the learnable query value weight parameter matrix of the mth agent, is the learnable key-value weight parameter matrix of the mth agent, for The transpose of .

[0038] For nodes that have been visited, their compatible values Set to -∞;

[0039] After that, based on the softmax operation, the probability that the decoder selects the i-th node as the next node to be visited is calculated:

[0040]

[0041] in, represents the probability that the decoder matches the i-th node to the m-th agent, and j represents the remaining nodes participating in the calculation The node number of represents the compatibility value of the decoder for matching the j-th node to the m-th agent.

[0042] Furthermore, the constructed path planning model based on the encoder-decoder structure calculates the number of all unvisited points in each time step according to the current time step. The node to be visited in the current time step is selected by sampling, and the planned path is updated until all nodes are visited to obtain the path set:

[0043] r[1,M]=(r[1],r[2],...,r[M]).

[0044] Furthermore, in S3, the designed rollback benchmark-based reinforcement learning framework structurally includes a backbone network and a benchmark network;

[0045] The path planning model constructed by S2 is used as the backbone network, and the backbone network selects nodes by sampling according to the probability values ​​generated by the decoder; the benchmark network selects the same structure as the backbone network, but selects nodes by greedily according to the probability values ​​generated by the decoder;

[0046] During the training process, the reinforcement learning framework calculates the respective reward functions based on the path planning results output by the backbone network and the benchmark network, and guides the backbone network parameter update with the goal of maximizing the backbone network's reward function;

[0047] At the end of each round of training, the output results of the backbone network are compared with those of the baseline network. If the output results of the backbone network are better at this time, the parameters of the baseline network are rolled back and updated to be the same as those of the backbone network. Otherwise, the parameters of the baseline network are kept unchanged until all rounds of training are completed.

[0048] Furthermore, in S3, the reward function is designed as follows:

[0049] R(r[1,M])=-Cost(r[1,M])-μM

[0050] Where R(r[1,M]) is the reward function, μ is the cost coefficient for each additional agent, M is the number of agents, and Cost(r[1,M]) is the path cost, which is the sum of the total lengths of all paths in the path set.

[0051] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0052] The multi-agent path planning method based on the attention mechanism provided by the present invention uses multiple encoders to extract node features and fuses multi-head attention operations with feedforward operations to obtain richer high-dimensional feature representations;

[0053] The multi-agent path planning method based on the attention mechanism provided by the present invention only considers the high-dimensional feature representation of unvisited nodes when designing the decoder, which overcomes the shortcomings of the original static embedding information and improves the dynamic perception ability of the model;

[0054] The multi-agent path planning method based on the attention mechanism provided by the present invention applies the rollback benchmark method to the reinforcement learning framework and designs a reward function for multi-agent tasks, thereby reducing the multi-agent path planning cost and improving the path planning quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Graph of the path planning model based on the encoder-decoder structure.

[0056] Figure 2This is a graph of a multi-agent path planning example. DETAILED DESCRIPTION

[0057] In order to make those skilled in the art better understand the scheme of the present invention, the technical scheme of the present invention is clearly and completely described below in conjunction with the embodiments. It should be understood that the embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] The present invention provides a multi-agent path planning method based on an attention mechanism, comprising the following steps:

[0059] S1: Select path planning training data and construct a training set;

[0060] The constructed training set includes several training samples, each of which is composed of a specified number of node coordinate information in the [0,1]×[0,1] two-dimensional area of ​​the sample, including the starting node and the remaining nodes.

[0061] In this example, the constructed training set includes three node sizes, and the number of samples of different sizes is shown in Table 1.

[0062] Table 1

[0063] Node scale 20 50 100 Sample size 409500 409500 386000

[0064] S2: Use the attention mechanism to design the encoder and decoder respectively, and build a path planning model based on the encoder-decoder structure;

[0065] Among them, each training sample is used as the input of the path planning model in the following matrix format:

[0066] X=[x 0 ;x 1 ;...x i ;...;x N ]

[0067] In the formula, x i represents the i-th node in the sample, which is composed of the coordinate data of the node; x 0 represents the starting node, and N is the number of nodes in the sample excluding the starting node.

[0068] like Figure 1 , the path planning model consists of an encoder and a decoder. For each training sample input, the encoder first processes the two-dimensional node information and outputs the node high-dimensional feature representation. The decoder combines the node high-dimensional feature representation and context embedding information to generate the path set planned by the path planning model for the sample:

[0069] r[1,M]=(r[1],r[2],...,r[M])

[0070] Where M is the number of agents, r[1, M] is the set of M independent paths of the M agents, all agents start from the starting node and eventually return to the starting node, and at the end of the task, all nodes except the starting node have been visited once by one of the M agents.

[0071] The encoder processes the two-dimensional node information as follows:

[0072] First, for the input node information matrix X, an embedding layer is used to calculate its high-dimensional node embedding:

[0073] E(0)=XW X +B X

[0074] Where E(0) is the high-dimensional node embedding of the input node information matrix X, W X is the learnable weight parameter matrix of the embedding layer, B X is the learnable bias parameter matrix of the embedding layer.

[0075] Secondly, E(0) is input into L end-to-end coding units, in this example, L is 3, and the final output E is obtained in the last coding unit. (l) , as the high-dimensional feature representation of the node. These encoding units have the same structure and do not share parameters. For the kth encoding unit, its input is E (k-1) The encoding unit first performs multi-head attention and normalization operations on E (k-1) Processing to obtain primary feature representation:

[0076]

[0077] in, For E (k-1) The primary feature representation is, BN() is the normalization operation, Sigmoid() is the nonlinear activation operation, MHA() is the multi-head attention operation, is the learnable attention weight parameter matrix of the kth encoding unit, is the learnable attention bias parameter matrix of the kth encoding unit, It is the corresponding element multiplication operation.

[0078] For the primary feature representation, the encoding unit processes it based on the feedforward operation and normalization operation to obtain the output feature representation of the kth encoding unit:

[0079]

[0080] Among them, E(k) is the output feature representation of the kth encoding unit, FF() is the feedforward operation, is the learnable feedforward weight parameter matrix of the kth encoding unit, is the learnable feedforward bias parameter matrix of the kth encoding unit.

[0081] After L end-to-end coding units, the final output of the encoder is:

[0082] E (l) =[e 0 ;e 1 ;...;e N ]

[0083] That is, the high-dimensional feature representation of each node.

[0084] The decoder constructs a context embedding vector based on the node high-dimensional feature representation and context embedding information:

[0085] c t =Concat(mean{e t:N},e 0 , e t-1 )W C

[0086] Among them, t is the current time step, c t is the context embedding vector, e 0 With e t-1 are high-dimensional feature representations of the starting node and the currently visited node, representing contextual embedding information. Concat() is a horizontal concatenation operation. C is the learnable context embedding weight parameter matrix, mean{e t:N} is the mean of the high-dimensional feature representation of the unvisited nodes, that is:

[0087]

[0088] After that, the compatibility value is calculated based on the attention mechanism:

[0089]

[0090] in, is the compatibility value of the decoder for matching the i-th node to the m-th agent, c t is the context embedding vector, e i is the high-dimensional feature representation of the i-th node, d k for e i The number of vector dimensions of is the learnable query value weight parameter matrix of the mth agent, is the learnable key-value weight parameter matrix of the mth agent, for The transpose of .

[0091] For nodes that have been visited, their compatible values Set to -∞.

[0092] After that, based on the softmax operation, the probability that the decoder selects the i-th node as the next node to be visited is calculated:

[0093]

[0094] in, represents the probability that the decoder matches the i-th node to the m-th agent, and j represents the remaining nodes participating in the calculation The node number of represents the compatibility value of the decoder for matching the j-th node to the m-th agent.

[0095] The constructed path planning model based on the encoder-decoder structure calculates the number of unvisited points in each time step according to the current time step. The node to be visited in this time step is selected by sampling, and the planned path is updated until all nodes are visited to obtain the path set:

[0096] r[1,M]=(r[1],r[2],...,r[M])

[0097] S3: Design a reinforcement learning framework based on rollback benchmarks and construct reward functions for multi-agent tasks;

[0098] The designed rollback benchmark-based reinforcement learning framework structurally includes a backbone network and a benchmark network.

[0099] Among them, the path planning model constructed by S2 is used as the backbone network, and the backbone network selects nodes by sampling according to the probability values ​​generated by the decoder; the benchmark network selects the same structure as the backbone network, but selects nodes by greedily according to the probability values ​​generated by the decoder.

[0100] During the training process, the reinforcement learning framework calculates the respective reward functions based on the path planning results output by the backbone network and the benchmark network, and guides the update of the backbone network parameters with the goal of maximizing the reward function of the backbone network.

[0101] At the end of each round of training, the output results of the backbone network are compared with those of the baseline network. If the output results of the backbone network are better at this time, the parameters of the baseline network are rolled back and updated to be the same as those of the backbone network. Otherwise, the parameters of the baseline network are kept unchanged until all rounds of training are completed.

[0102] The reward function is designed as follows:

[0103] R(r[1,M])=-Cost(r[1,M])-μM

[0104] Where R(r[1,M]) is the reward function, μ is the cost coefficient for each additional agent, M is the number of agents, and Cost(r[1,M]) is the path cost, which is the sum of the total lengths of all paths in the path set.

[0105] S4: Input the data set obtained by S1 into the path planning model constructed by S2, use the reinforcement learning framework designed by S3 to train the path planning model, and apply the trained model to multi-agent autonomous path planning.

[0106] In this example, the constructed path planning model is trained for 100 rounds. After the training, it is tested on node systems of different scales using a public data set. The results for each scale of path planning tasks are shown in Table 2.

[0107] Table 2

[0108] Node scale Total path distance Planning time 20 6.25m 4s 50 10.77m 8s 100 16.67m 19s

[0109] like Figure 2 As shown, the path planning is obtained by processing a test sample using the trained path planning model. From the obtained path planning results, it can be seen that the multi-agent path planning method based on the attention mechanism proposed in the present invention can reliably solve the multi-agent path planning task.

Claims

1. A multi-agent path planning method based on attention mechanism, characterized in that: The following steps are involved: S1: Select path planning training data and construct a training set; S2: Use the attention mechanism to design the encoder and decoder respectively, and build a path planning model based on the encoder-decoder structure; S3: Design a reinforcement learning framework based on rollback benchmarks and construct reward functions for multi-agent tasks; S4: Input the data set obtained by S1 into the path planning model constructed by S2, use the reinforcement learning framework designed by S3 to train the path planning model, and apply the trained model to multi-agent autonomous path planning.

2. The multi-agent path planning method based on the attention mechanism according to claim 1 is characterized in that: In S1, the constructed training set includes several training samples, each of which is composed of a specified number of node coordinate information in the [0,1]×[0,1] two-dimensional area of ​​the sample, including the starting node and the remaining nodes.

3. The multi-agent path planning method based on the attention mechanism according to claim 2, characterized in that: Each training sample is used as input to the path planning model in the following matrix format: X=[x0;x1;…x i ;…;x N ] Among them, x i represents the i-th node in the sample, which is composed of the coordinate data of the node; x0 represents the starting node, and N is the number of nodes in the sample excluding the starting node.

4. The multi-agent path planning method based on attention mechanism according to claim 1 is characterized in that: In S2, the path planning model consists of an encoder and a decoder. For each training sample input, the encoder first processes the two-dimensional node information and outputs the node high-dimensional feature representation. The decoder combines the node high-dimensional feature representation and context embedding information to generate the path set planned by the path planning model for the sample: r[1,M]=(r[1],r[2],…,r[M]) Where M is the number of agents, r[1,M] is the set of M independent paths of the M agents, all agents start from the starting node and eventually return to the starting node, and at the end of the task, all nodes except the starting node have been visited once by one of the M agents.

5. The multi-agent path planning method based on the attention mechanism according to claim 4 is characterized in that: The encoder processes the two-dimensional node information as follows: First, for the input node information matrix X, an embedding layer is used to calculate its high-dimensional node embedding: E (0) =XW X +B X Among them, E (0) is the high-dimensional node embedding of the input node information matrix X, W X is the learnable weight parameter matrix of the embedding layer, B X is the learnable bias parameter matrix of the embedding layer; Secondly, E (0) Input L end-to-end coding units and get the final output E at the last coding unit. (l) , as the high-dimensional feature representation of the node; these encoding units have the same structure and do not share parameters. For the kth encoding unit, its input is E (k-1) The encoding unit first performs multi-head attention and normalization operations on E (k-1) Processing to obtain primary feature representation: in, For E (k-1) The primary feature representation is, BN() is the normalization operation, Sigmoid() is the nonlinear activation operation, MHA() is the multi-head attention operation, is the learnable attention weight parameter matrix of the kth encoding unit, is the learnable attention bias parameter matrix of the kth encoding unit, is the corresponding element multiplication operation; For the primary feature representation, the encoding unit processes it based on the feedforward operation and normalization operation to obtain the output feature representation of the kth encoding unit: Among them, E (k) is the output feature representation of the kth encoding unit, FF() is the feedforward operation, is the learnable feedforward weight parameter matrix of the kth encoding unit, is the learnable feedforward bias parameter matrix of the kth encoding unit; After L end-to-end coding units, the final output of the encoder is: HAVE BEEN (l) =[e0;e1;…;e N ] That is, the high-dimensional feature representation of each node.

6. The multi-agent path planning method based on the attention mechanism according to claim 4, characterized in that: The decoder constructs a context embedding vector based on the node high-dimensional feature representation and context embedding information: c t =Concat(mean{e t:N },e0,e t-1 )W C Among them, t is the current time step, c t is the context embedding vector, e0 and e t-1 are high-dimensional feature representations of the starting node and the currently visited node, representing contextual embedding information. Concat() is a horizontal concatenation operation. C is the learnable context embedding weight parameter matrix, mean{e t:N } is the mean of the high-dimensional feature representation of the unvisited nodes, that is: Where N is the number of nodes in the sample excluding the starting node; After that, the compatibility value is calculated based on the attention mechanism: in, is the compatibility value of the decoder for matching the i-th node to the m-th agent, c t is the context embedding vector, e i is the high-dimensional feature representation of the i-th node, d k for e i The number of vector dimensions of is the learnable query value weight parameter matrix of the mth agent, is the learnable key-value weight parameter matrix of the mth agent, for The transpose of . For nodes that have been visited, their compatible values Set to -∞; After that, based on the softmax operation, the probability that the decoder selects the i-th node as the next node to be visited is calculated: in, represents the probability that the decoder matches the i-th node to the m-th agent, and j represents the remaining nodes participating in the calculation The node number of represents the compatibility value of the decoder for matching the j-th node to the m-th agent.

7. The multi-agent path planning method based on the attention mechanism according to claim 6, characterized in that: The constructed path planning model based on the encoder-decoder structure calculates the number of unvisited points in each time step according to the current time step. The node to be visited in the current time step is selected by sampling, and the planned path is updated until all nodes are visited to obtain the path set: r[1,M]=(r[1],r[2],…,r[M]).

8. The multi-agent path planning method based on attention mechanism according to claim 1 is characterized in that: In S3, the designed rollback benchmark-based reinforcement learning framework structurally includes a backbone network and a benchmark network; The path planning model constructed by S2 is used as the backbone network, and the backbone network selects nodes by sampling according to the probability values ​​generated by the decoder; the benchmark network selects the same structure as the backbone network, but selects nodes by greedily according to the probability values ​​generated by the decoder; During the training process, the reinforcement learning framework calculates the respective reward functions based on the path planning results output by the backbone network and the benchmark network, and guides the backbone network parameter update with the goal of maximizing the backbone network's reward function; At the end of each round of training, the output results of the backbone network are compared with those of the baseline network. If the output results of the backbone network are better at this time, the parameters of the baseline network are rolled back and updated to be the same as those of the backbone network. Otherwise, the parameters of the baseline network are kept unchanged until all rounds of training are completed.

9. The multi-agent path planning method based on attention mechanism according to claim 1, characterized in that: In S3, the reward function is designed as follows: R(r[1,M])=-Cost(r[1,M])-μM Among them, R(r[1,M]) is the reward function, μ is the cost coefficient for each additional agent, M is the number of agents, and Cost(r[1,M]) is the path cost, that is, the sum of the total lengths of all paths in the path set.

Citation Information

Cited By

  • Unmanned aerial vehicle inspection path planning method based on attention mechanism

    CN120576772A