Multi-agent formation control method and device based on dynamic graph attention network and deep reinforcement learning, and storage medium
By adopting dynamic graph attention network and deep reinforcement learning methods in multi-agent formation control, the problem of high complexity of collaborative control in the prior art is solved, and real-time and efficient control of multi-agent formations is achieved.
Patent Information
- Application Number
- CN202510623393.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing multi-agent formation control technology is highly complex in the collaborative control process, making it difficult to achieve real-time control and efficient collaboration.
The control method based on dynamic graph attention network and deep reinforcement learning is adopted, and the multi-agent formation is characterized by aggregation processing through the graph attention network, and combined with independent proximity strategy optimization model, the action probability distribution of the agent is generated, thereby achieving effective control of the multi-agent formation.
It improves the real-time and efficiency of multi-agent formation control, reduces the demand for global information, and reduces the possibility of control failure caused by some agent failures.
Smart Images

Figure CN120122722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of control technologies, and in particular to a multi-agent formation control method, device, and storage medium based on a dynamic graph attention network and deep reinforcement learning. Background Art
[0002] In fields such as industrial production, agricultural production, art performances, law enforcement processes, and military operations, it is usually necessary to control objects such as unmanned aerial vehicle formations, unmanned vehicle formations, unmanned boat formations, and robot formations to perform certain actions. In these formations, each unmanned aerial vehicle, unmanned vehicle, unmanned boat, robot, etc. is an entity that can perceive the environment and take actions to achieve specific goals. Therefore, each object such as an unmanned aerial vehicle, unmanned vehicle, unmanned boat, or robot can be called an agent, and the unmanned aerial vehicle formation, unmanned vehicle formation, unmanned boat formation, and robot formation form a multi-agent formation. The actions of each agent in the multi-agent formation may, on the one hand, affect the actions of other agents, and on the other hand, may also affect the overall actions of the multi-agent formation. Therefore, the control of the multi-agent formation needs to consider factors such as coordination, which is more complex than the control of a single agent. Summary of the Invention
[0003] Aiming at the technical problems such as the complex collaborative control process faced by the current multi-agent formation control technology, the purpose of the present invention is to provide a multi-agent formation control method, device, and storage medium based on a dynamic graph attention network and deep reinforcement learning.
[0004] On the one hand, an embodiment of the present invention includes a multi-agent formation control method based on a dynamic graph attention network and deep reinforcement learning. The multi-agent formation control method based on a dynamic graph attention network and deep reinforcement learning includes at least one processing process, and any one of the processing processes includes the following steps: Obtain a graph attention network; the graph attention network is used to perform feature aggregation processing on the data to be processed from the original node and the neighbor nodes to obtain an output vector; Obtain an independent proximal policy optimization model; the independent proximal policy optimization model includes a first neural network and a second neural network. The second neural network is used to process the output vector to obtain a state value; the first neural network is used to process the output vector to obtain an action probability distribution; Obtain the observation information of the first agent, use the first agent as the original node, use the second agent as the neighbor node, and use the observation information of the first agent as the data to be processed, and input it into the graph attention network and the independent proximal policy optimization model for processing; the first agent and the second agent are different agents in the same multi-agent formation; Control the first agent according to the action probability distribution optimized by the independent proximal policy optimization model.
[0005] Further, the feature aggregation process of the data to be processed from the original node and the neighbor nodes to obtain an output vector includes: Encode the data to be processed to obtain the features of the original node; the features of the original node include a vector matrix, an adjacency matrix, and stable features; For any neighbor node, determine the attention coefficient between the original node and the neighbor node according to the features of the original node and the features of the neighbor node; According to each of the attention coefficients, determine the feature update rules for each layer of the graph attention network through a multi-head attention mechanism; According to each of the feature update rules, perform a multi-layer mapping process on the vector matrix and the adjacency matrix to obtain node vectors; Concatenate the stable features and the node vectors to obtain the output vector.
[0006] Further, the determining the attention coefficient between the original node and the neighbor node according to the features of the original node and the features of the neighbor node includes: Calculate according to the formula
[0007] wherein, represents the features of the original node, represents the features of the neighbor node with the serial number , is a trainable vector, is a trainable matrix, represents traversing all values of to perform a softmax operation, represents the attention coefficient between the original node and the neighbor node with the serial number .
[0008] Further, the determining the feature update rules for each layer of the graph attention network through a multi-head attention mechanism according to each of the attention coefficients includes: Determine the feature update rule as
[0009] wherein, represents the output result of the -th layer in the graph attention network, is the number of attention heads, The set of all the neighbor nodes of the original node is the attention coefficient calculated by the th attention head, represents the concatenation operation in the feature dimension, represents the parameters of the th layer in the graph attention network, represents the output result of the th layer in the graph attention network,
[0010] Further, controlling the first agent according to the action probability distribution output by the independent proximal policy optimization model includes: Setting the reward function of the first agent; the reward function includes at least one of a proximity reward, a position reward, a direction reward, and a collision avoidance reward; Determining a target action with the goal of maximizing the discounted cumulative reward; the discounted cumulative reward is determined according to the action probability distribution and the reward function; Controlling the first agent to execute the target action.
[0011] Further, the multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning further includes: Executing at least one episode; each episode respectively includes a plurality of the processing procedures executed continuously; For any one of the episodes, when the episode is completed, determining a first loss function and a second loss function according to the execution results of the processing procedures in the episode; Updating the parameters of the first neural network with the goal of minimizing the first loss function; Updating the parameters of the second neural network with the goal of minimizing the second loss function; When the training end condition is not satisfied, executing the next episode, and when the training end condition is satisfied, ending the execution of all episodes; wherein, the training end condition is that the number of executed episodes reaches a preset upper limit, or the reward functions obtained in the executed episodes converge.
[0012] Further, determining the second loss function and the first loss function according to the execution results of the processing procedures in the episode includes: Determining the first loss function as
[0013] Determining the second loss function as .
[0015] Further, the parameter update of the first neural network includes: Updating the parameters of the first neural network using the policy gradient method; The parameter update of the second neural network includes: Updating the parameters of the second neural network using the temporal difference method.
[0016] On the other hand, an embodiment of the present invention further includes a computer device, including a memory and a processor. The memory is used to store at least one program, and the processor is used to load at least one program to execute the multi-agent formation control method based on a dynamic graph attention network and deep reinforcement learning in the embodiment.
[0017] On the other hand, an embodiment of the present invention further includes a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the multi-agent formation control method based on a dynamic graph attention network and deep reinforcement learning in the embodiment when executed by the processor.
[0018] The beneficial effects of the present invention are as follows: The multi-agent formation control method based on a dynamic graph attention network and deep reinforcement learning in the embodiment can process the multi-agent formation as a multi-node network by using a graph attention network and an independent proximal policy optimization model. Specifically, the observation information obtained by a specific agent (the first agent) in the multi-agent formation through perceiving the outside world (including other agents or obstacles) is processed to obtain the action probability distribution of the specific agent, thereby realizing the control of the specific agent; during the control process of the multi-agent formation, focusing on a specific agent (the first agent) and its neighbor nodes (the second agent), so the observation information detected by some agents can be used to realize the control of the multi-agent formation, which can reduce the demand for global information, thus facilitating the improvement of the real-time performance of the control, reducing the possibility of uncontrollability caused by the failure of some agents, and realizing the effective control of the multi-agent formation. Description of the Drawings
[0019] Figure 1 It is a schematic diagram of the structure and principle of the graph attention network and the independent proximal policy optimization model in the embodiment; Figure 2 It is a schematic diagram of the structure and principle of the perception compression model composed of the graph attention network in the embodiment; Figure 3 It is a schematic diagram of the structure and principle of the first neural network in the independent proximal policy optimization model in the embodiment; Figure 4 It is a schematic diagram of the structure and principle of the second neural network in the independent proximal policy optimization model in the embodiment; Figure 5 Schematic diagram of the steps of the multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning in the embodiment; Figure 6 Schematic diagram of the training effect of the graph attention network and the independent proximal policy optimization model in the embodiment; Figure 7 Schematic diagram of the test results of the data memory occupancy of the graph attention network and the independent proximal policy optimization model in the embodiment. Detailed implementation manners
[0020] Currently, large-scale multi-agent formation control technologies represented by centralized control rely on the global observation of all agents in the multi-agent formation to perform processes such as model training and agent control. Therefore, these technologies generally have problems such as high dependence on the input vector dimension, slow model training speed, and difficult convergence. In addition, since a single agent is difficult to obtain global information, and the method of collecting the information observed by all agents to obtain global information has problems such as serious time delay and easy error, global observation is often difficult to achieve in practical applications.
[0021] Based on the above principle, in this embodiment, a multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning is provided. In this embodiment, the formation of unmanned aerial vehicles is taken as an example of the multi-agent formation, and the unmanned aerial vehicle is taken as an example of the agent for illustration. Specifically, the multi-agent formation is regarded as a network, and each agent is a node in the network, and the nodes can be marked and distinguished by serial numbers.
[0022] In this embodiment, the node with the serial number is used as the observation subject and control target, and the node with the serial number is called the original node, which corresponds to a specific agent in the multi-agent formation, that is, the first agent. Other agents in the multi-agent formation, that is, nodes, especially the nodes close to the first agent, that is, the original node, have an adjacent relationship with the original node and belong to the neighbor nodes of the original node. The neighbor nodes correspond to other agents in the multi-agent formation except the first agent, which are called the second agents in this embodiment. The number of the second agents can be multiple.
[0023] In this embodiment, the multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning can be executed in a real application scenario to realize the control of physical agents (such as unmanned aerial vehicles, etc.). It can also be executed in a virtual environment to realize the simulation of the multi-agent formation control.
[0024] If you choose to execute the multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning in a virtual environment, you can use the Block scenario provided by the AirSim platform as the virtual scenario for agent activities. The Block scenario is a virtual environment built on the Unreal Engine, which provides a simplified but fully functional test site that can be used to simulate the behavior of typical agents such as drones in various typical tasks. Specifically, the virtual environment can be constructed through the following steps: (1-1) Modify the settings.json file to batch modify the number, type, and initial positions of agents in the Block environment. In this embodiment, a formation scenario including 1 leader drone and 10 follower drones can be set, that is, the total number of agents in the Block scenario is set to 11. Secondly, set the yaw angle of each follower drone to 0, the Z-axis height is fixed, and the X-axis and Y-axis coordinate values are determined according to the number and total number of follower drones. Specifically, on the XY plane, the leader drone uses a random position initially, and the follower drones should be evenly distributed at equal angular intervals on a circle centered on the leader with a radius of 2 times the expected distance between the follower and the leader. Set the initial angle offset , then the expected angle of each follower is:
[0025] where is the total number of drones in the sub-formation.
[0026] (1-2) Connect to, control the movement of, and obtain agent information for the drones in the Block environment through the interfaces provided by AirSim. Specifically, the MultirotorClient function provided by the airsim library can be used to establish a connection with the drones in the Block environment, the simGetGroundTruthKinematics function can be used to obtain the true state information of the drones (including their own state information and the observation information obtained by perceiving the environment and other drones), and the moveByVelocityZAsync function can be used to control the actions of the drones (including speed, etc.).
[0027] After constructing the virtual environment, the multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning can be executed in the virtual environment. In this embodiment, the multi-agent formation control method based on the dynamic graph attention network and deep reinforcement learning is executed using a graph attention network and an independent proximal policy optimization model. The overall structure composed of the graph attention network and the independent proximal policy optimization model is as Figure 1 shown.
[0028] Refer to Figure 1, the Graph Attention Network (GAT) and the Encoder together form the perception compression model DGAT. For the data to be processed from the original nodes (specifically, the observation information obtained by the original nodes through perceiving the environment or other nodes), it can be encoded by the Encoder to obtain and and other features of the original nodes. Based on the same principle, the features of neighbor nodes and etc. can also be obtained. The Graph Attention Network GAT performs feature aggregation on the features of the original nodes and the features of the neighbor nodes, and finally obtains the output vector .
[0029] In this embodiment, the specific structure of the perception compression model DGAT is as shown in Figure 2 . Referring to Figure 2 , first, the Encoder encodes the input data to be processed from the original nodes into the vector matrix of the original nodes, the adjacency matrix and the stable feature according to the preset encoding rules.
[0030] Specifically, the stable feature of the original nodes represents information such as the relative position of the leader UAV perceived and observed by the original nodes and the speed of the leader UAV; the vector matrix of the original nodes represents the relative position information between each other node (specifically, other UAVs or obstacles in the same multi-agent formation) perceived and observed by the original nodes and the original node itself; the size of the adjacency matrix of the original nodes is , where is the total number of agents (UAVs) in the same multi-agent formation, is the total number of obstacles, and the adjacency matrix contains multiple elements, and each element represents the adjacency relationship between the original node and a corresponding other node (other UAV or obstacle), specifically: if an other node (other UAV or obstacle) is within the perception range of the original node, that is, the original node can establish a connection with this other node, then the element value corresponding to this other node in the adjacency matrix of the original node is 1, conversely, that is, if this other node (other UAV or obstacle) is outside the perception range of the original node, that is, the original node cannot establish a connection with this other node, then the element value corresponding to this other node in the adjacency matrix The element value corresponding to this other node in the middle is 0.
[0031] In this embodiment, the vector matrix of the original node , the adjacency matrix and the stable features together constitute the features of the original node. Similarly, for other nodes, such as the node with the serial number , there is also its corresponding vector matrix , the adjacency matrix and the stable features , which together constitute the features of the node with the serial number .
[0032] In this embodiment, the attention coefficient is calculated for the original node (that is, the node with the serial number ) and its neighbor nodes. In this embodiment, the attention coefficient between the original node and the neighbor node with the serial number can be calculated according to the following formula: (1) Among them, represents the features of the original node (including the vector matrix , the adjacency matrix and the stable features , etc.), represents the features of the neighbor node with the serial number . is a trainable weight matrix that can map the features of the node to a new space. Therefore, is the feature obtained by mapping the features of the original node, is the feature obtained by mapping the features of the neighbor node with the serial number . is a trainable vector used to perform a linear transformation on the concatenated features . represents performing a softmax operation by traversing all the values of to ensure that all the are traversed, and the sum of
[0033] In this embodiment, the attention coefficient between the original node and the neighbor node with the serial number can also be calculated according to the following formula: (2) Among them represents the original node (with the serial number The set of all neighbor nodes of the node). Formula (2) is equivalent to Formula (1).
[0034] In this embodiment, the feature update rule of each layer of the graph attention network GAT can be expressed as: (3) To enhance the expressive power and stability of the model, the multi-head attention mechanism can be used in the graph attention network GAT. Under this mechanism, multiple attention heads independently calculate their respective attention coefficients and feature aggregations, and then the results are combined (usually concatenated or averaged). Using the multi-head attention mechanism, the feature update rule of each layer of the graph attention network GAT can be expressed as: (4) Among them, represents the output result of the -th layer in the graph attention network, is the number of attention heads, represents the set of all neighbor nodes of the original node, is the -th attention coefficient calculated by the attention head, represents the concatenation operation in the feature dimension, represents the parameters of the -th layer in the graph attention network, represents the output result of the -th layer in the graph attention network, represents an activation function such as tanh or ReLU.
[0037] In this embodiment, as Figure 2 shown, a graph attention network GAT with 2 layers is used to aggregate and update the features of the nodes, and 2 attention heads are used in the first layer of the graph attention network GAT for information perception and aggregation.
[0038] The features of the original node are represented as , and after being processed by the first layer (Layer1 in Figure 2 ) of the graph attention network GAT, it serves as the input node vector Figure 2 of the second layer (Layer2 in ) of the graph attention network GAT. This process can be expressed as: (5) The output dimension of the first layer of the graph attention network GAT is 64, and the ELU function is used as the activation function. The second layer of the graph attention network GAT serves as the output layer, with only one attention head, an input dimension of , and an output dimension of 4, using the ELU activation function.
[0040] Finally, the high-dimensional input feature vector, which is the feature of the original node , after being weighted and aggregated through the multi-head attention mechanism, is mapped to a lower-dimensional feature space. This process can be described as: (6) where is the node vector output by the graph attention network GAT, represents the processing performed by the graph attention network GAT, is the vector matrix of the original nodes, is the adjacency matrix of the original nodes, is the vector matrix the row vector feature dimension of, is the compressed node vector the dimension of. The graph attention network GAT in this embodiment can ensure that the output node vector contains sufficient information representation while significantly reducing data storage and transmission overhead.
[0042] In this embodiment, referring to Figure 1 and Figure 2 , the perception compression model DGAT concatenates the stable features of the original nodes directly output by the encoder Encoder, and the node vector obtained by the graph attention network GAT through feature aggregation processing of the original nodes to obtain the output vector corresponding to the original nodes, that is, (7) In this embodiment, referring to Figure 1 , after using the perception compression model DGAT to obtain the output vector corresponding to the original nodes, the output vector is input into the independent proximal policy optimization model for processing.
[0044] Referring to Figure 1 , the independent proximal policy optimization model (Independent Proximal Policy Optimization, IPPO) includes a first neural network Actor and a second neural network Critic.
[0045] In this embodiment, both the first neural network Actor and the second neural network Critic are neural networks.
[0046] Among them, the structure of the first neural network Actor is as Figure 3 shown. Referring to Figure 3 , the first neural network Actor consists of 2 fully connected layers ( Figure 3 Layer1 and Layer2 in ), both hidden layers contain 32 neurons, and the Tanh activation function is used between layers to enhance the non-linear expression ability; the last layer network in the first neural network Actor contains 5 neurons, and through the Softmax activation function, the vector input to the first neural network Actor is mapped into an action probability distribution .
[0047] The first neural network Actor implements a policy function , specifically, the policy function learns a mapping from the observation to a distribution in a range (discrete action space) or to the action mean and variance of a Gaussian function (continuous action space) for subsequent action sampling. That is, the first neural network Actor processes the input vector through the policy function and finally outputs the action probability distribution .
[0048] In this embodiment, the meaning of the action probability distribution is: the probability that the original node (the first agent) selects to execute the action (that is, the observed data to be processed ) in the state (specifically including values such as moving forward, backward, left, right, and not moving, etc.), where means the angle formed by the changed direction vector and the pre-changed direction vector if the original node (the first agent) selects to execute the action .
[0049] The structure of the second neural network Critic is as Figure 4 shown. Referring to Figure 4 , the second neural network Critic also includes 2 fully connected layers ( Figure 4 Layer1 and Layer2 in ), both hidden layers contain 32 neurons, and the Tanh activation function is used between layers to enhance the non-linear expression ability; the output layer of the second neural network Critic has only 1 neuron and no activation function is used. The input of the second neural network Critic is the output vector corresponding to the original node (the first agent) The output is the status value , and the status value is used to measure the overall benefit of the current strategy in a given state .
[0050] In this embodiment, referring to Figure 1 , the status value corresponding to the original node (the first agent) is and can be concatenated with the output vector and then input into the first neural network Actor for the first neural network Actor to process and obtain the action probability distribution
[0051] Based on Figures 1 - 4 the graph attention network and the independent proximal policy optimization model shown, as Figure 5 shown, the following steps can be executed S101. Obtain the graph attention network S102. Obtain the independent proximal policy optimization model S103. Obtain the observation information of the first agent. Taking the first agent as the original node, the second agent as the neighbor node, and the observation information of the first agent as the data to be processed, input them into the graph attention network and the independent proximal policy optimization model for processing S104. Control the first agent according to the action probability distribution output by the independent proximal policy optimization model
[0052] In steps S101 - S102, the graph attention network GAT and the independent proximal policy optimization model to be obtained are as Figure 1 shown, where the specific structures of the graph attention network GAT and its composed perception compression model DGAT are as Figure 2 shown, the specific structure of the first neural network Actor in the independent proximal policy optimization model is as Figure 3 shown, and the specific structure of the second neural network Critic in the independent proximal policy optimization model is as Figure 4 shown
[0053] In step S103, determine a specific agent in the multi - agent formation, that is, the first agent. The first agent detects and perceives its surrounding environment and the surrounding agents, so as to obtain the observation information of the first agent ; take the observation information of the first agent as the data to be processed and input it into the graph attention network and the independent proximal policy optimization model for processing. According to Figure 1According to the working principles of the graph attention network and the independent proximal policy optimization model shown, the first agent will serve as the original node to be processed when the graph attention network and the independent proximal policy optimization model are working, and other agents in the same multi-agent formation as the first agent, namely the second agent, will serve as the neighbor nodes to be processed when the graph attention network and the independent proximal policy optimization model are working. The observation information of the first agent is processed by the graph attention network GAT and the independent proximal policy optimization model and finally the independent proximal policy optimization model outputs the action probability distribution .
[0054] In this embodiment, steps S101 - S104 constitute one round of processing, and multiple rounds of processing can be executed. Different processing rounds can be distinguished by serial numbers . Since the execution speed of one round of processing is relatively fast, it can be considered that one round of processing is completed at one moment. Therefore it can also be understood as the execution moment of one round of processing and each step therein.
[0055] In this embodiment, when step S104 is executed in any round of processing, that is, when controlling the first agent according to the action probability distribution output by the independent proximal policy optimization model, the following steps can be specifically executed: S10401. Set the reward function of the first agent; S10402. Determine the target action with the goal of maximizing the discounted cumulative reward; the discounted cumulative reward is determined according to the action probability distribution and the reward function; S10403. Control the first agent to execute the target action.
[0056] Taking the execution of the th round of processing (or the processing at moment ) as an example, steps S10401 - S10403 are described.
[0057] In step S10401, a reward function is set for the original node (the first agent). In this embodiment, the set reward function includes at least one of proximity reward, position reward, orientation reward, and collision avoidance reward.
[0058] (2 - 1) Set the proximity reward: If the original node (the first agent) at moment (or when executing the th round of processing) compared to - 1 moment (or when executing the When it is closer to its desired position during the -1 round of processing, a positive reward that is positively correlated with its moving distance is obtained; conversely, if it is farther from the desired position, a negative penalty is obtained. Based on the above idea, the proximity reward is defined as: (8) where is the position of the original node (the first agent) at time -1 , is the position that the original node (the first agent) can reach by executing an action at based on time, is the desired position of the original node (the first agent) at time.
[0060] (2-2)Set the position reward to penalize the behavior of the original node (the first agent) staying stationary when it has not reached the target point. The farther the current position of the original node (the first agent) is from the target point, the greater the penalty. Based on the above idea, the position reward is defined as: (9) (2-3)Set the direction reward to penalize the behavior of the original node (the first agent) whose speed direction deviates from the desired direction. The greater the deviation angle, the greater the penalty. Based on the above idea, the direction reward is defined as: (10) where is the direction vector from the position of the original node (the first agent) at time -1 to its target point position, and the angle formed by the direction vector from the position of the original node (the first agent) at -1 to its position at time. Therefore can be defined as: (11) (2 - 4) Set a collision avoidance reward to punish the risk of collision between the original node (the first agent) and other nodes (the second agent) or obstacles. If other nodes (the second agent) or obstacles are within the danger radius of the original node (the first agent), the closer the original node (the first agent) is to other nodes (the second agent) or obstacles, the greater the punishment the original node (the first agent) receives. Based on the above idea, the collision avoidance reward is defined as: (12) Among them, and are coefficients for adjusting the punishment intensity. represents the set of all obstacles sensed by the original node (the first agent) at time , represents the distance between the original node (the first agent) at time and other nodes (the second agent) ; represents the distance between the original node (the first agent) at time and the obstacle ; represents the safe distance, which is a fixed value. When the distance or is less than the safe distance , the value of the exponential function increases, resulting in an increase in the collision punishment.
[0065] After setting each reward function in step S10401, execute step S10402. In step S10402, set the discounted cumulative reward as follows: (13) In formula (13), is the discount factor, which is used to control the influence degree of future expected rewards and can be set to 0.9; represents that if the original node (the first agent) (given that the observed information of the original node is ) at time , selects the action from the action probability distribution obtained according to step S103 and executes the action, generating the values of each reward function determined by and , that is, Specifically, it can be the proximity reward , position reward , orientation reward , collision avoidance reward , etc. indicates that the observation information is , and the action is in the case of the mathematical expectations of various forms (approach reward, position reward, orientation reward, collision avoidance reward, etc.), so as to obtain the discounted cumulative reward in the case where the action is .
[0067] In step S10402, aiming at maximizing the discounted cumulative reward, for example, traversing the action probability distribution obtained in step S103 for all actions , select the action that can make the discounted cumulative reward reach the maximum value as the target action. In step S10403, control the first agent to execute the target action determined in step S10402, such as moving forward, backward, left, right or not moving, to achieve the control of the first agent.
[0068] In this embodiment, by using the graph attention network and the independent proximal policy optimization model, the multi-agent formation can be processed as a multi-node network. Specifically, the observation information obtained by a specific agent (the first agent) in the multi-agent formation by perceiving the outside world (including other agents or obstacles) is processed, so as to obtain the action probability distribution of the specific agent, and the control of the specific agent is realized; in the process of controlling the multi-agent formation, focusing on a specific agent (the first agent) and its neighbor nodes (the second agent), so the observation information detected by some agents can be used to realize the control of the multi-agent formation, which can reduce the demand for global information, thus facilitating the improvement of the real-time performance of control and reducing the possibility of uncontrollability caused by the failure of some agents, and realizing the effective control of the multi-agent formation.
[0069] In the actual scenario or simulation environment, each time the processing process, i.e., steps S101 - S104, is executed, the first agent can be controlled to execute one step of action. By repeatedly executing the processing process multiple times, the first agent can be controlled to execute consecutive multi-step actions. Moreover, since the first agent can be arbitrarily selected in the multi-agent formation, the processing process can be respectively executed for some or all of the agents in the multi-agent formation, so as to realize the control of the multi-agent formation.
[0070] In this embodiment, before or after running the graph attention network and the independent proximal policy optimization model, the graph attention network and the independent proximal policy optimization model can be trained. The training of the graph attention network and the independent proximal policy optimization model includes the following steps: S201. Execute at least one episode; each episode includes a plurality of consecutive processing procedures; S202. For any one episode, after the episode is executed, determine the first loss function and the second loss function according to the execution results of the processing procedures in the episode; S203. With the goal of minimizing the first loss function, update the parameters of the first neural network; S204. With the goal of minimizing the second loss function, update the parameters of the second neural network; S205. When the training end condition is not met, execute the next episode, and when the training end condition is met, end the execution of all episodes; wherein, the training end condition is that the number of executed episodes reaches a preset upper limit, or the reward functions obtained in the executed episodes converge.
[0071] Before executing steps S201 - S205, first perform network initialization, and use Xavier to initialize the parameters of the Actor network and the Critic network to ensure the stability of the network in the initial stage of training; then, according to the characteristics of the perception compression model DGAT and the independent proximal policy optimization model IPPO, set necessary hyperparameters, including the learning rate, batch size, discount factor (discount factor), entropy coefficient and gradient clipping threshold, etc. For example, the initial learning rate is set to 1×10 -4 and gradually decreased to the minimum value of 1×10 -7 in an annealing decay manner to improve the stability in the later stage of training. The maximum number of steps of the agent within one episode is set to 50. During the training process, mini - batch is not used, but the parameters are updated with the complete data. The optimizer is selected as Adam to improve the training efficiency with its strong adaptability.
[0072] In step S201, execute multiple episodes, and set a maximum number of steps for each episode, that is, each episode executes processing procedures such as steps S101 - S104.
[0073] Taking one episode as an example for illustration. After executing After a processing procedure, step S202 is executed to determine a first loss function and a second loss function according to the execution results of the processing procedures in the round.
[0074] In this embodiment, the first loss function can be designed based on the policy gradient theory, and the form of the designed first loss function is: (14) In practical applications, the expected function can be expanded and calculated to obtain the first loss function in the following form: (15) In formula (15), is the preset number of training rounds, that is, the number of processing procedures executed in a round, and its value can be set to 50; is the rd processing procedure in the round (or the processing procedure executed at time), is the probability ratio of the new and old policies, satisfies (16) Among them, represents the probability of the action output with the updated parameters (new policy) if step S203 is executed to update the parameters of the first neural network, represents the probability of the action output with the original parameters (old policy) if the parameters of the first neural network are not updated; represents the execution of the clip operation to limit the amplitude of policy update and ensure the stability of the training process, where is a constant and can be set to 0.2; is the advantage function, and in this embodiment, the Generalized Advantage Function (GAE) is used to calculate , that is (17) In formula (17), is at a certain moment in a complete trajectory , , the trajectory can be expressed as , that is, the sequence of states and actions of the original node (the first intelligent agent) . is a GAE hyperparameter used to control the smoothness of the advantage function, and its value is set to 0.95 in this embodiment; is the discount factor used to control the influence degree of future expected rewards, and its value is set to 0.9 in this embodiment. is the temporal difference error at time (18) where is the original node (the first agent) at time, the immediate reward obtained, is the value estimate at time time, that is, the processing result of the second neural network Critic on the output vector at time is the entropy loss coefficient, and its initial value can be set to 0.5 and linearly decreased to 0.01 as the number of training times (the cumulative number of rounds of the execution process) increases; is the entropy of the policy, indicating the randomness of the policy distribution, which can be calculated by the following formula: (19) In this embodiment, a second loss function can be designed, and the mean square error is used to measure and minimize the difference between the predicted value and the true value. The form of the designed second loss function is: (20) In practice, can be replaced by because is the "correction" of the state value calculated from the TD error. In this way, the estimation of the value function can be made more accurate by minimizing the mean square error, and the following form of the second loss function can be obtained: (21) In formula (21), represents the state value (value estimate) output after processing with the updated parameters if the second neural network is updated with parameters in step S204, in represents the state value (value estimate) output after processing with the original parameters if the second neural network is not updated with parameters.
[0083] In step S203, the parameters of the first neural network Actor are updated using the policy gradient method. The update objective is to minimize the value of the first loss function calculated according to formula (15) after updating the parameters of the first neural network Actor.
[0084] In step S204, the parameters of the second neural network Critic are updated using the temporal difference method. The update objective is to minimize the value of the second loss function calculated according to formula (21) after updating the parameters of the second neural network Critic.
[0085] In step S205, it is judged whether the training end condition is satisfied after executing steps S201 - S204 of this round. In this embodiment, the training end condition is specifically "the number of executed rounds reaches the preset upper limit", or "the reward functions obtained in the executed rounds converge".
[0086] If the training end condition is satisfied, the execution of all rounds ends, the next round is no longer executed, and the training is completed; if the training end condition is not satisfied, the processing process of the next round is executed, and the training continues.
[0087] In this embodiment, by executing steps S201 - S205, the Figure 1 DGAT - IPPO model (composed of a graph attention network and an independent proximal policy optimization model) in Figure 6 is trained as shown by the green curve in Figure 6 . Referring to Figure 6 , under the condition that there are 10 follower drones in the multi - agent formation, it can be seen that by executing steps S201 - S205, as the number of episodes increases, the value of the obtained reward function also increases, and
[0088] In Figure 1 , the training result of a single IPPO model shown by the blue curve is set as a control. It can be seen that at the same number of episodes, the value of the reward function obtained by training the DGAT - IPPO model is significantly greater than the value of the reward function obtained by training a single IPPO model. ① Verification of the perception data compression ability: Run the DGAT model in different task scenarios, and record the memory occupation of the input data stream and the output data stream respectively.
[0089] ② Decision generation time test: The complete time from perception input to action generation of the Actor network of the DGAT-IPPO model and the IPPO model is tested respectively to evaluate the performance of the model in real-time response tasks. In this embodiment, the results of the data memory usage test on the DGAT-IPPO model are as follows: Figure 7 As shown. Figure 7 It can be seen that the memory usage of the DGAT-IPPO model is relatively stable and has good performance.
[0090] In this embodiment, the test results of the decision-making time of the DGAT-IPPO model are shown in Table 1.
[0091] Table 1
[0092] Table 1 provides the decision time of applying a single IPPO model. According to Table 1, it can be seen that the DGAT-IPPO model in this embodiment can significantly reduce the required decision time.
[0093] The DGAT-IPPO model that has been trained and verified can be used to execute steps S101-S104, thereby controlling multi-agent formations in actual applications or simulation scenarios.
[0094] A computer program that executes the multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning in this embodiment can be written and written into a computer device or storage medium. When the computer program is read out and executed, the multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning in this embodiment is executed, thereby achieving the same technical effect as the multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning in the embodiment.
[0095] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. In addition, the descriptions of up, down, left, right, etc. used in the present disclosure are only relative to the relative positional relationship of the components of the present disclosure in the accompanying drawings. The singular forms of "a", "" and "the" used in the present disclosure are also intended to include the plural forms, unless the context clearly indicates other meanings. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by those skilled in the art. The terms used in the specification of this embodiment are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any combination of one or more related listed items.
[0096] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided in this embodiment is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.
[0097] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.
[0098] In addition, the operations of the processes described in this embodiment can be performed in any suitable order, unless this embodiment otherwise indicates or is otherwise clearly contradictory to the context. The processes described in this embodiment (or variations and / or combinations thereof) can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed together on one or more processors, by hardware, or a combination thereof. A computer program includes a plurality of instructions executable by one or more processors.
[0099] Further, the method can be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and, when the storage medium or device is read by the computer, can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. When such media include instructions or programs that implement the above steps in conjunction with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0100] A computer program can be applied to input data to perform the functions of this embodiment, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0101] The above are only the preferred embodiments of the present invention. The present invention is not limited to the above embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.
Claims
1. A multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning, characterized in that: The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning includes at least one processing process, and any of the processing processes includes the following steps: Obtaining a graph attention network; the graph attention network is used to perform feature aggregation processing on the to-be-processed data from the original node and the neighboring nodes to obtain an output vector; Acquire an independent proximal policy optimization model; the independent proximal policy optimization model includes a first neural network and a second neural network, the second neural network is used to process the output vector to obtain a state value; the first neural network is used to process the output vector to obtain an action probability distribution; Obtaining observation information of a first agent, taking the first agent as the original node, taking the second agent as the neighbor node, and taking the observation information of the first agent as the data to be processed, inputting the data into the graph attention network and the independent proximal strategy optimization model for processing; the first agent and the second agent are different agents in the same multi-agent formation; The first agent is controlled according to the action probability distribution output by the independent proximal strategy optimization model.
2. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 1 is characterized in that: The process of performing feature aggregation processing on the data to be processed from the original node and the neighboring nodes to obtain an output vector includes: Encoding the data to be processed to obtain the features of the original node; the features of the original node include a vector matrix, an adjacency matrix and a stable feature; For any neighbor node, determining an attention coefficient between the original node and the neighbor node according to the characteristics of the original node and the characteristics of the neighbor node; According to each of the attention coefficients, determining the feature update rules of each layer of the graph attention network through a multi-head attention mechanism; According to each of the feature update rules, multi-layer mapping processing is performed on the vector matrix and the adjacency matrix to obtain a node vector; The stable feature and the node vector are concatenated to obtain the output vector.
3. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 2 is characterized in that: The determining, according to the feature of the original node and the feature of the neighboring node, an attention coefficient between the original node and the neighboring node comprises: According to the formula Calculate; where represents the characteristics of the original node, Indicates the serial number is The characteristics of the neighbor nodes, is a trainable vector, is a trainable matrix, Represents traversal All values of perform softmax operation, Indicates that the original node and the sequence number are The attention coefficient between the neighbor nodes.
4. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 2 is characterized in that: According to each of the attention coefficients, determining the feature update rules of each layer of the graph attention network through a multi-head attention mechanism includes: Determine the feature update rule as in, represents the first The output of the layer, is the number of attention heads, represents the set of all the neighbor nodes of the original node, It is The attention coefficient calculated by the attention head is represents the concatenation operation on the feature dimension, represents the first The parameters of the layer, represents the first The output of the layer, Represents the activation function.
5. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 1 is characterized in that: The controlling the first agent according to the action probability distribution output by the independent proximal strategy optimization model comprises: Setting a reward function for the first agent; the reward function includes at least one of a proximity reward, a position reward, a direction reward, and a collision avoidance reward; Determine a target action with the goal of maximizing the discounted cumulative reward; the discounted cumulative reward is determined according to the action probability distribution and the reward function; Control the first agent to perform the target action.
6. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 5 is characterized in that: The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning also includes: Execute at least one round; each of the rounds includes a plurality of the processing processes executed in succession; For any of the rounds, after the round is executed, a first loss function and a second loss function are determined according to the execution results of each of the processing processes in the round; Taking minimizing the first loss function as a goal, updating parameters of the first neural network; Taking minimizing the second loss function as a goal, updating parameters of the second neural network; When the training end condition is not met, the next round is executed, and when the training end condition is met, all rounds are terminated; wherein the training end condition is that the number of rounds executed reaches a preset upper limit, or the reward functions obtained in the executed rounds converge.
7. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 6 is characterized in that: Determining the second loss function and the first loss function according to the execution results of each of the processing processes in the round includes: Determine the first loss function as Determine the second loss function as 。 8. The multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning according to claim 6 or 7, characterized in that: The updating of parameters of the first neural network includes: Updating the parameters of the first neural network using a policy gradient method; The updating of parameters of the second neural network includes: The parameters of the second neural network are updated using a temporal difference method.
9. A computer device, characterized in that: It includes a memory and a processor, the memory is used to store at least one program, and the processor is used to load at least one program to execute the multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning as described in any one of claims 1-8.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor executable program is used to execute the multi-agent formation control method based on dynamic graph attention network and deep reinforcement learning as described in any one of claims 1-8 when executed by the processor.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle formation cluster control method based on multi-agent deep reinforcement learning
CN115755949A
Unmanned aerial vehicle cluster collaborative confrontation method based on graph attention reinforcement learning
CN116841317A
Unmanned aerial vehicle group cooperative combat method based on multi-agent reinforcement learning
CN117055623A
Heterogeneous unmanned aerial vehicle formation autonomous decision-making implementation method based on cooperation of units and knowledge enhancement
CN118838408A
Hypersonic flight vehicle adaptive attitude control method and system based on near-end strategy optimization
CN119270909A