Unmanned aerial vehicle cluster collaborative confrontation method introducing action interaction and credit distribution
By introducing action interaction and credit allocation methods, the problems of interaction and credit allocation in the drone cluster are solved, and effective simulation and collaboration efficiency of the drone cluster collaborative confrontation are achieved.
Patent Information
- Application Number
- CN202510102252.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In drone clusters, how to effectively manage interaction and credit distribution between drones, especially among agents with different capabilities and contribution degrees, traditional methods have problems of unfair credit distribution or inefficiency, which affects collaboration efficiency.
Introduce the method of action interaction and credit distribution, by constructing an agent network training model, using the self-attention mechanism to deal with the mismatch between observations and actions, dynamically adjust the agent's decisions, and use the double-layer hybrid network structure to perform credit distribution.
Effective simulation of the coordinated confrontation of drones clusters has been realized, the interactive coordination capabilities and collaboration efficiency between drones have been improved, and the cooperation efficiency of each drone has been ensured that each drone can receive appropriate rewards during the collaboration process, and promote the improvement of system performance.
Smart Images

Figure CN119937591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air-to-air confrontation game decision-making of drone clusters, and in particular to a drone cluster collaborative confrontation method introducing action interaction and credit allocation. Background Art
[0002] As a cutting-edge field, drone swarm technology is increasingly becoming a strategic high ground for international military powers to explore and deploy. China has also kicked off the research on the concept of drone swarm confrontation, fully invested in the wave of key technological breakthroughs, and continuously verified its feasibility and effectiveness through a series of swarm flight tests. At the same time, countries around the world have followed closely, and have carried out system integration and empirical exploration in response to the diverse needs of confrontation tasks, striving to make breakthrough progress in the application of drone swarm technology. The in-depth development of this series of research and experiments not only profoundly reveals the unlimited potential of drone swarm technology in modern warfare, but also indicates that drone technology has evolved from a traditional single confrontation platform to an intelligent, collaborative, and clustered confrontation system, and its importance at the strategic level is becoming increasingly prominent. With its high flexibility, autonomy, and survivability, drone swarm technology heralds a new change in the future form of warfare, which will bring unprecedented tactical advantages and strategic impacts to military confrontation.
[0003] In a drone swarm, there is cooperation and competition between multiple drones. These cooperation and competition are not simple addition and subtraction, but a highly complex and dynamic system. The drones in the system have different goals, capabilities and information resources. How to effectively manage the interaction between drones is a challenging problem. First, as the number of drones increases, the coordination and communication costs will increase rapidly; second, under incomplete information conditions, drones need to learn global information from local observations in order to make reasonable decisions. In the drone swarm system, determining the contribution of each drone to the global reward, that is, how to allocate credit, is a crucial issue. The purpose of credit allocation is to ensure that each drone can obtain appropriate rewards in the process of cooperation or common goal achievement to encourage cooperation and promote the improvement of system performance. However, traditional credit allocation methods often have problems of unfair or inefficient credit allocation, especially when facing agents with different capabilities and contribution levels. The importance of this problem lies in that without effective credit allocation, drones may lack the motivation to actively participate in cooperation or share resources, thereby affecting the collaborative efficiency of the drone swarm system. In addition, the interaction between drones is also an important challenge. Drones need to coordinate actions through effective interaction to achieve common goals. Different behaviors of drones will have different effects on other agents, which may be direct, indirect or even long-term. However, existing methods often perform poorly on drone interactions and fail to fully utilize the information of each agent, resulting in unsatisfactory task completion. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a drone cluster collaborative confrontation method that introduces action interaction and credit allocation in response to the above-mentioned deficiencies in the prior art, to realize drone cluster collaborative confrontation simulation, and to provide a method for establishing drone cluster confrontation game in a dynamic environment.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: a drone cluster collaborative confrontation method introducing action interaction and credit allocation, comprising the following steps:
[0006] Step 1: Model the confrontation space, action space, state space and observation space of the drone cluster respectively, and set the confrontation parties of the drone cluster;
[0007] The drone cluster is composed of multiple drones; in the confrontation scenario, the drones of both sides adopt a six-degree-of-freedom dynamics model, use dynamic equations to describe the movement of the drones in the continuous action space, adopt a continuous action set, and use vectors to represent the continuous action parameters of the drones;
[0008] Describe the drone swarm confrontation task as a tuple The decentralized part consists of an observable Markov decision process, where is the set of intelligent agents, is the state set, is the action set, P is the state transition function, r is the reward function, Ω is the observation result set, O is the observation function, γ is the discount factor, γ∈[0,1]; set To describe the global state of the environment, at each time step t, each agent Select an action And generate a joint action U is the set of joint actions; according to the state transfer function P(s′|s,u): The environment enters the next state, causing the joint action u to change, where s′ is the state of the agent at the next moment; all agents share the same reward function r(s,u): is a set of real numbers. The function of r(s,u) is to evaluate the state s and the joint action u, and return a real value as a reward to measure the contribution of the state and action. The common goal of all factors in the tuple is to maximize the expected reward. in Denotes the discounted sum of future rewards G t To find the expectation, the specific formula is:
[0009]
[0010] Among them, k is the step size after time t, r t+k is the instant reward obtained at time t+k;
[0011] Step 2: Consider the action interaction between drones, regard drones as independent and autonomous intelligent agents, build an intelligent agent network training model, determine whether the action of the intelligent agent only affects the environmental information of the adversarial space or whether the action of the intelligent agent also affects other intelligent agents, and dynamically adjust the decision of the intelligent agent according to the environmental information of the adversarial space; including the following steps:
[0012] Step 2.1: Build an agent network training model for each agent to learn the agent's strategy and decision-making;
[0013] Construct an agent network training model to enable agents to learn adaptive strategies in confrontation and achieve optimal strategies in an environment of mutual influence;
[0014] Divide the observations o and actions u input into the agent network training model, and divide the initial observations into three parts: the part related to the environment, the part related to itself, and the part related to other agents;
[0015] The observation value o of the agent is converted from sequence form to matrix form to make the representation of the observation-state space unified and comparable; the action u is divided into two parts: and action Only affects the environment information and the agent itself; action Influence other agents;
[0016] Step 2.2: Divide the agent network into different sub-modules, evaluate the local Q value of the agent action by the impact of the agent action on other agents, and calculate the global Q value of the agent network from the local Q values of all agents;
[0017] The agent network is divided into different sub-modules. By the influence of the agent action on other agents, it is judged whether the agent action only affects the environmental information of the adversarial space or the agent itself also affects other agents. If the agent action only affects the environmental information of the space and the agent itself, the initial observation value of the agent is input into the embedding layer in the agent network training model to evaluate the local Q value of the agent action. If the agent action affects other agents, the observation value o of agent i on agent j is i,j The action of agent i affecting agent j is input into the embedding layer of the agent network training model to evaluate the local Q value of the agent action; the global Q value of the agent network is calculated from the local Q values of all agents;
[0018] Introduce transformer into the agent network training model, use the self-attention mechanism to identify and filter out useful information, and weight the observations according to the useful information obtained; focus on the input features that are helpful to the agent, suppress attention to irrelevant features, and discard useless information that is not needed for decision making;
[0019] The useful information identified and filtered by the self-attention mechanism is observation data or features that are of substantial help and value to the agent in making decisions or performing tasks, and data perceived by the agent in the environment of the adversarial space that can directly or indirectly affect the agent's behavior, decision-making process or task completion;
[0020] Step 2.3: Use reinforcement learning algorithm to train the agent network training model;
[0021] Step 2.4: The agent perceives and understands the actions of other agents and dynamically adjusts the agent's decision based on the environmental information of the adversarial space;
[0022] Step 3: The two adversaries in the drone cluster confront each other, and at the same time evaluate and reason about the adversarial situation, receive the local Q value generated by the agent network training model as input, and establish a two-layer hybrid network structure using the monotonic value function decomposition method to obtain the global Q value Q of the agent network. tot , and then obtain the optimal joint action of the agent;
[0023] The two adversaries confront each other through the set adversarial space and intelligent agent network training model; both adversaries are committed to comprehensively collecting adversarial data, observing and recording the actions and environment of the adversary, estimating the state, and then evaluating the adversarial situation; based on the evaluation results of the adversarial situation, dynamically adjust the reasoning and decision-making of the intelligent agent;
[0024] The centralized training distributed execution paradigm is adopted, and the local Q values of each agent are integrated and constrained through a two-layer hybrid network to generate a global Q value. Each agent has an independent Q network and can make independent decisions based on its own local information. The output of the Q network is transmitted through a two-layer hybrid network to achieve the mapping of local Q values to global Q values.
[0025] The two-layer hybrid network is responsible for credit allocation, taking the local Q value generated by the agent network as input, and processing it through two separate super networks and activation functions contained in the two-layer hybrid network to output a global Q value Q tot , the specific formula is:
[0026] Q tot (τ,u,s;θ)=W2R eLU (W1Q+b1)+b2
[0027] in, is the local Q value output by the agent network, Q n (τ n ,u n ) is the local Q value of the nth agent, τ n ,u n are the observation history and action of the nth agent, τ is the observation history, s is the environment state, θ is the parameter of the target network, and R eLU is the activation function of W1Q+b1, and The weight matrices and bias terms generated for two separate hypernetworks satisfy: is a real matrix, m and n are real matrices respectively The row and column dimensions of the vector;
[0028] The loss function of the two-layer hybrid network is updated by minimizing the squared TD error. The loss function formula is as follows:
[0029]
[0030] in, is the TD target, τ′,u′,s′ are the observation-action history, joint action and state at the next moment, θ - are the target network parameters, which are periodically copied from θ;
[0031] The two separate hypernetworks dynamically adjust their own parameters or structures according to the input observation values to adapt to different environments;
[0032] Step 4: The two adversaries conduct collaborative task planning to provide the optimal strategy for multiple agents; update the local Q value of the agent and optimize the agent strategy through the credit allocation method. The specific method is as follows:
[0033] According to step 3, the global Q value Q tot , using the back propagation of the agent network training model to update the local Q value of the agent; using the gradient entropy to update the global Q value of the agent for distinguishable credit assignment and policy optimization;
[0034] The concept of normalized gradient entropy is introduced in the two-layer hybrid network to measure the distributability of credit; according to the global Q value Q tot The expanded result of , calculates the gradient:
[0035]
[0036] So we get:
[0037] dQ tot =W2FW1dQ
[0038]
[0039] in, is a diagonal matrix, and x c is the cth element output in the double-layer hybrid network, X is the input feature of the double-layer hybrid network, x1, x2, ..., x m Both are the outputs of the first layer linear transformation in the two-layer hybrid network, calculated by W1Q+b1;
[0040] Introducing loss functions for action interaction and credit assignment Including the original TD loss and regularized gradient entropy Two parts, as shown in the following formula:
[0041]
[0042] Among them, λ is a hyperparameter used to balance the TD loss and regularization term, and batchsize is the batch size of sampling;
[0043] Step 5: Repeat step 4 until the confrontation between the two parties ends, observe the confrontation performance of introducing action interaction and credit allocation, and complete the confrontation simulation of the drone cluster.
[0044] The beneficial effects of adopting the above technical solution are: the drone cluster collaborative confrontation method introducing action interaction and credit allocation provided by the present invention has the following advantages:
[0045] (1) By defining comprehensive adversarial space, action space, state space, and observation space, we ensure realistic and thorough modeling of UAV swarm adversarial scenarios; this detailed modeling facilitates more effective strategy development and accurate performance evaluation;
[0046] (2) The action matching mechanism proposed by action interaction can achieve the segmentation of action impact and use the self-attention mechanism to deal with the mismatch between observations and actions, so as to retain reliable information; it strengthens the ability of drones to autonomously identify and adapt to changing adversarial environments; it enables drones to interact effectively to coordinate actions to achieve common goals;
[0047] (3) By maximizing the joint action value of multiple agents, optimal cooperation and strategy between drones are ensured, leading to a more effective and unified adversarial approach;
[0048] (4) Through a unique credit allocation mechanism, drones can more accurately evaluate their respective contributions during the collaboration process;
[0049] (5) Observe that introducing action interactions and credit assignments can differentiate the performance of algorithms in adversarial simulations, leading to continuous improvement and adaptability of strategies;
[0050] Each step of the method of the present invention contributes uniquely to the overall effectiveness of drone swarms in confrontation; by integrating advanced algorithms and learning techniques, the proposed method provides a complex and adaptive approach to drone swarm confrontation, enhancing their strategic capabilities, adaptability and effectiveness in various confrontation scenarios; it is able to conduct drone swarm confrontation games under partial observability, improves the stability and confrontation performance of the agent learning strategy, and provides new methods and techniques for drone swarm confrontation. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A flowchart of a drone cluster collaborative confrontation method introducing action interaction and credit allocation provided by an embodiment of the present invention;
[0052] Figure 2 A structural diagram of a neural network model used in an agent network training model provided in an embodiment of the present invention;
[0053] Figure 3 A network structure diagram that can distinguish credit allocation provided by an embodiment of the present invention;
[0054] Figure 4 A network structure diagram that introduces distinguishable action interactions and credit allocations provided in an embodiment of the present invention;
[0055] Figure 5 A winning rate graph of a drone swarm confrontation game provided by an embodiment of the present invention;
[0056] Figure 6 A comparison chart of friendly force survival rates between a drone swarm collaborative confrontation method that introduces action interaction and credit allocation and other algorithms provided by an embodiment of the present invention;
[0057] Figure 7 A comparison chart of enemy mortality rates between a drone swarm collaborative confrontation method that introduces action interaction and credit allocation and other algorithms provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0059] In this embodiment, a drone cluster collaborative confrontation method is introduced that introduces action interaction and credit allocation, such as Figure 1 As shown, the following steps are included:
[0060] Step 1: Model the confrontation space, action space, state space and observation space of the drone cluster respectively, and set the two opposing parties of the drone cluster, namely the red party and the blue party;
[0061] Construct effective adversarial space, action space, state space and observation space for drone swarms, and mathematically describe the decision elements, attribute characteristics and interaction relationships of the adversarial space in order to design and evaluate different intelligent agent strategies;
[0062] The drone swarm is composed of multiple drones; they have the ability to coordinate and fight against each other; at preset time intervals, the drones can accurately launch missiles at enemy targets and carry out devastating strikes on enemy targets; once any drone is unfortunately hit by an enemy missile, it will be judged as destroyed and can no longer continue to participate in the battle; the core missions of both parties are clear, which is to completely destroy all drones of the other party and compete for air supremacy;
[0063] In the confrontation scenario, the drones of both sides (the red side and the blue side) adopt a six-degree-of-freedom dynamics model, using dynamic equations to describe the motion of the drone in the continuous action space, taking a continuous action set, and using vectors to represent the continuous action parameters of the drone. The dynamics model of each drone is expressed as:
[0064]
[0065] Among them, v is the flight speed of the drone, m is the mass of the drone, t is the time, and v x 、v y are the velocity components of the X-axis and Y-axis respectively, θ is the pitch angle, is the heading angle, φ is the roll angle, F is the resultant force of the drone, x t ,y t They are respectively the positions of the drone on the X and Y axes at time t, x t-1 ,y t-1 They are respectively the positions of the drone on the X and Y axes at time t-1, and r φ is the rolling angular velocity, r θ is the pitch angular velocity, is the heading angular velocity, all of these data types are floating point types;
[0066] The continuous action space is represented as a multi-dimensional Euclidean space, where each dimension corresponds to a range of action parameters. When an action needs to be selected, the continuous parameters in the action space are searched to find the optimal combination of action parameters to meet a specific goal.
[0067] The action space is a continuous action space, and the dynamic equation is used to describe the movement of the drone cluster in the continuous action space. A continuous action set is adopted, and a vector is used to represent the continuous action parameters of the drone cluster. Then the continuous action space is discretized to expand the dimension to 8 dimensions; the expression and control of the moving direction are more delicate and comprehensive;
[0068] In this embodiment, the continuous action space is fully covered in a complete 360-degree cycle; the number of dimensions of the action space is directly related to the degree of discretization of this continuous angle space; if the 360-degree omnidirectional space is subdivided into n specific, non-overlapping direction intervals, then the dimension of the action space is equal to n; in order to enhance the accuracy and flexibility of direction selection, this embodiment adopts a more refined discretization strategy, evenly dividing the original continuous 360-degree circle into eight main directions: east, northeast, north, northwest, west, southwest, south and southeast; this measure significantly expands the dimension of the action space to 8 dimensions, thereby achieving a more delicate and comprehensive expression and control of the moving direction;
[0069] In the centralized training phase, the state space and observation space fully capture and utilize the global state information of the drone cluster to optimize the collaborative performance of the entire drone cluster; in the execution phase, each drone needs to rely on its partial observable information to make independent decisions and perform corresponding actions;
[0070] In this embodiment, the global state information of the drone cluster is shown in Table 1, which covers the key elements such as the overall operation status of the drone cluster, environmental parameters, and the relative position and state of each member; each drone can only obtain partial observation information within its field of view as shown in Table 2, which constitutes the basis for the drone to make local decisions and actions;
[0071] Table 1 Global status information
[0072]
[0073] Table 2 Partial observation information
[0074]
[0075] In each time step, the drone will receive some observation data within its circular field of view; the radius of this circular field of view is equivalent to the maximum range of its line of sight, which limits the environmental boundaries that the drone can perceive; from the perspective of the drone, its field of view not only covers the terrain features, but also includes the last actions of all surviving friendly drones within its field of view; set the limitations of drones in obtaining information: they cannot directly judge the status of friendly drones outside their field of view, whether they are at a farther distance or have been destroyed; drones also have the ability to observe the surrounding terrain features, which helps them understand the environment more comprehensively and make action choices that are more adapted to the environment;
[0076] Describe the drone swarm (i.e. multi-agent) adversarial task as a tuple The decentralized part consists of an observable Markov decision process, where is the set of intelligent agents, is the state set, is the action set, P is the state transition function, r is the reward function, Ω is the observation result set, O is the observation function, γ is the discount factor, γ∈[0,1]; set To describe the global state of the environment, at each time step t, each agent Select an action And generate a joint action U is the set of joint actions; according to the state transfer function P(s′|s,u): The environment enters the next state, causing the joint action u to change, where s′ is the state of the agent at the next moment; all agents share the same reward function r(s,u): is a set of real numbers. The function of r(s,u) is to evaluate the state s and the joint action u, and return a real value as a reward to measure the contribution of the state and action. The common goal of all factors in the tuple is to maximize the expected reward. in Denotes the discounted sum of future rewards G t To find the expectation, the specific formula is:
[0077]
[0078] Among them, k is the step size after time t, r t+k is the instant reward obtained at time t+k;
[0079] Step 2: Consider the action interaction between drones, regard drones as independent and autonomous intelligent agents, build an intelligent agent network training model, determine whether the action of the intelligent agent only affects the environmental information of the adversarial space or whether the action of the intelligent agent also affects other intelligent agents, and dynamically adjust the decision of the intelligent agent according to the environmental information of the adversarial space; including the following steps:
[0080] Step 2.1: Build an agent network training model for each agent to learn the agent's strategy and decision-making;
[0081] Construct an agent network training model to enable the agent to learn adaptive strategies in confrontation and achieve the optimal strategy in an environment of mutual influence; the neural network model structure used in the agent network training model is as follows Figure 2 As shown;
[0082] Divide the observations o and actions u input into the agent network training model, and divide the initial observations into three parts: the part related to the environment, the part related to itself, and the part related to other agents;
[0083] At time t, the initial observation value of the i-th agent is in is the observed value related to the environment, is the inherent attribute value of the intelligent agent, such as attack, is the observation related to other agents, n is the number of agents;
[0084] The agent’s observation value o is converted from a sequence (concatenate) form to a matrix (matrix) form to make the representation of the observation-state space more unified and comparable; the action u is divided into two parts: and action Only affects the environment information and the agent itself, such as movement; action Influencing other agents, such as collaborative communication and attack;
[0085] Step 2.2: Divide the agent network into different sub-modules, evaluate the local Q value of the agent action by the impact of the agent action on other agents, and calculate the global Q value of the agent network from the local Q values of all agents;
[0086] The agent network is divided into different sub-modules. By the influence of the agent action on other agents, it is judged whether the agent action only affects the environmental information of the adversarial space or the agent itself also affects other agents. If the agent action only affects the environmental information of the space and the agent itself, the initial observation value of the agent is input into the embedding layer in the agent network training model to evaluate the local Q value of the agent action. If the agent action affects other agents, such as cooperative communication and attack, the observation value o of agent i on agent j is input into the embedding layer in the agent network training model to evaluate the local Q value of the agent action. i,j The action of agent i affecting agent j is input into the embedding layer of the agent network training model to evaluate the local Q value of the agent action; the global Q value of the agent network is calculated from the local Q values of all agents;
[0087] In this embodiment, the agent network is divided into different sub-modules, and the impact of the agent action on other agents is considered to determine whether the action is still part of ; if the action belongs to Then the initial observation value of the input is generated through the embedding layer E The local Q value of the action is then evaluated; if the action belongs to Then the observation value o of agent i on agent j is i,j Input into the embedding layer E, and input the action of agent i affecting agent j, and get the action belongs to The local Q value of agent i is The formula is:
[0088]
[0089] Among them, Q i is the local Q function of the ith agent, is the hidden state of agent i at the previous time step t-1, is the hidden state space, is the observation input, is the action at time step t;
[0090] Observation history It is the observation history space;
[0091] The global Q value Q of the agent network is calculated from the local Q values of all agents. tot ;
[0092]
[0093] Among them, τ t is the observation history at time t, u tis the joint action at time t, F(·) is the agent credit allocation function, and F is a monotonic function, which can be expressed as: Q i is the Q function of the ith agent, For the agent set An agent in
[0094] When matching the agent's observations with the agent's actions, there will always be a situation where the agent's observations do not match any agent's actions, that is, the agent's observations have no matching agent actions; if these agent's observations are discarded, it will cause information loss to the agent;
[0095] Transformer is introduced into the agent network training model, and the self-attention mechanism is used to identify and filter out useful information, and the observations are weighted according to the useful information obtained; the attention is focused on the input features that are helpful to the agent, the attention to irrelevant features is suppressed, and useless information not needed for decision-making is discarded; the self-attention mechanism enables the agent to consider all observations even when there is no direct match, that is, the observations have no direct match with the actions, and the agent can still obtain useful information from these observations and weight the observations according to their importance;
[0096] The useful information identified and filtered by the self-attention mechanism is observation data or features that are of substantial help and value to the agent in making decisions or performing tasks, and data perceived by the agent in the environment of the adversarial space that can directly or indirectly affect the agent's behavior, decision-making process or task completion;
[0097] The network structure of the transformer consists of a self-attention mechanism; first, each element in the input sequence is mapped to a query Q, a key K, and a value V through a linear transformation, and then the attention weights are weighted and summed, and the output is:
[0098]
[0099] Among them, d k is the scaling factor related to K;
[0100] Assume that the Q, K, and V matrices of each layer of the transformer are the same, that is, Where l is the number of transformer layers; the transformer is defined as follows:
[0101]
[0102] Among them, MLP is a linear function for calculating Q, K, and V. is the output calculated by the attention mechanism, using V i l To generate the input features of the next layer;
[0103] Use a layer of MLP to project the last layer of the transformer to the Q function Q of the i-th agent i The output space is:
[0104]
[0105] Among them, Q i is the Q function of the ith agent, is the hidden state of agent i at the previous time step t-1, is the hidden state space, is the observation input, is the action at time step t;
[0106] Step 2.3: Use reinforcement learning algorithm to train the agent network training model;
[0107] Step 2.4: The agent perceives and understands the actions of other agents and dynamically adjusts the agent's decision based on the environmental information of the adversarial space;
[0108] Step 3: The two adversarial parties (red party and blue party) in the drone cluster confront each other, and at the same time evaluate and reason about the adversarial situation, receive the local Q value generated by the agent network training model as input, and establish a two-layer hybrid network structure using the monotonic value function decomposition method to obtain the global Q value Q of the agent network tot , and then obtain the optimal joint action of the agent;
[0109] The two adversaries (the red team and the blue team) confront each other through the set adversarial space and the agent network training model; both adversaries are committed to comprehensively collecting adversarial data, observing and recording the actions and environment of the adversary, estimating the state, and then evaluating the adversarial situation; based on the evaluation results of the adversarial situation, dynamically adjust the reasoning and decision-making of the agent;
[0110] The centralized training distributed execution paradigm is adopted, and the local Q values of each agent are integrated and constrained through a two-layer hybrid network to generate a global Q value; considering the additional state information available under the centralized training distributed execution paradigm, the consistency and rationality of the value decomposition are ensured through monotonicity constraints; each agent has an independent Q network, which can make independent decisions based on its own local information, and the output of the Q network is transmitted through a two-layer hybrid network to realize the local Q value to the global Q value Q tot The mapping of
[0111] The two-layer hybrid network is responsible for credit allocation, taking the local Q value generated by the agent network as input, and processing it through two separate super networks and activation functions contained in the two-layer hybrid network to output a global Q value Q tot , the specific formula is:
[0112] The weights of the two-layer hybrid network are generated by independent super networks using the true state s as input. Each super network consists of a single linear layer followed by a ReLU function to ensure that the weights of the two-layer hybrid network are non-negative.
[0113] Q tot (τ,u,s;θ)=W2R eLU (W1Q+b1)+b2
[0114] in, is the Q value output by the agent network, Q n (τ n ,u n ) is the local Q value of the nth agent, τ n ,u n are the observation history and action of the nth agent, τ is the observation history, s is the environment state, θ is the parameter of the target network, R eLU is the activation function of W1Q+b1, and The weight matrices and bias terms generated for two separate hypernetworks satisfy: is a real matrix, m and n are real matrices respectively The row and column dimensions of the vector;
[0115] Since the monotone valued function decomposition algorithm QMIX is monotonic, all elements in W1 and W2 are non-negative, that is, they satisfy
[0116] By ensuring that the agents maximize local Q-values based only on their own local observation-action history, the global Q-value is maximized, thus ensuring that the joint action is the optimal action for the drone swarm;
[0117] In this embodiment, the loss function of the two-layer hybrid network is updated by minimizing the squared TD error. The loss function formula is as follows:
[0118]
[0119] in, is the TD target, τ′,u′,s′ are the observation-action history, joint action and state at the next moment, θ -are the target network parameters, which are periodically copied from θ;
[0120] The credit allocation is how to fairly and effectively allocate the global credit or reward of the agent to each agent or local agent network in the UAV cluster collaborative confrontation or multi-agent network training model; "credit" is regarded as a quantitative evaluation of the contribution or performance of the agent, which reflects the importance and effectiveness of the agent in completing tasks, collaborative operations or decision-making.
[0121] The two separate hypernetworks dynamically adjust their own parameters or structures according to the input observation values to adapt to different environments;
[0122] Step 4: The two adversaries (red and blue) conduct collaborative task planning to provide the optimal strategy for multiple agents; update the local Q value of the agent and optimize the agent strategy by credit allocation. The specific method is as follows:
[0123] According to step 3, the global Q value Q tot , using the reverse transfer of the agent network training model to update the local Q value of the agent; different agents have different degrees of influence on the adversarial game, and understanding the contribution of each agent to the global goal is the key to understanding credit allocation, that is, credit allocation is distinguishable; using gradient entropy to update the global Q value Q of the agent tot , for distinguishable credit allocation and strategy optimization;
[0124] In this embodiment, in the value function decomposition paradigm, according to the Taylor expansion, the global Q value Q tot Can be decomposed into:
[0125]
[0126] in, represents the initial global Q value, o(Q T Q) represents a high-order small quantity term, which is usually used to represent higher-order terms in Taylor expansion, [g1,…,g n ] T Q represents the gradient information of each agent network output value;
[0127] Local Q value vs global Q value Q tot The impact depends on the corresponding gradient. The larger the gradient, the greater the impact on the global Q value Q tot The greater the impact, the greater the contribution to the entire agent network; due to the different credit allocations of agents, the gradient distribution between agents is also uneven. In view of this situation, the gradient is normalized, and the normalized gradient entropy is calculated. The gradient entropy is used to measure the distinguishability of the agent's contribution; the higher the gradient entropy, the more uniform the gradient distribution, and the easier it is to distinguish the contributions between agents;
[0128] Gradient normalization g i ′ and the normalized gradient entropy is N orm :
[0129]
[0130] Among them, N orm As an evaluation metric to measure the discriminability of credit assignment, logn is the maximum entropy of n values; g i represents the global Q value Q of the i-th agent tot If the credit assignment is highly discriminable, the gradient distribution will not be uniform; is the sum of all overall gradients;
[0131] The gradient entropy is normalized using n random maximum entropies; entropy is inversely proportional to the degree of gradient discrimination, that is, the lower the gradient entropy, the more uneven the gradient distribution, which makes it more difficult to distinguish the contributions between agents; the concept of normalized gradient entropy is introduced in the two-layer hybrid network to measure the distributability of credit; according to the global Q value Q tot The expanded result of , calculates the gradient:
[0132]
[0133] So we get:
[0134] dQ tot =W2FW1dQ
[0135]
[0136] in, is a diagonal matrix, and x c is the cth element output in the double-layer hybrid network, X is the input feature of the double-layer hybrid network, x1, x2, ..., x m Both are the outputs of the first layer linear transformation in the two-layer hybrid network, calculated by W1Q+b1;
[0137] Introducing loss functions for action interaction and credit assignment It consists of two parts:
[0138] (1) Original TD loss That is, the formula By maximizing the global Q value Q tot , so that each agent can learn the best strategy;
[0139] (2) Regularized gradient entropy
[0140]
[0141] Among them, λ is a hyperparameter used to balance the TD loss and regularization term, and batchsize is the batch size of sampling;
[0142] When λ=0, no gradient entropy calculation is performed; the larger λ is, the stronger the penalty for indistinguishability is; the regularization term is optimized to update only the parameters in the hybrid network, so that the two-layer hybrid network can more accurately measure the contribution between agents; the obtained loss function is reversed to complete the update of the local Q value and strategy of the action value function, so that the agent can choose a better strategy and win the confrontation; by introducing gradient entropy for credit allocation, each agent can better understand the contribution of its actions to the global reward, which encourages the agents to cooperate better; the network structure setting for distinguishable credit allocation is as follows Figure 3 As shown, the network structure that introduces action interaction and credit assignment is distinguishable as follows Figure 4 As shown;
[0143] Step 5: Repeat step 4 until the confrontation between the two adversaries ends, observe the adversarial performance of introducing action interaction and credit allocation, and complete the adversarial simulation of the drone cluster;
[0144] In this embodiment, the winning rate curve of the confrontation game between the drone swarm and the two parties is as follows: Figure 5 As shown, through multiple experiments and observations, it is found that the winning rate curve has limitations in some cases; especially in some time steps, the winning rate may tend to be consistent, making it difficult to intuitively show the subtle differences and actual effects of the algorithm; in order to more accurately evaluate and compare the performance of different algorithms, this embodiment proposes to use friendly survival rate and enemy death rate as new evaluation indicators; where friendly survival rate = number of friendly survival / original number of friendly; enemy death rate = number of enemy deaths / original number of enemies;
[0145] In this embodiment, the two sides of the drone swarm confrontation use the method of the present invention and the multi-agent proximal strategy optimization algorithm and the monotonic value function decomposition algorithm. The friendly survival rate and the enemy death rate of other algorithms are as follows: Figure 6 , 7 As shown in the two figures, the data in the left column are respectively the friendly survival rate and the enemy death rate using the method of the present invention.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A drone swarm collaborative confrontation method that introduces action interaction and credit allocation, characterized by: The following steps are involved: Step 1: Model the confrontation space, action space, state space and observation space of the drone cluster respectively, and set the confrontation parties of the drone cluster; Step 2: Consider the action interaction between drones, regard drones as independent and autonomous intelligent agents, build an intelligent agent network training model, determine whether the action of the intelligent agent only affects the environmental information of the adversarial space or whether the action of the intelligent agent also affects other intelligent agents, and dynamically adjust the decision of the intelligent agent according to the environmental information of the adversarial space; Step 3: The two adversarial parties in the drone swarm confront each other, and at the same time, evaluate and reason about the confrontation situation, receive the local Q value generated by the agent network training model as input, and establish a two-layer hybrid network structure using the monotonic value function decomposition method to obtain the global Q value of the agent network, and then obtain the optimal joint action of the agent; Step 4: The two adversaries conduct collaborative task planning to provide the optimal strategy for multiple agents; update the local Q value of the agent and optimize the agent strategy through the credit allocation method; Step 5: Repeat step 4 until the confrontation between the two parties ends, observe the confrontation performance of introducing action interaction and credit allocation, and complete the confrontation simulation of the drone cluster.
2. The method for cooperative confrontation of drone swarms by introducing action interaction and credit allocation according to claim 1, characterized in that: The step 1 comprises: The drone cluster is composed of multiple drones; in the confrontation scenario, the drones of both sides adopt a six-degree-of-freedom dynamics model, use dynamic equations to describe the movement of the drones in the continuous action space, adopt a continuous action set, and use vectors to represent the continuous action parameters of the drones; Describe the drone swarm confrontation task as a tuple The decentralized part consists of an observable Markov decision process, where is the set of intelligent agents, is the state set, is the action set, P is the state transition function, r is the reward function, Ω is the observation result set, O is the observation function, γ is the discount factor, γ∈[0,1]; set To describe the global state of the environment, at each time step t, each agent Select an action And generate a joint action U is the joint action set; according to the state transfer function The environment enters the next state, causing the joint action u to change, where s′ is the state of the agent at the next moment; all agents share the same reward function is a set of real numbers. The function of r(s,u) is to evaluate the state s and the joint action u, and return a real value as a reward to measure the contribution of the state and action. The common goal of all factors in the tuple is to maximize the expected reward. in Denotes the discounted sum of future rewards G t To find the expectation, the specific formula is: Among them, k is the step size after time t, r t+k is the instant reward obtained at time t+k.
3. The method for cooperative confrontation of drone swarms by introducing action interaction and credit allocation according to claim 2, characterized in that: The step 2 comprises: Step 2.1: Build an agent network training model for each agent to learn the agent's strategy and decision-making; Step 2.2: Divide the agent network into different sub-modules, evaluate the local Q value of the agent action by the impact of the agent action on other agents, and calculate the global Q value of the agent network from the local Q values of all agents; Step 2.3: Use reinforcement learning algorithm to train the agent network training model; Step 2.4: The agent perceives and understands the actions of other agents and dynamically adjusts the agent's decision based on the environmental information of the adversarial space.
4. The method for cooperative confrontation of drone swarms with action interaction and credit allocation according to claim 3 is characterized by: The step 2.1 comprises: Construct an agent network training model to enable agents to learn adaptive strategies in confrontation and achieve optimal strategies in an environment of mutual influence; Divide the observations o and actions u input into the agent network training model, and divide the initial observations into three parts: the part related to the environment, the part related to itself, and the part related to other agents; The observation value o of the agent is converted from sequence form to matrix form to make the representation of the observation-state space unified and comparable; the action u is divided into two parts: and action Only affects the environment information and the agent itself; action Affect other agents.
5. The method for cooperative confrontation of drone swarms with action interaction and credit allocation according to claim 4, characterized in that: The step 2.2 comprises: The agent network is divided into different sub-modules. By the influence of the agent action on other agents, it is judged whether the agent action only affects the environmental information of the adversarial space or the agent itself also affects other agents. If the agent action only affects the environmental information of the space and the agent itself, the initial observation value of the agent is input into the embedding layer in the agent network training model to evaluate the local Q value of the agent action. If the agent action affects other agents, the observation value o of agent i on agent j is i,j The action of agent i affecting agent j is input into the embedding layer of the agent network training model to evaluate the local Q value of the agent action; the global Q value of the agent network is calculated from the local Q values of all agents; Introduce transformer into the agent network training model, use the self-attention mechanism to identify and filter out useful information, and weight the observations according to the useful information obtained; focus on the input features that are helpful to the agent, suppress attention to irrelevant features, and discard useless information that is not needed for decision making; The useful information identified and filtered using the self-attention mechanism is observation data or features that are of substantial help and value to the agent when making decisions or performing tasks, and data perceived by the agent in the environment of the adversarial space that can directly or indirectly affect the agent's behavior, decision-making process or task completion.
6. The method for cooperative confrontation of drone swarms with action interaction and credit allocation according to claim 5, characterized in that: The step 3 comprises: The two adversaries confront each other through the set adversarial space and intelligent agent network training model; both adversaries are committed to comprehensively collecting adversarial data, observing and recording the actions and environment of the adversary, estimating the state, and then evaluating the adversarial situation; based on the evaluation results of the adversarial situation, dynamically adjust the reasoning and decision-making of the intelligent agent; A centralized training and distributed execution paradigm is adopted, and the local Q-values of each intelligent agent are integrated and constrained through a two-layer hybrid network to generate a global Q-value; each intelligent agent has an independent Q-network and can make independent decisions based on its own local information. The output of the Q-network is transmitted through a two-layer hybrid network to realize the mapping of local Q-values to global Q-values.
7. The method for cooperative confrontation of drone swarms with action interaction and credit allocation according to claim 6, characterized in that: The two-layer hybrid network is responsible for credit allocation, taking the local Q value generated by the agent network as input, and processing it through two separate super networks and activation functions contained in the two-layer hybrid network to output a global Q value Q tot , the specific formula is: Q tot (τ,u,sθ)W2R eLU (W1Q+b1)+b2 in, is the local Q value output by the agent network, Q n (τ n ,u n ) is the local Q value of the nth agent, τ n ,u n are the observation history and action of the nth agent, τ is the observation history, s is the environment state, θ is the parameter of the target network, and R eLU is the activation function of W1Q+b1, and The weight matrices and bias terms generated for two separate hypernetworks satisfy: is a real number matrix, m and n are real number matrices The row and column dimensions of the vector; The loss function of the two-layer hybrid network is updated by minimizing the squared TD error. The loss function formula is as follows: in, is the TD target, τ′,u′,s′ are the observation-action history, joint action and state at the next moment, θ - are the target network parameters, which are periodically copied from θ; The two separate hypernetworks dynamically adjust their own parameters or structures according to the input observation values to adapt to different environments.
8. The method for cooperative confrontation of drone swarms with action interaction and credit allocation according to claim 7, characterized in that: The step 4 comprises: According to step 3, the global Q value Q tot , using the back propagation of the agent network training model to update the local Q value of the agent; using the gradient entropy to update the global Q value of the agent for distinguishable credit assignment and policy optimization; The concept of normalized gradient entropy is introduced in the two-layer hybrid network to measure the distributability of credit; according to the global Q value Q tot The expanded result of , calculates the gradient: So we get: dQ tot =W2FW1dQ in, is a diagonal matrix, and x c is the cth element output in the double-layer hybrid network, X is the input feature of the double-layer hybrid network, x1, x2, ..., x m Both are the outputs of the first layer linear transformation in the two-layer hybrid network, calculated by W1Q+b1; Introducing loss functions for action interaction and credit assignment Including the original TD loss and regularized gradient entropy Two parts, as shown in the following formula: Among them, λ is a hyperparameter used to balance the TD loss and the regularization term, and batchsize is the batch size of the sampling.
Citation Information
Patent Citations
Unmanned aerial vehicle cluster confrontation game simulation method based on anti-fact baseline
CN116136945A
Unmanned aerial vehicle group confrontation control method and system
CN117311392A
Unmanned aerial vehicle cluster collaborative confrontation method combined with self-attention mechanism
CN117608315A
Emergency guiding method for vehicle in mining area
CN118707948A
Unmanned aerial vehicle cluster near-end strategy optimization collaborative confrontation method based on anti-fact baseline
CN118778678A