Collaborative maneuvering decision-making method in air combat game confrontation of multiple unmanned aerial vehicles

By constructing a three-dimensional air combat simulation environment and a cognitive graph collaborative network, the problems of enemy strategy uncertainty and collaborative decision-making difficulties in multi-UAV air combat were solved, achieving efficient, robust and forward-looking decision-making under incomplete information conditions, and improving the collaborative combat capability of UAV swarms.

CN121997972APending Publication Date: 2026-05-08SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENYANG AEROSPACE UNIVERSITY
Filing Date
2025-12-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In multi-drone air combat games, existing methods are difficult to effectively cope with the uncertainty of the enemy's strategy and the difficulty of collaborative decision-making, especially the lack of robustness and foresight in the collaborative decision-making of drone swarms under conditions of incomplete information.

Method used

A three-dimensional continuous air combat simulation environment is constructed, a six-degree-of-freedom motion model is used to describe the dynamic characteristics of UAVs, a situation assessment model based on incomplete information is established, a multi-objective composite reward function is designed, a hypernetwork enemy strategy modeler and a cognitive graph collaborative network are introduced, information propagation and feature fusion are realized through a diffusion graph attention mechanism, and a centralized evaluator is combined to optimize decision-making.

Benefits of technology

It improves the foresight and adaptability of UAVs in decision-making under incomplete information conditions, enhances the effectiveness of group collaborative operations, and solves the shortcomings of traditional methods and existing deep reinforcement learning in policy mutation and collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997972A_ABST
    Figure CN121997972A_ABST
Patent Text Reader

Abstract

According to the collaborative maneuvering decision-making method in the air combat game confrontation of the multiple unmanned aerial vehicles disclosed by the invention, the dynamic response of the enemy is fully considered in the decision-making process by introducing the enemy behavior prediction model, so that the robustness and flexibility of decision-making are improved. Different from traditional multi-unmanned aerial vehicle game research, the method is characterized in that an enemy strategy is explicitly modeled, and online reasoning is carried out by adopting deep reinforcement learning. Specifically, the behavior of an enemy is predicted by establishing an enemy strategy modeling device, and perception and intention information is integrated in combination with a cognitive map collaborative network, so that the global visual field of decision making is improved. In addition, situation awareness information is processed by introducing a multi-head self-attention mechanism, prediction and evaluation of enemy response are optimized, and the adaptability of the decision making system is further enhanced. And in combination with centralized training and a distributed execution architecture, the problem of information asymmetry in collaborative decision making is solved, so that the unmanned aerial vehicle can make efficient decisions based on incomplete information in the execution stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a collaborative maneuver decision-making method in multi-UAV air combat game confrontation, belonging to the field of UAV collaborative combat and autonomous decision control technology. Background Technology

[0002] With the rapid development of UAV technology, its application in highly dynamic and intense air combat environments is becoming increasingly widespread. In multi-UAV cooperative air combat scenarios, the complexity of decision-making increases dramatically. Each UAV's decision depends not only on its own state and mission, but also on real-time prediction of enemy behavior, understanding of the overall situation, and efficient coordination with friendly forces. Traditional decision-making methods, such as expert systems based on static rules or domain knowledge, can execute pre-set tactics, but lack the flexibility to cope with sudden situations and unknown strategies. Methods based on optimization theory can find the best in high-dimensional spaces, but the computational cost is high, making it difficult to meet the needs of real-time decision-making. Game theory methods can model the strategic interactions between opposing sides, but in dynamic environments with incomplete information and rapidly evolving strategies, acquiring and processing enemy strategy information in real time faces significant challenges. Furthermore, the above methods generally struggle to effectively address the "curse of dimensionality" and information asymmetry problems in multi-UAV cooperation.

[0003] In recent years, deep reinforcement learning technology has provided a new approach to decision-making problems in complex and dynamic environments by enabling agents to learn optimal strategies through autonomous interaction with the environment. Existing research has applied it to UAV air combat, allowing UAVs to adaptively adjust their maneuver strategies. However, in complex games involving multiple aircraft, especially those with rapidly changing enemy strategies, existing deep reinforcement learning methods still have significant limitations: First, most methods implicitly learn opponent behavior, lacking explicit modeling and forward-looking prediction of enemy strategies, leading to slow response and insufficient decision robustness when facing policy changes; second, in a distributed execution framework, how to achieve efficient group collaboration based on locally incomplete information remains a key challenge; finally, existing methods are still incomplete in effectively integrating situational awareness, intent reasoning, and collaborative decision-making to form a global battlefield understanding. Summary of the Invention

[0004] The purpose of this invention is to address the problems of uncertain enemy strategy and difficulty in collaborative decision-making in multi-UAV air combat games, and to provide a collaborative maneuver decision-making method in multi-UAV air combat games, specifically including the following steps: A three-dimensional continuous air combat simulation environment is constructed, and the mission objectives of the opposing UAV swarms are defined. A six-degree-of-freedom motion model is used to describe the dynamic characteristics of the UAVs, and their state variables include at least position, velocity, attitude angle, and angular velocity. A UAV air combat situation assessment model based on incomplete information is constructed to assess the relative situation relationship between the enemy and ourselves at each time step. The calculation of the relative situation relationship is related to factors such as relative angle and relative distance. A global joint state space is constructed, which integrates the kinematic state and battlefield environment information of all UAVs. During the execution phase, each UAV makes decisions based on its local observations, which include its own state, the relative state information of friendly UAVs within the communication range and enemy UAVs within the perception range. Construct a joint action space and define the set of executable actions for each UAV, including speed control, attitude adjustment and attack tactical actions to support collaborative decision modeling; Design a multi-objective composite reward function, which includes at least survival reward, defeat reward, cooperation reward, boundary violation penalty and round win / loss reward. Through weight configuration, guide the drone to achieve a balance optimization between individual survival, attack effectiveness and group cooperation. An enemy strategy modeler based on a hypernetwork is established to predict the probability distribution of enemy drone actions in the future based on historical environmental states and enemy action sequences, and output the prediction confidence. A cognitive graph collaborative network is constructed to model the perception and intent interaction relationships between our UAVs using a graph structure, and multi-step information propagation and feature fusion between nodes are realized through a diffusion graph attention mechanism to improve the overall situational awareness and collaborative decision-making capabilities of the group. A centralized evaluator is constructed to jointly optimize the temporal difference loss, enemy prediction loss, and Actor policy loss of the Critic network, enabling the UAV to output the final maneuver decision based on local observation and cooperative network during the distributed execution phase.

[0005] Furthermore, let the set of Blue Team intelligent agents be... The enemy forces assembled as At time t, the global joint state space is constructed as follows:

[0006]

[0007] in, = [xj, yj, zj]T represents the position, = [vxj, vyj, vzj]T represents the velocity components, φ, θ, ψ represent the roll, pitch, and yaw angles respectively, and ωj = [ωxj, ωyj, ωzj] represents the angular velocity. For global timing information, During the execution phase, agent i can only rely on local observables. decision making

[0008] in, For the set of friendly / enemy neighbors within the communication range, This refers to the relative status of friendly aircraft. Let the enemy aircraft be in a relative state, expressed in the aircraft's coordinate system, and let... ,

[0009] Then for any target, define

[0010]

[0011] Based on this, construct the relative state vector.

[0012] in, The rotation matrix from inertia to the body. The distance is relative. It is the azimuth angle. Angle of elevation The angle between the enemy aircraft and the opposite direction. For hit indication of attack and no-escape zones, geometric determination is used. To characterize timing information, an observation stack of length H is employed. .

[0013] Furthermore, a joint action space is constructed, and the actions of all UAVs are modeled uniformly. The set of executable actions for each UAV at time t is denoted as .

[0014] in, Let j represent the j-th optional action of UAV i in the action space, including speed adjustment, attitude adjustment, and attack action. The joint action space is composed of the Cartesian product of the actions of all UAVs.

[0015] At time t, the joint action of the entire system is represented as:

[0016] In aerial combat scenarios, speed control, attitude adjustment, and tactical maneuver adjustments are performed on each drone.

[0017] Furthermore, when building an adversary policy modeler based on hypernetworks: Generate enemy strategy parameters using hypernetworks.

[0018] in, For the enemy's previous action, a predicted distribution of the enemy's actions is obtained based on reasoning, and a confidence level assessment is performed simultaneously.

[0019]

[0020] in, ∈[0,1] indicates the prediction confidence level, and a value close to 1 indicates that the prediction is reliable.

[0021] Furthermore, a cognitive graph collaborative network is constructed to model the collaborative relationships among multiple drones, establishing the following perceptual intent encoder:

[0022] in, For local observation of UAV i, Information is propagated through a diffusion graph attention network to predict the enemy's strategy distribution.

[0023] in, Represents a set of neighbors. The attention-based diffusion weight is defined as:

[0024] The final output projection yields the policy action. .

[0025] Furthermore, the Critic network is used to evaluate the state-action value function, and its loss is defined as:

[0026] Where D is the experience replay pool. These are the state, action, reward, and next state. The value of actions predicted by the Critic network. For the target value, Discount factor; The goal of the enemy modeler is to minimize the difference between the predicted enemy action distribution and the actual observed actions, using the mean squared error form:

[0027] in, For the enemy's observation information, This represents the enemy's actual actions at time t. Actions predicted by the enemy modeler; The goal of an Actor network is to maximize expected reward and improve robustness by using adversary policy predictions. Its loss is defined as...

[0028] in, For the Actor network in state The action distribution of the output, For the Critic network to evaluate the value of actions The final joint loss function is defined as a weighted sum of three parts, where, , , These are the weighting coefficients.

[0029] This method aims to improve the foresight, adaptability, and overall collaborative combat effectiveness of UAVs under incomplete information conditions by explicitly modeling and predicting enemy strategies, combined with an innovative swarm collaborative perception mechanism.

[0030] The invention employs a centralized training and distributed execution architecture, the core of which lies in simultaneously constructing an adversary strategy modeler and a cognitive graph collaborative network. The adversary strategy modeler, based on a hypernetwork and historical interaction information, infers the probability distribution of adversary behavior strategies in real time, providing forward-looking input for friendly decision-making. The cognitive graph collaborative network models the interaction relationships between friendly UAVs using a graph structure, and achieves multi-step propagation and fusion of situational and intentional information among multiple agents through a diffusion graph attention mechanism, thereby overcoming the limitations of local observation and improving the intelligence level of group collaboration.

[0031] Furthermore, this invention introduces a multi-head self-attention mechanism to deeply encode the original observations, enhancing the ability to extract environmental temporal and feature dependencies. Through the collaborative work of the above modules, this invention enables UAVs to jointly optimize strategies, value networks, and enemy models using global information during the training phase; during the execution phase, each UAV relies only on its own local observations to acquire group situational awareness through the collaborative network and, combined with enemy behavior predictions, generate rapid, collaborative, and adaptive maneuver decisions. This invention effectively addresses the shortcomings of traditional methods and existing deep reinforcement learning methods in responding to policy changes and achieving efficient collaboration, providing a new solution for multi-UAV intelligent air combat systems. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a flowchart of the collaborative maneuver decision-making method in multi-UAV air combat game confrontation in this invention.

[0034] Figure 2 This is a diagram of the six-degree-of-freedom motion model of the UAV in this invention.

[0035] Figure 3 This is a diagram of the situation assessment and decision-making model in this invention.

[0036] Figure 4 This is a diagram of the algorithm based on cognitive graph collaboration and adversary strategy modeling in this invention. Figure 5 The diagram shows the framework technology details in the embodiments of the present invention. Figure 6 This is a 3v3 game adversarial graph of drones in an embodiment of the present invention. Figure 7 This is a defeat rate data graph in an embodiment of the present invention. Figure 8 Attention data diagram in an embodiment of the present invention Figure 9 Direction error data in embodiments of the present invention Figure 10 This is an attack and defense energy efficiency diagram in an embodiment of the present invention. Detailed Implementation To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0038] like Figure 1 As shown, this invention provides a cooperative maneuver decision-making method in multi-UAV air combat game, specifically including the following steps: S1: A six-degree-of-freedom (6-DOF) model is adopted as the physical motion model of the UAV, specifically considering the UAV's position, velocity, attitude, and rotation in the air. The model uses an inertial coordinate system and a body coordinate system to define the global position and local motion of the UAV, respectively. Among them, (xu, yu, zu) represent the local coordinate axes in the UAV's body coordinate system, defining the forward, lateral, and vertical directions of the UAV. The motion of the UAV is described by its linear velocities (vx, vy, vz) in the x, y, and z directions, as well as its pitch, roll, and yaw angles (θ, μ, φ). The motion model of the UAV is shown in formula (1): (1) Where v=(vx,vy,vz) is the velocity vector of the UAV. , and Let v be the rate of change of position in the three directions. In the three-dimensional continuous air combat simulation environment, a six-degree-of-freedom motion model is used to describe the dynamic characteristics of the UAV, where the dynamic model of the UAV is shown in formula (2): (2) in, It is the projection of v onto the xoy plane. It is the pitch rate of the drone. This is the yaw rate of the drone, and g is the acceleration due to gravity. It is an overload in the velocity direction. It is an overload in the velocity normal direction.

[0039] S2: Construct a UAV air combat situation assessment model based on incomplete information-based air combat game theory. The calculation of the relative situation is mainly related to factors such as relative angle and relative distance. The relative situation relationship between the red and blue sides in the battle is as follows: Figure 3 As shown. Attack angle Calculated using formula (3).

[0040] (3) in, The relative distance between the red and blue teams. Let be the velocity vector of the red side. Whether the attack conditions can be met is determined by formula (4).

[0041] (4) in, and These are the shortest and longest attack ranges. This is the maximum attack angle, set to 45°. Escape angle. It can be calculated using formula (5).

[0042] (5) The no-escape zone is defined in air combat as the fatal zone in which a target UAV cannot escape the attack within a specified time, even with maximum maneuvering, once it enters a specific geometric area of ​​the attacker. Within this zone, the target UAV will inevitably be destroyed. It delineates the range of the attacker's advantage in air combat and can be expressed as formula (6). (6) in, and It is the minimum and least escape distance. It is the minimum escape angle in the no-escape zone, set to 45°.

[0043] S3: Construct a global joint state space to integrate the kinematic states and battlefield environment information of all UAVs. The joint state space includes the kinematic and dynamic states of each UAV, as well as environmental and enemy UAV information, thus providing complete state input for collaborative decision-making. To support CTDE, this section presents a unified modeling of the global state and individual observables. Let the set of blue team agents be denoted as... The enemy forces assembled as At time t, the global state is defined as (7) (8) in, = [xj, yj, zj]T represents the position, = [vxj, vyj, vzj]T is the velocity component, φ, θ, ψ are the roll, pitch, and yaw angles respectively, and ωj = [ωxj, ωyj, ωzj] is the angular velocity. Global timing information During the execution phase, agent i can only rely on local observables. decision making (9) in, For the set of friendly / enemy neighbors within the communication range, This refers to the relative status of friendly aircraft. This represents the relative state of the enemy aircraft. The relative state is expressed in the aircraft's coordinate system, let... , (10) Then for any target, define

[0044] (11) Based on this, construct the relative state vector. (12) in, The rotation matrix from inertia to the body. The distance is relative. It is the azimuth angle. Angle of elevation The angle between the enemy aircraft and the opposite direction. Hit indications for attack zones and no-escape zones are determined geometrically. A stack of observations of length H is used to characterize timing information. (13) S4: Construct a joint action space to model the actions of all UAVs in a unified manner to better describe the multi-agent collaborative decision-making process.

[0045] The set of executable actions for each drone at time t is denoted as . (14) in, Let represent the j-th optional action of UAV i in the action space, including speed adjustment, attitude adjustment, and attack actions. The joint action space is composed of the Cartesian product of the actions of all UAVs. (15) Therefore, at time t, the joint action of the entire system can be expressed as: (16) In the air combat scenario described in this article, the specific actions of each drone can be further subdivided into the following categories: 1. Speed ​​control actions: 1. **Deceleration, Hold, and Acceleration:** These correspond to deceleration, hold, and acceleration, respectively. 2. **Attitude Adjustment Maneuvers:** Maneuvering is achieved by adjusting pitch, roll, and yaw angles to control the UAV's spatial attitude. 3. **Tactical Maneuvers:** These include discrete maneuvers such as attacks.

[0046] S5: Design a multi-objective composite reward function, including survival reward, defeat reward, cooperative reward, boundary violation penalty, and round victory / defeat reward. This reward function guides drones to pursue individual survival and enemy elimination while also considering overall cooperative combat objectives. The designed reward function mainly consists of the following parts: Survival Rewards If the drone is still alive at the current time step, it will receive a persistent reward to encourage long-term survival. (17) a) Defeat Rewards Destroying an enemy drone grants a significant positive reward. If the drone is destroyed by the enemy, a substantial negative reward is given. (18) b) Collaborative Rewards To avoid individualistic behavior, a cooperative component is added to the reward function, rewarding multiple drones for coordinating attacks on the same enemy aircraft or for cooperative evasion. (19) c) Penalty for leaving the battlefield At each step, the drone's position information is checked for out-of-bounds errors. Once it exceeds the predetermined operational boundary, a corresponding penalty is triggered, thereby constraining the agent to limit its decisions to a reasonable operational airspace.

[0047] (20) d) Global Rewards Global rewards are determined at the end of each round. Based on the number of drones remaining on both sides at the end of the match, all friendly drones receive the same global reward or penalty.

[0048] (twenty one) f) Calculation of comprehensive reward The comprehensive reward consists of survival reward, defeat reward and cooperation reward and withdrawal penalty, as expressed by formula (22). (twenty two) The reward coefficients involved in the above reward function are respectively assigned values. =1、 =100、 =50、 =25、 =100 S6: Construct an enemy policy modeler based on a hypernetwork. The goal of the enemy policy modeler is to infer the distribution of future enemy actions based on observation sequences and historical actions. First, enemy policy parameters are generated using a hypernetwork. (twenty three) in, This represents the enemy's action in the previous step. Then, behavioral reasoning is performed to obtain the predicted distribution of the enemy's actions, along with a confidence assessment. (twenty four) (25) in, ∈[0,1] represents the prediction confidence level; the closer to 1, the more reliable the prediction.

[0049] S7: Construct a cognitive graph collaboration network to model the collaborative relationships between multiple drones. First, introduce a perception-intent encoder: (26) in, For local observation of UAV i, This represents the predicted distribution of enemy strategies. Information is then propagated through a diffusion graph attention network. (27) in, Represents a set of neighbors. The attention-based diffusion weight is defined as: (28) The final output projection yields the policy action. (29) S8: Construct a centralized evaluator to jointly optimize the temporal difference (TD) loss of the Critic, the prediction loss of the adversary policy modeling, and the policy loss of the Actor. This can enhance the ability to model adversary behavior while maintaining the stability of policy optimization.

[0050] Critic networks are used to evaluate the state-action value function, and their loss is defined as: (30) Where D is the experience replay pool. These are the state, action, reward, and next state. The value of actions predicted by the Critic network. For the target value, This is the discount factor.

[0051] The goal of the enemy modeler is to minimize the difference between the predicted enemy action distribution and the actual observed actions, using the mean squared error form: (31) in, For the enemy's observation information, This represents the enemy's actual actions at time t. Actions predicted by the enemy modeler.

[0052] The goal of an Actor network is to maximize expected reward while improving robustness by utilizing adversary policy predictions. Its loss is defined as... (32) in, For the Actor network in state The action distribution of the output, For the Critic network to evaluate the value of actions The final joint loss function is defined as a weighted sum of three parts, where, , , These are weighting coefficients, used in the first 20% of training rounds. =1.0、 =lin_ramp(0→0.5), =1.0, then keep =0.5. This "weighting" can stabilize Critic and Actor in the early stages, and then gradually emphasize the opponent's predictions.

[0053] (33) The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0054] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0055] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0056] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0057] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cooperative maneuver decision-making method in multi-UAV air combat game, characterized in that: A three-dimensional continuous air combat simulation environment is constructed, and the mission objectives of the opposing UAV swarms are defined. A six-degree-of-freedom motion model is used to describe the dynamic characteristics of the UAVs, and their state variables include at least position, velocity, attitude angle, and angular velocity. A UAV air combat situation assessment model based on incomplete information is constructed to assess the relative situation relationship between the enemy and ourselves at each time step. The calculation of the relative situation relationship is related to factors such as relative angle and relative distance. A global joint state space is constructed, which integrates the kinematic state and battlefield environment information of all UAVs. During the execution phase, each UAV makes decisions based on its local observations, which include its own state, the relative state information of friendly UAVs within the communication range and enemy UAVs within the perception range. Construct a joint action space and define the set of executable actions for each UAV, including speed control, attitude adjustment and attack tactical actions to support collaborative decision modeling; Design a multi-objective composite reward function, which includes at least survival reward, defeat reward, cooperation reward, boundary violation penalty and round win / loss reward. Through weight configuration, guide the drone to achieve a balance optimization between individual survival, attack effectiveness and group cooperation. An enemy strategy modeler based on a hypernetwork is established to predict the probability distribution of enemy drone actions in the future based on historical environmental states and enemy action sequences, and output the prediction confidence. A cognitive graph collaborative network is constructed to model the perception and intent interaction relationships between our UAVs using a graph structure, and multi-step information propagation and feature fusion between nodes are realized through a diffusion graph attention mechanism to improve the overall situational awareness and collaborative decision-making capabilities of the group. A centralized evaluator is constructed to jointly optimize the temporal difference loss of the Critic network, the enemy prediction loss, and the policy loss of the Actor, enabling the UAV to output the final maneuver decision based on local observation and cooperative network during the distributed execution phase.

2. The cooperative maneuver decision-making method in multi-UAV air combat game confrontation according to claim 1, characterized in that: Let the set of blue-side intelligent agents be... The enemy forces assembled as At time t, the global joint state space is constructed as follows: in, = [xj, yj, zj]T represents the position, = [vxj, vyj, vzj]T represents the velocity components, φ, θ, ψ represent the roll, pitch, and yaw angles respectively, and ωj = [ωxj, ωyj, ωzj] represents the angular velocity. For global timing information, During the execution phase, agent i can only rely on local observables. decision making in, For the set of friendly / enemy neighbors within the communication range, This refers to the relative status of friendly aircraft. Let the enemy aircraft be in a relative state, expressed in the aircraft's coordinate system, and let... , Then for any target, define Based on this, construct the relative state vector. in, The rotation matrix from inertia to the body. The distance is relative. It is the azimuth angle. Angle of elevation The angle between the enemy aircraft and the opposite direction. For hit indication of attack and no-escape zones, geometric determination is used. To characterize timing information, an observation stack of length H is employed. 。 3. The cooperative maneuver decision-making method in multi-UAV air combat game as described in claim 1, characterized in that: Construct a joint action space, model the actions of all UAVs in a unified manner, and denote the set of executable actions of each UAV at time t as . in, Let j represent the j-th optional action of UAV i in the action space, including speed adjustment, attitude adjustment, and attack action. The joint action space is composed of the Cartesian product of the actions of all UAVs. At time t, the joint action of the entire system is represented as: In aerial combat scenarios, speed control, attitude adjustment, and tactical maneuver adjustments are performed on each drone.

4. The cooperative maneuver decision-making method in multi-UAV air combat game confrontation according to claim 1, characterized in that: When building an adversary policy modeler based on hypernetworks: Generate enemy strategy parameters using hypernetworks. in, For the enemy's previous action, a predicted distribution of the enemy's actions is obtained based on reasoning, and a confidence level assessment is performed simultaneously. in, ∈[0,1] indicates the prediction confidence level, and a value close to 1 indicates that the prediction is reliable.

5. The cooperative maneuver decision-making method in multi-UAV air combat game confrontation according to claim 1, characterized in that: A cognitive graph collaborative network is constructed to model the collaborative relationships among multiple drones, and the following perceptual intent encoder is established: in, For local observation of UAV i, Information is propagated through a diffusion graph attention network to predict the enemy's strategy distribution. in, Represents a set of neighbors. The attention-based diffusion weight is defined as: The final output projection yields the policy action. 。 6. The cooperative maneuver decision-making method in multi-UAV air combat game confrontation according to claim 1, characterized in that: Critic networks are used to evaluate the state-action value function, and their loss is defined as: Where D is the experience replay pool. These are the state, action, reward, and next state. The value of actions predicted by the Critic network. For the target value, Discount factor; The goal of the enemy modeler is to minimize the difference between the predicted enemy action distribution and the actual observed actions, using the mean squared error form: in, For the enemy's observation information, This represents the enemy's actual actions at time t. Actions predicted by the enemy modeler; The goal of an Actor network is to maximize expected reward and improve robustness by using adversary policy predictions. Its loss is defined as... in, For the Actor network in state The action distribution of the output, For the Critic network to evaluate the value of actions The final joint loss function is defined as a weighted sum of three parts, where, , , These are the weighting coefficients. 。