A decision tree-based multi-round attack-defense game decision-making method for drones
By adopting a multi-round offensive and defensive game decision-making method based on decision trees, the optimal maneuver problem in multi-round confrontation in UAV air combat is solved, achieving efficient decision-making in complex air combat environments and improving the probability of victory and the accuracy of air combat strategies for UAVs.
Patent Information
- Application Number
- CN202310132309.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-02-20
AI Technical Summary
Existing UAV air combat decision-making methods are mainly applicable to single-round confrontations and are difficult to effectively solve the optimal maneuver problem in multi-round confrontations. In particular, they have high computational load and low accuracy in complex air combat environments, which cannot meet the actual air combat needs.
A multi-round offensive and defensive game decision-making method based on decision trees is adopted. By establishing an advantage function model and a payoff function model, and combining the decision tree algorithm, the Nash equilibrium solution and expected payoff value are solved, and the UAV maneuver strategy is dynamically adjusted until the final victory condition is reached.
This paper presents a multi-round air combat decision-making method with low computational cost and high classification accuracy. It can achieve optimal maneuvering of UAVs in complex air combat environments, improve the probability of victory, and is suitable for research on multi-aircraft collaborative decision-making.
Smart Images

Figure CN116306941B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV attack and defense decision-making, specifically: a multi-round attack and defense game decision-making method for UAVs based on decision trees. Background Technology
[0002] With the development of artificial intelligence technology, drones, with their advantages of zero casualties and high maneuverability, have become one of the main combat units in air combat. Autonomous air combat decision-making is key to drones' victory in air combat. However, most current air combat decision-making methods are only applicable to single-round engagements. In reality, actual air combat requires multiple rounds of engagement between opposing sides, yet research results on this type of problem are scarce. Therefore, it is essential to conduct research on dynamic air combat decision-making for multi-round drone air combat engagements.
[0003] Typically, offensive and defensive decision-making methods include matrix strategy methods and expert system methods. However, for complex air combat environments, these methods suffer from high computational complexity and low accuracy. Game theory decision-making primarily studies decision-making processes involving struggle or competition. It focuses on the strategic changes between the two sides, incorporating these changes into the decision-making process. In modern UAV air combat, the offensive and defensive decision-making processes of both sides often cannot be completed in a single round. Both sides will take a series of actions to pursue their own victory, which aligns with the characteristics of game theory decision-making. Therefore, air combat game theory decision-making is increasingly attracting the attention of some scholars. For UAV air combat, game theory algorithms are considered to obtain reasonable decisions based on the changing situation of both sides. Li Shih-hao proposed a fuzzy game model for air combat maneuvers, but the determination of fuzzy attribute weights relies on expert experience. Zhao Ming-ming applied the Quantum Particle Swarm Optimization (QPSO) algorithm to solve for the Nash equilibrium solution of the fuzzy matrix of the strategy set. Undoubtedly, the above research results have promoted the advancement of UAV intelligent decision-making technology to some extent; however, they only considered single-round games. In fact, multi-round UAV confrontations are more consistent with actual air combat.
[0004] Decision tree algorithms are efficient optimization algorithms in machine learning that predict the category of unknown data by constructing a decision tree from existing known data. Due to their advantages such as low computational cost, high classification accuracy, and simple logic, they have been widely applied in fields such as data mining and cancer analysis. These advantages are well-suited for solving the problem of determining the optimal maneuver in air combat. Zhao Kexin proposed an improved decision tree for situation assessment problems. However, due to the dynamic changes in UAV situation and the complexity of dynamic game decision-making problems in multi-round air combat, existing methods cannot be directly applied to solve the optimal maneuver problem in multi-round confrontations.
[0005] Therefore, it is necessary to study a new dynamic game theory decision-making method to deal with multi-round air combat adversarial problems. The main difficulty of this game theory decision-making problem lies in the fact that the UAV should perform the optimal maneuvers in the next round to gain a better situational advantage and ultimately win in multi-round combat. Summary of the Invention
[0006] Therefore, to address the aforementioned problems, this paper takes a decision tree-based multi-round attack-defense game decision-making method for UAVs as the research object, proposing a multi-round dynamic attack-defense game concept and combining it with the decision tree algorithm to study the UAV air combat situation. First, a multi-round UAV attack-defense game model is established using air combat information. Second, based on the air combat situation in each round, the Nash equilibrium solution and expected payoff value of the mixed strategies for both sides are calculated. Based on the Nash equilibrium solution and expected payoff value, it is determined whether to proceed to the next round. If a next round is required, the decision tree is used to take appropriate maneuvers based on the current air combat situation, and the Nash equilibrium solution and expected payoff value for the current round are recalculated. The above steps are repeated until the final winning condition is met. Finally, simulations verify the effectiveness of the algorithm. This method provides an important reference for UAV air combat analysis and subsequent research on multi-aircraft cooperative decision-making.
[0007] To solve its technical problem, the present invention adopts the following technical solution:
[0008] A decision-making method for multi-round attack-defense game of drones based on decision trees includes the following steps:
[0009] Step 1) Utilize the air combat information of the current round to establish an advantage function model and a drone payoff function model;
[0010] Step 2) Solve for our Nash equilibrium solution and the corresponding expected payoff value, and determine whether to proceed to the next round of the game based on the Nash equilibrium solution and the expected payoff value.
[0011] Step 3) Use the decision tree to take appropriate maneuvers;
[0012] Step 4) Obtain the air combat information of both sides' UAVs after the maneuver, and recalculate the hybrid Nash equilibrium solution and the corresponding expected payoff value after the maneuver;
[0013] Step 5) Repeat steps 1)-4) until the conditions for final victory are met.
[0014] Preferably, in step 1):
[0015] The UAV air combat situation advantage function model is expressed by equation (1):
[0016]
[0017] in, Let i be the velocity advantage function of our drone in round k. Let i be the velocity vector of our drone i in round k. Let $\mathbf{k}$ be the velocity vector of enemy drone $j$ in round $k$. Let R be the range advantage function of our drone i in round k. max R is the maximum launch range of the missile. min For the minimum launch distance of the missile, σ = 10(R) max -R min ), Let i be the distance between our drone i and the enemy drone j in round k. Let i be the altitude advantage function of our drone i in round k. V represents the altitude difference between our drone i and the enemy drone j in round k. Ai Let V be the velocity vector of our UAV i. Bj Let D be the velocity vector of the enemy drone j. ij Let i be the distance between our drone i and the enemy drone j. For V in round k Ai and D ij The angle between them It is the angle advantage function of our drone i in round k. For V in round k Bj and D ij The angle between them as well as These are the weighting coefficients of the speed, distance, altitude, and angle advantage functions in the k-th round, and they satisfy the following relationship:
[0018] The overall advantage function model of the UAV is expressed by equation (2):
[0019]
[0020]
[0021]
[0022]
[0023]
[0024] in, and These are the air-to-air combat capability indices of our drone i and the enemy drone j in round k, respectively. and Let be the maneuverability parameters of the two sides' drones in round k. and Let represent the firepower parameters of the drones of both sides in round k. and Here are the target detection capability parameters for both sides' UAVs in round k. and The parameters represent the operational performance of both sides' drones in round k. and Let be the survivability parameters of both sides' drones in round k. and Let be the range coefficient of both sides' drones in round k. and The electronic warfare capability coefficients of the drones of both sides in round k. Let i be the performance advantage function of our drone i in round k. Let v be the payoff function for our drone i in round k. dj For the value of enemy drone j, v dmax p represents the maximum value of the enemy drone. wij Let i be the probability that our drone i can hit the enemy drone j. Let be the overall advantage function of our drone i over the enemy drone j in round k. as well as These are the weight coefficients of the situational advantage function, performance advantage function, and payoff function in round k, and they satisfy the following relationship:
[0025] The drone revenue payment function model is expressed by equation (3):
[0026] In aerial combat, the set of strategies employed by our drones can be represented as S. A ={s A1 ,s A2 ,…,s Ap ,…,s Ar The set of strategies adopted by the enemy drones can be represented as S. B ={s B1 ,s B2 ,…,s Bq ,…,s Bl}, where r and l are the total number of strategies adopted by our drones and the enemy drones, respectively.
[0027] In a multi-drone combat scenario, our strategy set s at each stage Ap All of these are determined by the actions taken by our m drones. Therefore, any strategy set s Ap All can be expressed in the following forms: sAp ={s Ap1 ,s Ap2 ,…,s Api ,…,s Apm}, where s Api Represents the current strategy s Ap The actions taken by our i-th drone.
[0028] In the current strategy Ap In this scenario, our i-th drone can attack any enemy drone. Therefore, s Api It can also be further expressed as: s Api ={p i1 ,p i2 ,…p ij ,…p in}, where p ij Represents the current strategy s Ap The actions taken by our i-th drone against the j-th enemy drone. Similarly, the enemy drone strategy set can be represented in the following form: s Bq ={s Bq1 ,s Bq2 ,…,s Bqj ,…,s Bqn}, s Bqj ={q j1 ,q j2 ,…q ji ,…q jm}
[0029] When our side adopts strategy Ap The enemy adopted a strategy. Bq At that time, our revenue payout function is defined as follows:
[0030]
[0031] Among them, our strategy s Ap Enemy strategy Bq Down Paying value for our drones in air combat. Paying points for enemy drone air combat, u 2ji Let p be the overall advantage function of enemy UAV j against our UAV i. ij Indicating in strategy s Ap Did our drone i hit the enemy drone j, q? ji Indicating in strategy s Bq Did enemy drone j hit our drone i? ij =1 indicates that our drone is in strategy s Ap Next, our i-th drone attacks the enemy j-th drone, pij =0 indicates that our drone is in strategy s Ap In the following scenario, our i-th drone did not attack the enemy j-th drone; q ji =1 indicates that the enemy drone is in strategy s Bq Next, the enemy's j-th drone attacks our i-th drone; q ji =0 indicates that the enemy drone is in strategy s Bq The enemy's j-th drone did not attack our i-th drone.
[0032] Preferably, in step 2):
[0033] Using the payoff function combined with linear programming, the Nash equilibrium solution for a single-round game can be obtained through equation (2):
[0034]
[0035] The optimal solution obtained This is the Nash equilibrium solution in round k, which has a negative expected payoff. Our drone cannot win, so the game decision tree makes a maneuver decision to start the next round of the game.
[0036] Preferably, the specific process of step 3) is as follows:
[0037] Step 3-1) Both sides have seven maneuvers in their maneuver libraries. Based on the initial position, speed, pitch angle, and yaw angle of both sides, determine which maneuver our UAV should ultimately take under those initial conditions. Continuously change the initial air combat information of both sides' UAVs to obtain multiple sets of air combat data, ultimately forming an air combat data sample set D.
[0038] Step 3-2) Based on the input feature e in the dataset b To determine whether a certain possible value d is obtained, the Gini coefficient for each value of each input feature in the dataset D is calculated using the following formula.
[0039]
[0040] In this case, after partitioning the sample set D according to the value d, the dataset of the left node can be represented as D1, and the dataset of the right node can be represented as D2. and are the number of samples in D1 and D2, respectively, and Gini(D1) and Gini(D2) are the Gini coefficients corresponding to their samples.
[0041] Step 3-3) Among the calculated Gini coefficients of each feature, select the feature with the smallest Gini coefficient e. b The corresponding value d is used as the optimal feature and the optimal split point;
[0042] Steps 3-5) Calculate the Gini coefficients of the datasets D′1 and D′2 for the left and right nodes respectively. If the Gini coefficient of the current node is less than the threshold, stop the construction of the subtree for the current node; otherwise, continue.
[0043] Steps 3-6) Repeat steps 2)-4) for the left and right nodes to generate the preliminary CART classification tree model T0;
[0044] Steps 3-7) Calculate the surface error gain rate g(a) of all non-leaf nodes in the preliminary classification tree model T0 from bottom to top using the following formula;
[0045]
[0046] Where C(a) is the error cost of the subtree with a single node number of a, and C(T) is the error cost of the subtree with a single node number of a. a Let ) be the subtree T with node a as its root. a The cost of error For subtree T a The number of leaf nodes, r(a) is the error rate of node a, and p(a) is the proportion of data at node a to the total data.
[0047] Steps 3-8) Based on the calculation results, the smallest surface error gain ratio among the non-leaf nodes is denoted as α. Then, the preliminary classification tree is traversed from top to bottom to find the node with a surface error gain ratio of α and pruning is performed to obtain subtree T1.
[0048] Steps 3-9) Repeat steps 7)-8) until the root node is obtained. Assuming the tree has undergone n pruning operations, the following subtree sequence can be obtained: {T0, T1, ..., T n};
[0049] Step 3-10) For the subtree sequence {T0,T1,…,T…} n Using an independent validation dataset, the Gini index of each subtree is tested. The decision tree with the smallest Gini index is considered the optimal decision tree.
[0050] The leaf nodes of the optimal decision tree are the corresponding maneuver action numbers. During multi-round offensive and defensive confrontations, when the decision tree is invoked, the input features are given first. Based on the values of the input features, the optimal decision tree is traversed until a leaf node of one of the branches is reached. At this point, the traversal ends and the corresponding maneuver action number is output.
[0051] Preferably, the specific process of step 4) is as follows:
[0052] Step 4-1): Combining the decision results of the optimal decision tree, the maneuver action number is matched with the UAV control quantity. Combining the UAV motion equation and dynamic equation, the following formula is used to obtain the air combat information of the UAV in the next round.
[0053]
[0054] Where x, y, z represent the three-dimensional position of the UAV in the inertial coordinate system, and v, ψ, η represent the UAV's velocity, pitch angle, and yaw angle, respectively. These are the components of the drone's velocity v on the three coordinate axes, and n is the value of n. x ,n f These represent the tangential and normal overloads of the UAV, respectively; φ is the roll angle of the UAV around its velocity vector; and g is the acceleration due to gravity, taken as 9.8 m / s². 2 Based on actual flight dynamics and combined with the particle motion model, the range of attitude angles is determined to be [0, π], where yaw and roll angles are defined as positive to the right and pitch angle as positive upward.
[0055] Step 4-2): Based on the established situational advantage function model, overall advantage function model, and payoff function model, obtain the payoff matrix for the current round;
[0056] Step 4-3): Based on the payoff matrix of the current round, use linear programming to calculate the Nash equilibrium solution and expected payoff value again.
[0057] Step 4-4): Compare the expected payoff value of the current round with the final winning condition to determine whether the game can end in this round.
[0058] The beneficial effects of this invention are as follows:
[0059] 1. Taking into account the actual situation of air combat, the process of multi-round confrontation between UAVs was analyzed, the process of multi-round confrontation was given, and the advantage function model and the UAV payoff function model were established.
[0060] 2. By combining decision trees with the multi-round offensive and defensive confrontation process of UAVs, the relevant process of UAV maneuver decision-making and the relationship between the UAV motion states in each round were analyzed, providing a solution for actual UAV air combat game offense and defense confrontation. Attached Figure Description
[0061] Figure 1 This is a schematic diagram of the UAV's movement in the kth round of this invention;
[0062] Figure 2 This is a flowchart of the multi-round air combat game of the present invention;
[0063] Figure 3This is a flowchart of the air combat decision tree construction process of the present invention;
[0064] Figure 4 This is the particle motion model of the UAV of the present invention;
[0065] Figure 5 This is a schematic diagram of the UAV's movement in the (k+1)th round of this invention;
[0066] Figure 6 The curve showing the change in the expected revenue of our UAV in this invention (2v4);
[0067] Figure 7 The curves showing the position changes of both enemy and friendly UAVs in this invention (2v4). Detailed Implementation
[0068] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments;
[0069] The schematic diagram of the drone's movement in the kth round, upon which this patent research is based, is as follows: Figure 1 As shown, the UAV air combat situation advantage function model is established, as shown in equation (1).
[0070]
[0071] in, Let i be the velocity advantage function of our drone in round k. Let i be the velocity vector of our drone i in round k. Let $\mathbf{k}$ be the velocity vector of enemy drone $j$ in round $k$. Let R be the range advantage function of our drone i in round k. max R is the maximum launch range of the missile. min For the minimum launch distance of the missile, σ = 10(R) max -R min ), Let i be the distance between our drone i and the enemy drone j in round k. Let i be the altitude advantage function of our drone i in round k. V represents the altitude difference between our drone i and the enemy drone j in round k. Ai Let V be the velocity vector of our UAV i. Bj Let D be the velocity vector of the enemy drone j. ij Let i be the distance between our drone i and the enemy drone j. For V in round k Ai and D ij The angle between them It is the angle advantage function of our drone i in round k. For V in round k Bj and D ij The angle between them as well as These are the weighting coefficients of the speed, distance, altitude, and angle advantage functions in the k-th round, and they satisfy the following relationship:
[0072] The overall advantage function model of the UAV is established as shown in equation (2) below.
[0073]
[0074]
[0075]
[0076]
[0077]
[0078] in, and These are the air-to-air combat capability indices of our drone i and the enemy drone j in round k, respectively. and Let be the maneuverability parameters of the two sides' drones in round k. and Let represent the firepower parameters of the drones of both sides in round k. and Here are the target detection capability parameters for both sides' UAVs in round k. and The parameters represent the operational performance of both sides' drones in round k. and Let be the survivability parameters of both sides' drones in round k. and Let be the range coefficient of both sides' drones in round k. and The electronic warfare capability coefficients of the drones of both sides in round k. Let i be the performance advantage function of our drone i in round k. Let v be the payoff function for our drone i in round k. dj For the value of enemy drone j, v dmax p represents the maximum value of the enemy drone. wij Let i be the probability that our drone i can hit the enemy drone j. Let be the overall advantage function of our drone i over the enemy drone j in round k. as well as These are the weight coefficients of the situational advantage function, performance advantage function, and payoff function in round k, and they satisfy the following relationship:
[0079] The drone payment function model is established as shown in equation (3) below.
[0080] In aerial combat, the set of strategies employed by our drones can be represented as S. A ={s A1 ,s A2 ,…,s Ap ,…,s Ar The set of strategies adopted by the enemy drones can be represented as S. B ={s B1 ,s B2 ,…,s Bq ,…,s Bl}, where r and l are the total number of strategies adopted by our drones and the enemy drones, respectively.
[0081] In a multi-drone combat scenario, our strategy set s at each stage Ap All of these are determined by the actions taken by our m drones. Therefore, any strategy set s Ap All can be expressed in the following forms: s Ap ={s Ap1 ,s Ap2 ,…,s Api ,…,s Apm}, where s Api Represents the current strategy s Ap The actions taken by our i-th drone.
[0082] In the current strategy Ap In this scenario, our i-th drone can attack any enemy drone. Therefore, s Api It can also be further expressed as: s Api ={p i1 ,p i2 ,…p ij ,…p in}, where p ij Represents the current strategy s Ap The actions taken by our i-th drone against the j-th enemy drone. Similarly, the enemy drone strategy set can be represented in the following form: s Bq ={s Bq1 ,s Bq2 ,…,s Bqj ,…,s Bqn}, s Bqj ={q j1 ,q j2 ,…qji ,…q jm}
[0083] When our side adopts strategy Ap The enemy adopted a strategy. Bq At that time, our revenue payout function is defined as follows:
[0084]
[0085] Among them, our strategy s Ap Enemy strategy Bq Down Paying value for our drones in air combat. Paying points for enemy drone air combat, u 2ji Let p be the overall advantage function of enemy UAV j against our UAV i. ij Indicating in strategy s Ap Did our drone i hit the enemy drone j, q? ji Indicating in strategy s Bq Did enemy drone j hit our drone i? ij =1 indicates that our drone is in strategy s Ap Next, our i-th drone attacks the enemy j-th drone, p ij =0 indicates that our drone is in strategy s Ap In the following scenario, our i-th drone did not attack the enemy j-th drone; q ji =1 indicates that the enemy drone is in strategy s Bq Next, the enemy's j-th drone attacks our i-th drone; q ji =0 indicates that the enemy drone is in strategy s Bq The enemy's j-th drone did not attack our i-th drone.
[0086] The specific process of step 2) is as follows:
[0087] According to equation (3), we can obtain the expression for our air combat payout matrix in round k.
[0088]
[0089] in, This is our air combat payoff matrix for round k. Let be our payoff function in round k.
[0090] The equation for solving the Nash equilibrium is as follows. The Nash equilibrium solution for a single-round game can be obtained by combining linear programming.
[0091]
[0092] Where A is the set of real numbers with values in the range [0,1].
[0093] The optimal solution obtained This is the Nash equilibrium solution in round k, which has a negative expected payoff. Our drone cannot win, so the game decision tree makes a maneuver decision to start the next round of the game.
[0094] The specific process of step 3) is as follows:
[0095] Step 3-1) Both sides have seven maneuvers in their maneuver libraries. Based on the initial position, speed, pitch angle, and yaw angle of both sides, determine which maneuver our UAV should ultimately take under those initial conditions. Continuously change the initial air combat information of both sides' UAVs to obtain multiple sets of air combat data, ultimately forming an air combat data sample set D.
[0096] The correspondence between input features and air combat information is shown in Table 1, and the correspondence between output attributes and maneuver action numbers is shown in Table 2.
[0097] Table 1 Input Features of Decision Tree
[0098]
[0099] Step 3-2) Based on the input feature e in the dataset b To determine whether a certain possible value d is obtained, the Gini coefficient for each value of each input feature in the dataset D is calculated using the following formula.
[0100]
[0101] In this case, after partitioning the sample set D according to the value d, the dataset of the left node can be represented as D1, and the dataset of the right node can be represented as D2. and are the number of samples in D1 and D2, respectively, and Gini(D1) and Gini(D2) are the Gini coefficients corresponding to their samples.
[0102] Step 3-3) Among the calculated Gini coefficients of each feature, select the feature with the smallest Gini coefficient e. b The corresponding value d is used as the optimal feature and the optimal split point;
[0103] Steps 3-4) Based on the optimal feature e b Given the optimal split point d, the dataset D is divided into two parts D′1 and D′2. At the same time, two child nodes of the current node are generated, with the dataset of the left node being D′1 and the dataset of the right node being D′2.
[0104] Steps 3-5) Calculate the Gini coefficients of the datasets D′1 and D′2 for the left and right nodes respectively. If the Gini coefficient of the current node is less than the threshold, the construction of the subtree for the current node is stopped; otherwise, continue.
[0105] Steps 3-6) Repeat steps 2)-4) for the left and right nodes to generate the preliminary CART classification tree model T0;
[0106] Steps 3-7) Calculate the surface error gain rate g(a) of all non-leaf nodes in the preliminary classification tree model T0 from bottom to top using the following formula;
[0107]
[0108] Where C(a) is the error cost of the subtree with a single node number of a, and C(T) is the error cost of the subtree with a single node number of a. a Let ) be the subtree T with node a as its root. a The cost of error For subtree T a The number of leaf nodes, r(a) is the error rate of node a, and p(a) is the proportion of data at node a to the total data.
[0109] Steps 3-8) Based on the calculation results, the smallest surface error gain ratio among the non-leaf nodes is denoted as α. Then, the preliminary classification tree is traversed from top to bottom to find the node with a surface error gain ratio of α and pruning is performed to obtain subtree T1.
[0110] Steps 3-9) Repeat steps 7)-8) until the root node is obtained. Assuming the tree has undergone n pruning operations, the following subtree sequence can be obtained: {T0, T1, ..., T n};
[0111] Step 3-10) For the subtree sequence {T0,T1,…,T…} n Using an independent validation dataset, the Gini index of each subtree is tested. The decision tree with the smallest Gini index is considered the optimal decision tree.
[0112] The leaf nodes of the optimal decision tree are the corresponding maneuver action numbers. During multi-round offensive and defensive confrontations, when the decision tree is invoked, the input features are given first. Based on the values of the input features, the optimal decision tree is traversed until a leaf node of one of the branches is reached. At this point, the traversal ends and the corresponding maneuver action number is output.
[0113] The specific process of step 4) is as follows:
[0114] Step 4-1): Combining the decision results of the optimal decision tree, the maneuver action number is matched with the UAV control quantity. Combining the UAV motion equation and dynamic equation, the following formula is used to obtain the air combat information of the UAV in the next round.
[0115]
[0116] Where x, y, z represent the three-dimensional position of the UAV in the inertial coordinate system, and v, ψ, η represent the UAV's velocity, pitch angle, and yaw angle, respectively. These are the components of the drone's velocity v on the three coordinate axes, and n is the value of n. x ,n f These represent the tangential and normal overloads of the UAV, respectively; φ is the roll angle of the UAV around its velocity vector; and g is the acceleration due to gravity, taken as 9.8 m / s². 2 Based on actual flight dynamics and combined with the particle motion model, the range of attitude angles is determined to be [0, π], where yaw and roll angles are defined as positive to the right and pitch angle as positive upward.
[0117] Step 4-2): Based on the established situational advantage function model, overall advantage function model, and payoff function model, obtain the payoff matrix for the current round;
[0118] Step 4-3): Based on the payoff matrix of the current round, use linear programming to calculate the Nash equilibrium solution and expected payoff value again.
[0119] Step 4-4): Compare the expected payoff value of the current round with the final winning condition to determine whether the game can end in this round.
[0120] In summary, the multi-round air combat game analysis process is as follows: Figure 2 As shown, future research will incorporate incomplete information to conduct more in-depth research on game theory attack and defense decision-making.
[0121] The process of a decision tree-based multi-round attack-defense game decision-making method for drones is as follows:
[0122] Step 1: Utilize the air combat information of the current round to establish an advantage function model and a drone payoff function model.
[0123] Step 2: Solve for the Nash equilibrium solution of the mixed strategy of both parties and the corresponding expected payoff value. Based on the Nash equilibrium solution and expected payoff value, determine whether to proceed to the next round of the game.
[0124] Step 3: Use the decision tree to take appropriate maneuvers.
[0125] Step 4: Obtain the air combat information of both sides' drones after the maneuver, and recalculate the hybrid Nash equilibrium solution and corresponding expected payoff value after the maneuver.
[0126] Step 5: Repeat steps 1)-4) until the conditions for final victory are met.
[0127] To demonstrate the effectiveness of this method, simulation analysis was conducted. The correspondence between maneuvering actions and control variables in the simulation is shown in Table 3. The three-dimensional position, velocity, pitch angle, and yaw angle information of our UAV are shown in Table 4, and the three-dimensional position, velocity, pitch angle, and yaw angle information of the enemy UAV are shown in Table 5.
[0128] Table 3. Correspondence between maneuvering actions and control quantities
[0129]
[0130]
[0131] Table 4. Three-dimensional position, velocity, pitch angle, and yaw angle information of our UAV.
[0132]
[0133] Table 5. Enemy UAV 3D position, velocity, pitch angle, and yaw angle information. Remark: First, a preliminary air combat game decision tree is established, and then a simplified air combat game decision tree is obtained using a pruning algorithm. Based on the initial position, velocity, pitch angle, and yaw angle information of both sides in the table above, the initial air combat payoff matrix of our UAV is obtained, and the simulation results are shown below.
[0134]
[0135] Based on the obtained payoff matrix, the Nash equilibrium value and expected return value of the hybrid strategy are solved using a linear method. The Nash equilibrium solution for our UAV is: x * =00000000010,0,0,0,0,0), our expected gain is -0.321. At this point, our gain is negative, and we believe that we cannot achieve victory in the air battle, so we decide to engage in a multi-round offensive and defensive game.
[0136] The game confrontation time for each round is set to Δt = 1s, and the maximum simulation time is T = 50Δt. The two sides conduct simulated air combat confrontation, and the confrontation ends when the expected value of our drone is greater than 0 or the maximum simulation time is reached.
[0137] 1) Assuming we have two drones and the enemy has four drones, and an air combat engagement is underway, the expected payoff curve for our drones is as follows: Figure 6 As shown.
[0138] 2) Assuming our side has two drones and the enemy has four drones, and an air combat exercise is conducted, the position change curves of both sides' drones are as follows: Figure 7 As shown.
[0139] Under the aforementioned initial conditions, after 27 rounds of combat, our expected return is greater than 0, meaning we can achieve victory in the air battle. The expected return at the end is 0.02. During the air battle, when our drone sorties are fewer and our initial situation is not advantageous, we can employ a decision tree-based offensive and defensive game theory algorithm to analyze the situation of both sides in real time, maneuver multiple times to change our own motion state, thereby gaining a superior position and ultimately achieving victory with fewer troops.
[0140] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A decision-making method for multi-round attack-defense game of unmanned aerial vehicles based on decision trees, characterized in that, Includes the following steps: Step 1) Establish the UAV air combat situation advantage function model and the UAV payoff function model; Step 2) Solve for our Nash equilibrium solution and the corresponding expected payoff value, and determine whether to proceed to the next round of the game based on the Nash equilibrium solution and the expected payoff value. Step 3) Use the decision tree to take appropriate maneuvers; Step 4) Obtain the air combat information of both sides' UAVs after the maneuver, and recalculate the hybrid Nash equilibrium solution and the corresponding expected payoff value after the maneuver; Step 5) Repeat steps 1)-4) until the conditions for final victory are met; In step 2): According to equation (3), the expression for the payoff matrix of our UAV in the k-th round of air combat is obtained: in, This is our air combat payoff matrix for round k. Let this be our payoff function in round k; The Nash equilibrium solution for a single-round offensive and defensive game is obtained by combining linear programming: Where A is the set of real numbers with values in the range [0,1]; The optimal solution obtained This is the Nash equilibrium solution of our attack and defense game in round k. If the expected payoff corresponding to the Nash equilibrium solution of the attack and defense game in round k is negative, then our drone cannot win. Therefore, the decision tree adopts a flexible decision to carry out the next round of game confrontation. The specific implementation process of step 4) is as follows: Step 4-1): Combining the decision results of the optimal decision tree, the maneuver action number is matched with the UAV control quantity. Combining the UAV motion equation and dynamic equation, the air combat information of the UAV in the next round is obtained using equation (8). Where x, y, z represent the three-dimensional position of the UAV in the inertial coordinate system, and v, ψ, η represent the UAV's velocity, pitch angle, and yaw angle, respectively. These are the components of the drone's velocity v on the three coordinate axes, and n is the value of n. x ,n f These represent the tangential and normal overloads of the UAV, respectively; φ is the roll angle of the UAV around its velocity vector; and g is the gravitational acceleration, taken as 9.8 m / s². 2 The yaw and roll angles are defined as positive for right yaw and positive for pitch angle, and both are within the range of [0,π]. Step 4-2): Based on the established UAV air combat situation advantage function model, UAV overall advantage function model, and UAV payoff function model, obtain the payoff matrix for the current round; Step 4-3): Based on the payoff matrix of the current round, calculate the Nash equilibrium solution and expected payoff using linear programming. Step 4-4): Compare the expected payoff value of the current round with the final winning condition to determine whether the game can end in this round.
2. The decision-making method for multi-round attack and defense game of unmanned aerial vehicles based on decision trees according to claim 1, characterized in that, In step 1): Assume the number of rounds in the attack and defense game between the two sides' drones is {1,2,…,k,k+1,…,K}. Our drones did not win in the first k-1 rounds, and the position, speed, pitch angle, and yaw angle information of both sides' drones have been obtained in the kth round. The UAV air combat situation advantage function model is expressed by equation (1): in, Let i be the velocity advantage function of our drone in round k. Let i be the velocity vector of our drone i in round k. Let $\mathbf{k}$ be the velocity vector of enemy drone $j$ in round $k$. Let R be the range advantage function of our drone i in round k. max R is the maximum launch range of the missile. min For the minimum launch distance of the missile, σ = 10(R) max -R min ), Let i be the distance between our drone i and the enemy drone j in round k. Let i be the altitude advantage function of our drone i in round k. V represents the altitude difference between our drone i and the enemy drone j in round k. Ai Let V be the velocity vector of our UAV i. Bj Let D be the velocity vector of the enemy drone j. ij Let i be the distance between our drone i and the enemy drone j. For V in round k Ai and D ij The angle between them It is the angle advantage function of our drone i in round k. For V in round k Bj and D ij The angle between them as well as These are the weighting coefficients of the speed, distance, altitude, and angle advantage functions in the k-th round, and they satisfy the following relationship: The overall advantage function model of the UAV is expressed by equation (2): in, and These are the air-to-air combat capability indices of our drone i and the enemy drone j in round k, respectively. and Let be the maneuverability parameters of the two sides' drones in round k. and Let represent the firepower parameters of the drones of both sides in round k. and Here are the target detection capability parameters for both sides' UAVs in round k. and The parameters represent the operational performance of the drones used by both sides in round k. and Let be the survivability parameters of both sides' drones in round k. and Let be the range coefficient of both sides' drones in round k. and The electronic warfare capability coefficients of the drones of both sides in round k. Let i be the performance advantage function of our drone i in round k. Let v be the payoff function for our drone i in round k. dj For the value of enemy drone j, v dmax p represents the maximum value of the enemy drone. wij Let i be the probability that our drone i can hit the enemy drone j. Let be the overall advantage function of our drone i over the enemy drone j in round k. as well as These are the weight coefficients of the situational advantage function, performance advantage function, and payoff function in round k, and they satisfy the following relationship: In aerial combat, the set of strategies adopted by our drones is denoted as S. A ={s A1 ,s A2 ,…,s Ap ,…,s Ar }, s Ar S represents the set of strategies adopted by our drones and the set of strategies adopted by the enemy drones. B ={s B1 ,s B2 ,…,s Bq ,…,s Bl }, s Bl The strategies adopted by enemy drones, where r is the total number of strategies adopted by friendly drones and l is the total number of strategies adopted by enemy drones; The strategy adopted by our drones Ap Represented as: s Ap ={s Ap1 ,s Ap2 ,…,s Api ,…,s Apm }, where s Api This represents the current strategy adopted by our drones. Ap The actions taken by our i-th drone; The current strategy adopted by our drones Ap In the middle, our i-th drone takes action to attack any enemy drone, therefore, s Api Further expressed as: s Api ={p i1 ,p i2 ,…p ij ,…p in }, where p ij This represents the current strategy adopted by our drones. Ap The actions taken by our i-th drone against the enemy j-th drone; Similarly, the set of strategies adopted by the enemy drones is represented as: s Bq ={s Bq1 ,s Bq2 ,…,s Bqj ,…,s Bqn }, will s Bqj Further expressed as: s Bqj ={q j1 ,q j2 ,…q ji ,…q jm }; When our drones adopt strategy Ap The enemy drones adopted a strategy Bq At that time, the payment function f of our drone apq The model is defined as follows: Among them, our strategy s Ap Enemy strategy Bq Down Paying value for our drones in air combat. Paying points for enemy drone air combat, u 2ji Let p be the overall advantage function of enemy UAV j against our UAV i. ij Indicating in strategy s Ap Did our drone i hit the enemy drone j, q? ji Indicating in strategy s Bq Did the enemy drone J hit our drone i? ij =1 indicates that our drone is in strategy s Ap Next, our i-th drone attacks the enemy j-th drone, p ij =0 indicates that our drone is in strategy s Ap In the following scenario, our i-th drone did not attack the enemy j-th drone; q ji =1 indicates that the enemy drone is in strategy s Bq Next, the enemy's j-th drone attacks our i-th drone; q ji =0 indicates that the enemy drone is in strategy s Bq The enemy's j-th drone did not attack our i-th drone.
3. The decision-making method for multi-round attack and defense game of unmanned aerial vehicles based on decision trees according to claim 2, characterized in that, The specific implementation process of step 3) is as follows: Step 3-1) Both sides have seven maneuvering actions in their maneuvering action libraries. Based on the initial position, speed, pitch angle, and yaw angle information of the UAVs of both sides, we first calculate the current gain of our UAV, and then calculate the gain of our UAV under different maneuvers. The maneuvering action corresponding to the maximum gain is considered to be the final maneuvering action taken by our UAV under the initial conditions. We continuously change the initial air combat information of the UAVs of both sides. The air combat information includes position, speed, pitch angle, and yaw angle information to obtain multiple sets of air combat data, which finally constitute the air combat data sample set D. The position, speed, pitch angle, and yaw angle of the UAVs of both sides are the input features E = {e1, e2, e3, e4, e5, e6, e7, e8} of the sample set D, where e1-e8 are the numbers of the initial air combat information, and the output attribute W = {w1, w2, w3, w4, w5, w6, w7}, where w1-w7 are the numbers of the maneuvers, i.e., D = {E, W}. Step 3-2) Combine equation (6) to calculate the input feature e in the air combat data sample set D. b The Gini coefficient for the value d of (b = 1, 2, ..., 8): (6) D1={(e b ,w a )|value(e b )<d,a=1,2,…,7} D2={(e b ,w a )|value(e b )≥d,a=1,2,…,7} In this case, after partitioning the sample set D according to the value d, the dataset of the left node is represented as D1, and the dataset of the right node is represented as D2. and , respectively, are the number of samples in D1 and D2, and Gini(D1) and Gini(D2) are the Gini coefficients corresponding to their samples; Step 3-3) Among the Gini coefficients calculated in Step 3-2), select the input feature corresponding to the smallest Gini coefficient as the optimal input feature e. b ', optimal input feature e b The corresponding value is taken as the optimal split point d'; Steps 3-4) Based on the optimal input feature e b The optimal split point d divides the air combat data sample set D into two parts, D1' and D'2, and generates two child nodes for the current node. The dataset of the left node is D1' and the dataset of the right node is D'2. Steps 3-5) Calculate the Gini coefficient for the datasets D1' and D'2 of the left and right nodes respectively. If the Gini coefficient of the current node is less than the threshold, stop the construction of the subtree for the current node; otherwise, continue. Steps 3-6) Repeat steps 2)-4) for the left and right nodes to generate the preliminary CART classification tree model T0; Steps 3-7) Use equation (7) to calculate the surface error gain rate g(a) of all non-leaf nodes in the preliminary CART classification tree model T0 from bottom to top; C(a) = r(a)·p(a) Where C(a) is the error cost of the subtree with a single node number of a, and C(T) is the error cost of the subtree with a single node number of a. a Let ) be the subtree T with node a as its root. a The cost of error For subtree T a The number of leaf nodes, r(a) is the error rate of node a, and p(a) is the proportion of data at node a to the total data; Steps 3-8) Based on the calculation results, the smallest surface error gain ratio among the non-leaf nodes is denoted as α. Then, the preliminary classification tree is traversed from top to bottom to find the node with a surface error gain ratio of α and pruning is performed to obtain subtree T1. Steps 3-9) Repeat steps 7)-8) until the root node is obtained. Assuming the tree has undergone n pruning operations, the resulting subtree sequence is: {T0, T1, ..., T...} n }; Step 3-10) For the subtree sequence {T0,T1,…,T…} n Using an independent validation dataset, the Gini index of each subtree is tested, and the decision tree with the smallest Gini index is considered the optimal decision tree. The leaf nodes of the optimal decision tree are the corresponding maneuver action numbers. In the process of multi-round offensive and defensive confrontation game, when the decision tree is called, the input features are given first. The optimal decision tree is traversed according to the value of the input features until the leaf node of one of the branches is reached. Then the traversal ends and the corresponding maneuver action number is output.
Citation Information
Patent Citations
A dynamic game method for multi-unmanned aerial vehicle air battle under uncertain information
CN107463094A
UAV decision and control system
US20100163621A1