Multi-Aircraft Air Combat Decision-Making Method Based on General Experience Game Reinforcement Learning

The integration of deep reinforcement learning and game theory frameworks addresses the challenge of incomplete information in multi-aircraft combat by optimizing strategies through iterative updates, enhancing collaborative decision-making and strategy effectiveness.

CN116468121BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310232390.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-12
Publication Date
2025-07-15
Estimated Expiration
2043-03-12

AI Technical Summary

Technical Problem

The existing multi-agent method is difficult to train better agents in multi-air air combat due to conflicts of interest, resulting in poor air combat strategies and difficult to apply in actual combat environments.

Method used

A method based on general empirical game reinforcement learning is adopted, combined with LSTM neural network and fully connected neural network, an agent structure is designed, and a payment matrix and α-rank algorithm optimization strategy is formed to form a better synergistic air combat strategy.

Benefits of technology

In multi-air air combat, the independent coordinated decision-making ability of the agent is realized, the coordination and effectiveness of the air combat strategy are improved, and the complex air combat situation is adapted to complex air combat situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468121B_ABST
    Figure CN116468121B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-aircraft air combat decision-making method based on general experience game reinforcement learning. The state space vector of the fighter aircraft is selected as the input of the intelligent agent structure. The intelligent agent includes action and evaluation neural networks, which are LSTM neural networks and fully connected neural networks. The payoff matrix is calculated, and the Nash equilibrium under the current payoff matrix is solved to optimize the new strategy of the k-th fighter aircraft at present. The reinforcement learning algorithm is used for solving, and new intelligent agents are generated and added to the corresponding strategy sets, and the process is repeated until the algorithm reaches the specified number of training iterations. After the entire game framework is trained to the specified number of iterations, the latest trained intelligent agent is extracted, relevant information is input, and maneuvering actions can be output to complete one-step decision-making. The present invention can form complex maneuvering actions and cooperative tactics. Compared with the intelligent agents trained by general multi-intelligent agent reinforcement learning algorithms, the decision-making among multiple aircraft is more collaborative, which also demonstrates the effectiveness of the method of the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of game theory and multi - agent reinforcement learning, and relates to a multi - aircraft air combat decision - making method based on general experience game reinforcement learning. Background Art

[0002] At present, most of the research in the field of air combat decision - making considers problems under the overly idealized background of model certainty and full observability. However, in the actual combat environment, it is very difficult to obtain complete state information of enemy aircraft. If the non - integrity of information is directly ignored, the decision - making effect will be greatly reduced, and at the same time, it will also make it difficult for research results to be implemented. Modern air combat has gradually developed to cluster or even swarm combat. Multi - aircraft combat can achieve the effect of "1 + 1>2". It is necessary to further promote the previous single - aircraft combat decision - making research experience to the multi - aircraft combat scenario.

[0003] Autonomous decision - making is to give the maneuvering actions that the fighter should take according to the current air combat situation information, which puts high requirements on the real - time nature of decision - making. The deep reinforcement learning method adopted by the present invention can well solve the sequential decision - making problem such as air combat game. The introduction of a deep neural network can accurately perceive the situation information. Combining with the reinforcement learning algorithm and introducing the game theory framework to solve the strong antagonism in the air combat environment, finally, the optimal mapping relationship from the situation to the maneuver can be obtained. Summary of the Invention

[0004] Technical Problems to be Solved

[0005] In order to avoid the deficiencies of the prior art, the present invention proposes a multi - aircraft air combat decision - making method based on general experience game reinforcement learning. When solving multi - aircraft air combat with existing multi - agent methods, due to the interest conflicts among multi - aircraft, it is difficult to train better agents. On the basis of solving the air combat strategy using the commonly used reinforcement learning algorithm, an experience game framework is further added here, so that the capabilities of the agents can be steadily improved, and better coordination and more optimal air combat strategies can be obtained in multi - aircraft air combat.

[0006] Technical Solution

[0007] A multi - aircraft air combat decision - making method based on general experience game reinforcement learning is characterized by the following steps:

[0008] Step 1: The fighter in the simulation platform is a dynamic model in three - dimensional space, and the control input of each fighter is the roll angle and two control inputs N x , N z , and they are encoded; N x represents the tangential overload, and N z represents the normal overload; the maneuvering actions include seven types: constant, deceleration, acceleration, left turn, right turn, pull - up, and dive.

[0009] The state update equation of the simulation platform:

[0010] s t+1 = f(s t , u m1 , u m2 ,…, u mi ,…, u mN , T)

[0011] f represents the state update function composed of a system of differential equations, and u mi represents the basic maneuver of the i-th fighter jet, and T represents the simulation time step;

[0012] Step 2: Each fighter jet's agent includes an action neural network and an evaluation neural network. The neural network structure is an LSTM neural network connected to a fully connected neural network;

[0013] The air combat situation, that is, the sequence of observables, is the input of the LSTM neural network, and the global state information is extracted as the output;

[0014] This output is the input of the fully connected neural network. In the action network, the output layer of the fully connected neural network adds a softmax function to normalize the output value into a probability form, and the output is the 7 kinds of maneuvers in Step 1. The output of the evaluation network is the value of the value function;

[0015] The fully connected neural network uses the ReLu activation function;

[0016] Step 3: Each fighter jet has a policy set, and each policy set is equipped with multiple groups of action neural networks and evaluation neural networks; the input of each group of neural networks is the air combat situation sequence;

[0017] Add a group of action and evaluation neural networks with random parameters to the policy set of each fighter jet to initialize the policy sets of K fighter jets;

[0018] Step 4: For the policy sets of K fighter jets, the dimension of the payoff matrix changes with the number of iterations as:

[0019] n i1 ×n i2 ×…×n iK →n (i+1)1 ×n (i+1)2 ×...×n (i+1)K , n (i+1)j = n ij + 1

[0020] Where: n ijDenote the strategy of the $j$-th fighter plane in the $i$-th iteration. Since each fighter plane will generate a new optimal agent in one round of agent optimization, we have $n$ (i+1)j = $n$ ij + 1;

[0021] When adding a new strategy to the strategy set of each fighter plane, the missing terms of the payoff matrix are:

[0022]

[0023] $s$ nwk Denote the new strategy generated by the $k$-th fighter plane in the $w$-th iteration;

[0024] The value of the missing term of the payoff matrix is obtained through air combat simulation by the following formula;

[0025] $sim$ k $(s_1, s_2, \cdots, s$ k , \cdots, s$ K )

[0026] = $\sum$ j∈{1,2,,K},j≠k $eval\_oto$ k $(s$ k , s$ j )

[0027] In the $n$-player game, the strategy $s$ k is the agent decision-making model, and $eval\_oto$ k $(s$ k , s$ j ) is the one-on-one air combat situation assessment that regards the agent $s$ k as its own side and $s$ j as the enemy;

[0028] Select the air combat situation in the air combat situation assessment: And the win-loss judgment reward: The weighted sum of is used as the air combat ability assessment of the agent:

[0029] $R$ e = $\omega_1\times R$ a + $\omega_2\times R$ w , $\omega_1+\omega_2 = 1$;

[0030] Step 5: Use the $\alpha$-rank algorithm to obtain the approximate Nash equilibrium under the current payoff matrix, that is, each fighter plane selects a set of actions and evaluation neural networks from its strategy set, and the selected multiple neural networks form a strategy combination;

[0031] Step 6: Under the strategy combination obtained in Step 5, except for the k-th fighter jet, the other K-1 fighter jets use their respective selected action networks as decision-making mechanisms to control their own maneuvers. The PPO algorithm is used as the basic reinforcement learning algorithm to optimize the new strategy of the current k-th fighter jet, generate new action and evaluation networks, and add them to the strategy set of the k-th fighter jet. Optimize each fighter jet in turn, and at the same time add the generated new agents to the corresponding strategy sets, and return to Step 4 until the set number of training iterations is met;

[0032] Step 7: Select the latest intelligent agents of each game player as decision-makers, extract their intelligent agent action neural networks as decision-making networks, and the observed air combat situation state vector is the input of the LSTM neural network. After passing through the LSTM neural network and the global neural network, the maneuvering actions are output to complete one-step decision-making.

[0033] The tangential overload N x varies in the range of [-2, 2], and the normal overload N z varies in the range of [-5, 5], and the roll angle varies in the range of [-π / 3, π / 3].

[0034] The dynamic model of the fighter jet in the three-dimensional ground coordinate system:

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041] On the left side of each differential equation are the first-order derivatives of the flight speed, yaw angle, pitch angle, and three-dimensional position of the fighter jet respectively.

[0042] Beneficial effects

[0043] A multi-aircraft air combat decision-making method based on general experience game reinforcement learning proposed by the present invention first establishes a dynamic model of a fighter plane in three-dimensional space, and then designs a platform state transition module, selects the state space vector of the fighter plane, and designs an air combat training platform. Subsequent intelligent agent training and decision algorithm verification will be carried out on this platform. Design the intelligent agent structure. The intelligent agent includes action and evaluation neural networks. To adapt to the partially observable environment, the neural network is designed as an LSTM neural network and a fully connected neural network. Calculate the payoff matrix. A strategy is initialized in the strategy space of each fighter plane. The strategy is set as the action and evaluation neural networks, and the parameters are random. The meta-game payoff matrix needs to use the intelligent agents in the fighter plane strategy space to improve the payoff matrix, and fill in the missing items in the payoff matrix when there are new strategies in each round of iteration. When forming the new optimal strategy of the k-th fighter plane, it needs to be carried out under the fixed strategies of other fighter planes. Since the strategy of the meta-game in the present invention is an intelligent agent. Therefore, based on the intelligent agents fixed by other fighter planes, the k-th fighter plane generates a new optimal intelligent agent and adds it to the strategy set. To ensure that the newly obtained optimal intelligent agent is the best response BR, it is necessary to solve the Nash equilibrium under the current payoff matrix. Under the strategy combination of the Nash equilibrium, optimize the new strategy of the current k-th fighter plane. Enter the new intelligent agent expansion stage. In the previous stage, the Nash equilibrium has been solved. When solving the optimal new strategy for each fighter plane, new intelligent agents need to be further generated on the basis of the Nash equilibrium strategy combination fixed by other fighter planes. And when the decision-making models of other intelligent agents are fixed, the present invention uses a reinforcement learning algorithm to solve. Relative to the intelligent agent to be optimized currently, other intelligent agents or the air combat simulation environment all belong to the environment in reinforcement learning. Generate new intelligent agents and add them to the corresponding strategy sets, and repeat until the algorithm reaches the specified number of training iterations. When the entire game framework is trained to the specified number of iterations, extract the newly trained intelligent agent. Its action network directly uses it as the decision-making network. Input relevant information, and it can output maneuvering actions to complete one-step decision-making. Through multiple rounds of iterative learning, the friendly intelligent agents have highly autonomous cooperative decision-making capabilities when facing complex air combat situations.

[0044] Through the combination of maneuvering instructions among the fighter planes of the present invention, complex maneuvering actions and cooperative tactics can be formed. Compared with the intelligent agents trained by general multi-intelligent agent reinforcement learning algorithms, the decision-making among multiple aircraft is more collaborative, which also illustrates the effectiveness of the method of the present invention. Brief Description of the Drawings

[0045] Figure 1 Are schematic diagrams of seven maneuvering actions

[0046] Figure 2 Are schematic diagrams of the state update module

[0047] Figure 3 Is the interaction diagram of the population.

[0048] Figure 4 It is the structure diagram of the action network.

[0049] Figure 5 It is the structure diagram of the evaluation network.

[0050] Figure 6 It is the schematic diagram of the clip function clip.

[0051] Figure 7 It is the schematic diagram of the min function.

[0052] Figure 8 It is the flow chart of the intelligent agent optimization based on the PPO algorithm.

[0053] Figure 9 It is the reward accumulation curve.

[0054] Figure 10 It is the availability comparison curve.

[0055] Figure 11 It is the reward value of comparing the fully connected neural network and the LSTM neural network as the decision network.

[0056] Figure 12 It is the three-dimensional air combat trajectory map at the end of the 15th round of the training of the friendly intelligent agent.

[0057] Figure 13 It is the side view of the air combat trajectory at the end of the 15th round of the training of the friendly intelligent agent.

[0058] Figure 14 It is the top view of the air combat trajectory at the end of the 15th round of the training of the friendly intelligent agent. Specific implementation manners

[0059] The present invention will be further described below in combination with embodiments and the accompanying drawings:

[0060] The implementation manner of Step 1 is as follows:

[0061] (1) Design of the fighter aircraft mathematical model

[0062] Air combat is essentially a sequential decision-making problem. The simulation platform needs to calculate the state at the next moment according to the simulation step length to reflect the dynamic change process of the air combat state. The mathematical description of the state change is a system of differential equations. The object of study in this article is a fixed-wing fighter aircraft. This type of aircraft generally flies within the atmosphere, and its flight altitude is limited and the time and space span of air combat confrontation is small. In order to simplify the air combat problem studied, the following reasonable assumptions are necessary:

[0063] (1) Ignore the earth's curvature, that is, adopt the "flat earth hypothesis";

[0064] (2) Ignoring the rotation and revolution of the Earth, the ground coordinate system is considered as an inertial coordinate system;

[0065] (3) Ignoring the mass loss of the fighter jet caused by factors such as fuel consumption, that is, the mass of the fighter jet particle is constant;

[0066] (4) Ignoring the frictional resistance in the air, etc.;

[0067] (5) Ignoring the air flow velocity, that is, the airspeed is 0, and the airspeed of the aircraft is the speed of the aircraft in the ground coordinate system;

[0068] (6) Ignoring the sideslip angle of the fighter jet.

[0069] Based on the above assumptions, the dynamic model of the fighter jet as shown in Equation (1) can be obtained in the three-dimensional ground coordinate system.

[0070]

[0071] In the above formulas, the left sides of each differential equation are respectively the first derivatives of the flight speed, yaw angle, pitch angle, and three-dimensional position of the fighter jet. Some fixed-wing aircraft use the rudder to control the heading, but the turning efficiency is too low. Here, it is assumed that there is no sideslip angle, and the roll angle is used to control the coordinated turn of the fighter jet. Therefore, the roll angle is not considered as a state here and is regarded as one of the control inputs to control the turn of the fighter jet. Except for the roll angle In addition, there are two other control inputs N x , N z , N x represents the tangential overload, which can be calculated by the ratio of the thrust along the axis of the fighter jet to the gravity. It mainly controls the acceleration and deceleration of the fighter jet. N z represents the normal overload, which can be obtained by the ratio of the lift perpendicular to the plane of the fighter jet fuselage to the gravity. It mainly controls the pitch of the fighter jet.

[0072] (2) Maneuver design

[0073] In order to cope with the complex environment during air combat, the fighter jet needs to make corresponding air combat maneuvers according to different air combat situations, maintain the pursuit air combat maneuvers in the dominant air combat situation to expand the advantageous position, change the maneuver in the non-dominant air combat situation to weaken the disadvantage, and even reverse the air combat situation from defense to offense. Therefore, the reasonable use of maneuvers in air combat confrontation will bring great advantages to the air combat process.

[0074] Traditional basic air combat maneuvering actions include post-roll, high-power loop, low-power loop, loop, roll shear maneuver, etc., which can be selected by fighter pilots according to the current air combat situation. However, these maneuvering actions are based on human experience and have certain limitations theoretically. When using agent decision-making, more basic maneuvering actions should be selected, and the agent can learn to obtain advanced maneuvering actions. The following experiments prove that the agent can learn more advanced maneuvering actions during the learning process. This paper selects the set of basic fighter maneuvering actions (Basic Fight Maneuvering, BFM) proposed by NASA, which includes steady state, deceleration, acceleration, left turn, right turn, pull-up and dive. During air combat, the agent can form more advanced maneuvering actions by combining these seven basic maneuvers. For example Figure 1 are the schematic diagrams of the seven basic maneuvering actions.

[0075] Figure 1 In Figure 1, maneuver 1 is steady state, maneuver 2 is acceleration, maneuver 3 is deceleration, maneuver 4 is left turn, maneuver 5 is right turn, maneuver 6 is pull-up, and maneuver 7 is dive. Combining with the fighter mathematical model established in Subsection 3.3.1, the basic maneuvering actions can be mapped to the control inputs of the fighter. By maintaining the corresponding control inputs for a certain period of time, the corresponding basic maneuvering actions can be performed. For the control input triple encode, due to the physical characteristics limitations of the actual fighter, the available control inputs are limited. According to the characteristics of a certain type of fighter, here the tangential overload N x is set to vary in the range of [-2, 2], the normal overload N z is set to vary in the range of [-5, 5], and the roll angle is set to vary in the range of [-π / 3, π / 3]. Under the limited maneuvering actions, to maximize the maneuvering efficiency, each control input is discretized by selecting the boundary values and the intermediate value, that is, the tangential overload N x selects -2, 0, and 2, the normal overload N z selects -5, 1, and 5, and the roll angle selects -π / 3, 0, and π / 3. Table 1 shows the mapping relationship between the specific control input encoding and the basic maneuvering actions.

[0076] Table 1 Mapping relationship between control input encoding and basic maneuvering actions

[0077]

[0078] (3) Design of state transition module

[0079] Based on the control input encoding and the established fighter differential equation model (1), the state update equation of the simulation platform can be calculated by Equation (2).

[0080] s t+1 = f(s t , u m1 , u m2 , …, u mi , …, u mN , T) (2)

[0081] In the above formula, f represents the state update function formed by establishing a differential equation system, u mi represents the basic maneuver of the i-th fighter plane, and T represents the simulation time step. Using a certain numerical differential calculation method, the states of all fighter planes at the next simulation moment can be obtained based on the states of all current fighter planes and the input of basic maneuvers. The specific state update schematic diagram is as shown in Figure 2 shown.

[0082] In actual simulation experiments, different numerical solution methods can be selected according to needs. In the present invention, the fourth-order Runge-Kutta method is used as the solution tool.

[0083] The implementation manner of Step 2 is as follows:

[0084] The LSTM neural network has the information perception ability in a partially observable environment, which meets the requirements of the situation complexity perception in a multi-aircraft air combat environment. Secondly, the fully connected neural network has good decision-making ability. Therefore, in the design of the action and evaluation networks, a form combining the LSTM neural network and the fully connected neural network is adopted.

[0085] (1) Action network design

[0086] The input of the LSTM neural network is the partially observable air combat situation, generally the state of one's own fighter planes and partial states of enemy fighter planes, and the states are directly input into the LSTM network. Since the output of the LSTM is the extracted global state information, on the one hand, the output is only feature information, and on the other hand, its dimension is not the dimension of the maneuver. Therefore, a fully connected layer is added after the LSTM layer to further process the extracted global state information and map it to the maneuver to achieve the purpose of decision-making. Therefore, the output layer of the fully connected neural network has 7 output layer neurons. The fully connected neural network uses the ReLu activation function, which is the same as in the fourth chapter. After the output layer of the fully connected neural network, the softmax function is added to normalize the output value into a probability form. The specific action network structure diagram is as shown in Figure 4 shown, where S t represents the input state, and a t represents the input maneuver.

[0087] (2) Evaluation network design

[0088] Similar to the action network, since the air combat situation is partially observable, an LSTM structure needs to be added to the network structure. After inputting the partially observable air combat situation, the belief state is obtained through reasoning. A fully connected neural network is added at the output of the LSTM to map the belief network to the value corresponding to the state at that position. The entire fully connected network corresponds to the state value function, which will not be elaborated here. The network structure is as Figure 5 shown.

[0089] In the two network design frameworks, it can be understood that the LSTM accepts the sequence of partial observations of the agent and memorizes and infers the belief state as the output. According to the introduction in the second chapter, the belief state is the state in the full-dimensional observation. Taking the belief state as the input of the subsequent fully connected neural network, and then obtaining the output through the fully connected neural network. Therefore, the above-mentioned fully connected network plays a role in fitting the mapping from the state to the maneuver action or from the state to the value.

[0090] The implementation method of step three is as follows:

[0091] After the strategies of all K fighter jets are upgraded, the dimension change of the payoff matrix is as shown in Equation (3).

[0092] n i1 ×n i2 ×…×n iK →n (i+1)1 ×n (i+1)2 ×…×n (i+1)K ,n (i+1)j =n ij +1 (3)

[0093] The above-mentioned n ij represents the strategy owned by the j-th fighter jet in the i-th round of iteration. Since a new optimal agent will be generated for each fighter jet in one round of agent optimization, so there is n (i+1)j =n ij +1. And with the addition of matrix dimensions, in order to ensure the integrity of the payoff matrix, the values of the corresponding internal elements also need to be set. The improvement of the payoff matrix is completed through air combat confrontation game simulation. Since as the training progresses, it becomes increasingly difficult to distinguish the winner and loser in the confrontation between agents, so here the air combat situation is used as the simulation result of the air combat confrontation game. Since a new strategy is added to the strategy set of each fighter jet, the missing items of the matrix are as shown in Equation (4).

[0094]

[0095] In the above represents the new strategy generated by the k-th fighter jet in the w-th round of iteration. As more and more strategies are added to the strategy set, the missing items will eventually increase exponentially. After understanding the above symbol rules here, a group of strategies are extracted to illustrate the principle of the obtained simulation results, as shown in Equation (5).

[0096]

[0097] In the above meta-game, the strategy s k is the agent decision-making model, eval_oto k (s k , s j ) regards the agent s k as its own side and s j as the enemy side for one-on-one air combat situation assessment. The cumulative reward obtained through multi-step simulation on the simulation platform mentioned in Chapter 3 is used as the missing term, and the reward function is modified here to evaluate the air combat situation between two fighter jets.

[0098] Here, the weighted sum of the air combat situation (6) and the win / loss judgment reward (7) is selected as the evaluation of the agent's air combat ability, and its calculation is shown in Equation (8).

[0099]

[0100]

[0101] R e = ω1 × R a + ω2 × R w , ω1 + ω2 = 1 (8)

[0102] In the above formula, since the win / loss judgment reward is the most direct means of evaluating the agent's air combat ability, the value of ω2 should be relatively large among the two weights ω1 and ω2. When the agent cannot obtain the win / loss judgment reward, it is necessary to refer to how much advantageous situation the agent has occupied to judge its air combat ability. Therefore, the air combat situation reward is an auxiliary reference. Combining the two pieces of information can judge the current agent's air combat ability, thereby filling the missing term in the payoff matrix.

[0103] The implementation method of Step 4 is as follows:

[0104] Using the empirical game framework, the α-rank algorithm is used in each step to obtain the strategy of the enemy aircraft under the approximate Nash equilibrium, avoiding solving the Nash equilibrium in a multi-fighter and non-zero-sum environment. The newly trained agent is the best response, which can accelerate the convergence speed. The α-rank method benefits from its unique and effective calculation mode, so it can also effectively solve the approximate Nash equilibrium in the non-zero-sum game of multiple fighter jets. Its generality makes it perform better in multi-agent direct learning. The α-rank probability distribution is obtained by constructing the response graph of the game, that is, each strategy combination s ∈ S is a node of the graph, and there is a directed edge between the node s ∈ S and the node σ ∈ S if the following two conditions are met:

[0105] (1) There is only one difference in the strategies of fighter jet k between the two strategy combinations s and σ;

[0106] (2) M k (σ) > M k (s).

[0107] α-rank constructs random movements in a directed graph. Usually, the transfer is carried out according to the direction of the edges, but there is also a small probability of reverse transfer. This perturbation probability is controlled by the α parameter, which is preset in the algorithm at the beginning and is a hyperparameter of the algorithm. This way of directed graph movement ensures the irreducibility of the Markov Chain and the unique stationary distribution π ∈ Δ S exists, and this distribution is also called the α-rank distribution. The convergence of π is guaranteed by the Sink Strongly-Connected Components (SSCCs) of the response graph.

[0108] The core of the α-rank algorithm is population evolution, where the interactions between individuals influence each other. Strong individuals will replace weak individuals. Due to this nature of survival of the fittest, some transient individuals in the population evolution process will be filtered out, and finally, a steady-state individual ranking can be obtained. In this chapter, the selection of the enemy strategy set can be solved based on the obtained steady-state ranking. Figure 3 is a conceptual graph of population interaction. As shown by the blue arrows, there is a very small probability that mutants may occur, and individuals with stronger fitness will spread and replicate throughout the population.

[0109] Figure 3 In, the key terms are explained as follows:

[0110] (1) Strategy: The parameters of the agent that need to be trained in an empirical game;

[0111] (2) Individual: A member of a population, responsible for arranging strategies into specific slots of an empirical game and executing them;

[0112] (3) Population: A finite set of individuals;

[0113] (4) Homogeneous population: A population where all individuals are the same;

[0114] (5) Set of homogeneous populations: A finite set composed of homogeneous populations;

[0115] (6) Focus population: The population where a small probability mutation is currently occurring.

[0116] The advantage of the method in this section is that it can be used in the non-zero-sum game of multiple fighter jets, where the number of strategies can be extended to more than 4. The overall idea is to obtain the dynamic information of the whole system through the Markov chain established by the states of the isomorphic population set. By calculation, the transition probability between states can be obtained, and then the transition matrix can be obtained. This matrix represents the mutation direction and probability size of individuals in the population. In addition to the representation method of the transition matrix, the graph method can also be used, where the edges represent the transition probability between states and the nodes represent the states. The stationary distribution can be obtained from the transition matrix, which quantifies the time spent by each node in the state transition process.

[0117] Next, consider the interaction of K isomorphic populations. During the evolution process, it is necessary to evaluate the ability of individuals executing strategies, that is, the evolutionary intensity. Suppose there are m individuals in each population, and the individuals in population k will all execute the strategies in set S k in.

[0118] Figure 3 In, in each sampling period T, individuals are uniformly sampled in each population to obtain the interactive game of K individuals. The number of individuals executing strategy s k in population k is From the sampling method, the fitness calculation of the individuals executing strategy s k is as shown in Equation (9).

[0119]

[0120] To facilitate the description of the mutation and replication of individuals in the population, assume that there are two individuals in the population executing strategies τ and σ respectively, and their corresponding fitnesses are f k (τ, p -k ) and f k (σ, p -k ). Here, a discrete dynamics model is introduced. An individual in the population has a very small probability of mutating into an individual with a random strategy τ and a certain probability of replicating strategy σ, or persists in executing the mutated strategy. The above process reflects the replication and spread of individuals in the population. Through the following calculations, it can be obtained that only powerful individuals have such abilities.

[0121] Individuals in a population do not interact directly with each other, so the state of population k is independent of the fitness of individuals. However, in (9), the fitness of each population may be directly affected by other populations. The use of a mutation mechanism with an extremely low probability can significantly reduce the complexity analysis of the system. The present invention defines the probability of a strategy randomly mutating to another strategy as μ, and assumes that this probability is extremely low (for example: it can be assumed that this probability is close to 0). If the occurrence of mutations is considered negligible, then during the evolution of the population, it can be considered that there are always monomorphic individuals. On the contrary, if the probability of mutation is considered to be a very small probability but not ignored, then it means that mutations may occur. If the mutant individuals are relatively strong, they may be fixed and replicated and spread throughout the population. On the contrary, if their ability is relatively weak, they will be eliminated during the evolution of the population, and finally the population will return to the isomorphic state again. Since the mutation probability is very small, the current population will become an isomorphic population before new mutant individuals arrive. This means that any given population k will not contain more than two strategies at any evolutionary time point.

[0122] Using the same analysis method, there is also the above rule between different populations, that is, before mutant individuals appear in the next population, the current population will first return to the isomorphic state. Therefore, at any given evolutionary time point, mutant individuals will not appear in two populations simultaneously. So in any population c ∈ {1,..., K}\k, one of the strategies and the rest of the strategies are all 0. Given a sufficiently small mutation probability, the analysis on the focal population k only needs to be considered when other populations are in the isomorphic state. Equation (9) can be further simplified to Equation (10).

[0123] f k (s k ,s -k ) = M k (s k ,s -k ) (10)

[0124] where s -k represents the strategy combination of other populations. Let and represent the individuals executing strategies τ and σ in the focal population k respectively, and the relationship between the two is as shown in Equation (11). According to Equation (10), the fitness calculation methods of the two types of individuals can be further simplified to Equations (12) and (13).

[0125]

[0126] f k (τ,s -k ) = M k (τ,s -k ) (12)

[0127] f k (σ, s -k ) = M k (σ, s -k ) (13)

[0128] Randomly sample two individuals in population k. Let P(τ → σ, s -k ) be the logical selection function for one execution strategy τ to replicate another execution strategy σ, also known as the Fermi distribution, which controls the dynamic model in a finite population set, and the calculation method is shown in Equation (14).

[0129]

[0130] Among them, α is the ranking intensity, which controls the intensity of selection.

[0131] Based on the above premise, define a Markov Chain on the set of strategy combinations Π k S k . Here, there are a total of Π k |S k | states on the Markov chain. Any state s ∈ Π k S k represents the termination state of the isomorphic population set. The transition probability between these states is defined as the fixed probability when the mutant strategy is introduced into the focal population. Combining the number of states, there are a total of (Π k |S k |) 2 transition probabilities. Define as the probability that the mutant strategy τ is fixed in the focal population k with the individual's execution strategy being σ, while the remaining K - 1 populations remain in the isomorphic state s -k unchanged. Since the spread of mutant individuals only occurs in the current focal population, given a strategy combination, there are a total of ∑ k (|S k | - 1) next strategy combinations to be transferred. Therefore, let η = 1 / (∑ k (|S k | - 1)), is the probability that the combined population state transfers from (σ, s -k ) to (τ, s -k ) when the focal population k mutates. The steady-state distribution of the Markov chain is the average time spent at each state.

[0132] The fixed probability of the replication and spread of mutant individuals with execution strategy τ to the focal population k can be calculated by Equation (23), and the probability of the decrease / increase of one individual with execution strategy τ in the population is shown in Equation (15).

[0133]

[0134] Let \(u = f\) k (\(\tau, s\) -k ) - f k (\(\sigma, s\) -k ), among the \(m - 1\) individuals implementing the strategy \(\sigma\), where the fixed probability of a single mutant individual implementing the strategy \(\tau\) is specifically derived as shown in Eqs. (16) to (23).

[0135]

[0136]

[0137] The above results correspond to the \(m -\)step transition based on the logical selection function (14) on the Markov chain. The quotient \(T\) k(-1) (p k , \(\tau, \sigma, s\) -k ) / \(T\) k(+1) (p k , \(\tau, \sigma, s\) -k ) expresses the likelihood of the direction of the mutation process in the population. If the quotient is close to 0, then the likelihood of an increase in mutant individuals is high; if the quotient is large, then the likelihood of a decrease in mutant individuals is low; if the quotient is close to 1, then the likelihood of an increase and a decrease in mutant individuals is equal. The Markov transition probability matrix can be obtained through formula (23), from any strategy combination \(s\) i \(\in \Pi\) k \(S\) k to any strategy combination \(s\) j \(\in \Pi\) k \(S\) k The matrix elements are as in Eq. (24).

[0138]

[0139] The above \(i, j\in\{1, \ldots, \Pi\) k |\(S\) k \}.

[0140] According to the Markov transition probability matrix, its stationary distribution can be calculated. The following gives the stationary distribution theorem of the Markov chain.

[0141] Let the Markov transition probability matrix be \(C\), and any two states are connected, exists and is independent of \(i\), denoted as Then there are:

[0142] 1:

[0143] 2:

[0144] 3: π is the only non - negative solution of the equation πP = π;

[0145] where π = [π(1), π(2), …, π(j), …],

[0146] π is called the stationary distribution of the Markov chain. In the application of the present invention, the state set is finite. In the above definition, the summation sign is changed to a finite summation. After obtaining the stationary distribution, since each term in the stationary distribution corresponds to a state node, that is, a strategy combination, the proportion of each strategy combination can be obtained. where represents the i - th strategy in the strategy set of the k - th fighter plane. The strategy combinations can be sorted in descending order of magnitude to obtain rank, but the direct sorting of strategy combinations cannot be directly used and needs to be further transformed into the mixed strategy of each fighter plane. Therefore, it is necessary to further calculate the marginal probability of the stationary distribution to finally obtain the mixed strategy of each fighter plane. k The implementation manner of step five is as follows:

[0147] The action and evaluation networks need to optimize the network parameters according to the experience samples of the interaction between the agent and the environment. Since multiple rounds of iteration are required to optimize the agent in this chapter and the stability requirements for the reinforcement learning algorithm are relatively high, the PPO algorithm is used as the basic reinforcement learning algorithm. The PPO algorithm avoids large policy updates to improve the stability of training. From the perspective of deep learning, there are mainly two reasons:

[0148] (1) In training, smaller iterations of policy parameters are more likely to converge to the optimal solution;

[0149] (2) A large policy update during training will lead to "off the cliff", that is, the policy instantaneously obtains poor parameters, usually taking a long time to recover or not being able to recover to the original better parameters at all.

[0150] (2) A large policy update during training will lead to "off the cliff", that is, the policy instantaneously obtains poor parameters, usually taking a long time to recover or not being able to recover to the original better parameters at all.

[0151] The reinforcement learning objective function with the advantage function is shown in Equation (25).

[0152] L PG (θ) = E t [logπ θ (a t |s t ) * A t (25)

[0153] where A t is the advantage function, A t = Q(s t , a t) - V(s t ) Introducing the advantage function can reduce the variance during training, mitigate the training fluctuations caused by excessive variance, and thus weaken the overfitting problem. A value greater than 0 indicates that in the current state, action a t is better than other alternative actions, logπ(a t |s t ) is the logarithm of the probability of taking action a t in state s t . The model adopts the gradient descent strategy, changing the policy model parameters in the gradient direction to enable the agent to take actions with higher rewards in future decisions. In the parameter update of gradient descent, if the update amplitude of each step of the policy parameter is too small, the training time will be very long, and if the amplitude is too large, it will lead to large training fluctuations and even cause training instability.

[0154] The policy update in the PPO algorithm is relatively conservative. First, it is necessary to calculate the change magnitude of the current policy relative to the past policy, that is, define a ratio between the new policy and the old policy, and achieve stable policy update by restricting the ratio to the interval [1 - ε, 1 + ε], ultimately achieving the goal of stable change of the new and old policy parameters. In the algorithm implementation, it is necessary to use the clipping technique to constrain the policy parameter update of each step, and adopt the probability ratio constraint such as Equation (3) to update the front and back policies.

[0155]

[0156] Form a new objective function with policy update constraints, and calculate it as shown in Equation (27).

[0157]

[0158] Among them represents the expected value, min(a, b) represents choosing a minimum value between parameters a and b, represents the estimation of the advantage function, and the estimated value of the advantage function can be obtained using methods such as Monte Carlo. r t (θ) is the ratio function, indicating that in the current state s t and the current policy parameter θ, the probability of executing action a t is the probability ratio of executing action a t in the current state s old and the previous policy parameter θ t , reflecting the difference magnitude between the two policies. When r t (θ) > 1, it indicates that in the current policy, the probability of executing action a t in the current state s t is greater than the possibility of the previous policy. When 0 < r t (θ) < 1, it indicates that in the current policy, in the current state st The probability of performing action a t is smaller than that of the previous policy. Therefore, the probability ratio can estimate the difference between the new and old policies.

[0159] The min function contains two parts. is the unclipped part. This term only works when the changes between the new and old policies are small. When the two policies change significantly, it will lead to a large policy gradient update. The clipping function clip(·) is added to the objective function to constrain the probability ratio away from 1. This constraint method is simpler and more effective than using the KL divergence constraint outside the objective function in the TRPO (Trust Region Policy Optimization) method to limit the change of policy parameters.

[0160] clip(r t (θ), 1 - ε, 1 + ε) represents the restricted ratio function. The value range of r t (θ) is in (1 - ε, 1 + ε), that is, when r t (θ) is within the range of (1 - ε, 1 + ε), the clipping function outputs the ratio function itself. When r t (θ) ≤ 1 - ε, it outputs 1 - ε. When r t (θ) ≥ 1 + ε, it outputs 1 + ε. Figure 6 is the schematic diagram of the clip clipping function.

[0161] Combined with the full formula, as Figure 7 is the schematic diagram of the min function, which is discussed in two cases of the advantage function respectively. The green dashed line represents the first term of the min function without the clip clipping term, and the blue dashed line represents the second term of the min function with the clip clipping term. The resulting red part is the output of the final min function.

[0162] With the above constraints, it can be ensured that the amplitude of each step of policy update will not be too large. Here, the Actor-Critic structure is adopted, and the objective function of the evaluation network adopts the form of mean squared error, as shown in Equation (28).

[0163]

[0164] During gradient update, the objective functions of the action and evaluation networks can be integrated and trained together, and an entropy term is added to the objective function to improve the exploration ability of the agent. Equation (29) is the expression of the overall objective function.

[0165]

[0166] In the above formula, c1 and c2 are constant coefficients. is the mean squared error between the evaluation network and its target network, and S[π θ (s t ) represents the entropy of the current policy, and its calculation method is shown in Equation (30).

[0167]

[0168] The larger the entropy value of the policy, the more evenly the probabilities of selecting each action are, and the stronger the exploration ability of the policy. Generally, its coefficient c2 is set to 0.01. The parameters of the policy model need to be changed in the direction of the maximum value of the objective function, as shown in Equation (31).

[0169]

[0170] The training process of the entire algorithm is as Figure 8 shown.

[0171] The implementation method of Step 6 is as follows:

[0172] When the training algorithm has converged, select the latest agents of each game party as decision-makers, extract their agent action networks as decision networks, input the air combat situation state vector, and after passing through the network, the maneuver actions can be output to complete one step of decision-making.

[0173]

Experimental Verification

[0174] Set some observation conditions, that is, each agent can only observe all the states of its own fighter plane, but can only observe a part of the states of other fighter planes. To be more in line with the actual situation, set that only the spatial position information of the enemy plane can be observed. The decision network used in this chapter will automatically deduce the hidden state by receiving the time series information. Assume that there are a total of N fighter planes, and the fighter plane where the current agent is located is the i-th one. Then the state quantity that the current agent can receive is shown in Equation (32).

[0175]

[0176] Among them, s ip represents the partial observation quantity of the i-th agent, s i represents the own state quantity of the i-th fighter plane, s j represents the partial observation state quantity of the fighter planes except the j-th fighter plane. In this experiment, specifically set 4 fighter planes to participate in the battle, including 2 enemy fighter planes and 2 friendly fighter planes. The initial states of each fighter plane are shown in Table 1, where Fighter Plane No. 1 and Fighter Plane No. 2 are friendly fighter planes, and Fighter Plane No. 3 and Fighter Plane No. 4 are enemy fighter planes.

[0177] Table 1 Initial States of Both Sides in Flight

[0178]

[0179] In the simulation platform, the observed quantity is set to full-dimensional observation, that is, the actions input to the agent and the evaluation network during simulation are 12-dimensional air combat situation vectors. The simulation platform uses the fourth-order Runge-Kutta method to solve the differential equation system, and the gravitational acceleration is 10m / s 2 , the solution step size is 0.01s, the decision-making cycle of the fighter is set to 0.1s, the condition for the fighter to be judged crashed is that the altitude is less than 10 or it is hit 20 decision-making cycles cumulatively, and the maximum engagement time for one air combat is set to 100s, that is, if the red and blue sides do not determine the winner within 1000 simulation cycles, this round will be forced to end and the state will be forced to be initialized.

[0180] PPO algorithm parameters: the discount factor γ = 0.995, the learning rates of the action network and the evaluation neural network are both 0.0003, the clip parameter is 0.2, and the total training is 1×10 7 minimum step sizes, and the constant coefficients c1 and c2 in the loss function are 0.5 and 0.01 respectively.

[0181] The weights for calculating the missing terms in the payoff matrix are set to ω1 = 0.2 and ω2 = 0.8, the number of repeated simulations N for the missing terms is 500 times, and the α-rank hyperparameter α for the payoff matrix solution method is 1×10 -4 .

[0182] Due to the relatively low configuration of the simulation environment, in order to minimize the computational load as much as possible, a shared policy mechanism is adopted here, that is, the fighters on the same side share a set of policies. Therefore, 2 neural networks participate in the training during the training iteration process, and the maximum number of iterations is 15. In the action and evaluation networks, the number of neurons is uniformly set to 64, and the hidden unit dimension is 20.

[0183] Figure 9 To facilitate observing the change trend of the reward value, the average reward value of all training agents in each round is taken and plotted. (a) - (d) are the graphs of the average reward value obtained from the 1st, 5th, 10th, and 15th rounds of reinforcement learning training during the training iteration process versus the step size. The dark red line represents the value after taking the average of the reward values, and the light red line represents the true value of the reward. It can be found that in the partially observable air combat environment, the LSTM neural network structure can use the partial observation information received by the agent to infer the complete air combat situation, and input the situation into the fully connected neural network to output maneuvering actions. The value of the reward function increases steadily in each iteration, indicating that the agent can also learn a highly robust air combat strategy using the curriculum reward structure in the partially observable environment.

[0184] To further illustrate the effectiveness of the α-rank algorithm used in this chapter in the multi-aircraft air combat decision-making game, with other experimental conditions remaining unchanged, neural network decision-making mechanisms are set for all 4 fighter jets without sharing parameters with each other. It is compared with the replicate dynamic (RD) method. RD is proposed based on the biological evolution mechanism. The core idea is that strategies superior to the average level will gradually be adopted by individual members of the biological group (game participants), while strategies inferior to the average level will gradually be abandoned by individual members of the biological group (game participants). Its evolution mechanism is shown in Equation (31), which defines a dynamic system in the strategy distribution.

[0185]

[0186] By calculating the availability and plotting Figure 10 the availability comparison curve, it can be found that in the multi-player game, the α-rank algorithm converges faster than the RD method and converges faster towards the Nash equilibrium.

[0187] From Figure 11 it can be seen that using the LSTM neural network structure in a partially observable environment can effectively improve the agent's ability to perceive the environment, infer the current environmental belief state from the historical sequential partially observable information, and thus obtain a higher reward value under the same learning task. It can also be seen from the figure that although the fully connected neural network cannot infer the environmental belief state, it has a powerful decision mapping ability and still obtains a certain reward value during training and occupies a certain air combat advantage relying on the partially observable state information. Therefore, this illustrates the advantage of combining the LSTM neural network and the fully connected neural network structure in a partially observable environment. On the one hand, the LSTM neural network structure extracts the environmental belief state from the partially observable information, and on the other hand, the fully connected neural network structure maps to the optimal maneuvering action according to the perceived environmental belief state.

[0188] Figure 12 it can be seen that the multi-aircraft dogfighting characteristics shown among the fighter jets. In single-aircraft air combat, the game relationship between fighter jets is competition, mainly manifested as the interest conflict between fighter jets. In multi-aircraft air combat, since both the friendly and enemy fighter jet groups contain more than 1 fighter jet, the game relationship between fighter jets is competition and cooperation. At this time, the complexity of the game scenario increases, and the multi-aircraft training difficulty is more complex than that of single-aircraft air combat. After multiple decision steps, the enemy's No. 3 fighter jet is shot down first, followed by the enemy's No. 4 fighter jet, and finally the friendly side wins.

[0189] From the side view Figure 13It can be seen that all fighter jets choose to climb rapidly under the same speed and altitude advantages to improve their potential energy advantages. After a certain opportunity, they dive to convert the potential energy advantage into kinetic energy advantage so as to quickly reach the advantageous attack position. It can be found that when the fighter jets cannot obtain an advantageous attack or defensive position during the dive, the fighter jets will climb again, converting the internal energy and kinetic energy of the fighter jets into potential energy, thereby improving the potential energy advantage and preparing for the next round of seizing the advantageous attack position.

[0190] From the top view Figure 14 it can be seen more clearly the tactical intentions among the fighter jets. It can be seen in the figure the start of a multi-aircraft air combat dogfight. Two friendly fighter jets have clear targets, namely, friendly fighter jet No. 1 is engaged in a dogfight with enemy fighter jet No. 3, and friendly fighter jet No. 2 is engaged in a dogfight with enemy fighter jet No. 4. During the dogfight, the friendly aircraft group forces the enemy aircraft group to stay in the middle area through maneuvers. When friendly fighter jet No. 2 moves to the position of 2000 m on the X-axis, friendly fighter jet No. 2 gives up the dogfight with enemy fighter jet No. 4, and its strategy changes to cooperate with friendly fighter jet No. 1 to shoot down enemy fighter jet No. 3 together. After enemy fighter jet No. 3 is shot down, friendly fighter jet No. 2 and No. 1 together encircle enemy fighter jet No. 4 and continuously attack it until it is shot down. During this process, it can be found that in the cooperative air combat of the friendly fighter jet group, first, the targets are allocated one by one. In the one-on-one air combat, the positions of the enemy aircraft are suppressed. When some enemy aircraft fly into the common area of the friendly fighter jet group, the friendly fighter jet group cooperates to shoot down one of the enemy aircraft first, and then cooperates to shoot down the remaining enemy aircraft to achieve the breakthrough of the targets one by one.

[0191] It can be found from the plane trajectory diagrams of the side view and the top view that the intelligent agents have learned the multi-aircraft air combat strategy and have the ability of autonomous decision-making.

[0192] To sum up, through multiple rounds of iterative learning, the friendly intelligent agents have the highly autonomous cooperative decision-making ability when facing complex air combat situations. Through the combination of maneuver commands among the fighter jets, complex maneuver actions and cooperative tactics can be formed. Compared with the intelligent agents trained by general multi-intelligent agent reinforcement learning algorithms, the decision-making among multiple aircraft is more cooperative, which also shows the effectiveness of the method of the present invention.

Claims

1. A multi-aircraft air combat decision-making method based on general experience game reinforcement learning, characterized in that The steps are as follows: Step 1: The fighter jets in the simulation platform are dynamic models in three-dimensional space, and the control input for each fighter jet is the roll angle and two control inputs N x , N z , and they are encoded; N x represents tangential overload, and N z represents normal overload; The maneuvering actions include seven types: steady, decelerating, accelerating, turning left, turning right, pulling up, and diving; The state update equation of the simulation platform: s t+1 = f(s t , u m1 , u m2 , …, u mi , …, u mN , T) Let \(f\) denote the state update function formed by establishing a system of differential equations, and \(u\) mi represents the basic maneuver of the \(i\)-th fighter plane, and \(T\) represents the simulation time step; Step 2: The agent of each fighter plane includes an action neural network and an evaluation neural network. The neural network structure is an LSTM neural network connected to a fully connected neural network; The air combat situation, that is, the sequence of observables, is the input of the LSTM neural network, and the global state information is extracted as the output; This output is the input of the fully connected neural network. In the action network, the softmax function is added to the output layer of the fully connected neural network to normalize the output value into a probability form. The output is the 7 kinds of maneuvering actions in Step 1, and the output of the evaluation network is the value of the value function; The fully connected neural network uses the ReLu activation function; Step 3: Each fighter plane has a policy set, and multiple groups of action neural networks and evaluation neural networks are set in each policy set; the input of each group of neural networks is the air combat situation sequence; Add a group of action and evaluation neural networks with random parameters to the policy set of each fighter plane, and initialize the policy sets of K fighter planes; Step 4: For the policy sets of K fighter planes, the dimension of the payoff matrix changes with the number of iterations as follows: n i1 × n i2 × … × n iK → n (i+1)1 × n (i+1)2 × … × n (i+1)K , n (i+1)j = n ij + 1 Where: n ij represents the strategy owned by the j-th fighter in the i-th iteration. Since each fighter will generate a new optimal agent in one round of agent optimization, so there is n (i+1)j = n ij + 1; When adding a new policy to the policy set of each fighter plane, the missing items in the payoff matrix are: Denote the new strategy generated by the k-th fighter plane in the w-th iteration; The value of the missing item in the payoff matrix is obtained through the following formula for air combat simulation; The strategy s in the meta-game k is the agent decision-making model, eval_oto k (s k , s j ) is the one-on-one air combat situation assessment that regards agent s k as its own side and s j as the enemy side; Select the air combat situation in air combat situation assessment: And the win-loss judgment reward: The weighted sum of is used as the air combat ability assessment of the agent: R e = ω1 × R a + ω2 × R w , ω1 + ω2 = 1; Step 5: Use the α-rank algorithm to obtain an approximate Nash equilibrium under the current payoff matrix, that is, each fighter plane selects a group of action and evaluation neural networks from its policy set, and the selected multiple groups of neural networks form a policy combination; Under the policy combination obtained in Step 5, except for the k-th fighter plane, the other K-1 fighter planes use their corresponding selected action networks as decision-making agencies to control their own maneuvers, and use the PPO algorithm as the basic reinforcement learning algorithm to optimize the new policy of the current k-th fighter plane, generating new action and evaluation networks and adding them to the policy set of the k-th fighter plane; optimize each fighter plane in turn, and at the same time add the generated new agents to the corresponding policy sets, and return to Step 4 until the set number of training iterations is satisfied; Step 7: Select the latest agents of each game party as decision-makers, extract their agent action neural networks as decision-making networks, and the state vector of the air combat situation of the observables is the input of the LSTM neural network. After passing through the LSTM neural network and the global neural network, the maneuvering action is output to complete one-step decision-making.

2. The multi-aircraft air combat decision-making method based on general experience game reinforcement learning according to claim 1, wherein: The tangential overload N x varies in the range of [-2, 2], and the normal overload N z varies in the range of [-5, 5], and the roll angle varies in the range of [-π / 3, π / 3].

3. The multi-aircraft air combat decision-making method based on general experience game reinforcement learning according to claim 1, characterized in that: The fighter plane dynamics model in the three-dimensional ground coordinate system: The left sides of each differential equation are the first-order derivatives of the fighter plane's flight speed, yaw angle, pitch angle, and three-dimensional position respectively.

Citation Information

Patent Citations

  • A dynamic game method for multi-unmanned aerial vehicle air battle under uncertain information

    CN107463094A

  • Air combat maneuvering strategy generation technology based on deep random game

    CN112052511A