Air Combat Decision-Making Method Based on Attention Network Transfer for Multi-Agent Reinforcement Learning
Through the attention network migration method trained in small-scale air combat scenarios, the problem of low training efficiency in large-scale air combat scenarios is solved, the model is rapidly converged and the agent combat capability is improved, and the Actor-Critic architecture and Transformer module are used to process global observation information, which improves the decision-making efficiency and tactical strategies of multi-agent systems.
Patent Information
- Application Number
- CN202410888880.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-07-04
AI Technical Summary
The existing multi-agent reinforced learning air combat decision-making methods are inefficient in training in large-scale air combat scenarios, the model convergence speed is slow, and the agent's combat capability is insufficient.
The multi-agent reinforcement learning method based on attention network migration is adopted. The agent decision network trained in small-scale air combat scenarios is reused and reconstructed, and it is loaded into large-scale air combat scenarios for training. It combines the Actor-Critic architecture and the Transformer module to process global observation information, and uses the MAPPO algorithm for training.
It significantly improves the convergence speed and robustness of the model in large-scale air combat scenarios, and improves the combat capability and battlefield situation analysis capabilities of the agent.
Smart Images

Figure CN118917171B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer force generation and artificial intelligence, and specifically relates to a multi-agent reinforcement learning air combat decision-making method and its application based on attention network transfer constructed by the K multi-agent reinforcement learning method. Background Art
[0002] Tactical-level multi-UAV air combat missions have become a research hotspot in the current domestic and foreign military fields due to their potential application value.
[0003] Reinforcement-deep learning technology provides an effective way to achieve the autonomous mission capabilities of UAV swarms. A series of multi-Agent training frameworks based on reinforcement-deep learning have been invented, such as QMIX, Q-Trans, MAPPO, etc. These learning frameworks are oriented towards multi-Agent systems and have made a lot of innovations in basic principles or structural designs to improve the learning efficiency and quality of multi-Agents collaborative strategies. The QMIX network simplifies the value distribution problem between Agent individuals and collectives based on the value decomposition principle and simplifies the learning structure. The Q-Trans network introduces a transformation function to express how an individual Agent adjusts its strategy according to the behaviors of other Agents, thereby capturing the complex interactions between Agents. MAPPO uses the advantage function to estimate the expected return of each Agent's action relative to other possible actions, promoting the Agent to learn better independent strategies, and at the same time ensuring the consistency of the entire Agent system strategy through a coordination mechanism to achieve the balance between independence and coordination. In addition, the Transformer network has repeatedly proven its unique advantages in expressing high-dimensional situations in various applications, so it has also been used in the multi-aircraft air combat decision-making problem and achieved good results. Summary of the Invention
[0004] In order to solve the defects existing in the prior art, the present invention discloses a multi-agent reinforcement learning air combat decision-making system based on attention network transfer, and its technical solution is as follows:
[0005] It includes the following modules: a multi-aircraft short-range air combat simulation platform, a multi-agent reinforcement learning network design and small-scale multi-aircraft air combat scenario training module, large-scale multi-aircraft air combat scenario training based on attention network transfer, and a battlefield attention information analysis module, and is characterized in that:
[0006] Multi-aircraft short-range air combat simulation platform: Construct a multi-aircraft air combat scenario where aircraft obtain the firing advantage over targets through tactical maneuvers, with short-range air-to-air missiles and machine guns as the main weapons; The two sides of the air combat are defined as the red and blue sides; The red aircraft are our side, and the tactical strategies are obtained through learning; The blue aircraft are used as opponents and adopt tactics with different difficulty levels based on random action sampling, air combat scripts, and virtual self-play; Model the above scenario to define aircraft characteristics; At the same time, according to the theory of partially observable Markov games, calculate the state space, action space, and reward function through the position, speed, and weapon information of the aircraft.
[0007] Multi-agent reinforcement learning network design and small-scale (2.vs.2) multi-aircraft air combat scenario training module: Design the Actor-Critic network and reinforcement learning algorithm. For the multi-agent reinforcement learning air combat decision-making method based on the migration of the attention network, its characteristics are as follows: The step 2 includes the following content: The network adopted is the Actor-Critic architecture. For the Actor and Critic networks, the input information includes: its own information O t,i , friendly force information O t,f , enemy force information O t,j and global information O t,full . After passing O t,i , O t,f , O t,j through a fully connected network, a high-dimensional tensor is obtained. After the global information O t,full undergoes an embedding operation, it is input into the LSTM network for temporal processing, and the obtained tensor is connected to the Transformer module and concatenated with the high-dimensional tensor as the decision-making information for the Actor and Critic networks. Set up a 2.vs.2 air combat environment in the environment and adopt the self-play method to learn the red side's tactical strategies; The geographical scene size is 50x50 kilometers; A total of 10,000 episodes are set in the training process, and each episode is an air combat process with a maximum of 2,000 decision steps; The hyperparameters are set as follows: The learning rates of the Actor and Critic networks are lr = 0.0001, the discount factor γ = 0.95, the MAPPO hyperparameter clipping coefficient ∈ = 0.2, the optimizer is Adam, and the batch size = 2,000; The strategies adopted by the blue aircraft are divided into four levels (L1-L4), and the higher the level, the more complex and diverse the behavior of the blue side; The stronger the ability of the trained red agent to adapt to complex environments; The trained agent strategies π i , π j ; The specific classification of the blue side's strategy is as follows:
[0008] ·L1: static, the blue aircraft are in a stationary state;
[0009] L2: random, the blue plane takes random actions within the allowed action space;
[0010] L3: script, the blue aircraft engages the nearest red aircraft and after engaging, moves away from the red aircraft;
[0011] L4: Self-play, the Red aircraft uses the Blue strategy trained in L3 to fight against the opponent.
[0012] Large-scale (5.vs.5) multi-aircraft air combat scenario training and battlefield attention information analysis module based on attention network migration: The model trained in small-scale scenarios is reused and put into large-scale multi-level air combat scenarios for training. The output information from the attention network of the two scenarios is analyzed to obtain the battlefield attention level of each aircraft to the enemy and friendly forces.
[0013] The present invention discloses a multi-agent reinforcement learning air combat decision-making method based on attention network migration, and the method is based on the above system.
[0014] The present invention also discloses a non-volatile storage medium, characterized in that the non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above method.
[0015] The present invention also discloses an electronic device, characterized in that it comprises a processor and a memory; the memory stores computer-readable instructions, and the processor is used to execute the computer-readable instructions, wherein the computer-readable instructions execute the above method when they are executed.
[0016] Beneficial Effects
[0017] 1) Establish a multi-aircraft near-visual-range air combat simulation platform: --- Use the two-dimensional fixed-wing aircraft dynamics equation to model the aircraft. Based on Markov game theory, set the state space, action space and reward function of the dynamic characteristics and weapon carrying characteristics for the multi-aircraft air combat decision-making task.
[0018] 2) A multi-agent reinforcement learning algorithm framework based on attention network transfer is proposed. This method reuses and reconstructs the agent decision network that has been trained and converged in the 2.vs.2 air combat scenario, and loads it into the complex 5.vs.5 air combat scenario for training. It is proved that this method can improve the model convergence speed and robustness of reinforcement learning training, as well as the combat capability of the agent.
[0019] 3) A battlefield situation analysis method based on an attention network is proposed. By processing the global battlefield attention observation information output by the trained attention network, the battlefield attention of the agent to other friendly and enemy forces in the battlefield can be obtained, and then the battlefield situation can be analyzed. Description of the Drawings
[0020] Figure 1 It is a schematic diagram showing the description of each functional component of the present invention and their interrelationships;
[0021] Figure 2 It is a schematic diagram showing the description of the characteristics of two-dimensional aircraft close-range air combat;
[0022] Figure 3 It is a schematic diagram showing the description of a multi-aircraft air combat simulation platform;
[0023] Figure 4 It is a schematic diagram showing the description of a decision-making network;
[0024] Figure 5 It is a schematic diagram showing the description of the reward function curve of small-scale air combat training;
[0025] Figure 6 It is a schematic diagram showing the description of the test win rate graph of small-scale air combat training;
[0026] Figure 7 It is a schematic diagram showing the description of a small-scale air combat trajectory graph;
[0027] Figure 8 It is a schematic diagram showing the description of the battlefield attention heat map at a certain moment in air combat;
[0028] Figure 9 It is a schematic diagram showing the description of the training curve of a large-scale multi-aircraft air combat scenario for attention network migration;
[0029] Figure 10 It is a schematic diagram showing the description of the test win rate graph of a large-scale multi-aircraft air combat scenario for attention network migration;
[0030] Figure 11 It is a schematic diagram showing the description of a large-scale multi-aircraft scenario trajectory graph;
[0031] Figure 12 It is a schematic diagram showing the description of the battlefield attention heat map of a large-scale multi-aircraft scenario trajectory graph.
[0032] Figure 13 Schematic diagram of the method flow of the present invention. Detailed Implementation Manner
[0033] Example 1
[0034] Multi-agent Reinforcement Learning Air Combat Decision-making System Based on Attention Network Migration, including the following modules: Multi-aircraft Short-range Air Combat Simulation Platform, Multi-agent Reinforcement Learning Network Design and Small-scale Multi-aircraft Air Combat Scenario Training Module, Large-scale Multi-aircraft Air Combat Scenario Training Based on Attention Network Migration, and Battlefield Attention Information Analysis Module.
[0035] See Figure 1 as shown. Multi-aircraft Short-range Air Combat Simulation Platform: Construct a multi-aircraft air combat scenario. The aircraft obtain the firing advantage over the target through tactical maneuvers, and use short-range air-to-air missiles and machine guns as the main weapons. The two sides of the air combat are defined as the red side and the blue side. The red-side aircraft are our side, and the tactical strategy is obtained through learning. The blue-side aircraft are used as opponents and adopt tactics with different difficulty levels based on random action sampling, air combat scripts, and virtual self-play. Model the above scenario to define the aircraft characteristics. At the same time, according to the theory of partially observable Markov games, calculate the state space, action space, and reward function through the position, speed, and weapon information of the aircraft.
[0036] Multi-agent Reinforcement Learning Network Design and Small-scale Multi-aircraft Air Combat Scenario Training Module: Design the Actor-Critic network and reinforcement learning algorithm. For the multi-agent reinforcement learning air combat decision-making method based on attention network migration, the network adopted is the Actor-Critic architecture. For the Actor and Critic networks, the input information includes: its own information O t,i , friendly force information O t,f , enemy force information O t,j and global information O t,full . After passing O t,i , O t,f , O t,j through a fully connected network, a high-dimensional tensor is obtained. After embedding the global information O t,full , it is input into the LSTM network for temporal processing. The obtained tensor is connected to the Transformer module and concatenated with the high-dimensional tensor as the decision-making information for the Actor and Critic networks. Set up a 2 vs. 2 air combat environment in the environment and adopt the self-play method to learn the red-side tactical strategy. The geographical scene size is 50x50 kilometers. A total of 10,000 episodes are set in the training process. Each episode is an air combat process with a maximum of 2,000 decision steps. The hyperparameters are set as follows: The learning rates of the Actor and Critic networks lr = 0.0001, the discount factor γ = 0.95, the MAPPO hyperparameter clip coefficient ∈ = 0.2, the optimizer is Adam, and the batch size = 2000. The strategies adopted by the blue-side aircraft are divided into four levels (L1-L4). The higher the level, the more complex and diverse the behavior of the blue side. The stronger the ability of the trained red-side intelligent agent to adapt to complex environments. The trained intelligent agent strategy πi , π j ; The blue - side strategies are specifically classified as follows:
[0037] ·L1: static, the blue - side aircraft is in a stationary state;
[0038] ·L2: random, the blue - side aircraft randomly takes actions within the allowed action space;
[0039] ·L3: script, the blue - side aircraft engages in combat with the nearest red - side aircraft and, after the combat, moves away from the red - side aircraft;
[0040] ·L4: self - play, the red - side aircraft adopts the blue - side strategy trained by L3 to confront the opponent.
[0041] Large - scale multi - aircraft air combat scenario training and battlefield attention information analysis module based on attention network transfer: Large - scale (5 vs. 5) multi - aircraft air combat scenario training and battlefield attention information analysis module based on attention network transfer: From the designed air combat simulation platform, construct a close - range air combat scenario of five red - side aircraft and five blue - side aircraft. Randomly expand and increase the weights of the input layer of the decision network again for the model trained in the small - scale scenario, ensuring that its input dimension is consistent with that of the red - side aircraft. At the same time, reuse the parameters of the trained LSTM network and Transfomer network and put them into the large - scale multi - level air combat scenario for training. Analyze the output information from the attention networks of the two scenarios to obtain the degree of battlefield attention of each aircraft to friendly and enemy forces.
[0042] Embodiment 2
[0043] See Figures 2 - 13 As shown. The multi - agent reinforcement learning air combat decision - making method based on attention network transfer, this method is based on the following modules: multi - aircraft short - range air combat simulation platform, multi - agent reinforcement learning network design and small - scale multi - aircraft air combat scenario training, large - scale multi - aircraft air combat scenario training based on attention network transfer, and battlefield attention information analysis, including the following steps:
[0044] Step 1: Build a simulation environment for performing close - combat air combat tasks through the aircraft power equation. At the same time, read information such as the positions and weapons of the aircraft from the simulation environment, and define the state space, action space, and reward function of multi - agent reinforcement learning.
[0045] Multi-aircraft short-range air combat simulation platform: Construct a multi-aircraft air combat scenario where aircraft obtain the firing advantage over targets through tactical maneuvers, with short-range air-to-air missiles and machine guns as the main weapons; the two sides of the air combat are defined as the red and blue sides; the red aircraft are our side, and the tactical strategies are obtained through learning; the blue aircraft are used as opponents and adopt tactics with different difficulty levels based on random action sampling, air combat scripts, and virtual self-play to respond; model the above scenario to define aircraft characteristics; at the same time, according to the theory of partially observable Markov games, calculate the state space, action space, and reward function through the position, speed, and weapon information of the aircraft.
[0046] Step 2: Improve the well-known Actor-Critic network by adding an LSTM time-series processing module and a Transformer architecture based on global information, and use it as the decision network for multi-intelligent reinforcement learning. Then, train it with the MAPPO algorithm. Set up a small-scale multi-aircraft air combat scenario, and use the designed network and algorithm to train for the close-quarter combat air combat simulation scenario of two red aircraft and two blue aircraft.
[0047] Multi-agent reinforcement learning network design and small-scale (2 vs. 2) multi-aircraft air combat scenario training module: Design an Actor-Critic network and a reinforcement learning algorithm. For the multi-agent reinforcement learning air combat decision-making method based on attention network transfer. The network adopted is the Actor-Critic architecture. For the Actor and Critic networks, the input information includes: its own information O t,i , friendly force information O t,f , enemy force information O t,j and global information O t,full . After passing O t,i , O t,f , O t,j through a fully connected network, a high-dimensional tensor is obtained, and the global information O t,fullAfter the embedding operation, it is input into the LSTM network for sequential processing. The obtained tensor is connected to the Transformer module and concatenated with the high-dimensional tensor as the decision-making information for the Actor and Critic networks. A 2 vs. 2 air combat environment is set up in the environment, and the self-play method is used to learn the red side's tactical strategy; the geographical scene size is 50x50 kilometers; a total of 10,000 episodes are set in the training process, each episode is an air combat process with a maximum of 2,000 decision steps; the hyperparameters are set as follows: the learning rates of the Actor and Critic networks lr = 0.0001, the discount factor γ = 0.95, the MAPPO hyperparameter clipping coefficient ∈ = 0.2, the optimizer is Adam, batchsize = 2000; the strategies adopted by the blue side aircraft are divided into four levels (L1-L4), and the higher the level, the more complex and diverse the behavior of the blue side; the stronger the ability of the trained red side agent to adapt to complex environments; the agent strategy π i ,π j ; The specific classification of the blue side strategy is as follows:
[0048] ·L1: static, the blue side aircraft is in a static state;
[0049] ·L2: random, the blue side aircraft randomly takes actions within the allowed action space;
[0050] ·L3: script, the blue side aircraft engages in combat with the nearest red side aircraft and stays away from the red side aircraft after the combat;
[0051] ·L4: self-play, the red side aircraft adopts the blue side strategy trained in L3 to confront the opponent.
[0052] Step 3: Migrate the policy network trained in Step 2, reuse it in the close combat scenario of multiple red side aircraft and multiple blue side aircraft, extract the output information of the attention network, and perform attention analysis on this information to obtain the battlefield attention degree of each agent to the enemy and friendly forces.
[0053] Large-scale (5 vs. 5) multi-aircraft air combat scenario training and battlefield attention information analysis module based on attention network migration: Reuse the model trained in the small-scale scenario and put it into the large-scale multi-aircraft air combat scenario for training. Analyze the output information from the attention networks of the two scenarios to obtain the battlefield attention degree of each aircraft to the enemy and friendly forces. Specifically, the heat value range of the red side aircraft for each other aircraft in the battlefield is [-0.5, 0.5].
[0054] Next, we will further describe the working principle of the present invention.
[0055] 1: Multi - aircraft short - range air combat simulation platform
[0056] The multi - aircraft air combat scenario of the present invention refers to air combat within visual range, obtaining the firing advantage over the target through tactical maneuvers, and using short - range air - to - air missiles and cannons as the main weapons. The two sides of the air combat are defined as the red side and the blue side. Among them, the red - side aircraft is our side, and the tactical strategy is obtained through learning; the blue - side aircraft is used as the opponent and adopts tactics based on scripts with different difficulty levels to respond.
[0057] Based on the establishment of a 2D air combat aircraft model, air - to - air missiles and cannon weapons can be launched, and the characteristics are as follows:
[0058] Steering angular velocity range: [0, 5] degrees / second;
[0059] Speed range: [200, 500] m / s;
[0060] Short - range missile weapon: Maximum quantity: 8, range [0, 111] km; Launch cone angle (centered on the longitudinal axis of the aircraft) range [-60, 60] degrees; Single - shot kill probability: 0.75;
[0061] Cannon weapon: Maximum quantity: 400; Range [0, 2] km, launch cone angle range [-60, 60] degrees, maximum continuous firing time 200 seconds.
[0062] 2: Multi - agent reinforcement learning network design and small - scale (2 vs. 2) multi - aircraft air combat scenario training.
[0063] Multi - aircraft air combat is modeled as a partially observable Markov game (POMG) process. From the perspective of the red side, the air combat process can be represented as a six - tuple (S, O i (i ∈ N), A i (i ∈ n), P, R, γ). Among them, S is the air combat state jointly composed of the red and blue sides, Oi is the state observation of the ith red - side aircraft, and N >= 1 is the number of red - side aircraft. A i represents the action space of the ith red - side aircraft. The joint action A = A 1 ×…×A n , representing the joint action of all red - side aircraft at a certain decision step. P: S×A → Δ(S) is the state transition probability, indicating the transition probability of s → s' when the red - side aircraft group takes the joint action a in the state s ∈ S. represents the reward function, that is, the immediate reward obtained by the red - side aircraft group when taking the action a in the state s and entering the new state s’. γ ∈ [0, 1] is the discount factor, used to accumulate and calculate the long - term decision reward.
[0064] As a typical sequential decision-making problem, an air combat aircraft observes the environment, makes tactical decisions and executes them (in a simulation environment) (such as Figure 2 ), while receiving and accumulating rewards. In the partially observable Markov game model, the observation space of each aircraft consists of three parts:
[0065] 1) Its own information, expressed as:
[0066] O t,i _t^i = [x_t^i, y_t^i, v_t^i, h_t^i, A_{off_t}^{ij}, AA_t^{ij}, ATA_t^{ij}, A_{off_t}^{ji}, AA_t^{ji}, ATA_t^{ji}, d_t^i, \dot{d}_t^i, c_t^i, m_t^i, \Omega_t^i, l_t^i] i,j , AA i,j , ATA i,j , A_{off i,f , AA i,f , ATA i,f , d i,j , d i,f , c, m, \Omega, l]
[0067] Among them, the subscript t represents the current moment, the subscript i ∈ N represents the i-th aircraft of the red side, and N is the total number of red-side aircraft; the subscript j ∈ M represents the j-th aircraft of the blue side, and M is the total number of blue-side aircraft. The subscript f ∈ N -i represents the f-th friendly aircraft of aircraft i. (x, y) represents the 2D position (x, y) of the aircraft, v is the current speed, c is the remaining number of machine guns, and m is the remaining number of missiles. \Omega represents whether it is ready to launch the next missile, and l represents whether the aircraft is shooting. h represents the current turning angle of the aircraft, A_{off i,j represents the engagement angle of blue-side aircraft j relative to red-side aircraft i, AA i,j represents the azimuth angle of blue-side aircraft j relative to red-side aircraft i, and ATA i,j represents the radar tracking angle of blue-side aircraft j relative to red-side i. Other symbols with multiple subscripts can be understood in the same pattern.
[0068] 2) Friendly aircraft information, expressed as:
[0069] O t,f _t^{if} = [v_t^{if}, h_t^{if}, A_{off_t}^{ij}, AA_t^{ij}, ATA_t^{ij}, d_t^i, c_t^i, m_t^i] f,i , AA f,i , ATA f,i , d j,i , c, m]
[0070] 3) Air combat opponent information, expressed as:
[0071] O t,j _t^{oj} = [v_t^{oj}, h_t^{oj}, A_{off_t}^{ij}, AA_t^{ij}, ATA_t^{ij}, d_t^j, c_t^j, m_t^j] j,i , AA j,i , ATA j,i , d j,i , c, m]
[0072] Note: The original text seems to have some inconsistent or incomplete subscript notations in the formulas. I've tried to make sense of it and translate as accurately as possible. If there are any specific corrections or clarifications needed for the source text, it would help in providing a more precise translation. Also, the "<0000...>" tags are left unchanged as per the requirement.Connect the observation information of all red - side aircraft to obtain the global observation O t,full :
[0073] O t,full = O t,i ∪O t,f ∪O t,j
[0074] The action space is defined as: [h, v, c, r], where h represents the turning angle of the aircraft, v represents the flight speed of the aircraft, c ∈ [0, 1] represents whether to fire the cannon, and r ∈ [0, 1] represents whether to fire the missile. The actions h and v are discretized to simplify the action space, and we have:
[0075] h ∈ [-6, …, 6], which respectively represent the discrete action space of the aircraft turning angle in the range of [-180, 180] deg / s:
[0076] v = [0, …, 9]: which respectively represent the discrete action space of the aircraft speed in the range of [200, 500] m / s.
[0077] The reward function is defined as:
[0078] During air - to - air combat, facing the tail of the opponent is a relatively advantageous shooting scenario for oneself. Therefore, we use AA, D and the remaining amounts of cannon and missile of the participating attacking aircraft to encourage the agent to attack and improve its attack efficiency. AA is the azimuth angle of the red - side aircraft relative to the blue - side aircraft, and D is the distance between the two.
[0079]
[0080] Among them, C max , S max represent the maximum amounts of cannon and missile carried by the aircraft, and C remain , S remain represent the remaining amounts of cannon and missile after defeating the opponent. This reward function encourages the aircraft to turn towards the tail of the opponent to attack and destroy the opponent with the least amount of weapons. A penalty r f = - 5 is given when the aircraft flies out of the environmental boundary, and a penalty r m = - 2 is given when the aircraft destroys friendly aircraft. The total reward for each simulation step is: r = r a + r f + r m .
[0081] The adopted network is the Actor - Critic architecture. As Figure 3 shown, for the Critic network, in addition to O t,i , O t,f , O t,jIn addition to their respective individual input channels, the global observation O t,full is input as a separate channel into the Transformer module for processing. In this method, this structure is considered to be able to better express the relationship between individual and global situations, improve the sensitivity of individual decisions to the global situation, and enhance cooperation. A similar structure is also adopted in the Actor network. The Transformer has the advantage of strong expression ability for high-dimensional data space. Therefore, the training framework uses the Transformer to process the input global information, hoping to better learn the high-dimensional features closely related to tactical decisions from the high-dimensional observation data. The addition of the Long Short-Term Memory (LSTM) module is to further capture the temporal correlation of the observation information and enhance the historical cognitive ability. Finally, a shared fully connected layer is also set up in the network to share parameters between the actor and critic networks, further improving the consistency of multi-agent decisions.
[0082] Training is carried out for the 2vs.2 air combat scenario. Under the above-designed decision network, the MAPPO (Multi-Agent PPO) method is used for training. The blue side adopts four different strategies to train the red side aircraft. The geographical scene size is 50x50 kilometers. A total of 10,000 episodes are set in the training process. Each episode is an air combat process with a maximum of 2,000 decision steps. The Figure 4 training curve and Figure 5 the win rate graph are obtained. It can be seen from the two graphs that the rewards of the red side agents of the four strategies trained can all reach the convergence level. At the same time, the win rate is also as high as over 70%. This proves the effectiveness of this decision network.
[0083] 3: Training for large-scale (5vs.5) multi-aircraft air combat scenarios based on the transfer of attention networks and analysis of battlefield attention information.
[0084] In the Figure 3 shown structure, the Transformer module is used to process the global observation information, hoping to better capture the global features and provide support for the decisions of individual aircraft. Therefore, a simple idea is: if the Transformer module has learned to be able to better express the situation features of the 2vs.2 air combat scenario, then reusing the parameters of the Transformer module, can it speed up the learning of air combat strategies on a larger scale? Based on this idea, the present invention reuses the parameters of the Transformer module trained in the 2vs.2 air combat scenario in the 5vs.5 scenario, avoiding starting training from scratch, and hoping to improve the training efficiency on the premise of ensuring the training quality.
[0085] Analyze the attention information of the trained attention network and transfer it to a larger-scale battlefield environment. Through the transferred attention network, the agent can quickly obtain the battlefield situation information that needs to be focused on in different environments and tasks, improving the convergence efficiency and robustness of its model. Compare the model obtained through transfer training with the model trained from scratch to obtain Figure 8 and Figure 9 . Comparing the reward curves of training from scratch, the training curve based on the sharing of attention parameters of the 2.vs.2 model shows that the episode reward converges more quickly and the reward value after convergence is smoother. It is confirmed that the model has better convergence speed and robustness. At the same time, after testing the model trained from scratch, the winning rate is 71.0%, while the winning rate of the model obtained through transfer learning is as high as 82.0%, which confirms that the model has better combat capabilities.
[0086] From the Multihead_attention network in the above network structure, multiply o full as the network input by the trained attention weights to obtain the battlefield situation information with attention, and add up the situation information of each aircraft to obtain the attention information heat map of each aircraft. Obtain Figure 10 and Figure 11 . For the aircraft targets that have been destroyed, their attention is all 0, such as r1 in 10. For the red aircraft r5, the target with the highest attention is b7, so a missile was launched at this target; at the same time, the blue aircraft b9 is located at the tail of the aircraft r5 and is very close, so r5 also maintains a high level of attention to b9 (the attention value is 0.49).
[0087] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-agent reinforcement learning air combat decision-making system based on attention network migration, comprising: Multi-aircraft short-range air combat simulation platform, multi-agent reinforcement learning network design and small-scale multi-aircraft air combat scenario training module, large-scale multi-aircraft air combat scenario training and battlefield attention information analysis module based on attention network transfer, characterized in that: Multi-aircraft short-range air combat simulation platform: The aircraft dynamics and weapon characteristics are modeled using the two-dimensional fixed-wing aircraft dynamics equation. The flight speed and heading angle of the aircraft are set. At the same time, the weapon characteristics of the aircraft launching short-range air-to-air missiles and artillery are modeled. Air combat modeling is carried out based on the partially observable Markov game theory. The state space, action space and reward function of multi-aircraft air combat are set; Multi-agent reinforcement learning network design and small-scale multi-aircraft air combat scenario training module: It is improved on the basis of the well-known Actor-Critic network, and the LSTM network and multi-head attention mechanism network are introduced; The Mappo algorithm is adopted. In the multi-aircraft short-range air combat simulation platform, multiple red aircraft and multiple blue aircraft are set for close combat, and the improved Actor-Critic network is introduced; Large-scale multi-aircraft air combat scenario training and battlefield attention information analysis module based on attention network migration: Reuse the model trained in a small-scale scenario and place it in a large-scale multi-level air combat scenario for training; Analyze the output information in the attention networks of both the small-scale scenario and the large-scale multi-level air combat scenario to obtain the battlefield attention level of each aircraft towards friendly and enemy forces; The multi-agent reinforcement learning network design and small-scale multi-aircraft air combat scenario training module includes the following: The network adopted is the well-known Actor-Critic architecture, and the input information includes: its own information O t,i , friendly force information O t,f , enemy force information O t,j and global information O t,full ; After passing O t,i , O t,f , O t,j through a fully connected network, a high-dimensional tensor is obtained. After the global information Ot ,full undergoes an embedding operation, it is input into the LSTM network for sequential processing. The obtained tensor is connected to the Transformer module and concatenated with the high-dimensional tensor as the decision-making information for the Actor and Critic networks; Train for the close combat air combat simulation scenario of two red aircraft and two blue aircraft; Use the multi-agent proximal policy optimization algorithm MAPPO for training under the decision network framework.
2. The multi-agent reinforcement learning air combat decision-making system based on attention network migration according to claim 1, characterized in that: The multi-aircraft short-range air combat simulation platform uses two-dimensional fixed-wing aircraft dynamics and weapon characteristics for modeling, including the following contents: Construct a multi-aircraft air combat scenario. The aircraft obtains the firing advantage over the target through tactical maneuvers, and uses short-range air-to-air missiles and machine guns as the main weapons; The two sides of the air combat are defined as the red and blue sides; The red aircraft is our side, and the tactical strategy is obtained through learning; The blue aircraft is used as the opponent, and different tactics based on random action sampling, air combat scripts and virtual self-play with different difficulty levels are adopted to respond; The above scenario is modeled to define the aircraft characteristics; At the same time, according to the partially observable Markov game theory, the state space, action space and reward function are calculated through the position, speed and weapon information of the aircraft.
3. The multi-agent reinforcement learning air combat decision-making system based on attention network migration according to claim 1, characterized in that: The large-scale multi-aircraft air combat scenario training and attention analysis module based on attention network transfer includes the following contents: Reuse the parameters of the Transformer module trained in the close combat air combat scenario of two red aircraft and two blue aircraft in the close combat scenario of five red aircraft and five blue aircraft, and conduct training to improve the model convergence speed.
4. A multi-agent reinforcement learning air combat decision-making method based on attention network transfer, which can realize multi-agent air combat decision-making training, and through the transfer of the attention network, realize transfer learning of different tasks and battlefield attention analysis. This method is based on the system described in claim 1, characterized in that: Step 1: Build a simulation environment for performing close combat air combat tasks through the aircraft dynamic equation. At the same time, read the position and weapon information of the aircraft from the simulation environment, and define the state space, action space and reward function of multi-agent reinforcement learning; Step 2 improves the well-known Actor-Critic network by adding an LSTM time series processing module based on global information and a Transformer architecture, which is used as the decision-making network for multi-agent reinforcement learning, and then trained with the MAPPO algorithm; a small-scale multi-aircraft air combat scenario is set up, and the designed network and algorithm are used to train the close-range dogfight air combat simulation scenario of two red aircraft and two blue aircraft. Step 3: Migrate the policy network trained in Step 2, reuse and extract the output information of the attention network in the close-range dogfight scenario of five red aircraft and five blue aircraft, and perform attention analysis on this information to obtain the battlefield attention degree of each agent to the enemy and friendly forces.
5. The multi-agent reinforcement learning air combat decision-making method based on attention network migration according to claim 4, characterized in that: The content of Step 1 is as follows: Taking short-range air-to-air missiles and machine guns as the main weapons, the two sides of the air combat are defined as the red and blue sides, the red aircraft is our side, and the tactical strategy is obtained through learning; the blue aircraft is used as the opponent and adopts tactics with different difficulty levels based on random action sampling, air combat scripts and virtual self-play; the modeling equation of the aircraft in the two-dimensional plane is described as follows: ; where: x and y are the horizontal and vertical coordinates of the aircraft, u and v are the horizontal and vertical velocities of the aircraft, r, is the heading angular velocity and heading angle of the aircraft; F x , F y are the forces in the horizontal and vertical directions; A 2D air combat aircraft model is established, which can launch air-to-air missiles and cannon weapons, and the characteristics are as follows: · Steering angular velocity range: [0, 5] degrees per second; · Speed range: [200, 500] meters per second; · Short-range missile weapon: Maximum quantity: 8, range [0, 111] kilometers; Launch cone angle range [-60, 60] degrees centered on the longitudinal axis of the aircraft: Single-shot damage probability: 0.75; · Machine gun weapon: Maximum quantity: 400; Range [0, 2] kilometers, launch cone angle range [-60, 60] degrees, maximum continuous firing time 200 seconds.
6. The multi-agent reinforcement learning air combat decision-making method based on attention network migration according to claim 5, characterized in that: The content of Step 2 is as follows: The network adopted is the well-known Actor-Critic architecture; For the Actor and Critic networks, the input information includes: its own information O t,i , friendly force information O t,f , enemy force information O t,j and global information O t,full ; After passing O t,i , O t,f , O t,j through a fully connected network, a high-dimensional tensor is obtained. After passing the global information O t,full through an embedding operation, it is input into the LSTM network for sequential processing. The obtained tensor is connected to the Transformer module and concatenated with the high-dimensional tensor as the decision-making information for the Actor and Critic networks; A 2 vs. 2 air combat environment is set in the environment, and the self-play method is used to learn the red side's tactical strategies; The geographical scene size is 50x50 kilometers; A total of 10,000 episodes are set in the training process, and each episode is an air combat process with a maximum of 2,000 decision steps; The hyperparameters are set as follows: The learning rates of the Actor and Critic networks lr = 0.0001, the discount factor γ = 0.95, the MAPPO hyperparameter clipping coefficient ∈ = 0.2, the optimizer is Adam, and the batch size = 2000; The strategies adopted by the blue side aircraft are divided into four levels (L1 - L4), and the higher the level, the more complex and diverse the behavior of the blue side; The stronger the ability of the trained red side agent to adapt to complex environments; Extract the trained agent strategy π i , π j ; The specific classification of the blue side strategy is as follows: · L1: static, the blue aircraft is in a stationary state; · L2: random, the blue aircraft randomly takes actions within the allowed action space; · L3: script, the blue aircraft engages with the nearest red aircraft and moves away from the red aircraft after the engagement; · L4: self-play, the red aircraft uses the blue strategy trained in L3 to confront the opponent.
7. The multi-agent reinforcement learning air combat decision-making method based on attention network migration according to claim 5, characterized in that: Migrate the policy network trained in Step 3, reuse the parameters of the Transformer module trained in the close-range dogfight air combat scenario of two red aircraft and two blue aircraft in the close-range dogfight scenario of five red aircraft and five blue aircraft, and perform retraining to improve the model convergence speed; extract the output information of the attention network and perform attention analysis on this information to obtain the battlefield attention degree of each agent to the enemy and friendly forces.
8. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein the program controls the device where the non-volatile storage medium is located to execute the method described in any one of claims 5 to 7 when running.
9. An electronic device, characterized in that, It includes a processor and a memory; computer-readable instructions are stored in the memory, and the processor is used to run the computer-readable instructions, wherein the computer-readable instructions execute the method described in any one of claims 5 to 7 when running.
Citation Information
Patent Citations
Close-range air combat maneuver decision-making method based on improved DDPG
CN116661475A
Motorized penetration method combining transfer learning and deep reinforcement learning
CN117348391A