Multi-uav air combat decision-making method based on multi-agent hierarchical reinforcement learning

By employing a multi-agent hierarchical reinforcement learning method, the UAV decision-making process is divided into high-level and low-level policy agents. By combining heterogeneous policy synchronous learning and self-game mechanism, the problems of low sample learning efficiency and unstable state transition in multi-UAV air combat collaborative decision-making are solved, and efficient UAV collaborative control is achieved.

CN115291625BActive Publication Date: 2025-11-04TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210831674.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-11-04
Estimated Expiration
2042-07-15

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning algorithms have low sample learning efficiency, high computing power requirements, unstable state transitions, and complex communication architectures in multi-UAV air combat collaborative decision-making, making it difficult to balance exploration and utilization, and difficult to balance individual and team goals.

Method used

A multi-agent hierarchical reinforcement learning approach is adopted to abstract the UAV decision-making process into high-level and low-level policy agents. The high-level policy agents learn the target allocation strategy in the high-time dimension, while the low-level policy agents optimize the control strategy in the low-time dimension. By combining heterogeneous synchronous learning and self-game mechanism, the UAV is trained to make decisions based on only local observations.

Benefits of technology

It improves sample learning efficiency, reduces data storage and communication difficulty, has real-time and robust properties, avoids local optima, and can be effectively applied in real-world environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115291625B_ABST
    Figure CN115291625B_ABST
Patent Text Reader

Abstract

The application provides a multi-unmanned aerial vehicle air combat decision-making method based on multi-agent hierarchical reinforcement learning, the method comprising: constructing a simulation environment based on an actual multi-unmanned aerial vehicle air combat scene, including an environment constraint model, an unmanned aerial vehicle individual constraint model and an antagonistic influence rule; modeling the multi-unmanned aerial vehicle air combat problem as a semi-Markov game model, abstracting the decision-making process of a single unmanned aerial vehicle as two agents representing high-level and low-level strategies, and respectively defining the state space representation, action, reward function and action termination condition of each agent; adopting a multi-agent reinforcement learning algorithm combining heterogeneous strategy synchronous learning and self-game mechanism to train the unmanned aerial vehicle high-level and low-level strategy agents; after training, the unmanned aerial vehicle makes decisions based on the strategy network and local observation of the low-level strategy agent; the method can realize autonomous unmanned cooperative decision-making of the multi-unmanned aerial vehicle in the air combat environment, does not need human pilots to intervene, and has good instantaneity and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-unmanned aerial vehicle autonomous cooperative decision and control, in particular to a multi-unmanned aerial vehicle air combat decision method based on multi-agent hierarchical reinforcement learning. BACKGROUND

[0002] Multi-unmanned aerial vehicle cluster air combat generally refers to the fight between enemy and friendly unmanned aerial vehicles carrying weapons in a certain airspace, aiming to cooperatively attack the opponent while ensuring the survival of the self, and has the characteristics of three-dimensionality, large scale, high maneuverability, etc., which puts very high requirements on the immediacy, robustness and stability of unmanned aerial vehicle autonomous cooperative decision and control.

[0003] Multi-agent reinforcement learning integrates the perception, learning and decision-making of individuals into the same framework, and at the same time embodies the mutual cooperation between individuals, realizes complex swarm intelligence through multiple individuals with simple intelligence, and under the background of rapid progress of computing and storage technology in recent years, combined with deep learning, can realize the "end-to-end" decision from raw input to policy output, which is an effective method to solve the problem of multi-unmanned aerial vehicle cluster air combat autonomous cooperative decision, which is a high dynamic, nonlinear and strongly constrained system.

[0004] However, as a kind of method mainly driven by data, the existing multi-agent reinforcement learning algorithm often relies on a large amount of exploration of the environment when solving the complex problem of multi-unmanned aerial vehicle air combat cooperative decision, resulting in low sample learning efficiency and high demand for computing power. In order to solve such problems, some methods use the experience of human pilots for supervised pre-training, or combine expert rules to design actions to accelerate the convergence of the algorithm, but this brings the problem of easily making the strategy fall into local optimum.

[0005] Therefore, how to better balance exploration and utilization is a major difficulty in using deep multi-agent reinforcement learning method to solve such problems. In addition, the multi-agent environment also brings the problems of unstable state transition, complex communication architecture caused by partially observable states, and difficulty in balancing individual and team goals. SUMMARY

[0006] To address the aforementioned technical challenges, this application proposes a multi-agent hierarchical reinforcement learning-based multi-UAV air combat decision-making method. The UAV's decision-making process is abstracted into two agents representing high-level and low-level strategies, respectively. The high-level strategy agent learns target allocation strategies in a higher time dimension. First, it learns tactical-level strategies by combining situational estimation based on prior knowledge. Then, it further guides the low-level strategy agent to optimize basic control strategies in a lower time dimension, thereby avoiding excessive blind exploration in the continuous action space and improving sample utilization. A multi-agent reinforcement learning algorithm combining heterogeneous strategy synchronous learning and self-game mechanism is used for training. The centralized training and decentralized execution multi-agent reinforcement learning architecture allows the UAV to only need to port the strategy network of the low-level strategy agent and make decisions based on local observations, without relying on communication to obtain the global state, thus exhibiting good immediacy and robustness.

[0007] This application provides a multi-agent hierarchical reinforcement learning-based multi-UAV air combat decision-making method, the method comprising:

[0008] A simulated combat environment for multi-UAV collaborative air combat was constructed based on actual air combat scenarios.

[0009] The UAV air combat collaborative decision-making problem in the multi-UAV collaborative air combat simulation environment is constructed as a semi-Markov game model. Under the semi-Markov game model, the UAV decision-making process is abstracted into a high-level policy agent and a low-level policy agent.

[0010] A multi-agent reinforcement learning algorithm combining heterogeneous strategy synchronous learning and self-game mechanism is used to train the high-level policy agent and the low-level policy agent; wherein, the high-level policy agent learns the target allocation strategy based on the current situation and global state in a higher time dimension, and the low-level policy agent learns the optimal control strategy based on the current allocation target and local observations in a lower time dimension.

[0011] Decisions are made based on the policy network and local observations of the underlying policy agent.

[0012] Preferably, the construction of a multi-UAV cooperative air combat simulation environment based on actual air combat scenarios includes:

[0013] Based on actual air combat scenarios, a multi-UAV cooperative air combat simulation environment is constructed using computer simulation.

[0014] Preferably, the construction of a multi-UAV cooperative air combat simulation environment based on actual air combat scenarios using computer simulation includes:

[0015] Define an environmental constraint model, including the resistance to spatial regions and physical influencing factors;

[0016] Defining a UAV individual constraint model, including the motion ability, perception ability and firepower attack ability of the individual UAV;

[0017] Defining an antagonistic influence rule, including the antagonistic interaction mode, antagonistic target and win-lose condition of the enemy and friendly UAVs.

[0018] Preferably, the multi-UAV air combat cooperative decision-making problem is constructed as a semi-Markov game model, including:

[0019] The multi-UAV air combat cooperative decision-making problem is constructed as a semi-Markov game model by using a multi-agent hierarchical reinforcement learning method.

[0020] Preferably, the semi-Markov game model is described by a tuple .

[0021] wherein, is a finite set of all agents, including a subset of high-level strategy agents and a subset of low-level strategy agents is a joint state space; is a state transition probability; is a joint action space; is a reward; is an n-step termination condition of the upper-level decision-making.

[0022] Preferably, the multi-agent reinforcement learning algorithm combining the heterogeneous strategy synchronous learning and self-game mechanism is used to train the high-level strategy agents and the low-level strategy agents, including:

[0023] The high-level strategy agent H i is trained by using a double deep Q network algorithm, and a neural network and a Q B (s, a|θ B ) are calculated according to samples in an experience replay pool. A loss function is calculated and a gradient is back-propagated, and network parameters θ A and θ B are alternately updated; wherein S T and S T+1 are vectorized global states; and are the reward and action of H i .

[0024] The low-level strategy agent L i is trained by using a multi-agent proximal policy optimization algorithm, and a Critic neural network V i (S, a1, a2,..., a n |θ V) according to the sample The loss of the value function is calculated and the gradient is back-propagated to update the network parameters θ V ; wherein S t and S t+1 is the global state; is the high-level policy action at this time; and is the reward and action of L i ; the actor neural network π i (o i | θ π ) according to the sample The loss of the policy function is calculated and the gradient is back-propagated to update the network parameters θ π ; wherein, and are the vectorized local observations.

[0025] Preferably, the training of the high-level policy agent and the low-level policy agent comprises:

[0026] First stage: the enemy drone adopts a strategy based on expert rules: accelerate to the maximum speed after determining the target; maintain the same height as the target after determining the target; determine the attack target using the following priority function:

[0027]

[0028] wherein g ij is the priority factor of the drone i to the enemy drone j, and the target with the lowest priority factor is selected for attack; δ ij is the angle of the vector and projecting in the xy plane; ε max is the maximum change in heading angle in a single time step; h ij is the relative height of i and j; ζ max is the maximum change in height in a single time step; m j is the number of times the drone j has been assigned as a target;

[0029] Second stage: self-game training, the strategy network of the enemy and friendly drones trained in the first stage makes decisions, and the respective decision-making models are further trained based on the generated trajectory samples. A virtual self-game mechanism is adopted to avoid strategy loops.

[0030] Preferably, the method further comprises:

[0031] After the training is completed, the drone only retains the strategy network of the low-level policy agent, and outputs control actions by inputting local observations, which can be further migrated to actual scenes.

[0032] Compared with the prior art, the application has the beneficial technical effects:

[0033] 1) The decision-making process of a single unmanned aerial vehicle is abstracted into agents representing the strategies of the tactical level and the control level, respectively, an adaptive reward mechanism is designed to realize the synchronous training of the agents at different decision levels, the search of the solution space of the bottom-level strategy is guided by the high-level strategy, the sample learning efficiency is high, and the ability to jump out of the local optimum is also possessed, so that the exploration and utilization of the reinforcement learning algorithm can be well balanced.

[0034] 2) The training of the bottom-level strategy agent uses an algorithm with a centralized Critic and a decentralized Actor architecture, the actions of the high-level strategy agent are used as part of the input of the value network to evaluate the bottom-level strategy, and after the training is completed, the unmanned aerial vehicle only retains the policy network of the bottom-level strategy agent and makes decisions based on local observations, thereby reducing the difficulty of data storage, communication and calculation.

[0035] 3) The training process does not require the intervention of human pilots, and the simulation environment constructed based on a high-fidelity fixed-wing unmanned aerial vehicle model enables the method to be further migrated to a real environment. BRIEF DESCRIPTION OF DRAWINGS

[0036] Fig. 1 is a schematic diagram of a multi-unmanned aerial vehicle cooperative air combat decision-making model based on multi-agent hierarchical reinforcement learning provided by an embodiment of the application;

[0037] Fig. 2 is a process of an asynchronous strategy synchronous learning mechanism algorithm provided by an embodiment of the application;

[0038] Fig. 3 is a two-stage game training schematic diagram provided by an embodiment of the application. DETAILED DESCRIPTION

[0039] Referring to Figs. 1-3 , the application provides a multi-unmanned aerial vehicle air combat decision-making method based on multi-agent hierarchical reinforcement learning. First, a simulation environment is constructed based on an actual air combat scene: isomorphic symmetric capability fixed-wing unmanned aerial vehicles carrying limited missiles perform aerial combat confrontation in a certain three-dimensional airspace with the goal of eliminating opponents, and the control variables control the airspeed, heading angle, height and firing of the unmanned aerial vehicle, respectively, the environment constraint model, the unmanned aerial vehicle individual constraint model and the confrontation influence rule are defined, respectively.

[0040] Then, the multi-unmanned aerial vehicle air combat problem is modeled as a semi-Markov game (Semi-Markov Game) problem, which is described by a tuple , wherein is a finite set of all agents, including a subset representing high-level decision-making agents and a subset For a joint state space, Let be the state transition probability. For joint action space, As a reward, Let n be the termination condition for the higher-level decision-making process on the lower-level actions. The decision-making process of UAV i is abstracted into agents H representing the higher-level and lower-level strategies, respectively. i and L i H i In a higher time dimension, based on the current global state S T Execute strategic actions Return to the state S at the next moment. T+1 and rewards The state-space representation includes the following two parts:

[0041] (1) The three two-dimensional matrices represent the relative positions of the UAV on the xy, xz, and yz axes in three-dimensional space, respectively;

[0042] (2) A one-dimensional array of size 4*1 [v B , χ, γ, M] represent the current speed, heading angle, flight path angle, and remaining number of missiles of the machine, respectively.

[0043] Among them, the actions of high-level policy agents To select targets based on the current situation, the threat index σ of UAV i against j is calculated using the following formula. ij :

[0044]

[0045] Where α1, α2 and α3 are the weights of the corresponding distance, angle and speed threats, respectively, and should satisfy α1|α2|α3=1.

[0046] In this embodiment of the application, the threat index set Adv of UAV i against all n enemy aircraft is calculated. i -{σ i1 , ..., σ io Threat Index Set i ={σ 1i , ..., σ ni The high-level strategic agent has the following actions: 1) Attack the Advanced... i 1) Target and destroy the enemy aircraft with the highest threat index; 2) Attack the group of nearest friendly aircraft j (Thr) j 3) Attack the group of enemy aircraft with the highest threat index and destroy the target; j 4) Avoid Thr iIdentify enemy aircraft with the highest threat index and reduce their threat level.

[0047] In this embodiment of the application, the reward of the high-level policy agent This represents the cumulative reward obtained by the underlying policy agent within time step T. t0 and n are determined by the termination conditions: a) Drone i is determined to be dead; b) The currently selected target for attack or evasion changes.

[0048] In this embodiment of the application, the underlying policy agent L i Based on current local observations in a lower time dimension Execute action Return to local observation at the next time step and rewards Among the actions The control variables of the UAV within a unit time step

[0049] In this embodiment, the reward obtained by the underlying policy agent at time step t during its interaction with the environment is... The current action A of the high-level strategic agent i This is relevant in establishing the connection between the two-level decision-making models:

[0050]

[0051] in, This represents the distance between drone i and target drone j. Represents the velocity vector and relative pose vector The included angle, The scalar represents velocity, α and β are the weighting coefficients, which should satisfy α1+α2+α3=1 and β1+β2+β3=1 respectively, R0 is the basic reward, and R a With R d These are rewards for attacking and penalties for being hit.

[0052] Secondly, a multi-agent reinforcement learning algorithm combining heterogeneous policy synchronous learning and self-game mechanisms is used to train the high-level and low-level policy agents of the UAV. The high-level policy agent H... i The Double Deep Q Network (DDQN) algorithm is adopted, and the neural network Q... A (s,a|θ A ) and Q B (s,a|θ B Based on experience, the samples in the playback pool were replayed. Calculate the loss function and backpropagate the gradient, then alternately update the network parameters θ. A and θ B ST and S T+1 For vectorized global state, and For H i Rewards and actions.

[0053] In this embodiment of the application, the underlying policy agent L i Employing the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, the Critic neural network V... i (S, a1, a2, ..., a) n |θ V According to the sample Calculate the loss of the value function and backpropagate gradients to update the network parameters θ. V S i and S i+1 This is the global state. This is a strategic move by the top management at this time. and For L i Rewards and actions; Actor neural network π i (o i |θ π According to the sample Calculate the loss of the policy function and backpropagate gradients to update the network parameters θ. π ,in and For vectorized local observations.

[0054] In this embodiment, the algorithm training is divided into two stages. The first stage involves using a fixed rule strategy for the adversary drone: ① Accelerate to maximum speed after identifying the target; ② Maintain the same altitude as the target after identifying the target; ③ Use the following priority function to determine the attack target:

[0055]

[0056] Among them, g ij Let δ be the priority factor of drone i against enemy drone j, and select the target with the lowest priority factor to attack; ij For vectors and The angle between the projections onto the xy plane, ε max h is the maximum heading angle at a single time step. ij Let ζ be the relative height of i and j. max The maximum height change in a single time step; m j The number of times that drone j has been assigned as a target.

[0057] Train with fixed strategy opponent for a certain number of rounds N fix_pule After that, both sides make decisions based on the strategy network saved after the first stage training, and perform the second stage self-play training. In order to avoid the occurrence of strategy cycle, the fictitious self-play mechanism is adopted, and a certain number of rounds N are trained self_play .

[0058] Finally, the UAV only retains the strategy network of the bottom strategy agent after completing the training i (s|θ π ), input local observation Get action by sampling the output probability distribution

[0059] Please refer to Figs. 1-3 In other embodiments of the present application, the multi-agent hierarchical reinforcement learning-based multi-UAV air combat decision-making method comprises the following steps:

[0060] Step 1: Construct a visual simulation environment for multi-UAV air combat based on Matlab, and define environment constraint model, UAV individual constraint model and confrontation influence rule respectively. Among them, the environment constraint model defines the confrontation space area and physical influence factors; the individual constraint model defines the movement ability, perception ability and firepower striking ability of individual UAV; the confrontation influence rule defines the confrontation interaction mode, confrontation target and win-lose condition of enemy and friendly UAVs.

[0061] Step 1-1: Define the environment constraint model. The air combat area is a three-dimensional space with a unit length of 1000*1000*1000. The gravity acceleration is set as a constant g, the wind speed is 0, and there is no obstacle or no-fly area in the space.

[0062] Step 1-2: Define the individual constraint model, including the movement ability, perception ability and firepower striking ability of UAV, wherein the movement ability is described by the movement model of fixed-wing UAV:

[0063]

[0064] Among them, the velocity of displacement of UAV in the three-dimensional space defined by the x-y-z coordinate system is determined by its relative ground speed v g (airspeed v a =v g in a windless environment), heading angle χ and flight path angle γ; the control quantities u v , u φ , u h control the airspeed v a, heading angle χ and flight path angle γ, to realize the flight movement of the fixed-wing UAV in three-dimensional space, k * is the gain of the corresponding control variable.

[0065] wherein the perception capability is described by the detection range of the UAV radar:

[0066]

[0067] wherein, is the relative pose vector of the UAV i and j, d sen is the maximum perception radius of the UAV radar.

[0068] wherein the fire attack capability is the condition that needs to be met for the UAV to carry out effective attack:

[0069]

[0070] wherein d att is the maximum range of the UAV missile, ω ij is the speed vector of the UAV i and the angle between the relative pose vector d of the UAV j and the relative pose vector of the UAV i, ω max is the maximum launch angle of the UAV missile, and M is the current number of carried missiles.

[0071] Step 1-3: Define the rules of confrontation, including the confrontation interaction mode, confrontation target and winning and losing conditions of the enemy and friendly UAVs. Among them, the UAV can fire once under the condition of meeting the fire attack capability constraint, and the missile will hit the target with a certain probability:

[0072]

[0073] wherein α1 and α2 are weight coefficients, satisfying α1+α2=1, and the UAV is determined to be dead and exits the battlefield when it is hit by the missile, and cannot be perceived and attacked by other UAVs until the round ends.

[0074] In a feasible implementation manner, the confrontation targets of the enemy and friendly UAVs are to attack and destroy all opponents of the enemy side under the condition of ensuring the survival of the own side. When the maximum time step t max is reached or it is determined that one side wins, the current round ends. The conditions for determining that the red UAV wins are as follows: ① the number of surviving red UAVs n r ≥1, and the number of surviving blue UAVs n b =0; ② the maximum round time is reached, n r >n b ; ③ the maximum round time is reached and n r =n b , and the number of remaining missiles of the red side Mr Greater than the number of remaining missiles of the Blue Team (M) b .

[0075] A draw is declared if the following conditions are met: the maximum turn time is reached and n... r =n b M r =M b Except where the above conditions are met, the blue team's drone is deemed the winner.

[0076] Step 2: Employing a multi-agent hierarchical reinforcement learning method, the multi-UAV air combat cooperative decision-making problem described in Step 1 is modeled as a semi-Markov game problem, consisting of tuples. Describe it, in which, For a finite set of all intelligent agents, including a subset representing high-level decision-making agents. and a subset representing the underlying decision-making intelligent agent For a joint state space, Let be the state transition probability. For joint action space, As a reward, This is the termination condition for the n-step decision-making process at the higher level.

[0077] The general-purpose reinforcement learning library gym in Python encapsulates a Matlab-based environment to provide an interface for reinforcement learning algorithms.

[0078] Step 2-1: Execute the reset() command to initialize the environment and all underlying policy agents. and high-level strategic intelligent agents Return to the initial state of drone i. The state space representation includes the following two parts:

[0079] (1) The three two-dimensional matrices represent the relative positions of the UAV on the xy, xz, and yz axes in three-dimensional space, respectively. The global state matrix has a size of 1000*1000, and the local observation state matrix has a size of 2d. sen *2d sen Let the xy coordinates of UAV i be [x0, y0], the friendly UAV j within the perception range be [x1, y1], and the enemy UAV k be [x2, y2]. Let L represent the intelligent agent. i and H i The zero matrix of the global state The zero matrix B represents the local state. ix,iy =63, B fx,fy =127, B kx,ky =255.

[0080] Where ix = iy = d sen+1, jx=ix+(x1-x0), jy=iy+(y1-y0), kx=ix+(x2-x0), ky=iy+(y2-y0).

[0081] (2) A one-dimensional array of size 4*1 [v B [x, γ, M] represent other states of the machine.

[0082] Step 2-2: At each time step T, execute the step_H() command, and the high-level policy agent H... i State-Action Value Network Q i (s, a|θ) Input the current global state S T Output Action Interacting with the environment returns the global state S T+1 and rewards And The data is placed into the experience replay pool, where the action of the high-level policy agent is to select a target based on the current situation, and the threat index σ of UAV i against j is calculated according to the following formula. ij :

[0083]

[0084] Where α1, α2, and α3 are the weights of the corresponding distance, angle, and speed threats, respectively, and should satisfy α1 + α2 + α3 = 1. Calculate the threat index set Adv of UAV i against all n enemy aircraft. i ={σ i1 , ..., σ in Threat Index Set i ={σ 1i , ..., σ ni The high-level strategic agent has the following actions:

[0085]

[0086] Among them, the high-level strategic agent H i Reward at time step T for in For the lower-level policy agent L i The reward at time step t The actions specified for the termination conditions In the low-time dimension, the start and end times are defined as follows, where the termination condition is either one of the following: a) the drone i is determined to be dead; b) the currently selected target for attack or evasion changes.

[0087] Steps 2-3: Within each time step T defined by the termination condition of the upper-level policy, the step_L() command is executed at each time step t, and the lower-level policy agent L...i The policy network π(s|θ) takes the current local observation as input. Sampling actions based on the output probability distribution Return global state s i+1 Local observation and rewards Will and The actions of the underlying policy agent are placed into the experience replay pool. Defined as the basic control variables of a drone: v u x u h y F >, where u v Control the current airspeed, u x Control the current heading, u h Control the current height, u F It is a 0-1 variable that controls whether to fire.

[0088] Rewards for interacting with the environment It is expressed as follows:

[0089]

[0090] Among them, A i For the current high-level strategic intelligent agent H i The action, This represents the distance between drone i and target drone j. Represents the velocity vector and relative pose vector The included angle, The scalar represents velocity, α and β are the weighting coefficients, which should satisfy α1|α2|α3=1 and β1|β2|β3=1 respectively, R0 is the basic reward, and R a With R d The attack reward and the penalty for being hit are respectively opposite numbers and their absolute values ​​should be greater than R0. For example, if R0-1, then R a -100, R d --100.

[0091] Steps 2-4: When drone i is determined to be dead, agent L... i and H i Will exit the environment early; with L i Based on the time step of the action, when the maximum time step t is reached... max If one side is determined to be the winner, the current round ends, and the reset() command is executed to initialize the environment. The number of rounds is N. episode =N episode +1. ​

[0092] Step 3: A multi-agent reinforcement learning algorithm combining heterogeneous policy synchronous learning and self-game mechanisms is used to train the high-level and low-level policy agents of the UAV. The UAV's historical states, actions, and rewards are recorded.<s,a,T,s′> The state-action value network of the high-level policy agent and the policy network and state-value network of the low-level policy agent are trained using the form of samples. The training is carried out in two stages, in which the first stage is based on opponents with fixed rules and the second stage is self-game training.

[0093] Step 3-1: Use the DDQN algorithm to train two state-action value networks Q with the same hyperparameters on the high-level policy agent. A (s,a|θ A ) and Q B (s,a|θ B A batch of samples is sampled from the experience replay pool, and θ is updated alternately at a certain frequency according to the following loss function. A and θ B :

[0094]

[0095] Where b is the current batch sample size, a * The action corresponding to the current highest Q(s, a) value. For the kth sample in the current batch, corresponding to step 2-2 respectively

[0096] Step 3-2: Use the MAPPO algorithm to train the state-value network V(s|θ) on the underlying policy agent. V ) and policy network A batch of samples is sampled from the experience replay pool, where the state value network is updated according to the following loss function: in For the target network, clip is the cutoff function, s is the cutoff threshold, and s is the input state of the value network. k The corresponding global state S in steps 2-3 t And from the high-level policy agent H in step 2-2 i action The concatenation after vectorization.

[0097] In one feasible implementation, the policy network is updated according to the following loss function:

[0098] in, AG represents the action probability obtained from the old and new strategies under important sampling. kan advantage function obtained by the state value network output and the reward, represents the entropy of the strategy, and a is the weight coefficient of the term, the input state s of the strategy network k corresponding to the local observation in step 2-3 Since the UAVs in the environment are homogeneous, all UAVs share the agent H i and L i corresponding to the parameters of the neural network.

[0099] Step 3-3: The training process of steps 3-1 and 3-2 is divided into two stages, and the enemy UAVs in the first stage use a strategy based on expert rules: ① accelerate to the maximum speed after determining the target; ② maintain the same height as the target after determining the target; ③ determine the attack target using the following priority function:

[0100]

[0101] where g ij is the priority factor of UAV i to enemy UAV j, and the target with the lowest priority factor is selected for attack; δ ij is the angle between the x-y plane projections of and , and ε max is the maximum heading angle change in a single time step; h ij is the relative height of i and j, and ζ max is the maximum height change in a single time step; m j is the number of times UAV j has been assigned as a target.

[0102] The second stage is self-game training, and the strategy networks of enemy and friendly UAVs in the first stage training make decisions, and the generated trajectory samples are used to further train the respective decision models. To avoid strategy loops, a virtual self-game mechanism is used.

[0103] Step 4: After training, the UAVs only retain the strategy network of the bottom strategy agent, and output control actions by inputting local observations. They can be further migrated to actual scenarios.

Claims

1. A multi-agent hierarchical reinforcement learning-based multi-UAV air combat decision-making method, characterized in that, The method includes: constructing a multi-UAV collaborative air combat simulation environment based on actual air combat scenarios; The UAV air combat collaborative decision-making problem in the multi-UAV collaborative air combat simulation environment is constructed as a semi-Markov game model. Under the semi-Markov game model, the UAV decision-making process is abstracted into a high-level policy agent and a low-level policy agent. A multi-agent reinforcement learning algorithm combining heterogeneous strategy synchronous learning and self-game mechanism is used to train the high-level policy agent and the low-level policy agent; wherein, the high-level policy agent learns the target allocation strategy based on the current situation and global state in a higher time dimension, and the low-level policy agent learns the optimal control strategy based on the current allocation target and local observations in a lower time dimension. Decision-making is based on the policy network and local observations of the underlying policy agent; The aforementioned construction of a multi-UAV collaborative air combat simulation environment based on actual air combat scenarios includes: Based on actual air combat scenarios, a multi-UAV cooperative air combat simulation environment is constructed using computer simulation. Specifically, the multi-UAV air combat cooperative decision-making problem is constructed as a semi-Markov game model. Under this model, the UAV decision-making process is abstracted into a high-level policy agent and a low-level policy agent, including: A multi-agent hierarchical reinforcement learning method is employed to construct the multi-UAV air combat cooperative decision-making problem as a semi-Markov game model; the semi-Markov game model consists of tuples. Describe it; in, For a finite set of all intelligent agents, including a subset representing high-level policy agents. and a subset representing the underlying policy agent For a joint state space; The state transition probability; For joint action space; As a reward; The termination condition for the n-step decision-making process at the higher level; The multi-agent reinforcement learning algorithm, which combines heterogeneous policy synchronous learning with a self-game mechanism, trains the high-level policy agent and the low-level policy agent, including: The high-level strategic agent H i Training is performed using a dual-depth Q-network algorithm. The neural network Q... A (s,a|θ A ) and Q B (s,a|θ B Based on experience, the samples in the playback pool were replayed. Calculate the loss function and backpropagate the gradient, then alternately update the network parameters θ. A and θ B Among them, S T and S T+1 The global state is vectorized; and For H i Rewards and actions; The underlying policy agent L i The Critic neural network V is trained using a multi-agent proximal policy optimization algorithm. i (S, a1, a2, ..., a) n |θ V According to the sample Calculate the loss of the value function and backpropagate gradients to update the network parameters θ. V Among them, S t and S t+1 This is the global state; This is a strategic move by high-level personnel at this time; and For L i Rewards and actions; Actor neural network π i (o i |θ π According to the sample Calculate the loss of the policy function and backpropagate gradients to update the network parameters θ. π ;in, and For vectorized local observations; Phase 1: The enemy drone employs an expert rule-based strategy: after identifying the target, it accelerates to maximum speed; after identifying the target, it maintains the same altitude as the target; and it uses the following priority function to determine the attack target: Among them, g ij Let δ be the priority factor of drone i against enemy drone j, and select the target with the lowest priority factor to attack; ij For vectors and The angle between the projections onto the xy plane; ε max h is the maximum change in heading angle at a single time step. ij Let i be the relative height of i and j; ζ max The maximum height change in a single time step; m j The number of times drone j has been assigned as a target; Let i be the velocity vector of the drone. Let be the relative pose vectors of UAV i and UAV j; The second stage is self-game training. The policy networks trained in the first stage by both sides' drones make decisions, and further train their respective decision models based on the generated trajectory samples. To avoid policy loops, a virtual self-game mechanism is adopted.

2. The method according to claim 1, characterized in that, The aforementioned construction of a multi-UAV cooperative air combat simulation environment based on actual air combat scenarios, using computer simulation, includes: Define an environmental constraint model, including the resistance to spatial regions and physical influencing factors; Define an individual constraint model for unmanned aerial vehicles (UAVs), including the individual UAV's mobility, perception capabilities, and firepower capabilities; Define the rules of adversarial impact, including the interaction methods between enemy and friendly drones, adversarial targets, and conditions for victory or defeat.

3. The method according to claim 2, characterized in that, The method further includes: After training is completed, the drone retains only the policy network of the underlying policy agent. By inputting local observations and outputting control actions, it can be further transferred to real-world scenarios.

Citation Information

Patent Citations

  • Fixed-wing unmanned aerial vehicle autonomous control cooperation strategy training method

    CN112034888A

  • Multi-unmanned aerial vehicle cooperative air combat decision autonomous learning and semi-physical simulation verification method

    CN114167756A