Aircraft cluster cooperative control method based on MADRL
Through the MADRL-based aircraft cluster collaborative control method, the problem of unmanned combat aircraft relying on manual control and susceptible interference is solved, and the autonomous coordinated decision-making and rapid response of the aircraft cluster are realized, which improves air combat performance.
Patent Information
- Application Number
- CN202510641724.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-26
AI Technical Summary
Existing unmanned combat vehicles rely on operator manual control and remote communication, resulting in delayed communication control, susceptible to electronic warfare interference, and lack of autonomous intelligent decision-making capabilities.
The aircraft cluster collaborative control method based on multi-agent deep reinforcement learning (MADRL) is adopted. By constructing the multi-agent Markov decision-making process of the aircraft cluster, the deep reinforcement learning training cluster reward function and joint strategy are designed, and training is used using the Actor-Critic framework and the MATD3 algorithm to achieve autonomous collaborative control of the aircraft cluster.
It improves the robustness, rapidity and universality of coordinated control of aircraft clusters, enhances the strategy search efficiency and training stability in air combat, and ensures that aircraft clusters make better decisions in complex air combat environments.
Smart Images

Figure CN120540339A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aircraft controller design, and specifically relates to an aircraft cluster collaborative control method based on multi-agent deep reinforcement learning (MADRL). Background Art
[0002] With the development of autonomous navigation, remote flight control, and the aviation industry, unmanned combat air vehicles (UCAVs) have become widely used by militaries worldwide. Compared to manned UCAVs, UCAVs offer numerous advantages in areas such as covert operations, surveillance, and prolonged combat missions. However, current UCAVs rely heavily on manual operator control and remote communication technologies, resulting in communication control delays and making them more vulnerable to electronic warfare jamming and suppression. Therefore, multi-agent autonomous decision-making models that enable UCAVs to make autonomous and intelligent decisions during combat missions have become a current research hotspot in the aviation field.
[0003] Multi-agent autonomous decision-making models can autonomously learn and generate intelligent decisions for combat avoidance, circumventing remote control and communication by human operators, and can make collaborative decisions between aircraft clusters. Methods for autonomous air combat decision-making cover multiple fields, including game theory and deep reinforcement learning. Game theory is generally more suitable for solving discrete space problems, while problems in continuous action spaces such as air combat involve complex strategies and action search spaces, and require the exploration of more intelligent methods such as expert systems and deep reinforcement learning algorithms to make tactical decisions. However, because expert systems rely more on prior knowledge and rules, deep reinforcement learning algorithms are more suitable for solving air combat problems with complex uncertainties. In recent years, researchers have incorporated deep reinforcement learning tactical pursuit points into aircraft cluster air combat to build decision-making models, and this method has demonstrated excellent performance. Summary of the Invention
[0004] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide an aircraft cluster collaborative control method based on MADRL, which realizes the mapping of air combat battlefield state values to aircraft joystick and throttle action values, and has faster solution response capability, better versatility and robustness.
[0005] The specific technical solutions of the present invention are as follows:
[0006] A MADRL-based aircraft cluster collaborative control method includes the following steps:
[0007] S1. Obtain the aircraft cluster collaborative task, consider the aircraft simulation environment with sudden interference, and build the aircraft kinematic and dynamic models based on the task;
[0008] S2. Modeling the aircraft swarm task as a multi-agent Markov decision process, replacing the control process based on negative feedback with a decision process based on the aircraft environment state value;
[0009] S3. Based on a simulated multi-agent aircraft air combat scenario, design the cluster reward function and joint strategy required for deep reinforcement learning training. Perform deep reinforcement learning training on the multi-agent cluster system to generate throttle and three-axis rudder control commands for the aircraft based on the current aircraft environment state value.
[0010] S4. For variable-length trajectory scenarios where the length of observation sequences of different agents may vary, a deep reinforcement learning algorithm based on selective experience replay is designed to store the aircraft environment state values and the corresponding throttle and three-axis rudder control commands in the replay buffer pool.
[0011] Furthermore, the specific method of constructing the aircraft kinematic and dynamic models in step S1 is as follows:
[0012] Generate the environmental status value of the aircraft based on the three-dimensional position and attitude information of multiple aircraft, including position, attitude, speed, acceleration, running time, and whether it is alive;
[0013] Consider a cluster system consisting of n aircraft. The kinematic model of each aircraft is as follows:
[0014]
[0015] in, is the three-dimensional position coordinate of the i-th aircraft in the ground system, V i is the speed of the i-th aircraft, θ i is the trajectory inclination angle of the i-th aircraft, ψ Vi is the trajectory deviation angle of the i-th aircraft, η xi ,η yi ,η zi are random noises in three axis directions respectively;
[0016] Consider a cluster system consisting of n aircraft, and the 6-DOF dynamic model of each aircraft is:
[0017]
[0018] Among them, m i is the mass of the i-th aircraft, V i is the linear velocity of the i-th aircraft, ω i The angular velocity of the i-th aircraft, F pi ,F ai are the engine thrust and aerodynamic force on the i-th aircraft respectively. iis the moment of inertia of the i-th aircraft, M pi ,M ai are the torques generated by the engine thrust and aerodynamic force on the i-th aircraft, which are determined by the aircraft's throttle, elevator deflection, aileron roll deflection, and rudder deflection.
[0019] Furthermore, the specific method of modeling the aircraft cluster task as a multi-agent Markov decision process in step S2 is as follows:
[0020] The aircraft control problem is converted into a Markov decision process. The Markov decision is represented by an array {S, A, P, R}, where S represents the state space and s t ∈S,s t is the environmental state value of the aircraft at time t, including the aircraft position, attitude, speed, acceleration and running time; A represents the action space, a t ∈A,a t is the flight control instruction from the intelligent control platform, which includes the throttle and three-axis rudder control instructions; P represents the state transition probability of all states; R represents the reward value obtained from all state transitions; and the strategy π:S→A represents the mapping from the state space to the action space.
[0021] Furthermore, the Markov decision process is:
[0022] When the system state s is observed t After that, the agent starts from the action set A t Randomly select an output action a t And execute in the environment, the environment then transfers to the next state s t+1 At the same time, the reward value R t Feedback to the agent;
[0023] The Q value calculated by the Q function is used to evaluate the value of the output action taken, guiding the agent to continuously adjust the output action to obtain the maximum cumulative reward value and ultimately achieve the maximum winning rate. This interactive process will continue until the training is completed.
[0024] Furthermore, the algorithm used in step S3 to perform deep reinforcement learning training on the multi-agent cluster system is based on the Actor-Critic framework, that is, the neural network required for training consists of an Actor network and a Critic network. The Actor network is a policy network consisting of a main network with a parameter φ and a policy network with a parameter The target network is composed of a main network with a parameter of θ and a target network with a parameter of The target network is composed of the Actor network, which evaluates the value of the Actor network output action and generates the gradient of the Actor network update;
[0025] Both the Actor network and the Critic network are composed of fully connected neural networks. During the training process, the aircraft environment state value in the historical data set is input into the Actor network to generate throttle and three-axis rudder control instructions for controlling the aircraft. The aircraft environment state value, the throttle and three-axis rudder control instructions for controlling the aircraft, the aircraft environment state value generated by the flight environment according to the control instructions, and the reward function value are stored in a replay buffer pool; at the same time, the aircraft environment state value, the generated reward function value, and the corresponding throttle and three-axis rudder control instructions are input into the Critic network, and the target Q value is generated by the Critic network.
[0026] Furthermore, the joint strategy expression used in the deep reinforcement learning training in step S3 is:
[0027]
[0028] Where μ(a|s,ρ) is the joint strategy of the aircraft cluster, μ(a i |s i ,ρ i ) is the strategy of the i-th aircraft, a i is the action space of the i-th aircraft, s i is the state space of the i-th aircraft, ρ i is the strategy parameter of the i-th aircraft, and the joint strategy is used to achieve rapid learning and training of the aircraft cluster system;
[0029] In step S3, in order to achieve coordinated control of the aircraft cluster, the total reward of the system of n aircraft at time t is expressed as:
[0030]
[0031] Where, in is the cluster reward between the i-th aircraft and the other n-1 aircraft at time t; α jis the induction factor of the jth aircraft; σ is the baseline distance, which affects the starting point and intensity of the reward decay with distance. A larger σ means stronger interaction between distant agents and looser group behavior. β is the distance decay factor, which is used to adjust the decay rate of the reward as the distance between two agents changes. A larger β means faster decay, which means a more significant impact on agents that are closer.
[0032] The total return expression obtained by the joint strategy is:
[0033]
[0034] Where γ is the discount factor. Since the reward is bounded in the expression of the total reward, when the reward approaches 0 during the training process of the aircraft cluster system, the system is considered to have completed training.
[0035] Furthermore, the training algorithm of step S3 adopts a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm;
[0036] The MATD3 algorithm combines the "centralized training, distributed execution" framework of the dual-delay deep deterministic policy gradient algorithm (TD3) and the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient), so that each agent has its own actor network and critic network. The actor network includes an actor policy network and a target actor policy network, and the critic network includes two centralized independent evaluation critic value networks and a target critic value network. For each agent, the TD3 algorithm decouples the action value function and uses two Q networks to approximate the actor network and the critic network, which can effectively solve the problem of overestimation of Q value. The MATD3 algorithm adopts a policy delay update method.
[0037] Furthermore, the step S3 performs deep reinforcement learning training on the multi-agent cluster system, and the specific method of generating the throttle and three-axis rudder control instructions for controlling the aircraft from the current aircraft environment state value is as follows:
[0038] During the training process, for each agent, the neural network parameters are first initialized, including the Actor strategy network π φ , Critic value network Q θ ; Initialize the target network parameters, including the Target Actor strategy network π φand TargetCritic value network The target network parameters and the Actor strategy network π φ and Critic value network Q θ Same; randomly sample data from the single-step data playback buffer pool of each agent (s t ,a t ,r t ,s t+1 ), train the Actor policy network π through the designed loss function and the selected optimizer φ and Critic value network Q θ ;Target Actor strategy network and Target Critic Value Network Perform soft updates;
[0039] The critic network uses the mean square error formula to The value approaches the expected output y value of the Critic network, and the strategy π in the Actor network φ Maximization The direction of the value is approached, and after a large number of iterations, the strategy π φ A strategy to maximize the reward function;
[0040] After the training is completed, the aircraft environment state value generated by the aircraft flight simulation platform is input into the strategy π φ In the network, the output is the throttle and three-axis rudder control command value A π .
[0041] Furthermore, the expected output y value of the critic network is calculated according to the following formula:
[0042]
[0043] Among them, γ is the discount factor, r is the reward function value, is the target value of the Critic value network, s is the aircraft environment state value, The target value of the Actor strategy network.
[0044] Furthermore, step S4 designs a deep reinforcement learning algorithm based on selective experience replay, storing the aircraft environment state value and the corresponding throttle and three-axis rudder control instructions into a data replay buffer pool for multi-agent deep reinforcement learning training. The specific method is as follows:
[0045] The time series data of each agent in the air combat environment is collected as historical data and stored in the sequence data playback buffer pool. Each data consists of the last state, action value, and reward function of each aircraft. All data are segmented with the minimum time series data length h. The data format in the sequence data playback buffer pool is {S k ,A k ,R k L,S k+h ,A k+h ,R k+h} form, and each agent generates its own experience replay buffer based on its own and other agents' observation and action states, rather than sharing a common experience replay buffer; during the training process of multi-agent deep reinforcement learning, each agent only interacts with the data in its corresponding experience replay buffer;
[0046] Design a target performance comparator to compare the control performance of different time series data in the same environment and output a good or bad preference value. The target performance is designed based on the signal deviation value and the response time. Randomly sample two groups of sequences {σ0, σ1} from the sequence data cache replay pool. The target performance comparator generates a preference result space Y. The preference result space Y is matched with the sequence to convert it into a sequence data group {σ0, σ1, Y} containing preferences.
[0047] The agent action controller is designed as a deep neural network. The sequence data group {σ0, σ1, Y} of each agent is used to iterate the actor network and critic network parameters of each agent through the MATD3 algorithm, and the loss function is optimized to update the parameters. The "centralized training, distributed execution" method is used to input the observed state values of all agents into each agent, and the corresponding agent only outputs its own corresponding action value, and the optimal collaboration strategy between multiple agents is obtained through training.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] (1) The present invention provides a MADRL-based aircraft cluster collaborative control method, which obtains aircraft cluster collaborative control instructions by adopting a multi-agent deep reinforcement learning algorithm. Compared with existing methods, it improves the robustness, rapidity and versatility of aircraft cluster collaborative control.
[0050] (2) The present invention designs an aircraft controller based on a deep reinforcement learning algorithm, and designs cluster rewards and joint strategies based on the state values of the aircraft cluster, thereby improving the efficiency of strategy search in multi-agent air combat.
[0051] (3) The present invention designs a deep reinforcement learning algorithm based on selective experience replay. Compared with existing aircraft control algorithms, the MATD3 algorithm that applies the experience replay mechanism can enhance training stability and efficiency in air combat scenarios involving destroyed units, enabling aircraft clusters to make better decisions in air combat missions. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions disclosed in the present invention, the following briefly introduces the drawings required for use in some embodiments disclosed in the present invention. Obviously, the drawings described below are only drawings of some embodiments disclosed in the present invention, and those skilled in the art can also derive other drawings based on these drawings. Furthermore, the drawings described below should be viewed as schematic diagrams and are not intended to limit the actual dimensions of the products, actual processes of the methods, etc. involved in the embodiments disclosed in the present invention.
[0053] Figure 1 Flowchart of the MADRL-based aircraft cluster collaborative control method provided by the present invention;
[0054] Figure 2 Schematic diagram of Euler angles and rudder control values for a 6-DOF aircraft provided by the present invention;
[0055] Figure 3 Schematic diagram of the Actor-Critic framework provided by the present invention;
[0056] Figure 4 A schematic diagram of the network structure of the MATD3 algorithm provided by the present invention;
[0057] Figure 5 A schematic diagram of the single-step data playback caching mechanism provided by the present invention;
[0058] Figure 6 This is a schematic diagram of the data playback caching mechanism provided by the present invention.
[0059] Explanation of Figure Numbers
[0060] 1- Elevator; 2- Pitch; 3- Aileron; 4- Roll; 5- Rudder; 6- Yaw. DETAILED DESCRIPTION
[0061] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0062] The present invention provides a MADRL-based aircraft cluster collaborative control method, comprising the following steps:
[0063] S1. Obtain the aircraft cluster collaborative task, consider the aircraft simulation environment with sudden interference, and build the aircraft kinematic and dynamic models based on the task;
[0064] S2. Modeling the aircraft swarm task as a multi-agent Markov decision process, replacing the control process based on negative feedback with a decision process based on the aircraft environment state value;
[0065] S3. Based on a simulated multi-agent aircraft air combat scenario, design the cluster reward function and joint strategy required for deep reinforcement learning training. Perform deep reinforcement learning training on the multi-agent cluster system to generate throttle and three-axis rudder control commands for the aircraft based on the current aircraft environment state value.
[0066] S4. For variable-length trajectory scenarios where the length of observation sequences of different agents may vary, a deep reinforcement learning algorithm based on selective experience replay is designed to store the aircraft environment state values and the corresponding throttle and three-axis rudder control commands in the replay buffer pool.
[0067] Furthermore, the specific method of constructing the aircraft kinematic and dynamic models in step S1 is as follows:
[0068] Generate the environmental status value of the aircraft based on the three-dimensional position and attitude information of multiple aircraft, including position, attitude, speed, acceleration, running time, and whether it is alive;
[0069] Consider a cluster system consisting of n aircraft. The kinematic model of each aircraft is as follows:
[0070]
[0071] in, is the three-dimensional position coordinate of the i-th aircraft in the ground system, V i is the speed of the i-th aircraft, θ i is the trajectory inclination angle of the i-th aircraft, ψ Vi is the trajectory deviation angle of the i-th aircraft, η xi ,ηyi ,η zi are random noises in three axis directions respectively;
[0072] Consider a cluster system consisting of n aircraft, and the 6-DOF dynamic model of each aircraft is:
[0073]
[0074] Among them, m i is the mass of the i-th aircraft, V i is the linear velocity of the i-th aircraft, ω i The angular velocity of the i-th aircraft, F pi ,F ai are the engine thrust and aerodynamic force on the i-th aircraft respectively. i is the moment of inertia of the i-th aircraft, M pi ,M ai are the torques generated by the engine thrust and aerodynamic force on the i-th aircraft, which are determined by the aircraft's throttle, elevator deflection, aileron roll deflection, and rudder deflection.
[0075] Furthermore, the specific method of modeling the aircraft cluster task as a multi-agent Markov decision process in step S2 is as follows:
[0076] The aircraft control problem is converted into a Markov decision process. The Markov decision is represented by an array {S, A, P, R}, where S represents the state space and s t ∈S,s t is the environmental state value of the aircraft at time t, including the aircraft position, attitude, speed, acceleration and running time; A represents the action space, a t ∈A,a t is the flight control instruction from the intelligent control platform, which includes the throttle and three-axis rudder control instructions; P represents the state transition probability of all states; R represents the reward value obtained from all state transitions; and the strategy π:S→A represents the mapping from the state space to the action space.
[0077] Furthermore, the Markov decision process is:
[0078] When the system state s is observed t After that, the agent starts from the action set A t Randomly select an output action a t And execute in the environment, the environment then transfers to the next state s t+1 At the same time, the reward value R tFeedback to the agent;
[0079] The Q value calculated by the Q function is used to evaluate the value of the output action taken, guiding the agent to continuously adjust the output action to obtain the maximum cumulative reward value and ultimately achieve the maximum winning rate. This interactive process will continue until the training is completed.
[0080] Furthermore, the algorithm used in step S3 to perform deep reinforcement learning training on the multi-agent cluster system is based on the Actor-Critic framework, that is, the neural network required for training consists of an Actor network and a Critic network. The Actor network is a policy network consisting of a main network with a parameter φ and a policy network with a parameter The target network is composed of a main network with a parameter of θ and a target network with a parameter of The target network is composed of the Actor network, which evaluates the value of the Actor network output action and generates the gradient of the Actor network update;
[0081] Both the Actor network and the Critic network are composed of fully connected neural networks. During the training process, the aircraft environment state value in the historical data set is input into the Actor network to generate throttle and three-axis rudder control instructions for controlling the aircraft. The aircraft environment state value, the throttle and three-axis rudder control instructions for controlling the aircraft, the aircraft environment state value generated by the flight environment according to the control instructions, and the reward function value are stored in a replay buffer pool; at the same time, the aircraft environment state value, the generated reward function value, and the corresponding throttle and three-axis rudder control instructions are input into the Critic network, and the target Q value is generated by the Critic network.
[0082] Furthermore, the joint strategy expression used in the deep reinforcement learning training in step S3 is:
[0083]
[0084] Where μ(a|s,ρ) is the joint strategy of the aircraft cluster, μ(a i |s i ,ρ i ) is the strategy of the i-th aircraft, a i is the action space of the i-th aircraft, s i is the state space of the i-th aircraft, ρ i is the strategy parameter of the i-th aircraft, and the joint strategy is used to achieve rapid learning and training of the aircraft cluster system;
[0085] In step S3, in order to achieve coordinated control of the aircraft cluster, the total reward of the system of n aircraft at time t is expressed as:
[0086]
[0087] Where, in is the cluster reward between the i-th aircraft and the other n-1 aircraft at time t; α j is the induction factor of the jth aircraft; σ is the baseline distance, which affects the starting point and intensity of the reward decay with distance. A larger σ means stronger interaction between distant agents and looser group behavior. β is the distance decay factor, which is used to adjust the decay rate of the reward as the distance between two agents changes. A larger β means faster decay, which means a more significant impact on agents that are closer.
[0088] The total return expression obtained by the joint strategy is:
[0089]
[0090] Where γ is the discount factor. Since the reward is bounded in the expression of the total reward, when the reward approaches 0 during the training process of the aircraft cluster system, the system is considered to have completed training.
[0091] Furthermore, the training algorithm of step S3 adopts a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm;
[0092] The MATD3 algorithm combines the "centralized training, distributed execution" framework of the dual-delay deep deterministic policy gradient algorithm (TD3) and the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient), so that each agent has its own actor network and critic network. The actor network includes an actor policy network and a target actor policy network, and the critic network includes two centralized independent evaluation critic value networks and a target critic value network. For each agent, the TD3 algorithm decouples the action value function and uses two Q networks to approximate the actor network and the critic network, which can effectively solve the problem of overestimation of Q value. The MATD3 algorithm adopts a policy delay update method.
[0093] Furthermore, the step S3 performs deep reinforcement learning training on the multi-agent cluster system, and the specific method of generating the throttle and three-axis rudder control instructions for controlling the aircraft from the current aircraft environment state value is as follows:
[0094] During the training process, for each agent, the neural network parameters are first initialized, including the Actor strategy network π φ , Critic value network Q θ ; Initialize target network parameters, including Target Actor strategy network and TargetCritic value network The target network parameters and the Actor strategy network π φ and Critic value network Q θ Same; randomly sample data from the single-step data playback buffer pool of each agent (s t ,a t ,r t ,s t+1 ), train the Actor policy network π through the designed loss function and the selected optimizer φ and Critic value network Q θ ;Target Actor strategy network and Target Critic Value Network Perform soft updates;
[0095] The critic network uses the mean square error formula to The value approaches the expected output y value of the Critic network, and the strategy π in the Actor network φ Maximization The direction of the value is approached, and after a large number of iterations, the strategy π φ A strategy to maximize the reward function;
[0096] After the training is completed, the aircraft environment state value generated by the aircraft flight simulation platform is input into the strategy π φ In the network, the output is the throttle and three-axis rudder control command value A π .
[0097] Furthermore, the expected output y value of the critic network is calculated according to the following formula:
[0098]
[0099] Among them, γ is the discount factor, r is the reward function value, is the target value of the Critic value network, s is the aircraft environment state value, The target value of the Actor strategy network.
[0100] Furthermore, step S4 designs a deep reinforcement learning algorithm based on selective experience replay, storing the aircraft environment state value and the corresponding throttle and three-axis rudder control instructions into a data replay buffer pool for multi-agent reinforcement learning training. The specific method is as follows:
[0101] The time series data of each agent in the air combat environment is collected as historical data and stored in the sequence data playback buffer pool. Each data consists of the last state, action value, and reward function of each aircraft. All data are segmented with the minimum time series data length h. The data format in the sequence data playback buffer pool is {S k , A k , R k L, S k+h , A k+h , R k+h} form, and each agent generates its own experience replay buffer based on its own and other agents' observation and action states, rather than sharing a common experience replay buffer; during the training process of multi-agent reinforcement learning, each agent only interacts with the data in its corresponding experience replay buffer;
[0102] Design a target performance comparator to compare the control performance of different time series data in the same environment and output a good or bad preference value. The target performance is designed based on the signal deviation value and the response time. Randomly sample two groups of sequences {σ0, σ1} from the sequence data cache replay pool. The target performance comparator generates a preference result space Y. The preference result space Y is matched with the sequence to convert it into a sequence data group {σ0, σ1, Y} containing preferences.
[0103] The agent action controller is designed as a deep neural network. The sequence data group {σ0, σ1, Y} of each agent is used to iterate the actor network and critic network parameters of each agent through the MATD3 algorithm, and the loss function is optimized to update the parameters. The "centralized training, distributed execution" method is used to input the observed state values of all agents into each agent, and the corresponding agent only outputs its own corresponding action value, and the optimal collaboration strategy between multiple agents is obtained through training.
[0104] Example 1
[0105] This embodiment provides a MADRL-based aircraft cluster collaborative control method, the process is as follows: Figure 1 As shown, the following steps are included:
[0106] S1. Obtain the aircraft cluster collaborative task, consider the aircraft simulation environment with sudden interference, and build the aircraft kinematic and dynamic models based on the task;
[0107] S2. Modeling the aircraft swarm task as a multi-agent Markov decision process, replacing the control process based on negative feedback with a decision process based on the aircraft environment state value;
[0108] S3. Based on a simulated multi-agent aircraft air combat scenario, design the cluster reward function and joint strategy required for deep reinforcement learning training. Perform deep reinforcement learning training on the multi-agent cluster system to generate throttle and three-axis rudder control commands for the aircraft based on the current aircraft environment state value.
[0109] S4. For variable-length trajectory scenarios where the length of observation sequences of different agents may vary, a deep reinforcement learning algorithm based on selective experience replay is designed to store the aircraft environment state values and the corresponding throttle and three-axis rudder control commands in the replay buffer pool.
[0110] Furthermore, the specific method of constructing the aircraft kinematic and dynamic models in step S1 is as follows:
[0111] Generate the environmental status value of the aircraft based on the three-dimensional position and attitude information of multiple aircraft, including position, attitude, speed, acceleration, running time, and whether it is alive;
[0112] Consider a cluster system consisting of n aircraft. The kinematic model of each aircraft is as follows:
[0113]
[0114] in, is the three-dimensional position coordinate of the i-th aircraft in the ground system, V i is the speed of the i-th aircraft, θ i is the trajectory inclination angle of the i-th aircraft, ψ Vi is the trajectory deviation angle of the i-th aircraft, η xi ,η yi ,η zi are random noises in three axis directions respectively;
[0115] Consider a cluster system consisting of n aircraft, and the 6-DOF dynamic model of each aircraft is:
[0116]
[0117] Among them, m i is the mass of the i-th aircraft, V i is the linear velocity of the i-th aircraft, ω iThe angular velocity of the i-th aircraft, F pi ,F ai are the engine thrust and aerodynamic force on the i-th aircraft respectively. i is the moment of inertia of the i-th aircraft, M pi ,M ai are the torques generated by the engine thrust and aerodynamic force on the i-th aircraft, which are determined by the aircraft's throttle, elevator 1 deflection, aileron 3 roll rudder deflection, and rudder 5 deflection. The Euler angles and control variables of the 6-DOF aircraft are as follows: Figure 2 shown.
[0118] Furthermore, the specific method of modeling the aircraft cluster task as a multi-agent Markov decision process in step S2 is as follows:
[0119] The aircraft control problem is converted into a Markov decision process. The Markov decision is represented by an array {S, A, P, R}, where S represents the state space and s t ∈S,s t is the environmental state value of the aircraft at time t, including the aircraft position, attitude, speed, acceleration and running time; A represents the action space, a t ∈A,a t is the flight control instruction from the intelligent control platform, which includes the throttle and three-axis rudder control instructions; P represents the state transition probability of all states; R represents the reward value obtained from all state transitions; and the strategy π:S→A represents the mapping from the state space to the action space.
[0120] Furthermore, the Markov decision process is:
[0121] When the system state s is observed t After that, the agent starts from the action set A t Randomly select an output action a t And execute in the environment, the environment then transfers to the next state s t+1 At the same time, the reward value R t Feedback to the agent;
[0122] The Q value calculated by the Q function is used to evaluate the value of the output action taken, guiding the agent to continuously adjust the output action to obtain the maximum cumulative reward value and ultimately achieve the maximum winning rate. This interactive process will continue until the training is completed.
[0123] Furthermore, the algorithm used in step S3 to perform deep reinforcement learning training on the multi-agent cluster system is based on the Actor-Critic framework, such as Figure 3The neural network required for training consists of an Actor network and a Critic network. The Actor network is a policy network consisting of a main network with a parameter of φ and a Critic network with a parameter of The target network is composed of a main network with a parameter of θ and a target network with a parameter of The target network is composed of the Actor network, which evaluates the value of the Actor network output action and generates the gradient of the Actor network update;
[0124] Both the Actor network and the Critic network are composed of fully connected neural networks. During the training process, the aircraft environment state value in the historical data set is input into the Actor network to generate throttle and three-axis rudder control instructions for controlling the aircraft. The aircraft environment state value, the throttle and three-axis rudder control instructions for controlling the aircraft, the aircraft environment state value generated by the flight environment according to the control instructions, and the reward function value are stored in a replay buffer pool; at the same time, the aircraft environment state value, the generated reward function value, and the corresponding throttle and three-axis rudder control instructions are input into the Critic network, and the target Q value is generated by the Critic network.
[0125] Furthermore, the joint strategy expression used in the deep reinforcement learning training in step S3 is:
[0126]
[0127] Where μ(a|s,ρ) is the joint strategy of the aircraft cluster, μ(a i |s i ,ρ i ) is the strategy of the i-th aircraft, a i is the action space of the i-th aircraft, s i is the state space of the i-th aircraft, ρ i is the strategy parameter of the i-th aircraft, and the joint strategy is used to achieve rapid learning and training of the aircraft cluster system;
[0128] In step S3, in order to achieve coordinated control of the aircraft cluster, the total reward of the system of n aircraft at time t is expressed as:
[0129]
[0130] Where, in is the cluster reward between the i-th aircraft and the other n-1 aircraft at time t; α jis the induction factor of the jth aircraft; σ is the baseline distance, which affects the starting point and intensity of the reward decay with distance. A larger σ means stronger interaction between distant agents and looser group behavior. β is the distance decay factor, which is used to adjust the decay rate of the reward as the distance between two agents changes. A larger β means faster decay, which means a more significant impact on agents that are closer.
[0131] The total return expression obtained by the joint strategy is:
[0132]
[0133] Where γ is the discount factor. Since the reward is bounded in the expression of the total reward, when the reward approaches 0 during the training process of the aircraft cluster system, the system is considered to have completed training.
[0134] Furthermore, the training algorithm of step S3 adopts MATD3 algorithm, and the structure design of MATD3 algorithm is as follows: Figure 4 As shown;
[0135] The MATD3 algorithm combines the "centralized training, distributed execution" framework of the TD3 algorithm and the MADDPG algorithm, so that each agent has its own Actor network and Critic network. The Actor network includes an Actor policy network and a Target Actor policy network, and the Critic network includes two centralized independent evaluation Critic value networks and a Target Critic value network. For each agent, the TD3 algorithm decouples the action-value function and uses two Q networks to approximate the Actor network and the Critic network, which can effectively solve the problem of overestimation of Q values. The MATD3 algorithm adopts a policy delayed update method.
[0136] Furthermore, the step S3 performs deep reinforcement learning training on the multi-agent cluster system, and the specific method of generating the throttle and three-axis rudder control instructions for controlling the aircraft from the current aircraft environment state value is as follows:
[0137] During the training process, for each agent, the neural network parameters are first initialized, including the Actor strategy network π φ , Critic value network Q θ ; Initialize target network parameters, including Target Actor strategy network and TargetCritic value network The target network parameters and the Actor strategy network π φ and Critic value network Q θ Same; single-step data playback cache mechanism is as follows Figure 5 As shown, data is randomly sampled from the single-step data playback buffer pool of each agent (s t ,a t ,r t ,s t+1 ), train the Actor policy network π through the designed loss function and the selected optimizer φ and Critic value network Q θ ;Target Actor strategy network and Target Critic Value Network Perform soft updates;
[0138] The critic network uses the mean square error formula to The value approaches the expected output y value of the Critic network, and the strategy π in the Actor network φ Maximization The direction of the value is approached, and after a large number of iterations, the strategy π φ A strategy to maximize the reward function;
[0139] After the training is completed, the aircraft environment state value generated by the aircraft flight simulation platform is input into the strategy π φ In the network, the output is the throttle and three-axis rudder control command value A π .
[0140] Furthermore, the expected output y value of the critic network is calculated according to the following formula:
[0141]
[0142] Among them, γ is the discount factor, r is the reward function value, is the target value of the Critic value network, s is the aircraft environment state value, The target value of the Actor strategy network.
[0143] Furthermore, the step S4 designs a deep reinforcement learning algorithm based on selective experience replay, and stores the aircraft environment state value and the corresponding throttle and three-axis rudder control instructions into a data replay cache pool for multi-agent deep reinforcement learning training. The data replay cache mechanism is as follows: Figure 6 The specific method is as follows:
[0144] The time series data of each agent in the air combat environment is collected as historical data and stored in the sequence data playback buffer pool. Each data consists of the last state, action value, and reward function of each aircraft. All data are segmented with the minimum time series data length h. The data format in the sequence data playback buffer pool is {Sk ,A k ,R k L,S k+h ,A k+h ,R k+h} form, and each agent generates its own experience replay buffer based on its own and other agents' observation and action states, rather than sharing a common experience replay buffer; during the training process of multi-agent deep reinforcement learning, each agent only interacts with the data in its corresponding experience replay buffer;
[0145] Design a target performance comparator to compare the control performance of different time series data in the same environment and output a good or bad preference value. The target performance is designed based on the signal deviation value and the response time. Randomly sample two groups of sequences {σ0, σ1} from the sequence data cache replay pool. The target performance comparator generates a preference result space Y. The preference result space Y is matched with the sequence to convert it into a sequence data group {σ0, σ1, Y} containing preferences.
[0146] The agent action controller is designed as a deep neural network. The sequence data group {σ0, σ1, Y} of each agent is used to iterate the actor network and critic network parameters of each agent through the MATD3 algorithm, and the loss function is optimized to update the parameters. The "centralized training, distributed execution" method is used to input the observed state values of all agents into each agent, and the corresponding agent only outputs its own corresponding action value, and the optimal collaboration strategy between multiple agents is obtained through training.
[0147] This embodiment not only enables autonomous intelligent decision-making in the flight of a multi-agent aircraft cluster system in a complex battlefield environment, but also improves the robustness, rapidity and versatility of the collaborative control of the aircraft cluster by designing reward functions and joint strategies.
[0148] Although the present invention has been disclosed above with reference to preferred embodiments, this is not intended to limit the present invention. Any person skilled in the art may utilize the above disclosure to make possible changes and modifications to the technical solutions of the present invention without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent variations, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the scope of protection of the technical solutions of the present invention. The scope of protection of the present invention shall be based on the appended claims.
Claims
1. A MADRL-based aircraft cluster collaborative control method, characterized in that: The following steps are involved: S1. Obtain the aircraft cluster collaborative task, consider the aircraft simulation environment with sudden interference, and build the aircraft kinematic and dynamic models based on the task; S2. Modeling the aircraft swarm task as a multi-agent Markov decision process, replacing the control process based on negative feedback with a decision process based on the aircraft environment state value; S3. Based on a simulated multi-agent aircraft air combat scenario, design the cluster reward function and joint strategy required for deep reinforcement learning training. Perform deep reinforcement learning training on the multi-agent cluster system to generate throttle and three-axis rudder control commands for the aircraft based on the current aircraft environment state value. S4. For variable-length trajectory scenarios where the length of observation sequences of different agents may vary, a deep reinforcement learning algorithm based on selective experience replay is designed to store the aircraft environment state values and the corresponding throttle and three-axis rudder control commands in the replay buffer pool.
2. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: The specific method of constructing the aircraft kinematic and dynamic models in step S1 is as follows: Generate the environmental status value of the aircraft based on the three-dimensional position and attitude information of multiple aircraft, including position, attitude, speed, acceleration, running time, and whether it is alive; Consider a cluster system consisting of n aircraft. The kinematic model of each aircraft is as follows: in, is the three-dimensional position coordinate of the i-th aircraft in the ground system, V i is the speed of the i-th aircraft, θ i is the trajectory inclination angle of the i-th aircraft, ψ Vi is the trajectory deviation angle of the i-th aircraft, η xi ,η yi ,η zi are random noises in three axis directions respectively; Consider a cluster system consisting of n aircraft, and the 6-DOF dynamic model of each aircraft is: Where m i is the mass of the i-th aircraft, V i is the linear velocity of the i-th aircraft, ω i The angular velocity of the i-th aircraft, F pi ,F ai are the engine thrust and aerodynamic force on the i-th aircraft respectively. i is the moment of inertia of the i-th aircraft, M pi ,M ai are the torques generated by the engine thrust and aerodynamic force on the i-th aircraft, which are determined by the aircraft's throttle, elevator deflection, aileron roll deflection, and rudder deflection.
3. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: The specific method of step S2 for modeling the aircraft cluster task as a multi-agent Markov decision process is as follows: The aircraft control problem is converted into a Markov decision process. The Markov decision is represented by an array {S, A, P, R}, where S represents the state space and s t ∈S,s t is the environmental state value of the aircraft at time t, including the aircraft position, attitude, speed, acceleration and running time; A represents the action space, a t ∈A,a t is the flight control instruction from the intelligent control platform, which includes the throttle and three-axis rudder control instructions; P represents the state transition probability of all states; R represents the reward value obtained from all state transitions; and the strategy π:S→A represents the mapping from the state space to the action space.
4. The MADRL-based aircraft cluster collaborative control method according to claim 1 or 3, characterized in that: The Markov decision process is: When the system state s is observed t After that, the agent starts from the action set A t Randomly select an output action a t And execute in the environment, the environment then transfers to the next state s t+1 At the same time, the reward value R t Feedback to the agent; The Q value calculated by the Q function is used to evaluate the value of the output action taken, guiding the agent to continuously adjust the output action to obtain the maximum cumulative reward value and ultimately achieve the maximum winning rate. This interactive process will continue until the training is completed.
5. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: The algorithm used in step S3 to perform deep reinforcement learning training on the multi-agent cluster system is based on the Actor-Critic framework, that is, the neural network required for training consists of an Actor network and a Critic network; Both the Actor network and the Critic network are composed of fully connected neural networks. During the training process, the aircraft environment state value in the historical data set is input into the Actor network to generate throttle and three-axis rudder control instructions for controlling the aircraft. The aircraft environment state value, the throttle and three-axis rudder control instructions for controlling the aircraft, the aircraft environment state value generated by the flight environment according to the control instructions, and the reward function value are stored in a replay buffer pool; at the same time, the aircraft environment state value, the generated reward function value, and the corresponding throttle and three-axis rudder control instructions are input into the Critic network, and the target Q value is generated by the Critic network.
6. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: The joint strategy expression used in the deep reinforcement learning training in step S3 is: Where μ(a|s,ρ) is the joint strategy of the aircraft cluster, μ(a i |s i ,ρ i ) is the strategy of the i-th aircraft, a i is the action space of the i-th aircraft, s i is the state space of the i-th aircraft, ρ i is the strategy parameter of the i-th aircraft, and the joint strategy is used to achieve rapid learning and training of the aircraft cluster system; In step S3, in order to achieve coordinated control of the aircraft cluster, the total reward of the system of n aircraft at time t is expressed as: Where, in is the cluster reward between the i-th aircraft and the other n-1 aircraft at time t; α j is the induction factor of the jth aircraft; σ is the baseline distance, which affects the starting point and intensity of the reward decay with distance. A larger σ means stronger interaction between distant agents and looser group behavior. β is the distance decay factor, which is used to adjust the decay rate of the reward as the distance between two agents changes. A larger β means faster decay, which means a more significant impact on agents that are closer. The total return expression obtained by the joint strategy is: Where γ is the discount factor. Since the reward is bounded in the expression of the total reward, when the reward approaches 0 during the training process of the aircraft cluster system, the system is considered to have completed training.
7. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: The training algorithm of step S3 adopts MATD3 algorithm; The MATD3 algorithm combines the "centralized training, distributed execution" framework of the TD3 algorithm and the MADDPG algorithm, so that each agent has its own Actor network and Critic network. The Actor network includes an Actor policy network and a Target Actor policy network, and the Critic network includes two centralized independent evaluation Critic value networks and a Target Critic value network. In addition, for each agent, the TD3 algorithm decouples the action-value function and uses two Q networks to approximate the Actor network and the Critic network, which can effectively solve the problem of overestimation of Q values. The MATD3 algorithm adopts a policy delayed update method.
8. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: The specific method of performing deep reinforcement learning training on the multi-agent cluster system and generating the throttle and three-axis rudder control instructions for controlling the aircraft from the current aircraft environment state value is as follows: During the training process, for each agent, the neural network parameters are first initialized, including the Actor strategy network π φ , Critic value network Q θ ; Initialize target network parameters, including Target Actor strategy network and Target Critic Value Network The target network parameters and the Actor strategy network π φ and Critic value network Q θ Same; randomly sample data from the single-step data playback buffer pool of each agent (s t ,a t ,r t ,s t+1 ), train the Actor policy network π through the designed loss function and the selected optimizer φ and Critic value network Q θ ;Target Actor strategy network and Target Critic Value Network Perform soft updates; The critic network uses the mean square error formula to The value approaches the expected output y value of the Critic network, and the strategy π in the Actor network φ Maximization The direction of the value is approached, and after a large number of iterations, the strategy π φ A strategy to maximize the reward function; After the training is completed, the aircraft environment state value generated by the aircraft flight simulation platform is input into the strategy π φ In the network, the output is the throttle and three-axis rudder control command value A π .
9. The MADRL-based aircraft cluster collaborative control method according to claim 8, characterized in that: The expected output y value of the Critic network is calculated according to the following formula: Among them, γ is the discount factor, r is the reward function value, is the target value of the Critic value network, s is the aircraft environment state value, The target value of the Actor strategy network.
10. The MADRL-based aircraft cluster collaborative control method according to claim 1, characterized in that: Step S4 designs a deep reinforcement learning algorithm based on selective experience replay, storing the aircraft environment state value and the corresponding throttle and three-axis rudder control instructions in a data replay buffer pool for multi-agent deep reinforcement learning training. The specific method is as follows: The time series data of each agent in the air combat environment is collected as historical data and stored in the sequence data playback buffer pool. Each data consists of the last state, action value, and reward function of each aircraft. All data are segmented with the minimum time series data length h. The data format in the sequence data playback buffer pool is {S k ,A k ,R k L,S k+h ,A k+h ,R k+h } form, and each agent generates its own experience replay buffer based on its own and other agents' observation and action states, rather than sharing a common experience replay buffer; during the training process of multi-agent deep reinforcement learning, each agent only interacts with the data in its corresponding experience replay buffer; Design a target performance comparator to compare the control performance of different time series data in the same environment and output a good or bad preference value. The target performance is designed based on the signal deviation value and the response time. Randomly sample two groups of sequences {σ0, σ1} from the sequence data cache replay pool. The target performance comparator generates a preference result space Y. The preference result space Y is matched with the sequence to convert it into a sequence data group {σ0, σ1, Y} containing preferences. The agent action controller is designed as a deep neural network. Using the sequence data set {σ0, σ1, Y} of each agent, the actor network and critic network parameters of each agent are iterated through the MATD3 algorithm, and the loss function is optimized to update the parameters. Using the "centralized training, distributed execution" method, the observed state values of all agents are input into each agent, and the corresponding agent only outputs its own corresponding action value, and the optimal collaborative strategy between multiple agents is obtained through training.
Citation Information
Cited By
Large-scale complex group consensus decision-making method, terminal equipment and medium
CN120996079A
Aircraft envelope test method and device, storage medium and electronic equipment
CN121778183A