A system and method for defending against incoming missiles based on deep reinforcement learning

By building an autonomous decision-making model for aircraft through deep reinforcement learning algorithms, the problem of insufficient autonomy and adaptability of aircraft in defending against incoming missiles is solved, and the battlefield survivability and penetration capabilities are improved.

CN115562007BActive Publication Date: 2025-10-03NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211147804.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2025-10-03
Estimated Expiration
2042-09-20

AI Technical Summary

Technical Problem

Existing technologies mainly rely on manual decision-making when defending aircraft against incoming missiles. A single strategy is difficult to cope with complex and changing battlefield situations, resulting in insufficient autonomy and adaptability, and human errors may cause huge losses.

Method used

Using a deep reinforcement learning algorithm, a relative motion model of the carrier aircraft, decoys, and incoming missiles is constructed, a Markov decision process is established, and a neural network is built. Through training and simulation verification, autonomous decision-making plans are generated, including decoy missile deployment and carrier aircraft maneuver plans.

Benefits of technology

It enables autonomous and adaptive intelligent decision-making of aircraft in complex battlefield environments, improves battlefield survivability and penetration capabilities, and reduces the risks caused by human errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115562007B_ABST
    Figure CN115562007B_ABST
Patent Text Reader

Abstract

The present invention discloses a system and method for defending against incoming missiles based on deep reinforcement learning. The system includes a model building module, a Markov decision module, a neural network creation module, a training module, a simulation verification module, and a decision module. The method comprises the following steps: first, a relative motion model of the carrier aircraft, decoy, and incoming missile is built, and a Markov decision process is established. Then, a neural network is built according to a deep reinforcement learning algorithm, and the neural network is trained using the training module to obtain a trained neural network model. The simulation verification module then uses the trained neural network model for simulation verification. Finally, based on the simulated and verified neural network model, an interference delivery plan and a carrier aircraft maneuver plan are generated to defend against incoming missiles. The algorithm of the present invention is highly efficient and has strong autonomous decision-making capabilities. It enhances the autonomous, adaptive, and intelligent decision-making capabilities of aircraft for typical scenarios, thereby improving the aircraft's battlefield survivability and penetration capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aircraft decision-making confrontation, and in particular to a system and method for defending against incoming missiles based on deep reinforcement learning. Background Art

[0002] Currently, artificial intelligence technology is rapidly developing and maturing. Strategic research for both escaping and intercepting forces is shifting from traditional problem-solving methods to intelligent game-based confrontation methods. Intelligent game-based confrontation can make decisions based on its own state and the perceived state of the opponent. It learns through continuous interaction with the environment to compete with the opponent, without requiring knowledge of the opponent's strategy. This offers certain advantages in complex and ever-changing confrontation scenarios. Currently, intelligent game-based confrontation is rarely used in aircraft attack and defense confrontations. When faced with such problems, only fixed solutions are used, and only single strategies are considered, such as studying decoy strategies or maneuvering strategies. In complex combat scenarios, a single penetration strategy is unlikely to guarantee successful aircraft penetration. Therefore, a combination of multiple strategies is necessary to adapt to changing battlefield situations.

[0003] Furthermore, current decoy deployment by aircraft is primarily human-based. This highly dependent on manual effort to counter incoming missiles, and any human error can result in significant losses. Therefore, the early implementation of intelligent aircraft decision-making would be of great benefit to aircraft decision-making countermeasures. Summary of the Invention

[0004] The purpose of the present invention is to provide a deep reinforcement learning-based incoming missile defense countermeasure system and method with high algorithm efficiency, strong autonomous decision-making capability, the ability to enhance the autonomy and adaptive intelligent decision-making capability of aircraft, and the ability to improve the battlefield survivability and penetration capability of aircraft.

[0005] The technical solution for achieving the purpose of the present invention is: a deep reinforcement learning-based incoming missile defense countermeasure system, including a model building module, a Markov decision module, a neural network creation module, a training module, a simulation verification module and a decision module, wherein:

[0006] The model building module is used to construct the relative motion equations and models between the carrier aircraft, the decoy, and the incoming missile;

[0007] The Markov decision module is used to establish a Markov decision process, including a state space, an action space, a state transition equation, and a reward function;

[0008] The neural network creation module is used to build a neural network based on a deep reinforcement learning algorithm;

[0009] The training module is used to train the neural network to obtain a trained neural network model;

[0010] The simulation verification module is used to call the trained neural network model for simulation verification;

[0011] The decision module is used to generate a decoy missile delivery plan and an aircraft maneuvering plan.

[0012] A deep reinforcement learning-based approach to defend against incoming missiles. The steps are as follows:

[0013] Step 1: The model building module builds the relative motion model of the carrier aircraft, decoy and incoming missile;

[0014] Step 2: The Markov decision module establishes a Markov decision process based on the relative motion model of the carrier aircraft, decoy, and incoming missile, including the state space, action space, state transition equation, and reward function;

[0015] Step 3: The neural network creation module builds a neural network based on the deep reinforcement learning algorithm;

[0016] Step 4: The training module trains the neural network based on the relative motion model of the carrier aircraft, decoy and incoming missile, the state space, the action space, the state transition equation and the reward function to obtain a trained neural network model;

[0017] Step 5: The simulation verification module calls the trained neural network model for simulation verification;

[0018] Step 6: The decision module generates an interference delivery plan and an aircraft maneuvering plan based on the neural network model verified by simulation to defend against the incoming missile.

[0019] Compared with the existing technology, the present invention has the following significant advantages: (1) adopting the deep reinforcement learning DDPG algorithm, designing a deep reinforcement learning program for countering incoming missiles, using a neural network to fit the mapping relationship between the carrier aircraft and the decoy missile control, and training it, so that the carrier aircraft can use the trained neural network to make autonomous decisions; (2) by establishing the dynamic model and motion equation of the carrier aircraft, decoy and incoming missile, applying deep reinforcement learning and other methods to design and train the defense decision model, achieving rapid autonomous decision-making, and improving the autonomy and adaptability of the aircraft for typical scenarios; (3) establishing a simulation environment, exploring ways and methods to apply deep reinforcement learning and other methods to carry out self-learning defense decision-making technology against incoming missiles, by constructing a simulation environment model for deep reinforcement learning and using the training of the deep reinforcement learning algorithm, the algorithm efficiency is continuously improved, rapid autonomous decision-making is achieved, and the battlefield survivability and penetration capability of the aircraft are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1This is a flow chart of a method for defending against incoming missiles based on deep reinforcement learning according to the present invention.

[0021] Figure 2 This is a flowchart of the Actor-Critic algorithm in a specific embodiment of the present invention.

[0022] Figure 3 Schematic diagram of the structure of the decision neural network in a specific embodiment of the present invention.

[0023] Figure 4 Schematic diagram of the structure of the valuation neural network in a specific embodiment of the present invention.

[0024] Figure 5 2 is a graph showing the change of the reward function in an embodiment of the present invention.

[0025] Figure 6 This is a motion trajectory curve diagram of the carrier aircraft, decoy and incoming missile in an embodiment of the present invention.

[0026] Figure 7 1 is a curve diagram of the relative distance change between the carrier aircraft and the incoming missile in an embodiment of the present invention.

[0027] Figure 8 1 is a thrust variation curve diagram of the aircraft maneuvering device in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] The present invention provides an incoming missile defense countermeasure system based on deep reinforcement learning, comprising a model building module, a Markov decision module, a neural network creation module, a training module, a simulation verification module and a decision module;

[0030] The model building module is used to construct the relative motion equations and models between the carrier aircraft, the decoy, and the incoming missile;

[0031] The Markov decision module is used to establish a Markov decision process, including a state space, an action space, a state transition equation, and a reward function;

[0032] The neural network creation module is used to build a neural network based on a deep reinforcement learning algorithm;

[0033] The training module is used to train the neural network to obtain a trained neural network model;

[0034] The simulation verification module is used to call the trained neural network model for simulation verification;

[0035] The decision module is used to generate a decoy missile delivery plan and an aircraft maneuvering plan.

[0036] like Figure 1 As shown in Figure 1, a deep reinforcement learning-based defense countermeasure method for incoming missiles has the following steps:

[0037] Step 1: The model building module builds the relative motion model of the carrier aircraft, decoy and incoming missile;

[0038] Step 2: The Markov decision module establishes a Markov decision process based on the relative motion model of the carrier aircraft, decoy, and incoming missile, including the state space, action space, state transition equation, and reward function;

[0039] Step 3: The neural network creation module builds a neural network based on the deep reinforcement learning algorithm;

[0040] Step 4: The training module trains the neural network based on the relative motion model of the carrier aircraft, decoy and incoming missile, the state space, the action space, the state transition equation and the reward function to obtain a trained neural network model;

[0041] Step 5: The simulation verification module calls the trained neural network model for simulation verification;

[0042] Step 6: The decision module generates an interference delivery plan and an aircraft maneuvering plan based on the neural network model verified by simulation to defend against the incoming missile.

[0043] Furthermore, the model building module in step 1 builds a relative motion model of the carrier aircraft, the decoy, and the incoming missile, as follows:

[0044] First, a dynamic model of the carrier aircraft, decoy, and incoming missile is established to analyze the various forces acting on them. This model is used to establish the motion and dynamics of the aircraft in a complex force field environment. This provides a foundation for subsequent research on the model. Specifically,

[0045] The dynamic equation of the incoming missile in the inertial coordinate system is:

[0046]

[0047] Where: x m ,y m , z m is the missile's position vector; v x_m , v y_m , v z_m is the velocity vector of the missile; P x , P y , P z is the thrust vector of the missile; m is the mass of the missile; M is the mass of the earth; G is the gravitational constant; rm is the distance of the missile relative to the origin;

[0048] The motion equation of the carrier in the inertial coordinate system is:

[0049]

[0050]

[0051] In the formula, (x t ,y t , z t ) is the coordinate of the carrier in the inertial coordinate system, v t ,θ t 、φ t is the speed, track pitch angle and track deflection angle of the carrier aircraft, n tx The longitudinal control overload of the target, n ty 、n tz They are the turning control overloads in the target yaw and pitch directions respectively;

[0052] Assume that the position vector of the carrier aircraft relative to the missile is r, and the velocity vector is V. It can be expressed in the inertial coordinate system by (r, V, q ε ,q β ), then the relative motion equation between the missile and the carrier aircraft is:

[0053]

[0054] Where r x =x t -x m , r y =y t -y m , r z =z t -z m ; V x =v x_t -v x_m , V y =v y_t -v y_m , V z =v z_t -v z_m ;q β is the sight azimuth; q ε It is the sight height angle;

[0055] Derivative of Equation (4) with respect to time yields:

[0056]

[0057] The motion equation of the decoy bomb in the inertial coordinate system is:

[0058]

[0059]

[0060] In the formula, (x d ,y d , z d ) is the coordinate of the carrier in the inertial coordinate system, v d ,θ d 、φ d are the speed, track pitch angle and track deflection angle of the carrier aircraft.

[0061] Assume that the missile adopts proportional guidance to attack the carrier aircraft. Its guidance instructions are:

[0062]

[0063] Where: k y 、k z is the proportional guidance coefficient; Indicates the rate of change of the visual range between the missile and the carrier aircraft; They represent the rate of change of sight elevation angle and azimuth angle respectively;

[0064] The maneuvering device of the carrier aircraft is assumed to generate a fixed amount of thrust. Due to fuel limitations, the number of times the maneuvering device can operate is limited. The thrust transition process during the power-on and power-off of the maneuvering device is ignored. The thrust is 0 when the device is off, and the thrust is P when the device is on. The maneuvering command of the carrier aircraft is K, where K∈{-1,0,1}. The projected components of the force generated by the maneuvering device on each axis in the body coordinate system are as follows:

[0065]

[0066] Furthermore, the inertial coordinate system and the body coordinate system are specifically as follows:

[0067] The inertial coordinate system is defined as follows: The origin of the inertial coordinate system O G Take a point in the inertial space. Axis O G X G and axis O G Y G In the horizontal plane, axis O G X G Pointing due north, axis O G Y G Pointing due east, axis O G Z G Satisfy the right-hand rule and plumb downward;

[0068] The definition of the body coordinate system is: the body coordinate system is fixed to the aircraft, and the coordinate origin is at the center of mass of the aircraft O T Axis OT X T Located in the symmetry plane of the aircraft, parallel to the fuselage axis and pointing forward. T Y T Perpendicular to the aircraft's symmetry plane (i.e. O T X T Z T plane), pointing to the right; axis O T Z T Located in the symmetry plane of the aircraft, perpendicular to X T The axis points downward toward the belly of the vehicle.

[0069] Furthermore, the Markov decision module in step 2 establishes a Markov decision process based on the relative motion model of the carrier aircraft, the decoy, and the incoming missile, including the state space, the action space, the state transition equation, and the reward function, as follows:

[0070] The state space is:

[0071] S=[R,q ε ,q β ] (10)

[0072] Where: R is the relative distance between the missile and the carrier aircraft, q ε is the elevation angle of sight, q β Line of sight azimuth;

[0073] The aircraft carries a certain number of decoys. At each moment, it can choose to scatter or not scatter each decoy, adjust its posture, or choose to maneuver. Therefore, the designed action space is:

[0074]

[0075] Where: K1 is the instruction for the carrier aircraft to release the bait, ψ t , γ t is the aircraft attitude adjustment instruction, K2 is the aircraft maneuver instruction, K1, K2∈(-1, 0, 1). ;

[0076] The state transfer equation is:

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084]

[0085] The reward function is:

[0086]

[0087]

[0088]

[0089] Where n is the number of remaining baits, t1 is the remaining maneuvering time, and t2 is the moment of releasing the bait.

[0090] Furthermore, the neural network creation module in step 3 builds a neural network based on a deep reinforcement learning algorithm, as follows:

[0091] Step 3.1: Build a valuation neural network to update the evaluation of each state-action pair based on the reward information at that moment. Input the current and next state and output the corresponding state-action pair evaluation value.

[0092] Step 3.2: Build a strategy neural network and update the incoming missile countermeasure strategy based on the valuation neural network, so that each selected incoming missile countermeasure strategy always moves in the direction of the evaluation. Input the current state of the environment, including the relative distance between the missile and the carrier aircraft, the elevation angle of the line of sight, and the azimuth of the line of sight parameters, and output the strategy that the carrier aircraft should adopt;

[0093] Step 3.3: Design a loss function based on the rewards from the environment to update the valuation neural network and the policy neural network.

[0094] This method adopts the classic Actor-Critic architecture in deep reinforcement learning. Its basic network structure is as follows Figure 2 As shown in Figure 2, Actor-Critic combines the Policy Gradient (Actor) and Function Approximation (Critic) methods. After the state is input into the neural network, the parameters are updated. The Actor network outputs the action, i.e., the action probability; the Critic outputs the calculated Q-value, i.e., the TD-error.

[0095] The neural network includes a valuation neural network and a strategy neural network, combined with Figure 3 、 Figure 4The strategy neural network and the valuation neural network are both four-layer neural networks, with one input layer, two hidden layers, and the output layer. Relu is used as the activation function, the hidden layer contains 128 neurons, and the discount factor is set to 0.99.

[0096] Furthermore, the training module in step 4 trains the neural network based on the relative motion model of the carrier aircraft, decoy, and incoming missile, the state space, the action space, the state transition equation, and the reward function to obtain a trained neural network model, as follows:

[0097] Step 4.1, initialize the policy neural network parameters and the valuation neural network parameters;

[0098] Step 4.2: Initialize the state space to obtain the current state s t ;

[0099] Step 4.3: The relative motion model of the carrier aircraft, decoy, and incoming missile selects behavior a based on the action space according to the strategy output by the strategy neural network. t , execute the state transfer equations (1) to (9) to obtain the next state s t+1 , get the reward r according to the reward function t , calculate TD-error, and convert the experience (s t ,a t ,r t ,s t+1 ) is stored in the memory pool;

[0100] Step 4.4: Update the parameters of the policy neural network and the valuation neural network using the gradient descent method based on the loss function of the DDPG algorithm.

[0101] Step 4.5: The policy neural network outputs the new policy;

[0102] Step 4.6: Repeat steps 4.2 to 4.5 for more than 2000 times to complete the training of the neural network model and save the trained neural network model.

[0103] Step 5: The simulation verification module calls the trained neural network model for simulation verification;

[0104] Step 6: The decision module generates an interference delivery plan and an aircraft maneuvering plan based on the neural network model verified by simulation to defend against the incoming missile.

[0105] The convergence results of the simulated reward function are as follows Figure 5 As shown. Figure 5 It can be seen that the reward function has converged. The relative motion trajectories of the carrier, decoy and incoming missile are as follows: Figure 6 shown. Figure 7 The figure shows the relative distance change curve between the carrier aircraft and the missile. Figure 8 The figure shows the change in thrust of the aircraft's maneuvering device. The simulation results show that the aircraft has completed intelligent defense decisions against incoming missiles and output an effective and feasible evasion strategy. The output strategy of the aircraft after training is shown in Table 2:

[0106] Table 2 Output strategy of the carrier aircraft after training

[0107]

[0108] from Figure 6 、 7 As can be seen from Figures 8 and Table 2, the present invention, based on the deep reinforcement learning (DDPG) algorithm, designs a deep reinforcement learning program for self-learning defense decision-making against incoming missiles. This program uses a neural network to fit the mapping relationship between the environment and the agent's behavior and trains it, enabling the carrier aircraft to make autonomous decisions using the trained neural network. Furthermore, the present invention studies and establishes dynamic models and equations of motion for the carrier aircraft, decoy, and incoming missiles. It applies deep reinforcement learning and other methods to design and train a self-learning defense decision-making model against incoming missiles, enabling intelligent decision-making for the aircraft and significantly improving its autonomous and adaptive capabilities for typical scenarios.

Claims

1. A method for defending against incoming missiles based on deep reinforcement learning, characterized in that: The method uses an incoming missile defense countermeasure system based on deep reinforcement learning, including a model building module, a Markov decision module, a neural network creation module, a training module, a simulation verification module, and a decision module. The method steps are as follows: Step 1: The model building module builds the relative motion model of the carrier aircraft, decoy and incoming missile; Step 2: The Markov decision module establishes a Markov decision process based on the relative motion model of the carrier aircraft, decoy, and incoming missile, including the state space, action space, state transition equation, and reward function; Step 3: The neural network creation module builds a neural network based on the deep reinforcement learning algorithm; Step 4: The training module trains the neural network based on the relative motion model of the carrier aircraft, decoy and incoming missile, the state space, the action space, the state transition equation and the reward function to obtain a trained neural network model; Step 5: The simulation verification module calls the trained neural network model for simulation verification; Step 6: The decision module generates a jammer delivery plan and aircraft maneuvering plan based on the neural network model verified by simulation to defend against the incoming missile. The relative motion model of the carrier aircraft, decoy, and incoming missile described in step 1 is as follows: The dynamic equation of the incoming missile in the inertial coordinate system is: Where: x m ,y m , z m is the missile's position vector; v x_m , v y_m , v z_m is the velocity vector of the missile; P x , P y , P z is the thrust vector of the missile; m is the mass of the missile; M is the mass of the earth; G is the gravitational constant; r m is the distance of the missile relative to the origin; The motion equation of the carrier in the inertial coordinate system is: In the formula, (x t ,y t , z t ) is the coordinate of the carrier in the inertial coordinate system, v t ,θ t 、φ t is the speed, track pitch angle and track deflection angle of the carrier aircraft, n tx The longitudinal control overload of the target, n ty 、n tz They are the turning control overloads in the target yaw and pitch directions respectively; Assume that the position vector of the carrier aircraft relative to the missile is r, and the velocity vector is V. In the inertial coordinate system, (r, V, q ε ,q β ), then the relative motion equation between the missile and the carrier aircraft is: Where r x =x t -x m , r y =y t -y m , r z =z t -z m ; V x =v x_t -v x_m , V y =v y_t -v y_m , V z =v z_t -v z_m ;q β is the sight azimuth; q ε It is the sight height angle; Derivative of Equation (4) with respect to time yields: The motion equation of the decoy bomb in the inertial coordinate system is: In the formula, (x d ,y d , z d ) is the coordinate of the carrier in the inertial coordinate system, v d ,θ d 、φ d is the speed, track pitch angle and track deflection angle of the carrier aircraft; Assume that the missile adopts proportional guidance to attack the carrier aircraft. The guidance instructions are: Where k y 、k z is the proportional guidance coefficient; Indicates the rate of change of the visual range between the missile and the carrier aircraft; They represent the rate of change of sight elevation angle and azimuth angle respectively; The maneuvering device of the carrier aircraft is assumed to generate a fixed thrust. Due to fuel limitations, the number of times the maneuvering device can operate is limited. The thrust transition process during the power-on and power-off of the maneuvering device is ignored. The thrust is 0 when the device is turned off, and the thrust is P when the device is turned on. The maneuvering instruction of the carrier aircraft is K, where K∈{-1,0,1}. The projected components of the force generated by the maneuvering device on each axis in the aircraft coordinate system are: The inertial coordinate system and the body coordinate system are as follows: The inertial coordinate system is defined as follows: The origin of the inertial coordinate system O G Take a point in the inertial space; axis O G X G and axis O G Y G In the horizontal plane, axis O G X G Pointing due north, axis O G Y G Pointing due east, axis O G Z G Satisfy the right-hand rule and plumb downward; The definition of the body coordinate system is: the body coordinate system is fixed to the aircraft, and the coordinate origin is at the center of mass of the aircraft O T ; Axis O T X T Located in the symmetry plane of the aircraft, parallel to the fuselage axis and pointing forward; axis O T Y T Perpendicular to the aircraft's symmetry plane, i.e. O T X T Z T plane, pointing to the right; axis O T Z T Located in the symmetry plane of the aircraft, perpendicular to X T The axis points downward toward the belly of the vehicle; Based on the simulation model described in step 2, a Markov decision process is established, including the state space, action space, state transition equation, and reward function, as follows: The state space is: S=[R,q ε ,q β ] (10) Where: R is the relative distance between the missile and the carrier aircraft, q ε is the elevation angle of sight, q β Line of sight azimuth; The aircraft carries a set number of decoys. At each moment, it can choose to scatter or not scatter each decoy, make posture adjustments, or choose to maneuver. Therefore, the designed action space A is: Where: K1 is the instruction for the carrier aircraft to release the bait, ψ t , γ t is the aircraft attitude adjustment instruction, K2 is the aircraft maneuver instruction, K1, K2∈(-1, 0, 1); The state transfer equation is: The reward function is: Where n is the number of remaining baits, t1 is the remaining maneuvering time, and t2 is the moment of releasing the bait.

2. The incoming missile defense countermeasure method based on deep reinforcement learning according to claim 1 is characterized in that: The deep reinforcement learning algorithm described in step 3 is specifically the DDPG algorithm based on the Actor-Critic architecture.

3. The incoming missile defense countermeasure method based on deep reinforcement learning according to claim 2 is characterized in that: The neural network in step 3 includes the valuation neural network and the policy neural network; The strategy neural network and the valuation neural network are both four-layer neural networks, with one input layer, two hidden layers, and the output layer. Relu is used as the activation function, the hidden layer contains 128 neurons, and the discount factor is set to 0.

99.

4. The incoming missile defense countermeasure method based on deep reinforcement learning according to claim 3 is characterized in that: The neural network creation module described in step 3 builds a neural network based on the deep reinforcement learning algorithm, as follows: Step 3.1: Build a valuation neural network to update the evaluation of each state-action pair based on the reward information at that moment. Input the current and next state and output the corresponding state-action pair evaluation value. Step 3.2: The strategy neural network updates the incoming missile countermeasure strategy based on the valuation neural network, ensuring that the selected incoming missile countermeasure strategy always moves in the direction of the evaluation. The input is the current state of the environment, including the relative distance between the missile and the carrier aircraft, the elevation angle of the line of sight, and the azimuth of the line of sight, and the output is the strategy that the carrier aircraft should adopt. Step 3.3: Design a loss function based on the rewards from the environment to update the valuation neural network and the policy neural network.

5. The incoming missile defense countermeasure method based on deep reinforcement learning according to claim 4 is characterized in that: In step 4, the neural network is trained based on the relative motion model of the carrier aircraft, decoy, and incoming missile, the state space, the action space, the state transition equation, and the reward function to obtain a trained neural network model, as follows: Step 4.1, initialize the policy neural network parameters and the valuation neural network parameters; Step 4.2: Initialize the state space to obtain the current state s t ; Step 4.3: The relative motion model of the carrier aircraft, decoy, and incoming missile selects behavior a based on the action space according to the strategy output by the strategy neural network. t , execute the state transfer equations (1) to (9) to obtain the next state s t+1 , get the reward r according to the reward function t , calculate TD-error, and convert the experience (s t ,a t ,r t ,s t+1 ) is stored in the memory pool; Step 4.4: Update the parameters of the policy neural network and the valuation neural network using the gradient descent method based on the loss function of the DDPG algorithm. Step 4.5: The policy neural network outputs the new policy; Step 4.6: Repeat steps 4.2 to 4.5 for more than 2000 times to complete the training of the neural network model and save the trained neural network model.

6. The incoming missile defense countermeasure method based on deep reinforcement learning according to claim 5 is characterized in that: In step 6, based on the neural network model verified by simulation, an interference delivery plan and an aircraft maneuvering plan are generated to defend against the incoming missile, as follows: After simulation verification, the neural network model outputs the number, timing, angle, and maneuvering time of dropping decoys. The carrier aircraft adopts decoy dropping, attitude adjustment, and maneuvering adjustment strategies based on these control quantities to avoid incoming missiles.