An intelligent game method for attacking active defense targets based on deep reinforcement learning

By constructing a three-body game adversarial scenario using a deep reinforcement learning-based intelligent game approach and introducing adaptive step size sparsification sampling and model prior knowledge, the interception problem of traditional guidance methods when facing active defense targets is solved, and intelligent decision-making and precision strike of offensive missiles in complex environments are realized.

CN122449904APending Publication Date: 2026-07-24HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN ENG UNIV
Filing Date
2026-04-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional guidance methods lack predictive and game-theoretic capabilities when facing intelligent targets with active defense capabilities. They are unable to effectively counter the target's defense strategies, are easily intercepted, and heavily rely on accurate target motion models. Their performance degrades when information is uncertain or the model is mismatched.

Method used

By employing a deep reinforcement learning-based intelligent game theory approach, a three-body game adversarial scenario is constructed. Adaptive step-size sparsification sampling, model prior knowledge, and failure scenario compensation training are introduced to generate highly adaptive guidance commands, enabling intelligent decision-making for offensive missiles in complex environments.

Benefits of technology

It has improved the combat effectiveness and reliability of the guidance and control system, increased the mission success rate in diverse combat environments, and enhanced the ability to accurately strike targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122449904A_ABST
    Figure CN122449904A_ABST
Patent Text Reader

Abstract

The application relates to an intelligent game method for attacking an active defense target based on deep reinforcement learning, comprising the following steps: constructing a three-body game confrontation scene containing an attacking missile, a target and a defense missile according to a dynamic model; modeling a three-body game process as a Markov decision process, defining an observation space, an action space and a reward function; training the attacking missile by using a deep reinforcement learning algorithm, generating an intelligent game guidance law suitable for the three-body game confrontation scene, so that the attacking missile can accurately attack the target while avoiding interception by the defense missile, wherein an adaptive step thinning sampling mechanism, a model prior knowledge fusion mechanism and a failure scene compensation training mechanism are introduced in the training process. The application can cope with various target maneuvers and defense strategies, successfully avoid interception and accurately hit the target, and has engineering application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aerospace guidance and control technology, and in particular to an intelligent game theory method for attacking actively defended targets based on deep reinforcement learning. Background Technology

[0002] Traditional guidance laws, such as proportional navigation (PN), typically assume that the target is static or moving at a constant velocity in a straight line. When facing an actively defended target, the target will employ a series of complex tactical maneuvers, such as "evasion-launching an interceptor missile-evading again." Traditional guidance methods are essentially passive responses to the target's movement, lacking predictive and game-theoretic capabilities, and unable to understand the enemy's defensive intentions. They often fall into a passive "you maneuver, I track" dilemma, unable to effectively counter the target's active defense strategy. When the target launches defensive missiles for coordinated defense, the attacking missile needs to determine which task is more urgent—evasion or penetration—within a very short time and plan the optimal evasion / penetration strategy. However, traditional guidance systems lack dynamic, real-time threat assessment and are easily intercepted by defensive missiles. Some guidance methods may pre-set some maneuvering evasion strategies; however, they cannot adaptively generate maneuvers based on the real-time situation of the battlefield, lacking intelligence and innovation. Once these maneuvering evasion strategies are recognized by the target, their evasion effectiveness will be greatly reduced. Furthermore, traditional guidance methods rely heavily on accurate target motion models and measurement information. When information is uncertain or the model is mismatched, the performance will drop sharply, which can easily lead to guidance failure.

[0003] It is evident that traditional guidance methods, with their inherent deterministic and passive characteristics, are no longer adequate for the demands of modern, highly intelligent, and adversarial air combat environments when facing intelligent targets with active defense capabilities. Therefore, this invention proposes an intelligent game-theoretic method for attacking actively defended targets based on deep reinforcement learning. Summary of the Invention

[0004] The purpose of this invention is to provide an intelligent game theory method for attacking actively defended targets based on deep reinforcement learning, which overcomes the challenges of complex adversarial challenges. By integrating adaptive step size sparsified sampling, model prior knowledge, and failure scenario compensation training, the missile can make intelligent decisions based on the real-time battlefield situation when facing actively defended targets, generate smooth and highly adaptable guidance commands, break through target defenses and accurately strike targets, improve the mission success rate in diverse combat environments, and enhance the combat effectiveness and reliability of the entire guidance and control system.

[0005] To achieve the above objectives, the present invention provides the following solution: A deep reinforcement learning-based intelligent game theory method for attacking and actively defending targets includes: A three-body game-like adversarial scenario involving offensive missiles, targets, and defensive missiles is constructed based on a dynamic model. The three-body game process is modeled as a Markov decision process, and the observation space, action space and reward function are defined. The offensive missile is trained using a deep reinforcement learning algorithm to generate an intelligent game guidance law suitable for the three-body game confrontation scenario, enabling the offensive missile to accurately strike the target while evading the interception of the defensive missile. In the training process, an adaptive step size sparse sampling mechanism, a model prior knowledge fusion mechanism, and a failure scenario compensation training mechanism are introduced.

[0006] Optionally, constructing a three-body game-like adversarial scenario involving offensive missiles, targets, and defensive missiles based on a dynamic model includes: The air defense confrontation scenario, which includes offensive missiles, targets, and defensive missiles, is modeled as a geometric representation in an inertial coordinate system; Based on the geometric representation, the flight state and combat process of each aircraft are defined, the dynamic model of each aircraft is constructed, and the model is linearized and reduced in order to complete the construction of the three-body game confrontation scenario.

[0007] Optionally, the flight state describes the relationship between velocity tilt angle, flight speed, and actual acceleration; the engagement process includes a first engagement process between the offensive missile and the target and a second engagement process between the offensive missile and the defensive missile; the dynamic model adopts a first-order delayed dynamic model.

[0008] Optionally, model linearization and order reduction include: Based on the characteristics of linear time-invariant systems, the concept of zero-control miss quantity is introduced, and the performance index of the linear model is converted into the expression of zero-control miss quantity. The performance index adopts the form of linear quadratic form, which takes into account both miss quantity and energy consumption. By solving the state transition matrix through terminal projection transformation and inverse Laplace transform, the dynamic differential equation of the zero-control miss quantity is derived, thus completing the model order reduction.

[0009] Optionally, the observation space includes the distance, approach speed, and line-of-sight angular rate between the offensive missile and the target, as well as the distance, approach speed, and line-of-sight angular rate between the offensive missile and the defensive missile; the action space is the guidance command of the offensive missile; and the reward function integrates the miss distance, evasion behavior, and energy consumption design.

[0010] Optionally, the guidance command in the action space includes a first component and a second component, wherein the first component corresponds to the engagement process between the offensive missile and the target and adopts proportional guidance; the second component corresponds to the engagement process between the offensive missile and the defensive missile and adopts bias acceleration.

[0011] Optionally, the reward function includes intermediate process rewards and terminal rewards. The intermediate process rewards adopt an exponential function, which gives positive rewards for the behavior of the offensive missile approaching the target and negative rewards for the behavior of the defensive missile approaching the offensive missile. The terminal rewards adopt a stepped sparse reward, which gives a high positive reward when the miss distance between the offensive missile and the target is less than the kill radius and a high negative reward when the miss distance between the offensive missile and the defensive missile is less than the interception radius.

[0012] Optionally, the adaptive step size sparsification sampling mechanism is used to dynamically adjust the simulation step size according to the changes in the aircraft state, and to optimize the experience samples in combination with the data forgetting mechanism; the model prior knowledge fusion mechanism is used to introduce physical constraints of the guidance law into the state variables, action variables and reward functions; the failure scenario compensation training mechanism is used to dynamically adjust the distribution of training samples based on the PPO algorithm.

[0013] The beneficial effects of this invention are as follows: This invention proposes an intelligent game theory method for attacking and defending targets based on deep reinforcement learning. Significant technical improvements are achieved by integrating adaptive step-size sparsified sampling, model prior knowledge, and failure scenario compensation training. Adaptive step-size sparsified sampling improves sample efficiency, allowing the agent to acquire more rounds of data at the same number of rounds, resulting in faster convergence of the learning curve and higher cumulative rewards. Integrating model prior knowledge compresses the policy search space, reduces training difficulty, smooths guidance commands, and keeps misses within the kill radius, outperforming traditional guidance laws. Compensation training, by dynamically adjusting the sampling weights of failure scenarios, significantly improves the interception success rate and is less affected by target acceleration measurement errors, demonstrating significantly enhanced robustness and stability. This guidance law can cope with various target maneuvers and defense strategies, successfully evading interception and achieving accurate hits, possessing potential for engineering applications. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of an intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning, according to an embodiment of the present invention.

[0016] Figure 2 This is a schematic diagram of a three-body game air combat scenario according to an embodiment of the present invention; Figure 3 This is a geometrical diagram of a three-body game battle according to an embodiment of the present invention; Figure 4This is a block diagram of a proportional guidance system according to an embodiment of the present invention; Figure 5 This is a block diagram of a proportional guidance and accompaniment system according to an embodiment of the present invention; Figure 6 This is a comparison chart of the training and learning curves of the agent under adaptive step size sparse sampling according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the learning curve of the three-body game guidance law of the offensive missile according to an embodiment of the present invention. Figure 8 This is a schematic diagram of the median off-target distance as a function of training steps in an embodiment of the present invention. Figure 9 This is a simulation verification diagram of the three-body game guidance law for offensive missiles according to an embodiment of the present invention, wherein (a) is the combat trajectory, (b) is the velocity of each aircraft, (c) is the zero-control miss distance of the two sets of combat processes, and (d) is the command acceleration of the offensive missile and the defensive missile. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] This embodiment provides an intelligent game theory method for attacking actively defended targets based on deep reinforcement learning, including: A three-body game-like adversarial scenario involving offensive missiles, targets, and defensive missiles is constructed based on a dynamic model. The three-body game process is modeled as a Markov decision process, and the observation space, action space and reward function are defined. The offensive missile is trained using a deep reinforcement learning algorithm to generate an intelligent game guidance law suitable for the three-body game confrontation scenario, enabling the offensive missile to accurately strike the target while evading the interception of the defensive missile. In the training process, an adaptive step size sparse sampling mechanism, a model prior knowledge fusion mechanism, and a failure scenario compensation training mechanism are introduced.

[0020] Specifically, such as Figure 1As shown, this embodiment first clarifies the three-body game adversarial scenario consisting of offensive missiles, targets, and defensive missiles through dynamic environment modeling, establishes dynamic equations, and performs model linearization and order reduction to lay the model foundation for subsequent training. Next, it defines a Markov decision process, and designs an observation space containing key physical quantities, action quantities that meet engineering constraints, and a reward function to guide learning, based on the model's prior knowledge, to build a reinforcement learning framework. Subsequently, through simulation verification, it sets hyperparameters and environmental parameters, and improves the training process of the agent based on deep reinforcement learning by adopting adaptive step size thinning sampling and error compensation training under failure scenarios. The training effect is evaluated by combining indicators such as learning curve and miss rate. Finally, through typical combat scenarios, it verifies the performance of the guidance law under different strategy combinations, confirming that it can guide offensive missiles to accurately strike targets while evading interception by defensive missiles, thereby improving combat effectiveness and reliability in complex adversarial environments.

[0021] In actual deployment, this guidance law receives a series of real-time observation data as input. These inputs correspond to the "observation space" defined during the training phase, including relative position information and relative velocity information. Relative position information refers to the relative distance and azimuth between the attacking missile and the target, such as the position components in the missile-target line-of-sight coordinate system, as well as the relative distance and azimuth between the attacking and defending missiles. Relative velocity information mainly includes the relative velocity between the attacking and target, and the relative velocity between the attacking and defending missiles.

[0022] The guidance law, based on the received real-time input, performs calculations through its internal deep neural network and outputs control commands. These commands correspond to the action quantities defined during the training phase, i.e., maneuver overload commands. The guidance law calculates the magnitude and direction of the normal overload that the attacking missile should apply at the current moment. This command is directly transmitted to the missile's autopilot, which adjusts the control surface deflection accordingly to generate the required maneuver overload, altering the trajectory to achieve the dual objectives of evading interception and striking the target.

[0023] Specifically, it includes the following: 1. Dynamic environment modeling; Dynamic environment modeling defines the core tasks that the agent needs to handle, ranging from intercepting targets with simple trajectories to dealing with complex situations involving high maneuverability and strong adversarial forces. This sets clear objective boundaries for strategy training, ensuring that the learning process revolves around actual combat needs. The specific process in this embodiment is as follows: Figure 1 As shown, the main steps are as follows: Step 1: Define the combat scenario; This embodiment mainly focuses on combat scenarios such as Figure 2As shown, this is a typical air defense combat scenario, with the dashed box indicating the operational scenario primarily focused on in this embodiment. Generally, the target aircraft has limited maneuverability, making it difficult to evade missiles using traditional maneuvering strategies. Therefore, related research has begun to focus on enabling targets to adopt active defense strategies to enhance their battlefield survivability. For example, in this scenario, Wingman-I can launch a defensive missile to intercept an enemy attack missile. Furthermore, unmanned wingmen can also act as loyal wingmen, sacrificing themselves if necessary to protect the human pilot's aircraft. In this case, the pilot's aircraft is the protected target, while Wingman-II plays the role of the defender.

[0024] This embodiment will design an intelligent game-theoretic guidance law for attacking actively defended targets from the perspective of an offensive missile. The offensive missile needs to evade interception by the defensive missile, essentially a penetration problem, while simultaneously attacking a mobile target. In the three-body game problem, because it must simultaneously consider both mobile penetration and precision strikes, the offensive missile facing the active defense target faces a much greater challenge than the target / defense missile. The offensive missile must balance both aspects in a complex game, avoiding losing the strike window due to excessive evasion or being intercepted by the defensive missile due to haste in attacking, ultimately achieving an optimal strategy that balances mobile penetration and precision strikes.

[0025] Step 2: Modeling the Three-Body Problem Battle Scenario; based on Figure 2 The battle scenes depicted Figure 3 The battle is distilled into an inertial coordinate system. Geometric representation in [the language].

[0026] The model includes three participants: offensive missiles. (Missile) The target that the offensive missile is intended to attack. (Target), and defensive missiles protecting the target. (Defender). The flight state of each aircraft is determined by its velocity angle. Flight speed and actual acceleration Descriptions, each using subscripts , and This indicates the offensive missile, the target, and the defensive missile. The engagement process is divided into two related but independent parts: the attack of the offensive missile on the target (…). ) and defensive missiles intercepting offensive missiles to protect the target ( In these kinds of battles, Indicates the distance between each aircraft. Indicates line of sight. Represents the line-of-sight angle. Initial conditions at the start of combat are indicated by the subscript 0.

[0027] like Figure 3 As shown, the nonlinear kinematic equations for the engagement between offensive missile and target, and between offensive missile and defensive missile, can be written as: ; ; in Indicates enemy attack missiles. This refers to our carrier aircraft (i.e., the target in the Three-Body Problem) or defensive missiles. .

[0028] The rate of change of the velocity tilt angle can be expressed as: ; Furthermore, it is assumed that the dynamic model of each aircraft can be expressed as a linear system of arbitrary order: ; in Represents the internal state quantities of the aircraft. The system output represents the actual acceleration of the aircraft in the vertical velocity direction. The input to the system represents the corresponding control commands under different modeling methods. Equation (4) represents the modeling methods with different granularities. Traditional methods generally assume a zero-order ideal dynamic model, in which case... This means that the system input is a command acceleration, and the aircraft's acceleration response has no delay, thus facilitating the derivation of the subsequent analytical guidance law.

[0029] Step 3: Model linearization and order reduction; Although many studies have formally preserved the arbitrary-order dynamic characteristics shown in Equation (4), it is worth noting that the final solution often simplifies to an idealized dynamic model in order to derive analytical expressions or introduce integral terms while preserving dynamic delays. However, idealized dynamics assumes an instantaneous response, which does not reflect the characteristics of the real system and may lead to performance degradation in practical applications. The strategy of introducing integral terms, on the other hand, puts pressure on the real-time calculation of the onboard computer.

[0030] Given these limitations, this embodiment employs first-order delay dynamics from the outset. First-order dynamics agrees well with the response characteristics of the actual system, achieving a good balance between representativeness and simplicity. Specifically, in equation (4), first-order delay dynamics means... ,and ,in Represents the time constant. Select state variables. The equation of motion for an offensive missile attacking a target or a defensive missile can be expressed as: ; The system matrix and control matrix are as follows: ; The system's control input is the command acceleration of each aircraft: ; The remaining flight time of the attack missile during its engagement with the target is used as... This indicates the remaining flight time of the engagement between offensive and defensive missiles. In other words, under the online assumption, the engagement flight time can be defined as: ; ; Therefore, the remaining flight times can be calculated as follows: .

[0031] Assuming the engagement between the offensive and defensive missiles occurs before the engagement between the offensive missile and the target, that is, on the timeline, and Satisfying the time interval between the two groups of combat This is because in terminal guidance engagements, the linearization assumption can be used to assume that the target... and The estimation has high accuracy, therefore and It exhibits a linear change. Once the offensive missile can hit or miss the target before the defensive missile can intercept it, it means that the entire engagement process is over, and the defensive missile no longer needs to participate in the battle.

[0032] This situation can be described as a three-body game problem where opposing sides have different optimal performance indices. The performance index used combines miss distance and energy consumption, presented in a linear-quadratic (LQ) form, thus achieving a balance between the two. Compared to norm-based performance indices, LQ-based performance indices have the advantage of avoiding frequent control input switching or even chattering, while also resulting in lower energy consumption and a smoother control trajectory. This also means that LQ-based performance indices ensure that control commands do not exceed the aircraft's maneuverability limits through a soft constraint. The specific form of the performance index under the linearized model is as follows: ; in, and It is a non-negative weighting coefficient related to the miss distance. and These respectively reflect the maneuverability of the target and the defensive missile relative to the offensive missile. When the target's maneuverability relative to the offensive missile is relatively weak, Take the larger value, and This indicates that the target is immobile. When the defensive missile is more mobile than the offensive missile, Take the smaller value, and This indicates that the defensive missile has an absolute advantage in terms of high mobility.

[0033] Nevertheless, directly solving the analytical solution to the aforementioned linear optimal control problem remains quite difficult; therefore, reducing the order of the model is crucial. In guidance analysis, the zero-effort-miss (ZEM) is a widely used concept, representing the miss distance if the interceptor performs no maneuvers—that is, the control command is zero from the current moment until the end of engagement. From the perspective of signal processing and linear system analysis, the zero-effort-miss is essentially equivalent to the zero-input response of a linear system. Using the terminal projection transformation of a linear system, the zero-effort-miss can be... and Represented as: ; in It is a constant coefficient matrix. Let be the state transition matrix.

[0034] Due to the system matrix Since the state transition matrix is ​​a linear time-invariant matrix, it can be directly obtained through the inverse Laplace transform. ; Therefore, we can obtain... and The parsing expression: ; ; in: ; ; However, and Still related to linearized system variables Coupled. Furthermore, for the variables of interest (i.e. and To solve these problems, their dynamic differential equations need to be obtained so that variational methods can be applied within the framework of differential games. Therefore, by differentiating the equations and expanding the expression for the state transition matrix, we can obtain... and The derivatives with respect to time are as follows: ; in: ; Accordingly, the performance metrics of the linearization problem can be transformed into a form expressed in terms of zero-control miss quantities: ; Furthermore, based on differential game theory, the closed-loop expressions for the optimal control of the target / defense missile and the attack missile can be derived using the variational method as follows: ; Each term of each control variable can be written in a form similar to proportional guidance, forming a closed-loop feedback from the zero-control miss distance. Since this form of guidance law is derived based on performance indices of linear quadratic form, this type of guidance law can be called a linear quadratic differential game guidance law. This theoretical knowledge constitutes a priori understanding of the model, and its reasonable integration into the construction of Markov decision processes can effectively improve the efficiency of agent training.

[0035] 2. Definition of Markov Decision Process; To achieve adaptive learning of interception strategies and robust control in highly dynamic environments, this embodiment models the missile defense guidance process as a Markov Decision Process (MDP). This model can effectively represent the policy evolution process of an agent under limited information and dynamic interaction, and is the foundation for building a deep reinforcement learning framework. The specific steps are as follows: Step 1: Observation space design; The observation space focuses on selecting physical quantities that are critical to the interception decision and are readily available. The agent's observations are defined as follows: ; That is, the observations simultaneously considered two sets of combat processes. This indicates that the offensive missile struck the target. (This indicates that offensive missiles evade defensive missiles), and all of them consider distance, approach speed, and line-of-sight angular rate as strategy inputs.

[0036] From the perspective of reinforcement learning, adding reasonable observations helps improve the accuracy of estimating the state value function through partial observations, which in turn allows for a more accurate assessment of the current situation's favorableness to the missile, thereby better optimizing the agent's strategy.

[0037] Step 2: Setting the amount of exercise; To further improve training efficiency and stability, a thorough analysis of the positive and negative values ​​of the equivalent guidance ratio is first conducted. From a control perspective, proportional guidance can be viewed as a feedback control system whose goal is to adjust the zero-control miss distance to zero. Therefore, only a negative feedback system can avoid divergence, such as... Figure 4As shown. The simplest step maneuver is often used to analyze the performance of guidance systems; the conclusion given in the literature is that the miss distance tends to zero as the flight time increases.

[0038] like Figure 5 The accompanying system of the negative feedback guidance system is given. For convenience, it is referred to as... replace And from the convolution integral, we obtain: ; Transforming from the time domain to the frequency domain, we obtain: ; Next, integrating the previous equation, we get: ; When the guidance system is a first-order system, it means: ; The final expression for the miss distance of the negative feedback guidance system in the frequency domain can be obtained as follows: ; Applying the final value theorem, it can be found that as flight time increases, the miss distance will tend to zero: ; This means the guidance system is stable and controllable. Similarly, the miss distance expression for a positive feedback guidance system in the frequency domain can be found as follows: ; Applying the final value theorem again, we find that the miss distance does not converge with increasing flight time, but rather tends to infinity: ; From a control perspective, the above conclusion is obvious, as positive feedback systems are typically avoided due to their divergent characteristics. Therefore, positive feedback is never used in proportional guidance, and the guidance ratio is never set negative. However, the situation we now face is... Hope to reduce , and Hope to reduce ,and Hope to increase ,at the same time and Hope to increase Therefore, combining the characteristics of negative and positive feedback systems, a direct approach is to... , and The function is set as positive feedback, while... , and Its function is set as negative feedback.

[0039] However, on the other hand, combined with the analysis, it can be seen that because the performance indicators take energy consumption into account, and the scenario setting requires that the escapee's maneuverability must be weaker than the pursuer's (otherwise the problem is meaningless and the saddle point solution does not exist), this leads to... In Item and In The zero-control miss distance becomes meaningless when it is small (meaning a low probability of escape). Based on the above considerations, the guidance law of the offensive missile can still be broken down into two terms, corresponding to the engagement process with the target and the defensive missile respectively, but the evasion of the defensive missile is no longer used. Instead of using the proportional feedback form, it was directly changed to an offset acceleration term, that is: ; In summary, the action amount of the agent is selected as... and The former is taken as positive feedback, and the latter is taken as the range of maneuverability of the offensive missile.

[0040] Step 3: Reward function design; The reward function is a crucial factor in reinforcement learning, directly impacting the agent's learning efficiency and policy quality. Many problems can be solved by continuously adjusting the manifold of the reward function to train a feasible policy. Ideally, by fully utilizing model knowledge and optimizing the training process, agent training can be made insensitive to the hyperparameters in the reward function, thus significantly reducing the burden of hyperparameter tuning.

[0041] In research, exponential reward functions are often used to balance the requirements of the final task metrics with the guidance of rewards during the intermediate process. However, how to adjust the specific shape of the reward function often becomes an obstacle that perplexes researchers. The advantage of choosing an exponential reward function is that its gradient changes continuously and significantly, but the key is to determine where the inflection point where the exponent begins to rise significantly is selected.

[0042] From the start of engagement to miss, there is always a process where the missile and the target are constantly closing in relative distance. The reward for this process should be designed to be in the phase where the exponential function changes gradually. What is truly meaningful for strategy optimization is to explore a better strategy based on the fixed reward obtained from this initial random strategy, and then guide the better strategy with a steeper reward gradient. This design concept itself reflects the designer's understanding of the engagement process.

[0043] The final reward function is designed as a combination of a dense exponential reward in the middle and a sparse tiered reward at the terminal. ; ; ; In the formula Its function is to guide the attack missile to continuously approach the target; parameters and The values ​​were set to 0.02 and 0.01 respectively, thus reducing the bonus gained from the natural proximity of aircraft flying towards each other in the early stages of combat, and distributing more of the bonus within the zero-control miss range. This is to ensure that the offensive missiles can hit the target with high accuracy. A tiered terminal reward system is used to reward strategies that bring targets into the kill radius. , , as well as Similarly, This allows offensive missiles to remain alert to the approach of defensive missiles, parameters and They are taken as 0.05 and 0.1 respectively, which is more formally than The steeper slope means that a defensive missile only poses a threat when it gets close enough. Meanwhile, it is designed with... In , , as well as This prevents offensive missiles from being intercepted with high precision by defensive missiles.

[0044] 3. Agent training and result analysis; Simulation verification is a crucial step in evaluating the effectiveness of the algorithm. By setting training hyperparameters, environmental parameters, and combinations of target and defense missile strategies, a deep learning algorithm is used for training. The learning curve, median miss distance, and other metrics are analyzed to verify the actual effects of adaptive step-size thinning sampling, model prior knowledge fusion, and imputation training, providing data support for strategy optimization. The specific steps are as follows: Step 1: Environment status and hyperparameter settings; The hyperparameters and network structure settings involved in the reinforcement learning algorithm are shown in Table 1. The parameter settings of each aircraft involved in the training environment are shown in Table 2. The defensive missile is launched from the target, and its initial altitude has a certain degree of randomness relative to the offensive missile. At the same time, the strategy combinations that the target and the defensive missile can adopt are listed in Table 3, where PN is proportional guidance, APN is enhanced proportional guidance, and CLQDG is linear quadratic differential game guidance law.

[0045] Table 1 Table 2 Table 3 It is important to note that, according to the analysis in dynamic environment modeling, the performance index of the linear quadratic form, due to its consideration of energy consumption, necessitates that the pursuing player's maneuverability be superior to the escaping player's; otherwise, the pursuit-escape problem is meaningless. Once the defensive missile fails to intercept the attacking missile, the target's maneuvering strategy derived from the linear quadratic differential game is likely to reduce maneuverability to "save" fuel, thus lowering its survival probability. Therefore, strategies five and six in Table 3 represent the target switching to step maneuvers or serpentine maneuvers after the defensive missile misses its target, respectively. During training, the target / defensive missile combination randomly employs any of the strategies in Table 3 in different rounds. This means the attacking missile faces an opponent with diverse strategies, introducing more uncertainty and greater difficulty to the training.

[0046] Step 2: Adaptive step-size thinning sampling; In aerospace guidance and control research, the training environment for intelligent agents is typically defined by a system of differential equations, which are then solved numerically to continuously update the flight state. To ensure the trained strategy can be ported to hardware in the future, the training environment setup and solution results must be sufficiently reliable to migrate to the real environment, achieving what is known as Sim2Real. Setting a sufficiently small integration step size is crucial for ensuring reliable state updates without altering the dynamics model. However, an excessively small integration step size can lead to longer program execution times and excessively long trajectories and denser sampling data in a single cycle.

[0047] The adaptive step-size integration method dynamically adjusts the step size for the next time step based on changes in the system state and the error estimate of the current step size. A smaller step size is used when state changes drastically to ensure integration accuracy, while a larger step size is used when state changes are gradual to reduce the generation of redundant samples. This dynamic adjustment not only improves computational efficiency but also generates more representative samples, enhancing sample diversity and effectiveness. However, simply adjusting the step size is insufficient to solve this problem; therefore, a data forgetting mechanism is further introduced. After completing one round of learning, the generated trajectory samples undergo random forgetting, i.e., some sampling points are deleted to shorten the trajectory length. The aim is to maintain information diversity while reducing the storage and processing of highly similar samples, thereby increasing the capacity of the experience replay buffer to accommodate different trajectories and enhancing the effectiveness of policy updates.

[0048] Figure 6 Table 4 and its schematic diagram and pseudocode are respectively provided for this technology. Figure 6 It can be seen that the state trajectory The degree of change within a single round is not uniform. Therefore, if a fixed interval is used as the sampling step size, data redundancy may occur in the slow-changing region, while dynamic information cannot be effectively captured in the fast-changing region. The adaptive step size sampling method can effectively overcome this shortcoming. At the same time, forgetting can further shorten the trajectory length, thereby allowing the buffer pool to accommodate more data samples under different conditions.

[0049] Table 4 Step 3: Inferential training in failure scenarios; Table 5 presents the pseudocode for the imputation training process based on the PPO algorithm. PPO was chosen as the model because it uses methods such as pruning the objective function to limit the magnitude of policy updates, thus maintaining stability during the update process. Specifically, it introduces a threshold; when the ratio of the new to the old policy exceeds this threshold, the loss function is pruned to prevent over-updates.

[0050] By setting constraints on the objective function, PPO can remain within the "trust region" with each update, ensuring that policy changes are gradual rather than drastic. This characteristic makes PPO particularly suitable for combination with impaired training, avoiding situations where policy updates are too large and cause previously effective scenarios to fail. The iterative framework of impaired training adopts a closed-loop structure of "training-evaluation-adjustment," and the specific process is as follows: In the initial training phase, the agent is placed in a simulated environment for preliminary training to learn basic strategies and behavioral patterns. After the preliminary training, the agent's performance is comprehensively evaluated to identify its weaknesses in certain specific scenarios. These weak scenarios may include extreme conditions, edge cases, anomalous events, or complex tasks. These weak scenarios will be the focus of subsequent remedial training.

[0051] The core of gap-filling training lies in designing a gap-filling training set that includes these weak scenarios and continuously adjusting the training frequency of these scenarios during training. Specifically, each time the environment is reset, the distribution of the initial state is adjusted so that the proportion of training in the weak scenarios gradually increases. In this way, the agent can gain more experience in these specific scenarios, thereby improving its performance. During this process, the learning rate and exploration strategy can be adjusted in a timely manner to ensure that the agent stably consolidates its knowledge of the known strategy while maintaining open exploration of new strategies, thus promoting continuous optimization of the strategy.

[0052] As the training progresses, the agent's performance in weak scenarios gradually improves. The training is considered complete when its success rate across all randomly initialized states reaches a predetermined standard. In this way, training effectively enhances the agent's overall performance, giving it greater adaptability and stability in complex and changing environments.

[0053] Table 5 Step 4: Analysis of training results; The attack missile was trained using a deep learning algorithm, and the training results are as follows: Figure 7 and Figure 8 As shown. From Figure 7 It can be seen that the agent's policy converges after approximately 3e6-4e6 training steps. Meanwhile, in Figure 8 The curves showing the median miss distance (denoted by and ) as a function of training steps are presented. The median is unaffected by extreme values ​​(outliers) and better represents the central position of the data when the data distribution is asymmetrical. Therefore, it better represents the general level of the policy when there are diverse policy variations and a large scattered miss distance distribution. As the agent's policy training converges and stabilizes below 2m, it stabilizes around 26m, meaning that at this point the policy can evade interception by defensive missiles while hitting the target with high accuracy.

[0054] Furthermore, through supplementary training, the success rate of the agent's policy in randomly initialized scenarios increased from 92% to 98%. To further verify the effectiveness of the trained policy, simulations are conducted in two specific scenarios below.

[0055] 4. Combat scenario verification and explanation; Verification in actual combat scenarios is the final step in testing the robustness of the strategy. By constructing typical adversarial scenarios, simulating the penetration and strike process of offensive missiles, and analyzing details such as trajectory, acceleration, and zero-control miss distance, the adaptability and reliability of the guidance law in complex dynamic games are verified, confirming its potential for engineering applications, as detailed below.

[0056] In this scenario, the defensive missile uses APN to intercept the attacking missile, and the target performs serpentine maneuvers at a frequency of 1 rad / s. The simulation results are as follows: Figure 9 As shown. Figure 9 (a) shows the trajectories of each aircraft during the engagement. The defensive missile missed its target at 5.19s with a miss distance of 39.52m, allowing the offensive missile to successfully hit the target at 6.41s with a miss distance of 0.18m. The offensive missile's strategy was clearly purposeful: firstly, to maneuver away from the defensive missile's threat, and secondly, to provide favorable handover conditions for subsequent attacks when changing maneuvering directions. From Figure 9As can be seen from the acceleration curve in (b), the attacking missile rapidly increases its maneuverability after 2 seconds, and the direction of maneuver is opposite to the initial maneuver direction. This causes the defensive missile to quickly reach overload saturation around 4 seconds. This process corresponds to... Figure 9 (a) shows the first inflection point of the offensive missile's trajectory. However, as the defensive missile's overload reaches saturation, the offensive missile changes its maneuvering direction again, ultimately causing the defensive missile to fail to keep up with the offensive missile's adjustment and miss the target.

[0057] The above-mentioned battle process from Figure 9 (c) Zero-control miss distance curve and Figure 9 The command acceleration curve in (d) also reflects this. Looking at the change in zero-control miss distance, the attack missile does not rush to lock onto the interception triangle at the beginning of the engagement. Instead, it first uses significant maneuvers to evade the defensive missile, thus failing to converge to zero. Although the attack missile employs two large-G maneuvers in opposite directions while attempting to evade the defensive missile, it adjusts to near zero while the defensive missile misses, thus providing favorable conditions for subsequent attacks.

[0058] exist Figure 9 The command acceleration curve in (d) more clearly shows the process by which the offensive missile adjusts its maneuvering direction, causing the defensive missile to miss its target due to overload saturation. The offensive missile's strategy is similar to bang-bang maneuvers, both switching directions at the overload capacity boundary, causing the interceptor to miss its target. However, the difference is that the offensive missile must also consider attacking the target at the same time, so its guidance commands have more complex behavior.

[0059] The simulation examples above demonstrate that the intelligent game-theoretic guidance law for attacking and actively defending targets based on deep reinforcement learning can successfully evade interception while simultaneously hitting the target. The proposed strategy is reasonable and interpretable, and exhibits good adaptability to different scenarios. Simulation results prove the effectiveness of this embodiment.

[0060] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. An intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning, characterized in that, include: A three-body game-like adversarial scenario involving offensive missiles, targets, and defensive missiles is constructed based on a dynamic model. The three-body game process is modeled as a Markov decision process, and the observation space, action space and reward function are defined. The offensive missile is trained using a deep reinforcement learning algorithm to generate an intelligent game guidance law suitable for the three-body game confrontation scenario, enabling the offensive missile to accurately strike the target while evading the interception of the defensive missile. In the training process, an adaptive step size sparse sampling mechanism, a model prior knowledge fusion mechanism, and a failure scenario compensation training mechanism are introduced.

2. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 1, characterized in that, Based on the dynamic model, a three-body game-like adversarial scenario involving offensive missiles, targets, and defensive missiles is constructed, including: The air defense confrontation scenario, which includes offensive missiles, targets, and defensive missiles, is modeled as a geometric representation in an inertial coordinate system; Based on the geometric representation, the flight state and combat process of each aircraft are defined, the dynamic model of each aircraft is constructed, and the model is linearized and reduced in order to complete the construction of the three-body game confrontation scenario.

3. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 2, characterized in that, The flight state describes the relationship between velocity tilt angle, flight speed, and actual acceleration; the engagement process includes the first engagement between the offensive missile and the target and the second engagement between the offensive missile and the defensive missile; the dynamic model adopts a first-order delayed dynamic model.

4. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 2, characterized in that, Model linearization and order reduction include: Based on the characteristics of linear time-invariant systems, the concept of zero-control miss quantity is introduced, and the performance index of the linear model is converted into the expression of zero-control miss quantity. The performance index adopts the form of linear quadratic form, which takes into account both miss quantity and energy consumption. By solving the state transition matrix through terminal projection transformation and inverse Laplace transform, the dynamic differential equation of the zero-control miss quantity is derived, thus completing the model order reduction.

5. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 1, characterized in that, The observation space includes the distance, approach speed, and line-of-sight angular rate between the offensive missile and the target, as well as the distance, approach speed, and line-of-sight angular rate between the offensive missile and the defensive missile; the action space is the guidance command of the offensive missile; the reward function integrates the miss distance, evasion behavior, and energy consumption design.

6. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 5, characterized in that, The guidance commands in the action space include a first component and a second component. The first component corresponds to the engagement process between the offensive missile and the target and adopts proportional guidance. The second component corresponds to the engagement process between the offensive missile and the defensive missile and adopts bias acceleration.

7. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 5, characterized in that, The reward function includes intermediate process rewards and terminal rewards. The intermediate process rewards adopt an exponential function, which gives positive rewards for the behavior of offensive missiles approaching the target and negative rewards for the behavior of defensive missiles approaching offensive missiles. The terminal rewards adopt a stepped sparse reward, which gives a high positive reward when the miss distance between the offensive missile and the target is less than the kill radius and a high negative reward when the miss distance between the offensive missile and the defensive missile is less than the interception radius.

8. The intelligent game theory method for attacking and actively defending targets based on deep reinforcement learning according to claim 1, characterized in that, The adaptive step size thinning sampling mechanism is used to dynamically adjust the simulation step size according to the changes in the aircraft state, and to optimize the experience samples in combination with the data forgetting mechanism; the model prior knowledge fusion mechanism is used to introduce physical constraints of the guidance law into the state variables, action variables and reward function. The failure scenario compensation training mechanism is used to dynamically adjust the distribution of training samples based on the PPO algorithm.