Aircraft collision angle defense method and system based on reinforcement learning and sliding mode control
By combining reinforcement learning and sliding mode control methods, a three-body dynamic model is constructed and GAIL-SAC learning is carried out, which solves the safety hazards of drone intrusion interception and achieves fast and high-precision interception of incoming aircraft.
Patent Information
- Application Number
- CN202510423681.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art is difficult to effectively intercept drone invasions, especially in airport approach areas or waterways, which pose safety hazards.
Combining reinforcement learning and sliding mode control, a three-body dynamic model of incoming aircraft-target-defense aircraft is constructed, guiding laws are designed and defense models are generated through GAIL-SAC learning to achieve rapid interception of incoming aircraft.
It realizes high-precision interception of incoming aircraft in complex environments, ensures the safety of targets, and improves the interception success rate and time efficiency.
Smart Images

Figure CN119916699B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aircraft interception, and provides an aircraft collision angle defense method and system based on reinforcement learning and sliding mode control. Background Art
[0002] As the use of drones becomes more and more common, the flight of drones may be unrestricted due to drone malfunctions. If a drone invades the approach area of an aircraft or even the aircraft's route, the consequences will be disastrous. Therefore, the protection of airports is imminent.
[0003] Current approaches to UAV interception include proportional navigation guidance (PNG), optimal control, sliding mode control, and differential game methods. Proportional navigation, in which the UAV's acceleration is proportional to the angular velocity of the line of sight between the aircraft and the target, is widely used in guidance due to its simple structure.
[0004] Optimal control methods design guidance laws for UAVs to intercept targets by selecting optimal performance metrics for different performance factors. Differential game methods, based on optimal control theory, consider the optimal control problem for both the UAV and the target to achieve optimal goals for both. Sliding mode control theory has been widely used in nonlinear systems because it is unaffected by model uncertainty and external disturbances and does not require precise maneuvering information for guidance.
[0005] With the development of artificial intelligence technology, deep reinforcement learning has been used in the guidance field. The guidance problem is regarded as a Markov decision process. By setting the reward function and optimizing the strategy network, the adaptive online output of the acceleration of the UAV aircraft to intercept the target is achieved. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention provides an aircraft collision angle defense method and system based on reinforcement learning and sliding mode control to achieve rapid interception of incoming aircraft.
[0007] The present invention provides an aircraft collision angle defense method based on reinforcement learning and sliding mode control, comprising the following steps:
[0008] S1: Construct a three-body dynamics model of the attacking aircraft, the target, and the defensive aircraft;
[0009] S2: Constructing guidance law design objectives based on the three-body dynamics model;
[0010] S3: constructing a sliding mode guidance law for the defense aircraft according to the three-body dynamics model and the guidance law design goal;
[0011] S4: performing GAIL-SAC learning on the sliding mode guidance law of the defense aircraft to obtain a defense model;
[0012] S5: Inputting the position, velocity and acceleration of the attacking aircraft; the position, velocity and acceleration of the target; and the position, velocity and acceleration of the defending aircraft into the defense model to intercept the aircraft.
[0013] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S1 includes:
[0014] S11: Construct a two-dimensional kinematic model between the incoming aircraft and the target:
[0015]
[0016] in, is the relative speed between the incoming aircraft and the target, is the speed of the incoming aircraft, is the heading angle of the incoming aircraft relative to the inertial coordinate system, is the line of sight angle between the incoming aircraft and the target relative to the inertial coordinate system, is the heading angle of the target velocity relative to the inertial coordinate system, is the target speed, for The first derivative of for The second derivative of is the acceleration of the incoming aircraft, is the relative distance between the incoming aircraft and the target, is the acceleration of the target, for The first derivative of for The first derivative of
[0017] S12: Construct a two-dimensional kinematic model between the attacking aircraft and the defending aircraft:
[0018]
[0019] in, Indicates the relative speed between the attacking aircraft and the defending aircraft, It represents the line of sight angle of the line connecting the attacking aircraft and the defending aircraft relative to the inertial coordinate system, Indicates the speed of the defensive aircraft, Indicates the heading angle of the defense aircraft relative to the inertial coordinate system, for The first derivative of for The second derivative of Indicates the relative distance between the attacking aircraft and the defending aircraft, represents the acceleration of the defense aircraft, for The first derivative of
[0020] S13: The two-dimensional kinematic model between the incoming aircraft and the target and the two-dimensional kinematic model between the incoming aircraft and the defending aircraft are combined to form a three-body dynamic model of the incoming aircraft-target-defending aircraft.
[0021] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S2 includes:
[0022] S21: Design the interception conditions for the defensive aircraft to intercept the incoming aircraft:
[0023]
[0024] in, The desired sight angle for the defending aircraft;
[0025] S22: Define guidance law design objectives based on the interception conditions:
[0026]
[0027] in, To prevent the aircraft’s sight angle tracking error, for The first derivative of .
[0028] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S3 includes:
[0029] S31: Define the sliding surface:
[0030]
[0031]
[0032] in, is the designed sliding surface, is the first adjustment parameter, is the second adjustment parameter, is the first exponential parameter, is the second exponential parameter, To define the symbolic function, for The independent variable of the function, To take the absolute value, is a symbolic function;
[0033] S32: Design convergence law:
[0034]
[0035] in, is the first derivative of the sliding surface, is a gain that takes a positive value in the reaching law;
[0036] S33: Design a sliding mode guidance law for the defense aircraft based on the sliding mode surface and the reaching law:
[0037]
[0038]
[0039] in, is the scale factor of the proportional guidance.
[0040] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S31 includes: The value range of is: ; The value range of is: ; The value range of is: ; .
[0041] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S4 includes:
[0042] S41: constructing a SAC algorithm model according to the sliding mode guidance law of the defensive aircraft, and generating a strategy according to the SAC algorithm model;
[0043] S42: using the generated strategy as an expert strategy, inputting the state space and action space of the SAC algorithm model and the expert strategy into a discriminator to obtain a probability derived from the state space and action space;
[0044] S43: Update the SAC algorithm model according to the possibilities derived from the state space and the action space to obtain the defense model.
[0045] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S41 includes:
[0046] S411: Set the state space, action space and reward function of the SAC algorithm:
[0047]
[0048]
[0049]
[0050] in, is the state space, is the action space, is the reward function, is the first reward parameter, is the distance between the defending aircraft and the attacking aircraft at the initial moment, is the second reward parameter, is the third reward parameter, Indicates the distance between the incoming aircraft and the target at the initial moment, is the fourth reward parameter, is the fifth reward parameter, is the sixth reward parameter;
[0051] S412: Design the optimal strategy function of the SAC algorithm:
[0052]
[0053]
[0054] in, the strategy to adopt for the next moment; The strategy adopted at the current moment, To find the maximum function, For expectations, For the current moment, is the expected reward obtained after taking an action in the state following the strategy at the current moment, is the current state, is the action at the current moment, is the regularization coefficient, is the information entropy, The state of following the policy Take action ;
[0055] S413: Constructing the value of the objective function and value function :
[0056]
[0057]
[0058] in, The expected objective function value of the state at the current moment, The expected value obtained after taking an action in the current state of following the strategy, is the discount factor, for Always follow the strategy The expected value of a state after taking an action; To follow the strategy state Take action ;
[0059] S414: Calculate the loss function of the objective function value and the loss function of the value function :
[0060]
[0061]
[0062] in, is the reward at the current moment, To get the minimum value, is the mark number, is the expected value corrected by the loss function of the value function, is the Gaussian distribution function, is the standard normally distributed noise.
[0063] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S42 includes:
[0064] S421: and As the input of the discriminator, As a generator, it is trained to obtain the possibility from the state space and action space;
[0065] S422: According to the loss function of the discriminator Modify the training:
[0066]
[0067] in, is the input function of the discriminator.
[0068] According to an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention, step S43 includes:
[0069] S431: Activate the possibilities from the state space and action space through the activation function to obtain the automatic adjustment entropy regularization term ;
[0070] S432: Modify the reinforcement learning objective and rewrite it as:
[0071]
[0072] in, To obtain the maximum value, is the target entropy, Indicates restrictions;
[0073] S433: Loss function based on automatic adjustment of entropy regularization term Modify the training:
[0074] .
[0075] The present invention also provides an aircraft collision angle defense system based on reinforcement learning and sliding mode control, comprising: a model building module for constructing a three-body dynamics model of an attacking aircraft, a target, and a defense aircraft; constructing a guidance law design target based on the three-body dynamics model; and constructing a sliding mode guidance law for the defense aircraft based on the three-body dynamics model and the guidance law design target.
[0076] Model learning module: performing deep learning on the sliding mode guidance law of the defensive aircraft to obtain a defense model;
[0077] Aircraft interception module: The position, velocity and acceleration of the attacking aircraft; the position, velocity and acceleration of the target; and the position, velocity and acceleration of the defending aircraft are input into the defense model to intercept the aircraft.
[0078] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:
[0079] The present invention provides a method and system for aircraft collision angle defense based on reinforcement learning and sliding mode control. By combining deep reinforcement learning with traditional sliding mode control, the intelligent agent is guided to conduct non-blind exploration in the early stage. In the design of the reward function, an integer reward function is added to achieve rapid convergence of reinforcement learning, complete the active defense guidance of the defensive aircraft in the three-body confrontation, and intercept the incoming aircraft with high precision and angle constraints.
[0080] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0082] Figure 1 This is a flow chart of an aircraft collision angle defense method based on reinforcement learning and sliding mode control provided by the present invention;
[0083] Figure 2 This is a schematic diagram of GAIL-SAC learning;
[0084] Figure 3 This is a schematic diagram of the active defense trajectory of an embodiment of the present invention;
[0085] Figure 4 is a graph showing changes in the sight angle between the defending aircraft and the attacking aircraft over time according to an embodiment of the present invention;
[0086] Figure 5 This is a structural block diagram of an aircraft collision angle defense device based on reinforcement learning and sliding mode control provided by the present invention;
[0087] Reference numerals:
[0088] 101. Model building module; 102. Model learning module; 103. Aircraft interception module. DETAILED DESCRIPTION
[0089] To make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0090] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0091] The following combination Figures 1 to 5 The present invention is described.
[0092] Example
[0093] like Figure 1 As shown, Figure 1 The present invention provides a flowchart of an aircraft collision angle defense method based on reinforcement learning and sliding mode control, which includes the following steps:
[0094] S1: Construct a three-body dynamics model of the attacking aircraft, the target, and the defensive aircraft;
[0095] S2: Constructing guidance law design objectives based on the three-body dynamics model;
[0096] S3: constructing a sliding mode guidance law for the defense aircraft according to the three-body dynamics model and the guidance law design goal;
[0097] S4: performing GAIL-SAC learning on the sliding mode guidance law of the defense aircraft to obtain a defense model;
[0098] S5: Inputting the position, velocity and acceleration of the attacking aircraft; the position, velocity and acceleration of the target; and the position, velocity and acceleration of the defending aircraft into the defense model to intercept the aircraft.
[0099] Specifically, step S1 includes:
[0100] S11: Construct a two-dimensional kinematic model between the incoming aircraft and the target:
[0101]
[0102] in, is the relative speed between the incoming aircraft and the target, is the speed of the incoming aircraft, is the heading angle of the incoming aircraft relative to the inertial coordinate system, is the line of sight angle between the incoming aircraft and the target relative to the inertial coordinate system, is the heading angle of the target velocity relative to the inertial coordinate system, is the target speed, for The first derivative of for The second derivative of is the acceleration of the incoming aircraft, is the relative distance between the incoming aircraft and the target, is the acceleration of the target, for The first derivative of for The first derivative of
[0103] S12: Construct a two-dimensional kinematic model between the attacking aircraft and the defending aircraft:
[0104]
[0105] in, Indicates the relative speed between the attacking aircraft and the defending aircraft, It represents the line of sight angle of the line connecting the attacking aircraft and the defending aircraft relative to the inertial coordinate system, Indicates the speed of the defensive aircraft, Indicates the heading angle of the defense aircraft relative to the inertial coordinate system, for The first derivative of for The second derivative of Indicates the relative distance between the attacking aircraft and the defending aircraft, represents the acceleration of the defense aircraft, for The first derivative of
[0106] S13: The two-dimensional kinematic model between the incoming aircraft and the target and the two-dimensional kinematic model between the incoming aircraft and the defending aircraft are combined to form a three-body dynamic model of the incoming aircraft-target-defending aircraft.
[0107] Specifically, step S2 includes:
[0108] S21: Design the interception conditions for the defensive aircraft to intercept the incoming aircraft:
[0109]
[0110] in, The desired sight angle for the defending aircraft;
[0111] S22: Define guidance law design objectives based on the interception conditions:
[0112]
[0113] in, To prevent the aircraft’s sight angle tracking error, for The first derivative of .
[0114] Since the present invention takes into account the problem of limited terminal acceleration of the aircraft, it is defined as follows:
[0115]
[0116] in, is the acceleration ordinal number, , It is the maximum maneuverability that each aircraft actuator can provide during the guidance process.
[0117] Specifically, step S3 includes:
[0118] S31: Define the sliding surface:
[0119]
[0120]
[0121] in, is the designed sliding surface, is the first adjustment parameter, is the second adjustment parameter, is the first exponential parameter, is the second exponential parameter, To define the symbolic function, for The independent variable of the function, To take the absolute value, is a symbolic function; The value range of is: ; The value range of is: ; The value range of is: ; .
[0122] S32: Design convergence law:
[0123]
[0124] in, is the first derivative of the sliding surface, is a gain that takes a positive value in the reaching law;
[0125] S33: Design a sliding mode guidance law for the defense aircraft based on the sliding mode surface and the reaching law:
[0126]
[0127]
[0128] in, is the scale factor of the proportional guidance.
[0129] In the three-body active defense problem, the motion status information of the three aircraft is abstracted into a Markov decision process. The intelligent agent selects actions by continuously interacting with the environment and optimizes its own strategy by maximizing rewards.
[0130] In this paper, we use the Generative Adversarial Imitation Learning (GAIL) in reinforcement learning combined with the traditional SAC (Soft Actor-Critic) to construct an intelligent agent for defending aircraft in a three-body confrontation environment. Figure 1 As shown, the generator in GAIL-SAC consists of a strategy network constructed using the SAC algorithm for the defensive aircraft, and the discriminator is composed of a neural network. Based on the guidance law designed for the defensive aircraft using traditional sliding mode control in step S3, the initial positions and performance states of the three aircraft in the random initial adversarial environment are randomly determined using Monte Carlo theory, and an expert dataset is collected for network training.
[0131] In the GAIL-SAC algorithm, the generator uses the set state as the network input and the action to be performed by the defense aircraft as the output. The generated data trajectory is collected into memory and the generator's network parameters are updated to make the data generated by the generator as close as possible to the expert data. In the discriminator network, the expert dataset and the three-body environment aircraft generated by the generator use the agent state and action as the neural network input. The neural network uses the Sigmoid function as the activation function of the network output layer. The data generated by the generator is labeled as 0, and the expert data is labeled as 1. The loss of the discriminator model is calculated using the BCELoss function. The RMSprop algorithm is used to optimize the network parameters, allowing the discriminator to accurately assess the gap between the generator and the expert experience, thereby guiding the training of the generator.
[0132] like Figure 2 As shown, Figure 2 This is a schematic diagram of GAIL-SAC learning. Specifically, step S4 includes:
[0133] S41: constructing a SAC algorithm model according to the sliding mode guidance law of the defensive aircraft, and generating a strategy according to the SAC algorithm model;
[0134] S411: Set the state space, action space and reward function of the SAC algorithm:
[0135]
[0136]
[0137]
[0138] in, is the state space, is the action space, is a reward function, which is an integer reward function, is the first reward parameter, is the distance between the defending aircraft and the attacking aircraft at the initial moment, is the second reward parameter, is the third reward parameter, Indicates the distance between the incoming aircraft and the target at the initial moment, is the fourth reward parameter, is the fifth reward parameter, is the sixth reward parameter;
[0139] S412: Design the optimal strategy function of the SAC algorithm:
[0140]
[0141]
[0142] in, the strategy to adopt for the next moment; The strategy adopted at the current moment, To find the maximum function, For expectations, For the current moment, is the expected reward obtained after taking an action in the state following the strategy at the current moment, is the current state, is the action at the current moment, is the regularization coefficient, is the information entropy, The state of following the policy Take action ;
[0143] S413: Constructing the value of the objective function and value function :
[0144]
[0145]
[0146] in, The expected objective function value of the state at the current moment, The expected value obtained after taking an action in the current state of following the strategy, is the discount factor, for Always follow the strategy The expected value of a state after taking an action; To follow the strategy state Take action ;
[0147] S414: Calculate the loss function of the objective function value and the loss function of the value function :
[0148]
[0149]
[0150] in, is the reward at the current moment, To get the minimum value, is the mark number, is the expected value corrected by the loss function of the value function, is the Gaussian distribution function, is the standard normally distributed noise.
[0151] In reinforcement learning, the Proximal Policy Optimization (PPO) algorithm, a representative online policy algorithm, requires a large number of samples for learning, resulting in low sample efficiency. The deep deterministic policy gradient (DDPG) algorithm, an offline policy algorithm, improves sample utilization. However, DDPG is a deterministic policy algorithm, considering only the best action in each state. This leads to insufficient exploration capabilities, unstable training, poor convergence, sensitivity to hyperparameters, and difficulty adapting to diverse and complex environments. The SAC algorithm, an offline policy algorithm with a stochastic policy, has higher sample efficiency than PPO and stronger exploration capabilities than DDPG.
[0152] S42: using the generated strategy as an expert strategy, inputting the state space and action space of the SAC algorithm model and the expert strategy into a discriminator to obtain a probability derived from the state space and action space;
[0153] S421: and As the input of the discriminator, As a generator, it is trained to obtain the possibility from the state space and action space;
[0154] S422: According to the loss function of the discriminator Modify the training:
[0155]
[0156] in, is the input function of the discriminator.
[0157] S43: Update the SAC algorithm model according to the possibilities derived from the state space and the action space to obtain the defense model.
[0158] S431: Activate the possibilities from the state space and action space through the activation function to obtain the automatic adjustment entropy regularization term ;
[0159] S432: Modify the reinforcement learning objective and rewrite it as:
[0160]
[0161] in, To obtain the maximum value, is the target entropy, Indicates restrictions;
[0162] S433: Loss function based on automatic adjustment of entropy regularization term Modify the training:
[0163] .
[0164] The effectiveness of the embodiments of the present invention is demonstrated below:
[0165] First, a three-system guidance simulation environment is constructed, and the initial state parameters are shown in Table 1.
[0166] Table 1 Monte Carlo simulation environment parameters
[0167]
[0168] in, is the position coordinate of the target aircraft, are the coordinates of the incoming aircraft, is the acceleration due to gravity.
[0169] In the Monte Carlo experiment, the defense aircraft initially adopted the sliding mode guidance law designed by the present invention to collect expert data sets. The range is , velocity inclination , the defense aircraft overload range is . A dataset of 3,000 experts is constructed and used in the GAIL-SAC algorithm to guide the intelligent agent to imitate the expert experience, thereby protecting the target and achieving active defense.
[0170] Table 2 shows the settings of network parameters in the intelligent network GAIL-SAC algorithm.
[0171] Table 2 GAIL-SAC network parameter settings
[0172]
[0173] In order to verify the algorithm proposed above, the optimal agent neural network parameters and weights obtained by GAIL-SAC training are saved, the initial state of the three bodies is reinitialized, and a simulation test experiment is carried out. In the test experiment, the defense aircraft and the target start from the same position at the initial moment, the initial position is (0, 0), and the initial speed of the defense aircraft is 750 , the heading angle is , the target's initial speed is 400 , the heading angle is , the initial position of the incoming aircraft is (10000, 0), and the initial speed is 400 , the heading angle is The target acceleration is , the attacking aircraft intercepts the target with proportional guidance, where the proportional coefficient is 3. The expected sight angle of the defensive aircraft intercepting the attacking aircraft is .
[0174] like Figure 3 As shown, Figure 3 This is the active defense trajectory diagram of the embodiment of the present invention, from Figure 3 It can be seen that the defensive aircraft can intercept the incoming aircraft within the specified time and ensure that the target is attacked. Its active defense miss margin is 0.0924, which meets the requirements.
[0175] like Figure 4 As shown, Figure 4 The graph of the sight angle between the defensive aircraft and the attacking aircraft changes with time is shown in the figure. As can be seen from the figure, the defensive aircraft is at the desired sight angle value. , to intercept the incoming aircraft.
[0176] like Figure 5 As shown, the following describes an aircraft collision angle defense system based on reinforcement learning and sliding mode control provided by the present invention. The aircraft collision angle defense system based on reinforcement learning and sliding mode control described below and the aircraft collision angle defense method based on reinforcement learning and sliding mode control described above can correspond to each other.
[0177] Model building module 101: constructing a three-body dynamics model of an attacking aircraft, a target, and a defensive aircraft; constructing a guidance law design target based on the three-body dynamics model; and constructing a sliding mode guidance law for the defensive aircraft based on the three-body dynamics model and the guidance law design target.
[0178] Model learning module 102: performing deep learning on the sliding mode guidance law of the defense aircraft to obtain a defense model;
[0179] Aircraft interception module 103: The position, velocity, and acceleration of the attacking aircraft; the position, velocity, and acceleration of the target; and the position, velocity, and acceleration of the defending aircraft are input into the defense model to intercept the aircraft. Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art will appreciate that the technical solutions described in the above embodiments may be modified or some of the technical features thereof may be replaced by equivalents. Such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
[0180] The present invention combines traditional sliding mode guidance with reinforcement learning to propose an active defense guidance method based on reinforcement learning GAIL-SAC combined with sliding mode control for complex three-body confrontation environments. Guided by an expert strategy based on sliding mode control, this guidance method ensures that, in complex three-body confrontation environments, the defending aircraft can intercept the incoming aircraft within the specified maximum time, subject to the terminal angle constraint, with high accuracy, thus protecting the target from attack by the incoming aircraft.
[0181] The present invention can accelerate the convergence speed of the network, realize the online autonomous decision-making of the defense aircraft, and improve the success rate of the defense aircraft in intercepting the incoming aircraft.
Claims
1. A method for preventing aircraft collision angles based on reinforcement learning and sliding mode control, characterized in that: The following steps are involved: S1: Construct a three-body dynamics model of the attacking aircraft, the target, and the defensive aircraft; S2: Constructing guidance law design objectives based on the three-body dynamics model; S3: constructing a sliding mode guidance law for the defense aircraft according to the three-body dynamics model and the guidance law design goal; S4: Perform GAIL-SAC learning on the sliding mode guidance law of the defensive aircraft to obtain a defense model. When constructing a SAC algorithm model based on the sliding mode guidance law of the defensive aircraft, set the reward function as: in, is a reward function, which is an integer reward function, is the first reward parameter, Indicates the relative distance between the attacking aircraft and the defending aircraft, is the distance between the defending aircraft and the attacking aircraft at the initial moment, is the second reward parameter, It represents the line of sight angle of the line connecting the attacking aircraft and the defending aircraft relative to the inertial coordinate system, is the third reward parameter, is the relative distance between the incoming aircraft and the target, Indicates the distance between the incoming aircraft and the target at the initial moment, is the fourth reward parameter, is the fifth reward parameter, is the sixth reward parameter; S5: Inputting the position, velocity and acceleration of the attacking aircraft; the position, velocity and acceleration of the target; and the position, velocity and acceleration of the defending aircraft into the defense model to intercept the aircraft.
2. The method for preventing collision angles of aircraft based on reinforcement learning and sliding mode control according to claim 1, characterized in that: Step S1 includes: S11: Construct a two-dimensional kinematic model between the incoming aircraft and the target: in, is the relative speed between the incoming aircraft and the target, is the speed of the incoming aircraft, is the heading angle of the incoming aircraft relative to the inertial coordinate system, is the line of sight angle between the incoming aircraft and the target relative to the inertial coordinate system, is the target speed, is the heading angle of the target velocity relative to the inertial coordinate system, for The first derivative of for The second derivative of is the acceleration of the incoming aircraft, is the acceleration of the target, for The first derivative of for The first derivative of S12: Construct a two-dimensional kinematic model between the attacking aircraft and the defending aircraft: in, Indicates the relative speed between the attacking aircraft and the defending aircraft, Indicates the speed of the defending aircraft, Indicates the heading angle of the defense aircraft relative to the inertial coordinate system, for The first derivative of for The second derivative of represents the acceleration of the defense aircraft, for The first derivative of S13: The two-dimensional kinematic model between the incoming aircraft and the target and the two-dimensional kinematic model between the incoming aircraft and the defending aircraft are combined to form a three-body dynamic model of the incoming aircraft-target-defending aircraft.
3. The method for preventing collision angles of aircraft based on reinforcement learning and sliding mode control according to claim 2, characterized in that: Step S2 includes: S21: Design the interception conditions for the defensive aircraft to intercept the incoming aircraft: in, The desired sight angle for the defending aircraft; S22: Define guidance law design objectives based on the interception conditions: in, To prevent the aircraft’s sight angle tracking error, for The first derivative of .
4. The method for preventing collision angles of aircraft based on reinforcement learning and sliding mode control according to claim 3, characterized in that: Step S3 includes: S31: Define the sliding surface: in, is the designed sliding surface, is the first adjustment parameter, is the second adjustment parameter, is the first exponential parameter, is the second exponential parameter, To define the symbolic function, for The independent variable of the function, To take the absolute value, is a symbolic function; S32: Design convergence law: in, is the first derivative of the sliding surface, is a gain that takes a positive value in the reaching law; S33: Design a sliding mode guidance law for the defense aircraft based on the sliding mode surface and the reaching law: in, is the scale factor of the proportional guidance.
5. The method for preventing collision angle of aircraft based on reinforcement learning and sliding mode control according to claim 4, characterized in that: Step S31 includes: The value range of is: ; The value range of is: ; The value range of is: ; .
6. The method for preventing collision angles of aircraft based on reinforcement learning and sliding mode control according to claim 4, characterized in that: Step S4 includes: S41: constructing a SAC algorithm model according to the sliding mode guidance law of the defensive aircraft, and generating a strategy according to the SAC algorithm model; S42: using the generated strategy as an expert strategy, inputting the state space and action space of the SAC algorithm model and the expert strategy into a discriminator to obtain a probability derived from the state space and action space; S43: Update the SAC algorithm model according to the possibilities derived from the state space and the action space to obtain the defense model.
7. The method for preventing collision angle of aircraft based on reinforcement learning and sliding mode control according to claim 6, characterized in that: Step S41 includes: S411: Set the state space, action space and reward function of the SAC algorithm: in, is the state space, is the action space; S412: Design the optimal strategy function of the SAC algorithm: in, the strategy to adopt for the next moment; The strategy adopted at the current moment, To find the maximum function, For expectations, For the current moment, is the expected reward obtained after taking an action in the state following the strategy at the current moment, is the current state, is the action at the current moment, is the regularization coefficient, is the information entropy, The state of following the policy Take action ; S413: Constructing the value of the objective function and value function : in, The expected objective function value of the state at the current moment, The expected value obtained after taking an action in the current state of following the strategy, is the discount factor, for Always follow the strategy The expected value of a state after taking an action; To follow the strategy state Take action ; S414: Calculate the loss function of the objective function value and the loss function of the value function : in, is the reward at the current moment, To get the minimum value, is the mark number, is the expected value corrected by the loss function of the value function, is the Gaussian distribution function, is the standard normally distributed noise.
8. The method for preventing collision angles of aircraft based on reinforcement learning and sliding mode control according to claim 7, characterized in that: Step S42 includes: S421: and As the input of the discriminator, As a generator, it is trained to obtain the possibility from the state space and action space; S422: According to the loss function of the discriminator Modify the training: in, is the input function of the discriminator.
9. The method for preventing collision angles of aircraft based on reinforcement learning and sliding mode control according to claim 8, characterized in that: Step S43 includes: S431: Activate the possibilities from the state space and action space through the activation function to obtain the automatic adjustment entropy regularization term ; S432: Modify the reinforcement learning objective and rewrite it as: in, To obtain the maximum value, is the target entropy, Indicates restrictions; S433: Loss function based on automatic adjustment of entropy regularization term Modify the training: 。 10. An aircraft collision angle defense system based on reinforcement learning and sliding mode control, used to execute the aircraft collision angle defense method based on reinforcement learning and sliding mode control as claimed in any one of claims 1 to 9, characterized in that: include: Model building module: constructs a three-body dynamics model of the attacking aircraft, the target, and the defensive aircraft; and constructs the guidance law design target based on the three-body dynamics model; Constructing a sliding mode guidance law for a defense aircraft based on the three-body dynamics model and the guidance law design goal; Model learning module: performing deep learning on the sliding mode guidance law of the defensive aircraft to obtain a defense model; Aircraft interception module: The position, velocity and acceleration of the attacking aircraft; the position, velocity and acceleration of the target; and the position, velocity and acceleration of the defending aircraft are input into the defense model to intercept the aircraft.