A reinforcement learning guidance and control integrated method for intercepting three-dimensional maneuvering targets
Through deep reinforcement learning technology, an integrated guidance and control method based on the Actor-Critic framework was designed, which solved the problem of difficulty in collaborative cooperation between guidance and control loops in three-dimensional maneuvering target interception, realized the aircraft's independent decision-making and intelligent control, and improved the interception accuracy and efficiency.
Patent Information
- Application Number
- CN202411012483.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-26
AI Technical Summary
When facing three-dimensional maneuvering targets, the existing integrated guidance and control methods are difficult to achieve autonomous decision-making and intelligent control of the aircraft, resulting in difficulty in collaborative cooperation between guidance and control loops, significant impact on time coupling, affecting interception accuracy and efficiency.
The integrated guidance and control method based on deep reinforcement learning is adopted, and the neural network structure of the Actor-Critic framework is designed by constructing a six-degree of freedom interceptor missile model and target maneuver mode, so as to realize the agent's independent decision-making and action learning, and directly generate the rudder surface deflection control instructions.
It effectively avoids the problem of reducing guidance accuracy caused by inconsistent time constants of guidance and control loops, improves the aircraft's autonomous perception and intelligent learning capabilities, and improves the robustness and scope of application of guidance and control algorithms.
Smart Images

Figure CN118938676B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft guidance and control, and in particular to a reinforcement learning guidance and control integrated method for intercepting a three-dimensional maneuvering target. Background Art
[0002] In the design of traditional aircraft guidance and control algorithms, the guidance loop and the control loop are usually designed separately. This method is applicable when the aircraft is flying at a low speed. In modern air combat, the increase in the flight speed and maneuverability of interceptors and targets has made the traditional idea of designing the outer loop guidance loop and the inner loop attitude control loop separately no longer meet the spectrum separation conditions. The use of an integrated guidance and control design method can effectively avoid the problem of delayed response of interceptor guidance instructions, and solve the channel coupling effect between the interceptor guidance loop and the attitude control loop. With the development of nonlinear control theory, constrained optimization methods, and disturbance observers, a large number of research results on the integration of guidance and control have emerged. The integrated guidance and control strategy based on the sliding mode control theory can achieve the interception of maneuvering targets at a given strike angle. In order to solve the system interference problem caused by target maneuvers, a first-order auxiliary system is designed to realize the adaptive estimation and compensation of the upper bound of external disturbances, and then an adaptive robust guidance and control integration strategy is designed. In order to make a more accurate estimate of external disturbances, a fast terminal sliding mode integrated method based on an extended state observer was proposed. The external disturbances were regarded as extended states, and the fixed time convergence theory was used to realize the fixed time estimation of the disturbance state. In order to further reduce the flight time of the interceptor missile, some scholars proposed a preset time guidance and control integrated strategy based on the time generator function (TBG), which realized that the line of sight angular rate converged to zero quickly within the preset time, effectively reducing the miss amount for maneuvering targets. In view of the high-order and strong nonlinear characteristics of the guidance and control integrated model, the incremental guidance and control integrated design method based on the backstepping control theory has been widely studied, and the robust interception of three-dimensional maneuvering targets has been realized. The above analysis results show that most of the existing guidance and control integrated methods are based on traditional nonlinear control theory. For example, although the sliding mode variable structure control method has strong system robustness, the problem of control signal jitter cannot be avoided. Although the backstepping control theory can solve the high-order system control problem, the controller parameters are complex and the parameter selection is difficult. Although the observer technology can estimate the system disturbance caused by the target maneuver, the estimation effect is seriously dependent on the accuracy of the integrated guidance model.
[0003] In order to solve the problems of complex controller design process, low environmental adaptability of control law, and high model dependence caused by traditional control theory, a class of guidance and control methods based on reinforcement learning theory have been proposed one after another. By constructing a reinforcement learning guidance strategy composed of multi-layer neural networks, it is possible to effectively intercept maneuvering targets in two-dimensional planes or three-dimensional space. In view of the strong nonlinearity and multivariable characteristics of three-dimensional guidance models, guidance and control methods based on deep reinforcement learning theory are widely used. At present, the mainstream reinforcement learning algorithms include proximal policy optimization algorithm (PPO) and deterministic policy gradient algorithm (DDPG). This type of method has become one of the mainstream algorithms for solving continuous aircraft guidance and control problems due to its unique advantages in solving intelligent decision-making problems in continuous state and action space. Some scholars have adopted reinforcement learning methods to realize intelligent guidance under the condition of only providing seeker angle information. In order to improve the efficiency of interaction between the intelligent agent and the environment, relevant scholars have designed a new orbital vehicle guidance law using the actor-critic framework to realize intelligent obstacle avoidance and collaborative guidance of multiple spacecraft.
[0004] However, most of the existing reinforcement learning guidance and control research results use a three-dimensional guidance model to achieve the attack or interception of fixed targets or three-dimensional maneuvering targets, generate overload instructions through the action network, and then the inner-loop attitude control system tracks the outer-loop guidance instructions to complete the update of the space orientation of the aircraft. Although this method has been proven to be effective, the coordination between subsystems is difficult to proceed smoothly. In particular, when the relative distance between the interceptor and the target becomes smaller, the rapid change of relative geometry may cause the separation design method to fail. The guidance and control loops cannot work together, resulting in the inability to give full play to the inherent maneuverability and flight characteristics of the aircraft. In order to avoid these shortcomings and improve the autonomy and intelligence of the aircraft at the same time, it is necessary to study the six-degree-of-freedom guidance and control method based on reinforcement learning. In view of the problem of three-dimensional maneuvering targets in the interception space, designing a guidance and control integrated strategy with autonomous decision-making ability, reducing the time coupling effect between the outer-loop guidance loop and the inner-loop attitude control loop of the aircraft, and realizing effective interception of three-dimensional maneuvering targets is one of the key scientific problems that need to be solved urgently. Based on this, the present invention proposes a reinforcement learning guidance and control integrated method for intercepting three-dimensional maneuvering targets to solve the above problems. Summary of the invention
[0005] The purpose of the present invention is to provide an integrated method of reinforcement learning guidance and control for intercepting three-dimensional maneuvering targets to solve the problems raised in the above-mentioned background technology.
[0006] To achieve the above purpose, a reinforcement learning guidance and control integrated method for intercepting a three-dimensional maneuvering target includes the following steps:
[0007] Step S1, establishing a three-dimensional missile-target relative kinematics model, and constructing a six-degree-of-freedom interceptor missile model for guidance and control based on the common maneuvering mode of the target;
[0008] Step S2, determining a deep reinforcement learning method based on the existing Actor-Critic framework structure of deep reinforcement learning theory;
[0009] Step S3, based on the integrated design of guidance and control of deep reinforcement learning, build a neural network structure, and design the state space, action space, reward function, etc. for intelligent agent training and learning.
[0010] Preferably: before the model is established in step S1, the coordinate system needs to be defined, mainly the ground coordinate system, the coordinate system Axyz fixedly connected to the earth's surface is the ground coordinate system; the missile body coordinate system, which is fixedly connected to the missile body and moves and rotates in space with the missile body, is a dynamic coordinate system; the ballistic coordinate system, the origin O of the coordinate system is selected at the center of mass of the interceptor missile, and it is a dynamic coordinate system that moves in space with the center of mass of the interceptor missile; the line of sight coordinate system, the missile-target line of sight is selected as the Ox4 axis, where the coordinate origin O is selected at the center of mass of the interceptor missile.
[0011] Preferably: the process of establishing the three-dimensional guidance model in step S1 is as follows: R is the distance between the projectile and the target, θ L and φ L Represents the line of sight inclination and line of sight deflection, V m and V T are the speeds of the interceptor missile and the target, θ m and φ m is the lead angle of the interceptor missile in the line of sight, θ t and φ t is the lead angle of the target in the line of sight, let a m =[a mR a mθ a mφ ] T and a t =[a tR a tθ a tφ ] T are the components of the interceptor missile and target acceleration vector in the line of sight, respectively. The missile-target relative kinematic model is described as
[0012]
[0013] By taking the derivatives of equations (1) to (3), and combining equations (4) and (7), we can obtain the following three-dimensional guidance model:
[0014]
[0015] The acceleration vector a of the interceptor missile and the target in the velocity system m and a t Projected to the line of sight coordinate system respectively, the three-dimensional guidance model can be rewritten as
[0016]
[0017] Where Y and Z are the interception lift and lateral force, and the expression is:
[0018]
[0019] Where: m, S ref , q are the interceptor mass, reference area and flight pressure, α, β are the angle of attack and sideslip angle, and
[0020]
[0021] Preferably, the process of constructing the target maneuvering mode in step S1 is as follows: in the target track coordinate system, the target dynamic equation is:
[0022]
[0023] Among them, a tx ,a ty ,a tz is the component of the target acceleration vector in each axis direction in the target track coordinate system, θ t , is the target's track pitch angle and track azimuth relative to the reference coordinate system. According to the transformation relationship between the track coordinate system and the reference coordinate system,
[0024]
[0025] Among them, T TC Transformation matrix from reference coordinate system to target coordinate system
[0026]
[0027] The three-dimensional kinematic equation of the target is:
[0028]
[0029] Preferably: the six-degree-of-freedom interceptor missile model in step S1 is established as follows. The attitude dynamics of the interceptor missile in the STT configuration can be described as:
[0030]
[0031] Where: γ V , w i , J i , δ i, i = x, y, z are respectively the interceptor missile roll angle, the missile system's attitude angular velocity in three directions, the three-axis moment of inertia and the rudder surface deflection angle, M i , i = x, y, z is the aerodynamic moment, let L be the moment reference length, then the aerodynamic moment expression is,
[0032]
[0033] Let θ Lf and φ Lf To intercept the target’s desired sight angle, define the following system state variables:
[0034]
[0035] Combining the above three-dimensional guidance kinematic model and the aircraft attitude dynamics model, the following six-degree-of-freedom guidance and control integrated model of the interceptor missile can be obtained:
[0036]
[0037] Where: d i , i=2,3,4 is the total interference of the system, Sat(·) is the saturation function, and,
[0038]
[0039] Preferably: The deep reinforcement learning method of step S2 is determined as follows. A typical reinforcement learning process can usually be represented by a 5-tuple: <S,A,P,R,γ RL >, where S, A, P, and R are the system state space, action space, system state transition probability, and reward function space, respectively, and γ RL is the discount factor, and the goal of reinforcement learning is to find the optimal strategy π * :S→A maximizes the following expected return
[0040]
[0041] The optimal strategy in reinforcement learning is expressed as
[0042]
[0043] Q-value function Q π (s t ,a t ) is used to indicate that in state s t Next, take action a t Expected value of revenue that can be achieved
[0044] Q π (s t ,a t )=E π[R t |s t ,a t ] (27)
[0045] The V value function is used to represent the t Under this condition, the expected return value that can be achieved by adopting strategy π is
[0046] V π (s t )=E π [R t |s t ] (28)
[0047] A deep neural network is used to approximate the Q value and V value in equations (27) and (28), thus forming a deep reinforcement learning theory. Let θ Q and φ V are the hyperparameters of the Q-value network and the V-value network respectively, then,
[0048]
[0049] Taking the deterministic policy gradient algorithm as an example, the target network Q value in the Actor-Critic framework is expressed as,
[0050] y t =r t +γQ(s t+1 ,π(s t+1 |φ π ′)|θ′ Q ) (30)
[0051] Where: θ′ Q and φ π ′ is the hyperparameter of the target Q-value network and V-value network. The following performance index function is minimized through the experience replay mechanism, and the parameters of the Critic network are updated at the same time.
[0052]
[0053] Where: N is the number of samples in the experience pool, and the policy gradient algorithm is used to optimize the Actor network parameters.
[0054]
[0055] Use soft update to update the target Actor network and target Critic network
[0056]
[0057] Where: τ is the soft update factor.
[0058] Preferably: the process of designing the action space and state space of the intelligent agent in step S3 is as follows: the flight state and relative state information of the interceptor missile are selected as the intelligent agent observation variable s t ,
[0059]
[0060] The rudder angles of the pitch channel and the yaw channel are selected as the action variables a t ,
[0061] a t =[δ y ,δ z ] T (35)
[0062] Using the above state variable s t and action variable a t When training an agent, you need to t and a t Perform normalization.
[0063] Preferably, the agent reward function of step S3 is specifically designed as follows: the agent reward function is composed of the following four parts:
[0064] a) Distance reward function
[0065]
[0066] Where: ω1, R(t-1), R(t) are weight coefficients, respectively, the missile-target distance at time t-1, the missile-target distance at time t, and the distance reward function divided by the interceptor missile velocity vector size It plays a role in normalizing the distance reward function;
[0067] b) Heading angle reward function
[0068]
[0069] Where: V (t) and are respectively the ballistic deviation angle at time t and the expected ballistic inclination angle at time t;
[0070] c) Energy consumption reward function
[0071]
[0072] Where: ω3>0, k i ,i∈3,4 is the weight coefficient;
[0073] d) Training end reward function
[0074]
[0075] Where: t max , R max is the maximum training step length and the maximum allowed miss amount. The training end reward function is triggered when the interceptor successfully intercepts the target or other termination conditions occur.
[0076] The comprehensive reward function of the agent training process is
[0077]
[0078] Preferably: the neural network structure design in step S3 is specifically the Actor network and Critic network structure in the DDPG deep reinforcement learning framework, wherein the Actor action network and the Critic evaluation network both adopt a 4-layer linear BP neural network structure, the hidden layer network both adopts the tanh (·) function as the activation function, and the network output layer both adopts linear (·) as the activation function. In the neural network design, the Actor network input dimension is designed to be 12, the network structure is 12-120-50-20-2, the Critic network input dimension is 14, and the network structure is 14-120-25-5-1.
[0079] Preferably: the specific training process of the agent in step S3 is as follows:
[0080] 1) Set the three-dimensional maneuvering target type and the basic attribute parameters of the six-degree-of-freedom interceptor missile;
[0081] 2) Design the DDPG deep reinforcement learning framework and initialize the network parameters;
[0082] 3) Randomly initialize the initial state of the interceptor missile and the target point position;
[0083] 4) Using the current target and interceptor missile relative state as the Actor network input, calculate the current interceptor missile rudder deviation command;
[0084] 5) Update the interceptor missile status and target space position according to the guidance and control integrated instructions;
[0085] 6) Calculate the reward function and store the current state, action, reward value and next state in the experience pool;
[0086] 7) Randomly sample historical data from the experience pool and update the parameters of the reinforcement learning network through certain learning rules;
[0087] 8) Determine whether the relative state of the interceptor missile meets the simulation end condition. If so, re-initialize the initial state and spatial orientation of the target and interceptor missile, otherwise proceed to the next step;
[0088] 9) Determine whether the agent reward function converges. If so, stop training the agent, otherwise go to step 4).
[0089] Compared with the prior art, the present invention has the following beneficial effects:
[0090] 1) Different from the three-degree-of-freedom guidance model used in the prior art, the present invention uses a six-degree-of-freedom interceptor missile model as the research object, effectively avoiding the problem of reduced guidance accuracy caused by inconsistent time constants of the guidance and control loops;
[0091] 2) Different from the existing guidance law literature that uses overload or acceleration as the interceptor missile input command, the guidance and control integration method proposed in the present invention directly generates the control command of the rudder surface deflection, which can give full play to the performance advantages of the six-degree-of-freedom interceptor missile and has a wider range of model application;
[0092] 3) The guidance and control strategy designed by the present invention draws on a large number of interactive learning ideas between the intelligent agent and the environment, which improves the interceptor missile's autonomous perception and intelligent learning capabilities of the complex external flight environment, and effectively improves the robustness of the guidance and control algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 It is a schematic diagram of the three-dimensional interception geometry of the present invention;
[0094] Figure 2 It is a schematic diagram of the average reward function of the intelligent agent of the present invention after 2500 trainings;
[0095] Figure 3 It is a schematic diagram of the miss amount of the intelligent agent of the present invention after 2500 trainings;
[0096] Figure 4 It is a three-dimensional trajectory diagram of 50 Monte Carlo simulations of the present invention;
[0097] Figure 5 It is the rudder deviation diagram of the yaw channel simulated by 50 times of Monte Carlo simulation of the present invention;
[0098] Figure 6 It is the rudder deflection diagram of the pitch channel simulated by 50 times of Monte Carlo simulation of the present invention;
[0099] Figure 7 It is the angle of attack diagram of 50 Monte Carlo simulation flights of the present invention;
[0100] Figure 8 It is a side slip angle diagram of 50 Monte Carlo simulation flights of the present invention;
[0101] Fig. 9 It is the yaw rate diagram of 50 Monte Carlo simulations of the present invention;
[0102] Fig.10It is the pitch rate diagram of 50 Monte Carlo simulations of the present invention;
[0103] Fig.11 It is the projectile-target distance diagram of 50 Monte Carlo simulations of the present invention;
[0104] Fig.12 This is a graph of miss distances obtained by 50 Monte Carlo simulations of the present invention. DETAILED DESCRIPTION
[0105] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the implementation regulations described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0106] Example
[0107] See also Figure 1-Figure 4 , FIG. 1 is a preferred embodiment of the present invention, a reinforcement learning guidance and control integrated method for intercepting a three-dimensional maneuvering target, comprising the following steps:
[0108] Step S1, establishing a three-dimensional missile-target relative kinematics model, and constructing a six-degree-of-freedom interceptor missile model for guidance and control based on the common maneuvering mode of the target;
[0109] Step S2, determining a deep reinforcement learning method based on the existing Actor-Critic framework structure of deep reinforcement learning theory;
[0110] Step S3, based on the integrated design of guidance and control of deep reinforcement learning, build a neural network structure, and design the state space, action space, reward function, etc. for intelligent agent training and learning.
[0111] Further, specifically including:
[0112] (1) Establish a three-dimensional guidance model and a six-degree-of-freedom model of the interceptor missile:
[0113] 1) Define the coordinate system, including,
[0114] Ground coordinate system: The coordinate system Axyz fixed to the earth's surface is the ground coordinate system. The center of mass of the launched instantaneous interceptor missile is selected as the coordinate origin A, and the intersection of the ballistic plane and the horizontal plane is selected as the Ax axis. The direction pointing to the target is positive; the Ay axis is perpendicular to the Ax axis and points upward, and the Az axis is given according to the right-hand rule, thereby determining the coordinate system Axyz.
[0115] Projectile body coordinate system: It is fixed to the projectile body, moves and rotates in space with the projectile body, and is a dynamic coordinate system. Usually, its coordinate origin O is selected at the center of mass of the interceptor missile. The Ox1 axis is generally parallel to the axis of symmetry of the projectile body or parallel to the average aerodynamic chord of the missile wing, and is positive when pointing to the head; the Oy1 axis is located in the left-right symmetry plane of the projectile body, the Oy1 axis is perpendicular to the Ox1 axis and contains the origin, and the Oy1 axis is positive upward; the Oz1 axis is determined according to the right-hand rule.
[0116] Ballistic coordinate system: The origin O of the coordinate system is selected at the center of mass of the interceptor missile (target). It is a dynamic coordinate system that moves in space with the center of mass of the interceptor missile. The Ox2 axis is in the same direction as the velocity vector V of the center of mass of the interceptor missile. The Oy2 axis is in the vertical plane passing through the Ox2 axis and is perpendicular to Ox2. The direction away from the center of the earth is positive; the Oz2 axis is determined by the right-hand rule.
[0117] Line of sight coordinate system: The line of sight between the missile and the target is selected as the Ox4 axis, where the coordinate origin O is selected at the center of mass of the interceptor missile, and the direction of the interceptor missile pointing to the target is positive. The Oy4 axis is perpendicular to the Ox4 axis and is located on the plumb plane containing the Ox4 axis. The Oy4 axis is positive when it points upward. The Oz4 axis is determined according to the right-hand rule. The relative motion equation between the interceptor missile and the target in the line of sight direction is mainly established in the line of sight coordinate system. The interceptor missile and the target are fixed to the line of sight coordinate system, which is a moving coordinate system.
[0118] 2) Establish a 3D guidance model
[0119] The relative kinematic relationship between the interceptor missile and the target is usually described in the line of sight coordinate system, where R is the missile-target distance, θ L and φ L Represents the line of sight inclination and line of sight deflection, V m and V T are the speeds of the interceptor missile and the target, θ m and φ m is the lead angle of the interceptor missile in the line of sight, θ t and φ t is the lead angle of the target in the line of sight, let a m =[a mR a mθ a mφ ] T and a t =[a tR a tθ a tφ ] T are the components of the interceptor missile and target acceleration vector in the line of sight, respectively. Then the missile-target relative kinematic model is described as:
[0120]
[0121]
[0122] By taking the derivatives of equations (1) to (3), and combining equations (4) and (7), we can obtain the following three-dimensional guidance model:
[0123]
[0124] The acceleration vector a of the interceptor missile and the target in the velocity system m and a t Projected to the line of sight coordinate system respectively, the three-dimensional guidance model can be rewritten as:
[0125]
[0126] Where Y and Z are the interception lift and lateral force, and the expression is:
[0127]
[0128] Where: m, S ref , q are the interceptor mass, reference area and flight pressure, α, β are the angle of attack and sideslip angle, and,
[0129]
[0130] 3) Constructing the target maneuvering pattern
[0131] In the target track coordinate system, the three-dimensional space dynamics and kinematics model of the target can be obtained, and the dynamic equation of the target is:
[0132]
[0133] Among them, a tx ,a ty ,a tz θ is the component of the target acceleration vector in each axis direction in the target track coordinate system. t , is the target's track pitch angle and track azimuth relative to the reference coordinate system. According to the transformation relationship between the track coordinate system and the reference coordinate system,
[0134]
[0135] Among them, T TC Transformation matrix from reference coordinate system to target coordinate system
[0136]
[0137] The three-dimensional kinematic equation of the target is:
[0138]
[0139] The interception process of an interceptor missile on a target is generally short. The impact of target motion on guidance is mainly reflected in the motion of the target as a particle. Target maneuvering generally refers to the motion of a particle perpendicular to its velocity direction, which is expressed by maneuvering overload, and its unit is generally expressed in gravitational acceleration g. Common target maneuvers are shown below.
[0140] ①Constant linear motion
[0141] The target moves in a straight line at a constant speed, that is, the target acceleration is zero, the speed magnitude and direction remain constant, and the target motion trajectory is a straight line motion. This is the most common type of maneuver, and its motion equation is described as:
[0142] a tx =a ty =a tz =0 (16)
[0143] Where V t is a constant.
[0144] ②S maneuver (snake maneuver)
[0145] The serpentine maneuver is an anti-radar, anti-aircraft gun, and anti-interceptor missile maneuver. Its purpose is to increase the radar tracking error, increase the pre-calculation error of the anti-aircraft gun and the flight control overload of the interceptor missile, thereby reducing the interception effectiveness of the air defense weapon system. The enemy target is within the firepower area of our ground air defense weapons and performs this maneuver before entering the bombing route or when leaving after dropping the bomb. The trajectory of the serpentine maneuver can be approximately represented by a sine curve, that is, when the target performs a maneuvering turn according to the maximum acceleration to increase the trajectory angle to 180°, it performs a maneuvering turn according to the opposite acceleration to increase the trajectory angle to 180°, and alternates until the maneuver ends. The maneuver equation is described as,
[0146]
[0147] Where k = 0, 1, 2, 3, ..., a ty0 、V t are all constants.
[0148] ③ The target performs a horizontal maximum maneuvering turn flight, that is, the target performs a maneuvering turn flight, and after reaching a certain moment (such as the trajectory deflection angle increases to 90° or 180°), it then turns to a uniform straight flight at a speed of 180°.
[0149] For example, the maneuver equation is described as
[0150]
[0151] Among them, a ty0 、V t are all constants.
[0152] ④ The target performs vertical maximum maneuvering climb flight, that is, the target performs maneuvering climb flight (such as the trajectory inclination increases to 90° and then decreases to 0°). The maneuvering equation is described as follows:
[0153]
[0154] Among them, a tz0 、V t are all constants.
[0155] ⑤ Barrel movement (screw movement)
[0156] The spiral maneuver can be regarded as the result of the superposition of turning in the horizontal plane and climbing in the vertical plane. The law of the spiral maneuver can be expressed as follows:
[0157]
[0158] 4) Establish a six-degree-of-freedom interceptor missile model
[0159] The attitude dynamics of the interceptor missile in the STT configuration can be described as:
[0160]
[0161] Where: γ V , w i , J i , δ i , i = x, y, z are respectively the interceptor missile roll angle, the missile system's attitude angular velocity in three directions, the three-axis moment of inertia and the rudder surface deflection angle, M i , i = x, y, z is the aerodynamic moment, let L be the moment reference length, then the aerodynamic moment expression is,
[0162]
[0163] Let θ Lf and φ Lf To intercept the target’s desired sight angle, define the following system state variables:
[0164]
[0165] Combining the above three-dimensional guidance kinematic model and the aircraft attitude dynamics model, the following six-degree-of-freedom guidance and control integrated model of the interceptor missile can be obtained:
[0166]
[0167] Where: d i , i=2,3,4 is the total interference of the system, Sat(·) is the saturation function, and,
[0168]
[0169] In order to facilitate the design of guidance and control laws, the present invention makes the following assumptions:
[0170] Assumption 1: The STT interceptor missile has an axisymmetric shape, and the inertia product of the interceptor missile about each axis of the missile body coordinate system is approximately zero;
[0171] Assumption 2: Considering that the terminal guidance process is short, the change in the interceptor missile speed during the terminal guidance process is not considered;
[0172] Assumption 3: Assume that the STT interceptor missile roll channel is in a well-controlled state during the terminal guidance phase, and the interceptor missile roll angle and roll angle rate are approximately zero.
[0173] (2) Reinforcement learning theory under the Actor-Critic framework:
[0174] Reinforcement learning theory is an intelligent heuristic optimization algorithm that aims to maximize expected returns. By simulating human trial and error behavior, it obtains a large amount of historical experience data on the interaction between the intelligent agent and the external environment, and generates the optimal strategy through certain optimization rules.
[0175] A typical reinforcement learning process can usually be represented by a 5-tuple <S,A,P,R,γ RL >, where S, A, P, and R are the system state space, action space, system state transition probability, and reward function space, respectively, and γ RL is the discount factor, and the goal of reinforcement learning is to find the optimal strategy π * :S→A maximizes the following expected return,
[0176]
[0177] The optimal strategy in reinforcement learning is expressed as,
[0178]
[0179] Q-value function Q π (s t ,a t ) is used to indicate that in state s t Next, take action a t The expected value of the profit that can be achieved,
[0180] Q π (s t ,a t )=E π [R t |s t ,a t ] (27)
[0181] The V value function is used to represent the tUnder this condition, the expected return value that can be achieved by adopting strategy π is
[0182] V π (s t )=E π [R t |s t ] (28)
[0183] In order to solve the dimensionality curse problem that occurs in classical reinforcement learning as the system state dimension or action dimension increases, deep neural networks are usually used to approximate the Q value and V value in equations (27) and (28), thus forming the deep reinforcement learning theory. Let θ Q and φ V are the hyperparameters of the Q-value network and the V-value network respectively, then,
[0184]
[0185] The actor-critic framework (Actor-Critic) can effectively improve the interaction efficiency and sample utilization rate between the intelligent agent and the external environment. It is one of the current mainstream deep reinforcement learning frameworks. Typical algorithms include AC, A3C, DPG and DDPG. In the Actor-Critic learning framework, Actor is a policy network from state input to action output, and Critic is a Q-value network from the state and action input of the intelligent agent to the expected return output. Taking the deterministic policy gradient algorithm (DDPG) as an example, the Q value of the target network in the Actor-Critic framework is expressed as,
[0186] y t =r t +γQ(s t+1 ,π(s t+1 |φ π ′)|θ′ Q ) (30)
[0187] Where: θ′ Q and φ π ′ is the hyperparameter of the target Q-value network and V-value network. The following performance index function is minimized through the experience replay mechanism, and the parameters of the Critic network are updated at the same time.
[0188]
[0189] Where: N is the number of samples in the experience pool, and the policy gradient algorithm is used to optimize the Actor network parameters.
[0190]
[0191] The target actor network and target critic network are updated by soft update.
[0192]
[0193] Where: τ is the soft update factor.
[0194] (3) Integrated design of guidance and control based on reinforcement learning theory:
[0195] Compared with the traditional integrated guidance and control design method, reinforcement learning theory is used to solve the guidance and control problems of aircraft. It has unique advantages such as weak model dependence, simple controller parameter selection, performance optimization under complex constraints, and dynamic adaptive target maneuvering. The DDPG algorithm is used to generate integrated guidance and control instructions, which includes two stages: offline training of the interceptor model and online use of the strategy network. In the offline training stage of the interceptor agent, a training environment for real flight environment and tasks is constructed, including the six-degree-of-freedom model of the interceptor, the target dynamics model, the interactive environment design, the reward function design, and the Actor-Critic deep network framework construction. In the online use stage of the integrated guidance and control model, it is only necessary to propose the converged intelligent agent Actor network, and input the current six-degree-of-freedom flight state of the interceptor and the relative information of the missile and the target to realize the online generation of integrated guidance instructions.
[0196] ① Agent action space and state space selection
[0197] In order to realize the mission of intercepting three-dimensional maneuvering targets, the flight state and relative state information of the interceptor missile are selected as the agent observation variable s. t ,
[0198]
[0199] Since the change of the interceptor missile roll channel is not considered, the rudder angles of the pitch channel and the yaw channel are selected as the action variable a. t ,
[0200] a t =[δ y ,δ z ] T (35)
[0201] Using the above state variable s t and action variable a t When training an agent, in order to avoid the problem of training stability caused by the difference in the order of magnitude of different state variables, it is necessary to adjust the variable s t and a t Perform normalization.
[0202] ②Agent reward function design
[0203] Reinforcement learning aims to maximize the cumulative reward of the agent. The optimized strategy is closely related to the design of the reward function. Considering the interception effect, energy consumption, trajectory smoothness and other aspects, the agent reward function designed by the present invention consists of the following four parts:
[0204] a) Distance reward function
[0205]
[0206] Where: ω1, R(t-1), R(t) are weight coefficients, respectively, the missile-target distance at time t-1, and the missile-target distance at time t. The distance reward function is divided by the interceptor missile velocity vector size It plays a role in normalizing the distance reward function.
[0207] b) Heading angle reward function
[0208]
[0209] Where: V (t) and are the ballistic deviation angle at time t and the expected ballistic inclination angle at time t respectively.
[0210] c) Energy consumption reward function
[0211]
[0212] Where: ω3>0, k i ,i∈3,4 is the weight coefficient.
[0213] d) Training end reward function
[0214]
[0215] Where: t max , R max is the maximum training step and the maximum allowed miss amount. The training end reward function is triggered when the interceptor missile successfully intercepts the target or other termination conditions (such as training step timeout, vehicle state singularity, etc.) occur.
[0216] The comprehensive reward function of the agent training process is:
[0217]
[0218] ③Neural network structure design
[0219] The Actor network and Critic network structure in the DDPG deep reinforcement learning framework, where the Actor action network and the Critic evaluation network both use a 4-layer linear BP neural network structure, the hidden layer network uses the tanh(·) function as the activation function, and the network output layer uses linear(·) as the activation function.
[0220] In the neural network design, the Actor network input dimension is designed to be 12, the network structure is 12-120-50-20-2, the Critic network input dimension is designed to be 14, and the network structure is 14-120-25-5-1.
[0221] ④Agent training process
[0222] The integrated design of guidance and control based on reinforcement learning theory can intercept three-dimensional maneuvering targets. The algorithm input is the six-degree-of-freedom interceptor model, target model, Actor and its target network, Critic and its target network and other information. The algorithm output is the Actor network, which is used to generate integrated control instructions for intercepting three-dimensional maneuvering targets online. The specific training process of the intelligent agent is as follows:
[0223] 1) Set the three-dimensional maneuvering target type and the basic attribute parameters of the six-degree-of-freedom interceptor missile;
[0224] 2) Design the DDPG deep reinforcement learning framework and initialize the network parameters;
[0225] 3) Randomly initialize the initial state of the interceptor missile and the target point position;
[0226] 4) Using the current target and interceptor missile relative state as the Actor network input, calculate the current interceptor missile rudder deviation command;
[0227] 5) Update the interceptor missile status and target space position according to the guidance and control integrated instructions;
[0228] 6) Calculate the reward function and store the current state, action, reward value and next state in the experience pool;
[0229] 7) Randomly sample historical data from the experience pool and update the parameters of the reinforcement learning network through certain learning rules;
[0230] 8) Determine whether the relative state of the interceptor missile meets the simulation end condition. If so, re-initialize the initial state and spatial orientation of the target and interceptor missile, otherwise proceed to the next step;
[0231] 9) Determine whether the agent reward function converges. If so, stop training the agent, otherwise go to step 4).
[0232] The offline training of the six-degree-of-freedom interceptor model is completed through reinforcement learning, and the converged action network is saved and output for online generation of integrated guidance and control instructions.
[0233] In order to verify the effectiveness of the algorithm proposed in the present invention, a terminal guidance simulation scene as shown in Table 1 is set in three-dimensional space. The interceptor missile properties and aerodynamic parameters refer to publicly published literature, and the maximum rudder deflection angle is set to 10°.
[0234] Table 1 Interceptor missile and target initial state settings
[0235]
[0236] In order to improve the interception probability of maneuvering targets, the following two types of target maneuvers are used in the agent training process:
[0237] Sine maneuver: a yθ =a tφ =5gsin(π / 4t)m / s 2
[0238] Constant value maneuver: a tθ =a tφ =5g, g=9.8m / s 2
[0239] The DDPG reinforcement learning network and training parameters are shown in Table 2. The simulation step size is set to 0.01s, the discount rate is 0.99, the Actor network learning rate is 0.003, the Critic network learning rate is 0.003, and the exploration noise model parameter μ OU =0.0,σ OU =0.10, the number of single sample sampling N is 256, the maximum training round M is set to 2500 times, the maximum single simulation step L is 4000 steps, the experience pool capacity D is 1e7, the soft update constant τ is 0.005, the training and simulation results are shown in Figure 2-Figure 12 shown.
[0240] Table 2 Reinforcement learning network parameter settings
[0241]
[0242] The above content is a further detailed description of the present invention in combination with specific implementation methods. It cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, some simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the scope of protection determined by the claims submitted for the present invention.
Claims
1. A reinforcement learning guidance and control integrated method for intercepting three-dimensional maneuvering targets, characterized in that: The steps include: Step S1, establishing a three-dimensional missile-target relative kinematics model, and constructing a six-degree-of-freedom interceptor missile model for guidance and control based on the common maneuvering mode of the target; Step S2, determining a deep reinforcement learning method based on the existing Actor-Critic framework structure of deep reinforcement learning theory; Step S3, based on deep reinforcement learning, integrated design of guidance and control, building a neural network structure, and designing the state space, action space, and reward function for agent training and learning; The six-degree-of-freedom interceptor missile model in step S1 is established as follows. The attitude dynamics of the interceptor missile in the STT configuration can be described as follows: Where: γ V , w i , J i , δ i , i = x, y, z are respectively the interceptor missile roll angle, the missile system's attitude angular velocity in three directions, the three-axis moment of inertia and the rudder surface deflection angle, M i , i = x, y, z is the aerodynamic moment, let L be the moment reference length, then the aerodynamic moment expression is, Let θ Lf and φ Lf To intercept the target’s desired sight angle, define the following system state variables: Combining the above three-dimensional guidance kinematic model and the aircraft attitude dynamics model, the following six-degree-of-freedom guidance and control integrated model of the interceptor missile can be obtained: Where: d i , i=2,3,4 is the total interference of the system, Sat(·) is the saturation function, and, 2. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 1, characterized in that: Before the model is established in step S1, the coordinate system needs to be defined, mainly the ground coordinate system, and the coordinate system Axyz fixed to the earth's surface is the ground coordinate system; the missile body coordinate system is fixed to the missile body, moves and rotates in space with the missile body, and is a dynamic coordinate system; the ballistic coordinate system, the origin O of the coordinate system is selected at the center of mass of the interceptor missile, and it is a dynamic coordinate system that moves in space with the center of mass of the interceptor missile; the line of sight coordinate system, the missile-target line of sight is selected as the Ox4 axis, and the coordinate origin O is selected at the center of mass of the interceptor missile.
3. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 2, characterized in that: The process of establishing the three-dimensional guidance model in step S1 is as follows: R is the distance between the projectile and the target, θ L and φ L Represents the line of sight inclination and line of sight deflection, V m and V t are the speeds of the interceptor missile and the target, θ m and φ m is the lead angle of the interceptor missile in the line of sight, θ t and φ t is the lead angle of the target in the line of sight, let a m =[a mR a mθ a mφ ] T and a t =[a tR a tθ a tφ ] T are the components of the interceptor missile and target acceleration vector in the line of sight, respectively. The missile-target relative kinematic model is described as By taking the derivatives of equations (1) to (3), and combining equations (4) and (7), we can obtain the following three-dimensional guidance model: The acceleration vector a of the interceptor missile and the target in the velocity system m and a t Projected to the line of sight coordinate system respectively, the three-dimensional guidance model can be rewritten as Where Y and Z are the interception lift and lateral force, and the expression is: Where: m, S ref , q are the interceptor mass, reference area and flight pressure, α, β are the angle of attack and sideslip angle, and 4. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 3, characterized in that: The process of constructing the target maneuvering mode in step S1 is as follows: In the target track coordinate system, the target dynamic equation is: Among them, a tx ,a ty ,a tz is the component of the target acceleration vector in each axis direction in the target track coordinate system, θ t , is the target's track pitch angle and track azimuth relative to the reference coordinate system. According to the transformation relationship between the track coordinate system and the reference coordinate system, Among them, T TC Transformation matrix from reference coordinate system to target coordinate system The three-dimensional kinematic equation of the target is:
5. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 1, characterized in that: The process of determining the deep reinforcement learning method in step S2 is as follows: a typical reinforcement learning process can be represented by a 5-tuple. <S,A,P,R,γ RL >, where S, A, P, and R are the system state space, action space, system state transition probability, and reward function space, respectively, and γ RL is the discount factor, and the goal of reinforcement learning is to find the optimal strategy π * :S→A maximizes the following expected return The optimal strategy in reinforcement learning is expressed as Q-value function Q π (s t ,a t ) is used to represent the state variable s t Next, take action a t Expected value of revenue that can be achieved Q π (s t ,a t )=E π [R t |s t ,a t ] (27) The V value function is used to represent the t Under this condition, the expected return value that can be achieved by adopting strategy π is V π (s t )=E π [R t |s t ] (28) A deep neural network is used to approximate the Q value and V value in equations (27) and (28), thus forming a deep reinforcement learning theory. Let θ Q and φ V are the hyperparameters of the Q-value network and the V-value network respectively, then, Taking the deterministic policy gradient algorithm as an example, the target network Q value in the Actor-Critic framework is expressed as, y t =r t +γQ(s t+1 ,π(s t+1 |f π ′)|θ′ Q ) (30) Where: θ′ Q and φ π ′ is the hyperparameter of the target Q-value network and V-value network. The following performance index function is minimized through the experience replay mechanism, and the parameters of the Critic network are updated at the same time. Where: N is the number of samples in the experience pool, and the policy gradient algorithm is used to optimize the Actor network parameters. Use soft update to update the target Actor network and target Critic network Where: τ is the soft update factor.
6. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 5, characterized in that: The design process of the agent action space and state space in step S3 is as follows: t , The rudder angles of the pitch channel and the yaw channel are selected as the action variables a t , a t =[δ y ,d z ] T (35) Using the above state variable s t and action variable a t When training an agent, you need to t and a t Perform normalization.
7. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 6, characterized in that: The agent reward function design in step S3 is specifically as follows: the agent reward function is composed of the following four parts: a) Distance reward function Where: ω1, R(t-1), R(t) are weight coefficients, respectively, the missile-target distance at time t-1, the missile-target distance at time t, and the distance reward function divided by the interceptor missile velocity vector size It plays a role in normalizing the distance reward function; b) Heading angle reward function Where: V (t) and are respectively the ballistic deviation angle at time t and the expected ballistic inclination angle at time t; c) Energy consumption reward function Where: ω3>0, k i ,i∈3,4 is the weight coefficient; d) Training end reward function Where: t max , R max is the maximum training step length and the maximum allowed miss amount. The training end reward function is triggered when the interceptor successfully intercepts the target or other termination conditions occur. The comprehensive reward function of the agent training process is 8. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 7, characterized in that: The neural network structure design in step S3 is specifically the Actor network and Critic network structure in the DDPG deep reinforcement learning framework, wherein the Actor action network and the Critic evaluation network both adopt a 4-layer linear BP neural network structure, the hidden layer network both adopts the tanh(·) function as the activation function, and the network output layer both adopts linear(·) as the activation function. In the neural network design, the Actor network input dimension is designed to be 12, and the network structure is 12-120-50-20-2, the Critic network input dimension is designed to be 14, and the network structure is 14-120-25-5-1.
9. The method of integrated reinforcement learning guidance and control for intercepting a three-dimensional maneuvering target according to claim 8, characterized in that: The specific training process of the agent in step S3 is as follows: 1) Set the three-dimensional maneuvering target type and the basic attribute parameters of the six-degree-of-freedom interceptor missile; 2) Design the DDPG deep reinforcement learning framework and initialize the network parameters; 3) Randomly initialize the initial state of the interceptor missile and the target point position; 4) Using the current target and interceptor missile relative state as the Actor network input, calculate the current interceptor missile rudder deviation command; 5) Update the interceptor missile status and target space position according to the guidance and control integrated instructions; 6) Calculate the reward function and store the current state, action, reward value and next state in the experience pool; 7) Randomly sample historical data from the experience pool and update the parameters of the reinforcement learning network through certain learning rules; 8) Determine whether the relative state of the interceptor missile meets the simulation end condition. If so, reinitialize the initial state and spatial orientation of the target and interceptor missile, otherwise proceed to the next step; 9) Determine whether the agent reward function converges. If so, stop training the agent, otherwise go to step 4).
Citation Information
Patent Citations
Forward guidance-based three-dimensional guidance law design method for intercepting hypersonic aircrafts
CN106934120A
Remote interception launch data acquisition method and system based on single-sided launch table
CN113642122A
Terminal guidance law design method based on deep reinforcement learning
CN115857548A