A design method for multi-constraint guidance laws driven by deep deterministic policy gradients
By combining the improved A3DPG algorithm with the self-attention mechanism and LSTM, the multi-constraint problem in missile guidance was solved, and precise interception of missiles under the constraints of field of view angle and overload changes was achieved, thereby improving the accuracy and stability of missile interception.
Patent Information
- Application Number
- CN202411791783.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing missile guidance algorithms find it difficult to simultaneously achieve attack angle constraints, overload change constraints, and terminal accuracy requirements while ensuring that the target is within the seeker's field of view, resulting in problems such as increased overload and field of view angle exceeding the range.
The DDPG algorithm is improved by combining the self-attention mechanism and LSTM algorithm, and the A3DPG algorithm is designed. By introducing the fractional-order sliding mode control law and reward mechanism into the three-dimensional relative kinematic model, the guidance law is optimized to meet multiple constraints such as attack angle constraint, field of view angle constraint, and suppression of terminal overload changes.
A multi-constraint guidance law with shortest time, small miss distance, field of view (FOV) constraint, acceleration change suppression, sliding surface convergence and terminal attack angle constraint is realized during the missile interception process, which improves the accuracy and stability of missile interception.
Smart Images

Figure CN119644742B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of missile guidance technology, and in particular relates to a design method for a multi-constraint guidance law driven by a deep deterministic policy gradient. Background Art
[0002] To accurately and effectively intercept targets and improve the performance and accuracy of missile terminal guidance, it is necessary to consider realistic constraints such as angle of attack, overload variation, and the seeker's field of view. Traditional guidance algorithms, such as biased proportional guidance and sliding mode control, can achieve angle of attack constraints but struggle to ensure the target is within the seeker's field of view. With the current trend toward intelligent aircraft guidance, data-driven aircraft guidance laws have become a hot research topic due to their flexible design and ease of implementing multiple constraints. However, purely data-driven approaches lack stability and reliability compared to traditional methods.
[0003] Reference [1] designed a deep reinforcement learning guidance method under field angle constraints. Taking a two-dimensional dynamic model as an example, the bias term of the biased proportional guidance is learned to adjust the guidance law. By suppressing the angular velocity of the missile-target line of sight, quasi-parallel approach interception is achieved, indirectly realizing the constraint of the missile's field of view angle. Since the designed guidance law is not implemented with a three-dimensional dynamic model, it has certain simplifications, and its guidance law has a huge overload when approaching the target at the end. This sharp increase in overload puts a lot of pressure on the missile controller implementation, so it has certain limitations in practical terms.
[0004] Reference [2] designed a guidance method with attack angle constraints based on distributed reinforcement learning. It not only considered the convergence of the sight angle but also directly set up a reward mechanism for the field of view angle constraint. For moving targets, it can be seen that the terminal acceleration is still too large, and the field of view angle of the scheme itself is within the set range. The maximum value of the field of view angle is larger than the maximum value of the field of view angle of the sliding mode guidance law, which does not reflect the advantages of sliding mode control.
[0005] The terminal constrained angle guidance law proposed in the literature [3] is a three-dimensional dynamic model and is also a bias term used to learn proportional guidance. However, the overload generated by the guidance law changes dramatically, and the reward mechanism only uses angle convergence as a dense reward. The design of the reward mechanism is too simplified.
[0006] Reference [4] designs a three-dimensional guidance law based on deep reinforcement meta-learning and attack angle constraint to solve the problem of intercepting maneuvering targets with time-varying speed missiles with partial actuator failure. The terminal overload does not cause a sudden change, but the sudden change in the terminal line of sight rate may cause the target to fly out of the field of view, resulting in insufficient actual terminal accuracy.
[0007] Reference [5] proposed a deep reinforcement learning guidance law based on the double-delayed deep deterministic policy gradient (TD3) algorithm to solve the problem of high-speed maneuvering target interception. It directly used the overload instruction as the action and verified the advantages of its algorithm. However, it did not consider the limitation of the field of view angle range during the interception of high-speed maneuvering targets, which has certain practical limitations.
[0008] The present invention improves the traditional reinforcement learning (DDPG) algorithm based on a guidance law designed based on a three-dimensional relative kinematic model. It adds a self-attention mechanism as the state feature extraction layer and uses LSTM as the first layer of the policy network and value network, making it more suitable for learning sequence data with long round cycles. It also introduces an exponential decay exploration strategy (step S323) and changes the training strategy, which can achieve precise guidance under realistic constraints such as attack angle constraints, field of view angle constraints, and suppression of terminal overload transformation.
[0009] [1] Zhang Qinglong, Zhao Bin, Xu Xinpeng. Deep reinforcement learning guidance method under the constraint of passive detection field of view angle [J]. Journal of Astronautics, 2024, Vol. 45(8): 1281-1289;
[0010] [2] Li Bohao, An Xuman, Yang Xiaofei, Wu Yunjie, Li Guofei. Distributed reinforcement learning guidance method under attack angle constraint[J]. Acta Astronautics, 2022, Vol. 43(8): 1061-1069;
[0011] [3] Zheng Chengchen, Li Hui, Tao Wei, Liu Sicheng, Wu Fengguo, He Li. Missile terminal constraint angle guidance law based on deep reinforcement learning[J]. Tactical Missile Technology, 2022, (6): 93-102, 110;
[0012] [4] Liang Chen, Wang Weihong, Lai Chao. Deep reinforcement meta-learning guidance law with attack angle constraint[J]. Journal of Astronautics, 2021, Vol. 42(5): 611-620;
[0013] [5]Qiu Xiaoqi, Gao Changsheng, Jing Wuxing. Deep reinforcement learning guidance law for intercepting maneuvering targets in the atmosphere[J]. Acta Astronautics, 2022, Vol. 43(5): 685-695 Summary of the Invention
[0014] This paper proposes a method for designing a multi-constraint guidance law driven by a deep deterministic policy gradient. This method combines the self-attention mechanism and the LSTM algorithm to improve the policy network and value network in the traditional DDPG deep reinforcement learning algorithm. The proposed algorithm, A3DPG, is an attention-driven deterministic policy gradient reinforcement learning algorithm. The attention layer serves as a feature extraction layer and is shared by the policy network and the value network. This method addresses the problem of high-precision three-dimensional target interception by considering constraints such as angle of attack, overload variation, and seeker field of view angle, enabling guidance law design.
[0015] To achieve the above object, the technical solution adopted by the present invention is:
[0016] A method for designing a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm, the method comprising the following steps:
[0017] S1: Establish a three-dimensional relative kinematic model of the missile interception target;
[0018] S2: Based on S1, design the fractional-order sliding mode control law and analyze its stability;
[0019] S3: Design and optimization of A3DPG algorithm;
[0020] S4: In a dynamic environment, the reward mechanism is used to induce the strategy network to learn the action strategy, achieving a multi-constraint guidance law with the shortest time, small miss distance, field of view angle (FOV) constraint, acceleration change suppression, sliding surface convergence reward, terminal attack angle constraint, and total reward.
[0021] Furthermore, in S1, the step of establishing a three-dimensional relative kinematic model of the missile interception target comprises the following steps:
[0022] S11: Establish a kinematic mathematical model for missile interception targets;
[0023]
[0024] Where: r is the relative distance between the projectile and the target; V is the relative velocity between the projectile and the target; m is the missile's speed; θ m is the ballistic inclination angle of the missile; is the ballistic deviation angle of the missile; V t is the missile's speed; θ t is the target's ballistic inclination; is the target's trajectory angle; θ los is the high and low angle; is the rate of change of high and low angles; is the azimuth angle change rate;
[0025] S12: Establish the missile's center of mass kinematic model;
[0026]
[0027] Where: a x ,a y ,a z are the acceleration components in the x-axis, y-axis, and z-axis directions in the ballistic coordinate system; x, y, and z are the positions of the missile in the launch inertial coordinate system; V is the velocity of the missile's center of mass; θ is the velocity inclination angle of the missile; is the missile's velocity deflection angle.
[0028] Furthermore, in S2, the design of the fractional-order sliding mode control law and its stability analysis include:
[0029] S21: Derivation of fractional-order sliding mode control law;
[0030] The Caputo-type fractional integral is defined as:
[0031]
[0032] Where: t represents the current time; t0 represents the initial time;
[0033] The Caputo-type fractional derivative is defined as:
[0034]
[0035] Where: t represents the current time; t0 represents the initial time; m is the reaching law parameter;
[0036] Design sliding surface;
[0037]
[0038] Where: S is the sliding surface; D α is the fractional derivative of the power of α; sign is the sign function; k1, k2, p, q are the sliding surface parameters;
[0039] Design convergence law;
[0040]
[0041] Where: S is the sliding surface; v is the reaching law; η1, η2, m, n are the reaching law parameters; sign is the sign function;
[0042] S22: Proof of stability;
[0043] S221: Proof of stability of approaching modes;
[0044] For S, consider choosing the following Lyapunov function The designed composite sliding surface, if the control law satisfies the conditions, makes When the reaching law is satisfied:
[0045]
[0046] Ability to demonstrate stability;
[0047] Where: V1 is the Lyapunov energy function; S is the sliding surface; D α is the fractional derivative of the power of α; sign is the sign function; k1, k2, p, q are the sliding surface parameters; η1, η2, m, n are the reaching law parameters;
[0048] S222: Proof of finite-time convergence of approaching modes;
[0049]
[0050] Let y = ln|s|, then s = e y
[0051]
[0052]
[0053] Where: S is the designed sliding surface; η1, η2, m, n are reaching law parameters; sign is the sign function; t0 is the initial time; t s is the sliding surface convergence time;
[0054] S223: Convergence proof of sliding mode;
[0055]
[0056] When the sliding mode variables converge, the system state satisfies the following sliding surface equation:
[0057]
[0058] Further analysis of the global asymptotic stability of the above equation shows that, according to the fractional-order integral operator distribution model, the above equation is equivalent to the following infinite-dimensional integer-order differential equation:
[0059]
[0060] Where:
[0061]
[0062] In the above frequency distribution model, ω is the frequency; z(ω,t) is the real state variable of the fractional-order system; x(t) is the pseudo-state variable; the following two Lyapunov functions are constructed;
[0063] Select the Lyapunov function Taking the derivative of V2 and simplifying it, we get:
[0064]
[0065] Where: V2 is the Lyapunov energy function;
[0066] S23: Sliding surface design;
[0067]
[0068]
[0069] Where: θ los is the high and low angle; is the desired terminal angle of attack; is the azimuth; is the desired terminal azimuth angle; S y is the sliding surface used to control the trajectory inclination; S z is the sliding surface used to control the trajectory angle; D α is the fractional derivative of the power of α; k1, k2, k3, k4, p1, q1, p2, q2 are the sliding surface parameters, and sign is the sign function;
[0070] S24: Convergence law design;
[0071]
[0072] roll out:
[0073]
[0074] Where: S is the sliding surface; v is the reaching law; r is the relative distance between the projectile and the target; is the relative velocity between the projectile and the target; η1, η2, m, n are the reaching law parameters; a y is the normal overload of the vertical plane; a z is the normal overload in the lateral plane.
[0075] Furthermore, in S3, the design and optimization of the A3DPG algorithm includes the following steps:
[0076] S31: Build network structure;
[0077] S311: shared feature extraction layer;
[0078] The input state passes through a self-attention layer, which is shared by the policy network and the value network. The self-attention layer generates a representation that contains global information by calculating the weighted average of the state.
[0079] S312: Policy Network;
[0080] The first layer of the policy network uses an LSTM layer, which is used to capture long-term dependencies in time series.
[0081] S313: Value Network;
[0082] The first layer of the value network also uses an LSTM layer to evaluate the current state and the actions generated by the policy network;
[0083] S32: Algorithm flow;
[0084] S321: Training strategy;
[0085] After each round, all the data of that round is used for training;
[0086] S322: Optimizer update mechanism;
[0087] Updates to the policy network, value network, and attention mechanism;
[0088] S323: Adaptive exploration mechanism;
[0089] According to the total number of training rounds, the variance of Gaussian noise is exponentially reduced, so that the training results tend to be stable as the number of training rounds increases;
[0090]
[0091] In the formula: Episodes is the total number of training times; episode i is the current round number;
[0092] S33: training process;
[0093] S331: Initialize the replay pool and select the required capacity to meet the sequence length of at least 5 rounds;
[0094] S332: Initialization;
[0095] S3321: Online Network Model;
[0096] Including self-attention network K(s t |θ k ), policy network u(s t |θ u ), value network Q(s t ,a t |θQ );
[0097] Among them, s t is the current state of the environment; θ k is the self-attention network parameter; θ u is the policy network parameter; θ Q is the value network parameter; a t The action generated by the policy network at the current moment;
[0098] S3322: target network model;
[0099] Including target self-attention network, target strategy network, target value network K′,u′,Q′;
[0100] Parameter update: K′←K,u′←u,Q′←Q;
[0101] S3323: Set training parameters;
[0102] α1=1×10 -2 , α2=1×10 -2 ,γ=0.95,N~N(0,σ 2 );
[0103] Where: α1 is the learning rate of the policy network; α2 is the learning rate of the value network; γ is the discount factor; N is the random noise used to improve the exploration ability;
[0104] S333: Training Episodes rounds, the training process of each round is as follows:
[0105] The following steps are repeated from the initial moment to the end of the round until the missile successfully hits the target or misses the target;
[0106] S3331: Value Network Update:
[0107] During training, the goal of the value network is to minimize the TD error, and its objectives are as follows:
[0108] y i =r i +γQ′(s i+1 ,μ′(s i+1 ∣θ μ′ )|θ Q′ ,θ K′ )
[0109] Evaluation indicators of value network:
[0110]
[0111] The derivative is:
[0112]
[0113] Update the value network by backpropagating the evaluation indicators:
[0114]
[0115] Where: y i is the expected value of the state at step i; α2 is the learning rate of the value network; L is the evaluation performance function of the value network; is the derivative of L; θ Q is the value network parameter; θ K is the self-attention network parameter; Q(s i ,a i ∣θ Q ,θ K ) is the value of the state-action pair estimated by the value network;
[0116] S3332: Policy Network Update:
[0117] The performance index function of the policy network is:
[0118]
[0119] The derivative is:
[0120]
[0121] Update the policy network by backpropagating the evaluation indicators:
[0122]
[0123] Where: Q(s i ,a i ∣θ Q ,θ K ) is the value of the state-action pair estimated by the value network; α1 is the learning rate of the policy network; J is the evaluation performance function of the value network; is the derivative of J; θ μ is the value network parameter; θ K are the parameters of the self-attention network.
[0124] Furthermore, in S4, the policy network is induced to learn the action policy through the reward mechanism in the dynamic environment; the steps include:
[0125] S41: State Space Design;
[0126] S42: Action Space Design;
[0127] S43: Strategy for reward function design;
[0128] S431: Enhance reward feedback; accelerate the learning process by introducing intermediate rewards;
[0129] S432: Introduce a penalty mechanism to prevent the agent from falling into an invalid action loop;
[0130] S433: Sparsity and density; Balancing sparse and dense rewards to optimize learning efficiency;
[0131] S434: Reward scale; ensure that the range of reward values is reasonable and avoid gradient problems;
[0132] S435: Reward stability; ensuring that the reward function is not affected by environmental changes or noise;
[0133] S44: Introducing the self-attention mechanism;
[0134] S45: Feature extraction layer; its goal is to extract useful features from the environment state to help the agent make better decisions.
[0135] Furthermore, in S434, the reward scale includes:
[0136] S4341: Reasonable scale; ensure that the range of reward values is reasonable to avoid vanishing or exploding gradients during training;
[0137] S4342: Positive reward; represents the desired behavior so that the agent learns the correct strategy.
[0138] Furthermore, in step S435, the reward stability includes:
[0139] S4351: The reward function is stable; the agent will not be unable to learn effectively due to environmental changes or noise interference;
[0140] S4352: Reward Anatomy; Break down rewards into components to simplify task understanding.
[0141] Furthermore, in S43, the reward function design strategy further includes:
[0142] S436: Multi-target reward;
[0143] S437: Exploration-based rewards;
[0144] S4371: Novelty bonus;
[0145] S4372: Adaptive reward mechanism.
[0146] Furthermore, in S4, the reward function expression of the field of view angle FOV constraint is:
[0147] R1=-1·GELU(fov-fovlim ,0.01,2,10)
[0148] Where: R1 is the reward value for the field of view constraint; fov is the missile field of view; fov lim is the seeker field of view angle constraint range;
[0149]
[0150] The reward function expression for acceleration change suppression is:
[0151] R2=0.1·GELU(Δacc,1,1,0.1)
[0152] Δacc=A max ·dt-(|Δa y |+|Δa z |)
[0153] Where: R2 is the reward value for overload change suppression; |Δa y | is the absolute value of the rate of change of normal acceleration of the vertical plane; |Δa z | is the absolute value of the rate of change of normal acceleration of the horizontal plane; A max is the maximum overload value of the missile; dt is the integration step length;
[0154] The reward function expression of the sliding surface convergence reward is:
[0155]
[0156] Where: R3 is the reward value for the convergence of the sliding surface; S y is the sliding surface used to control the trajectory inclination; S z is the sliding surface used to control the trajectory angle;
[0157]
[0158] The reward function expression for small miss distance is:
[0159]
[0160] Where: R4 is the reward value for the miss amount; M dis is the off-target amount; K r is the missile killing radius;
[0161] The reward function expression with the shortest time is:
[0162] R5=-t f
[0163] Where: R5 is the reward value for the guidance process time;
[0164] The reward function expression of the terminal attack angle constraint is:
[0165]
[0166] Where: R6 is the reward value for the attack angle constraint of the terminal;
[0167] The reward function expression of the total reward is:
[0168] Return=λ1R1+λ2R2+λ3R3+(R4+R5+R6)·done
[0169] Where: λ1 = 0.3, λ2 = 0.4, λ3 = 0.3, which are the weight distributions of rewards and punishments under different constraints.
[0170] Further, in S41, the state space is designed;
[0171] In order to fully describe the relative kinematics information, the following state vector is selected:
[0172]
[0173] In S42, the action space is designed;
[0174] Select the fractional sliding surface variables and parameters in the reaching law as actions;
[0175] A t ={K1,K2,η1,η2}
[0176] Among them: K1∈(9.5,10.5),K2∈(9.5,10.5),η1∈(0.05,0.15),η2∈(0.05,0.15).
[0177] The beneficial effects of the present invention over the prior art are as follows: the design method of the multi-constraint guidance law proposed by the present invention combines the traditional fractional-order sliding mode guidance algorithm with deep reinforcement learning, and utilizes their various advantages to cleverly implement multiple constraints. The present invention uses fractional-order sliding mode theory to design a guidance law with convergence of the sight angle (i.e., the angle of attack), and then uses the reinforcement learning algorithm to dynamically learn its sliding surface parameters and convergence law parameters, thereby realizing a dynamically adjusted guidance law. And through a reward mechanism in a dynamic environment, taking into account factors such as the possible fast time, the minimum miss distance, the field of view angle FOV constraint, the acceleration change suppression, and the sliding surface convergence reward, the policy network is induced to learn the optimal guidance law under the evaluation of the above indicators. BRIEF DESCRIPTION OF THE DRAWINGS
[0178] Figure 1 It is a schematic diagram of the three-dimensional relative kinematic model;
[0179] Figure 2 This is the schematic diagram of the A3DPG algorithm;
[0180] Figure 3 This is the structural diagram of the self-attention network, policy network and value network;
[0181] Figure 4 It is a trajectory curve diagram of different guidance laws for chasing a uniform linear target;
[0182] Figure 5 This is a graph showing the longitudinal overload variation for different guidance laws when chasing a uniform linear target.
[0183] Figure 6 This is a graph showing the lateral overload variation for different guidance laws when chasing a uniform linear target.
[0184] Figure 7 It is a graph showing the elevation and depression angle variations for different guidance laws;
[0185] Figure 8 It is the azimuth angle variation curve of different guidance laws;
[0186] Figure 9 It is a curve diagram of field of view angle change for different guidance laws;
[0187] Figure 10 It is a trajectory diagram of different guidance laws for chasing strong maneuvering targets;
[0188] Figure 11 This is a graph showing the longitudinal overload variation for different guidance laws when pursuing a highly maneuvering target.
[0189] Figure 12 This is a graph showing the lateral overload variation for different guidance laws when pursuing a strong maneuvering target.
[0190] Figure 13 It is a graph showing the elevation and depression angle variations for different guidance laws;
[0191] Figure 14 It is the azimuth angle variation curve of different guidance laws;
[0192] Figure 15 This is a curve diagram of the field of view angle change for different guidance laws. DETAILED DESCRIPTION
[0193] This embodiment discloses a method for designing a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm, the method comprising the following steps:
[0194] S1: Establish a three-dimensional relative kinematic model of the missile interception target;
[0195] S2: Based on S1, design the fractional-order sliding mode control law and perform stability analysis (this step ensures that the missile can be effectively guided and controlled according to the relative kinematic model);
[0196] S3: Design and optimization of the A3DPG algorithm (the A3DPG algorithm uses reinforcement learning to optimize the guidance law so that it can meet multiple constraints);
[0197] S4: In a dynamic environment, the policy network is induced to learn action strategies through a reward mechanism to achieve a multi-constraint guidance law with the shortest time, small miss distance, field of view (FOV) constraint, acceleration change suppression, sliding surface convergence reward, terminal attack angle constraint, and total reward (this step uses the A3DPG algorithm optimized in S3 to train the policy network through a reward mechanism, so that it can learn action strategies that meet multiple constraints in a dynamic environment).
[0198] The current aircraft guidance law has the following problems:
[0199] 1. Consider the line of sight angle / attack angle constraints;
[0200] 2. Consider the seeker field of view constraints;
[0201] 3. Suppression of terminal G-load surges on non-maneuvering targets;
[0202] 4. Limitations of traditional DDPG.
[0203] Furthermore, in S1, the step of establishing a three-dimensional relative kinematic model of the missile interception target includes the following steps:
[0204] S11: Establish a kinematic mathematical model for missile interception targets;
[0205]
[0206] Where: r is the relative distance between the projectile and the target; V is the relative velocity between the projectile and the target; m is the missile's speed; θ m is the ballistic inclination angle of the missile; is the ballistic deviation angle of the missile; V t is the missile's speed; θ t is the target's ballistic inclination; is the target's trajectory angle; θ los is the high and low angle; is the rate of change of high and low angles; is the azimuth angle change rate;
[0207] S12: Establish the missile's center of mass kinematic model;
[0208]
[0209] Where: a x ,a y ,a z are the acceleration components in the x-axis, y-axis, and z-axis directions in the ballistic coordinate system; x, y, and z are the positions of the missile in the launch inertial coordinate system; V is the velocity of the missile's center of mass; θ is the velocity inclination angle of the missile; is the missile's velocity deflection angle.
[0210] Furthermore, in S2, the design of the fractional-order sliding mode control law and its stability analysis include:
[0211] S21: Derivation of fractional-order sliding mode control law;
[0212] Fractional Calculus (FOC) Theory:
[0213] Definition of Caputo-type fractional calculus:
[0214] The Grünwald-Letnikov and Riemann-Liouville definitions require that the initial values of the non-zero value problem have fractional derivatives, which is difficult to obtain in actual engineering. Therefore, the present invention adopts the Caputo-type fractional definition widely used in the engineering field. Unless otherwise specified, the subsequent schemes all use D λ f(t) substitution
[0215] D λ f(t) represents the derivative of f(t) of order λ, represents the fractional differential defined by the Caputo type, t0 is the initial time of integration, and t is the current time.
[0216] The Caputo-type fractional integral is defined as:
[0217]
[0218] Where: t represents the current time; t0 represents the initial time; this formula is quoted from Xue Dingyu. Fractional Calculus and Fractional Order Control[M]. Beijing: Science Press, 2018.
[0219] The Caputo-type fractional derivative is defined as:
[0220]
[0221] Where: t represents the current time; t0 represents the initial time; m is the reaching law parameter;
[0222] Quoted from the same source as above.
[0223] Design sliding surface;
[0224]
[0225] Where: S is the sliding surface; D α is the fractional derivative of the power of α; sign is the sign function; k1, k2, p, q are the sliding surface parameters;
[0226] Design convergence law;
[0227]
[0228] Where: S is the sliding surface; v is the reaching law; η1, η2, m, n are the reaching law parameters; sign is the sign function;
[0229] S22: Proof of stability;
[0230] S221: Proof of stability of approaching modes;
[0231] For S, consider choosing the following Lyapunov function The designed composite sliding surface, if the control law satisfies the conditions, makes When the reaching law is satisfied:
[0232]
[0233] Ability to demonstrate stability;
[0234] Where: V1 is the Lyapunov energy function; S is the sliding surface; D α is the fractional derivative of the power of α; sign is the sign function; k1, k2, p, q are the sliding surface parameters; η1, η2, m, n are the reaching law parameters.
[0235] Take the pitch plane as an example:
[0236]
[0237]
[0238] Substitute the control quantity a y get:
[0239]
[0240] Where: θ los It is the high and low sight angle; is the desired terminal sight angle (angle of attack); S y is the designed sliding surface; a y is the acceleration component in the x-axis direction in the ballistic coordinate system;
[0241] Stability is proven.
[0242] S222: Proof of finite-time convergence of approaching modes;
[0243]
[0244] Let y = ln|s|, then s = e y
[0245]
[0246] Where: S is the designed sliding surface; η1, η2, m, n are reaching law parameters; sign is the sign function; t0 is the initial time; t s is the sliding surface convergence time;
[0247] S223: Convergence proof of sliding mode;
[0248]
[0249] When the sliding mode variables converge, the system state satisfies the following sliding surface equation:
[0250]
[0251] Further analysis of the global asymptotic stability of the above equation shows that, according to the fractional-order integral operator distribution model (the model is cited from: Trigeassou JC, Maamri N, Sabatier J, et al. A Lyapunov Ap proach to the Stability of Fractional Differential Equations[J]. Signal Processing, 2011, 91(3): 437-445), the above equation is equivalent to the following infinite-dimensional integer-order differential equation:
[0252]
[0253] Where:
[0254]
[0255] In the above frequency distribution model, ω is the frequency; z(ω,t) is the real state variable of the fractional-order system; x(t) is the pseudo-state variable; the following two Lyapunov functions are constructed;
[0256] Select the Lyapunov function Taking the derivative of V2 and simplifying it, we can obtain (cited from: Tang Xiao, Ye Jikun. Fractional-order sliding mode guidance law for high-speed maneuvering targets [J]. Aviation Weaponry, 2021, Vol. 28(2): 21-26):
[0257]
[0258] Where: V2 is the Lyapunov energy function;
[0259] According to the Lyapunov stability theorem, the system state will converge to zero in a finite time, and the convergence speed is different for different fractional orders;
[0260] S23: Sliding surface design (the sliding surface is used to ensure that the missile's trajectory can quickly and accurately converge to the target trajectory);
[0261]
[0262] Where: θ los is the high and low angle; is the desired terminal angle of attack; is the azimuth; is the desired terminal azimuth angle; S y is the sliding surface used to control the trajectory inclination; S z is the sliding surface used to control the trajectory angle; D α is the fractional derivative of the power of α; k1, k2, k3, k4, p1, q1, p2, q2 are the sliding surface parameters, and sign is the sign function;
[0263] S24: Converging law design (the convergence law is used to ensure that the missile's line of sight angle (i.e., attack angle) converges to the designed desired line of sight angle at a certain convergence speed);
[0264]
[0265] roll out:
[0266]
[0267] Where: S is the sliding surface; v is the reaching law; r is the relative distance between the projectile and the target; is the relative velocity between the projectile and the target; η1, η2, m, n are the reaching law parameters; a y is the normal overload of the vertical plane; a z is the normal overload in the lateral plane (the other variables are the same as described above);
[0268] In order to suppress the chattering of the sliding mode variable along the sliding surface, the function To increase the stability of the control law.
[0269] Comparison algorithms:
[0270] 1. Guidance law of linear sliding surface based on exponential reaching law:
[0271] 1.1 Linear sliding surface
[0272]
[0273] 1.2 Exponential convergence law
[0274] v=-η1S-η2·sign(S)1.3 Guidance law:
[0275]
[0276] 2. Using the optimal guidance law with a navigation ratio of 3:
[0277] 2.1 Proportional guidance law,
[0278]
[0279] Furthermore, in S3, the design and optimization of the A3DPG algorithm includes the following steps:
[0280] Basic principles of the A3DPG algorithm:
[0281] Based on the classic Deep Deterministic Policy Gradient (DDPG) algorithm, it improves on it by introducing LSTM and self-attention mechanism to enhance the performance of policy and value networks.
[0282] A3DPG is also based on the reinforcement learning algorithm of the actor-critic framework. It is applicable to continuous action spaces. Its online network and target network are composed of three networks respectively:
[0283] Self-Attention Network: Extract features from the features generated by the environment;
[0284] Policy network (Actor): generates actions for a given state;
[0285] Value Network (Critic): Evaluates the value of a given state-action pair;
[0286] The A3DPG model has been improved as follows: the policy network and value network share a feature extraction layer (a self-attention network), and the first layer of each network is implemented using LSTM. The exploration mechanism has been improved. As the model explores a sufficient number of states with increasing training times and the network parameters gradually stabilize, an adaptive noise exploration mechanism has been designed. Furthermore, the traditional DDPG training method based on random sampling has been abandoned, and a new training mechanism has been proposed that uses episode data for training.
[0287] S31: Build network structure;
[0288] S311: shared feature extraction layer;
[0289] The input state passes through a self-attention layer, which helps the network focus on key parts of the state and weight different features to obtain a more meaningful representation. The self-attention mechanism, by weighted aggregation of multiple dimensions of the state, helps capture dependencies in the state sequence, which is particularly important in tasks with temporal dependencies (such as trajectory prediction and target tracking).
[0290] The self-attention layer is shared by the policy network and the value network (this improves network training efficiency, reduces parameter redundancy, and ensures that both learn similar temporal dependency information);
[0291] The self-attention layer generates a representation containing global information by calculating the weighted average of the states, so that the policy network and value network can make better decisions and evaluations from a global perspective.
[0292] S312: Policy Network;
[0293] The first layer of the policy network (Actor) uses an LSTM layer, which is used to capture long-term dependencies in time series. (Since environmental states in reinforcement learning are often temporal, LSTM can help the policy and value networks remember past states and understand long-term relationships between states, thereby improving decision-making and evaluation.)
[0294] The goal of the policy network is to generate a reasonable action based on the current state to maximize the long-term reward; the policy network gradually adjusts the parameters of the online self-attention network and online policy network models through interaction with the value network to improve the quality of decision-making. After reaching the set number of training times (3 times or more), the target self-attention network and target policy network are updated through soft updates.
[0295] S313: Value Network;
[0296] The first layer of the value network (Critic) also uses an LSTM layer to evaluate the current state and the actions generated by the policy network. The Q-values generated by the target value network and the Q-values of the state-action pairs estimated by the online value network are used to guide the optimization of the parameters of the online self-attention network and the online policy network. After reaching the set number of training times (3 times or more), the target self-attention network and the target policy network are updated through soft updates.
[0297] S32: Algorithm flow;
[0298] S321: Training strategy;
[0299] After each episode, all data from that episode is used for training. Unlike traditional DDPG, which samples from the experience replay pool for training at each step, the algorithm in this paper uses all data from that episode for training after each episode. This is because the relationship between "state-action" pairs in reinforcement learning is often temporal, and actions at a given moment and future rewards may have long-term dependencies. Therefore, using data from the entire episode for training can better capture this temporal information and avoid the destruction of dependencies between data during random sampling.
[0300] S322: Optimizer update mechanism;
[0301] The policy network and attention mechanism are updated; the policy network optimizer not only updates the policy network but also updates the parameters of the self-attention layer. This ensures that while optimizing the policy, the network's weighting of states (i.e., the attention mechanism) is also optimized accordingly, allowing the entire network to more accurately select and evaluate actions.
[0302] Value Network and Attention Mechanism Updates: Similarly, the value network optimizer not only updates the parameters of the value network, but also updates the shared self-attention mechanism. This allows the value network to adjust its feature representation based on the self-attention mechanism to better estimate the value of state-action pairs.
[0303] S323: Adaptive exploration mechanism;
[0304] According to the total number of training rounds, the variance of Gaussian noise is exponentially reduced, so that the training results tend to be stable as the number of training rounds increases;
[0305]
[0306] In the formula: Episodes is the total number of training times; episode i is the current round number;
[0307] S33: training process;
[0308] S331: Initialize the replay pool and select the required capacity to meet the sequence length of at least 5 rounds;
[0309] S332: Initialization;
[0310] S3321: Online Network Model;
[0311] Including self-attention network K(s t |θ k ), policy network u(s t |θ u ), value network Q(st ,a t |θ Q );
[0312] Among them, s t is the current state of the environment; θ k is the self-attention network parameter; θ u is the policy network parameter; θ Q is the value network parameter; a t The action generated by the policy network at the current moment;
[0313] S3322: target network model;
[0314] Including target self-attention network, target strategy network, target value network K′,u′,Q′;
[0315] Parameter update: K′←K,u′←u,Q′←Q;
[0316] S3323: Set training parameters;
[0317] α1=1×10 -2 , α2=1×10 -2 ,γ=0.95,N~N(0,σ 2 );
[0318] Where: α1 is the learning rate of the policy network; α2 is the learning rate of the value network; γ is the discount factor; N is the random noise used to improve the exploration ability;
[0319] S333: Training Episodes rounds, the training process of each round is as follows:
[0320] The following steps are repeated from the initial moment to the end of the round (the policy network takes an action based on the current state. The dynamic environment enters the next state space based on the current state and the action taken by the model) until the missile successfully hits the target or misses the target;
[0321] S3331: Value Network Update:
[0322] During training, the goal of the value network is to minimize the TD error, and its objectives are as follows:
[0323] y i =r i +γQ′(s i+1 ,μ′(s i+1 ∣θ μ′ )|θ Q′ ,θ K′ )
[0324] Evaluation indicators of value network:
[0325]
[0326] The derivative is:
[0327]
[0328] Update the value network by backpropagating the evaluation indicators:
[0329]
[0330] Where: y i is the expected value of the state at step i; α2 is the learning rate of the value network; L is the evaluation performance function of the value network; is the derivative of L; θ Q is the value network parameter; θ K is the self-attention network parameter; Q(s i ,a i ∣θ Q ,θ K ) is the value of the state-action pair estimated by the value network;
[0331] S3332: Policy Network Update:
[0332] The purpose of the policy network is to make the actions generated as highly evaluated by the value network as possible.
[0333] The performance index function of the policy network is:
[0334]
[0335] The derivative is:
[0336]
[0337] Update the policy network by backpropagating the evaluation indicators:
[0338]
[0339] Where: Q(s i ,a i ∣θ Q ,θ K ) is the value of the state-action pair estimated by the value network; α1 is the learning rate of the policy network; J is the evaluation performance function of the value network; is the derivative of J; θ μ is the value network parameter; θ K is the self-attention network parameter;
[0340] During the training process, the optimization of the policy network and the value network is carried out alternately, and the goal of maximizing long-term rewards is ultimately achieved by gradually adjusting the network parameters.
[0341] Algorithm advantages:
[0342] 1. Capture timing dependencies;
[0343] Introducing LSTM can help reinforcement learning models remember historical information when handling tasks with long-term dependencies and adjust future decisions based on this information. This is particularly important for tasks where the environment states have a clear temporal relationship (for example, videos, trajectories, game rounds, etc.).
[0344] 2. Self-attention mechanism;
[0345] The self-attention mechanism allows the network to dynamically focus on important parts of the state. This mechanism makes the model more flexible in decision-making, especially in complex state spaces, by focusing on more important state features. This mechanism helps the model understand the state holistically rather than relying solely on local features, enhancing the model's expressive power.
[0346] 3. Round data training;
[0347] By using all data after an episode for training, the model can better understand the temporal dependencies between states, rather than just the decisions made at a single time step. This training approach is closer to the actual conditions of reinforcement learning tasks and can improve the model's performance in long-term tasks.
[0348] 4. Parameter sharing;
[0349] The shared design of the LSTM and self-attention mechanisms reduces the number of model parameters and improves training efficiency. The shared feature extraction layer can effectively transfer temporal and global information between the policy network and the value network, helping to accelerate the training process and improve generalization capabilities.
[0350] 5. Robust training process;
[0351] Because the policy and value network optimizers operate on the policy / value network and the self-attention mechanism, respectively, the training process is more stable, avoiding the instability between the policy and value networks common in traditional methods. The optimization of the self-attention mechanism gradually guides the entire model to more precise focus and decision-making in the state space.
[0352] In summary, the proposed A3DPG algorithm, by incorporating LSTM and self-attention mechanisms into the classic DDPG framework, enables the model to better handle reinforcement learning tasks with long-term dependencies and complex structures. By training and sharing LSTM layers after each episode, the algorithm's performance on sequential tasks is expected to be significantly improved, particularly in tasks requiring memory, sequence dependencies, and global context. Furthermore, by independently updating the optimizers for the policy and value networks, network updates are more robust and training efficiency and performance are enhanced.
[0353] Furthermore, in S4, the policy network is induced to learn the action policy through the reward mechanism in the dynamic environment; the steps include:
[0354] S41: State Space Design;
[0355] S42: Action Space Design;
[0356] S43: Strategy for reward function design;
[0357] S431: Enhance reward feedback; accelerate the learning process by introducing intermediate rewards (sparse rewards often slow down the agent's learning process. To avoid this, introduce more intermediate rewards so that the agent can receive some positive feedback as it gradually approaches the goal);
[0358] S432: Introduce a penalty mechanism to prevent the agent from falling into an invalid action loop (the agent may fall into an invalid action loop, such as switching back and forth between the same states; to avoid this situation, a penalty mechanism is introduced);
[0359] S433: Sparsity and density; Balancing sparse and dense rewards to optimize learning efficiency;
[0360] S434: Reward scale; ensure that the range of reward values is reasonable and avoid gradient problems;
[0361] S435: Reward stability; ensuring that the reward function is not affected by environmental changes or noise;
[0362] S44: Introducing the self-attention mechanism;
[0363] It is a technique that can dynamically capture the relationships between elements in sequential data. Its application in the DDPG algorithm can effectively enhance feature extraction capabilities, helping agents better understand state relationships in complex environments and improve learning outcomes. This makes self-attention a worthy feature extraction layer choice for reinforcement learning.
[0364] S45: Feature Extraction Layer; its goal is to extract useful features (abstract features without physical meaning) from the environmental state (including six variables: relative distance, relative velocity, elevation and elevation angles, azimuth and elevation angles, interceptor missile trajectory inclination, and interceptor missile trajectory deviation) to help the agent make better decisions. Self-attention is used to process time series state inputs, helping the model understand the relationship between states. The self-attention mechanism can dynamically focus on important state information, thereby generating more effective state representations and helping the policy network learn better.
[0365] Furthermore, in S434, the reward scale includes:
[0366] S4341: Reasonable scale; ensure that the range of reward values is reasonable to avoid vanishing or exploding gradients during training;
[0367] S4342: Positive reward; represents the desired behavior so that the agent learns the correct strategy.
[0368] Furthermore, in step S435, the reward stability includes:
[0369] S4351: The reward function is stable; the agent will not be unable to learn effectively due to environmental changes or noise interference;
[0370] S4352: Reward Parsing; breaking down rewards into multiple components to simplify task understanding (making it easier for the agent to understand the structure and requirements of the task).
[0371] Clarify the goal; set appropriate reward values for the agent (based on the constraints and expected results, avoid sparse rewards, ensure smooth numerical changes, and ensure that the rewards across constraints are of the same magnitude) (reward design should clarify the agent's learning goals so that appropriate reward values can be set for the agent's various behaviors);
[0372] Guide learning: guide the agent's learning through positive and negative rewards (reward design should be able to guide the agent's learning, that is, provide positive rewards for the agent's performance of behaviors that help achieve the learning goal, and provide negative rewards or no rewards for harmful behaviors);
[0373] Balance: Ensure that the reward design does not favor a certain behavior (the reward design should be balanced, that is, it should not be too biased towards a certain behavior, otherwise it will cause the agent to learn more slowly or even fail to learn).
[0374] Furthermore, in S43, the reward function design strategy further includes:
[0375] S436: Multi-target reward; In complex tasks (in multi-constraint guidance laws, not only field of view constraints but also overload change suppression are considered, the time is as short as possible, and the miss distance is as small as possible), it is necessary to optimize multiple targets at the same time. In this case, a combined reward is designed to reflect the importance of each target, and different reward signals are integrated through a weighted method.
[0376] S437: Exploration-based rewards;
[0377] S4371: Novelty reward: rewards for exploring new states or performing new actions, encouraging agents to explore unseen environments;
[0378] S4372: Adaptive reward mechanism; the reward mechanism is dynamically adjusted based on the learning progress of the agent (for example, as the agent's performance improves, the reward is gradually reduced to maintain the challenge. The reward is increased or decreased in a timely manner based on the learning efficiency of the agent).
[0379] Furthermore, in S4, the reward function expression of the field of view angle FOV constraint is:
[0380] R1=-1·GELU(fov-fov lim ,0.01,2,10)
[0381] Where: R1 is the reward value for the field of view constraint; fov is the missile field of view; fov lim is the seeker field of view angle constraint range;
[0382] This paper proposes the "generalized exponential linear unit" function, which is a generalization of the ELU function, adding parameters and providing higher flexibility.
[0383]
[0384] The reward function expression for acceleration change suppression is:
[0385] R2=0.1·GELU(Δacc,1,1,0.1)
[0386] Δacc=A max ·dt-(|Δa y |+|Δa z |)
[0387] Where: R2 is the reward value for overload change suppression; |Δa y | is the absolute value of the rate of change of normal acceleration of the vertical plane; |Δa z | is the absolute value of the rate of change of normal acceleration of the horizontal plane; A max is the maximum overload value of the missile; dt is the integration step length;
[0388] The reward function expression of the sliding surface convergence reward is:
[0389]
[0390] Where: R3 is the reward value for the convergence of the sliding surface; S y is the sliding surface used to control the trajectory inclination; S z is the sliding surface used to control the trajectory angle;
[0391]
[0392] The reward function expression for small miss distance is:
[0393]
[0394] Where: R4 is the reward value for the miss amount; M dis is the off-target amount; K r is the missile killing radius;
[0395] The reward function expression with the shortest time is:
[0396] R5=-t f
[0397] Where: R5 is the reward value for the guidance process time;
[0398] The reward function expression of the terminal attack angle constraint is:
[0399]
[0400] R6 is the reward value for the attack angle constraint of the terminal.
[0401] Where:
[0402] The reward function expression of the total reward is:
[0403] Return=λ1R1+λ2R2+λ3R3+(R4+R5+R6)·done
[0404] Where: λ1 = 0.3, λ2 = 0.4, λ3 = 0.3, which are the weight distributions of rewards and punishments under different constraints.
[0405] Further, in S41, the state space is designed;
[0406] In order to fully describe the relative kinematics information, the following state vector is selected:
[0407]
[0408] In S42, the action space is designed;
[0409] Select the fractional sliding surface variables and parameters in the reaching law as actions;
[0410] A t ={K1,K2,η1,η2}
[0411] Among them: K1∈(9.5,10.5),K2∈(9.5,10.5),η1∈(0.05,0.15),η2∈(0.05,0.15).
[0412] Design simulation cases, compare with traditional methods, and verify the effectiveness of the invention.
[0413] Example 1:
[0414] The goal of uniform linear motion:
[0415]
[0416] Example 2:
[0417] The goal of strong mobility:
[0418] The target performs an 'S' shape maneuver in the pitch plane and a 'C' shape maneuver in the yaw plane.
[0419]
[0420] The comparison algorithm in the embodiment:
[0421] 1. Proportional guidance law:
[0422] Take the simplified optimal guidance law with a navigation ratio of 3 as an example;
[0423]
[0424] 2. Guidance law for linear sliding surface based on exponential reaching law;
[0425]
[0426] For Example 1, the parameters are selected as follows:
[0427] k1=3, k2=3, k3=0.5, k4=0.5, ε1=0.001, ε2=0.001
[0428] For Case 2, the parameters are selected as follows:
[0429] k1=1, k2=1, k3=0.05, k4=0.05, ε1=0.001, ε2=0.001.
[0430] like Figure 2 As shown in the figure, the A3DPG (attention-driven deep deterministic policy gradient algorithm) algorithm principle diagram shows that the online network continuously interacts with the dynamic environment. The value network in the online network performs temporal difference between the state-action values generated by the value network of the target network, and then updates the parameters of the online value network. The online policy network also updates its parameters based on the evaluation value of the value network. After a certain step size, the online network soft-updates the target network.
[0431] like Figure 3 Figure 2 shows the structure of the self-attention network, policy network, and value network. The policy network and value network use the same network architecture. The first layer consists of an LSTM, followed by a normalization layer, a linear layer, a normalization layer, a random dropout layer, and a linear layer. External information first passes through the self-attention layer before passing through the value network or policy network to produce an output.
[0432] Simulation case 1: Tracking a uniform linear motion target;
[0433] like Figure 4 As shown, from the three-dimensional ballistic curves of the method of the present invention and the other two classic methods in striking a uniform linear target, it can be seen that the ballistic curve based on the A3DPG guidance law of the present invention is obviously smoother than that of the other two algorithms, and is faster and hits the target earlier.
[0434] like Figure 5 As shown in the figure, it can be clearly seen from the acceleration curves of the three guidance methods in the y-axis direction of the ballistic coordinate system that the A3DPG algorithm guidance law proposed in the present invention is relatively small, that is, the fuel consumption is small; and the terminal overload is significantly smaller than the other two methods, which can improve the hit rate of the interceptor missile.
[0435] like Figure 6 As shown in the figure, the acceleration curves of the three guidance methods in the z-axis direction of the ballistic coordinate system clearly show that the A3DPG algorithm proposed in this invention has a relatively small total amount of guidance law, which means less fuel consumption. Furthermore, the terminal overload is significantly less than that of the other two methods, which can improve the hit rate of the interceptor missile.
[0436] like Figure 7 As shown in the figure, the missile-target line-of-sight elevation curves for the three guidance methods show that the classic proportional guidance method cannot achieve line-of-sight elevation constraint, that is, attack elevation constraint. The A3DPG algorithm and sliding mode guidance method proposed in this invention can both achieve attack elevation constraint, but the sliding mode method's premature convergence of line-of-sight elevation may make the interceptor missile sensitive to target maneuvers, resulting in unstable and sudden overload commands.
[0437] like Figure 8As shown in the figure, the missile-target line-of-sight azimuth curves for the three guidance methods show that the classic proportional guidance method cannot achieve line-of-sight azimuth constraint, that is, attack azimuth constraint. The A3DPG algorithm and sliding mode guidance method proposed in this invention can both achieve attack azimuth constraint, but the sliding mode method's premature convergence of line-of-sight azimuth may make the interceptor missile sensitive to target maneuvers, and the overload command may be unstable and surge.
[0438] like Figure 9 As shown in the figure, the field-of-view constraint curves for the three guidance laws clearly show that, given a set maximum field-of-view angle, only the A3DPG algorithm can maintain this angle throughout the entire guidance process. The other two methods all exceed this angle at certain stages. In practical applications, exceeding this angle can cause the seeker to lose its target, making it difficult to intercept the target.
[0439] Simulation Case 2: Tracking a target with strong maneuverability;
[0440] like Figure 10 As shown, from the three-dimensional ballistic curves of the method of the present invention and the other two classical methods in striking highly maneuverable targets, it can be seen that the ballistic curve based on the A3DPG guidance law of the present invention is obviously smoother than that of the other two algorithms, and is faster and hits the target earlier.
[0441] like Figure 11 As shown in FIG, it can be clearly seen from the acceleration curves of the three guidance methods in the y-axis direction of the ballistic coordinate system that the total amount of the A3DPG algorithm guidance law proposed in the present invention is relatively small, that is, the fuel consumption is small; and the terminal overload is significantly smaller than that of the other two methods, which can improve the hit rate of the interceptor missile.
[0442] like Figure 12 As shown in the figure, it can be clearly seen from the acceleration curves of the three guidance methods in the z-axis direction of the ballistic coordinate system that the total amount of the A3DPG algorithm guidance law proposed in the present invention is relatively small, that is, the fuel consumption is small; and the terminal overload is significantly smaller than the other two methods, which can improve the hit rate of the interceptor missile.
[0443] like Figure 13 As shown in the figure, the elevation curves of the missile-target line-of-sight angle for the three guidance methods show that when tracking a highly maneuvering target, the classic proportional guidance and sliding mode methods cannot achieve line-of-sight angle constraints, that is, attack angle constraints. Due to the strong maneuverability of the target, the proposed A3DPG algorithm is not completely accurate in achieving attack angle constraints, but it does provide a significant improvement over the other two algorithms.
[0444] like Figure 14As shown in the figure, the missile-target line-of-sight azimuth curves for the three guidance methods show that when tracking a highly maneuvering target, the classic proportional guidance and sliding mode guidance methods cannot achieve line-of-sight azimuth constraint, that is, attack azimuth constraint. The A3DPG algorithm proposed in this paper is not completely accurate in achieving attack azimuth constraint, but it is significantly improved compared to the other two algorithms.
[0445] like Figure 15 As shown in Figure 3, the field of view constraint curves for the three guidance laws show that a maximum field of view is set. All three methods can achieve field of view constraints throughout the entire process, but this constraint is caused by the coincidence of the relative motion scene, rather than the algorithm's constraints.
[0446] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A design method for a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm, characterized by: The method comprises the following steps: S1: Establish a three-dimensional relative kinematic model of the missile interception target; S2: Based on S1, design the fractional-order sliding mode control law and analyze its stability; including: S21: Derivation of fractional-order sliding mode control law; The Caputo-type fractional integral is defined as: Where: t represents the current time; t0 represents the initial time; The Caputo-type fractional derivative is defined as: Where: t represents the current time; t0 represents the initial time; m is the reaching law parameter; Design sliding surface; Where: S is the sliding surface; D α is the fractional derivative of the power of α; sign is the sign function; k1, k2, p, q are the sliding surface parameters; Design convergence law; Where: S is the sliding surface; v is the reaching law; η1, η2, m, n are the reaching law parameters; sign is the sign function; S22: Proof of stability; S221: Proof of stability of approaching modes; For S, consider choosing the following Lyapunov function The designed composite sliding surface, if the control law satisfies the conditions, makes When the reaching law is satisfied: Ability to demonstrate stability; Where: V1 is the Lyapunov energy function; S is the sliding surface; D α is the fractional derivative of the power of α; sign is the sign function; k1, k2, p, q are the sliding surface parameters; η1, η2, m, n are the reaching law parameters; S222: Proof of finite-time convergence of approaching modes; Let y = ln|s|, then s = e y Where: S is the designed sliding surface; η1, η2, m, n are reaching law parameters; sign is the sign function; t0 is the initial time; t s is the sliding surface convergence time; S223: Convergence proof of sliding mode; When the sliding mode variables converge, the system state satisfies the following sliding surface equation: Further analysis of the global asymptotic stability of the above equation shows that, according to the fractional-order integral operator distribution model, the above equation is equivalent to the following infinite-dimensional integer-order differential equation: Where: In the above frequency distribution model, ω is the frequency; z(ω,t) is the real state variable of the fractional-order system; x(t) is the pseudo-state variable; the following two Lyapunov functions are constructed; Select the Lyapunov function Taking the derivative of V2 and simplifying it, we get: Where: V2 is the Lyapunov energy function; S23: Sliding surface design; Where: θ los is the high and low angle; is the desired terminal angle of attack; is the azimuth; is the desired terminal azimuth angle; S y is the sliding surface used to control the trajectory inclination; S z is the sliding surface used to control the trajectory angle; D α is the fractional derivative of the power of α; k1, k2, k3, k4, p1, q1, p2, q2 are the sliding surface parameters, and sign is the sign function; S24: Convergence law design; roll out: Where: S is the sliding surface; v is the reaching law; r is the relative distance between the projectile and the target; is the relative velocity between the projectile and the target; η1, η2, m, n are the reaching law parameters; a y is the normal overload of the vertical plane; a z is the normal overload in the lateral plane; S3: Design and optimization of the A3DPG algorithm; including the following steps: S31: Build network structure; S311: shared feature extraction layer; The input state passes through a self-attention layer, which is shared by the policy network and the value network. The self-attention layer generates a representation that contains global information by calculating the weighted average of the state. S312: Policy Network; The first layer of the policy network uses an LSTM layer, which is used to capture long-term dependencies in time series. S313: Value Network; The first layer of the value network also uses an LSTM layer to evaluate the current state and the actions generated by the policy network; S32: Algorithm flow; S321: Training strategy; After each round, all the data of that round is used for training; S322: Optimizer update mechanism; Updates to the policy network, value network, and attention mechanism; S323: Adaptive exploration mechanism; According to the total number of training rounds, the variance of Gaussian noise is exponentially reduced, so that the training results tend to be stable as the number of training rounds increases; In the formula: Episodes is the total number of training times; episode i is the current round number; S33: training process; S331: Initialize the replay pool and select the required capacity to meet the sequence length of at least 5 rounds; S332: Initialization; S3321: Online Network Model; Including self-attention network K(s t |θ k ), policy network u(s t |θ u ), value network Q(s t ,a t |θ Q ); Among them, s t is the current state of the environment; θ k is the self-attention network parameter; θ u is the policy network parameter; θ Q is the value network parameter; a t The action generated by the policy network at the current moment; S3322: target network model; Including target self-attention network, target strategy network, target value network K′,u′,Q′; Parameter update: K′←K,u′←u,Q′←Q; S3323: Set training parameters; α1=1×10 -2 ,α2=1×10 -2 ,γ=0.95,N~N(0,σ 2 ); Where: α1 is the learning rate of the policy network; α2 is the learning rate of the value network; γ is the discount factor; N is the random noise used to improve the exploration ability; S333: Training Episodes rounds, the training process of each round is as follows: The following steps are repeated from the initial moment to the end of the round until the missile successfully hits the target or misses the target; S3331: Value Network Update: During training, the goal of the value network is to minimize the TD error, and its objectives are as follows: y i =r i +γQ′(s i+1 ,μ′(s i+1 ∣θ μ′ )∣θ Q′ ,θ K′ ) Evaluation indicators of value network: The derivative is: Update the value network by backpropagating the evaluation indicators: (i Q ,i K )←(θ Q ,i K )-α2▽L Where: y i is the expected value of the state at step i; α2 is the learning rate of the value network; L is the evaluation performance function of the value network; ▽L is the derivative of L; θ Q is the value network parameter; θ K is the self-attention network parameter; Q(s i ,a i ∣θ Q ,θ K ) is the value of the state-action pair estimated by the value network; S3332: Policy Network Update: The performance index function of the policy network is: The derivative is: Update the policy network by backpropagating the evaluation indicators: (i μ ,i K )←(θ μ ,i K )+α1▽J Where: Q(s i ,a i ∣θ Q ,θ K ) is the value of the state-action pair estimated by the value network; α1 is the learning rate of the policy network; J is the evaluation performance function of the value network; ▽J is the derivative of J; θ μ is the value network parameter; θ K is the self-attention network parameter; S4: In a dynamic environment, a reward mechanism is used to induce the strategy network to learn the action strategy, achieving a multi-constraint guidance law with the shortest time, small miss distance, field of view (FOV) constraint, acceleration change suppression, sliding surface convergence reward, terminal attack angle constraint, and total reward. The reward function expression for the field of view (FOV) constraint is: R1=-1·GELU(fov-fov lim ,0.01,2,10) Where: R1 is the reward value for the field of view constraint; fov is the missile field of view; fov lim is the seeker field of view angle constraint range; The reward function expression for acceleration change suppression is: R2=0.1·GELU(Δacc,1,1,0.1) Δacc=A max ·dt-(|Δa y |+|Δa z |) Where: R2 is the reward value for overload change suppression; |Δa y | is the absolute value of the rate of change of normal acceleration of the vertical plane; |Δa z | is the absolute value of the rate of change of normal acceleration of the horizontal plane; A max is the maximum overload value of the missile; dt is the integration step length; The reward function expression of the sliding surface convergence reward is: Where: R3 is the reward value for the convergence of the sliding surface; S y is the sliding surface used to control the trajectory inclination; S z is the sliding surface used to control the trajectory angle; The reward function expression for small miss distance is: Where: R4 is the reward value for the miss amount; M dis is the off-target amount; K r is the missile killing radius; The reward function expression with the shortest time is: R5=-t f Where: R5 is the reward value for the guidance process time; The reward function expression of the terminal attack angle constraint is: Where: R6 is the reward value for the attack angle constraint of the terminal; The reward function expression of the total reward is: Return=λ1R1+λ2R2+λ3R3+(R4+R5+R6)·done Where: λ1 = 0.3, λ2 = 0.4, λ3 = 0.3, which are the weight distributions of rewards and punishments under different constraints.
2. The design method of a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm according to claim 1 is characterized by: In S1, the process of establishing a three-dimensional relative kinematic model of a missile intercepting a target comprises the following steps: S11: Establish a kinematic mathematical model for missile interception targets; Where: r is the relative distance between the projectile and the target; V is the relative velocity between the projectile and the target; m is the missile's speed; θ m is the ballistic inclination angle of the missile; is the ballistic deviation angle of the missile; V t is the missile's speed; θ t is the target's ballistic inclination; is the target's trajectory angle; θ los is the high and low angle; is the rate of change of high and low angles; is the azimuth angle change rate; S12: Establish the missile's center of mass kinematic model; Where: a x ,a y ,a z are the acceleration components in the x-axis, y-axis, and z-axis directions in the ballistic coordinate system; x, y, and z are the positions of the missile in the launch inertial coordinate system; V is the velocity of the missile's center of mass; θ is the velocity inclination angle of the missile; is the missile's velocity deflection angle.
3. The design method of a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm according to claim 1 is characterized by: In S4, the policy network is induced to learn the action policy through the reward mechanism in the dynamic environment; the following steps are included: S41: State Space Design; S42: Action Space Design; S43: Strategy for reward function design; S431: Enhance reward feedback; accelerate the learning process by introducing intermediate rewards; S432: Introduce a penalty mechanism to prevent the agent from falling into an invalid action loop; S433: Sparsity and density; Balancing sparse and dense rewards to optimize learning efficiency; S434: Reward scale; ensure that the range of reward values is reasonable and avoid gradient problems; S435: Reward stability; ensuring that the reward function is not affected by environmental changes or noise; S44: Introducing the self-attention mechanism; S45: Feature extraction layer; its goal is to extract useful features from the environment state to help the agent make better decisions.
4. The design method of a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm according to claim 3 is characterized by: In S434, the reward scale includes: S4341: Reasonable scale; ensure that the range of reward values is reasonable to avoid vanishing or exploding gradients during training; S4342: Positive reward; represents the desired behavior so that the agent learns the correct strategy.
5. The design method of a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm according to claim 3 is characterized by: In step S435, the reward stability includes: S4351: The reward function is stable; the agent will not be unable to learn effectively due to environmental changes or noise interference; S4352: Reward Anatomy; Break down rewards into components to simplify task understanding.
6. The design method of a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm according to claim 3, characterized in that: In S43, the strategy for designing the reward function further includes: S436: Multi-target reward; S437: Exploration-based rewards; S4371: Novelty bonus; S4372: Adaptive reward mechanism.
7. The design method of a multi-constraint guidance law driven by a deep deterministic policy gradient algorithm according to claim 3, characterized in that: In S41, the state space design; In order to fully describe the relative kinematics information, the following state vector is selected: In S42, the action space is designed; Select the fractional sliding surface variables and parameters in the reaching law as actions; A t ={K1,K2,η1,η2} Among them: K1∈(9.5,10.5),K2∈(9.5,10.5),η1∈(0.05,0.15),η2∈(0.05,0.15).
Citation Information
Patent Citations
Impact angle constraint guidance method based on fractional order time-varying sliding mode preset time convergence
CN114706309A
Reinforcement learning guidance control integration method for intercepting three-dimensional maneuvering target
CN118938676A