A reinforcement learning enhanced generalized guidance method and system with impact time constraint for intercepting maneuvering targets
By combining backstepping control and reinforcement learning, a generalized strike time guidance law was designed, which solved the problem of intercepting maneuvering targets in three-dimensional space and achieved effective interception and optimal energy control of maneuvering targets.
Patent Information
- Application Number
- CN202510246049.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Existing technologies struggle to effectively design time-of-attack guidance laws for intercepting maneuvering targets in a three-dimensional combat space. Especially when considering target maneuverability, existing methods require complex numerical optimization and manual adjustments, and cannot guarantee optimal system energy consumption.
A generalized strike time guidance law is designed using a backstepping control method. Combined with a numerical iterative algorithm for predicting future relative average velocity and a reinforcement learning algorithm, the time-varying control law gain is obtained through agent training, thereby achieving the interception of maneuvering targets.
It can clearly and definitively intercept stationary, moving, and maneuvering targets in a three-dimensional combat space, has versatility, and has optimal energy consumption.
Smart Images

Figure CN120085546B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent guidance control of aircraft, in particular to a reinforcement learning enhanced generalized guidance method and system for intercepting a maneuvering target with strike time constraint. BACKGROUND
[0002] Proportional navigation guidance law is widely recognized in academia and industry due to its simplicity and effectiveness. However, with the increasing deployment of advanced anti-missile systems on high-value targets, traditional PNG control law, which mainly focuses on minimizing the terminal miss distance, cannot threaten high-value targets. To solve this problem, salvo attack emerges as a promising strategy, which improves the survivability of missile clusters and the penetration ability of defense systems. Under the premise of setting the same strike time for all missiles, individual strike time guidance law can achieve salvo attack tasks and saturate enemy defense systems.
[0003] Individual strike time guidance law can be divided into guidance law requiring residual flight time and guidance law not requiring residual flight time according to whether residual flight time is needed. The research on strike time guidance law for stationary targets is already mature, but the research on strike time guidance law for intercepting maneuvering targets in three-dimensional combat space is less, and most of the strike time guidance law for intercepting maneuvering targets is developed in two-dimensional space.
[0004] Patent publication CN118534915A discloses a method of forming a line-of-sight angular rate in a two-dimensional engagement plane to achieve a desired attack time on a maneuvering target, but this guidance law needs to repeatedly adjust one parameter in the desired line-of-sight angle to achieve the desired attack time, which is a predictive generalized guidance method. Patent publication CN108416098B discloses an attack time constraint guidance law for intercepting a maneuvering target in a two-dimensional engagement space, which designs a desired line-of-sight angular rate and uses a sliding mode control method to track the desired line-of-sight angular rate to achieve time constraint on intercepting a maneuvering target, but needs to repeatedly adjust and optimize the parameters in the desired line-of-sight angular rate profile to achieve attack time constraint, which is a predictive guidance law. K.S.Erer, M, et al. introduced a relative guidance system and designed a strike time constraint guidance law for intercepting a maneuvering target in a two-dimensional plane.
[0005] The reasons for this situation are as follows: firstly, the two-dimensional strike time guidance problem is simpler than the three-dimensional strike time guidance problem, without considering the complex coupling problem of pitch and yaw channels. Secondly, for a stationary target, an exact analytical form of the remaining flight time can be given, but for a mobile target, it is difficult to give an estimate of the remaining flight time of the vehicle. The estimate of the remaining flight time has a great influence on the accuracy of the guidance law. In addition, the controller gain of the existing strike time guidance law needs to be manually adjusted artificially, and the use process is complex, which cannot guarantee the optimal energy consumption of the system.
[0006] The above existing patents and papers only develop in the two-dimensional engagement plane, and need repeated numerical optimization. K.S.Erer, M, etc. also only develop in the two-dimensional engagement plane. Therefore, it is necessary to design a generalized strike time constraint guidance law in the three-dimensional combat space considering target maneuvering. To this end, the present application provides a reinforcement learning enhanced strike time constraint generalized guidance method and system for intercepting mobile targets. SUMMARY
[0007] The present application aims to provide a reinforcement learning enhanced strike time constraint generalized guidance method and system for intercepting mobile targets to solve the problems raised in the background art.
[0008] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a reinforcement learning enhanced strike time constraint generalized guidance method for intercepting mobile targets, comprising the following steps:
[0009] Step 1, based on the backstepping control method, a generalized strike time constraint guidance law for intercepting mobile targets is determined;
[0010] Step 2, a numerical iterative algorithm for predicting future relative average speed is proposed;
[0011] Step 3, based on the reinforcement learning algorithm, a strike time guidance law with time-varying control law gain is designed;
[0012] Step 4, based on the generalized strike time guidance law of step 1, the future relative average speed prediction algorithm of step 2 and the agent with time-varying control law gain trained in step 3, a strike time guidance law for intercepting mobile targets is designed.
[0013] Preferably, the generalized strike time guidance law of step 1 is designed in the relative vector guidance coordinate system, and the specific engagement model can be described as: the velocity vector and control acceleration vector of the missile are V M and A M , and the velocity vector and control acceleration of the target are VT and A T , the line-of-sight vector between the missile and the target is R, the vector and is a unit vector, the relative velocity vector is V R = V M - V T , the unit vector of the relative velocity is the relative acceleration vector is A R = A M - A T , the relative acceleration vector is perpendicular to the relative velocity vector, then the following formula is established:
[0014]
[0015] wherein, Ω L and are the rotation angular velocity of the line-of-sight vector and the rotation angular velocity of the relative velocity vector respectively, the relative lead angle is defined as σ R = a cos (r·v R ), the derivative of the relative lead angle is,
[0016]
[0017] wherein
[0018] Preferably: the step 1 based on the backstepping control method determines a generalized attack time guidance law of the interception maneuvering target, and specifically comprises:
[0019] Step 1-1, two error variables are defined as
[0020]
[0021] wherein, T d is the expected attack time, t is the actual flight time of the missile, is the future relative average velocity;
[0022] Step 1-2, two new error variables are defined as,
[0023]
[0024] wherein φ (x 1e ) is a virtual control input quantity to realize the convergence of x 1e ;
[0025] Step 1-3, the derivative of formula (6) can be obtained,
[0026]
[0027] Step 1-4, the guidance law is designed as:
[0028]
[0029] wherein τ>0 is a controller gain;
[0030] Step 1-5, the function φ(x 1e ) used in step 1-4 needs to have the following properties: ① φ(x 1e ) and are continuous on the interval [0, +∞); ② when x 1e > 0, φ(x 1e ) is strictly monotonically increasing, and φ(0) = 0.
[0031] Preferably, the numerical iterative algorithm for predicting the future relative average speed in step 2 specifically comprises:
[0032] Step 2-1, when the impact time error x 1e is zero, the impact time constraint guidance law is transformed into
[0033]
[0034] Step 2-2, define a new guidance state vector as
[0035]
[0036] Step 2-3, take the derivative of the guidance state vector in step 2-2 with respect to the missile-target distance R, and have
[0037]
[0038] wherein
[0039]
[0040] Step 2-4, the superscript C represents the current state of the aircraft, H represents the number of numerical iterations, and the iteration step size is Δ h = -R C / H; the state vector x and the missile-target distance R at the discrete nodes are represented as x h and R h , respectively, wherein h = {0, 1,..., H-1}; the future relative average speed prediction algorithm is as follows
[0041] Step 2-4-1, set Δ h = -R C / H;
[0042] Step 2-4-2, set h = 0;
[0043] Step 2-4-3, when h < H, execute Step 2-4-4 to Step 2-4-7;
[0044] Step 2-4-4, calculate f(x
[0045] Step 2-4-5, calculate f(x h ) according to formula (12);
[0046] Step 2-4-6, calculate x h+1 = x h + Δ h f(x h );
[0047] Step 2-4-7, calculate h = h + 1;
[0048] Step 2-4-8, calculate D go = x H (2)-x C (2);
[0049] Step 2-4-9, calculate t go = x H (1)-x C (1);
[0050] Step 2-4-10, calculate the predicted future average speed
[0051] Preferably, the step 3 is based on a reinforcement learning algorithm, and the time-varying control law gain is designed to be a time-to-impact guidance law, specifically comprising:
[0052] Step 3-1, design the agent state space as (R, σ R , x 1e , x 2e );
[0053] Step 3-2, design the action space as the gain τ of the control law;
[0054] Step 3-3, design the reward function as
[0055]
[0056] wherein
[0057]
[0058] wherein ω1 and ω2 are weights.
[0059] The generalized guidance system of the reinforced learning enhanced strike time constraint for intercepting a maneuvering target according to the above comprises a vehicle and target kinematics calculation unit, a vehicle real-time guidance law calculation unit, an intelligent agent real-time adaptive control law gain calculation unit and a vehicle guidance control unit, wherein:
[0060] The vehicle and target kinematics calculation unit is used for extracting guidance kinematics information of the vehicle and the target.
[0061] The vehicle real-time guidance law calculation unit is used for calculating a control overload amount of the vehicle according to a designed strike time guidance law in real time.
[0062] The intelligent agent real-time adaptive control law gain calculation unit is used for calculating an adaptive gain of the control law according to a real-time state vector of the vehicle.
[0063] The vehicle guidance control unit is used for guiding the vehicle to intercept the maneuvering target at a desired strike time by using the generalized strike time guidance law of the reinforced learning enhanced intercepting maneuvering target.
[0064] Compared with the prior art, the method has the following beneficial effects:
[0065] The generalized guidance method and system of the reinforced learning enhanced strike time constraint for intercepting a maneuvering target provided by the application obtain a generalized strike time constraint guidance law and a future relative speed prediction method based on a backstepping control method, obtain a control gain of the guidance law by using an intelligent agent trained by inputting a real-time guidance state, and then obtain a control amount for driving the vehicle, so that the maneuvering target can be intercepted at a desired time. The solving process is clear and definite, and the method can be applied to intercepting stationary, moving and maneuvering targets in a three-dimensional combat space. The guidance law is a generalized strike time constraint guidance law, and a user can design a suitable virtual function φ(x 1e ) according to the needs, and the method has universality. BRIEF DESCRIPTION OF DRAWINGS
[0066] Figure 1 It is a three-dimensional vector guidance coordinate system of a missile and a target in the application;
[0067] Figure 2 It is a three-dimensional relative vector guidance coordinate system of a missile and a target in the application;
[0068] Figure 3 It is a flowchart of the generalized guidance method and system of the reinforced learning enhanced strike time constraint in the application;
[0069] Figure 4 It is a change curve of a training reward of the intelligent agent in the application;
[0070] Figure 5 for the aircraft and target motion trajectory plot;
[0071] Figure 6 for the aircraft and target motion trajectory plot;
[0072] Figure 7 for the aircraft and target motion trajectory plot;
[0073] Figure 8 for the aircraft and target motion trajectory plot.
[0074] Figure 9 for the aircraft and target motion trajectory plot. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0076] EMBODIMENT
[0077] Please refer to Figures 1-9 A reinforcement learning enhanced generalized guidance method of strike time constraint for intercepting a maneuvering target, comprising the following steps:
[0078] Step 1, a generalized strike time guidance law for intercepting a maneuvering target is derived based on the backstepping control method;
[0079] Step 2, a numerical iteration algorithm for predicting the future relative average speed is proposed;
[0080] Step 3, a strike time guidance law with time-varying control law gain is designed based on the reinforcement learning algorithm;
[0081] Step 4, a typical strike time guidance law for intercepting a maneuvering target is designed based on the generalized strike time guidance law in step 1, the future relative average speed prediction algorithm in step 2 and the time-varying control law gain of the agent trained in step 3.
[0082] Figure 1 for the three-dimensional relative guidance coordinate system of the missile and the target, combined with Figure 1 , the engagement model between the missile and the moving target is specifically described as:
[0083] The velocity vector and the control acceleration vector of the missile are V M and A M , and the velocity vector and the control acceleration of the target are V TAnd A T , the line-of-sight vector between the missile and the target is R, the vector And is a unit vector, the relative velocity vector is V R = V M -V T , the unit vector of the relative velocity is The relative acceleration vector is A R = A M -A T , the relative acceleration vector is perpendicular to the relative acceleration vector, and the following relationship holds
[0084]
[0085] Where, Ω L And are the rotation rate of the line-of-sight vector and the rotation rate of the relative velocity vector, respectively, and the relative lead angle is defined as σ R = a cos (r·v R ), the derivative of the relative lead angle is
[0086]
[0087] Where
[0088] Further, the step 1 based on the backstepping control method, a generalized strike time guidance law for intercepting maneuvering targets is derived, which includes:
[0089] Step 1-1, define two error variables as
[0090]
[0091] Where, T d is the desired strike time, t is the actual flight time of the missile, is the future relative average speed.
[0092] Step 1-2, define two new error variables as
[0093]
[0094] Where φ(x 1e ) is a virtual control input to achieve the convergence of x 1e .
[0095] Step 1-3, the derivative of formula (6) can be obtained
[0096]
[0097] Step 1-4, in order to ensure x 1e and x 2e convergence, the guidance law is designed as:
[0098]
[0099] Where τ>0 is the controller gain.
[0100] Step 1-5, the φ(x 1e ) function used in step 1-4 needs to have the following properties: ① φ(x 1e ) and are continuous on the interval [0, +∞); ② when x 1e > 0, φ(x 1e ) is strictly monotonically increasing, and φ(0) = 0.
[0101] Further, a numerical iterative algorithm for predicting future relative average speed in step 2, specifically includes:
[0102] Step 2-1, when the impact time error x 1e is zero, the impact time constraint guidance law converges to
[0103]
[0104] Step 2-2, define a new guidance state vector as
[0105]
[0106] Step 2-3, the guidance state vector in step 2-2 is differentiated with respect to the missile-target distance R, and has
[0107]
[0108] Where,
[0109]
[0110] Step 2-4, the superscript C represents the current state of the aircraft, H represents the number of numerical iterations, and the iteration step size is Δ h = -R C / H, the state vector x and the missile-target distance R at the discrete nodes are denoted as x h and R h , where h = {0, 1, …, H-1}, then the future relative average speed prediction algorithm is as follows:
[0111] Step 2-4-1, set Δ h = -R C / H;
[0112] Step 2-4-2, set h = 0;
[0113] Step 2-4-3, when h < H, execute Step 2-4-4 to Step 2-4-7;
[0114] Step 2-4-4, calculate according to formula (9)
[0115] Step 2-4-5, calculate f(x h ) according to formula (12)
[0116] Step 2-4-6, calculate x h+1 = x h + Δ h f(x h );
[0117] Step 2-4-7, calculate h = h + 1;
[0118] Step 2-4-8, calculate D go = x H (2) - x C (2);
[0119] Step 2-4-9, calculate t go = x H (1) - x C (1);
[0120] Step 2-4-10, calculate the predicted future average speed
[0121] Further, the reinforcement learning algorithm-based Step 3 designs a time-varying control law gain impact time guidance law, which specifically includes:
[0122] Step 3-1, design the agent state space as (R, σ R , x 1e , x 2e );
[0123] Step 3-2, design the action space as the gain τ of the control law;
[0124] Step 3-3, design the reward function as
[0125]
[0126] wherein
[0127]
[0128] wherein ω1 and ω2 are weights.
[0129] Further, the reinforcement learning algorithm-based step 3 specifically adopts the proximal policy optimization algorithm (PPO).
[0130] Further, based on the step 1-based generalized strike time guidance law, the step 2 future relative average speed prediction algorithm and the step 3 trained time-varying control law gain of the agent of step 4, a strike time guidance law for intercepting a maneuvering target is designed, specifically including:
[0131] Step 4-1, a φ(x 1e ) meeting the properties of step 1-5 is designed;
[0132] Step 4-2, the expected strike time T d of the aircraft is designed;
[0133] Step 4-3, based on the step 1-based generalized strike time guidance law and the step 2 future relative average speed prediction algorithm, the flight state (R, σ R , x 1e , x 2e ) of the aircraft is input to the step 3 trained time-varying control law gain of the agent to obtain real-time aircraft control, which is input to the kinematic iteration equation of the aircraft;
[0134] Step 4-4, repeat step 4-3 until the aircraft successfully intercepts the maneuvering target.
[0135] A reinforcement learning enhanced strike time constrained guidance system for intercepting a maneuvering target, comprising: an aircraft and target kinematics calculation unit, an aircraft real-time guidance law calculation unit, an agent real-time adaptive control law gain calculation unit and an aircraft guidance control unit; wherein:
[0136] The aircraft and target kinematics calculation unit is used to extract the guidance kinematics information of the aircraft and the target;
[0137] The aircraft real-time guidance law calculation unit calculates the aircraft control overload in real time according to the designed strike time guidance law;
[0138] The agent real-time adaptive control law gain calculation unit calculates the adaptive gain of the control law in real time according to the real-time state vector of the aircraft;
[0139] The aircraft guidance control unit uses the reinforcement learning enhanced generalized strike time guidance law for intercepting a maneuvering target to guide the aircraft to intercept the maneuvering target at the expected strike time, and establishes a strike time constrained guidance law.
[0140] Then, taking a single aircraft intercepting a maneuvering target as an example, φ(x 1e ) is designed as Where k1 = 0.005 and k2 = 0.22, the expected strike time is set to 60s, the missile's initial velocity is 300m / s, and the velocity direction is (cos-10°cos35°, sin-10°cos35°, sin35°), the target's velocity is 15m / s, and the velocity direction is (1,0,0). The target's maneuver is A. T = 2cos(0.2t)m / s 2 This simulation case only tests the impact time T. d =60s, other impact time constraints only require changing the impact time T of the binding. d That's all. Figures 5-9 Simulation results for guidance of a single missile.
[0141] like Figure 4 The curves showing the change in reward during the training of the controller gain in a reinforcement learning agent are presented. It can be seen that the agent's reward can increase and remain convergent. Figure 6 The curve depicting the change in missile-target distance shows that the aircraft was able to successfully intercept the maneuvering target within 60 seconds, thus achieving the strike time constraint. Figure 8 The curve of the gain of the real-time controller of the reinforcement learning agent is given. It first increases and then decreases to optimize the energy consumption of the system.
[0142] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0143] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A reinforcement learning enhanced strike time constrained generalized guidance method for intercepting a maneuvering target, characterized by, The method comprises the following steps: Step 1, a generalized strike time constraint guidance law for intercepting a maneuvering target is determined based on a backstepping control method; Step 2, a numerical iterative algorithm for predicting future relative average speed is provided; Step 3, a strike time guidance law with time-varying control law gain is designed based on a reinforcement learning algorithm; Step 4, a strike time guidance law for intercepting a maneuvering target is designed based on the generalized strike time guidance law of step 1, the future relative average speed prediction algorithm of step 2 and the agent with trained time-varying control law gain of step 3; The generalized strike time guidance law for intercepting a maneuvering target determined based on the backstepping control method of step 1 specifically comprises: Step 1-1, two error variables are defined as , wherein, , is the desired impact time, is the actual flight time of the missile, is the future relative average speed; is the missile-target distance; Step 1-2, two new error variables are defined as , wherein is a virtual control input quantity to achieve convergence; Steps 1-3, Derivation of the formula Taking the derivative of the formula, we get, , Step 1-4, the guidance law is designed as , wherein, Kc is the controller gain; Step 1-5, Step 1-4 The function needs to have the following properties: ① on the interval , and are continuous; ② when , is strictly monotonically increasing, and ; the relative velocity vector is ; and are the rotation rates of the line-of-sight vector and the relative velocity vector, respectively, is the relative lead angle; , is the unit vector of the relative velocity, is the unit vector, is the line-of-sight vector between the missile and the target.
2. The reinforcement learning enhanced, strike time constrained, generalized guidance method of intercepting a maneuvering target of claim 1, wherein: The generalized hit time guidance law of step 1 is designed in the relative vector guidance coordinate system, and a specific engagement model can be described as follows: the velocity vector and the control acceleration vector of the missile are and respectively, the velocity vector and the control acceleration of the target are and respectively, the line-of-sight vector between the missile and the target is , the vectors , and are unit vectors, the relative velocity vector is , the unit vector of the relative velocity is , the relative acceleration vector is , and the relative acceleration vector is perpendicular to the relative acceleration vector, so the following formula is established: , , , where the relative lead angle is defined as Taking the derivative of the relative lead angle, we have, 。 3. The reinforcement learning enhanced, strike time constrained, generalized guidance method of intercepting a maneuvering target of claim 1, wherein: The numerical iterative algorithm for predicting future relative average speed of step 2 specifically comprises: Step 2-1, When the impact time error is zero, the impact time constrained guidance law is transformed into , Step 2-2, a new guidance state vector is defined as , Step 2-3, Derivative of the guidance state vector with respect to the missile-target range Taking the derivative, we have , Wherein, , Step 2-4, superscript representing the current state of the aircraft, denoting the number of value iterations, the iteration step size is , the state vector and the missile-target distance are respectively denoted as and , where , then the future relative average speed prediction algorithm is as follows Step 2 - 4 - 1, Set up ; Step 2 - 4 - 2, Setting up ; Step 2-4-3, when Steps 2-4-4 to 2-4-7 are performed. Step 2 - 4-4, according to the formula Calculation ; Step 2 - 4-5, according to the formula calculation ; Step 2 - 4-6, calculation ; Step 2 - 4-7, calculation ; Step 2 - 4-8, calculation ; Step 2 - 4-9, calculation ; Step 2 - 4-10, calculate predicted future average speed .
4. The reinforcement learning enhanced, strike time constrained, generalized guidance method of intercepting a maneuvering target of claim 3, wherein: The strike time guidance law with time-varying control law gain designed based on the reinforcement learning algorithm of step 3 specifically comprises: Step 3-1, design the agent state space as ; Step 3-2, design action space as gain for control law ; Step 3-3, a reward function is designed as , Wherein , wherein and are weights.
5. A guidance system employing the reinforcement learning enhanced time-to- strike constrained generalized guidance method of intercepting a maneuvering target according to any one of claims 1-4, characterized in that, It comprises: An aircraft and target kinematics calculation unit, an aircraft real-time guidance law calculation unit, an agent real-time adaptive control law gain calculation unit and an aircraft guidance control unit; wherein: The aircraft and target kinematics calculation unit is used for extracting guidance kinematics information of the aircraft and the target; The aircraft real-time guidance law calculation unit is used for calculating a control overload of the aircraft in real time according to the designed strike time guidance law; The agent real-time adaptive control law gain calculation unit is used for calculating an adaptive gain of the control law in real time according to a real-time state vector of the aircraft; The aircraft guidance control unit is used for guiding the aircraft to intercept the maneuvering target at a desired strike time by using the generalized strike time guidance law for intercepting a maneuvering target enhanced by reinforcement learning, so as to establish a strike time constraint guidance law.
Citation Information
Patent Citations
A method for designing attack time-constrained guidance laws for intercepting maneuvering targets
CN108416098B
Attack time constraint guidance method based on line-of-sight angular rate shaping
CN118534915A
Attack-time-constraint guidance-law design method of intercepting maneuvering target
CN108416098A
Nonlinear optimal flight time control guidance method for intercepting maneuvering target
CN116679743A