Reinforced learning enhanced strike time constrained generalized guidance method and system for intercepting maneuvering target

Through reinforcement learning-enhanced strike time constraint guidance method, combined with reverse step control and numerical iteration algorithm, a generalized strike time guidance law suitable for three-dimensional combat space was designed, which solved the problem of insufficient research on the strike time guidance law of maneuvering targets in the existing technology, achieved interception of maneuvering targets within the expected time, and optimized system energy consumption.

CN120085546AActive Publication Date: 2025-06-03HARBIN INST OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510246049.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing technology has little research on the strike time guidance law of intercepting maneuverable targets in three-dimensional combat space, and the controller gain of the existing strike time guidance law requires manual adjustment, and the use process is complicated and the system energy consumption cannot be guaranteed.

Method used

A generalized guidance method for enhanced strike time constraints for intercepting maneuverable targets is proposed, including generalized strike time guidance law based on reverse step control method, numerical iterative algorithm for predicting future relative average velocity, and time-varying control law gain design based on reinforcement learning algorithm.

Benefits of technology

It realizes intercepting maneuverable targets within the expected time, which is suitable for three-dimensional combat space, with clear solution process, optimal energy consumption, and universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085546A_ABST
    Figure CN120085546A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent guidance control of aircrafts, and discloses a generalized guidance method and system of strike time constraint with reinforcement learning enhancement for intercepting a maneuvering target, and the method comprises the following specific steps: firstly, designing a generalized strike time guidance law for intercepting the maneuvering target based on a backstepping method; then designing a numerical iterative algorithm capable of predicting future average speed, designing a strike time constraint guidance law of adaptive controller gain based on a reinforcement learning algorithm, and finally designing a guidance law for intercepting a maneuvering target based on a trained adaptive gain intelligent agent. The aircraft is guided to intercept the maneuvering target at a desired time. Guidance control is realized by adopting a method of combining a reinforcement learning algorithm and a traditional control algorithm, and the method has very important practical significance for realizing a guidance target for intercepting a maneuvering target at specified time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent guidance control of aircraft, and specifically to a generalized guidance method and system for enhancing the strike time constraint of intercepting maneuvering targets by reinforcement learning. Background Technique

[0002] The proportional navigation guidance law has been widely recognized in both academia and industry due to its simplicity and effectiveness. However, with the increasing deployment of advanced anti-missile systems on high-value targets, the traditional PNG control law, which mainly focuses on minimizing the terminal miss distance, can no longer pose a threat to high-value targets. To solve this problem, salvo attack has emerged as a promising strategy, which improves the survivability of missile clusters and the penetration ability against defense systems. On the premise of setting the same strike time for all missiles, the individual strike time guidance law can achieve the salvo attack mission and conduct a saturation strike on the enemy's defense system.

[0003] The individual strike time guidance law can be divided into the guidance law that requires the remaining flight time and the guidance law that does not require the remaining flight time according to whether the remaining flight time is needed. Currently, the research on the strike time guidance law for stationary targets has been very mature. However, in the three-dimensional combat space, there is less research on the strike time guidance law for intercepting maneuvering targets, and most of the strike time guidance laws for intercepting maneuvering targets are carried out in two-dimensional space.

[0004] Patent Publication No. CN118534915A discloses a method for shaping the line-of-sight angular rate in a two-dimensional engagement plane to achieve the strike of a maneuvering target at the desired attack time. However, this guidance law needs to repeatedly adjust a parameter in the desired line-of-sight angle to achieve the desired strike time, which is a predictive generalized guidance method. Patent Publication No. CN108416098B discloses an attack time constraint guidance law for intercepting maneuvering targets in a two-dimensional engagement space. In this guidance law, the desired line-of-sight angular rate is designed, and the sliding mode control method is used to track the desired line-of-sight angular rate to achieve the interception time constraint of the maneuvering target. However, it is necessary to repeatedly adjust and optimize the parameters in the desired line-of-sight angular rate profile to achieve the strike time constraint, which is a predictive guidance law. K.S. Erer, M et al. introduced a relative guidance system and designed a strike time constraint guidance law for intercepting maneuvering targets in a two-dimensional plane.

[0005] There are two reasons for this situation: first, the two-dimensional strike time guidance problem is simpler than the three-dimensional strike time guidance problem, and there is no need to consider the complex coupling problem of pitch and yaw channels. Second, for stationary targets, the precise analytical form of the remaining flight time can be given, but for maneuvering targets, it is difficult to give an estimate of the remaining flight time of the aircraft. The estimation of the remaining flight time has a great impact on the accuracy of the guidance law. In addition, the controller gain of the existing strike time guidance law needs to be manually adjusted, the use process is relatively complicated, and the system energy consumption cannot be guaranteed to be optimal.

[0006] The above-mentioned existing patents and papers only constrain the strike time of the target. CN118534915A and CN108416098B are only developed in the two-dimensional engagement plane, and repeated numerical optimization is required. KSErer, M et al. also only develop in the two-dimensional engagement plane. Therefore, it is necessary to design a generalized strike time constraint guidance law in the three-dimensional combat space while considering the target maneuverability. To this end, the present invention proposes a reinforcement learning enhanced strike time constraint generalized guidance method and system for intercepting maneuvering targets. Summary of the invention

[0007] The object of the present invention is to provide a reinforcement learning enhanced strike time constrained generalized guidance method and system for intercepting maneuvering targets, so as to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solution: a reinforcement learning enhanced strike time constraint generalized guidance method for intercepting maneuvering targets, comprising the following steps:

[0009] Step 1: Based on the backstepping control method, a generalized strike time constraint guidance law for intercepting maneuvering targets is determined;

[0010] Step 2, a numerical iterative algorithm for predicting the future relative average speed is proposed;

[0011] Step 3: Based on the reinforcement learning algorithm, design the strike time guidance law with time-varying control law gain;

[0012] Step 4: Based on the generalized strike time guidance law of step 1, the future relative average speed prediction algorithm of step 2, and the intelligent agent of the time-varying control law gain trained in step 3, design the strike time guidance law for intercepting maneuvering targets.

[0013] Preferably, the generalized strike time guidance law of step 1 is designed in the relative vector guidance coordinate system, and the specific engagement model can be described as follows: the velocity vector and the control acceleration vector of the missile are V M and A M , the target velocity vector and control acceleration are VT and A T , the line-of-sight vector between the missile and the target is R, and the vectors and are unit vectors, and the relative velocity vector is V R = V M -V T , and the unit vector of the relative velocity is The relative acceleration vector is A R = A M -A T , if the relative acceleration vector is perpendicular to the relative acceleration vector, then the following equation holds:

[0014]

[0015] where, Ω L and are the angular rates of rotation of the line-of-sight vector and the relative velocity vector respectively, and the relative lead angle is defined as σ R = acos(r·v R ), taking the derivative of the relative lead angle, we have,

[0016]

[0017] where

[0018] Preferably: The generalized strike-time guidance law for intercepting a maneuvering target determined by the backstepping control method in step 1 specifically includes:

[0019] Step 1-1: Define two error variables as

[0020]

[0021] where, T d is the desired strike time, t is the actual flight time of the missile, is the future relative average velocity;

[0022] Step 1-2: Define two new error variables as,

[0023]

[0024] where φ(x 1e ) is a virtual control input quantity to achieve the convergence of x 1e ;

[0025] Step 1-3: Taking the derivative of formula (6), we can obtain,

[0026]

[0027] Steps 1-4, the guidance law is designed as:

[0028]

[0029] where τ > 0 is the controller gain;

[0030] Steps 1-5, the φ(x 1e ) function used in Step 1-4 needs to have the following properties: ① On the interval [0, +∞), φ(x 1e ) and are continuous; ② When x 1e > 0, φ(x 1e ) is strictly monotonically increasing, and φ(0) = 0.

[0031] Preferably: The numerical iteration algorithm for predicting the future relative average speed described in Step 2 specifically includes:

[0032] Step 2-1, when the impact time error x 1e is zero, the impact time constraint guidance law is transformed into

[0033]

[0034] Step 2-2, define the new guidance state vector as

[0035]

[0036] Step 2-3, take the derivative of the guidance state vector in Step 2-2 with respect to the missile-to-target distance R, and there is

[0037]

[0038] where,

[0039]

[0040] Step 2-4, the superscript C represents the current state of the aircraft, H represents the number of numerical iterations, the iteration step size is Δ h =-R C / H, the state vector x and the missile-to-target distance R are respectively represented as x h and R h at the discrete nodes, where h = {0, 1,..., H-1}, then the future relative average speed prediction algorithm is as follows

[0041] Step 2-4-1, set Δ h =-R C / H;

[0042] Step 2-4-2, set h = 0;

[0043] Step 2-4-3: When h < H, execute Steps 2-4-4 to 2-4-7;

[0044] Step 2-4-4: Calculate according to formula (9)

[0045] Step 2-4-5: Calculate f(x h ) according to formula (12);

[0046] Step 2-4-6: Calculate x h+1 = x h + Δ h f(x h );

[0047] Step 2-4-7: Calculate h = h + 1;

[0048] Step 2-4-8: Calculate D go = x H (2) - x C (2);

[0049] Step 2-4-9: Calculate t go = x H (1) - x C (1);

[0050] Step 2-4-10: Calculate the predicted future average speed

[0051] Preferably: The strike time guidance law with time-varying control law gain designed based on the reinforcement learning algorithm in Step 3 specifically includes:

[0052] Step 3-1: Design the agent state space as (R, σ R , x 1e , x 2e );

[0053] Step 3-2: Design the action space as the gain τ of the control law;

[0054] Step 3-3: Design the reward function as

[0055]

[0056] where

[0057]

[0058] where ω 1 and ω 2 are weights.

[0059] According to the above generalized guidance system with enhanced strike time constraint for intercepting maneuvering targets based on reinforcement learning, it includes: an aircraft and target kinematics calculation unit, an aircraft real-time guidance law calculation unit, an agent real-time adaptive control law gain calculation unit, and an aircraft guidance and control unit; where:

[0060] The aircraft and target kinematics calculation unit is used to extract the guidance kinematic information of the aircraft and the target;

[0061] The aircraft real-time guidance law calculation unit calculates the aircraft control overload in real time according to the designed strike time guidance law;

[0062] The agent real-time adaptive control law gain calculation unit calculates the adaptive gain of the control law in real time according to the real-time state vector of the aircraft;

[0063] The aircraft guidance and control unit uses the generalized strike time guidance law for intercepting maneuvering targets enhanced by reinforcement learning to guide the aircraft to intercept the maneuvering target at the expected strike time and establish a strike time constraint guidance law.

[0064] Compared with the prior art, the beneficial effects of the present invention are:

[0065] The generalized guidance method and system with enhanced strike time constraint for intercepting maneuvering targets based on reinforcement learning proposed by the present invention obtains the generalized strike time constraint guidance law and the future relative velocity prediction method based on the backstepping control method. By inputting the real-time guidance state value into the trained agent, the control gain of the guidance law is obtained, and then the control amount for driving the aircraft is obtained, realizing the interception of the maneuvering target within the expected time. The solution process is clear and definite, and it can be adapted to the interception of stationary, moving, and maneuvering targets in a three-dimensional engagement space. Moreover, this guidance law is a generalized strike time constraint guidance law, and users can design a suitable virtual function φ(x 1e ) according to their own needs, which has universality. Description of the Drawings

[0066] Figure 1 It is the three-dimensional vector guidance coordinate system of the missile and the target of the present invention;

[0067] Figure 2 It is the three-dimensional relative vector guidance coordinate system of the missile and the target of the present invention;

[0068] Figure 3 It is the flow chart of the generalized guidance method and system with enhanced strike time constraint for intercepting maneuvering targets based on reinforcement learning of the present invention;

[0069] Figure 4 It is the change curve of the reinforcement learning agent training reward of the present invention;

[0070] Figure 5 It is a trajectory diagram of the aircraft and the target;

[0071] Figure 6 It is a curve diagram of the change in the missile-target distance;

[0072] Figure 7 It is a curve diagram of the overload in the pitch and yaw directions of the aircraft;

[0073] Figure 8 It is a curve diagram of the controller gain calculated in real time by the reinforcement learning agent.

[0074] Figure 9 It is a curve diagram of the energy consumption of the aircraft. Specific implementation manner

[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0076] Embodiment

[0077] Please refer to Figures 1-9 , a generalized guidance method for enhancing the strike time constraint of intercepting a maneuvering target based on reinforcement learning, including the following steps:

[0078] Step 1: Based on the backstepping control method, a generalized strike time guidance law for intercepting a maneuvering target is derived;

[0079] Step 2: A numerical iterative algorithm for predicting the future relative average speed is proposed;

[0080] Step 3: Based on the reinforcement learning algorithm, a strike time guidance law with time-varying control law gain is designed;

[0081] Step 4: Based on the generalized strike time guidance law in Step 1, the future relative average speed prediction algorithm in Step 2, and the agent with time-varying control law gain trained in Step 3, a typical strike time guidance law for intercepting a maneuvering target is designed.

[0082] Figure 1 This is the three-dimensional relative guidance coordinate system of the missile and the target of the present invention. Combining Figure 1 , the engagement model between the missile and the moving target is specifically described as:

[0083] The velocity vector and control acceleration vector of the missile are V M and A M , the velocity vector and control acceleration of the target are V Tand A T , the line-of-sight vector between the missile and the target is R, and the vectors and are unit vectors, and the relative velocity vector is V R = V M - V T , and the unit vector of the relative velocity is The relative acceleration vector is A R = A M - A T , the relative acceleration vector is perpendicular to the relative acceleration vector, and the following relationship holds

[0084]

[0085] where, Ω L and are the angular rate of rotation of the line-of-sight vector and the angular rate of rotation of the relative velocity vector respectively, and the relative lead angle is defined as σ R = acos(r·v R ), taking the derivative of the relative lead angle, we have

[0086]

[0087] where

[0088] Furthermore, based on the backstepping control method in step 1, a generalized hitting-time guidance law for intercepting maneuvering targets is derived, which specifically includes:

[0089] Step 1-1: Define two error variables as

[0090]

[0091] where, T d is the expected hitting time, t is the actual flight time of the missile, is the future relative average velocity.

[0092] Step 1-2: Define two new error variables as

[0093]

[0094] where φ(x 1e ) is the virtual control input quantity to achieve the convergence of x 1e .

[0095] Step 1-3: Taking the derivative of formula (6), we can obtain

[0096]

[0097] Steps 1-4. To ensure the convergence of x 1e and x 2e , the guidance law is designed as:

[0098]

[0099] where τ > 0 is the controller gain.

[0100] Steps 1-5. The φ(x 1e ) function used in Step 1-4 needs to have the following properties: ① On the interval [0, +∞), φ(x 1e ) and are continuous; ② When x 1e > 0, φ(x 1e ) is strictly monotonically increasing, and φ(0) = 0.

[0101] Furthermore, a numerical iteration algorithm for predicting the future relative average velocity in Step 2 specifically includes:

[0102] Steps 2-1. When the impact time error x 1e is zero, the impact time constraint guidance law converges to

[0103]

[0104] Steps 2-2. Define a new guidance state vector as

[0105]

[0106] Steps 2-3. Take the derivative of the guidance state vector in Step 2-2 with respect to the missile-to-target distance R, and we have

[0107]

[0108] where

[0109]

[0110] Steps 2-4. The superscript C represents the current state of the aircraft, H represents the number of numerical iterations, the iteration step size is Δ h = -R C / H, the state vector x and the missile-to-target distance R are respectively represented as x h and R h at discrete nodes, where h = {0, 1,..., H - 1}, then the future relative average velocity prediction algorithm is as follows:

[0111] Steps 2-4-1. Set Δ h = -R C / H;

[0112] Step 2-4-2: Set h = 0;

[0113] Step 2-4-3: When h < H, execute Steps 2-4-4 to 2-4-7;

[0114] Step 2-4-4: Calculate according to formula (9)

[0115] Step 2-4-5: Calculate f(x h );

[0116] Step 2-4-6: Calculate x h+1 = x h + Δ h f(x h );

[0117] Step 2-4-7: Calculate h = h + 1;

[0118] Step 2-4-8: Calculate D go = x H (2) - x C (2);

[0119] Step 2-4-9: Calculate t go = x H (1) - x C (1);

[0120] Step 2-4-10: Calculate the predicted future average speed

[0121] Furthermore, based on the reinforcement learning algorithm in Step 3, a hitting time guidance law with time-varying control law gain is designed, specifically including:

[0122] Step 3-1: Design the agent state space as (R, σ R , x 1e , x 2e );

[0123] Step 3-2: Design the action space as the gain τ of the control law;

[0124] Step 3-3: Design the reward function as

[0125]

[0126] where

[0127]

[0128] where, ω 1 and ω 2 are weights.

[0129] Furthermore, for the reinforcement learning algorithm in step 3, the Proximal Policy Optimization algorithm (PPO) is specifically adopted.

[0130] Furthermore, based on the intelligent agent with the generalized strike time guidance law in step 1, the future relative average velocity prediction algorithm in step 2, and the time-varying control law gain completed in step 3 training, a strike time guidance law for intercepting a maneuvering target is designed, specifically including:

[0131] Step 4-1: Design a φ(x that conforms to the properties of steps 1-5 1e );

[0132] Step 4-2: Design the desired strike time T of the aircraft d ;

[0133] Step 4-3: Based on the generalized strike time guidance law in step 1 and the future relative average velocity prediction algorithm in step 2, input the obtained flight state (R, σ R , x 1e , x 2e ) of the aircraft into the intelligent agent with the time-varying control law gain completed in step 3 training to obtain the real-time aircraft control amount, and input it into the kinematic iteration equation of the aircraft;

[0134] Step 4-4: Repeat step 4-3 until the aircraft successfully intercepts the maneuvering target.

[0135] A guidance system for intercepting a maneuvering target with a strike time constraint enhanced by reinforcement learning includes: an aircraft and target kinematic calculation unit, an aircraft real-time guidance law calculation unit, an intelligent agent real-time adaptive control law gain calculation unit, and an aircraft guidance control unit; where:

[0136] The aircraft and target kinematic calculation unit is used to extract the guidance kinematic information of the aircraft and the target;

[0137] The aircraft real-time guidance law calculation unit calculates the control overload of the aircraft in real time according to the designed strike time guidance law;

[0138] The intelligent agent real-time adaptive control law gain calculation unit calculates the adaptive gain of the control law in real time according to the state vector of the real-time aircraft;

[0139] The aircraft guidance control unit uses the generalized strike time guidance law for intercepting a maneuvering target enhanced by reinforcement learning to guide the aircraft to intercept the maneuvering target at the desired strike time and establish a strike time constraint guidance law.

[0140] Then, taking a single aircraft intercepting a maneuvering target as an example, φ(x 1e ) is designed as where k1 = 0.005 and k 2 =0.22, the expected strike time is set to 60s, the missile's initial speed is 300m / s, the speed direction is (cos-10°cos35°, sin-10°cos35°, sin35°), the target speed is 15m / s, the speed direction is (1,0,0), and the target maneuver is A T =2cos(0.2t)m / s 2 , this simulation case only tests the strike time T d = 60s, other strike time constraints only need to change the binding strike time T d That's it, Figures 5-9 This is the simulation result of a single missile guidance.

[0141] like Figure 4 The reward curve of the reinforcement learning agent during the training controller gain process is given. It can be seen that the agent reward can increase and maintain convergence. Figure 6 The variation curve of the missile-target distance is described. It can be seen that the aircraft can successfully intercept the maneuvering target within 60 seconds, achieving the strike time constraint. Figure 8 The change curve of the gain of the real-time controller of the reinforcement learning agent is given, which increases first and then decreases, which can optimize the energy consumption of the system.

[0142] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0143] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A reinforcement learning enhanced strike time-constrained generalized guidance method for intercepting maneuvering targets, characterized in that: The steps include: Step 1: Based on the backstepping control method, a generalized strike time constraint guidance law for intercepting maneuvering targets is determined; Step 2: A numerical iterative algorithm is proposed to predict the future relative average speed; Step 3: Based on the reinforcement learning algorithm, design the strike time guidance law with time-varying control law gain; Step 4: Based on the generalized strike time guidance law of step 1, the future relative average speed prediction algorithm of step 2, and the intelligent agent of the time-varying control law gain trained in step 3, design the strike time guidance law for intercepting maneuvering targets.

2. The reinforcement learning enhanced strike time constraint generalized guidance method for intercepting maneuvering targets according to claim 1, characterized in that: The generalized strike time guidance law in step 1 is designed in the relative vector guidance coordinate system. The specific engagement model can be described as follows: the missile velocity vector and the control acceleration vector are V M and A M , the target velocity vector and control acceleration are V T and A T , the line of sight vector between the missile and the target is R, and the vector and is a unit vector, and the relative velocity vector is V R =V M -V T , the unit vector of relative velocity is The relative acceleration vector is A R =A M -A T , the relative acceleration vector is perpendicular to the relative acceleration vector, then the following formula holds: Among them, Ω L and are the angular rates of rotation of the sight vector and the relative velocity vector, respectively. The relative lead angle is defined as σ R =acos(r·v R ), and taking the derivative of the relative lead angle, we have, in 3. The reinforcement learning enhanced strike time constraint generalized guidance method for intercepting maneuvering targets according to claim 2 is characterized by: The step 1, based on the backstepping control method, determines a generalized strike time guidance law for intercepting a maneuvering target, specifically including: Step 1-1, define two error variables respectively in, T d is the expected attack time, t is the actual flight time of the missile, is the future relative average speed; Step 1-2: Define two new error variables: Where φ(x 1e ) is the virtual control input to realize x 1e Convergence of Step 1-3: Derivative formula (6) to obtain: Step 1-4: Guidance law design is: Where τ>0 is the controller gain; Step 1-5, φ(x used in step 1-4 1e ) function needs to have the following properties: ① In the interval [0, +∞), φ(x 1e )and is continuous; ② when x 1e >0, φ(x 1e ) is strictly monotonically increasing, and φ(0)=0.

4. The reinforcement learning enhanced strike time constraint generalized guidance method for intercepting maneuvering targets according to claim 3 is characterized by: The numerical iterative algorithm for predicting the future relative average speed described in step 2 specifically includes: Step 2-1, when the strike time error x 1e When it is zero, the strike time constraint guidance law is transformed into Step 2-2, define a new guidance state vector as Step 2-3: Derivation of the guidance state vector in step 2-2 with respect to the missile-target distance R yields in, Step 2-4, the superscript C represents the current state of the aircraft, H represents the number of numerical iterations, and the iteration step is Δ h =-R C / H, the state vector x and the missile-target distance R are expressed as x in discrete nodes respectively. h and R h , where h = {0, 1, ..., H-1}, then the future relative average speed prediction algorithm is as follows Step 2-4-1, set Δ h =-R C / H; Step 2-4-2, set h = 0; Step 2-4-3, when h<H, execute steps 2-4-4 to 2-4-7; Step 2-4-4: Calculate according to formula (9) Step 2-4-5: Calculate f(x) according to formula (12) h ); Step 2-4-6, calculate x h+1 =x h +Δ h f(x h ); Step 2-4-7, calculate h=h+1; Step 2-4-8, calculate D go =x H (2)-x C (2); Step 2-4-9, calculate t go =x H (1)-x C (1); Step 2-4-10: Calculate the predicted future average speed 5. The reinforcement learning enhanced strike time constraint generalized guidance method for intercepting maneuvering targets according to claim 4 is characterized by: The strike time guidance law of the time-varying control law gain designed based on the reinforcement learning algorithm in step 3 specifically includes: Step 3-1: Design the agent state space as (R,σ R ,x 1e ,x 2e ); Step 3-2, design the action space as the gain τ of the control law; Step 3-3, design the reward function as in Among them, ω1 and ω2 are weights.

6. A reinforcement learning enhanced strike time-constrained generalized guidance system for intercepting maneuvering targets according to any one of claims 1 to 5, characterized in that: include: Aircraft and target kinematics calculation unit, aircraft real-time guidance law calculation unit, intelligent agent real-time adaptive control law gain calculation unit and aircraft guidance control unit; wherein: Aircraft and target kinematics calculation unit, used to extract guidance kinematics information of aircraft and target; The aircraft real-time guidance law calculation unit calculates the aircraft control overload in real time according to the designed strike time guidance law; The intelligent agent real-time adaptive control law gain calculation unit calculates the adaptive gain of the control law in real time according to the state vector of the real-time aircraft; The aircraft guidance control unit uses the generalized strike time guidance law for intercepting maneuvering targets enhanced by reinforcement learning to guide the aircraft to intercept the maneuvering target at the desired strike time and establish a strike time constraint guidance law.

Citation Information

Patent Citations

  • A method for designing attack time-constrained guidance laws for intercepting maneuvering targets

    CN108416098B

  • Attack time constraint guidance method based on line-of-sight angular rate shaping

    CN118534915A

  • Attack-time-constraint guidance-law design method of intercepting maneuvering target

    CN108416098A

  • Nonlinear optimal flight time control guidance method for intercepting maneuvering target

    CN116679743A

  • Hypersonic aircraft three-dimensional guidance method based on interference utilization technology

    CN117348402A