Attack time constraint guidance method based on hierarchical reinforcement learning and related device
Through a two-layer intelligent agent model based on hierarchical reinforcement learning, the line of sight angle curve and line of sight angular rate are adaptively adjusted to generate guidance acceleration instructions, which solves the guidance problem of high-speed maneuvering targets and achieves accurate hitting and overload control under high-speed maneuvering targets.
Patent Information
- Application Number
- CN202511011340.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-23
AI Technical Summary
When dealing with high-speed maneuvering targets, existing methods make it difficult to accurately estimate the remaining flight time, resulting in interception failures, overload saturation and other problems.
An attack time-constrained guidance method based on hierarchical reinforcement learning is adopted, and flight guidance is performed through a two-layer intelligent agent model. The upper-layer deep reinforcement learning intelligent agent adaptively adjusts the expected line of sight angle curve and line of sight angular rate, and the lower-layer deep reinforcement learning intelligent agent generates guidance acceleration commands to achieve tracking control of the reference trajectory.
There is no need to estimate the remaining flight time, and the attack time constraints can be accurately met, which improves the guidance performance, reduces the risk of overload saturation, adapts to different mission requirements, and enhances the flexibility and accuracy of the guidance system.
Smart Images

Figure CN120686587A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aerospace technology, and in particular relates to an attack time constraint guidance method based on hierarchical reinforcement learning and a related device. Background Art
[0002] With the rapid development of computer, communication and control technologies, multi-aircraft formation coordinated attack targets are attracting attention from academia and engineering fields.
[0003] Currently, the main research approach for attack time-constrained guidance laws is to achieve target impact at the desired attack time by estimating the remaining flight time and controlling the remaining flight time error to zero. However, due to the assumption of small-angle linearization, these existing methods are generally only suitable for attacking stationary or low-speed targets. When dealing with high-speed, highly maneuvering targets, they often suffer from large remaining flight time estimation errors, leading to overload saturation and interception failure. To explain the terminology, an attack time-constrained guidance law independently controls the flight time of each aircraft, ensuring that the aircraft cluster simultaneously impacts the target at the predetermined attack time. This achieves instantaneous saturation attack on the target and maximizes the killing efficiency.
[0004] In summary, the existing methods still have the technical difficulty of accurately estimating the remaining flight time when dealing with high-speed maneuvering targets, and there is an urgent need to develop a new attack time-constrained guidance scheme. Summary of the Invention
[0005] The present invention aims to provide a hierarchical reinforcement learning-based attack time constraint guidance method and related apparatus to address one or more of the aforementioned technical problems. The disclosed technical solution eliminates the need to estimate the remaining flight time, accurately meets the attack time constraint, and provides excellent guidance performance when intercepting high-speed maneuvering targets. This effectively addresses the technical challenge of existing methods that hinder accurate estimation of the remaining flight time when engaging high-speed maneuvering targets.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides an attack time constraint guidance method based on hierarchical reinforcement learning, comprising the following steps:
[0008] Obtain measurement information of the aircraft's seeker on the maneuvering target;
[0009] Based on the acquired measurement information, the trained two-layer intelligent agent model is used for flight guidance to output guidance acceleration commands.
[0010] The training steps of the two-layer agent model include:
[0011] Establishing a relative motion model between the aircraft and the maneuvering target, and converting the guidance problem with an attack time constraint into a mathematical model representation based on the relative motion model; constructing an expected line-of-sight angular rate curve and an expected line-of-sight angle curve based on the established mathematical model representation; and establishing a Markov decision process model for the guidance problem with an attack time constraint based on the mathematical model representation, the expected line-of-sight angular rate curve, and the expected line-of-sight angle curve;
[0012] The two-layer agent model is trained based on the Markov decision process model to obtain a trained two-layer agent model; wherein, the two-layer agent model includes an upper-layer deep reinforcement learning agent and a lower-layer deep reinforcement learning agent; the upper-layer deep reinforcement learning agent is used to serve as a task planner, adaptively adjusting the expected line of sight angle curve and the line of sight angular rate curve according to the seeker measurement information to generate a reference trajectory; the lower-layer deep reinforcement learning agent is used to receive the reference trajectory and the seeker measurement information to generate a guidance acceleration instruction.
[0013] A further improvement of the technical solution of the present invention is that, in the step of constructing the expected sight line angular rate curve and the expected sight line angle curve based on the established mathematical model,
[0014] The expected line of sight angular rate is expressed as:
[0015]
[0016] Where, is the desired line-of-sight angular velocity; δ1 and p are constants, p>1; t is time; T d is the expected attack time; is the initial line of sight angular velocity;
[0017] The desired sight angle is expressed as:
[0018]
[0019] Where q d is the desired sight angle; q0 is the initial sight angle;
[0020] in,
[0021]
[0022] Where δ2 and δ3 are both constants.
[0023] A further improvement of the technical solution of the present invention is that in the step of establishing a Markov decision process model for the guidance problem with attack time constraints based on the mathematical model representation, the expected sight angle rate curve and the expected sight angle curve,
[0024] Design state space for:
[0025]
[0026] Where r,q, They are the relative distance between the target and the aircraft, the sight angle and the sight angular rate respectively;
[0027] For the upper-level deep reinforcement learning agent, design the action space for:
[0028]
[0029] Where, the value range of δ1 is [-0.1, 0.1]; the value range of p is [2, 3];
[0030] Designing the action space for the underlying deep reinforcement learning agent for:
[0031]
[0032] Where a M is the normal acceleration of the aircraft, and its value range is is the maximum normal acceleration;
[0033] Design the terminal reward function R(t f )for:
[0034]
[0035] Where, k1>0 is a constant; t f is the actual attack time; T d is the expected attack time; r f is the terminal miss distance;
[0036] Design the immediate reward function R(t) as:
[0037]
[0038] Wherein, k2 and k3 are both constants, k2>0, k3>0.
[0039] A further improvement of the technical solution of the present invention is that the measurement information includes at least: the relative distance between the maneuvering target and the aircraft, the line of sight angle and the line of sight angular rate.
[0040] A further improvement of the technical solution of the present invention is that the guidance acceleration instruction includes at least: a normal acceleration instruction of the aircraft.
[0041] A further improvement of the technical solution of the present invention is that in the two-layer intelligent agent model, the upper-layer deep reinforcement learning intelligent agent and the lower-layer deep reinforcement learning intelligent agent both have a policy network and a pair of value function networks; the policy network generates action values based on the environmental state input, and the value function network is used to evaluate the value of the action value; wherein, the policy network and the value function network each have a target network with the same structure.
[0042] A further improvement of the technical solution of the present invention is that the step of training the two-layer agent model based on the Markov decision process model to obtain the trained two-layer agent model includes:
[0043] Initialize the upper-layer deep reinforcement learning agent, the lower-layer deep reinforcement learning agent and the sample pool, collect samples and update the network;
[0044] For the upper-level deep reinforcement learning agent, randomly select the sample pool B sup Select Group sample Update the value function network parameters to minimize the error:
[0045]
[0046] Where, Indicates the batch sampling size; Represents the value function network of the upper-level deep reinforcement learning agent; s t represents the state of the environment at time t; represents the action of the upper-level deep reinforcement learning agent at time t; represents the cumulative reward value from time t to t+c, c is a constant; γ represents the discount factor; The target network representing the upper-level deep reinforcement learning agent value function network; s t+1 represents the environmental state at time t+1; represents the target network of the upper-layer deep reinforcement learning agent policy network; ε represents the policy noise;
[0047] If t mod c = 0, update the policy network parameters using policy gradient:
[0048]
[0049] Where, The policy network representing the upper-level deep reinforcement learning agent;
[0050] Update the objective function network:
[0051]
[0052] Where, φ sup′represents the parameters of the target network of the upper-layer deep reinforcement learning agent policy network; κ is the soft update rate; φ sup Represents the parameters of the upper-level deep reinforcement learning agent policy network; Parameters of the target network representing the upper-level deep reinforcement learning agent value function network; Represents the parameters of the upper-level deep reinforcement learning agent value function network;
[0053] For the lower-level deep reinforcement learning agent, randomly select the sample pool B sub Select Group sample Update the value function network parameters to minimize the error:
[0054]
[0055] Where, Indicates the batch sampling size; A value function network representing the underlying deep reinforcement learning agent; represents the action of the underlying deep reinforcement learning agent at time t; R t represents the reward value at time t; A target network representing the underlying deep reinforcement learning agent value function network; represents the action of the upper-level deep reinforcement learning agent at time t+1; A target network representing the underlying deep reinforcement learning agent policy network;
[0056] If t mod c = 0, update the policy network parameters using policy gradient:
[0057]
[0058] Where, A policy network representing the underlying deep reinforcement learning agent;
[0059] Update the objective function network:
[0060]
[0061] Where, φ sub′ represents the parameters of the target network of the underlying deep reinforcement learning agent policy network; φ sub Represents the parameters of the underlying deep reinforcement learning agent policy network; Parameters of the target network representing the underlying deep reinforcement learning agent value function network; Represents the parameters of the underlying deep reinforcement learning agent value function network.
[0062] In a second aspect, the present invention provides an attack time constraint guidance system based on hierarchical reinforcement learning, comprising:
[0063] An information acquisition module is used to obtain measurement information of the aircraft's seeker on the maneuvering target;
[0064] The guidance module is used to perform flight guidance based on the acquired measurement information using the trained two-layer intelligent agent model and output guidance acceleration commands;
[0065] The training steps of the two-layer agent model include:
[0066] Establishing a relative motion model between the aircraft and the maneuvering target, and converting the guidance problem with an attack time constraint into a mathematical model representation based on the relative motion model; constructing an expected line-of-sight angular rate curve and an expected line-of-sight angle curve based on the established mathematical model representation; and establishing a Markov decision process model for the guidance problem with an attack time constraint based on the mathematical model representation, the expected line-of-sight angular rate curve, and the expected line-of-sight angle curve;
[0067] The two-layer agent model is trained based on the Markov decision process model to obtain a trained two-layer agent model; wherein, the two-layer agent model includes an upper-layer deep reinforcement learning agent and a lower-layer deep reinforcement learning agent; the upper-layer deep reinforcement learning agent is used to serve as a task planner, adaptively adjusting the expected line of sight angle curve and the line of sight angular rate curve according to the seeker measurement information to generate a reference trajectory; the lower-layer deep reinforcement learning agent is used to receive the reference trajectory and the seeker measurement information to generate a guidance acceleration instruction.
[0068] In a third aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the program, the attack time constraint guidance method based on hierarchical reinforcement learning as described in any one of the first aspects of the present invention is implemented.
[0069] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the attack time constraint guidance method based on hierarchical reinforcement learning as described in any one of the first aspects of the present invention is implemented.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] The disclosed hierarchical reinforcement learning-based attack time-constrained guidance method designs a line-of-sight angular rate reference trajectory that satisfies boundary conditions such as initial state and terminal constraints. Using a two-layer agent model structure, the upper-layer deep reinforcement learning agent adaptively adjusts the reference trajectory, while the lower-layer deep reinforcement learning agent tracks and controls the reference trajectory. This eliminates the need to estimate the remaining flight time, accurately meeting the attack time constraint and achieving excellent guidance performance when intercepting high-speed, maneuvering targets. Specifically, existing attack time-constrained guidance methods rely on accurate estimation of the remaining flight time, typically using a small-angle linearization assumption. This can lead to large estimation errors when dealing with high-speed, highly maneuvering targets, resulting in overload saturation and poor guidance performance. To address this issue, the present invention designs a line-of-sight angular rate reference trajectory that satisfies boundary conditions such as initial state and terminal constraints. Tracking the reference trajectory enables attack time-constrained guidance, resolving the overload saturation problem caused by large estimation errors and fully and rationally utilizing the available maneuverability of the aircraft to complete combat missions.
[0072] In addition, when the existing technology tracks and controls the reference trajectory, the performance of the controller is greatly affected by the design parameters of the reference trajectory, and improper selection of design parameters will affect the guidance performance; to address this problem, the present invention provides a hierarchical reinforcement learning attack time constraint guidance method that combines a task planning layer with an execution layer. The upper-layer deep reinforcement learning agent optimizes the reference trajectory by adjusting parameters, and the lower-layer deep reinforcement learning agent tracks and controls the reference trajectory. After training and learning, the agent can autonomously optimize the reference trajectory design parameters without human intervention, and has strong adaptability to different task requirements. The hierarchical reinforcement learning attack time constraint guidance method provided by the present invention can achieve hitting high-speed and highly maneuverable targets at the expected attack time in a balanced overload confrontation scenario, and has good guidance performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below; obviously, the drawings described below are some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0074] Figure 1 1 is a flow chart of an attack time constraint guidance method based on hierarchical reinforcement learning in an embodiment of the present invention;
[0075] Figure 2 is a schematic diagram of the relative motion relationship between the aircraft and the target in an embodiment of the present invention;
[0076] Figure 3Schematic diagram of a hierarchical reinforcement learning attack time constraint guidance method according to an embodiment of the present invention;
[0077] Figure 4 1 is a schematic diagram of an intelligent agent structure in an embodiment of the present invention;
[0078] Figure 5 This is a schematic diagram of a sample collection process in an embodiment of the present invention;
[0079] Figure 6 is a schematic diagram of the motion trajectories of an aircraft and a target in an embodiment of the present invention;
[0080] Figure 7 is a schematic diagram of a curve showing a change in the relative distance between an aircraft and a target in an embodiment of the present invention;
[0081] Figure 8 2 is a schematic diagram of an attack time constraint guidance system based on hierarchical reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention; it is obvious that the described embodiments and technical solutions are only part of the embodiments of the present invention, not all of the embodiments.
[0083] All other embodiments obtained by persons of ordinary skill in the art based on the technical solutions disclosed in the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.
[0084] See also Figure 1 The embodiment of the present invention provides an attack time constraint guidance method based on hierarchical reinforcement learning, comprising the following steps:
[0085] Step 1: obtaining measurement information of a maneuvering target by a seeker of an aircraft; in a specific exemplary technical solution, the measurement information includes at least: a relative distance between the maneuvering target and the aircraft, a line of sight angle, and a line of sight angular rate;
[0086] Step 2: Based on the measurement information obtained in step 1, the trained two-layer agent model is used to perform flight guidance and output a guidance acceleration instruction. In a specific exemplary technical solution, the guidance acceleration instruction includes at least: a normal acceleration instruction of the aircraft;
[0087] The training steps of the two-layer agent model include:
[0088] Establishing a relative motion model between the aircraft and the maneuvering target, and converting the guidance problem with attack time constraint into a mathematical model representation based on the relative motion model;
[0089] Based on the established mathematical model representation, constructing an expected sight line angular rate curve and an expected sight line angle curve;
[0090] Establishing a Markov decision process model for a guidance problem with attack time constraints based on the mathematical model representation, the expected line-of-sight angular rate curve, and the expected line-of-sight angle curve;
[0091] The two-layer agent model is trained based on the Markov decision process model to obtain a trained two-layer agent model; wherein, the two-layer agent model includes an upper-layer deep reinforcement learning agent and a lower-layer deep reinforcement learning agent; the upper-layer deep reinforcement learning agent is used to serve as a task planner, adaptively adjust the expected line of sight angle curve and the line of sight angular rate curve according to the seeker measurement information, and generate a reference trajectory; the lower-layer deep reinforcement learning agent is used to receive the reference trajectory and the seeker measurement information, and generate a guidance acceleration instruction.
[0092] In the technical solution disclosed in the embodiment of the present invention, a two-layer intelligent agent model is adopted, in which the upper layer acts as a task planner and the lower layer is responsible for specific execution. This layered structure makes the guidance system more modular and hierarchical. Compared with the traditional single intelligent agent or simple control method, it can more effectively handle complex guidance problems with attack time constraints, thereby improving the overall performance and adaptability of the system. In addition, the present invention provides a suitable theoretical framework for deep reinforcement learning by converting the guidance problem with attack time constraints into a Markov decision process model, making the training and decision-making process of the intelligent agent more standardized and explainable, which is conducive to in-depth understanding and optimization of the guidance strategy. In the present invention, the upper-layer intelligent agent adaptively adjusts the expected line of sight angle curve and line of sight angular rate curve according to the measurement information of the seeker to generate a reference trajectory that is more in line with the actual flight situation; the lower-layer intelligent agent generates guidance acceleration instructions based on the reference trajectory and measurement information, which can more accurately control the flight trajectory of the aircraft, thereby improving the guidance accuracy of the maneuvering target and effectively improving the hit rate of the attack. Taking the attack time constraint into consideration during the guidance process, the collaborative work of the two-layer intelligent agent model can rationally plan the flight path of the aircraft to ensure accurate attack on the target within the specified time, meet the time requirements of the combat mission, and enhance the timeliness and flexibility of the operation.
[0093] See also Figures 2 to 7As a preferred embodiment of the technical solution of the present invention, the training process of the two-layer agent model specifically includes the following steps:
[0094] Step 1: Establish a relative motion model between the aircraft and the maneuvering target, and transform the guidance problem with attack time constraint into a mathematical model based on the relative motion model.
[0095] like Figure 2 As shown, in the specific exemplary technical solution, Figure 2 Center axis OX I Y I represents the reference inertial system, T and M represent the maneuvering target and the aircraft respectively. When gravity and drag are neglected and both parties are assumed to be flying at a constant speed, the relative motion model between the aircraft and the maneuvering target is established as follows:
[0096]
[0097] Where r and q represent the relative distance and line of sight angle between the target and the aircraft respectively; V T and θ T Respectively represent the target's speed and trajectory inclination; V M and θ M Respectively represent the speed and ballistic inclination of the aircraft; a T and a M denote the normal acceleration of the target and the aircraft respectively; σ T =θ T -q and σ M =θ M -q represents the lead angle of the target and the aircraft, that is, the angle between the velocity vector and the line of sight.
[0098] In addition, the guidance problem with attack time constraint is expressed as:
[0099]
[0100] Where r f is the terminal miss distance; r(t f ) is t f The off-target amount function; t f is the actual attack time; T d is the expected attack time.
[0101] Step 2, based on the mathematical model established in step 1, an embodiment of the present invention proposes a line of sight angular rate shaping method, constructs an expected line of sight angular rate curve and an expected line of sight angle curve to meet boundary conditions such as initial state and terminal constraints, and achieves hitting the target at the expected attack time.
[0102] Specific exemplary technical solutions include:
[0103] The desired sight line angular rate is designed as a function of time t, expressed as:
[0104]
[0105] Where, is the desired line-of-sight angular rate; δ1, δ2, δ3 and p are all constants, and p>1.
[0106] The boundary conditions are constructed based on the initial state and terminal constraints and are expressed as:
[0107]
[0108] Where t0 is the initial time; is the initial line of sight angular velocity; To ensure the terminal conditions for the aircraft to successfully hit the target.
[0109] The values of δ2 and δ3 can be determined according to the boundary conditions:
[0110]
[0111] Substituting equation (5) into equation (3) yields the expected line-of-sight angular velocity curve:
[0112]
[0113] Integrating both sides of equation (6) simultaneously yields the desired sight angle curve:
[0114]
[0115] Where q0 is the initial sight angle.
[0116] In the technical solution disclosed in the embodiment of the present invention, by adjusting the values of p and δ1 and designing a hierarchical reinforcement learning attack time constraint guidance algorithm, the desired line of sight angle and angular rate curve are tracked, without estimating the remaining flight time, so as to achieve hitting the maneuvering target at the expected attack time, and avoid overload saturation caused by excessive error in the remaining flight time estimation.
[0117] Step 3: Establish an attack time constraint guidance algorithm framework based on hierarchical reinforcement learning, that is, build a two-layer intelligent agent model.
[0118] like Figure 3 As shown in the specific exemplary technical solution, the designed algorithm framework has a two-layer structure, involving two deep reinforcement learning agents; wherein,
[0119] The two-layer agent model includes:
[0120] The upper-layer deep reinforcement learning agent is used as a task planner to adaptively adjust the desired sight angle and sight angular rate curve by changing the values of p and δ1 based on the measurement information of the seeker to generate a reference trajectory;
[0121] The lower-level deep reinforcement learning agent is used to receive the reference trajectory and seeker measurement information, generate guidance acceleration commands to track the reference trajectory, and achieve terminal constraints such as zero miss distance and attack time.
[0122] The technical solution of the embodiment of the present invention realizes adaptive adjustment and tracking of the desired line of sight angle and angular rate curve by designing a two-layer intelligent model structure, which can reduce the uncertainty caused by manual parameter adjustment and enhance the interpretability of deep reinforcement learning guidance strategy.
[0123] like Figure 4 As shown, in the specific exemplary technical solution, Figure 4 The structure of the upper and lower deep reinforcement learning agents is shown, each of which has a policy network μ φ and a pair of value function networks The policy network generates action values based on the environment state input, and the value function network is used to evaluate the value of the action value; the policy network and the value function network each have a target network μ with the same structure. φ′ 、 and Used to improve the smoothness of training.
[0124] Step 4: Based on the mathematical model established in steps 1 and 2 and the expected sight angle and sight angle rate curves, a Markov decision process model for the guidance problem with attack time constraints is established.
[0125] In a specific exemplary technical solution,
[0126] Design state space for:
[0127]
[0128] Where r,q, They are the relative distance between the target and the aircraft, the line of sight angle and the line of sight angular rate respectively.
[0129] For the upper-level deep reinforcement learning agent, design the action space for:
[0130]
[0131] Where, the value range of δ1 is [-0.1, 0.1]; the value range of p is [2, 3].
[0132] Designing the action space for the underlying deep reinforcement learning agent for:
[0133]
[0134] Where a M is the normal acceleration of the aircraft, and its value range is
[0135] The terminal reward function is designed as:
[0136]
[0137] In the formula, k1>0 is a constant.
[0138] According to the expected gaze angle and gaze angle rate curve established in step 2, the immediate reward function is designed as follows:
[0139]
[0140] Wherein, k2>0 and k3>0 are constants.
[0141] Step 5: Based on the algorithm framework and Markov decision process model designed in Steps 3 and 4, a hierarchical reinforcement learning attack time constraint guidance strategy training algorithm is established; wherein,
[0142] The specific implementation process of the training algorithm is as follows:
[0143] 1) Algorithm initialization, including:
[0144] Initialize the upper-level deep reinforcement learning agent policy network Parameter φ sup , value function network and parameter Initialize the target network parameters φ sup′ ←φ sup ,
[0145] Initialize the policy network of the underlying deep reinforcement learning agent Parameter φ sub , value function network and parameter Initialize the target network parameters φ sub′ ←φ sub , Initialize sample pool B sup and B sub .
[0146] 2) Collect samples, including:
[0147] like Figure 5 As shown, at time t = 0, the upper-level deep reinforcement learning agent is based on the environment state s0 and the policy network Output Action Generate the desired sight angle q by determining the values of δ1 and p d and line-of-sight angular rate Curve; The lower-level deep reinforcement learning agent receives the desired sight angle and sight angular rate curve instructions, based on the environment state s0 and the upper-level deep reinforcement learning agent action As input, output action Generate guidance instructions. After the interceptor executes the instructions, the environment state changes to s1, and an immediate reward value R0 is fed back. When 0≤t<c (c>0 is a constant), the upper-level deep reinforcement learning agent action remains unchanged, that is, The lower-level deep reinforcement learning agent continues to generate actions based on the environment state s1 After interacting with the environment, the environment state is transferred to s2, and the reward value R1 is fed back. This cycle repeats until t mod c = 0, when the upper-layer deep reinforcement learning agent updates the action Then update the expected sight angle q d and line-of-sight angular rate curve; the underlying deep reinforcement learning agent interacts with the environment based on the new reference trajectory signal.
[0148] The upper-level deep reinforcement learning agent collects samples every c time steps:
[0149]
[0150] The underlying deep reinforcement learning agent collects samples at each time step:
[0151]
[0152] 3) Update the network: For the upper-level deep reinforcement learning agent, randomly select the sample pool B sup Select Group sample Update the value function network parameters to minimize the error:
[0153]
[0154] Where γ represents the discount factor and ε represents the policy noise.
[0155] If t mod c = 0, update the policy network parameters using policy gradient:
[0156]
[0157] Update the objective function network:
[0158]
[0159] Where κ is the soft update rate.
[0160] For the lower-level deep reinforcement learning agent, randomly select the sample pool B sub Select Group sample Update the value function network parameters to minimize the error:
[0161]
[0162] If t mod c = 0, update the policy network parameters using policy gradient:
[0163]
[0164] Update the objective function network:
[0165]
[0166] In a specific embodiment of the present invention, an attack time constraint guidance method based on hierarchical reinforcement learning is provided, which can effectively solve the difficulty of accurately estimating the remaining flight time when dealing with high-speed maneuvering targets in existing methods, and realize attack time constraint guidance for high-speed maneuvering targets in a balanced overload confrontation scenario.
[0167] See also Figure 6 and Figure 7 , specifically, for example, training is performed under the scenarios shown in Table 1, and the target takes a sinusoidal maneuver The maximum available maneuvering acceleration of the aircraft is Expected attack time T d The value is taken every 0.1s in the range [38,42]s. Figure 6 is the motion trajectory of the aircraft and the target in this embodiment, Figure 7 The relative distance variation curve between the aircraft and the target in this embodiment is shown in the figure. It can be seen from the figure that the aircraft can d The expected attack time of =[38,42]s hits the maneuvering target, with an average miss distance of 0.01m and an average attack time error of 0.076s, which meets the guidance accuracy requirements.
[0168] Table 1. Scene parameters
[0169] parameter unit Value Relative distance m 80000 Sight angle deg 0 Aircraft speed m / s 2000 Aircraft trajectory inclination deg 0 Target speed m / s 1500 Target trajectory inclination deg 120
[0170] The following are device embodiments of the present invention, which can be used to perform the method embodiments of the present invention. For details not disclosed in the device embodiments, please refer to the method embodiments of the present invention.
[0171] See also Figure 8 In an embodiment of the present invention, an attack time constraint guidance system based on hierarchical reinforcement learning is provided, comprising:
[0172] An information acquisition module is used to obtain measurement information of the aircraft's seeker on the maneuvering target;
[0173] The guidance module is used to perform flight guidance based on the acquired measurement information using the trained two-layer intelligent agent model and output guidance acceleration commands;
[0174] The training steps of the two-layer agent model include:
[0175] Establishing a relative motion model between the aircraft and the maneuvering target, and converting the guidance problem with an attack time constraint into a mathematical model representation based on the relative motion model; constructing an expected line-of-sight angular rate curve and an expected line-of-sight angle curve based on the established mathematical model representation; and establishing a Markov decision process model for the guidance problem with an attack time constraint based on the mathematical model representation, the expected line-of-sight angular rate curve, and the expected line-of-sight angle curve;
[0176] The two-layer agent model is trained based on the Markov decision process model to obtain a trained two-layer agent model; wherein, the two-layer agent model includes an upper-layer deep reinforcement learning agent and a lower-layer deep reinforcement learning agent; the upper-layer deep reinforcement learning agent is used to serve as a task planner, adaptively adjusting the expected line of sight angle curve and the line of sight angular rate curve according to the seeker measurement information to generate a reference trajectory; the lower-layer deep reinforcement learning agent is used to receive the reference trajectory and the seeker measurement information to generate a guidance acceleration instruction.
[0177] In one embodiment of the present invention, a computer device is provided, comprising a processor and a memory, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to perform the operation of the attack time constraint guidance method based on hierarchical reinforcement learning.
[0178] In one embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device, used to store programs and data. It is understood that the computer-readable storage medium herein may include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides storage space, which stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for being loaded and executed by a processor. These instructions may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium herein may be a high-speed random access memory (RAM) or a non-volatile memory, such as at least one disk storage device. The processor may load and execute the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the attack time constraint guidance method based on hierarchical reinforcement learning in the above-mentioned embodiment.
[0179] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) that contain computer-usable program code.
[0180] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0181] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. An attack time constraint guidance method based on hierarchical reinforcement learning, characterized in that: The following steps are involved: Obtain measurement information of the aircraft's seeker on the maneuvering target; Based on the acquired measurement information, the trained two-layer intelligent agent model is used for flight guidance to output guidance acceleration commands. The training steps of the two-layer agent model include: Establishing a relative motion model between the aircraft and the maneuvering target, and converting the guidance problem with an attack time constraint into a mathematical model representation based on the relative motion model; constructing an expected line-of-sight angular rate curve and an expected line-of-sight angle curve based on the established mathematical model representation; and establishing a Markov decision process model for the guidance problem with an attack time constraint based on the mathematical model representation, the expected line-of-sight angular rate curve, and the expected line-of-sight angle curve; The two-layer agent model is trained based on the Markov decision process model to obtain a trained two-layer agent model; wherein, the two-layer agent model includes an upper-layer deep reinforcement learning agent and a lower-layer deep reinforcement learning agent; the upper-layer deep reinforcement learning agent is used to serve as a task planner, adaptively adjusting the expected line of sight angle curve and the line of sight angular rate curve according to the seeker measurement information to generate a reference trajectory; the lower-layer deep reinforcement learning agent is used to receive the reference trajectory and the seeker measurement information to generate a guidance acceleration instruction.
2. The attack time constraint guidance method based on hierarchical reinforcement learning according to claim 1 is characterized in that: In the step of constructing the expected sight line angular rate curve and the expected sight line angle curve based on the established mathematical model, The expected line of sight angular rate is expressed as: Where, is the desired line-of-sight angular velocity; δ1 and p are constants, p>1; t is time; T d is the expected attack time; is the initial line of sight angular velocity; The desired sight angle is expressed as: Where q d is the desired sight angle; q0 is the initial sight angle; in, Where δ2 and δ3 are both constants.
3. The attack time constraint guidance method based on hierarchical reinforcement learning according to claim 1 is characterized in that: In the step of establishing a Markov decision process model for the guidance problem with attack time constraints based on the mathematical model representation, the expected sight angle rate curve and the expected sight angle curve, Design state space for: Where r,q, They are the relative distance between the target and the aircraft, the sight angle and the sight angular rate respectively; For the upper-level deep reinforcement learning agent, design the action space for: Where, the value range of δ1 is [-0.1, 0.1]; the value range of p is [2, 3]; Designing the action space for the underlying deep reinforcement learning agent for: Where a M is the normal acceleration of the aircraft, and its value range is is the maximum normal acceleration; Design the terminal reward function R(t f )for: Where, k1>0 is a constant; t f is the actual attack time; T d is the expected attack time; r f is the terminal miss distance; Design the immediate reward function R(t) as: Wherein, k2 and k3 are both constants, k2>0, k3>0.
4. The attack time constraint guidance method based on hierarchical reinforcement learning according to claim 3 is characterized in that: The measurement information includes at least: the relative distance between the maneuvering target and the aircraft, the sight angle, and the sight angular rate.
5. The attack time constraint guidance method based on hierarchical reinforcement learning according to claim 3 is characterized in that: The guidance acceleration instruction at least includes: a normal acceleration instruction of the aircraft.
6. The attack time constraint guidance method based on hierarchical reinforcement learning according to claim 1 is characterized in that: In the two-layer agent model, both the upper-layer deep reinforcement learning agent and the lower-layer deep reinforcement learning agent have a policy network and a pair of value function networks; the policy network generates action values based on the environment state input, and the value function network is used to evaluate the value of the action value; wherein, the policy network and the value function network each have a target network with the same structure.
7. The attack time constraint guidance method based on hierarchical reinforcement learning according to claim 6, characterized in that: The step of training the two-layer agent model based on the Markov decision process model to obtain the trained two-layer agent model includes: Initialize the upper-layer deep reinforcement learning agent, the lower-layer deep reinforcement learning agent and the sample pool, collect samples and update the network; For the upper-level deep reinforcement learning agent, randomly select the sample pool B sup Select Group sample Update the value function network parameters to minimize the error: Where, Indicates the batch sampling size; Represents the value function network of the upper-level deep reinforcement learning agent; s t represents the state of the environment at time t; represents the action of the upper-level deep reinforcement learning agent at time t; represents the cumulative reward value from time t to t+c, c is a constant; γ represents the discount factor; The target network representing the upper-level deep reinforcement learning agent value function network; s t+1 represents the environmental state at time t+1; represents the target network of the upper-layer deep reinforcement learning agent policy network; ε represents the policy noise; If t mod c = 0, update the policy network parameters using policy gradient: Where, The policy network representing the upper-level deep reinforcement learning agent; Update the objective function network: Where, φ sup′ represents the parameters of the target network of the upper-layer deep reinforcement learning agent policy network; κ is the soft update rate; φ sup Represents the parameters of the upper-level deep reinforcement learning agent policy network; Parameters of the target network representing the upper-level deep reinforcement learning agent value function network; Represents the parameters of the upper-level deep reinforcement learning agent value function network; For the lower-level deep reinforcement learning agent, randomly select the sample pool B sub Select Group sample Update the value function network parameters to minimize the error: Where, Indicates the batch sampling size; A value function network representing the underlying deep reinforcement learning agent; represents the action of the underlying deep reinforcement learning agent at time t; R t represents the reward value at time t; A target network representing the underlying deep reinforcement learning agent value function network; represents the action of the upper-level deep reinforcement learning agent at time t+1; A target network representing the underlying deep reinforcement learning agent policy network; If t mod c = 0, update the policy network parameters using policy gradient: Where, A policy network representing the underlying deep reinforcement learning agent; Update the objective function network: Where, φ sub′ represents the parameters of the target network of the underlying deep reinforcement learning agent policy network; φ sub Represents the parameters of the underlying deep reinforcement learning agent policy network; Parameters of the target network representing the underlying deep reinforcement learning agent value function network; Represents the parameters of the underlying deep reinforcement learning agent value function network.
8. An attack time constraint guidance system based on hierarchical reinforcement learning, characterized in that: include: An information acquisition module is used to obtain measurement information of the aircraft's seeker on the maneuvering target; The guidance module is used to perform flight guidance based on the acquired measurement information using the trained two-layer intelligent agent model and output guidance acceleration commands; The training steps of the two-layer agent model include: Establishing a relative motion model between the aircraft and the maneuvering target, and converting the guidance problem with an attack time constraint into a mathematical model representation based on the relative motion model; constructing an expected line-of-sight angular rate curve and an expected line-of-sight angle curve based on the established mathematical model representation; and establishing a Markov decision process model for the guidance problem with an attack time constraint based on the mathematical model representation, the expected line-of-sight angular rate curve, and the expected line-of-sight angle curve; The two-layer agent model is trained based on the Markov decision process model to obtain a trained two-layer agent model; wherein, the two-layer agent model includes an upper-layer deep reinforcement learning agent and a lower-layer deep reinforcement learning agent; the upper-layer deep reinforcement learning agent is used to serve as a task planner, adaptively adjusting the expected line of sight angle curve and the line of sight angular rate curve according to the seeker measurement information to generate a reference trajectory; the lower-layer deep reinforcement learning agent is used to receive the reference trajectory and the seeker measurement information to generate a guidance acceleration instruction.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the attack time constraint guidance method based on hierarchical reinforcement learning according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the attack time constraint guidance method based on hierarchical reinforcement learning according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Attack-time-constraint guidance-law design method of intercepting maneuvering target
CN108416098A
Multi-unmanned aerial vehicle air combat decision-making method based on multi-agent layered reinforcement learning
CN115291625A
Cooperative guidance method containing time and angle constraints based on reinforcement learning
CN118210229A
Attack time constraint guidance method based on line-of-sight angular rate shaping
CN118534915A
Unmanned aerial vehicle cluster task planning algorithm based on hierarchical multi-agent deep reinforcement learning and evaluation method thereof
CN119088073A