A guidance method and system based on time-angle constraints for predictive correction guidance

By combining deep learning and reinforcement learning models, time- and angle-coordinated guidance for the aircraft was achieved, solving the problem of poor guidance law performance under constant velocity assumption and improving the accuracy and robustness of guidance and control.

CN117302554BActive Publication Date: 2026-03-24BEIJING INST OF TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Most existing guidance laws are based on the assumption of constant velocity, which leads to poor time and angle constraint control and makes it difficult to meet the requirements of precise time and angle coordination constraints.

Method used

A deep learning network model is used to predict the terminal state of the aircraft when it reaches the target, and a deep reinforcement learning model is used to continuously correct the error between the predicted state and the desired state, so as to achieve the coordinated constraint of time and angle. The guidance command is corrected by using deep neural networks and deep reinforcement learning models.

Benefits of technology

It achieves precise time and angle control of the aircraft, improves the control effect and robustness of the guidance law, is applicable to different types of aircraft, and has low computational complexity and high computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117302554B_ABST
    Figure CN117302554B_ABST
Patent Text Reader

Abstract

The application discloses a time-angle constraint guidance method and system based on a predictive correction guidance. The method comprises the following steps: S101, before the flight of an aircraft, the expected flight time and the expected landing angle of the aircraft are preset; S102, during the flight of the aircraft, the flight time error and the landing angle error of the aircraft are obtained in real time, wherein the flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle; and S103, the guidance instruction is corrected based on the flight time error and the landing angle error of the aircraft. The time-angle constraint guidance method can overcome the defects of the constant speed assumption and has good control effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aircraft control technology, and in particular to a guidance method and system based on time angle constraints of predictive correction guidance. Background Technology

[0002] Guidance laws are directly related to an aircraft's range, firing accuracy, survivability, and damage effects, and are one of the important aspects of aircraft design.

[0003] Currently, there is a lot of research on guidance laws with only angle constraints, but relatively little research on guidance laws with only time constraints, and even less research on guidance laws with both time and angle constraints.

[0004] Existing time-angle coordination constraints mainly fall into two categories: (1) coordinating the predicted arrival time and predicted arrival angle of each aircraft through inter-aircraft communication; and (2) setting equal expected arrival time and expected arrival angle for each aircraft before launch. However, regardless of the approach taken, it is necessary to accurately control the remaining flight time and angle of each aircraft. To address this issue, most existing guidance laws are based on the assumption of constant velocity, transforming the time constraint into a constraint on the remaining flight path length, and achieving angle control on this basis. However, the remaining flight time is related to the aircraft velocity, and the guidance law proposed based on the assumption of constant velocity has poor practical application results. Summary of the Invention

[0005] To address the problems existing in the prior art, the present invention provides a guidance method and system based on time angle constraints of predictive correction guidance.

[0006] To achieve the above objectives, in a first aspect, the present invention provides a guidance method based on time angle constraints of predictive correction guidance, comprising the following steps:

[0007] Step S101: Before the aircraft takes flight, the desired flight time and desired landing angle are pre-set on the aircraft;

[0008] Step S102: During the flight of the aircraft, the flight time error and landing angle error of the aircraft are acquired in real time, wherein the flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle.

[0009] Step S103: Correct the guidance command based on the flight time error and landing angle error of the aircraft.

[0010] Secondly, the present invention provides a guidance system based on time angle constraints for predictive correction guidance, comprising:

[0011] The preset module is used to pre-set the desired flight time and desired landing angle of the aircraft before it takes flight;

[0012] The error acquisition module is used to acquire the flight time error and landing angle error of the aircraft in real time during the flight process. The flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle.

[0013] The correction module is used to correct guidance commands based on the aircraft's flight time error and landing angle error.

[0014] The beneficial effects of the guidance method and system based on time angle constraints for predictive correction guidance of the present invention include:

[0015] (1) The time-angle constrained guidance method of the present invention can overcome the defects of relying on constant velocity assumption and has good control effect;

[0016] (2) The time-angle constrained guidance method of the present invention can be applied to different types of aircraft and has broad application prospects;

[0017] (3) The time-angle constraint guidance method of the present invention has low computational complexity and high computational efficiency, and can be used for online implementation. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a guidance method based on time angle constraints for predictive correction guidance according to the present invention.

[0019] Figure 2 This is a schematic diagram of the deep neural network model structure in this invention;

[0020] Figure 3 This is a schematic diagram of the structure of a guidance system based on time angle constraints for predictive correction guidance according to the present invention;

[0021] Figure 4 A simulation diagram showing the constraint-time reward obtained after training a deep reinforcement learning model 100 times;

[0022] Figure 5 A simulation diagram showing the constraint angle reward obtained after training a deep reinforcement learning model 100 times;

[0023] Figure 6 The image shows the simulation results of 1000 simulations for Embodiment 1 of the present invention. Detailed Implementation

[0024] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.

[0025] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0026] Time-angle coordinated guidance enables multiple aircraft to reach a target from different directions. This process requires precise control of the remaining flight time and angle of impact for each aircraft. Most existing guidance laws are based on the assumption of constant velocity, transforming the time constraint into a constraint on the remaining flight path length, and then controlling the angle of impact. However, the remaining flight time is related to the aircraft's velocity, and guidance laws based on the constant velocity assumption often result in poor control performance.

[0027] To address the problems of existing technologies, this invention provides a guidance method based on time-angle constraints in predictive correction guidance. This method utilizes a deep learning network model to predict the terminal state of the aircraft upon reaching the target, and then uses a deep reinforcement learning model to continuously correct the error between the predicted state and the desired state until both time and angle constraints are simultaneously satisfied, thereby obtaining a guidance command that meets the constraints and minimizes energy. Simulations in various scenarios have verified that the guidance command exhibits good performance and strong robustness.

[0028] In this invention, flight simulation programs are used to obtain various data of the aircraft during flight.

[0029] On the one hand, simulated flight tests of aircraft can employ a hardware-in-the-loop (HIL) platform, where the aircraft's flight control system is a physical entity, including a flight control computer and inertial measurement units (accelerometers, gyroscopes, and magnetometers, etc.), while the aircraft's GPS and target detection sensors (such as electro-optical pods and radars) and flight environment (i.e., atmosphere, terrain, etc.) are entirely virtual. On the other hand, the simulation environment in simulated flight tests can also be completely virtual, meaning both the aircraft's flight environment and flight control system are virtual.

[0030] Specifically, the aircraft's position is obtained by the GPS positioning system, which includes the aircraft's current altitude and lateral position; the aircraft's velocity vector is obtained by the inertial measurement unit and magnetometer, which includes the aircraft's current velocity and velocity direction.

[0031] In a first aspect, the present invention provides a guidance method based on time angle constraints of predictive correction guidance, the flowchart of which is as follows: Figure 1 As shown. This method mainly includes the following steps:

[0032] Step S101: Before the aircraft takes flight, the desired flight time and desired landing angle are pre-set on the aircraft.

[0033] Specifically, before the aircraft is launched, the operator can predict the expected flight time and expected angle of impact when the aircraft reaches the target based on the actual situation of the aircraft and the target, so that the aircraft can accurately attack the target.

[0034] Step S102: During the flight of the aircraft, the flight time error and landing angle error of the aircraft are acquired in real time. The flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle.

[0035] In a preferred embodiment of the present invention, step S102 may include:

[0036] Step S102-1: Construct two deep neural network models. Train the two deep neural network models respectively based on the multiple flight trajectory data of the aircraft, and determine the weights and offsets in each deep neural network model.

[0037] Preferably, the two deep neural network models have the same structure.

[0038] Specifically, for a single flight trajectory, its initial velocity, initial velocity direction, initial position, and altitude are recorded. At any given moment within the flight trajectory, the aircraft's velocity *v*, the tilt angle *θ* of its velocity in inertial space, the relative horizontal position *x* between the aircraft and the target, and the relative altitude *y* between the aircraft and the target are recorded. After the aircraft completes its mission, the total flight time is subtracted from the current moment to obtain the actual remaining flight time for that trajectory, and the velocity direction angle upon reaching the target is taken as the actual angle of impact. The aircraft's velocity *v*, tilt angle *θ* of its velocity in inertial space, relative horizontal position *x* between the aircraft and the target, relative altitude *y* between the aircraft and the target, actual flight time, and actual angle of impact at the current moment are used as training data.

[0039] The above process is repeated multiple times to obtain multiple training data for different flight trajectories. These multiple training data are combined into training samples to train the deep neural network model.

[0040] More preferably, the initial values ​​of the flight trajectory corresponding to each training data point in the training samples can be different, thereby increasing the randomness of the training samples. These initial values ​​include initial velocity, initial velocity direction, initial position, and altitude.

[0041] In this invention, the deep neural network model includes an input layer, multiple hidden layers, and an output layer. The input layer contains the current velocity v of the aircraft, the tilt angle θ of the aircraft in inertial space, the relative horizontal position x between the aircraft and the target, and the relative altitude y between the aircraft and the target. The output layer contains the aircraft's flight time and landing angle (angle), respectively. Figure 2 As shown.

[0042] In a preferred embodiment of the present invention, the deep neural network model has three hidden layers, each containing 100 neurons. Research has shown that determining the above parameters can guarantee the prediction accuracy and prediction speed of the deep neural network model.

[0043] In a preferred embodiment of the present invention, the hidden layer of the deep neural network model is represented by Equation 1:

[0044]

[0045] Where R represents the input of the hidden layer, C represents the output of the hidden layer, w1 represents the weight of the hidden layer, b1 represents the offset of the hidden layer, subscript j represents the number of the hidden layer, and subscript i represents the i-th neuron in hidden layer j.

[0046] Each neuron has a nonlinear activation function f, which is preferably represented by Equation 2:

[0047] f(x1) = max(0,x1) (Equation 2)

[0048] Here, x1 represents the input of each neuron.

[0049] In this invention, the training process for the deep neural network model is the same as that for traditional deep neural networks, and those skilled in the art can perform it based on experience.

[0050] According to a preferred embodiment of the present invention, when training a deep neural network model, parameters are updated using Equation 3:

[0051] β new =β old -cg t Formula 3

[0052] Where, β old Indicates the parameter before the update; β new represents the updated parameters; c represents the initial learning rate of the deep neural network model; preferably c is 0.01 to 0.0001; g t This represents the gradient of the loss function. β = {w1, b1}, where β includes the weights w1 and offsets b1 of the hidden layers in the deep neural network model; the loss function is expressed by Equation 4:

[0053]

[0054] Where M represents the number of training samples in the deep neural network model, o represents the training sample number, and y represents the actual output value corresponding to the training sample. This represents the predicted value output by the deep neural network model based on the training samples.

[0055] Preferably, in each training cycle, a certain proportion of samples are randomly selected from multiple flight trajectory data samples as training samples to train the deep neural network model. For example, the samples are divided into two parts: 80% are used as training samples to train the deep neural network model; and 20% are used as test samples to test the performance of the deep neural network model. The test samples are used to test the accuracy of the deep neural network model after the training samples are completed. If the test samples are accurate, the deep neural network model can be output. An overall accuracy of over 90% for the test samples is considered accurate.

[0056] More preferably, 50 training samples are selected in each training cycle, and the training cycle is set to 100,000 times. This allocation process can increase the robustness of the neural network model.

[0057] Therefore, it can be seen that by inputting the current velocity v of the aircraft, the tilt angle θ of the aircraft velocity in inertial space, the relative horizontal position x between the aircraft and the target, and the relative altitude y between the aircraft and the target into the deep neural network model, the predicted flight time and predicted landing angle of the aircraft are output. The predicted flight time and predicted landing angle are compared with the actual flight time and actual landing angle. The weights w1 and offsets b1 of the two deep neural network models are adjusted accordingly. This process is repeated multiple times until the predicted flight time and predicted landing angle are within a threshold range compared to the actual flight time and actual landing angle, thus completing the training of the deep neural network model.

[0058] According to a preferred embodiment of the present invention, the threshold is set according to the application scenario, taking into full account the impact of the error between the predicted value and the actual value on the guidance and control accuracy of the aircraft. For example, the threshold can be 0.05 to 0.3%, thereby ensuring the accuracy of training.

[0059] Specifically, if the error between the predicted value and the actual value is greater than a threshold, it is determined that the value does not meet the standard and training needs to continue; if the error between the predicted value and the actual value is less than or equal to the threshold, training is considered complete.

[0060] Step S102-2: Use the two trained deep neural network models to predict the flight time and landing angle of the aircraft.

[0061] Specifically, a precise deep neural network model can be obtained according to step S102-1. When the aircraft takes flight again, the trained deep neural network model can accurately predict the flight time and landing angle of the aircraft.

[0062] Step S103: Correct the guidance command based on the flight time error and landing angle error of the aircraft.

[0063] In a preferred embodiment of the present invention, in step S103, the guidance command is represented by Formula 5;

[0064]

[0065] Among them, a m Indicates the guidance command; N is a constant; λ represents the line-of-sight angle between the aircraft and the target; v represents the aircraft speed; a t As a bias term, a t =a bt +a ba a bt Indicates the time bias term; a ba This indicates the corner offset term.

[0066] In this invention, an offset proportional guidance command is used as the guidance law for the aircraft. The proportional guidance main term ensures that the aircraft can always hit the target, while the offset term achieves time and angle coordinated constraints.

[0067] In a preferred embodiment of the present invention, step S103 may include:

[0068] Step S103-1: Train the deep reinforcement learning model based on the aircraft's flight state, bias term, and reward obtained from executing the bias term.

[0069] Preferably, the flight state s of the aircraft at time t t Including flight time status and landing angle state

[0070] Among them, flight time status Where t eThe difference between the expected flight time and the predicted flight time, i.e., the flight time error, is represented by t. go This represents the actual remaining time, which is the difference between the actual flight time and the current flight time.

[0071] Among them, the landing angle state Where a e This represents the difference between the expected landing angle and the predicted landing angle, i.e., the landing angle error.

[0072] In this invention, the bias term a t For the time bias term a bt and the angle offset term a ba .

[0073] In this invention, the reward δ obtained by performing the bias term is... t Includes: Constraint time reward δ t1 and constraint landing angle reward δ t2 The value of the time constraint reward is The value of the constraint corner reward is

[0074] Preferably, the reward obtained from executing the bias term also includes: a target hit reward δ t3 The reward for hitting the target is δ. t3 =-R 2 R represents the distance between the aircraft and the target.

[0075] The goal of reinforcement learning is to maximize rewards through trial and error. Therefore, when rewards are set in this form, the goal of reinforcement learning becomes to make these rewards close to their maximum value of 0, thereby minimizing flight time error, landing angle error, and miss distance.

[0076] In this invention, the state value function V is utilized. π (s t ) indicates the flight status s t The potential value of this invention lies in its improved bias term generation strategy. The goal is to find a strategy that maximizes the total reward for the aircraft in an unknown environment. However, the total reward includes future rewards and cannot be directly calculated; therefore, a state-value function V is used. π (s t Approximate estimate of total reward.

[0077] In a preferred embodiment of the present invention, the deep reinforcement learning model is learned through the proximal policy optimization (PPO) method, and therefore the deep reinforcement learning model includes: behavior units and evaluation units.

[0078] Specifically, the behavioral unit refers to the flight state s of the aircraft at time t. t As input, with bias term a t As output and stored.

[0079] Where strategy τ represents the flight state s at time t. t Lower bias term a t The probability distribution. Due to the trial-and-error nature of deep reinforcement learning, the bias term a t It is generally not a definite value. The strategy τ takes the form of a normal distribution τ ~ N(μ,σ), and the bias term a t The probability density function is expressed by Equation 6:

[0080]

[0081] In this invention, the behavioral unit is preferably a neural network with two fully connected layers as hidden layers, containing 400 neurons. The activation function of the hidden layers is ReLU, and the activation function of the output layer is tanh. Based on the input flight state s... t First, output the intermediate variables μ and σ. Then, construct a normal distribution τ ~ N(μ,σ), randomly sample from it (i.e., randomly select values ​​from the normal distribution), and output the sampling result as the bias term a. t .

[0082] In this invention, the fully connected layer of the behavioral unit is represented by Equation 7:

[0083]

[0084] Where l1 represents the output of the fully connected layer of the behavioral unit, u1 represents the input of the fully connected layer of the behavioral unit, w2 represents the weight of the fully connected layer of the behavioral unit, b2 represents the offset of the fully connected layer of the behavioral unit, n represents the number of the connected layer, m represents the m-th neuron in connected layer n, and x2 represents the input of the ReLU activation function.

[0085] Preferably, the objective function of the behavioral unit is expressed by Equation 8:

[0086]

[0087] Where, N s The capacity of the experience pool for the deep reinforcement learning model is ω = {w2, b2}, where ω includes the weights w2 and offsets b2 of the fully connected layers of the behavioral units.

[0088] Where, r t (ω) represents the improvement strategy τ ω (a t |s t ) and the old strategy τ old (a t |s t The ratio between ) is

[0089] Among them, A τ (a t |s t () represents the dominance function, obtained through Equation 9:

[0090] A τ (a t |s t )=-V τ (s t )+δ t +γδ t+1 +γ 2 δ t+2 +…+γ k V τ (s t+k Formula Nine

[0091] Where k represents the number of rewards; V τ (s t ) indicates the flight status s t The corresponding state value function; V τ (s t+k ) represents the flight state s after obtaining k rewards. t+k The corresponding state-value function; δ t γ represents the reward at time t; γ is the discount factor, preferably set to 0.99.

[0092] Where clip represents the clipping function, obtained through equation ten:

[0093]

[0094] Where ∈ is the shearing parameter for the update magnitude of the behavior unit.

[0095] That is, the current strategy τ is based on the flight state s t Get a t and corresponding reward δ t Interact with the aircraft to obtain the state s of the next reward t+1 Then, based on the state of the next reward s t+1 The bias term a that will receive the next reward t+1 and corresponding reward δ t+1 Repeat this sampling process to obtain reward sequence data.

[0096] Among them, the element group (s) t ,a t ,δ t The data is stored in the experience pool to improve the generation strategy of bias terms. At the same time, the aircraft updates its current state to the successor state.

[0097] This shows that when A τ (at |s t When ) is positive, the flight state s at time t needs to be increased. t The probability when A τ (a t |s t When ) is negative, the flight state s at time t needs to be reduced. t The probability of.

[0098] In a further preferred embodiment, the gradient of the objective function of the behavioral unit is obtained by using the backpropagation algorithm, and the objective function of the behavioral unit is optimized to update the parameters of the behavioral unit.

[0099] The optimization of the objective function of the behavioral unit is to maximize it, which can be done using methods commonly used in existing technologies, such as stochastic gradient descent.

[0100] Specifically, the parameters in the behavioral unit are updated using Equation 11:

[0101]

[0102] Where, ω old Indicates the parameter before the update; ω new Indicates the updated parameter; α ω Indicates the parameter update rate in the behavioral unit; This represents the gradient of the objective function of the behavioral unit.

[0103] α ω The preferred value is 0.0001 to 0.00001, which allows for rapid parameter updates.

[0104] In summary, the behavioral unit and the aircraft used the old strategy τ old (a t |s t Interaction N s Next, the interaction time series generated during the interaction process is stored in a buffer. When updating the behavioral unit, the advantage function is first used. Then, based on the probability density function of the normal distribution, the probability τ of the previously executed behavior in the experience pool within the old policy is calculated. old (a t |s t Behavioral unit generation improvement strategy τ ω τ is calculated later. ω (a t |s t Then, the objective function is calculated, and the gradient of the objective function with respect to ω is obtained using the gradient descent method. The behavioral units are then updated to maximize the objective function.

[0105] Finally, the evaluation unit will evaluate the flight state s of the aircraft at time t.t As input, with the advantage function A τ (a t |s t As output, through the advantage function A τ (a t |s t Improved behavioral unit output bias term a t The strategy.

[0106] In this invention, the evaluation unit is preferably a neural network with two fully connected layers as hidden layers, containing 400 neurons. The activation function of the hidden layers is ReLU, and the activation function of the output layer is tanh. The evaluation unit estimates the advantage function A. τ (a t |s t The two state-value functions V in ) τ (s t ) and V τ (s t+k ).

[0107] In this invention, the fully connected layer of the evaluation unit is represented by Equation XII:

[0108]

[0109] Where l2 represents the output of the fully connected layer of the evaluation unit, u2 represents the input of the fully connected layer of the evaluation unit, w3 represents the weight of the fully connected layer of the evaluation unit, b3 represents the offset of the fully connected layer of the evaluation unit, q represents the number of the connected layer in the evaluation unit, p represents the p-th neuron in the connected layer q, and x3 represents the input of the ReLU activation function.

[0110] Preferably, the objective function of the evaluation unit is expressed by Equation Thirteen:

[0111]

[0112] Where ξ = {w3, b3}, ξ includes the weight w3 and offset b3 of the fully connected layer of the evaluation unit.

[0113] In a further preferred embodiment, the gradient of the objective function of the evaluation unit is obtained by using the backpropagation algorithm, and the objective function of the evaluation unit is optimized to update the parameters of the evaluation unit.

[0114] The optimization of the objective function of the evaluation unit is to minimize it, which can be done using methods commonly used in existing technologies, such as stochastic gradient descent.

[0115] Specifically, the parameters in the evaluation unit are updated using Equation 14:

[0116]

[0117] Where, ξ old Indicates the parameter before the update; ξ new Indicates the updated parameter; α ξ Indicates the parameter update rate in the evaluation unit; This represents the gradient of the objective function of the evaluation unit.

[0118] In summary, when updating the evaluation unit, the advantage function A obtained when updating the behavior unit should be used. τ (a t |s t The objective function of the evaluation unit can be obtained. Gradient descent is used to optimize the objective function of the evaluation unit, updating the parameters ξ of the evaluation unit to minimize the objective function.

[0119] After the action unit and evaluation unit are updated, the post-trained improvement strategy τ is used. ω (a t |s t Interaction with the aircraft N s Repeat this process until the simulation experiment is completed.

[0120] Step S103-2: When the rate of change of the reward obtained by executing the bias term is less than the preset threshold, the training of the deep reinforcement learning model is completed.

[0121] Preferably, when the reward δ obtained from executing the bias term is... t When the rate of change of the average value is less than a preset threshold, it is determined to be converged, the training of multiple aircraft is terminated, and the trained deep reinforcement learning model is saved.

[0122] Research has shown that when the preset threshold is set to 1-8%, preferably 2-5%, the trained deep reinforcement learning model can keep the flight time error and landing angle error of the aircraft close to zero.

[0123] According to the present invention, the trained deep neural network model and deep reinforcement learning model are loaded into the aircraft's onboard computer. Since the deep neural network model and deep reinforcement learning model have simple structures, the optimal time angle constraint guidance command can be quickly output after the aircraft takes off and is tested.

[0124] Secondly, the present invention provides a guidance system based on time angle constraints of predictive correction guidance, such as... Figure 3 As shown, it includes:

[0125] The preset module 301 is used to pre-set the desired flight time and desired landing angle of the aircraft before the aircraft takes flight;

[0126] Error acquisition module 302 is used to acquire the flight time error and landing angle error of the aircraft in real time during the flight of the aircraft. The flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle.

[0127] The correction module 303 is used to correct guidance commands based on the flight time error and landing angle error of the aircraft.

[0128] The guidance system based on predictive correction guidance with time angle constraints provided by the present invention can be used to execute the guidance method based on predictive correction guidance with time angle constraints described in the first aspect above. Its implementation principle and technical effect are similar, and will not be repeated here.

[0129] Preferably, in the guidance system based on predictive correction guidance with time angle constraints of the present invention, each module can be directly in hardware, in a software module executed by a processor, or in a combination of both.

[0130] Software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in this art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium.

[0131] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, but alternatively, it can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors incorporating a DSP core, or any other such configuration. Alternatively, the storage medium can be integrated with the processor. The processor and storage medium can reside in an ASIC. The ASIC can reside in the user terminal. Alternatively, the processor and storage medium can reside as discrete components in the user terminal.

[0132] Example

[0133] Example 1

[0134] Before the aircraft takes off, the expected flight time is set to 100 seconds and the expected landing angle to be 60 degrees.

[0135] During the flight of the aircraft, the flight time error and landing angle error of the aircraft are acquired in real time. The flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle.

[0136] Specifically, two deep neural network models were trained based on 1,000 flight trajectory data of the aircraft to determine the weights and offsets in each deep neural network model.

[0137] When training both deep neural network models, the parameters are updated using Equation 3:

[0138] β new =β old -cg t Formula 3

[0139] Where, β old Indicates the parameter before the update; β new Indicates the updated parameters; c = 0.001; g t This represents the gradient of the loss function. β = {w1, b1}, where β includes the weights w1 and offsets b1 of the hidden layers in the deep neural network model; the loss function is expressed by Equation 4:

[0140]

[0141] Where M represents the number of training samples in the deep neural network model, o represents the training sample number, and y represents the actual output value corresponding to the training sample. This represents the predicted value output by the deep neural network model based on the training samples.

[0142] Two trained deep neural network models were used to predict the flight time and angle of impact of the aircraft.

[0143] The guidance commands are corrected based on the aircraft's flight time error and landing angle error.

[0144] Specifically, the guidance command is represented by Equation 5;

[0145]

[0146] Among them, a mIndicates guidance command; N = 3; λ represents the line-of-sight angle between the aircraft and the target; v represents the aircraft speed; a t As a bias term, a t =a bt +a ba a bt Indicates the time bias term; a ba This indicates the corner offset term.

[0147] The deep reinforcement learning model is trained based on the aircraft's flight state, bias terms, and the rewards obtained from executing the bias terms.

[0148] Flight state s of the aircraft at time t t Including flight time status and landing angle state

[0149] Among them, flight time status Where t e The difference between the expected flight time and the predicted flight time, i.e., the flight time error, is represented by t. go Indicates the actual remaining time;

[0150] Among them, the landing angle state Where a e This represents the difference between the expected landing angle and the predicted landing angle, i.e., the landing angle error.

[0151] Bias term a t For the time bias term a bt and the angle offset term a ba .

[0152] The reward δ obtained from executing the bias term t Includes: Constraint time reward δ t1 Constraint landing angle reward δ t2 And the reward for hitting the target δ t3 The value of the time constraint reward is The value of the constraint corner reward is The reward for hitting the target is δ. t3 =-R 2 R represents the distance between the aircraft and the target.

[0153] Deep reinforcement learning models are trained using the proximal policy optimization (PPO) method and consist of action units and evaluation units.

[0154] The behavioral unit is based on the flight state s of the aircraft at time t. t As input, with bias term a t As output and stored.

[0155] The objective function of the behavioral unit is expressed by Equation 8:

[0156]

[0157] Where, N s The capacity of the experience pool for deep reinforcement learning models.

[0158] Where, r t (ω) represents the improvement strategy τ ω (a t |s t ) and the old strategy τ old (a t |s t The ratio between ) is

[0159] Among them, A τ (a t |s t () represents the dominance function, obtained through Equation 9:

[0160] A τ (a t |s t )=-V τ (s t )+δ t +γδ t+1 +γ 2 δ t+2 +…+γ k V τ (s t+k Formula Nine

[0161] Where k represents the number of rewards; V τ (s t ) indicates the flight status s t The corresponding state value function; V τ (s t+k ) represents the flight state s after obtaining k rewards. t+k The corresponding state-value function; δ t Let represent the reward at time t; γ is 0.99.

[0162] Where clip represents the clipping function, obtained through equation ten:

[0163]

[0164] Where ∈ is the shearing parameter for the update magnitude of the behavior unit.

[0165] The fully connected layer of the behavioral unit is represented by Equation 7:

[0166]

[0167] Where l1 represents the output of the fully connected layer of the behavioral unit, u1 represents the input of the fully connected layer of the behavioral unit, w2 represents the weight of the fully connected layer of the behavioral unit, b2 represents the offset of the fully connected layer of the behavioral unit, n represents the number of the connected layer, m represents the m-th neuron in connected layer n, and x2 represents the input of the ReLU activation function.

[0168] The parameters are updated in the behavioral unit using Equation 11:

[0169]

[0170] Where, ω old Indicates the parameter before the update; ω new Indicates the updated parameter; α ω =0.0001; Let ω represent the gradient of the objective function of the behavioral unit, where ω = {w2, b2}.

[0171] Finally, the evaluation unit uses the above data as input, with the advantage function A... τ (a t |s t As output, through the advantage function A τ (a t |s t Improved behavioral unit output bias term a t The strategy.

[0172] The objective function of the evaluation unit is expressed by Equation Thirteen:

[0173]

[0174] The fully connected layer of the evaluation unit is represented by Equation Twelve:

[0175]

[0176] Where l2 represents the output of the fully connected layer of the evaluation unit, u2 represents the input of the fully connected layer of the evaluation unit, w3 represents the weight of the fully connected layer of the evaluation unit, b3 represents the offset of the fully connected layer of the evaluation unit, q represents the number of the connected layer in the evaluation unit, p represents the p-th neuron in the connected layer q, and x3 represents the input of the ReLU activation function.

[0177] The parameters in the evaluation unit are updated using Equation 14:

[0178]

[0179] Where, ξ old Indicates the parameter before the update; ξ new Indicates the updated parameter; α ξ=0.0001; Let ξ represent the gradient of the objective function of the evaluation unit, where ξ = {w3, b3}.

[0180] The reward δ obtained from executing the bias term t When the rate of change of the average value is less than 2%, it is considered converged, the training of the aircraft is terminated, and the trained deep reinforcement learning model is saved.

[0181] Figure 4 and 5 The images show simulation results of the constrained time reward and constrained landing angle reward obtained after training the deep reinforcement learning model 100 times.

[0182] from Figure 4 and 5 As can be seen, the time-bound reward stabilized after fluctuating for 20 cycles; the corner-bound reward began to rise monotonically after fluctuating for 100 cycles, and then stabilized after the 200th cycle.

[0183] Figure 6 The aircraft was launched under random initial conditions, and different expected flight times and expected landing angles were set. The simulation results of 1000 simulations were obtained using a trained deep neural network model and a deep reinforcement learning model.

[0184] from Figure 6 As can be seen, the landing angle error and flight time error are basically close to zero, which verifies the effectiveness of the method of the present invention.

[0185] The present invention has been described in detail above with reference to specific embodiments and exemplary examples. However, these descriptions should not be construed as limiting the present invention. Those skilled in the art will understand that various equivalent substitutions, modifications, or improvements can be made to the technical solutions and embodiments of the present invention without departing from the spirit and scope of the present invention, and all such modifications and improvements fall within the scope of the present invention.

Claims

1. A guidance method based on time angle constraints in predictive correction guidance, characterized in that, Includes the following steps: Step S101: Before the aircraft takes flight, the desired flight time and desired landing angle are pre-set on the aircraft; Step S102: During the flight of the aircraft, the flight time error and landing angle error of the aircraft are acquired in real time, wherein the flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle. Step S103: Correct the guidance command based on the flight time error and landing angle error of the aircraft; The process of step S102 includes: Step S102-1: Construct two deep neural network models. Train the two deep neural network models respectively based on the multiple flight trajectory data of the aircraft, and determine the weights and offsets in each deep neural network model. Step S102-2: Predict the flight time and landing angle of the aircraft using the two trained deep neural network models respectively; When training two deep neural network models, the parameters are updated using Equation 3: β new = β old -cg t Formula III Where, β old Indicates the parameter before the update; β new Indicates the updated parameters; c represents the initial learning rate of the deep neural network model; g t This represents the gradient of the loss function. β = {w1, b1}, where β includes the weights w1 and biases b1 corresponding to the deep neural network model; the loss function is expressed by Equation 4: Where M represents the number of training samples in the deep neural network model, o represents the training sample number, and y represents the actual output value corresponding to the training sample. This represents the predicted value output by the deep neural network model based on the training samples.

2. The guidance method based on time angle constraints for predictive correction guidance according to claim 1, characterized in that, In step S103, the guidance command is represented by Equation 5: Among them, a m Indicates the guidance command; N is a constant; λ represents the line-of-sight angle between the aircraft and the target; v represents the aircraft speed; a t As a bias term, a t =a bt +a ba a bt Indicates the time bias term; a ba This indicates the corner offset term.

3. The guidance method based on time angle constraints for predictive correction guidance according to claim 2, characterized in that, The process of step S103 includes: Step S103-1: Train the deep reinforcement learning model based on the flight state of the aircraft, the bias term, and the reward obtained by executing the bias term; Step S103-2: When the rate of change of the reward obtained by executing the bias term is less than a preset threshold, the training of the deep reinforcement learning model is completed.

4. The guidance method based on time angle constraints for predictive correction guidance according to claim 3, characterized in that, The rewards obtained from executing the bias term include: constraint time reward and constraint landing angle reward, wherein the value of the constraint time reward is... t e This represents the flight time error; the constraint angle of landing reward value is... a e This indicates the angle error.

5. The guidance method based on time angle constraints for predictive correction guidance according to claim 3 or 4, characterized in that, The reward obtained by performing the bias term also includes: a target hit reward, wherein the value of the target hit reward is -R. 2 R represents the distance between the aircraft and the target.

6. The guidance method based on time angle constraints of predictive correction guidance according to claim 3, characterized in that, The advantage function in the deep reinforcement learning model is obtained through Equation 9: A τ (a t |s t )=-V τ (s t )+δ t +γδ t+1 +γ 2 δ t+2 +…+γ k V τ (s t+k ) Equation (9) Among them, A τ (a t |s t ) represents the advantage function; k represents the number of rewards; V τ (s t ) indicates the flight status s t The corresponding state value function; V τ (s t+k ) represents the flight state s after obtaining k rewards. t+k The corresponding state-value function; δ t Let t represent the reward at time t; γ is the discount factor.

7. A guidance system implementing the guidance method based on predictive correction guidance with time angle constraints as described in any one of claims 1 to 6, characterized in that, include: The preset module is used to pre-set the desired flight time and desired landing angle of the aircraft before it takes flight; The error acquisition module is used to acquire the flight time error and landing angle error of the aircraft in real time during the flight process, wherein the flight time error is the difference between the expected flight time and the predicted flight time, and the landing angle error is the difference between the expected landing angle and the predicted landing angle. The correction module is used to correct guidance commands based on the aircraft's flight time error and landing angle error.

Citation Information

Patent Citations

  • Reentry prediction-correction guidance method for predicting voyage based on BP neural network

    CN111813146A