An aircraft intelligent guidance method based on reinforcement learning
By using a reinforcement learning-based intelligent guidance method for aircraft, the tilt angle command is generated online using real-time status information. This solves the problem of traditional guidance methods relying on pre-launch design of the flight corridor, and enables the aircraft to make autonomous decisions and improve its adaptability in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional aircraft guidance methods rely on pre-launch design of flight corridors, which cannot effectively adapt to complex environments and lack autonomy and real-time decision-making capabilities.
A reinforcement learning-based intelligent guidance method for aircraft is adopted. By optimizing the reinforcement learning model through near-end strategy, the roll angle command is generated online using real-time state information. Combined with heat flow rate, overload and dynamic pressure constraints, intelligent guidance of the aircraft is realized.
It enables autonomous decision-making for aircraft in complex environments, breaks through the limitations of traditional methods, forms a new type of flight profile, and improves the autonomy and adaptability of guidance methods.
Smart Images

Figure CN119596971B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aircraft guidance and control, and relates to an intelligent guidance method for aircraft based on reinforcement learning. It is used to solve the problem that traditional guidance methods rely on pre-launch design of flight corridors, and innovatively forms a new type of flight profile. It enables aircraft to make online decisions on tilt angle commands based on real-time status information, and supports the leapfrog development of my country's aircraft towards intelligence. Background Technology
[0002] Aircraft, with their high speed, wide airspace, long range, and strong maneuverability, play a vital role in both military and civilian fields. During reentry, the external environment, such as atmospheric density and pressure, undergoes drastic changes, resulting in high complexity and uncertainty in the aircraft's dynamics and aerodynamic characteristics. Constraints related to process and control variables must be met during flight, which places higher demands on reentry guidance technology.
[0003] Currently, reentry guidance methods for aircraft mainly include two types: nominal trajectory-based guidance and predictive correction guidance. Nominal trajectory-based guidance involves offline design of drag-acceleration-energy and altitude-velocity profiles, and real-time tracking of the nominal trajectory during flight for guidance. This method is simple in principle and highly feasible in engineering, but it has significant shortcomings in adaptability and robustness to complex environments. Predictive correction guidance integrates the dynamic equations based on the aircraft's current state to calculate and predict the aircraft's terminal state. It then corrects the guidance command based on the deviation between the terminal state and the target point, thereby achieving precise guidance of the aircraft. Compared to nominal trajectory-based tracking guidance, predictive correction guidance offers greater autonomy and adaptability to complex environments.
[0004] However, traditional reentry guidance and predictive correction guidance based on nominal trajectories mostly require designing flight corridor parameters based on human experience. Guidance methods based on nominal trajectories need to transform constraints such as thermal flux and dynamic pressure during the reentry process into the upper and lower boundaries of the flight corridor, and then carefully design a nominal flight trajectory that meets the terminal constraints within the boundary range. Predictive correction guidance methods need to determine the sign of the aircraft's roll angle by designing a suitable heading angle error corridor. When the aircraft's heading angle error reaches the corridor boundary, the sign of the roll angle is flipped. Summary of the Invention
[0005] The technical problem solved by this invention is to overcome the shortcomings of the prior art and provide an intelligent guidance method for aircraft based on reinforcement learning, which solves the problem that traditional guidance methods rely on pre-launch design of flight corridors and enables aircraft to make online decisions on roll angle commands based on real-time status information.
[0006] The technical solution of this invention is: an intelligent guidance method for aircraft based on reinforcement learning, which includes the following steps:
[0007] S1. Define and obtain the observation space of the aircraft, which includes the aircraft state quantity, the aircraft state differential component, the difference between the aircraft state and the state at the mid-terminal handover point, the Euclidean distance between the aircraft latitude and longitude and the mid-terminal handover point latitude and longitude, and the difference between the aircraft energy and the terminal energy.
[0008] S2. Define the action space of the reinforcement learning policy model as the aircraft's roll angle reference value, use a near-end policy to optimize the reinforcement learning model, and generate the aircraft's roll angle reference value online based on the real-time state space.
[0009] S3. Map the aircraft's tilt angle reference value output by the reinforcement learning policy model to obtain the aircraft's theoretical tilt angle;
[0010] S4. Truncate the theoretical roll angle of the aircraft to obtain the final roll angle command actually used to control the aircraft.
[0011] S5, adjust the yaw angle command. Substitute the dynamic model into the aircraft state update and repeat steps S1 to S5.
[0012] Preferably, the calculation method for truncating the theoretical tilt angle is as follows:
[0013]
[0014] Where, σ max The η is the constraint on the tilt angle amplitude, and η is the safety factor.
[0015] Preferably, the constraint σ of the tilt angle amplitude max for:
[0016]
[0017]
[0018] in, The threshold for the tilt angle amplitude under heat flow rate constraints; The tilt angle amplitude threshold under overload constraints; The tilt angle amplitude threshold under dynamic pressure constraint; m is the vehicle mass; g is the gravitational acceleration; V is the vehicle velocity; r is the distance from the Earth's center; S ref k represents the characteristic area of the aircraft. Q These are the coefficients of the heat flow rate model; C is the maximum allowable heat flux at the stagnation point. L and C D These are the lift coefficient and drag coefficient, respectively; n maxThe maximum permissible overload; g0 is the gravitational acceleration at sea level; q max This represents the maximum dynamic pressure.
[0019] Preferably, the near-end policy optimization reinforcement learning model includes an Actor model and a Critic model. The Actor model outputs the aircraft's actions based on the aircraft's observation space, and the Critic model calculates the reward based on the aircraft's observation space, and optimizes the Actor model with the objective function of maximizing the expected reward.
[0020] Preferably, the objective function is calculated using the following formula:
[0021]
[0022] In the formula, R(s) t ,a t Let be the reward function, representing the state of the aircraft in state s. t Perform action a t The expected reward of the immediate reward received, including the lateral reward R. h (s t ,a t Vertical Rewards R v (s t ,a t ) and final reward R f (s t ,a t ), κ t This is the discount factor for the current calculation period, with a value ranging from 0 to 1.
[0023] Preferably, the lateral reward R h (s t ,a t () is the sum of latitude and longitude bonus and direction bonus.
[0024] Preferably, the directional reward r dir By calculating the aircraft's actual velocity direction ψ and line-of-sight angle ψ →tgt The difference between the two points guides the aircraft to fly along the line-of-sight direction towards the mid-to-late shift handover point. The calculation formula is as follows:
[0025]
[0026] In the formula, L togo This refers to the remaining range of the aircraft. This refers to the hyperparameter used to regulate weight changes in directional rewards; This is used to adjust the weight of the directional reward's impact on the aircraft; Represents the heading angle ψ t With line of sight difference;
[0027]
[0028] in, Let θ be the expected longitude of the mid-to-latest shift change point, φ be the actual longitude of the mid-to-latest shift change point, and φ be the actual longitude of the mid-to-latest shift change point. t represents the expected latitude of the mid-to-late handover point, t represents the current calculation cycle, and t-1 represents the previous calculation cycle.
[0029] The latitude and longitude bonus serves to guide the aircraft closer to the mid-to-late shift handover point, and the calculation formula is as follows:
[0030]
[0031] Among them, L togo For the remaining range of the aircraft, L * The set distance threshold; This refers to the hyperparameters used to regulate weight changes in latitude and longitude rewards; θ is the weight used to adjust the impact of latitude and longitude rewards on the aircraft; φ is the actual longitude value of the mid-to-latest handover point; These represent the Euclidean distances between the latitude and longitude of the aircraft in the current calculation cycle and the latitude and longitude of the mid-to-latest handover point in the previous calculation cycle.
[0032] Preferably, the vertical reward R v (s t ,a t () is the sum of energy reward, speed reward, and altitude reward.
[0033] Preferably, the energy reward r E The calculation formula is as follows:
[0034]
[0035]
[0036] In the formula, The mean is E μ The standard deviation is δ E Take from Gaussian distribution The probability, b V b is the bias amount used to adjust the speed error. h Here, V is the offset used to adjust altitude error, R0 is the Earth's radius, and GM represents the gravitational constant.
[0037] The speed reward r V The calculation formula is as follows:
[0038]
[0039] In the formula, This indicates that the mean is V μ The standard deviation is δ V V+b is obtained from a Gaussian distribution V The probability, b V This is the offset amount used to adjust the speed error;
[0040] The height reward r h The calculation formula is as follows:
[0041]
[0042] In the formula, This indicates that the mean is h μ The standard deviation is δ h h+b is obtained from a Gaussian distribution h The probability, b h This is the offset amount used to adjust the height error; This is the hyperparameter used to adjust the weight changes in this reward, where h is the flight altitude of the aircraft.
[0043] Preferably, the final reward r tgt The calculation formula is as follows when the aircraft arrives at the final handover point:
[0044]
[0045] In the formula, ω r ,ω θ ,ω φ ,ω V These are the influence coefficients ε, used to balance the effects of distance, longitude, latitude, and speed amplitude on the aircraft. r ,ε θ ,ε φ ,ε V These are hyperparameters used to calculate the relative errors of the aircraft's distance, longitude, latitude, and speed. The bonus at the end of each subsequent time period is 0.
[0046] The advantages of this invention compared to the prior art are:
[0047] (1) The intelligent guidance method for aircraft based on reinforcement learning proposed in this invention guides the aircraft to explore and experiment in the environment by setting effective rewards, and uses the PPO algorithm in reinforcement learning to train the aircraft tilt angle guidance model, giving full play to the wide-range flight advantage of the aircraft, breaking through the limitations of traditional methods, and creating a new quality flight profile to support the leapfrog development of my country's aircraft towards intelligence.
[0048] (2) The intelligent guidance method for aircraft based on reinforcement learning proposed in this invention can enable aircraft to make online decisions on tilt angle commands based on real-time state information, effectively improving the autonomy of the aircraft guidance method and its adaptability to complex environments. Attached Figure Description
[0049] Figure 1 The comparison results between the PPO strategy and the predictive correction guidance method (altitude versus velocity curves);
[0050] Figure 2 The comparison results between the PPO strategy and the prediction correction guidance method (latitude and longitude variation curves);
[0051] Figure 3 The comparison results between the PPO strategy and the predictive correction guidance method (tilt angle versus velocity curve);
[0052] Figure 4 The comparison results between the PPO strategy and the predictive correction guidance method (curve of heading angle versus time). Detailed Implementation
[0053] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0054] At present, preliminary research has been carried out on artificial intelligence-based aircraft guidance and control technology. It has the characteristics of low computational load and high precision, and can realize the rapid mapping of the real-time status of the aircraft to guidance commands, which has broad application prospects.
[0055] This invention designs an intelligent guidance method for aircraft based on reinforcement learning, which solves the problem of traditional guidance methods relying on pre-launch design of flight corridors. Under the constraints of dynamic equations, heat flow rate, overload, dynamic pressure, etc., the method guides the aircraft to explore the environment by setting effective rewards, optimizes the reinforcement learning model using near-end strategy, and generates roll angle commands online based on real-time state information. This method subverts the constraints of traditional guidance methods in terms of lateral corridors, explores completely different flight trajectories, and realizes the aircraft's online decision-making on roll angle commands based on real-time state information.
[0056] The invention will now be described from two aspects: aircraft modeling and a reinforcement learning-based intelligent guidance method for aircraft.
[0057] (1) Aircraft Modeling
[0058] (a) Aircraft dynamics modeling:
[0059]
[0060] In the formula, r is the distance from the Earth's center; θ and φ represent the longitude and latitude of the aircraft, respectively; V is the aircraft velocity; γ and ψ are the aircraft's track angle and heading angle, respectively; σ is the aircraft's roll angle; m is the aircraft's mass; g is the acceleration due to gravity; L and D are the aircraft's lift and drag; t is time.
[0061]
[0062] In the formula, ρ=ρ0e -βh ρ is the atmospheric density at the altitude h where the aircraft is located, and ρ0 is the atmospheric density at sea level; β = 1 / H MCP H MCP =7.11km is the reference altitude; S ref C represents the characteristic area of the aircraft. L and C D These are the lift coefficient and drag coefficient, respectively, which are related to the aircraft's angle of attack α and flight speed V. In a specific embodiment of this invention, taking the American general aviation aircraft CAV-H as the research object, the expressions for the lift coefficient and drag coefficient are:
[0063]
[0064] In the formula, C L0 C L1 C L2 and C L3 C represents the fitting coefficients for the lift expression; D0 C D1 C D2 and C D3 The coefficients are the fitting coefficients for the resistance expression.
[0065] (b) Process constraint modeling: During high-speed flight, the high-speed airflow generates a large amount of heat energy in the stagnation area of the glider. From the perspective of aircraft safety, the heat flux constraint should be satisfied:
[0066]
[0067] In the formula, Indicates the heat flux at the stagnation point; k Q These are the coefficients of the heat flow rate model; This represents the maximum permissible heat flux at the stagnation point.
[0068] Overload refers to the ratio of the resultant force of aerodynamic forces and thrust acting on an aircraft to its weight. Due to the limited structural strength of aircraft, and considering safety issues, the overload of an aircraft cannot exceed the maximum overload limit. Since there is no thrust during the gliding phase of the aircraft in this invention, its overload constraint is...
[0069]
[0070] In the formula, n represents the aircraft overload; n max g is the maximum permissible overload; g0 is the gravitational acceleration at sea level.
[0071] Similarly, considering the structural strength of the aircraft, it must also meet dynamic pressure constraints. During high-speed flight, air flows at high speed in the opposite direction to the fuselage. At this time, the air itself has kinetic energy, and the pressure presented by this kinetic energy is called dynamic pressure.
[0072]
[0073] In the formula, q represents the dynamic pressure of the aircraft; max This represents the maximum dynamic pressure.
[0074] Heat flow rate, overload, and dynamic pressure constraints all relate to aircraft safety and are considered hard constraints. Quasi-equilibrium gliding conditions, on the other hand, are soft constraints designed to ensure stable altitude and flight path angles, preventing large changes. The mathematical form of this constraint is as follows:
[0075]
[0076] Substituting these equations into the dynamic model's expression for the track angle, we obtain the equations of motion describing quasi-equilibrium gliding:
[0077]
[0078] As can be seen from the above equation, when the aircraft satisfies the above equation, the resultant force of its gravity and aerodynamic lift provides the required centripetal force. At this time, the change in flight altitude is very small, and the rate of change of the flight path angle γ is very small. Approaching zero.
[0079] (c) Terminal constraint modeling: The aircraft's terminal constraints include altitude, speed, and latitude / longitude constraints at the mid-to-final handover point.
[0080]
[0081] In the formula, h f V f θ f φ f These are the actual values of altitude, speed, longitude, and latitude of the aircraft when it reached the mid-to-late handover point; It represents the expected values of altitude, speed, longitude, and latitude at the mid-to-latest shift handover point; ε h ε V ε θ ε φ These represent the maximum allowable errors in terminal altitude, speed, longitude, and latitude, respectively.
[0082] (d) Modeling of roll angle magnitude constraints: During flight, the aircraft needs to satisfy the aforementioned hard constraints such as heat flux, overload, and dynamic pressure. However, repeatedly calculating the heat flux, overload, and dynamic pressure of the aircraft would result in a large computational load. Therefore, the hard constraints of heat flux, overload, and dynamic pressure can be transformed into constraints on the roll angle magnitude, i.e.:
[0083]
[0084] The constraint form for the tilt angle magnitude is:
[0085]
[0086] in, The threshold for the tilt angle amplitude under heat flow rate constraints; The tilt angle amplitude threshold under overload constraints; The tilt angle amplitude threshold under dynamic pressure constraint; m is the vehicle mass; g is the gravitational acceleration; V is the vehicle velocity; r is the distance from the Earth's center; S ref k represents the characteristic area of the aircraft. Q These are the coefficients of the heat flow rate model; C is the maximum allowable heat flux at the stagnation point. L and C D These are the lift coefficient and drag coefficient, respectively; n max The maximum permissible overload; g0 is the gravitational acceleration at sea level; q max This represents the maximum dynamic pressure.
[0087] e) Control variable change rate constraint modeling: Due to the action of the internal control mechanism of the aircraft, the change of the control variable requires a certain change time and rate, and cannot change to the specified value instantaneously. Since the angle of attack α adopts the standard angle of attack profile, the constraint on the control variable change rate mainly targets the change rate of the roll angle σ, that is:
[0088]
[0089] In the formula, This represents the upper bound of the rate of change of the tilt angle.
[0090] (2) Intelligent guidance method for aircraft based on reinforcement learning
[0091] Reinforcement learning algorithms consist of five basic elements: agent, environment, state, action, and reward. The agent takes actions, obtains reward values through interaction with the environment, and enters the next state. This process is repeated continuously to learn the policy, and eventually the agent tends to make actions with higher cumulative reward values. The observation space and action space are defined below, and reward functions are designed from multiple perspectives. Finally, the corresponding policy model structure is designed.
[0092] (a) Observation space: The observation space of the strategy model in this invention is defined as follows:
[0093] o(t)=[r,θ,φ,V,γ,ψ,dr,dθ,dφ,dV,Δr,Δθ,Δφ,ΔV,l2(θ,φ),ΔE]∈R 16
[0094] It mainly consists of the following parts, the aircraft state variables: [r,θ,φ,V,γ,ψ]∈R 6 Let R represent the distance, longitude, latitude, speed, track angle, and heading angle of the aircraft, respectively. The differential components of the aircraft state are: [dr,dθ,dφ,dV]∈R 4 Let be the differential components of the aircraft's distance, longitude, latitude, and speed, respectively, and let [Δr,Δθ,Δφ,ΔV]∈R be the difference between the aircraft's state and the state at the mid-to-late handover point. 4 These represent the differences in distance, longitude, latitude, and speed of the aircraft, respectively, and the Euclidean distance between the aircraft's latitude and longitude and the latitude and longitude of the mid-to-latest handover point: The difference between the spacecraft's energy and the terminal energy: The energy of the aircraft This represents the energy at the last handover point in the spacecraft; GM represents the gravitational constant, with a value of 3.986 × 10⁻⁶. 14 N·m 2 / kg, and These are the expected longitude and latitude values of the mid-to-late handover points, respectively.
[0095] (b) Action space: The strategy model needs to output the aircraft's roll angle reference (including the magnitude and sign of the roll angle reference) as a control variable, and needs to satisfy the constraint of the roll angle reference magnitude during the aircraft's reentry process.
[0096] (c) Intelligent guidance methods for aircraft
[0097] In its specific implementation, the tilt angle is calculated according to the following process:
[0098] S1. Define and obtain the observation space of the aircraft, which includes the aircraft state quantity, the aircraft state differential component, the difference between the aircraft state and the state at the mid-terminal handover point, the Euclidean distance between the aircraft latitude and longitude and the mid-terminal handover point latitude and longitude, and the difference between the aircraft energy and the terminal energy.
[0099] S2. Define the action space of the reinforcement learning policy model as the aircraft's roll angle reference value, use a near-end policy to optimize the reinforcement learning model, and generate the aircraft's roll angle reference value online based on the real-time state space.
[0100] The output of a near-end policy optimization reinforcement learning model follows a Gaussian distribution with mean μ and standard deviation, i.e., output δ.
[0101] During training, a real number is randomly sampled from the Gaussian distribution as the output, and the aircraft's tilt angle reference z∈R. During testing, the mean of the Gaussian distribution is directly taken as the output, and the aircraft's tilt angle reference z=μ∈R.
[0102] S3. Map the aircraft's tilt angle reference z, output by the reinforcement learning policy model, to obtain the aircraft's theoretical tilt angle;
[0103] The mapping relationship between the aircraft's tilt angle reference value z and the aircraft's theoretical tilt angle σ is as follows:
[0104] σ=σ lim ×tanhz
[0105] Where tanh represents the hyperbolic tangent function, σ lim As a defined tilt angle boundary, in a specific embodiment of the present invention, σ lim The value is 80°.
[0106] S4. Truncate the theoretical roll angle of the aircraft to obtain the final roll angle command actually used to control the aircraft.
[0107] The calculation method for truncating the theoretical tilt angle is as follows:
[0108]
[0109] Where, σ max η is the threshold for the tilt angle amplitude under the combined constraints of heat flow rate, overload, and dynamic pressure, and is a safety factor, generally recommended to be 0.8. max This is a constraint on the tilt angle amplitude.
[0110] S5, adjust the yaw angle command. Substitute the dynamic model into the aircraft state update and repeat steps S1 to S5.
[0111] (c) Proximity policy optimization reinforcement learning model
[0112] The PPO algorithm used in the proximal policy optimization reinforcement learning model of this invention belongs to the Actor-Critic algorithm. The proximal policy optimization reinforcement learning model includes an Actor model and a Critic model. The Actor model outputs the aircraft's actions based on the aircraft's observation space, and the Critic model calculates the reward based on the aircraft's observation space and optimizes the Actor model with the objective function of maximizing the expected reward.
[0113] In this invention, both the Actor and Critic models use three-layer fully connected networks, with the hidden variables in the middle layers having a length of 64. The final output of the Actor model is used to obtain the final control quantity through a tilt angle calculation process in the action space; the output of the Critic model, V(s)∈R, is the value function of the current aircraft state.
[0114] The formula for calculating the value function is as follows:
[0115]
[0116] In the formula, R(s) t ,a t Let be the reward function, representing the state of the aircraft in state s. t Perform action a t The expected reward of the immediate reward received, including the lateral reward R. h (s t ,a t Vertical Rewards R v (s t ,a t ) and final reward R f (s t ,a t ), κ t This is the discount factor for the current calculation period, with a value ranging from 0 to 1.
[0117] (d) Reward function design: Based on the task requirements, this invention designs multiple reward functions for PPO algorithm training from three aspects, namely horizontal reward, vertical reward and end-stage reward.
[0118] Lateral reward: This serves to guide the aircraft towards the mid-to-late handover point. To enhance the exploratory nature of the aircraft's interaction with the environment, an exponentially varying reward weight is designed. When the aircraft is far from the target point, the lateral reward is 0, allowing for flexible maneuvering. As the aircraft approaches the mid-to-late handover point to a certain distance, the weight coefficient increases, strengthening the guidance effect and ensuring the aircraft accurately reaches the late handover point. The lateral reward function will be designed from the perspectives of aircraft direction and latitude / longitude, i.e., the lateral reward R... h (s t ,a t () is the sum of latitude and longitude bonus and direction bonus.
[0119] Directional reward r dir It is done by calculating the aircraft's actual velocity direction angle ψ and line-of-sight angle ψ. →tgt The difference between them guides the aircraft to fly towards the mid-to-late shift handover point along the line of sight:
[0120]
[0121] In the formula, L togo This refers to the remaining range of the aircraft. This refers to the hyperparameter used to regulate weight changes in directional rewards; This is the weight used to adjust the impact of directional rewards on the aircraft; this weight increases as the remaining range decreases. Represents the heading angle ψ t With line of sight difference.
[0122]
[0123] in, Let θ be the expected longitude of the mid-to-latest shift change point, φ be the actual longitude of the mid-to-latest shift change point, and φ be the actual longitude of the mid-to-latest shift change point. t represents the expected latitude of the mid-to-late handover point, t represents the current calculation cycle, and t-1 represents the previous calculation cycle.
[0124] Latitude and longitude rewards are the same as direction rewards, with latitude and longitude rewards r dist It also serves to guide the aircraft closer to the mid-to-late shift handover point. Furthermore, when the aircraft is within range of the mid-to-late shift handover point but begins to move away, a significant penalty will be imposed through a weighting change. This negative reward will quickly offset the positive reward gained during the approach to the mid-to-late shift handover point. The calculation formula is as follows:
[0125]
[0126] In the formula, where L togo L represents the remaining range of the aircraft. * The set distance threshold is typically 2000km. This refers to the hyperparameters used to regulate weight changes in latitude and longitude rewards; The weight used to adjust the impact of latitude and longitude rewards on the aircraft increases as the remaining range decreases; θ is the actual longitude of the mid-to-latest handover point; φ is the actual longitude of the mid-to-latest handover point. These represent the Euclidean distances between the latitude and longitude of the aircraft in the current calculation cycle and the latitude and longitude of the mid-to-latest handover point in the previous calculation cycle.
[0127] The longitudinal reward serves to constrain the changes in energy (velocity and altitude) during the spacecraft's reentry. To obtain the expected value at each moment, a linear approximation method is used, assuming that energy / velocity / altitude changes approximately linearly with a certain variable from the starting point to the ending point. Since linear approximation obviously has errors, this invention adds biases to velocity and altitude and increases weight control during implementation to ensure the actual constraint effect. The longitudinal reward R... v (s t ,at () is the sum of energy reward, speed reward, and altitude reward.
[0128] Energy reward r E The design involves observing the change in the aircraft's mechanical energy with the remaining range. It was found that the mechanical energy changes approximately linearly with the remaining range; therefore, the remaining range was chosen to make a linear approximation of the energy change. The expected energy value E is calculated using the following formula. μ Based on the Gaussian distribution, the error between the actual energy and the expected energy is calculated to provide a reward, and the standard deviation of the Gaussian distribution is used as a hyperparameter to adjust the tolerance for error.
[0129]
[0130]
[0131] In the formula, E0 and L0 represent the energy and remaining range of the aircraft at the reentry point, respectively; The mean is E μ The standard deviation is δ E Take from Gaussian distribution The probability, b V b is the bias amount used to adjust the speed error. h Here, V is the offset used to adjust altitude error, V is the speed of the aircraft, and R0 is the Earth's radius.
[0132] Speed reward r V The design also observed the change in aircraft speed over time and found that the speed changes approximately linearly with time. Therefore, a linear approximation of the speed change was made based on the flight time. The expected value V of the speed was calculated using the following formula. μ Based on a Gaussian distribution, the error between the actual speed and the expected speed is calculated to provide a reward, and the standard deviation of the Gaussian distribution is used as a hyperparameter to adjust the tolerance for error.
[0133]
[0134]
[0135] In the formula, V0 represents the velocity of the aircraft at the reentry point; For the expected flight time; This indicates that the mean is V μ The standard deviation is δ V V+b is obtained from a Gaussian distribution V The probability, b V This is the offset used to adjust the speed error.
[0136] High reward r hBy observing the changes in aircraft altitude with speed, the design reveals that altitude fluctuates significantly in the early stages of flight, while exhibiting an approximately linear change in the later stages. Therefore, an exponential weight, similar to that of the lateral reward, is incorporated into the altitude reward. This weight is almost zero in the early stages of flight, increasing in weight in the later stages to constrain the aircraft's motion. The increased altitude reward prevents excessively rapid altitude descent. Since aircraft speed is significantly affected by control variables, this invention ultimately chooses a similar approach to the speed reward, using time to perform a linear approximation of altitude estimation, as shown in the following formula.
[0137]
[0138]
[0139] In the formula, h0 represents the altitude of the aircraft at the reentry point; This indicates that the mean is h μ The standard deviation is δ h h+b is obtained from a Gaussian distribution h The probability, b h This is the offset amount used to adjust the height error; The hyperparameter used to adjust the weight changes in this reward is set to 2000km, where h is the current flight altitude of the aircraft.
[0140] Final reward: Depending on the needs, the mission will be considered complete based on the aircraft's energy level. When the aircraft's energy drops to the level at the mid-to-final handover point, the mission is considered complete.
[0141] At the end of the mission, the error between the aircraft's final state and the target state will be calculated, and the aircraft will be given an appropriate reward accordingly. The smaller the error, the higher the reward, calculated using the following formula.
[0142]
[0143] In the formula, ω r ,ω θ ,ω φ ,ω V These are the influence coefficients ε, used to balance the effects of distance, longitude, latitude, and speed amplitude on the aircraft. r ,ε θ ,ε φ ,ε V These are hyperparameters used to calculate the relative errors of the aircraft's distance, longitude, latitude, and speed. In this invention, ε is taken as... r =1.5km,ε θ =0.15°, ε φ =0.15°, ε V =100m / s.
[0144] Simulation verification
[0145] To verify the effectiveness of the PPO algorithm, this invention conducted simulations under three scenarios. First, the algorithm was tested under undisturbed reentry point conditions, obtaining the change curves of each state variable during the spacecraft's reentry process. Then, each state variable at the reentry point was perturbed individually (i.e., only a single state variable was perturbed at a time), and the state error upon arrival at the final handover point was compared. Finally, the reentry point state was perturbed as a whole, and the state error upon arrival at the final handover point was also compared. In the experiment, PPO was trained for 800 episodes, with each episode consisting of 4096 parallel environments performing 400 environmental interactions. The optimal model was selected from the last 100 episodes for testing, yielding the following results.
[0146] The aircraft departs from a given reentry point, with the mission objective of flying to a pre-defined mid-to-late shift handover point. The aircraft's target state at the mid-to-late shift handover point is:
[0147] Table 1 Target Status of the Aircraft at the Mid-to-Late Shift Change Point
[0148]
[0149] During the training process, the parameters of the aircraft reentry point are randomly and uniformly sampled according to the parameter range shown in Table 2.
[0150] Table 2. Parameter range for the initial reentry point of the aircraft.
[0151]
[0152] (1) No disturbance at reentry point
[0153] The strategy trained using the PPO algorithm is compared with the predictive correction guidance method. Using the median of the parameter range in Table 2 as the reentry point state, the state change curves of the spacecraft during reentry are obtained as follows: Figures 1-4 As shown, the strategy obtained by the PPO algorithm maintains target accuracy similar to or even better than that of prediction correction guidance, while having a smoother tilt angle curve and exploring a completely different flight trajectory, exhibiting significantly different characteristics in latitude, longitude, and altitude changes compared to the classic flight trajectory.
[0154] The error between the spacecraft's state and the target's state at the end of the mission is shown in the table below. Compared with the predictive correction guidance method, the PPO algorithm achieves better speed control accuracy while maintaining similar position control (latitude, longitude, and altitude).
[0155] Table 3 shows the error between the aircraft status and the target status at the final handover point.
[0156]
[0157] (2) Reentry point single-state quantity disturbance
[0158] To test the impact of different state variable changes on the accuracy of the mid-to-final handover point, single-state variable perturbations were applied to the state variables of the aircraft at the reentry point. For the six state variables shown in Table 2, 11 points were evenly selected within their parameter ranges for testing, resulting in a total of 66 experiments. The average state error at the mid-to-final handover point is shown in the table below. It can be seen that the PPO algorithm has excellent generalization ability for single-state variable perturbations at the reentry point.
[0159] Table 4. Experimental results of single-state disturbance at the reentry point.
[0160]
[0161] (3) Total state disturbance
[0162] To test the impact of simultaneous changes in multiple state variables on the accuracy of the mid-to-final handover point, the state variables of the aircraft at the reentry point were subjected to a comprehensive perturbation. The maximum, minimum, and median values of the six state variables within the parameter range shown in Table 2 were combined to obtain 3... 6 =729 sets of tests, and the average state error at the mid-to-late shift handover point is shown in the table below.
[0163] Table 5 Results of the All-Variable Perturbation Experiment
[0164]
[0165]
[0166] This invention designs an intelligent guidance method for aircraft based on reinforcement learning. Under the constraints of dynamic equations, heat flow rate, overload, dynamic pressure, etc., the method guides the aircraft to explore the environment by setting effective rewards. The method uses a near-end strategy optimization algorithm in reinforcement learning to train the aircraft's tilt angle guidance model and generates tilt angle commands online based on real-time state information. This method subverts the constraints of traditional guidance methods in the lateral corridor and explores a completely different flight trajectory.
[0167] Simulation results show that the reinforcement learning-based intelligent guidance method for aircraft can achieve control accuracy comparable to predictive correction guidance algorithms, fully leverage the wide-range flight advantages of aircraft, further expand the flight profile, and also possess strong robustness to state disturbances, achieving good control performance under various conditions.
[0168] The contents not described in detail in this specification are common knowledge to those skilled in the art.
Claims
1. A method for intelligent guidance of an aircraft based on reinforcement learning, characterized in that Includes the following steps: S1. Define and obtain the observation space of the aircraft, which includes the aircraft state quantity, the aircraft state differential component, the difference between the aircraft state and the state at the mid-terminal handover point, the Euclidean distance between the aircraft latitude and longitude and the mid-terminal handover point latitude and longitude, and the difference between the aircraft energy and the terminal energy. S2. Define the action space of the reinforcement learning policy model as the aircraft's roll angle reference value, use a near-end policy to optimize the reinforcement learning model, and generate the aircraft's roll angle reference value online based on the real-time state space. S3. Map the aircraft's tilt angle reference value output by the reinforcement learning policy model to obtain the aircraft's theoretical tilt angle; S4. Truncating the theoretical roll angle of the aircraft, resulting in a final actual roll angle command used to control the aircraft ; S5, adjust the yaw angle command. Substitute the dynamic model into the aircraft state update and re-execute steps S1 to S5; The near-end policy optimization reinforcement learning model includes an Actor model and a Critic model. The Actor model outputs the aircraft's actions based on the aircraft's observation space, and the Critic model calculates the reward based on the aircraft's observation space and optimizes the Actor model with the objective function of maximizing the expected reward. The objective function is calculated using the following formula: wherein is a reward function, representing the reward of the aircraft in state performing action The expected reward of the immediate reward obtained includes lateral reward , longitudinal reward and terminal segment reward , is a discount factor of the current calculation period, and the value range is 0-1.
2. The intelligent guidance method for aircraft based on reinforcement learning according to claim 1, characterized in that, The calculation method for truncating the theoretical tilt angle is as follows: in, As a constraint on the tilt angle magnitude, This is for the safety factor.
3. The intelligent guidance method for aircraft based on reinforcement learning according to claim 2, characterized in that... The constraint on the tilt angle magnitude for: in, The threshold for the tilt angle amplitude under heat flow rate constraints; The tilt angle amplitude threshold under overload constraints; This is the threshold for the tilt angle amplitude under dynamic pressure constraints; For the mass of the aircraft; It is the acceleration due to gravity; For the speed of the aircraft; It is the distance from the Earth's center; The characteristic area of the aircraft; These are the coefficients of the heat flow rate model; This represents the maximum permissible heat flux at the stagnation point. and These are the lift coefficient and the drag coefficient, respectively. Maximum permissible overload; The gravitational acceleration at sea level; This represents the maximum dynamic pressure.
4. The intelligent guidance method for aircraft based on reinforcement learning according to claim 1, characterized in that... The horizontal reward It is the sum of latitude and longitude bonuses and direction bonuses.
5. The intelligent guidance method for aircraft based on reinforcement learning according to claim 4, characterized in that, The direction of reward By calculating the actual velocity direction of the aircraft and line of sight The difference between the two points guides the aircraft to fly along the line-of-sight direction towards the mid-to-late shift handover point. The calculation formula is as follows: In the formula, This refers to the remaining range of the aircraft. This refers to the hyperparameter used to regulate weight changes in directional rewards; This is used to adjust the weight of the directional reward's impact on the aircraft; Indicates heading angle With line of sight difference; in, This represents the expected longitude of the final handover point. This is the actual longitude value of the mid-to-late shift handover point; This is the actual longitude value of the mid-to-late shift handover point. This represents the expected latitude of the mid-to-late handover point. Indicates the current calculation period. Previous calculation cycle; The latitude and longitude bonus serves to guide the aircraft closer to the mid-to-late shift handover point, and the calculation formula is as follows: in, For the remaining range of the aircraft, The set distance threshold; This refers to the hyperparameters used to regulate weight changes in latitude and longitude rewards; This is used to adjust the weight of latitude and longitude rewards on the aircraft. This is the actual longitude value of the mid-to-late shift handover point; This is the actual longitude value of the mid-to-late shift handover point; , These represent the Euclidean distances between the latitude and longitude of the aircraft in the current calculation cycle and the latitude and longitude of the mid-to-latest handover point in the previous calculation cycle.
6. The intelligent guidance method for aircraft based on reinforcement learning according to claim 1, characterized in that... The vertical reward It is the sum of energy reward, speed reward, and altitude reward.
7. The intelligent guidance method for aircraft based on reinforcement learning according to claim 6, characterized in that, The energy reward The calculation formula is as follows: In the formula, The mean is The standard deviation is Take from Gaussian distribution The probability, This is the offset amount used to adjust the speed error. The offset amount used to adjust the height error. For the speed of the aircraft, For the Earth's radius, Represents the gravitational constant; The speed reward The calculation formula is as follows: In the formula, This indicates that the mean is The standard deviation is Take from Gaussian distribution The probability, This is the offset amount used to adjust the speed error; The height reward The calculation formula is as follows: In the formula, This indicates that the mean is The standard deviation is Take from Gaussian distribution The probability, This is the offset amount used to adjust the height error; This is the hyperparameter used to adjust the weight changes in this reward. This refers to the flight altitude of the aircraft.
8. The intelligent guidance method for aircraft based on reinforcement learning according to claim 1, characterized in that... The final reward The calculation formula is as follows when the aircraft arrives at the final handover point: In the formula, These are the influence coefficients used to balance the effects of distance, longitude, latitude, and speed amplitude on the aircraft. These are hyperparameters used to calculate the relative errors of the aircraft's distance, longitude, latitude, and speed, respectively, with the final reward being 0 at other times.