Intelligent control method for attitude of hypersonic aircraft based on global quadratic heuristic programming
The mathematical model of the hypersonic aircraft is converted into an unconstrained form through global quadratic heuristic planning and obstacle conversion function. Combined with the dynamic programming algorithm of the incremental model, an intelligent flight controller is designed to solve the problems of attitude control accuracy and stability of hypersonic aircraft in complex environments, and achieve high-precision and safe attitude control.
Patent Information
- Application Number
- CN202510814697.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
Existing hypersonic vehicle attitude control methods are difficult to ensure high precision and stability in complex flight environments, especially under conditions of model uncertainty, strong nonlinearity and strong coupling, traditional methods cannot effectively improve control accuracy.
An intelligent control method based on global quadratic heuristic planning is adopted. By establishing a mathematical model of the hypersonic aircraft and converting it into an unconstrained form using an obstacle conversion function, an intelligent flight controller is designed in combination with a global quadratic heuristic dynamic programming algorithm of the incremental model to achieve attitude control of the hypersonic aircraft.
It significantly improves the attitude control accuracy and safety of hypersonic aircraft under complex dynamic conditions, can achieve high-precision attitude control under the conditions of aerodynamic parameters and center of mass offset, suppress the adverse effects of parameter perturbations and center of mass offset, and ensure the stability and robustness of attitude control.
Smart Images

Figure CN120669732A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of hypersonic aircraft attitude control, and in particular to a hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning. Background Art
[0002] Due to the complex flight environment and special mission requirements of hypersonic aircraft, the strong environmental interference, strong nonlinearity, strong coupling, and strong uncertainty of the model pose daunting challenges to the selection and design of its attitude controller. Currently, traditional attitude control methods for hypersonic aircraft mainly include linear control methods based on small-disturbance linearization, split-channel design controller methods, and model-based control methods. Among them, the linear control method based on small-disturbance linearization is difficult to guarantee the control accuracy of hypersonic aircraft. The split-channel design controller method is difficult to simultaneously guarantee longitudinal and lateral tracking accuracy, and therefore cannot guarantee the control accuracy of hypersonic aircraft. Model-based control methods cannot effectively handle model uncertainty and therefore cannot guarantee the control accuracy of hypersonic aircraft.
[0003] Based on this, how to provide an attitude control method for hypersonic aircraft with higher control accuracy has become a technical problem that needs to be solved urgently in this field. Summary of the Invention
[0004] The purpose of this application is to provide a hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning, which can improve the attitude control accuracy of hypersonic aircraft.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] The present application provides a hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning, which specifically includes the following steps.
[0007] Acquire aerodynamic data of target hypersonic aircraft.
[0008] A hypersonic aircraft mathematical model is established based on the aerodynamic data of the target hypersonic aircraft; the hypersonic aircraft mathematical model is a mathematical model with state constraints.
[0009] The obstacle conversion function is used to convert the hypersonic aircraft mathematical model into an unconstrained hypersonic aircraft conversion system, and the tracking error and system dynamic information of the hypersonic aircraft conversion system are determined.
[0010] According to the attitude tracking error and system dynamic information of the hypersonic aircraft conversion system, a global quadratic heuristic dynamic programming algorithm based on an incremental model is adopted to perform attitude control on the hypersonic aircraft conversion system.
[0011] According to the specific embodiments provided in this application, this application has the following technical effects.
[0012] The present application provides a hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming. By establishing a hypersonic aircraft mathematical model and using an obstacle conversion function to convert the state-constrained hypersonic aircraft mathematical model into an unconstrained hypersonic aircraft conversion system, the state constraint problem is converted into an unconstrained optimization problem. Combined with the global quadratic heuristic dynamic programming algorithm based on the incremental model, the attitude control of the hypersonic aircraft conversion system is performed. It can realize real-time strategy optimization in a complex dynamic environment, effectively suppress the adverse effects of parameter perturbations and center of mass shifts, and at the same time ensure the stability and safety of attitude control, significantly improving the attitude control accuracy, safety and robustness of hypersonic aircraft under complex dynamic conditions, and fully demonstrating the advantages of the global quadratic heuristic dynamic programming algorithm based on the incremental model, an online learning algorithm, in solving state-constrained optimal control problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0014] Figure 1 A flowchart of a hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning is provided in accordance with an embodiment of the present application.
[0015] Figure 2 A schematic diagram of the structure of an intelligent flight controller provided in one embodiment of the present application.
[0016] Figure 3 A schematic diagram of the structure of a critic network provided in one embodiment of the present application.
[0017] Figure 4 A schematic diagram of the structure of the Actor network provided in one embodiment of the present application.
[0018] Figure 5 A graph showing the angle of attack tracking effect in an attitude intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0019] Figure 6 A graph showing the angle of attack tracking error in a simulation of attitude intelligent control verification based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0020] Figure 7 A graph showing the sideslip angle tracking effect in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0021] Figure 8 A graph showing the sideslip angle tracking error in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0022] Figure 9 A graph showing the roll angle tracking effect in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0023] Figure 10 A graph showing the roll angle tracking error in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0024] Figure 11 This is a graph showing the roll angular velocity variation in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0025] Figure 12 A graph showing the yaw rate variation in the attitude intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0026] Figure 13 A graph showing the pitch angular velocity variation in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0027] Figure 14 A schematic diagram of control input in a posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in an embodiment of the present application.
[0028] Figure 15 This is a graph showing the weight change from the critic hidden layer to the output layer in the posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0029] Figure 16 This is a graph showing the weight change from the critic input layer to the hidden layer in the posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0030] Figure 17A graph showing the weight change from the actor hidden layer to the output layer in the posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0031] Figure 18 This is a graph showing the weight change from the Actor input layer to the hidden layer in the posture intelligent control verification simulation based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0032] Figure 19 A graph showing the angle of attack tracking effect in a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0033] Figure 20 A graph showing the angle of attack tracking error in a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0034] Figure 21 A graph showing the sideslip angle tracking effect in a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0035] Figure 22 A graph showing the sideslip angle tracking error in a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0036] Figure 23 A graph showing the roll angle tracking effect in a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming provided in one embodiment of the present application.
[0037] Figure 24 A graph showing the roll angle tracking error in a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming provided in one embodiment of the present application. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0039] This embodiment proposes a hypersonic vehicle attitude intelligent control method based on global quadratic heuristic programming. This method aims to leverage the powerful learning capabilities of neural networks by studying attitude controller design methods for hypersonic aircraft and combining them with intelligent algorithms such as deep reinforcement learning. The method uses an incremental-model-based globalized dual heuristic programming (IGDHP) algorithm to achieve high-precision attitude control in the presence of aerodynamic parameter and center of mass offsets. Furthermore, considering that excessive changes in state and tracking errors during the online learning process of the IGDHP algorithm may cause instability in the attitude of the hypersonic aircraft, a safe intelligent control method based on global quadratic heuristic programming (SIGDHP) is introduced. By using an obstacle conversion function to transform the input state of the IGDHP algorithm, the SIGDHP algorithm is transformed into an unconstrained optimization problem. This method can suppress the adverse effects of parameter perturbations, improve the control accuracy of the hypersonic aircraft, and achieve safe, intelligent, and precise attitude control of the hypersonic aircraft in the face of parameter perturbations.
[0040] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0041] This embodiment proposes a hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming. This method is applied to hypersonic aircraft attitude control scenarios and primarily aims to improve the accuracy of hypersonic aircraft attitude intelligent control. By precisely controlling the hypersonic aircraft's attitude, the adverse effects of parameter perturbations and center of mass shifts are suppressed, improving the control stability and safety of the hypersonic aircraft under parameter perturbations and center of mass shifts. This ensures the accuracy, stability, and safety of the hypersonic aircraft's attitude control in complex flight environments and special missions, thus meeting the complex flight environment and special mission requirements of the hypersonic aircraft. The method specifically includes the following steps.
[0042] Step S1: Acquire aerodynamic data of a target hypersonic aircraft.
[0043] Hypersonic aircraft generally refer to any aircraft, artillery shells, missiles, or other winged or wingless aircraft with a flight speed of Mach 5 or higher. In this embodiment, the target hypersonic aircraft is a CAV-H (High-Performance Common Aero Vehicle) hypersonic aircraft, which is used as the simulation research object.
[0044] Step S2: Establishing a hypersonic aircraft mathematical model based on the aerodynamic data of the target hypersonic aircraft, wherein the hypersonic aircraft mathematical model is a mathematical model with state constraints.
[0045] Step S3: using an obstacle conversion function to convert the state-constrained hypersonic aircraft mathematical model into an unconstrained hypersonic aircraft conversion system, and determining the tracking error and system dynamic information of the hypersonic aircraft conversion system.
[0046] This embodiment considers modeling the attitude control problem as an optimization problem. However, a purely intelligent approach requires certain constraints on the state, making the attitude control problem a state-constrained optimization problem. However, the IGDHP algorithm cannot be directly applied to solve state-constrained optimization problems. Therefore, this embodiment introduces an obstacle conversion function. This function converts the tracking error and angular velocity state quantities of the target hypersonic aircraft, thereby transforming the state-constrained optimization problem into an unconstrained optimization problem. This allows the IGDHP algorithm to be directly applied to intelligent attitude control of the target hypersonic aircraft.
[0047] Step S4: Based on the attitude tracking error and system dynamic information of the hypersonic aircraft conversion system, the IGDHP algorithm is used to perform attitude control on the hypersonic aircraft conversion system, thereby realizing intelligent attitude control of the target hypersonic aircraft.
[0048] This embodiment establishes a mathematical model of a hypersonic aircraft, thereby modeling the attitude control problem of a hypersonic aircraft including aerodynamic parameter perturbations and center of mass offset as a constrained optimal control problem. By introducing an obstacle conversion function, the mathematical model of the hypersonic aircraft is converted into an unconstrained hypersonic aircraft conversion system, thereby converting the state constraint problem into an unconstrained optimization problem. Then, the attitude control of the hypersonic aircraft is studied based on the IGDHP algorithm. The structure of the complete intelligent flight controller is as follows: Figure 2 shown.
[0049] In this embodiment, the mathematical model of the hypersonic aircraft is converted into an unconstrained hypersonic aircraft conversion system through the obstacle conversion function, so that the control problem of the hypersonic aircraft with full state constraints can be converted into an unconstrained optimal control problem, which can be expressed as:
[0050]
[0051] Among them, J * (z k ,u k ) is the optimal cost function at time k, is the state quantity of the equivalent system, is the control variable of the hypersonic vehicle; η is the discount factor; L k (z k ,u k ) is the single-step cost, satisfying is a positive definite diagonal matrix, and the sensitivity to the control quantity can be changed by adjusting R. t ) is related to z t A quadratic function of is the optimal control at time k.
[0052] This embodiment takes into account that it is almost impossible to directly solve the above-mentioned optimal control problem. Therefore, this embodiment adopts the IGDHP method based on the Actor-Critic architecture to approximate the solution. The Critic network of IGDHP can simultaneously estimate the performance index function and the co-state function, reducing the error caused by back propagation. Therefore, the performance is far superior to IHDP and IDHP in dealing with the optimal control problem of high-dimensional nonlinear strongly coupled systems. Simultaneously estimating the performance index function and the co-state function further increases the computational complexity, and relying solely on a single Critic network to directly output two estimators cannot guarantee the derivative of the performance index function estimate with respect to the state. Estimation of the coordinated state Therefore, this embodiment proposes a novel dual critic network structure, that is, two critic networks (critic network and companion critic network) with shared weights output the estimation of performance index function and co-state function respectively, and selecting a suitable activation function can satisfy This structure not only reduces estimation errors but also lowers computational complexity, further improving performance and computational speed in complex situations. The Actor Network is a three-layer, fully connected neural network whose goal is to find a strategy that minimizes a defined performance function.
[0053] The structure of the intelligent flight controller is as follows Figure 2 As shown, Figure 2 The black solid line in the middle represents signal transmission, the green dotted line represents backpropagation to update the model, and the orange dotted line represents the calculation process of the TD error between the Actor network and the Critic network.
[0054] This embodiment uses the obstacle conversion function to convert the attitude tracking error Ω of the hypersonic aircraft at time t-1 into t-1 -Ω d,t-1 and angular velocity ω t-1 Transform to the state z of the equivalent unconstrained system t-1 , as the input of IGDHP. IGDHP consists of two critic networks (critic network and companion critic network), an actor network and an incremental model, where the critic network and the companion critic network share network weights and output performance indicator functions J(z t-1 ,u t-1 ) and the co-state function λ(z t-1 ) and Combined with the state matrix of the error system obtained by incremental model identification and To calculate the TD error of the Critic network and the Actor network, and then guide the gradient descent update of the network parameters.
[0055] This embodiment takes the attitude control of a hypersonic aircraft as its goal, fully considering the high-precision attitude control problem when its aerodynamic parameters and center of mass are offset. The attitude control problem of a hypersonic aircraft involving aerodynamic parameter perturbations and center of mass offsets is modeled as a constrained optimal control problem. By introducing an obstacle conversion function, the state constraint problem is converted into an unconstrained optimization problem, thereby effectively avoiding direct conflicts of state constraints. A global quadratic heuristic dynamic programming algorithm based on an incremental model is adopted to directly utilize the tracking error of the conversion system and the system dynamic information to achieve precise attitude control of the aircraft. Specifically, the following steps are included:
[0056] Step 1: Establish a mathematical model of hypersonic aircraft. It can be expressed as:
[0057]
[0058] Where Ω represents the state of the attitude loop, Ω=[μ,β,α] T , μ is the roll angle, β is the sideslip angle, α is the angle of attack; ω represents the state of the angular velocity loop, ω=[ω x ,ω y ,ω z ] T ,ω x is the roll angular velocity, ω y is the yaw angular velocity, ω z is the pitch angular velocity; δ represents the control input, δ=[δ x ,δ y ,δ z ] T , δx is the aileron deflection angle, δ y is the rudder deflection angle, δ z is the elevator deflection angle; f Ω 、f ω are the lumped disturbances of the attitude loop and angular velocity loop respectively; g Ω 、g ω Represent the control input matrices in the attitude loop and angular velocity loop respectively.
[0059] Step 2: Utilize the barrier conversion function Where τ represents the inverse function of the obstacle transfer function, a is the set state quantity boundary limit, and a>0; l is the state quantity, and -a<l<a. For the constrained hypersonic aircraft mathematical model in step 1, let the expected instruction be Ω d =[μ d ,β d ,α d ] T , where μ d , β d , α d Denote the roll angle command, sideslip angle command and angle of attack command respectively. Then the attitude tracking error can be written as e Ω =Ω-Ω d =[e μ ,e β ,e α ] T , where Ω represents the state quantity of the attitude loop; e μ 、e β 、e α Denote the tracking error of the roll angle command, the sideslip angle command, and the angle of attack command, respectively.
[0060]
[0061] in, They represent the conversion amount of the roll angle tracking error, the conversion amount of the sideslip angle tracking error, and the conversion amount of the angle of attack tracking error respectively; Respectively represent the conversion amount of roll angular velocity, the conversion amount of yaw angular velocity, and the conversion amount of pitch angular velocity; represents the attitude tracking error of the conversion system; Z ω Indicates the state of the conversion system; represents the barrier transfer function; Indicates definition equal to; superscript T indicates transposition; is the boundary constraint limit of the posture tracking error; l ω is the boundary constraint of the angular velocity state. The mathematical model of the hypersonic aircraft can then be converted into the following formula:
[0062]
[0063] in, represents the derivative of the state quantity of the conversion system; F(z) represents the lumped disturbance matrix of the conversion system; G(z) represents the control matrix of the conversion system. F(z) and G(z) satisfy the following equation:
[0064]
[0065] where ε(·) represents the inverse function of the derivative of the barrier transfer function with respect to the boundary l.
[0066] Using the obstacle transfer function, the attitude tracking error Ω of the CAV hypersonic aircraft at time t-1 is converted to t-1 -Ω d,t-1 and angular velocity ω t-1 Transform to the state z of the equivalent unconstrained system t-1 , as the input of IGDHP.
[0067] In this embodiment, IGDHP consists of two Critic networks, an Actor network and an incremental model, where the two Critic networks share network weights and output performance indicator functions J(z t-1 ,u t-1 ) Co-state function λ(z t-1 ) Combined with the state matrix of the error system obtained by incremental model identification and To calculate the TD error of the Critic network and the Actor network, and then guide the gradient descent update of the network parameters.
[0068] Step 3: In order to simplify the gradient descent update strategy of the IGDHP algorithm, the state z of the system is converted according to the hypersonic vehicle t-1 With control input u t-1 , the parameter matrix is identified using an incremental model based on recursive least squares. The update steps are as follows:
[0069]
[0070] Among them, ε t is the state estimation error at time t, is the parameter matrix of the incremental model; Δz t =z t -z t-1 is the increment of the equivalent system state at time t; Δu t =u t -u t-1is the increment of the control input at time t; F t-1 and G t-1 are all parameter matrices estimated by the incremental model; is the covariance matrix at time t, and satisfies is the covariance matrix at time t0, ρ is a large positive number, I is the n+m dimensional identity matrix; k is the forgetting factor.
[0071] According to the parameter matrix at time t-1 The state quantity z of the hypersonic aircraft conversion system at time t+1 can be predicted t+1 , expressed as:
[0072]
[0073] Among them, z t+1 is the state quantity at time t+1; z t is the state quantity at time t.
[0074] Step 4: According to the conversion system state z at time t t , control quantity u t , determine the performance indicator function. The performance indicator function is as follows:
[0075]
[0076] Among them, J(z t ,u t ) is the state quantity z at time t t and control quantity u t The corresponding performance index function value; J(z t+1 ,u t+1 ) is the state quantity z at time t+1 t+1 and control quantity u t+1 Corresponding performance index function; L l (z l ,u l ) represents the state quantity z at time l l and control quantity u l The corresponding single-step cost; L t (z t ,u t ) represents the state quantity z at time t t and control quantity u t The corresponding single-step cost is, is the integral term of the tracking error, t0 is the time when the simulation starts, is the equivalent system tracking error obtained after transformation by the obstacle conversion function at time l; is a positive definite diagonal matrix. By adjusting Q and R, the sensitivity to the state quantity, control quantity and tracking error integral term is changed; η is the learning rate.
[0077] Step 5: Design Figure 3 The Critic network shown is similar to Figure 4 The Actor network shown, according to the state z of the transition system t , the estimation of performance cost function, estimation of co-state function and control variables are as follows:
[0078]
[0079] in, is the estimation of the performance cost function, w c1 is the network weight from the input layer to the hidden layer; w c2 is the network weight from the hidden layer to the output layer; σ(·) is the Softplus activation function, satisfying σ(x)=ln(1+e x ); is an estimate of the co-state function, satisfying A(x) is the Sigmoid activation function, satisfying u IGDHP,t+1 is the control variable; w a2 is the network weight from the hidden layer to the output layer; is the Tanh activation function; Y a1 is the input of the hidden layer node; w a1 is the network weight from the input layer to the hidden layer, z t is the state quantity at time t. IGDHP,t+1 Input the hypersonic aircraft system model to obtain the state quantity Ω at the next moment t+1 With ω t+1 .
[0080] Step 6: Estimate the system parameter matrix F based on the incremental model t-1 With G t-1 , combined with the network weight w of the Critic network c1 With w c2 , the network weight w of the Actor network a1 With w a2 And the state quantity z of the conversion system at time t t , control quantity u t , we can get the weight update law of the Critic network and the Actor network, which is expressed as follows:
[0081]
[0082] Where Δw c (t) represents the gradient of the critic network weight; Δw a (t) represents the gradient of the Actor network weight; η aRepresents the learning rate of the Actor network; η c Represents the learning rate of the Critic network; Represents the forgetting factor of the Ctitic network; represents the forgetting factor of the critic companion network; L t-1 represents the single-step cost at time t-1; It represents the product of the control matrix obtained by incremental model identification at time t-1 and the partial derivative of the conversion system with respect to the control input.
[0083] Through the above steps, an intelligent flight control algorithm based on global quadratic heuristic dynamic programming can be obtained.
[0084] In this embodiment, an intelligent flight control algorithm based on global quadratic heuristic dynamic programming is used to perform attitude intelligent control simulation, which mainly includes the following steps:
[0085] Step (1): Initialize the network weights w of the Actor network and the Critic network a (t0), w c (t0) and various hyperparameters, where t0 is the time when the simulation starts.
[0086] Step (2): The posture tracking error at time t and angular velocity ω t After the transformation of the barrier transfer function, it is transformed into the state z of the equivalent system t , as the input of the Actor network to obtain the control amount u IGDHP,t+1 .
[0087] Step (3): Use the incremental model to calculate the increment Δz of the state of the equivalent system at time t t With input Δu IGDHP,t+1 Identify the parameter matrix Θ of the incremental model t-1 , and then calculate the gradients of the Critic network and the Actor network respectively according to the formula.
[0088] Step (4): Update the incremental model, Critic network and Actor network respectively according to the formula.
[0089] Step (5): Loop steps (2) to (4) until the simulation ends.
[0090] In this embodiment, when aerodynamic parameters are offset and there is a mismatched disturbance, the effectiveness of the intelligent flight control algorithm based on global quadratic heuristic dynamic programming is verified by giving sinusoidal bank angle and angle of attack commands.
[0091] In this embodiment, the speed V, altitude h, angle of attack α, sideslip angle β, bank angle μ, and roll angular velocity ω of the hypersonic aircraft are:x , yaw angular velocity ω y and the pitch angular velocity ω z The initial state value is set as: [V0,h0,α0,β0,μ0,ω x0 ,ω y0 ,ω z0 ]=[3500m / s,30000m,0°,0°,0°,0° / s,0° / s,0° / s]; a non-matching disturbance d(t)=0.5°sin(4t) is given in each attitude channel; considering that the parameter perturbation starts at t=10s and ends at t=15s, the size of the parameter perturbation is as shown; at t=0s, a sinusoidal angle of attack command signal with an amplitude of 5° and a frequency of 1 / 3 is given, and a sinusoidal roll angle command signal with an amplitude of 5° and a frequency of 1 / 3 is given.
[0092] In order to verify the performance and applicability of the hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning in this embodiment, a semi-physical (Hardware In the Loop, HIL) simulation verification platform for the hypersonic aircraft control algorithm was built in this embodiment. Semi-physical simulation, also known as semi-physical simulation or hardware-in-the-loop simulation, refers to a real-time simulation that replaces a part of the physical object with a digital model, and then forms a simulation loop with other physical hardware. Compared with full-physical testing, semi-physical simulation can effectively reduce the cost of the test, reduce the risk of the test, and thus shorten the research cycle; compared with full digital simulation, semi-physical simulation can fully verify the performance and applicability of the algorithm, and further improve the authenticity and credibility of the simulation results.
[0093] The physical-in-the-loop (HIP) simulation and verification platform for hypersonic vehicle control algorithms built in this example includes an STM32F767-based flight control computer, an Nvidia Jetson Orin Nx high-performance computer (16Gb, 100TOPS computing power at maximum power mode), a Lingsi Chuangqi rapid control prototyping real-time simulator, a host computer, and a ground station. Each component of the HIP simulation and verification platform uses the RS422 serial port protocol for asynchronous communication, and all transmitted data is serial data frames.
[0094] This embodiment takes the CAV-H hypersonic aircraft as an example, and the semi-physical simulation verification process is as follows.
[0095] (1) First, the hypersonic aircraft mathematical model of the CAV-H hypersonic aircraft is imported into the rapid control prototype through the host computer. Secondly, the nominal flight controller algorithm is downloaded to the flight control computer through the host computer, and the environment is configured in Nvidia Jetson and the intelligent algorithm is deployed.
[0096] (2) The various components in the semi-physical simulation verification platform of the hypersonic aircraft control algorithm are connected through the RS422 serial port line, and the simulation parameters, including the initial state, pull-off coefficient, etc., are set in the host computer.
[0097] (3) Turn on the power of the Nvidia Jetson OrinNx high-performance computer and the flight control computer, activate the rapid control prototype in the host computer, and observe the command tracking status in the ground station and the host computer respectively to verify the effectiveness and applicability of the attitude intelligent control method of this embodiment.
[0098] Based on the hypersonic vehicle control algorithm semi-physical simulation verification platform and the above-mentioned technical solution, this embodiment carried out a global quadratic heuristic dynamic programming-based attitude intelligent control verification simulation to verify the effectiveness of the above-mentioned method. The simulation parameter settings are shown in Table 1.
[0099] Table 1 Simulation parameter settings
[0100] symbol significance Numerical ρ Covariance matrix parameters <![CDATA[1×10 8 ]]> k Incremental model forgetting factor 1.01 Q Cost function state gain matrix diag(10,10,10,10,10,10) R The cost function controls the input gain matrix diag(10,10,10) P Cost function error integral gain matrix diag(0.01,0.01,0.01) η Discount Factor 0.9 <![CDATA[σ c ]]> Critic learning rate <![CDATA[3×10 -3 ]]> <![CDATA[σ a ]]> Actor learning rate <![CDATA[3×10 -3 ]]> Number of Actor and Critic hidden layer nodes 100 Δt Simulation step 0.01s <![CDATA[δ max d Vmax ]]> Rudder deflection limit and speed limit 20°,65° / s <![CDATA[l eΩ ]]> Attitude angle tracking error boundary constraint 2° <![CDATA[l ω ]]> Angular velocity state constraint 30° / s
[0101] The simulation parameters set in this embodiment are shown in Table 2.
[0102] Table 2 Simulation pull-off parameter settings
[0103]
[0104]
[0105] Figure 5 This is the angle of attack tracking effect curve; Figure 6 is the angle of attack tracking error curve; Figure 7 It is a curve diagram of the sideslip angle tracking effect; Figure 8 is the sideslip angle tracking error curve; Figure 9 This is a curve diagram of the roll angle tracking effect; Figure 10 This is the curve of the roll angle tracking error. Figure 11 It is the curve diagram of the roll angular velocity change. Figure 12 It is the curve diagram of yaw angular velocity change. Figure 13 is the pitch angular velocity variation curve. Figures 5 to 13 It can be seen that the attitude angle tracking error and angular velocity throughout the entire process are both within the state constraint boundaries. By updating the network weights of the Actor network and the Critic network through online learning, the attitude tracking error gradually converges. When the aerodynamic force, torque coefficient, and center of mass offset occur, the controller can quickly suppress the attitude angle jump and achieve high-precision tracking of the attitude command.
[0106] Figure 14 is the control input diagram, Figure 14Indicates that the rudder deflection command curve is smooth.
[0107] Figure 15 This is a curve diagram of the weight change from the Critic hidden layer to the output layer. Figure 16 This is a curve diagram of the weight change from the Critic input layer to the hidden layer. Figure 17 This is a graph showing the weight change from the Actor's hidden layer to the output layer. Figure 18 This is a graph showing the weight change from the Actor input layer to the hidden layer. Figures 15 to 18 This shows that the Critic network and the Actor network continuously update their weights to improve tracking accuracy in different environments.
[0108] In order to further verify the superiority of the technical solution designed in this embodiment, it is compared with the improved PID controller based on deep reinforcement learning, which is also model-free, and a comparative simulation of attitude intelligent control based on global quadratic heuristic dynamic programming is carried out. The initial state value, pull-off parameter setting and tracking curve of the hypersonic aircraft are the same as those described above. The hypersonic aircraft attitude angle tracking performance and control input curve are shown in Figure 2. Figures 19 to 24 As shown, Figure 19 This is the angle of attack tracking effect curve; Figure 20 is the angle of attack tracking error curve; Figure 21 It is a curve diagram of the sideslip angle tracking effect; Figure 22 is the sideslip angle tracking error curve; Figure 23 This is a curve diagram of the roll angle tracking effect; Figure 24 This is the curve of the roll angle tracking error. Figures 19 to 24 It shows that when the aerodynamic force, torque parameters and center of mass offset range, the intelligent flight control algorithm based on global quadratic heuristic dynamic programming sacrifices some angle of attack control accuracy, effectively suppresses the jump of bank angle and sideslip angle, and shows higher tracking accuracy and faster error convergence response speed.
[0109] This embodiment addresses the problem of high-precision attitude control of hypersonic vehicles under uncertain conditions such as aerodynamic parameter changes and center of mass shifts. By combining an obstacle transfer function, an intelligent control algorithm based on global quadratic heuristic dynamic programming is designed, achieving safe and efficient control under state constraints. This method enables real-time policy optimization in complex dynamic environments, effectively suppressing the adverse effects of parameter perturbations and center of mass shifts while ensuring the stability and safety of attitude control. This embodiment significantly improves the accuracy, safety, and robustness of attitude control for hypersonic vehicles under complex dynamic conditions, fully demonstrating the advantages of the IGDHP algorithm, an online learning algorithm, in solving state-constrained optimal control problems.
[0110] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0111] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A hypersonic vehicle attitude intelligent control method based on global quadratic heuristic programming, characterized in that: The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning includes: Acquire aerodynamic data of target hypersonic aircraft; Establishing a mathematical model of the hypersonic aircraft based on the aerodynamic data of the target hypersonic aircraft; the mathematical model of the hypersonic aircraft is a mathematical model with state constraints; Using an obstacle conversion function, converting the hypersonic aircraft mathematical model into an unconstrained hypersonic aircraft conversion system, and determining a tracking error and system dynamic information of the hypersonic aircraft conversion system; According to the attitude tracking error and system dynamic information of the hypersonic aircraft conversion system, a global quadratic heuristic dynamic programming algorithm based on an incremental model is adopted to perform attitude control on the hypersonic aircraft conversion system.
2. The hypersonic vehicle attitude intelligent control method based on global quadratic heuristic programming according to claim 1 is characterized in that: The mathematical model of the hypersonic aircraft is expressed as: Where Ω represents the state of the attitude loop, Ω=[μ,β,α] T , μ is the roll angle, β is the sideslip angle, α is the angle of attack; ω represents the state of the angular velocity loop, ω=[ω x ,ω y ,ω z ] T ,ω x is the roll angular velocity, ω y is the yaw angular velocity, ω z is the pitch angular velocity; δ represents the control input, δ=[δ x ,δ y ,δ z ] T , δ x is the aileron deflection angle, δ y is the rudder deflection angle, δ z is the elevator deflection angle; f Ω 、f ω are the lumped disturbances of the attitude loop and angular velocity loop respectively; g Ω 、g ω Represent the control input matrices in the attitude loop and angular velocity loop respectively.
3. The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming according to claim 2 is characterized in that: The obstacle conversion function is used to convert the hypersonic aircraft mathematical model into an unconstrained hypersonic aircraft conversion system, and the tracking error and system dynamic information of the hypersonic aircraft conversion system are determined, specifically including: Set the expected instruction to Ω d =[μ d ,β d ,α d ] T , where μ d , β d , α d They represent the roll angle command, sideslip angle command and angle of attack command respectively; the attitude tracking error is e Ω =Ω-Ω d =[e μ ,e β ,e α ] T , where Ω represents the state quantity of the attitude loop; e μ 、e β 、e α They represent the roll angle command tracking error, sideslip angle command tracking error, and angle of attack command tracking error respectively; then: in, They represent the conversion amount of the roll angle tracking error, the conversion amount of the sideslip angle tracking error, and the conversion amount of the angle of attack tracking error respectively; Respectively represent the conversion amount of roll angular velocity, the conversion amount of yaw angular velocity, and the conversion amount of pitch angular velocity; represents the attitude tracking error of the conversion system; Z ω Indicates the state of the conversion system; represents the barrier transfer function; Indicates definition equal to; superscript T indicates transposition; is the boundary constraint limit of the posture tracking error; l ω is the boundary constraint limit of the angular velocity state; Establish the barrier conversion function, expressed as: Wherein, τ represents the inverse function of the barrier conversion function, a is the set state quantity boundary limit, and a>0; l is the state quantity, and -a<l<a; Using the obstacle conversion function, the hypersonic aircraft mathematical model is converted into an unconstrained hypersonic aircraft conversion system, which is expressed as: in, represents the derivative of the state quantity of the conversion system; F(z) represents the lumped disturbance matrix of the conversion system; G(z) represents the control matrix of the conversion system; F(z) and G(z) satisfy: where ε(·) represents the inverse function of the derivative of the barrier transfer function of the boundary l; Tracking error and system dynamic information are determined based on the hypersonic vehicle conversion system.
4. The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming according to claim 3 is characterized in that: The global quadratic heuristic dynamic programming algorithm based on the incremental model includes a critic network, an accompanying critic network, an actor network and an incremental model; The critic network and the accompanying critic network share network weights, and the critic network and the accompanying critic network are used to output performance indicator functions J(z t-1 ,u t-1 ) and the co-state function λ(z t-1 ) The incremental model is used to calculate the performance index function J(z t-1 ,u t-1 ) and the co-state function λ(z t-1 ) The identified state matrix and The TD errors of the critic network, the companion critic network, and the actor network are calculated to guide the gradient descent update of the network parameters of the global quadratic heuristic dynamic programming algorithm based on the incremental model.
5. The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming according to claim 4 is characterized in that: According to the attitude tracking error and system dynamic information of the hypersonic aircraft conversion system, a global quadratic heuristic dynamic programming algorithm based on an incremental model is used to perform attitude control on the hypersonic aircraft conversion system, specifically including: According to the state quantity z of the hypersonic aircraft conversion system at t-1 t-1 and control quantity u t-1 , using the incremental model to identify the parameter matrix, expressed as: Among them, ε t is the state estimation error at time t, is the parameter matrix of the incremental model; Δz t =z t -z t-1 is the increment of the equivalent system state at time t; Δu t =u t -u t-1 is the increment of the control input at time t; F t-1 and G t-1 are all parameter matrices estimated by the incremental model; is the covariance matrix at time t, and satisfies is the covariance matrix at time t0, ρ is a maximum positive number, I is the n+m-dimensional identity matrix; k is the forgetting factor; According to the parameter matrix at time t-1 Predict the state quantity z of the hypersonic aircraft conversion system at time t+1 t+1 , expressed as: Among them, z t+1 is the state quantity at time t+1; z t is the state quantity at time t; According to the state quantity z of the hypersonic aircraft conversion system at time t t and control quantity u t , determine the performance index function, expressed as: Among them, J(z t ,u t ) is the state quantity z at time t t and control quantity u t The corresponding performance index function value; J(z t+1 ,u t+1 ) is the state quantity z at time t+1 t+1 and control quantity u t+1 Corresponding performance index function; L l (z l ,u l ) represents the state quantity z at time l l and control quantity u l The corresponding single-step cost; L t (z t ,u t ) represents the state quantity z at time t t and control quantity u t The corresponding single-step cost is, is the integral term of the tracking error, t0 is the time when the simulation starts, is the equivalent system tracking error obtained after transformation by the obstacle conversion function at time l; is a positive definite diagonal matrix. By adjusting Q and R, the sensitivity to the state quantity, control quantity and tracking error integral term is changed; η is the learning rate; According to the state quantity z of the hypersonic aircraft conversion system at time t t and the performance indicator function, calculate the estimation of the performance cost function, the estimation of the co-state function and the control variables, expressed as: in, is the estimation of the performance cost function, w c1 is the network weight from the input layer to the hidden layer; w c2 is the network weight from the hidden layer to the output layer; σ(·) is the Softplus activation function, satisfying σ(x)=ln(1+e x ); is an estimate of the co-state function, satisfying A(x) is the Sigmoid activation function, satisfying u IGDHP,t+1 is the control variable; w a2 is the network weight from the hidden layer to the output layer; is the Tanh activation function; Y a1 is the input of the hidden layer node; w a1 is the network weight from the input layer to the hidden layer, z t is the state quantity at time t; The parameter matrix F estimated according to the incremental model t-1 With G t-1 , the network weights w of the critic network and the associated critic network c1 With w c2 , the network weight w of the Actor network a1 With w a2 and the state quantity z of the hypersonic aircraft conversion system at time t t and control quantity u t , determine the weight update law of the Critic network, the accompanying Critic network, and the Actor network, expressed as: Where Δw c (t) represents the gradient of the critic network weight; Δw a (t) represents the gradient of the Actor network weight; η a Represents the learning rate of the Actor network; η c Represents the learning rate of the Critic network; Represents the forgetting factor of the Ctitic network; represents the forgetting factor of the critic companion network; L t-1 represents the single-step cost at time t-1; It represents the product of the control matrix obtained by incremental model identification at time t-1 and the partial derivative of the conversion system with respect to the control input.
6. The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming according to claim 5 is characterized in that: The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic planning also includes: An intelligent flight control algorithm based on global quadratic heuristic dynamic programming is used to simulate attitude intelligent control.
7. The hypersonic aircraft attitude intelligent control method based on global quadratic heuristic programming according to claim 6 is characterized in that: An intelligent flight control algorithm based on global quadratic heuristic dynamic programming is used to simulate attitude intelligent control, including: Initialize the network weights w of the Actor network, the Critic network, and the companion Critic network a (t0), w c (t0) and various hyperparameters, where t0 is the time when the simulation starts; The posture tracking error at time t With angular velocity ω t The state quantity z of the equivalent system is converted into the state quantity z by the barrier transfer function t , as the input of the Actor network to obtain the control amount u IGDHP,t+1 ; According to the increment Δz of the state of the equivalent system at time t t With input Δu IGDHP,t+1 , the parameter matrix Θ of the incremental model is obtained by using the incremental model identification t-1 , and calculate the gradients of the Critic network, the accompanying Critic network, and the Actor network respectively; Update the incremental model, the critic network, the accompanying critic network, and the actor network according to the gradients of the critic network, the accompanying critic network, and the actor network; Repeat step "to set the posture tracking error at time t With angular velocity ω t The state quantity z of the equivalent system is converted into the state quantity z by the barrier transfer function t , as the input of the Actor network to obtain the control amount u IGDHP,t+1 "To step "update the incremental model, the critic network, the associated critic network, and the actor network according to the gradients of the critic network, the associated critic network, and the actor network" until the simulation ends.
Citation Information
Cited By
Control-surface-free high-dynamic mass eccentric self-adaptive control method and system
CN121455200A