An intelligent decision-making and control integrated method for morphing aircraft based on reinforcement learning

By employing a reinforcement learning-based intelligent decision-making and control method, combined with adaptive dynamic programming and deep deterministic policy gradient algorithm, the problem of the coupling effect between attitude motion and deformation decision of morphing aircraft is solved. This enables real-time deformation and stable attitude tracking of morphing aircraft during the gliding phase, thereby improving flight performance.

CN119105289BActive Publication Date: 2025-11-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411389089.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-11-21
Estimated Expiration
2044-10-08

Smart Images

  • Figure CN119105289B_ABST
    Figure CN119105289B_ABST
Patent Text Reader

Abstract

The application provides a variable-span aircraft intelligent decision control integrated design method based on reinforcement learning. The method designs an attitude control law, formulates a variable shape decision index, and establishes an intelligent variable shape decision model from the perspective of improving the comprehensive flight performance index of the aircraft in the gliding phase, and solves the variable shape decision and coordinated control problems of the variable-span aircraft in the gliding phase. Compared with a fixed shape, the real-time adjustment of the aircraft configuration according to the action strategy given by the intelligent agent can effectively improve the lift-drag ratio, reduce the attitude tracking error, and realize the optimal comprehensive performance index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making and control for variable-span and length aircraft, specifically to an integrated method for intelligent decision-making and control of variable-span and length aircraft based on reinforcement learning. Background Technology

[0002] Deformable aircraft can autonomously change their aerodynamic configuration according to different flight stages and missions, thereby improving their dynamic characteristics and enhancing flight performance. Therefore, compared with fixed-shape aircraft, deformable aircraft have advantages such as adapting to more complex environments, performing more diverse missions, having a larger flight envelope, and greater economic benefits, making them one of the research hotspots in the aerospace field. Among various deformable design schemes, the telescopic wing method provides a large range of wing area and aspect ratio changes with a single degree of freedom, and the deformable mechanism design is simple and reliable, thus it is widely used in the overall design of deformable aircraft.

[0003] Deformation capability, including how to formulate deformation strategies, when to deform, and the degree of deformation, is a crucial factor in achieving optimal global performance for an aircraft. Deformation decisions require information from sensors mounted on the aircraft regarding its own motion state and the flight environment, and are formulated based on performance indicators derived from mission requirements during the gliding phase. The traditional, naive approach to deformation decision-making involves offline calibration during the flight mission, deploying corresponding configurations at different stages. However, this method lacks real-time capability and is ill-suited to unforeseen circumstances. Existing literature includes many numerical optimization algorithms such as PSO and GPM applied to the formulation of deformation strategies for morphing aircraft. However, these methods simplify the aircraft model to some extent, neglecting the coupling effect between deformation decisions and attitude control—that is, ignoring the direct or indirect relationship between flight performance and attitude motion. Summary of the Invention

[0004] To address the problem that existing technologies cannot comprehensively consider decision-making and control effects, and considering that the attitude motion of morphing aircraft is affected by changes in shape, and that the target of morphing decision is related to the attitude state, this invention proposes an integrated intelligent decision-making and control method for morphing aircraft based on reinforcement learning. This method can provide effective morphing strategies and stable attitude tracking control during the gliding phase of the aircraft, thus solving the problem of improving the overall performance of morphing aircraft during the gliding phase.

[0005] The technical solution of this invention is as follows:

[0006] A method for intelligent decision-making and control integration for variable-length aircraft based on reinforcement learning includes the following steps:

[0007] Step 1: For a variable span aircraft, establish an attitude dynamics model, using rudder deflection as the control variable, and design the optimal tracking control law using adaptive dynamic programming.

[0008] Step 1.1: Establish the attitude dynamics model of the variable span and length aircraft as follows:

[0009]

[0010] In the formula, α,β,γ V These are the angle of attack, sideslip angle, and roll angle, respectively; ω x ,ω y ,ω z The three-axis rotational angular velocity; θ represents the track inclination angle; F a =[L,D,N] T These are lift, drag, and lateral force, respectively; M a =[M ax M ay M az ] T For the aerodynamic torque of the three axes; F s =[F sx ,F sy ,F sz ] T M represents the additional deformation force along the three axes; s =[M sx M sy M sz ] T Add a torque to the deformation of the three axes; J x J y J z denoted by , where is the principal moment of inertia of the three axes; m is the total mass of the aircraft; V is the flight velocity; and g is the acceleration due to gravity.

[0011] Step 1.2: Rewrite the attitude dynamics model of the variable span and length aircraft into an attitude control system model:

[0012]

[0013] In the formula, the state variables are chosen as x1=[α,β,γ] V ] T x2=[ω x ,ω y ,ω z ] T Control input u = [δ x ,δ y ,δ z ] T ;

[0014] Step 1.3: Divide the attitude control system model into an angle slow loop and an angular velocity fast loop. First, design a virtual controller for the slow loop and use it as the desired tracking trajectory for the angular velocity fast loop. Then, design the controller for the fast loop.

[0015] Slow loop controller design:

[0016] Assume the angle subsystem satisfies the following desired trajectory:

[0017]

[0018] Where x 1c Let x1 be the expected value of the state variable. The steady-state portion of the virtual control input. This is an estimate of the combined disturbance d1(x1,x2,ξ);

[0019] Obtain the steady-state part of the virtual control input as follows:

[0020]

[0021] Define the optimal feedback part of the slow-loop virtual controller as follows: x 2d (t) is the virtual control input of the angle subsystem, and the tracking error is e1(t) = x1(t) - x 1c (t), then:

[0022]

[0023] Among them, f1 e (e1)=f1(x1)-f1(x 1c );

[0024] Fast-loop controller design:

[0025] Assume the desired trajectory of the angular velocity subsystem satisfies the following equation:

[0026]

[0027] Where u d This refers to the steady-state portion of the actual control input. This is an estimate of the combined disturbance d2(x1,x2,ξ);

[0028] The steady-state control input for rudder deflection is as follows:

[0029]

[0030] Define the optimal feedback part of the fast-loop controller as u e (t)=u(t)-u d (t), the tracking error is e2(t)=x2(t)-x 2d The tracking error equation for the fast loop (t) can be written as:

[0031]

[0032] in,

[0033] Step 1.4: Perform approximate optimal control based on adaptive dynamic programming:

[0034] The tracking error equations for the slow loop and the fast loop are unified as follows:

[0035]

[0036] The optimal cost function is expressed as follows:

[0037]

[0038] In the formula, Q, It is a positive definite symmetric matrix;

[0039] According to optimal control theory, by solving the HJB equation

[0040]

[0041] Obtain the optimal control input u * ;

[0042] Step 2: Based on the mission requirements of the variable span and length aircraft during the gliding phase, and addressing the deformation decision problem of the deformable aircraft, performance indicators are set for the comprehensive lift-to-drag ratio and attitude tracking error. A deep deterministic policy gradient algorithm is used to realize the intelligent deformation decision of the variable span and length aircraft. The Actor-Critic framework in deep reinforcement learning is used to adjust the deformation rate in real time according to the flight status and flight mission, thereby realizing the integration of decision-making and control.

[0043] Furthermore, the variable span aircraft is a variable span aircraft with symmetrical and continuous extension and retraction on both sides of the wing.

[0044] Furthermore, in the attitude dynamics model of the variable span-length aircraft in step 1.1, the calculation forms of the forces and moments acting on the aircraft are as follows:

[0045] (1) Aerodynamic F a

[0046] The lift L, drag D, and lateral force N are calculated as follows:

[0047]

[0048] In the formula, q is the dynamic pressure; S is the reference area of ​​the wing when it is not deformed; The equivalent lift coefficient, This is the equivalent drag coefficient. This is the equivalent lateral force coefficient;

[0049] (2) Deformation additional force F s

[0050] The analytical expression for the additional force generated by deformation in body coordinates is:

[0051]

[0052] In the formula, ω=[ω x ,ω y ,ω z ] T m i Indicates the mass of the left or right wing of an aircraft; v si The velocity vector of the deformable wing in the body coordinate system; s i This indicates the coordinates of the center of mass of the left or right wing of the aircraft in the body coordinate system.

[0053] (3) Aerodynamic torque M a

[0054] The calculation methods for roll moment, yaw moment, and pitch moment are as follows:

[0055]

[0056] In the formula, b and c are the reference lengths of the aircraft in the lateral and longitudinal directions, respectively. This is the equivalent rolling moment coefficient. This is the equivalent yaw moment coefficient. This is the equivalent yaw moment coefficient;

[0057] (4) Deformation additional torque M s

[0058] The deformation-induced additional moment characterizes the effect of the relative motion of the wing on the fuselage, and its analytical expression is as follows:

[0059]

[0060] Where v f This is the velocity vector of the aircraft.

[0061] Furthermore, the equivalent aerodynamic coefficients and equivalent aerodynamic moment coefficients are expressed as polynomial functions of the deformation rate ξ and the flight state, as shown below:

[0062]

[0063] Where δ x ,δ y ,δ z The rudder deflection angles in the three axes are represented by ξ; Ma is the Mach number; S(ξ) is a function reflecting the change in the reference area; C LC is the lift coefficient. D C is the drag coefficient. N m is the lateral force coefficient. x m is the rolling moment coefficient. y The yaw moment coefficient is m. z This is the yaw moment coefficient.

[0064] Furthermore, in step 1.3, This is obtained through the following nonlinear disturbance observer:

[0065]

[0066] This is obtained through the following nonlinear disturbance observer:

[0067]

[0068] Where p 11 ,p 12 ,p 21 ,p 22 L is an auxiliary variable for the observer. 11 ,L 12 ,L 21 ,L 22 Design parameters for the interference observer; L ij =diag(l ijk ),l ijk >0, where k = 1, 2, 3.

[0069] Furthermore, in step 1.4, an approximate optimal cost function for the evaluation network is adopted, consisting of a single-layer neural network.

[0070]

[0071] Where W represents the ideal weights for evaluating the network; Φ(e) represents the activation function to be designed; ε(e) represents the approximation error of the neural network; and the network is approximated using the following form:

[0072]

[0073] In the formula This represents an approximation of the ideal weights for evaluating the network;

[0074] The final control input is obtained based on the approximate results of the evaluation network:

[0075]

[0076] Furthermore, the evaluation network is trained with the goal of minimizing the loss function. According to the gradient descent method, the weights of the evaluation network are updated according to the following formula.

[0077]

[0078] In the formula, α c ,α s >0 represents the learning rate; And it is normalized. J s (e) is a Lyapunov function; The definition is as follows:

[0079]

[0080] Furthermore, in step 2, the performance indicators for the combined lift-to-drag ratio and attitude tracking error are as follows:

[0081]

[0082] In the formula, γ∈[0,1] represents the discount factor; e λ =λ-λ c The lift-to-drag ratio λ is the maximum lift-to-drag ratio λ that the aircraft can achieve. c The error; e α =α-α c For angle-of-attack tracking error; R t (·) represents the reward function.

[0083] Furthermore, in step 2, the gradient algorithm based on deep deterministic policy includes:

[0084] The state data in the observation space are:

[0085] S t =[s t-4 ,s t-3 ,s t-2 ,s t-1 ,s t ] T

[0086] Where s t =[e λ ,e α ,ω z ] T ;

[0087] The action space employs position-based commands to determine the desired deformation rate ξ of the variable-length aircraft. c As an action output;

[0088] The reward function takes the following form:

[0089]

[0090] in:

[0091]

[0092] In the formula, These are the weights of each term in the reward function.

[0093] Beneficial effects

[0094] This invention proposes an integrated intelligent decision-making and control design method for variable-span aircraft based on reinforcement learning. It designs attitude control laws, formulates deformation decision indicators, and establishes an intelligent deformation decision model from the perspective of improving the overall flight performance indicators during the gliding phase, thus solving the deformation decision and coordinated control problems of variable-span aircraft during the gliding phase. Compared to a fixed shape, this invention effectively improves the lift-to-drag ratio and reduces attitude tracking errors by adjusting the aircraft configuration in real time according to the action strategy given by the intelligent agent, achieving optimal overall performance indicators.

[0095] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0096] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0097] Figure 1 The flowchart for the integrated decision-making and control method for variable-length aircraft of this invention is shown below.

[0098] Figure 2 This is a schematic diagram of the variable span and length aircraft of the present invention.

[0099] Figure 3 Configure the structure of the Actor network;

[0100] Figure 4 Configure the structure of the Critic network;

[0101] Figure 5 This is a training result diagram of the intelligent deformation decision-making and control integration of the present invention;

[0102] Figure 6 This is a diagram showing the intelligent deformation decision output results of the present invention;

[0103] Figure 7 This is a graph showing the lift-to-drag ratio variation of the variable-span aircraft of this invention.

[0104] Figure 8 This is a diagram showing the change in attitude tracking error during the deformation decision-making stage of this invention.

[0105] Figure 9This is a diagram showing the rudder deflection change results during the deformation decision-making stage of this invention. Detailed Implementation

[0106] The embodiments of the present invention are described in detail below. These embodiments are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0107] This embodiment provides an integrated decision-making and control method for variable-span aircraft based on reinforcement learning, and provides a reference for the decision-making and control design process of variable-span aircraft. Figure 1 As shown, the integrated decision-making and control method for variable span length aircraft based on reinforcement learning includes:

[0108] Step 1: Facing such Figure 2 For the variable span aircraft shown, an attitude dynamics model is established, and the optimal tracking control law is designed using rudder deflection as the control variable and combined with adaptive dynamic programming.

[0109] Step 1.1: Based on the multi-rigid-body dynamics modeling method and the Newton-Euler method, the attitude dynamics model of a variable-span aircraft with symmetrical and continuously extending wings can be expressed as follows:

[0110]

[0111] In the formula, α,β,γ V These are the angle of attack, sideslip angle, and roll angle, respectively; ω x ,ω y ,ω z The three-axis rotational angular velocity; θ represents the track inclination angle; F a =[L,D,N] T These are lift, drag, and lateral force, respectively; M a =[M ax M ay M az ] T For the aerodynamic torque of the three axes; F s =[F sx ,F sy ,F sz ] T M represents the additional deformation force along the three axes; s =[M sx M sy M sz ] T Add a torque to the deformation of the three axes; J x J y J z denoted as , where is the principal rotational inertia of the three axes; m is the total mass of the aircraft; V is the flight speed; and g is the acceleration due to gravity.

[0112] The calculation methods for the forces and moments acting on the aircraft are as follows:

[0113] (1) Aerodynamic F a calculate

[0114] The lift L, drag D, and lateral force N are calculated as follows:

[0115]

[0116] In the formula, q represents dynamic pressure; S is the reference area of ​​the wing when it is not deformed. For the sake of convenience in the equation representation, the effects of deformation are uniformly considered in the force coefficients, resulting in the following expression for the equivalent aerodynamic force coefficients:

[0117]

[0118] In the formula, δ x ,δ y ,δ z Indicates the rudder deflection angle in the three axes; Ma is the Mach number; The equivalent lift coefficient, This is the equivalent drag coefficient. The equivalent lateral force coefficient; the change in reference area is represented by S(ξ), a function of the deformation rate ξ, which is defined as the deformation state quantity, i.e., the normalized parameter of the wing span, namely:

[0119]

[0120] In the formula, b represents the wingspan, and the subscripts represent the maximum and minimum values ​​of the wingspan.

[0121] (2) Deformation additional force F s calculate

[0122] The analytical expression for the additional force generated by deformation in body coordinates is:

[0123]

[0124] In the formula, ω=[ω x ,ω y ,ω z ] T m i Indicates the mass of the left or right wing of an aircraft; v si The velocity vector of the deformable wing in the body coordinate system; s i This indicates the coordinates of the center of mass of the left or right wing of the aircraft in the body coordinate system.

[0125] (3) Aerodynamic torque M a calculate

[0126] Similar to aerodynamic calculations, the calculation methods for roll moment, yaw moment, and pitch moment are as follows:

[0127]

[0128] In the formula, b and c are the reference lengths of the aircraft in the lateral and longitudinal directions, respectively. This is the equivalent rolling moment coefficient. This is the equivalent yaw moment coefficient. This is the equivalent yaw moment coefficient.

[0129] The equivalent aerodynamic moment coefficient is expressed as a polynomial function of the deformation rate and flight state, as shown below:

[0130]

[0131] (4) Deformation additional torque M s calculate

[0132] The deformation-induced additional moment characterizes the effect of the relative motion of the wing on the fuselage, and its analytical expression is as follows:

[0133]

[0134] Where v f This is the velocity vector of the aircraft.

[0135] Step 1.2: Rewrite the attitude dynamics model of the variable span and length aircraft as an attitude control system model in the following form:

[0136]

[0137] In the formula, the state variables are chosen as x1=[α,β,γ] V ] T x2=[ω x ,ω y ,ω z ] T Control input u = [δ x ,δ y ,δ z ] T .

[0138] Step 1.3: The attitude control system shown in formula (9) can be divided into an angle slow loop and an angular velocity fast loop. Based on the backstepping method, a virtual controller is first designed for the slow loop and used as the expected tracking trajectory of the angular velocity fast loop. Then, the controller of the fast loop is designed.

[0139] Slow loop controller design:

[0140] Assume the angle subsystem satisfies the following desired trajectory:

[0141]

[0142] Where x 1c Let x1 be the expected value of the state variable. The steady-state portion of the virtual control input. The estimated value of the combined disturbance d1(x1,x2,ξ) is obtained through the following nonlinear disturbance observer:

[0143]

[0144] Where p 11 ,p 12 L is an auxiliary variable for the observer. 11 ,L 12 For the design parameters of the interference observer, L ij =diag(l ijk ),l ijk >0, where k = 1, 2, 3.

[0145] Therefore, the steady-state part of the virtual control input can be obtained. as follows:

[0146]

[0147] Define the optimal feedback part of the slow-loop virtual controller as follows: x 2d (t) is the virtual control input of the angle subsystem, and the tracking error is e1(t) = x1(t) - x 1c (t), then:

[0148]

[0149] Among them, f1 e (e1)=f1(x1)-f1(x 1c The subscripts d and e represent the optimal feedback part.

[0150] Fast-loop controller design:

[0151] Assume the desired trajectory of the angular velocity subsystem satisfies the following equation:

[0152]

[0153] Where u d This refers to the steady-state portion of the actual control input. The estimated value of the combined disturbance d2(x1,x2,ξ) is obtained through the following nonlinear disturbance observer:

[0154]

[0155] Where p 21 ,p22 L is an auxiliary variable for the observer. 21 ,L 22 For the design parameters of the interference observer,

[0156] L ij =diag(l ijk ),l ijk >0, where k = 1, 2, 3.

[0157] Therefore, the steady-state control input for rudder deflection can be obtained as follows:

[0158]

[0159] Define the optimal feedback part of the fast-loop controller as u e (t)=u(t)-u d (t), the tracking error is e2(t)=x2(t)-x 2d The tracking error equation for the fast loop (t) can be written as:

[0160]

[0161] in,

[0162] Step 1.4: Approximate optimal control based on adaptive dynamic programming:

[0163] The error systems shown in formulas (13) and (17) can be uniformly expressed as:

[0164]

[0165] The optimal control input ensures the stability of the error system (18) while minimizing the cost function V(e,u). The optimal cost function is expressed as follows:

[0166]

[0167] In the formula, Q, It is a positive definite symmetric matrix.

[0168] According to optimal control theory, the optimal control input u * This can be obtained by solving the following HJB equation:

[0169]

[0170] The HJB equation is a nonlinear differential equation, which is usually difficult to solve analytically. In this embodiment, an evaluation network composed of a single-layer neural network is used to approximate V. * (e):

[0171]

[0172] in, This represents the ideal weights for evaluating the network; Let represent the activation function to be designed; l is the number of neurons; ε(e) represents the approximation error of the neural network.

[0173] The ideal weights of the neural network are unknown. The network is approximated using the following method:

[0174]

[0175] In the formula, This represents an approximation of the ideal weights for evaluating the network.

[0176] The final control input can be obtained based on the approximate results of the evaluation network:

[0177]

[0178] Therefore, the approximate value of the Hamiltonian function is:

[0179]

[0180] The training objective of evaluating a network is to minimize the loss function. According to the gradient descent method, the weights of the evaluation network are updated according to the following formula:

[0181]

[0182] In the formula, α c ,α s >0 represents the learning rate; And it is normalized. J s (e) represents the Lyapunov function to be designed; The definition is as follows:

[0183]

[0184] The stability analysis is given below:

[0185] Lemma: There exists a positive definite matrix and This makes the following equation true:

[0186]

[0187] Define the network weight estimation error as The Lyapunov function is chosen as follows:

[0188]

[0189] but:

[0190]

[0191] in, B=g(x)R -1 g T (x); ε H The residual caused by the neural network approximation is shown in (29).

[0192]

[0193] Considering (28) can be written as the following formula:

[0194]

[0195] Assume a m ≤||A||≤a M b m ≤||B||≤b M , ||ε(e)||≤E M , ||Φ(e)||≤Φ M ,||f(e)+g(x)u * ||≤δ(e), and the following inequality holds:

[0196]

[0197] Based on (31) and (32), we can obtain:

[0198]

[0199] in,

[0200] Scenario 1: When Then,

[0201] Suppose there exists K > 0 such that And combined We can obtain:

[0202]

[0203] Therefore, when When it is established and (35) or (36) is established,

[0204]

[0205] Scenario 2: When Then,

[0206] Based on (33) and (26), we can obtain:

[0207]

[0208] Therefore, when the following inequality holds,

[0209]

[0210] According to Lyapunov theory, attitude tracking error and weight estimation error are eventually bounded.

[0211] Step 2: Based on the mission requirements of the variable span and length aircraft during the gliding phase, and addressing the deformation decision problem of the deformable aircraft, an intelligent deformation decision-making system based on a deep deterministic policy gradient algorithm is adopted to realize the deformation decision of the variable span and length aircraft. The Actor-Critic framework in deep reinforcement learning is used to adjust the deformation rate in real time according to the flight status and flight mission, thereby realizing the integration of decision-making and control.

[0212] Lift-to-drag ratio is a key flight performance indicator for aircraft during the gliding phase. Generally, a higher lift-to-drag ratio aerodynamic shape can achieve a longer range, thereby increasing the aircraft's mission flexibility. Therefore, appropriate wing deformation can ensure the highest possible lift-to-drag ratio while the aircraft is tracking commands. Traditional deformation decision-making methods typically only consider the lift-to-drag ratio as a single indicator, neglecting the impact of deformation on attitude tracking accuracy. Therefore, the goal of this deformation decision-making step is to maximize the lift-to-drag ratio while minimizing attitude tracking error. This embodiment selects a performance indicator that integrates lift-to-drag ratio and attitude tracking error, and focuses more on performance over a specific flight time period than on performance at a single moment.

[0213]

[0214] In the formula, γ∈[0,1] represents the discount factor; e λ =λ-λ c The lift-to-drag ratio λ is the maximum lift-to-drag ratio λ that the aircraft can achieve. c The error; e α =α-α c For angle-of-attack tracking error; R t (·) represents the reward function.

[0215] Based on the deformation decision objective shown in formula (40), the intelligent deformation decision design of the variable span and length aircraft is carried out by the deep deterministic policy gradient algorithm:

[0216] (1) Observation space

[0217] The wing deformation decision aims to obtain the optimal lift-to-drag ratio while minimizing attitude tracking error. Furthermore, as a continuous system, the aircraft system has a certain degree of inertia. Including observations of error changes or state changes can help improve the fitting effect of the neural network on the actual situation. Therefore, the state variables in the observation space should include lift-to-drag ratio error, attitude tracking error, and three-axis rotational angular velocity.

[0218] In conjunction with the decision-making task objectives, the observation space contains the following state information:

[0219] s t =[e λ ,e α ,ω z ] T (41)

[0220] Considering the inertia of the spacecraft system, further, state data from five historical cycles are introduced into the observation space:

[0221] S t =[s t-4 ,s t-3 ,s t-2 ,s t-1 ,s t ] T (42)

[0222] When training a neural network, in order to avoid problems such as different magnitudes of different features and gradient explosion, the above state variables are normalized to [0,1].

[0223] (2) Action space

[0224] The agent's action output is the wing deformation command for the hypersonic deformable aircraft. In this embodiment, the action space adopts the form of position-based commands, i.e., the desired deformation rate of the variable span aircraft:

[0225] a t =ξ c (43)

[0226] (3) Reward function

[0227] The reward function takes the following form:

[0228]

[0229] in:

[0230]

[0231] In the formula, These are the weights of each term in the reward function. The first and second terms on the right-hand side of the reward function represent the penalties for angle-of-attack tracking error and lift-to-drag ratio error, respectively. Since these errors vary within different ranges during deformation, to facilitate weight adjustment and ensure the Q-function remains within a reasonable range to improve the fitting effect of the critic network, therefore... The first term represents the absolute value of the error normalized to [0,1]. The third term is used to reward situations where the agent tends to reduce the angle-of-attack tracking error when it takes action according to the policy, so as to incentivize the agent to make decisions in the direction of stable attitude tracking.

[0232] The execution flow of the deep deterministic policy gradient algorithm is shown in Table 1:

[0233] Table 1 Execution flow of the depth deterministic strategy gradient algorithm

[0234]

[0235]

[0236] (4) Neural network structure design

[0237] Both the Actor and Critic networks employ multi-hidden-layer backpropagation feedforward neural networks, with the following structural configuration: Figure 3 and Figure 4 As shown.

[0238] The Actor network's input layer consists of 15 neurons corresponding to 15-dimensional environmental input; it passes through three fully connected hidden layers, each with 64 neurons and ReLU activation function; the output layer's single neuron corresponds to the 1-dimensional agent's action, i.e., the wing deformation command, and its activation function is set to tanh type. Adding a bias ensures that the output action is within a set range, which helps the training converge quickly.

[0239] The Critic network's input layer has 26 neurons, including 15-dimensional environmental input and 1-dimensional action. The observation state is passed through two fully connected layers of 64 neurons each, and then summed with the action input through one fully connected layer of 64 neurons each, with ReLU activation function for both. After passing through another fully connected layer of 64 neurons, the output layer is a fully connected layer that outputs a 1-dimensional action cost function corresponding to the initial state and policy, i.e., the Q function.

[0240] For the integrated decision-making and control method in this embodiment, the effectiveness of the agent strategy is verified using MATLAB through the following simulation scenario.

[0241] The simulation scenario is set as follows: the period from 0 to 5 seconds is the attitude stabilization phase of the aircraft, and intelligent deformation decision-making is added starting from the 5th second.

[0242] Figure 5 This represents the change in the average reward curve during the training of a gradient agent with a deep deterministic policy.

[0243] Figure 6 The deformation rate change results of the variable span length aircraft are presented. Within 0 to 5 seconds, the attitude control tracking process of the variable span length aircraft during the change of wingspan is verified by a predetermined deformation command. After 5 seconds, the deformation rate is output online by the intelligent deformation decision.

[0244] Figure 7 This indicates the change in the lift-to-drag ratio of an aircraft; intelligent deformation decision-making can effectively improve the lift-to-drag ratio.

[0245] Figure 8 The results of attitude tracking error comparison in the intelligent deformation decision stage are presented. Compared with the basic decision command (uniform deformation), the design of comprehensive performance indicators effectively improves the attitude tracking accuracy of the aircraft.

[0246] Figure 9 This indicates the rudder deflection change result during the intelligent deformation decision-making stage.

[0247] Based on the above technical solutions, the reinforcement learning-based intelligent decision-making and control integrated method for variable span aircraft provided by this invention improves the overall performance of the aircraft during the gliding phase by designing attitude control laws and deformation decision-making methods.

[0248] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. A method for integrated intelligent decision-making and control of a variable-length aircraft based on reinforcement learning, characterized in that: Includes the following steps: Step 1: For a variable span aircraft, establish an attitude dynamics model, using rudder deflection as the control variable, and design the optimal tracking control law using adaptive dynamic programming. Step 1.1: Establish the attitude dynamics model of the variable span and length aircraft as follows: In the formula, α,β,γ V These are the angle of attack, sideslip angle, and roll angle, respectively; ω x ,ω y ,ω z The three-axis rotational angular velocity; θ represents the track inclination angle; F a =[L,D,N] T These are lift, drag, and lateral force, respectively; M a =[M ax M ay M az ] T For the aerodynamic torque of the three axes; F s =[F sx ,F sy ,F sz ] T M represents the additional deformation force along the three axes; s =[M sx M sy M sz ] T Add a torque to the deformation of the three axes; J x J y J z denoted by , where is the principal moment of inertia of the three axes; m is the total mass of the aircraft; V is the flight velocity; and g is the acceleration due to gravity. Step 1.2: Rewrite the attitude dynamics model of the variable span and length aircraft into an attitude control system model: In the formula, the state variables are chosen as x1=[α,β,γ] V ] T x2=[ω x ,ω y ,ω z ] T Control input u = [δ x ,δ y ,δ z ] T ; Step 1.3: Divide the attitude control system model into an angle slow loop and an angular velocity fast loop. First, design a virtual controller for the slow loop and use it as the desired tracking trajectory for the angular velocity fast loop. Then, design the controller for the fast loop. Slow loop controller design: Assume the angle subsystem satisfies the following desired trajectory: Where x 1c Let x1 be the expected value of the state variable. The steady-state portion of the virtual control input. This is an estimate of the combined disturbance d1(x1,x2,ξ); Obtain the steady-state part of the virtual control input as follows: Define the optimal feedback part of the slow-loop virtual controller as follows: x 2d (t) is the virtual control input of the angle subsystem, and the tracking error is e1(t) = x1(t) - x 1c (t), then: in, Fast-loop controller design: Assume the desired trajectory of the angular velocity subsystem satisfies the following equation: Where u d This refers to the steady-state portion of the actual control input. This is an estimate of the combined disturbance d2(x1,x2,ξ); The steady-state control input for rudder deflection is as follows: Define the optimal feedback part of the fast-loop controller as u e (t)=u(t)-u d (t), the tracking error is e2(t)=x2(t)-x 2d The tracking error equation for the fast loop is written as (t). in, Step 1.4: Perform approximate optimal control based on adaptive dynamic programming: The tracking error equations for the slow loop and the fast loop are unified as follows: The optimal cost function is expressed as follows: In the formula, It is a positive definite symmetric matrix; According to optimal control theory, by solving the HJB equation Obtain the optimal control input u * ; Step 2: Based on the mission requirements of the variable span and length aircraft during the gliding phase, and addressing the deformation decision problem of the deformable aircraft, performance indicators are set for the comprehensive lift-to-drag ratio and attitude tracking error. A deep deterministic policy gradient algorithm is used to realize the intelligent deformation decision of the variable span and length aircraft. The Actor-Critic framework in deep reinforcement learning is used to adjust the deformation rate in real time according to the flight status and flight mission, thereby realizing the integration of decision-making and control.

2. The integrated intelligent decision-making and control method for variable-span aircraft based on reinforcement learning according to claim 1, characterized in that: The variable span aircraft is a variable span aircraft with symmetrical and continuous extension and retraction on both sides of the wing.

3. The integrated intelligent decision-making and control method for variable-span aircraft based on reinforcement learning according to claim 1, characterized in that: In the attitude dynamics model of the variable span-length aircraft in step 1.1, the calculation forms of the forces and moments acting on the aircraft are as follows: (1) Aerodynamic F a The lift L, drag D, and lateral force N are calculated as follows: In the formula, q is the dynamic pressure; S is the reference area of ​​the wing when it is not deformed; The equivalent lift coefficient, This is the equivalent drag coefficient. This is the equivalent lateral force coefficient; (2) Deformation additional force F s The analytical expression for the additional force generated by deformation in body coordinates is: In the formula, ω=[ω x ,ω y ,ω z ] T ;m i Indicates the mass of the left or right wing of an aircraft; v si The velocity vector of the deformable wing in the body coordinate system; s i This indicates the coordinates of the center of mass of the left or right wing of the aircraft in the body coordinate system. (3) Aerodynamic torque M a The calculation methods for roll moment, yaw moment, and pitch moment are as follows: In the formula, b and c are the reference lengths of the aircraft in the lateral and longitudinal directions, respectively. This is the equivalent rolling moment coefficient. This is the equivalent yaw moment coefficient. This is the equivalent yaw moment coefficient; (4) Deformation additional torque M s The deformation-induced additional moment characterizes the effect of the relative motion of the wing on the fuselage, and its analytical expression is as follows: Where v f This is the velocity vector of the aircraft.

4. The integrated intelligent decision-making and control method for a variable-span aircraft based on reinforcement learning according to claim 3, characterized in that: The equivalent aerodynamic force coefficient and the equivalent aerodynamic moment coefficient are expressed as polynomial functions of the deformation rate ξ and the flight state, as shown below: Where δ x ,δ y ,δ z The rudder deflection angles in the three axes are represented; Ma is the Mach number; S(ξ) is a function reflecting the change in the reference area; C L C is the lift coefficient. D C is the drag coefficient. N m is the lateral force coefficient. x m is the rolling moment coefficient. y The yaw moment coefficient is m. z This is the yaw moment coefficient.

5. The integrated intelligent decision-making and control method for variable-span aircraft based on reinforcement learning according to claim 1, characterized in that: In step 1.3, This is obtained through the following nonlinear disturbance observer: This is obtained through the following nonlinear disturbance observer: Where p 11 ,p 12 ,p 21 ,p 22 L is an auxiliary variable for the observer. 11 ,L 12 ,L 21 ,L 22 Design parameters for the interference observer; L ij =diag(l ijk ),l ijk >0, where k = 1, 2, 3.

6. The integrated intelligent decision-making and control method for a variable-span aircraft based on reinforcement learning according to claim 1, characterized in that: In step 1.4, an approximate optimal cost function is used to evaluate the network, which is composed of a single-layer neural network. Where W represents the ideal weights for evaluating the network; Φ(e) represents the activation function to be designed; ε(e) represents the approximation error of the neural network; and the network is approximated using the following form: In the formula This represents an approximation of the ideal weights for evaluating the network; The final control input is obtained based on the approximate results of the evaluation network:

7. The integrated intelligent decision-making and control method for a variable-span aircraft based on reinforcement learning according to claim 6, characterized in that: The evaluation network is trained with the goal of minimizing the loss function. According to the gradient descent method, the weights of the evaluation network are updated according to the following formula. In the formula, α c ,α s >0 represents the learning rate; And it is normalized. J s (e) is a Lyapunov function; The definition is as follows:

8. The integrated intelligent decision-making and control method for a variable-span aircraft based on reinforcement learning according to claim 1, characterized in that: In step 2, the performance indicators for the combined lift-to-drag ratio and attitude tracking error are: In the formula, γ∈[0,1] represents the discount factor; e λ =λ-λ c The lift-to-drag ratio λ is the maximum lift-to-drag ratio λ that the aircraft can achieve. c The error; e α =α-α c For angle of attack tracking error; R t (·) represents the reward function.

9. The integrated intelligent decision-making and control method for variable-span aircraft based on reinforcement learning according to claim 1, characterized in that: In step 2, the gradient algorithm based on deep deterministic policy is as follows: The state data in the observation space are: S t =[s t-4 ,s t-3 ,s t-2 ,s t-1 ,s t ] T Where s t =[e λ ,e α ,ω z ] T ; The action space employs position-based commands to determine the desired deformation rate ξ of the variable-length aircraft. c As an action output; The reward function takes the following form: in: In the formula, These are the weights of each term in the reward function.

Citation Information

Patent Citations

  • Intelligent control method for knowledge and data hybrid-driven hypersonic-speed deformation aircraft

    CN118732589A

  • Tandem Rotor Unmanned Aerial Vehicle and Attitude Adjustment Control Method

    US20230312143A1