Reinforcement learning based morphing aircraft adaptive disturbance rejection control method and system

CN122776843APending Publication Date: 2026-09-18HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611123546.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

然而,这些方法严重依赖专家经验,规则库设计复杂,且只能在有限的离散状态间切换,无法实现高维参数空间内的连续、最优整定

Benefits of technology

[0017]The beneficial effects of this invention are as follows: By constructing an adaptive control architecture combining reinforcement learning and LADRC, this invention improves the high dynamic tracking accuracy of deformable aircraft; it overcomes the limitations of traditional LADRC, which relies on manual parameter tuning and fixed gains, and can continuously adjust the control bandwidth parameters of the three channels online based on real-time flight tracking errors and angular velocity states, greatly enhancing the controller's adaptive capability under complex flight conditions (such as structural deformation); through continuously optimized control actions, it avoids the parameter jump problems existing in methods such as fuzzy control, and compared with fixed parameter controllers, it significantly reduces the integral absolute error (IAE) of the attitude channel when facing external disturbances and command steps, improving the transient response speed and steady-state tracking accuracy of the closed-loop system. This invention overcomes the limitations of manual parameter tuning, and can continuously optimize control parameters online to avoid high-frequency chattering, ensuring the absolute stability of the closed-loop system while reducing attitude tracking errors, making it particularly suitable for high dynamic intelligent adaptive control scenarios of deformable aircraft in complex time-varying environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122776843A_ABST
    Figure CN122776843A_ABST
Patent Text Reader

Abstract

The present application relates to a morphing aircraft adaptive disturbance rejection control method and system based on reinforcement learning, the control method comprising: S1: constructing a continuous observation state space of an agent based on deep reinforcement learning; S2: determining a continuous action space of the agent for adaptive disturbance rejection control; S3: designing a reward function for accurate attitude tracking; S4: training the policy network of the agent using a soft action evaluation algorithm; S5: generating an initial control instruction online based on the current real-time state of the morphing aircraft; S6: constraining the initial control instruction based on the closed-loop stability of the variable parameter Lyapunov function to generate a final control instruction, and performing attitude control on the morphing aircraft based on the final control instruction. The control system uses the control method to perform attitude control on the morphing aircraft. The present application can continuously optimize the control parameters online to avoid high-frequency chattering, reduce the attitude tracking error, and ensure the absolute stability of the closed-loop system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an adaptive disturbance rejection control method and system for deformable aircraft based on reinforcement learning, belonging to the field of intelligent control technology for aircraft. Background Technology

[0002] In aircraft attitude control engineering practice, Linear Active Disturbance Rejection Control (LADRC) has been widely used due to its simple parameter tuning and lack of reliance on an accurate controlled object model. However, existing LADRC-based aircraft control systems typically employ fixed parameter designs. Engineers often tune parameters around a few preset nominal flight conditions. These fixed bandwidths and feedback gains are difficult to maintain optimal control performance when encountering rapid changes in the system's dynamic characteristics.

[0003] This problem is particularly prominent for morphing aircraft. Morphing aircraft change their structure in real time through wing folding and span variations, resulting in flight conditions and aerodynamic parameters that are subject to wide-range, high-frequency dynamic changes. Fixed-parameter LADRC often fails to adapt to this rapidly changing dynamic environment, leading to increased tracking errors and even system oscillations. To address the parameter adaptation problem, existing research has proposed some online adjustment schemes, such as parameter switching based on fuzzy logic or gain scheduling based on lookup tables. However, these methods heavily rely on expert experience, have complex rule base designs, and can only switch between a limited number of discrete states, failing to achieve continuous, optimal tuning in high-dimensional parameter spaces.

[0004] In recent years, reinforcement learning (RL), as a model-free optimization technique, has provided a new approach for real-time optimization of controller parameters. However, how to deeply integrate reinforcement learning algorithms with the aircraft disturbance rejection control framework, construct a reasonable observation state and continuous action space to achieve adaptive parameter tuning and ensure flight stability, remains a pressing technical challenge. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, this invention provides an adaptive disturbance rejection control method and system for deformable aircraft based on reinforcement learning. This method can continuously optimize control parameters online to avoid high-frequency chattering and ensure the absolute stability of the closed-loop system while reducing attitude tracking errors.

[0006] The technical solution adopted in this invention is: an adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning, comprising the following steps: S1: Construct a continuous observation state space for an agent based on deep reinforcement learning; S2: Determine the continuous action space of the agent for adaptive disturbance rejection control; S3: Design a reward function for accurate attitude tracking; S4: The agent's policy network is trained using a soft-action evaluation algorithm based on a continuous observation state space, a continuous action space, and a reward function. S5: Embed the trained policy network into the LADRC control loop and generate initial control commands online based on the current real-time state of the deformable aircraft in the continuously observed state space; S6: Constrain the initial control command based on the closed-loop stability of the variable parameter Lyapunov function, generate the final control command, and perform attitude control on the deformable aircraft based on the final control command.

[0007] Preferably, the continuous observation state space in S1 includes the flight attitude angle vector, attitude tracking error vector, and three-axis angular velocity vector.

[0008] Preferably, in S1, a continuous state vector of attitude angle, attitude tracking error and three-axis angular velocity is constructed. The hyperbolic tangent function is used in conjunction with the physical boundary of the deformable aircraft to perform nonlinear scaling and normalization on all state vectors, so that they are softly constrained within a specified range.

[0009] Furthermore, the attitude angle components in the state vector and its normalization process are normalized based on the physical boundary, the attitude tracking error components are based on the set maximum allowable error, and the normalization reference of the three-axis angular velocity components is uniformly set to a fixed value.

[0010] Preferably, the continuous motion space in S2 is the linear active disturbance rejection control feedback bandwidth of the three channels: roll, yaw, and pitch.

[0011] Preferably, the action output of the continuous action space is limited to a compact real number interval.

[0012] Preferably, in S3, the reward function introduces a preset reasonable weight coefficient to strengthen the lateral stability constraint and assigns a penalty weight to the sideslip angle error, so as to guide the strategy network to prioritize the convergence of the sideslip angle in multi-channel coupled excitation, thereby ensuring the aerodynamic symmetry and overall flight safety of the deformable aircraft.

[0013] Preferably, in step S4, an evaluation network and a policy network are constructed using a maximum entropy reinforcement learning framework. The evaluation network is used to map the states in the continuously observed state space and the actions in the continuously observed action space to reward functions. The policy network is used to map the states in the continuously observed state space to actions in the continuously observed action space. The evaluation network is used to train the policy network. The parameters of the evaluation network are updated by minimizing the Bellman residual. The policy network is updated by gradient ascent by maximizing the expected reward while taking into account the action entropy.

[0014] Preferably, in S5, during the control cycle, the policy network propagates forward based on the real-time perceived state of the deformable aircraft, outputting the dynamic bandwidth of three channels. Based on the principle of bandwidth parameterization, it is seamlessly mapped to the dynamic proportional and differential feedback gain of the LADRC control law. The dynamic proportional and differential feedback gain are substituted into the error feedback control law to obtain the initial command equation. The initial anti-interference control output (initial control command) is generated by combining the real-time estimate of the total internal and external disturbances by the extended state observer.

[0015] Preferably, in step S6, in the smoothing filtering stage between the agent's action output and the underlying controller, a hard constraint condition for the variable parameter bandwidth change rate based on the cascaded input state stability theory is established. First, the actual tracking error state vector of the channel is defined, and the error dynamics state space equation is established under time-varying bandwidth drive. Second, a time-varying Lyapunov candidate function containing bandwidth cross terms is constructed, and the absolute safe change rate boundary of the bandwidth adaptive adjustment that satisfies the cascaded input state stability requirement is obtained. The absolute safe change rate boundary is used to constrain the upper limit of the maximum change rate of the bandwidth, and the final control command is generated.

[0016] The reinforcement learning-based adaptive disturbance rejection control system for morphing aircraft uses any of the reinforcement learning-based adaptive disturbance rejection control methods disclosed in this invention to perform attitude control on the morphing aircraft, including: Continuous observation state space construction module: used to construct the continuous observation state space of an agent based on deep reinforcement learning; Continuous Action Space Determination Module: Used to determine the continuous action space of an agent oriented towards adaptive disturbance rejection control; Reward function design module: Used to design reward functions for accurate pose tracking; Policy Network Training Module: Used to train the agent's policy network based on the continuous observation state space, continuous action space, and reward function using a soft action evaluation algorithm; Initial control command generation module: used to embed the trained policy network into the LADRC control loop and generate initial control commands online based on the current real-time state of the deformable aircraft in the continuously observed state space; Final control command generation module: used to constrain the initial control command based on the closed-loop stability of the variable parameter Lyapunov function, and generate the final control command.

[0017] The beneficial effects of this invention are as follows: By constructing an adaptive control architecture combining reinforcement learning and LADRC, this invention improves the high dynamic tracking accuracy of deformable aircraft; it overcomes the limitations of traditional LADRC, which relies on manual parameter tuning and fixed gains, and can continuously adjust the control bandwidth parameters of the three channels online based on real-time flight tracking errors and angular velocity states, greatly enhancing the controller's adaptive capability under complex flight conditions (such as structural deformation); through continuously optimized control actions, it avoids the parameter jump problems existing in methods such as fuzzy control, and compared with fixed parameter controllers, it significantly reduces the integral absolute error (IAE) of the attitude channel when facing external disturbances and command steps, improving the transient response speed and steady-state tracking accuracy of the closed-loop system. This invention overcomes the limitations of manual parameter tuning, and can continuously optimize control parameters online to avoid high-frequency chattering, ensuring the absolute stability of the closed-loop system while reducing attitude tracking errors, making it particularly suitable for high dynamic intelligent adaptive control scenarios of deformable aircraft in complex time-varying environments. Attached Figure Description

[0018] Figure 1 This is a flowchart of the adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to the present invention; Figure 2 This is a diagram of the overall architecture of the adaptive control system based on the combination of reinforcement learning and LADRC of the present invention. Figure 3 This is a reward convergence curve during the training process of the reinforcement learning strategy network of this invention; Figure 4 This is a real-time change curve of the three-channel control bandwidth for online adaptive adjustment of reinforcement learning according to an embodiment of the present invention; Figure 5 These are comparison curves of aircraft attitude tracking under different control algorithms according to embodiments of the present invention; Figure 6 This is a detailed comparison chart of the three-channel attitude tracking errors according to an embodiment of the present invention; Figure 7 This is a comparison chart of the smoothness of the three-axis angular velocity response under adaptive adjustment according to an embodiment of the present invention; Figure 8 This is an attitude tracking adaptability diagram of an aircraft under different deformation angle trajectories according to an embodiment of the present invention. Detailed Implementation

[0019] See Figures 1-3This invention discloses an adaptive disturbance rejection control method for morphing aircraft based on reinforcement learning. First, a normalized continuous observation state space is constructed, including flight attitude angles, attitude tracking errors, and three-axis angular velocities. Next, the active disturbance rejection control feedback bandwidth of the LADRC in the three attitude channels is used as the continuous action output of the agent. Then, using the Soft Actor-Critic (SAC) algorithm framework from deep reinforcement learning, a reward function based on the weighted sum of the three-channel attitude tracking errors is designed, and the policy network and evaluation network are trained in a simulation environment. Finally, the trained policy network is embedded into the LADRC control loop, and the control bandwidth is continuously adjusted online according to the real-time state of the morphing aircraft, achieving adaptive disturbance rejection control of the morphing aircraft in complex time-varying environments. This method utilizes the online continuous output of the optimal control bandwidth by the reinforcement learning agent and combines Lyapunov stability theory to strictly constrain the rate of change of parameters, thereby achieving adaptive disturbance rejection control while ensuring the absolute stability of the closed-loop system.

[0020] The control method specifically includes the following steps: S1: Construct a continuous observation state space for an agent based on deep reinforcement learning.

[0021] To enable the agent to fully perceive the attitude dynamics of the deformable aircraft and the deviations from the control target, and to make optimal decisions in a Markov Decision Process (MDP), At time t, construct a continuous state vector with dimension 9. ,in, express The state at any given moment, This represents the complete nine-dimensional state space of the morphing aircraft. Because the state of the morphing aircraft fluctuates drastically when performing large maneuvers or variant missions, a hyperbolic tangent function is used to prevent gradient explosion during neural network training and to accelerate convergence. To align with the physical boundaries of the deformable aircraft, all state variables are nonlinearly scaled and normalized, softly constraining them to the range [-1, 1]. The state vector and its normalization calculation expression are as follows: ; in, In the state of tilt angle, In the sideslip angle state, At the angle of attack. This is the tilt angle error state. This is the side slip angle error state. This is the angle of attack error state. This refers to the roll channel angular acceleration state. This refers to the yaw channel angular acceleration state. This refers to the pitch channel angular acceleration state. To eliminate the influence of dimensions, attitude angular components are normalized constants based on physical boundaries. For example, the roll angle reference is 40°, the yaw angle is 20°, and the pitch angle is 30°. The calculation method is as follows: etc., among which The actual tilt angle; the tracking error component is based on the set maximum allowable error, for example... etc., among which This is the desired roll angle command in real time; the normalization reference for the three-axis angular velocity components is uniformly set to 200° / s, i.e. , This provides the real-time angular velocity of the roll channel. This soft constraint mechanism not only preserves gradient information under extreme conditions but also avoids data loss caused by direct truncation.

[0022] S2: Determine the continuous action space of the agent for adaptive disturbance rejection control.

[0023] The agent's action vector ( Indicates real-time actions. The control bandwidth of the linear active disturbance rejection control (LADRC) error feedback controller (representing the complete continuous three-dimensional motion space) is defined as follows: ; in, The LADRC controller, representing the roll-through channel, controls the bandwidth. The LADRC controller controls the bandwidth representing the yaw channel. The LADRC controller controls the bandwidth of the pitch channel.

[0024] The advantage of using the bandwidth parameterization method is that it reduces the originally high-dimensional controller parameter optimization to three intuitive physical quantities. Considering the physical deflection rate limitations of the deformable aircraft's control surface actuators (such as servo saturation effects) and the sampling noise limitations of the sensors, the action output domain is limited to... Within a compact real number interval, the bandwidth can be selected as [10, 30] radians / second. Too low a bandwidth will result in insufficient system noise immunity, while too high a bandwidth may amplify high-frequency measurement noise and excite unmodeled higher-order dynamics.

[0025] S3: Design a reward function for precise attitude tracking.

[0026] An instantaneous reward function is designed to minimize the three-channel attitude tracking error, and reasonable weighting coefficients are introduced to strengthen the lateral stability constraint: ; in, For real-time tilt angle error, For real-time sideslip angle error, For real-time angle of attack error, This represents the reward function.

[0027] When deformable aircraft undergo changes in airfoil structure, they are highly susceptible to inducing strong lateral aerodynamic coupling and sideslip divergence. Therefore, the sideslip angle error... Assigning a high penalty weight (weight value of 3) aims to guide the policy network to prioritize the convergence of the sideslip angle during multi-channel coupled maneuvers, ensuring the aerodynamic symmetry and overall flight safety of the aircraft. The overall negative reward mechanism motivates the agent to find the optimal control strategy that minimizes the error during exploration as quickly as possible.

[0028] S4: The agent's policy network is trained using a soft-action evaluation algorithm based on the continuous observation state space, continuous action space, and reward function.

[0029] We employ the maximum entropy reinforcement learning framework (Soft Actor-Critic, SAC) to construct an action-state evaluation network. With random policy networks This framework encourages the agent to explore a wider range of actions by introducing policy entropy into the objective function, effectively avoiding getting trapped in local optima. The evaluation network parameters are updated by minimizing the Bellman residual. : ; in, To evaluate the objective function of the network, These are the neural network weight parameters for the current evaluation network (Critic network). Indicates the state at the current moment. and actions From the experience replay pool Obtained from sampling, This is the current evaluation network, used to output the estimated value (Q-value) of the action to be taken in the current state. This indicates the action performed by the agent at the current moment. Then, the instantaneous reward value from environmental feedback, This is a discount factor used to measure the diminishing weight of future rewards on the current total value. Indicates the action at the next moment. It is based on the current parameters. Policy network (Actor network) obtained from sampling The output value of the Target Critic network is used to estimate the state at the next time step. and actions Target value The parameters of the target evaluation network are represented (usually obtained through a soft-update moving average of the current network parameters to ensure training stability). The adaptive temperature coefficient is used to adjust the proportion of policy entropy in the objective function, which determines the balance between exploration and exploitation in reinforcement learning.

[0030] The policy network updates by gradient ascent while maximizing expected reward and taking into account action entropy.

[0031] The network structure employs a multilayer perceptron (MLP) and breaks down the correlation of time series data through an experience replay buffer, significantly improving sample utilization.

[0032] S5: Embed the trained policy network into the LADRC control loop and generate initial control commands online based on the current real-time state of the deformable aircraft in the continuously observed state space.

[0033] During the control period, the policy network propagates forward based on the real-time perceived state and outputs the dynamic bandwidth of the three channels. Based on the principle of bandwidth parameterization, it is seamlessly mapped to the dynamic proportional and differential feedback gain of the LADRC control law: ; in, It is a dynamic scale; This is the differential feedback gain.

[0034] Will and Substituting the error feedback control law, we obtain the original command equation without disturbance compensation: ; in, and These are the observed values ​​of the system state. Let be the desired state variable.

[0035] Subsequently, the extended state observer (ESO) is used to obtain real-time estimates of the total internal and external disturbances of the system. This generates disturbance rejection control output (initial control command).

[0036] S6: Constrain the initial control command based on the closed-loop stability of the variable parameter Lyapunov function, generate the final control command, and perform attitude control on the deformable aircraft based on the final control command.

[0037] Traditional reinforcement learning methods that directly output control commands often lack strict safety guarantees. To prevent the closed-loop system from becoming unstable due to excessively rapid changes in the bandwidth parameter of the reinforcement learning output, a hard constraint condition for the variable parameter bandwidth change rate based on the Cascaded Input State Stability (ISS) theory is established in the smoothing filtering stage between the agent's action output and the underlying controller.

[0038] First, define the first The actual tracking error state vector of the channel is ,in The error of the i-th channel (i=1, 2, 3, representing roll, sideslip, and pitch channels respectively). Let be the error derivative of the i-th channel. The constructed error vector, in the time-varying bandwidth Under the driving force, the error dynamics state-space equation is expressed as: ; Wherein, the state matrix is , The total perturbation estimation error of the extended state observer, The perturbation transfer matrix, It is the derivative of the error vector.

[0039] Secondly, to strictly guarantee the stability of the variable parameter system, a time-varying Lyapunov candidate function containing bandwidth cross terms is constructed. Its positive definite matrix is ​​designed as follows: ; right Differentiating along time, the stability of an unforced system depends on the matrix Is it strictly positive definite? Through derivation, to make... For this to hold, the following quadratic inequality regarding the rate of change of bandwidth must be satisfied: ; Solving this inequality yields the absolute safety rate of change boundary for bandwidth adaptive adjustment: ; In the motion output smoothing mechanism, a maximum upper limit is set for the bandwidth change rate. To satisfy This generates the final control commands. This constraint ensures, from a rigorous mathematical perspective, that... This demonstrates that the tracking error subsystem under time-varying bandwidth possesses cascaded input state stability (ISS). This means that regardless of how the reinforcement learning agent explores under unknown conditions, all internal signals and tracking errors are uniformly eventually bounded (UUB), completely eliminating the risk of divergence in variable parameter systems.

[0040] Finally, attitude control of the deformable aircraft is performed based on the final control commands.

[0041] This invention also discloses an adaptive disturbance rejection control system for morphing aircraft based on reinforcement learning. The system employs any of the reinforcement learning-based adaptive disturbance rejection control methods disclosed in this invention to perform attitude control on the morphing aircraft, including: Continuous observation state space construction module: used to construct the continuous observation state space of an agent based on deep reinforcement learning; Continuous Action Space Determination Module: Used to determine the continuous action space of an agent oriented towards adaptive disturbance rejection control; Reward function design module: Used to design reward functions for accurate pose tracking; Policy Network Training Module: Used to train the agent's policy network based on the continuous observation state space, continuous action space, and reward function using a soft action evaluation algorithm; Initial control command generation module: used to embed the trained policy network into the LADRC control loop and generate initial control commands online based on the current real-time state of the deformable aircraft in the continuously observed state space; Final control command generation module: used to constrain the initial control command based on the closed-loop stability of the variable parameter Lyapunov function, and generate the final control command.

[0042] Example (verifying the technical effectiveness of the reinforcement learning-based adaptive disturbance rejection control method for deformable aircraft of the present invention): (1) Experimental environment and model deployment: The training and simulation environment of this invention is built based on Python. The numerical integration step size and control sampling period are both set to 0.01 seconds. Each training round corresponds to a real flight time of 2 seconds, and a total of 15,000 training rounds are conducted. Hardware deployment tests show that the neural network's single forward inference time is approximately 1.106 milliseconds, far less than the control period, fully meeting the real-time control requirements of highly dynamic aircraft. During policy execution, the stability boundary derived in S6 (i.e., ...) is strictly followed. This provides physical-level bandwidth limiting protection for the increase in network output bandwidth.

[0043] (2) Experimental results and analysis: The real-time variation curves of the three-channel control bandwidth under online adaptive adjustment of reinforcement learning, the comparison curves of aircraft attitude tracking under different control algorithms, the detailed comparison of three-channel attitude tracking errors, the comparison of the smoothness of the three-axis angular velocity response under adaptive adjustment, and the attitude tracking adaptability of the aircraft under different deformation angle trajectories are respectively as follows: Figure 4 , Figure 5 , Figure 6 , Figure 7 and Figure 8 As shown.

[0044] Figure 4 The three-channel control bandwidth variation curves of the policy network output are shown when facing dynamic reference commands and strong aerodynamic disturbances. The bandwidth of each channel achieves smooth and continuous dynamic adjustment within the range [10, 30]. This is because the underlying algorithm strictly adheres to... Rate of change limit constraint. Figure 7 The three-axis angular velocity comparison curves show that the continuous injection of time-varying parameters did not destroy the negative definiteness of the Lyapunov function of the system, and the closed-loop system maintained excellent transient stability, effectively avoiding the high-frequency oscillation or chattering phenomenon that often occurs in traditional gain scheduling control or reinforcement learning direct control.

[0045] Figure 5 and Figure 6 Attitude tracking error data demonstrates that, compared to manually preset fixed parameters LADRC, the reinforcement learning adaptive method proposed in this invention exhibits a more agile transient response when faced with continuous step commands or complex maneuvers. The system's integral absolute error (IAE) is significantly reduced, with the maximum tracking error successfully suppressed to an extremely small range of 0.2°, and no significant overshoot occurs during the transient transition phase. This perfectly validates the theoretical effectiveness of ISS stability in engineering practice.

[0046] exist Figure 8 In a high-dynamic deformation scenario where the wing folding angle continuously expands from 155° to 0° within 0 to 20 seconds, the aerodynamic focus and moment of inertia of the aircraft undergo drastic changes. Test results show that the method of this invention can still accurately capture and quickly compensate for changes in the aerodynamic moment derivative caused by the large-scale structural changes. Throughout the entire variant process, the attitude response curves of each group almost completely coincide with the reference command, fully demonstrating that the method of this invention has excellent adaptability to a wide flight envelope and extremely high engineering application value.

[0047] Unless otherwise specified or further limited to one preferred or optional technical means being another, the preferred and optional technical means disclosed in this invention can be arbitrarily combined to form several different technical solutions.

Claims

1. A morphing aircraft adaptive disturbance rejection control method based on reinforcement learning, characterized in that Includes the following steps: S1: Construct a continuous observation state space for an agent based on deep reinforcement learning; S2: Determine the continuous action space of the agent for adaptive disturbance rejection control; S3: Design a reward function for accurate attitude tracking; S4: The agent's policy network is trained using a soft-action evaluation algorithm based on a continuous observation state space, a continuous action space, and a reward function. S5: Embed the trained policy network into the LADRC control loop and generate initial control commands online based on the current real-time state of the deformable aircraft in the continuously observed state space; S6: Constrain the initial control command based on the closed-loop stability of the variable parameter Lyapunov function, generate the final control command, and perform attitude control on the deformable aircraft based on the final control command.

2. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 1, characterized in that... The continuous observation state space in S1 includes the flight attitude angle vector, attitude tracking error vector, and three-axis angular velocity vector.

3. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 2, characterized in that... In S1, a continuous state vector of attitude angle, attitude tracking error and three-axis angular velocity is constructed. The hyperbolic tangent function is used in conjunction with the physical boundary of the deformable aircraft to perform nonlinear scaling and normalization on all state vectors, so that they are softly constrained within a specified range.

4. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 1, characterized in that... The continuous motion space in S2 is the linear active disturbance rejection control feedback bandwidth of the three channels: roll, yaw, and pitch.

5. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 4, characterized in that... The action output of the continuous action space is limited to a compact range of real numbers.

6. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 1, characterized in that... In S3, the reward function introduces a preset reasonable weight coefficient to strengthen the lateral stability constraint and assigns a penalty weight to the sideslip angle error, so as to guide the strategy network to prioritize the convergence of the sideslip angle in multi-channel coupled excitation, thereby ensuring the aerodynamic symmetry and overall flight safety of the deformable aircraft.

7. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 1, characterized in that... In S4, an evaluation network and a policy network are constructed using a maximum entropy reinforcement learning framework. The evaluation network maps the states in the continuously observed state space and the actions in the continuously observed action space to reward functions. The policy network maps the states in the continuously observed state space to actions in the continuously observed action space. The evaluation network is used to train the policy network. The parameters of the evaluation network are updated by minimizing the Bellman residual. The policy network is updated by gradient ascent by maximizing the expected reward while taking into account the action entropy.

8. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 1, characterized in that... In S5, during the control cycle, the policy network propagates forward based on the real-time perceived state of the deformable aircraft, outputting the dynamic bandwidth of three channels. Based on the principle of bandwidth parameterization, it is seamlessly mapped to the dynamic proportional and differential feedback gain of the LADRC control law. The dynamic proportional and differential feedback gain are substituted into the error feedback control law to obtain the initial command equation. Combined with the real-time estimate of the total internal and external disturbances by the extended state observer, the initial anti-interference control output is generated.

9. The adaptive disturbance rejection control method for deformable aircraft based on reinforcement learning according to claim 1, characterized in that... In S6, in the smoothing filtering stage between the agent's action output and the underlying controller, a hard constraint condition for the variable parameter bandwidth change rate based on the cascaded input state stability theory is established. First, the actual tracking error state vector of the channel is defined, and the error dynamic state space equation is established under the time-varying bandwidth drive. Second, a time-varying Lyapunov candidate function containing bandwidth cross terms is constructed, and the absolute safe change rate boundary of the bandwidth adaptive adjustment that satisfies the cascaded input state stability requirement is obtained. The absolute safe change rate boundary is used to constrain the upper limit of the maximum change rate of the bandwidth, and the final control command is generated.

10. An adaptive disturbance rejection control system for deformable aircraft based on reinforcement learning, characterized in that... The attitude control of the morphing aircraft is performed using the reinforcement learning-based adaptive disturbance rejection control method for morphing aircraft as described in any one of claims 1-9, including: Continuous observation state space construction module: used to construct the continuous observation state space of an agent based on deep reinforcement learning; Continuous Action Space Determination Module: Used to determine the continuous action space of an agent oriented towards adaptive disturbance rejection control; Reward function design module: Used to design reward functions for accurate pose tracking; Policy Network Training Module: Used to train the agent's policy network based on the continuous observation state space, continuous action space, and reward function using a soft action evaluation algorithm; Initial control command generation module: used to embed the trained policy network into the LADRC control loop and generate initial control commands online based on the current real-time state of the deformable aircraft in the continuously observed state space; Final control command generation module: used to constrain the initial control command based on the closed-loop stability of the variable parameter Lyapunov function, and generate the final control command.