An ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints

By constructing a restricted Markov decision process to optimize ADRC parameters, the problem of controller performance degradation during the deformation process of variable sweep aircraft was solved, and efficient flight control under safety constraints was achieved.

CN122308365APending Publication Date: 2026-06-30BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2026-03-30
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Traditional parameter scheduling methods cannot meet the real-time control requirements under variable sweep flight conditions. Fixed parameter ADRC controllers are difficult to adapt to the dynamic characteristics changes of variable sweep aircraft during multi-stage flight, resulting in a decline in control quality.

Method used

By constructing a restricted Markov decision process, the desired velocity, desired altitude, and current flight state of the variable-sweep aircraft are used, combined with safety constraint information and deformable configuration perception information, to optimize the ADRC parameters. The ADRC parameters are autonomously learned and optimized by using a preset reward function and cost function, as well as a pre-trained policy network, reward evaluation network, and cost evaluation network.

Benefits of technology

It alleviates the performance degradation of the fixed-parameter ADRC controller caused by drastic changes in aerodynamic parameters during the deformation process of variable-sweep aircraft, ensuring that the controller can achieve efficient and stable flight control while meeting safety constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122308365A_ABST
    Figure CN122308365A_ABST
Patent Text Reader

Abstract

This invention provides an ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints. Based on the desired velocity, desired altitude, and current flight state of the variable-sweep aircraft, the current deviation information and current control variables of the variable-sweep aircraft are determined through an ADRC simulation model and a longitudinal motion model. A current state vector is constructed based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft. Based on the current state vector and current control variables, the ADRC parameters of the ADRC simulation model are optimized through a preset reward function, a preset cost function, and pre-trained policy network, reward evaluation network, and cost evaluation network. This invention can alleviate the performance degradation problem of fixed-parameter ADRC controllers caused by drastic changes in aerodynamic parameters during the deformation process of variable-sweep aircraft.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aircraft control technology, and in particular to an ADRC parameter optimization method for variable sweep aircraft that takes into account deformation safety constraints. Background Technology

[0002] Variable-sweep aircraft achieve cross-speed range flight capability by dynamically adjusting the wing sweep angle. Their core advantage lies in balancing high-speed penetration and low-speed maneuverability, making them valuable for military reconnaissance, rapid strikes, and flexible deployment in complex battlefield environments. The autopilot, as the core guarantee of flight stability, must cope with nonlinear aerodynamic disturbances caused by continuous wing deformation to maintain attitude stability.

[0003] Active Disturbance Rejection Control (ADRC) demonstrates robustness in highly uncertain scenarios by estimating and compensating for internal and external disturbances to the aircraft in real time. This technology does not rely on precise mathematical models and can effectively handle the coupling effect of time-varying aerodynamic parameters and external disturbances. However, traditional parameter scheduling methods face challenges under variable sweep flight conditions. Rapid changes in aerodynamic characteristics lead to a mismatch between controller parameters and the dynamic model, and synchronous parameter update strategies cannot meet real-time control requirements, posing greater challenges to autopilot parameter settings.

[0004] During the deformation process, the aerodynamic center and moment of inertia of a variable-sweep aircraft undergo drastic time-varying changes, exhibiting highly nonlinear and strongly coupled characteristics. Although ADRC (Advanced Dynamic Control) has strong robustness, its controller structure contains numerous parameters such as observer bandwidth, controller bandwidth, and nonlinear factors, resulting in significant coupling effects between parameters and making manual tuning extremely complex. Existing fixed-parameter ADRC controllers are ill-suited to adapting to the dynamic characteristic changes of variable-sweep aircraft during multi-stage flight, easily leading to a decline in control quality. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an ADRC parameter optimization method for variable sweep aircraft that takes into account deformation safety constraints, so as to alleviate the above-mentioned problems existing in related technologies.

[0006] In a first aspect, embodiments of the present invention provide an ADRC parameter optimization method for a variable-sweep aircraft considering deformation safety constraints, comprising: determining the current deviation information and current control quantity of the variable-sweep aircraft based on the desired velocity, desired altitude, and current flight state of the variable-sweep aircraft through an ADRC simulation model and a longitudinal motion model of the variable-sweep aircraft; constructing a current state vector based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft; wherein the safety constraint information includes the current angle of attack, current elevator deflection angle, and current normal overload of the variable-sweep aircraft, and the deformation configuration perception information includes the current sweep angle and its rate of change of the variable-sweep aircraft; optimizing the ADRC parameters of the ADRC simulation model based on the current state vector and the current control quantity through a preset reward function and a preset cost function, as well as a pre-trained policy network, reward evaluation network, and cost evaluation network; wherein the policy network is used to calculate a corresponding action vector based on the corresponding state vector, and the action vector includes the corresponding ADRC parameter set of the ADRC simulation model.

[0007] Secondly, embodiments of the present invention also provide an ADRC parameter optimization device for a variable-sweep aircraft considering deformation safety constraints, comprising: a determination module, used to determine the current deviation information and current control quantity of the variable-sweep aircraft based on the desired speed, desired altitude, and current flight state of the variable-sweep aircraft, through an ADRC simulation model and a longitudinal motion model of the variable-sweep aircraft; a construction module, used to construct a current state vector based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft; wherein the safety constraint information includes the current angle of attack, current elevator deflection angle, and current normal overload of the variable-sweep aircraft, and the deformation configuration perception information includes the current sweep angle and its rate of change of the variable-sweep aircraft; and an optimization module, used to optimize the ADRC parameters of the ADRC simulation model based on the current state vector and the current control quantity, through a preset reward function and a preset cost function, as well as a pre-trained policy network, reward evaluation network, and cost evaluation network; wherein the policy network is used to calculate a corresponding action vector based on the corresponding state vector, and the action vector includes the corresponding ADRC parameter set of the ADRC simulation model.

[0008] Thirdly, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints as described in the first aspect above.

[0009] This invention provides an ADRC parameter optimization method for a variable-sweep aircraft considering deformation safety constraints. Based on the desired velocity, desired altitude, and current flight state of the variable-sweep aircraft, the current deviation information and current control input are determined through an ADRC simulation model and a longitudinal motion model of the variable-sweep aircraft. A current state vector is constructed based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft. Based on the current state vector and current control input, the ADRC parameters of the ADRC simulation model are optimized through a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network. Using this technique, the ADRC parameter optimization problem can be modeled as a restricted Markov decision process. By utilizing the desired velocity, desired altitude, current flight state, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft, the ADRC parameters are optimized through autonomous learning via interaction with the environment, thereby mitigating the performance degradation problem of the fixed-parameter ADRC controller caused by drastic changes in aerodynamic parameters during the deformation process of the variable-sweep aircraft.

[0010] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0013] Figure 1 This is a flowchart illustrating an ADRC parameter optimization method for a variable-sweep aircraft considering deformation safety constraints, as described in an embodiment of the present invention. Figure 2 This is a block diagram illustrating the principle of the active disturbance rejection controller in an embodiment of the present invention. Figure 3 This is a block diagram illustrating the principle of the active disturbance rejection control framework based on PPO reinforcement learning in this embodiment of the invention. Figure 4 This is a flowchart of the training process for the Lagrange near-end policy optimization algorithm in an embodiment of the present invention; Figure 5This is a schematic diagram of the ADRC parameter optimization device for a variable sweep aircraft that considers deformation safety constraints in an embodiment of the present invention. Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] Currently, traditional parameter scheduling methods face challenges under variable-sweep flight conditions. Rapid changes in aerodynamic characteristics lead to a mismatch between controller parameters and the dynamic model, and synchronous parameter update strategies cannot meet real-time control requirements, posing greater challenges to autopilot parameter settings. During the deformation process, the aerodynamic center and moment of inertia of a variable-sweep aircraft undergo drastic time-varying changes, exhibiting highly nonlinear and strongly coupled characteristics. Existing fixed-parameter ADRC methods are ill-suited to adapting to the dynamic characteristic changes of variable-sweep aircraft during multi-stage flight, easily leading to a decline in control quality.

[0016] Based on this, the present invention provides an ADRC parameter optimization method for variable sweep aircraft that considers deformation safety constraints, which can alleviate the above-mentioned problems existing in related technologies.

[0017] To facilitate understanding of this embodiment, a detailed description of the ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints, as disclosed in this embodiment of the invention, will be provided first. (See [link to relevant documentation]). Figure 1 As shown, the method may include the following steps: Step S102: Based on the desired speed, desired altitude and current flight state of the variable-sweep aircraft, determine the current deviation information and current control quantity of the variable-sweep aircraft through the ADRC simulation model and longitudinal motion model of the variable-sweep aircraft.

[0018] Step S104: Construct the current state vector based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable sweep aircraft.

[0019] Among them, the safety constraint information may include the current angle of attack, current elevator deflection angle and current normal overload of the variable sweep aircraft, and the deformation configuration perception information may include the current sweep angle and its rate of change of the variable sweep aircraft.

[0020] Step S106: Based on the current state vector and the current control variable, optimize the ADRC parameters of the ADRC simulation model by using a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network.

[0021] The policy network can be used to calculate the corresponding action vector based on the corresponding state vector. The action vector can include the corresponding ADRC parameter set of the ADRC simulation model.

[0022] This invention provides an ADRC parameter optimization method for a variable-sweep aircraft considering deformation safety constraints. Based on the desired velocity, desired altitude, and current flight state of the variable-sweep aircraft, the current deviation information and current control input are determined through an ADRC simulation model and a longitudinal motion model of the variable-sweep aircraft. A current state vector is constructed based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft. Based on the current state vector and current control input, the ADRC parameters of the ADRC simulation model are optimized through a preset reward function, a preset cost function, and pre-trained policy network, reward evaluation network, and cost evaluation network. Using this approach, the ADRC parameter optimization problem can be modeled as a restricted Markov decision process. By utilizing the desired velocity, desired altitude, current flight state, safety constraint information, and deformation configuration perception information of the variable-sweep aircraft, the ADRC parameters are optimized through autonomous learning via interaction with the environment, thereby mitigating the performance degradation problem of the fixed-parameter ADRC controller caused by drastic changes in aerodynamic parameters during the deformation process of the variable-sweep aircraft.

[0023] As one possible implementation, the current flight state may include the current speed, current altitude, and current pitch angle of the variable-sweep aircraft; based on this, step S102 (i.e., determining the current deviation information and current control variables of the variable-sweep aircraft based on its desired speed, desired altitude, and current flight state using the ADRC simulation model and longitudinal motion model of the variable-sweep aircraft) may include: Step A1: Based on the desired speed, desired altitude, and current flight status, determine the desired pitch angle and current control variables of the variable sweep aircraft using the ADRC simulation model.

[0024] For example, the ADRC simulation model may include a first-order active disturbance rejection controller for speed, a second-order active disturbance rejection controller for altitude, and a second-order active disturbance rejection controller for pitch angle. The current control variables may include the current throttle opening and the current elevator deflection of the variable sweep aircraft. Step A1 above may include the following operation: determining the current throttle opening based on the desired speed and the current speed using the first-order active disturbance rejection controller for speed; determining the desired pitch angle based on the desired altitude and the current altitude using the second-order active disturbance rejection controller for altitude; and determining the current elevator deflection based on the desired pitch angle and the current pitch angle using the second-order active disturbance rejection controller for pitch angle.

[0025] Step A2: Based on the current control variables, control the current flight state using the longitudinal motion model.

[0026] Step A3: Based on the desired speed, desired altitude, desired pitch angle, and current flight status, determine the current deviation information using the ADRC simulation model.

[0027] The current deviation information may include the current speed error, current altitude error, and current pitch angle error of the variable sweep aircraft.

[0028] A longitudinal nonlinear dynamic model of the variable-sweep aircraft can be pre-established as shown in equation (1), which includes a sweep angle that varies with time. And its time-varying effect on aerodynamic coefficients.

[0029] (1) In equation (1), It is the total mass of the variable-sweep aircraft. It is aircraft drag. It is engine thrust. It is the lift of the aircraft. For longitudinal pitching moment, It's speed. It's an attack angle. It is the track angle. It is the pitch angle. It is the pitch angular velocity. For the pitch inertia of the aircraft, The flight altitude of an aircraft; lift ,resistance and pitch moment As shown in equation (2).

[0030] (2) In equation (2), For wing reference area, Atmospheric density, The aerodynamic chord of the wing; They are the lift coefficient, pitching moment coefficient, and drag coefficient, respectively, as shown in equation (3).

[0031] (3) In the formula, It's the elevator deflection angle. It's a sweep angle. , and These represent the lift coefficient, drag coefficient, and pitch moment coefficient at zero angle of attack, respectively. , and These are the first derivatives of the lift coefficient, drag coefficient, and pitch moment coefficient with respect to the angle of attack, respectively. This is the second derivative of the drag coefficient with respect to the angle of attack. and These are the first derivatives of the lift coefficient and pitch moment coefficient with respect to the elevator deflection angle, respectively.

[0032] Assuming the thrust of the aircraft It is controllable, and the thrust can be described by a simple linear model, as shown in equation (4).

[0033] (4) In equation (4), This is the engine thrust coefficient. This refers to the throttle opening.

[0034] The longitudinal dynamics of a variable-sweep aircraft can be transformed into the following strict feedback form: (5) in, , , , , , As shown in equation (6).

[0035] (6) Assuming uncertainties Continuously differentiable and satisfying ,in .

[0036] See Figure 2 As shown, it can be used for The first-order system design of the first-order active disturbance rejection controller (i.e., the velocity loop active disturbance rejection controller, including the velocity channel tracking differentiator, the velocity channel extended state observer, and the velocity channel state error feedback controller) is shown in the following equation (7): (7) In equation (7), and These are the hyperparameters of the velocity channel tracking differentiator. It is the speed factor (its function is to determine the speed of tracking commands, and physically corresponds to the maximum acceleration allowed by the system). This is the filtering factor (which determines the ability to filter out noise). A larger value indicates a better filtering effect and a less sensitive system to noise. (Take a multiple of the simulation step size). It is the expected speed. yes Smoothness (continuous and smooth). yes The smoothed derivative can be used as input to the velocity channel tracking differentiator to determine the desired velocity. The output of the differentiator is tracked through the speed channel. and output ; It is the current speed measured by the sensor. It's the throttle opening. This is an estimate of the current speed. It is an estimate of the derivative of the current velocity. It is a velocity error, which can be input into the velocity channel extended state observer. And input Output via velocity channel expansion state observer and output Input can be sent to the speed channel state error feedback controller. , and Output via speed channel state error feedback controller (This is the actual amount under control); and It is the bandwidth hyperparameter of the velocity channel expansion state observer. It is the nominal control gain of the velocity channel extended state observer (this parameter is for) (estimates) It is the hyperparameter of the speed channel state error feedback controller. The parameter tuning can be simplified by using the linear active disturbance rejection controller tuning method as shown in equation (8).

[0037] (8) In equation (8), It is the bandwidth of the velocity channel expansion state observer. It is the bandwidth of the speed channel error feedback controller.

[0038] See Figure 2 As shown, it can be used for The second-order system design of the second-order active disturbance rejection controller (i.e., the height loop active disturbance rejection controller, including the height channel tracking differentiator, the height channel extended state observer, and the height channel state error feedback controller) is shown in the following equation (9): (9) In the formula, and It is the hyperparameter of the height channel tracking differentiator. As a height factor, The filter factor; It is the expected level. yes Smoothness (continuous and smooth). yes The smoothed derivative can be used as input to the height channel tracking differentiator to determine the desired height. The output of the differentiator is tracked through the height channel. and output ; It is the current height measured by the sensor. It is the pitch angle. This is an estimate of the current altitude. It is an estimate of the current height differential. This is an estimate of the total disturbance in the height channel. It is an altitude error, which can be input into the altitude channel extended state observer. And input Output via height channel extended state observer , and Input can be sent to the height channel status error feedback controller. , , , and Output via height channel state error feedback controller (This is a virtual control variable); , and It is the bandwidth hyperparameter of the height channel extended state observer. It is the nominal control gain of the height channel extended state observer (this parameter is for) (estimates) and The hyperparameters of the high-level channel state error feedback controller are simplified by using the linear active disturbance rejection controller parameter tuning method as shown in equation (10).

[0039] (10) In equation (10), It is the bandwidth of the high-channel extended state observer. It is the bandwidth of the high-channel error feedback controller.

[0040] See Figure 2 As shown, it can be used for The second-order system design and the second-order active disturbance rejection controller are shown in equation (11): (11) In the formula, and It is the hyperparameter of the pitch angle channel tracking differentiator. The pitch angle factor, The filter factor; It is the expected pitch angle. yes Smoothness (continuous and smooth). yes The smoothed derivative can be used as input to the pitch angle channel tracking differentiator to obtain the desired pitch angle. The output of the tracking differentiator is achieved through the pitch angle channel. and output ; It is the current pitch angle measured by the sensor. It's the elevator deflection angle. This is an estimate of the current pitch angle. This is an estimate of the current pitch angular velocity. It is an estimate of the total disturbance in the pitch channel. It is the pitch angle error, which can be input into the pitch angle channel extended state observer. And input The output of the extended state observer is achieved through the pitch angle channel. , and ; Input can be sent to the pitch angle channel state error feedback controller , , , and The output is controlled by the pitch angle channel state error feedback controller. (This is the actual amount under control); , and It is the bandwidth hyperparameter of the pitch angle channel extended state observer. It is the nominal control gain of the pitch angle channel extended state observer (this parameter is for) (estimates) and The hyperparameters of the pitch angle channel state error feedback controller are simplified using the linear active disturbance rejection controller parameter tuning method, as shown in the following formula.

[0041] (12) In the formula, It is the bandwidth of the pitch angle channel extended state observer. It is the bandwidth of the pitch angle channel error feedback controller.

[0042] Specifically, for the construction such as Figure 2 The speed active disturbance rejection control loop, altitude active disturbance rejection control loop, and pitch angle active disturbance rejection control loop shown have a total of fifteen undetermined hyperparameters. To reduce the optimization dimensionality and improve computational efficiency, the parameters of these three active disturbance rejection control loops can be classified and processed, as shown in Table 1 below.

[0043] Table 1. Overall Hyperparameter Description of the Three Active Disturbance Rejection Control Loops

[0044] On the one hand, the tracking differentiator (TD) for the three channels (i.e., velocity, altitude, and pitch angle) primarily functions to arrange the transient response and extract the differential signal, having a relatively small impact on the system's steady-state performance. Therefore, the six parameters involved in this part are fixedly tuned using empirical engineering values ​​and do not participate in online optimization. On the other hand, the extended state observer (ESO) and error feedback control law (LSEF) directly determine the system's disturbance rejection capability and dynamic response quality, and are extremely sensitive to the time-varying characteristics of aerodynamic parameters during the variable sweep process. Therefore, the controller bandwidth, observer bandwidth, and nominal control gain of the three channels can be selected as variables to be optimized, and then adaptively adjusted in real time based on the current flight state, sweep angle, and safety constraints. In summary, the action space input to the reinforcement learning policy network... It can be composed of the nine core parameters of the above three channels, namely .

[0045] As one possible implementation, step S104 (i.e., constructing the current state vector based on the current deviation information, safety constraint information, and deformable configuration perception information of the variable-sweep aircraft) may include: forming a current tracking error term from the current velocity error, current altitude error, and current pitch angle error; forming a current error differential term from the rates of change of the current velocity error, current altitude error, and current pitch angle error; forming a current safety constraint term from the current angle of attack, current elevator deflection angle, and current normal overload; forming a current deformable configuration perception term from the current sweep angle and its rate of change; and forming the current state vector from the current tracking error term, current error differential term, current safety constraint term, and current deformable configuration perception term. Accordingly, the ADRC parameter set may include the ADRC parameters to be optimized for each of the first-order active disturbance rejection controller (ADRC) for velocity, the second-order ADRC controller for altitude, and the second-order ADRC controller for pitch angle.

[0046] Following the previous example, based on the aforementioned variable-sweep aircraft model and the aforementioned three-channel active disturbance rejection control loop, it is also necessary to construct a safety reinforcement learning agent. This will enable the aforementioned variable-sweep aircraft model, the aforementioned three-channel active disturbance rejection control loop, and the safety reinforcement learning agent to form a closed-loop control system. The overall control block diagram is as follows: Figure 3 As shown.

[0047] To achieve adaptive optimization of ADRC parameters, the Proximal Policy Optimization (PPO) algorithm can be used as the basic framework. PPO is a gradient-based online reinforcement learning algorithm whose core idea is to introduce a truncation mechanism when optimizing the policy network, limiting the magnitude of each policy update to prevent drastic changes in policy parameters that could lead to training failure.

[0048] Based on the PPO algorithm, the state space can be defined. Action space Reward function and safety cost function .

[0049] The state space design aims to comprehensively describe the aircraft's instantaneous motion state, tracking error, and current aerodynamic configuration. At the time step... The state vector is input into the security reinforcement learning agent. The definition is shown in equation (13): (13) The components in equation (13) are defined as follows: Tracking error term Including speed error Height error and pitch angle error This term reflects the current control accuracy; the error differential term. The rate of change of each of the above errors Used to reflect the error convergence trend; safety constraint terms Including the aircraft's current angle of attack Elevator deflection and normal overload Used to characterize the flight safety of aircraft; Deformation configuration perception item Including the current sweep angle and its rate of change The introduction of a deformation configuration perception term enables the safety reinforcement learning agent to perceive the dynamic changes in the deformation process, thereby outputting appropriate control parameters for different deformation stages.

[0050] The action space directly corresponds to the control parameters to be optimized in the three active disturbance rejection control loops. At time step... Action vector It consists of nine independent control parameters: (14) in, These represent the controller bandwidths for the speed channel, altitude channel, and pitch angle channel, respectively. These represent the observer bandwidths for the velocity, altitude, and pitch angle channels, respectively. These represent the nominal control gain for the speed channel, altitude channel, and pitch channel, respectively.

[0051] The reward function aims to guide the tracking performance and control stability of a safe reinforcement learning agent. At time steps... The single-step reward function is defined as shown in equation (15): (15) in, It is a penalty for tracking errors in speed, altitude, and pitch angle. It is a punishment for controlling input. It is a penalty for large changes in control parameters. It is a reward for steady-state accuracy. , and It is the steady-state accuracy threshold. These are the corresponding weighting coefficients.

[0052] To address the risk of aerodynamic instability caused by drastic changes in aerodynamic characteristics during the variable sweep process of variable-sweep wing aircraft, safety constraints can be set in three aspects: angle of attack, elevator deflection angle, and normal overload. These constraints prevent structural overload and avoid angle-of-attack divergence and control surface saturation due to the rearward shift of the aerodynamic center. (Time step) Single-step safety cost function The definition is as follows.

[0053] (16) in, It is a penalty for overloading beyond safety constraints. It is a penalty for exceeding the angle of attack constraints. It is a penalty for the elevator deflection angle exceeding the safety constraints. It is the overload safety constraint threshold. It is the angle of attack safety constraint threshold. It is the elevator deflection angle safety constraint threshold. These are the corresponding weighting coefficients.

[0054] As one possible implementation, step S106 (i.e., optimizing the ADRC parameters of the ADRC simulation model based on the current state vector and the current control quantity through a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network) may include: calculating the current action vector through the policy network based on the current state vector; calculating the current reward value through the preset reward function based on the current speed error, current altitude error, current pitch angle error, current control quantity, and the action vector calculated by the policy network; calculating the current cost value through the preset cost function based on the current angle of attack, current elevator deflection angle, and current normal overload; iteratively updating the parameters of the policy network through the reward evaluation network and cost evaluation network based on the current state vector, current action vector, current reward value, and current cost value; and determining the current ADRC parameters of the ADRC simulation model based on the action vector calculated by the policy network.

[0055] As one possible implementation, the above-mentioned ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints may further include: randomly sampling a pre-established experience pool to obtain an experience sample set; wherein each experience sample in the experience sample set includes the state vector, action vector, reward value, and cost value corresponding to the corresponding target time, as well as the state vector corresponding to the previous target time.

[0056] Accordingly, the training steps for the reward evaluation network may include: calculating the reward value of each experience sample in the experience sample set using the current reward evaluation network, calculating the reward advantage function value of the experience sample set, and then iteratively updating the parameters of the current reward evaluation network based on the reward value calculated by the current reward evaluation network and the reward advantage function value of the experience sample set, thus obtaining the reward evaluation network. The training steps for the cost evaluation network may include: calculating the cost value of each experience sample in the experience sample set using the current cost evaluation network, calculating the cost advantage function value of the experience sample set, and then iteratively updating the parameters of the current cost evaluation network based on the cost value calculated by the current cost evaluation network and the cost advantage function value of the experience sample set, thus obtaining the cost evaluation network.

[0057] As one possible implementation, the steps of iteratively updating the parameters of the policy network based on the current state vector, current action vector, current reward value, and current cost value through a reward evaluation network and a cost evaluation network may include: iteratively updating the Lagrange operator of a preset Lagrange dominance function based on the current state vector, current action vector, current reward value, and current cost value; calculating the current Lagrange dominance function value based on the current reward value and current cost value through the preset Lagrange dominance function; calculating the current reward value through the reward evaluation network and the current cost value through the cost evaluation network based on the current state vector; and iteratively updating the parameters of the policy network based on the current reward value, current cost value, and current Lagrange dominance function value with the objective of maximizing the Lagrange dominance function value.

[0058] For example, the expression for the preset Lagrange dominance function can be:

[0059] in, express The value of the Lagrange dominance function at time t; and They represent The reward advantage function value and cost advantage function value at each moment; Indicates the first Lagrange multipliers in the next iteration.

[0060] Continuing from the previous example, based on the PPO algorithm framework, the Lagrange multiplier method can be further introduced to handle safety constraints, modeling the aircraft control parameter optimization problem as a Restricted Markov Decision Process (CMDP). This process consists of quaternions. Definition, including state space Action space Reward function and safety cost function Four parts. The Lagrange proximal policy optimization algorithm can be used to dynamically adjust the penalty intensity using Lagrange multipliers, constrain the expected value of the cumulative discount cost, and force the safe reinforcement learning agent to directly search for the optimal ADRC parameter combination under the premise of strictly satisfying the physical safety boundary.

[0061] To solve the constructed restricted Markov decision process, a neural network model based on the Actor-Critic architecture can be established, and the Lagrange multiplier method can be used to handle the security constraints.

[0062] Specifically, three independent neural networks are constructed, namely the policy network (Actor, Instruction, Producer ... ), and reward rating network (Reward Critic, and the cost evaluation network Cost Critic, Policy network (parameters are...) The input is a state vector. The output is an action vector. ) is used to directly generate ADRC control parameters; reward evaluation network (parameters are The input is a state vector. The output is a scalar) used to estimate the expected cumulative reward value in the current state; the cost evaluation network (parameters are...) The input is a state vector. The output (a scalar) is used to estimate the expected cumulative default cost in the current state. The structures of the three neural networks are shown in Table 2.

[0063] Table 2. Explanation of the structure of the strategy network, reward evaluation network, and cost evaluation network.

[0064] In reinforcement learning algorithms, the advantage function is a core metric for evaluating the value of an action. It characterizes the superiority of taking a specific action in a given state compared to the average expected performance in that state. A positive advantage function means that the action is better than the average, and the probability of sampling that action should be increased in policy updates; conversely, a negative advantage function means that the probability of sampling that action should be decreased. Traditional proximal policy optimization (PPO) algorithms only construct and rely on the reward advantage function. The mechanism of updating the strategy has significant hidden dangers in the control of variable sweep aircraft. In order to pursue the maximum cumulative reward, the agent tends to output aggressive control parameters. However, due to the lack of awareness of physical boundaries, such aggressive strategies can easily cause the aircraft to exceed the overload limit, enter the stall angle of attack region, or cause control saturation during the deformation process, thereby causing serious safety accidents.

[0065] To overcome the above shortcomings, a Lagrangian dominance function that integrates performance incentives and security constraints is constructed. This function weights and couples reward advantage and cost advantage through Lagrange multipliers, serving as an indicator for policy network updates. Its calculation formula is defined as follows: (17) In equation (17), and They represent the first During the nth iteration and the 1st iteration Lagrange multipliers in the next iteration; This represents the learning rate of the Lagrange multiplier; This represents the average cumulative security cost of the current batch of data collection trajectories. This indicates the preset safety constraint threshold.

[0066] Following the neural network architecture and Lagrange dominance function design described above, a parameter optimization network model based on Lagrange proximal policy optimization was constructed. Then, the training loop and parameter updates were executed; see [link to documentation]. Figure 4 As shown, the training process of the Lagrange near-end policy optimization algorithm mainly includes: interactive sampling and advantage function calculation, and updates of the reward evaluation network, cost evaluation network and policy network.

[0067] a) Interactive sampling and dominance function calculation: like Figure 4 As shown, execute arrive Iterative learning in batches, each batch containing 3 trajectories, each trajectory includes arrive The time step. For the first Batch, policy network The parameters are Rewards and Evaluation Network The parameters are Cost evaluation network The parameters are Lagrange multipliers are Execute the first batch The simulation program for the step altitude and step velocity of the variable-sweep aircraft corresponding to the trajectory (including: inputting the state vector into the policy network, the policy network calculating and outputting the action vector, running the ADRC controller to calculate the actual control quantity, calculating the flight state using the longitudinal motion equation of the variable-sweep aircraft, calculating the reward function value, and calculating the safety cost function value) yields the [number]th [track]. batch Experience group, calculation Reward timing difference error at time step Cost and timing difference error As shown in equation (18): (18) In equation (18), Discount factor (range of values) (used to measure the weight of future rewards on current decisions). based on and calculate Time-based reward advantage function Cost advantage function and Lagrange dominance function As shown in equation (19): (19) In equation (19), Represents the GAE smoothing coefficient (range of values) (This is used to adjust the variance and bias balance of the dominance estimate).

[0068] b) Updates to the reward evaluation network, cost evaluation network, and policy network: The reward evaluation network update and cost evaluation network update aim to find the optimal evaluation network parameters and minimize the mean square error between the network's predicted value and the target value. Based on this, the two evaluation networks are updated respectively.

[0069] like Figure 4 As shown, each batch includes The network updates the round number and executes the process. arrive The network parameters are updated in each round. For the first round... The first batch Wheel, from the first Randomly selected from batch experience groups One experience group (denoted as) arrive The network parameters were updated. Before optimization, the reward evaluation network parameters were: Loss function of reward evaluation network With parameters The update is shown in equation (20): (20) In equation (20), It is the first The reward target value for an experience sample. It is a reward evaluation network about parameters The prediction function, It is the learning rate of the reward evaluation network. and They are the first The first batch After the first update and the second Updated reward evaluation network parameters; Definition of the first The first batch The updated cost evaluation network parameters are as follows: Loss function of cost evaluation network Construction and parameters The update is shown in equation (21): (twenty one) In equation (21), It is the first The cost target value for an empirical sample It is a cost evaluation network regarding parameters The prediction function, It is the learning rate of the cost evaluation network. It is the first The first batch The updated cost evaluation network parameters.

[0070] The policy network update aims to find the optimal policy network parameters. This allows for maximizing the Lagrange dominance function while satisfying safety constraints. A truncation mechanism from the near-end policy optimization algorithm can be used to limit the policy update magnitude, and an entropy regularization term can be introduced to enhance exploration capabilities.

[0071] like Figure 4 As shown, the first is defined The first batch The updated policy network parameters are as follows: The loss function construction and parameter update of the policy network are shown in Equation (22): (twenty two) In equation (22), This represents the probability ratio between the old and new strategies. Based on parameters The calculated action probability density, Based on the parameters to be optimized The calculated action probability density function; A cutoff function (used to limit the probability ratio to a certain value). (within the interval) This is a preset cutoff factor (used to prevent drastic changes in strategy parameters). This is the core objective function of the PPO algorithm; Entropy of the policy distribution (used to measure the randomness of the policy); is the entropy regularization coefficient (used to encourage agents to maintain a certain level of exploration ability and prevent premature convergence to local optima). The total loss function of the policy network; The learning rate of the policy network. For the first The first batch The updated policy network parameters.

[0072] like Figure 4 As shown, after executing the first... batch After the network parameters are updated, the updated network parameters are assigned to the first... Batch network parameters are used for the first The generation of batch experience groups can be represented as follows: .

[0073] c) Update of Lagrange multipliers: This step aims to adaptively adjust the penalty intensity of security constraints based on the actual security breaches in the current batch of data. (Lagrange multipliers) The update is independent of the neural network parameter update and is performed once after all the data in a batch has been processed.

[0074] like Figure 4 As shown, firstly, the average cumulative security cost of the current batch of collected trajectories is calculated. As shown in equation (23): (twenty three) In equation (23), For the first Number of tracks within a batch This represents the number of steps for a single trajectory. For the first The trajectory in The cost of constant security; Subsequently, the dual gradient ascent method is used to update the Lagrange multipliers. The calculation formula is as follows:

[0075] In this formula, and These represent the Lagrange multipliers before and after the update, respectively. Defined as cost overrun; when Update item For positive, Increasing the value of the penalty term in the Lagrange dominance function leads to an increase in its weight, forcing the policy network to converge to a safe region; when (Safety) When updating items Negative, As the value decreases to zero, the system gradually relaxes its focus on safety and instead concentrates on optimizing control performance. This is the projection operator (used to ensure that the Lagrange multipliers are always non-negative).

[0076] The training process of the Lagrange near-end policy optimization algorithm described above targets equation (17), and achieves optimal performance exploration within the safety constraint boundary by maximizing the Lagrange dominance function. When the aircraft is within the safety envelope (i.e., the predicted cumulative cost)... Below the preset threshold When ), according to the Lagrange multiplier update rule mentioned above, the update term is... A negative value prompts the Lagrange multiplier to... Decrease and approach 0, at which point equation (17) satisfies Penalty items Ineffective; the policy network is entirely driven by reward advantage, and the agent will focus on searching for ADRC parameters that improve response speed and steady-state accuracy, thereby fully exploiting the aircraft's maneuverability. When the agent discovers that aggressive parameters cause the aircraft to approach or violate physical constraints, the cost advantage function... It shows a positive value, but at the same time, because the actual risk exceeds the threshold ( Lagrange multipliers The cost will increase rapidly due to overspending, at which point the penalty term in equation (17) will increase. The weights increase significantly and become dominant, making the Lagrange dominance function... The value becomes extremely small and negative, which sends a strong negative feedback signal to the policy network, forcing the agent to significantly reduce the sampling probability of such high-risk actions in subsequent iterations, and forcibly pull the search path back to the safe region, thus achieving an adaptive trade-off of "safety first, performance second".

[0077] Through the above update process (such as...) Figure 4 (As shown), policy network It can iterate continuously within the parameter space and finally output a control strategy that satisfies both physical safety constraints and achieves optimal configuration of ADRC parameters.

[0078] In practical applications, the aforementioned autopilot for variable-sweep aircraft can be replaced with an autopilot for any object to adjust the ADRC parameters.

[0079] The aforementioned ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints actually provides an adaptive tuning scheme for ADRC parameters of variable-sweep aircraft that considers deformation safety constraints. This method models the control parameter optimization problem of variable-sweep aircraft as a restricted Markov decision process and uses the Lagrange near-end strategy optimization algorithm to optimize ADRC control parameters through interactive learning with the environment, under the premise that the flight state is explicitly constrained to meet the physical safety boundary. This solves the problem of performance degradation of fixed-parameter ADRC controller caused by drastic changes in aerodynamic parameters during the deformation process of variable-sweep aircraft.

[0080] Compared with existing technologies, the above-mentioned ADRC parameter optimization method for variable-sweep aircraft that considers deformation safety constraints has the following beneficial effects: (1) The parameter adaptive matching of the entire variable sweep process is realized: Compared with the traditional fixed gain ADRC, the sweep angle sensing term in the reinforcement learning state space is introduced, which enables the controller to automatically adjust the bandwidth and gain according to the current deformation stage, effectively solving the problem of control quality deterioration caused by aerodynamic center drift in the variable sweep process.

[0081] (2) Ensures the safety of the parameter optimization process: Compared with traditional reinforcement learning algorithms that only pursue reward maximization, the Lagrange dominance function is used to explicitly introduce physical safety constraints such as overload, angle of attack and elevator saturation. Risk is predicted by the cost evaluation network and the optimization direction is dynamically adjusted by the Lagrange multiplier, which effectively avoids the agent outputting aggressive parameters during the exploration process, causing the aircraft to become unstable or the structure to be damaged, and greatly improves the engineering usability of the algorithm.

[0082] (3) Improved overall performance of the control system: By designing a multi-objective reward function that includes tracking error, control energy consumption and motion smoothness, the optimized ADRC parameters not only have high-precision command tracking capability, but also effectively suppress the chatter of the actuator, achieving the best balance between dynamic response and steady-state accuracy.

[0083] Based on the above-described method for optimizing ADRC parameters of a variable-sweep aircraft considering deformation safety constraints, this invention also provides an apparatus for optimizing ADRC parameters of a variable-sweep aircraft considering deformation safety constraints. (See [link to relevant documentation]). Figure 5 As shown, the device may include the following modules: The determination module 502 is used to determine the current deviation information and current control variables of the variable-sweep aircraft based on the desired speed, desired altitude and current flight state of the variable-sweep aircraft, through the ADRC simulation model and longitudinal motion model of the variable-sweep aircraft.

[0084] The construction module 504 is used to construct the current state vector based on the current deviation information, safety constraint information and deformation configuration perception information of the variable sweep aircraft; wherein, the safety constraint information includes the current angle of attack, current elevator deflection angle and current normal overload of the variable sweep aircraft, and the deformation configuration perception information includes the current sweep angle and its rate of change of the variable sweep aircraft.

[0085] The optimization module 506 is used to optimize the ADRC parameters of the ADRC simulation model based on the current state vector and the current control variable, by using a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network; wherein, the policy network is used to calculate the corresponding action vector according to the corresponding state vector, and the action vector includes the corresponding ADRC parameter set of the ADRC simulation model.

[0086] By employing the aforementioned ADRC parameter optimization device for variable-sweep aircraft that considers deformation safety constraints, the ADRC parameter optimization problem can be modeled as a restricted Markov decision process. Utilizing the variable-sweep aircraft's desired velocity, desired altitude, current flight state, safety constraint information, and deformation configuration perception information, the device can autonomously learn through interaction with the environment to optimize ADRC parameters. This alleviates the performance degradation problem of the fixed-parameter ADRC controller caused by drastic changes in aerodynamic parameters during the deformation process of the variable-sweep aircraft.

[0087] The ADRC parameter optimization device for variable-sweep aircraft considering deformation safety constraints provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0088] This invention also provides an electronic device, such as... Figure 6 The diagram shows the structure of the electronic device, which includes a processor 61 and a memory 60. The memory 60 stores computer-executable instructions that can be executed by the processor 61. The processor 61 executes the computer-executable instructions to implement the above-mentioned ADRC parameter optimization method for variable-sweep aircraft that takes into account deformation safety constraints.

[0089] exist Figure 6 In the illustrated embodiment, the electronic device further includes a bus 62 and a communication interface 63, wherein the processor 61, the communication interface 63, and the memory 60 are connected via the bus 62.

[0090] The memory 60 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 62 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 62 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0091] Processor 61 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the aforementioned ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints can be completed through integrated logic circuits in the hardware of processor 61 or through software instructions. Processor 61 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints disclosed in this embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in the memory, and the processor 61 reads the information in the memory and, in conjunction with its hardware, completes the steps of the ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints as described in the aforementioned embodiments.

[0092] Unless otherwise specifically stated, the relative steps, numerical expressions, and values ​​of the components and steps described in these embodiments do not limit the scope of the invention.

[0093] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0094] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0095] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for optimizing ADRC parameters of a variable-sweep aircraft considering deformation safety constraints, characterized in that, include: Based on the desired speed, desired altitude and current flight state of the variable-sweep aircraft, the current deviation information and current control variables of the variable-sweep aircraft are determined by the ADRC simulation model and longitudinal motion model of the variable-sweep aircraft. The current state vector is constructed based on the current deviation information, safety constraint information, and deformable configuration perception information of the variable sweep aircraft; wherein, the safety constraint information includes the current angle of attack, current elevator deflection angle, and current normal overload of the variable sweep aircraft, and the deformable configuration perception information includes the current sweep angle of the variable sweep aircraft and its rate of change; Based on the current state vector and the current control variable, the ADRC parameters of the ADRC simulation model are optimized by using a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network. The policy network is used to calculate the corresponding action vector based on the corresponding state vector, and the action vector includes the corresponding ADRC parameter set of the ADRC simulation model.

2. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 1, characterized in that, The current flight state includes the current speed, current altitude, and current pitch angle of the variable-sweep aircraft; based on the desired speed, desired altitude, and current flight state of the variable-sweep aircraft, the current deviation information and current control variables of the variable-sweep aircraft are determined through the ADRC simulation model and longitudinal motion model of the variable-sweep aircraft, including: Based on the desired speed, the desired altitude, and the current flight state, the desired pitch angle and current control variables of the variable sweep aircraft are determined using the ADRC simulation model. Based on the current control quantity, the current flight state is controlled by the longitudinal motion model; Based on the desired speed, desired altitude, desired pitch angle, and current flight state, the current deviation information is determined using the ADRC simulation model.

3. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 2, characterized in that, The ADRC simulation model includes a first-order active disturbance rejection controller for velocity, a second-order active disturbance rejection controller for altitude, and a second-order active disturbance rejection controller for pitch angle. The current control variables include the current throttle opening and the current elevator deflection of the variable sweep aircraft. Based on the desired speed, desired altitude, and current flight state, the desired pitch angle and current control variables of the variable-sweep aircraft are determined using an ADRC simulation model, including: Based on the desired speed and the current speed, the current throttle opening is determined by a first-order active disturbance rejection controller. Based on the desired altitude and the current altitude, the desired pitch angle is determined by a second-order altitude active disturbance rejection controller. Based on the desired pitch angle and the current pitch angle, the current elevator deflection angle is determined by the second-order active disturbance rejection controller for pitch angle.

4. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 3, characterized in that, The current deviation information includes the current velocity error, current altitude error, and current pitch angle error of the variable sweep aircraft; The current state vector is constructed based on the current deviation information, safety constraint information, and deformable configuration perception information of the variable-sweep aircraft, including: The current speed error, current altitude error, and current pitch angle error are combined to form the current tracking error term, and the rate of change of each of the current speed error, current altitude error, and current pitch angle error is combined to form the current error differential term; The current angle of attack, the current elevator deflection angle, and the current normal overload are combined to form the current safety constraint term, and the current sweep angle and its rate of change are combined to form the current deformation configuration perception term. The current tracking error term, the current error differential term, the current safety constraint term, and the current deformation configuration perception term are combined to form the current state vector.

5. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 4, characterized in that, The ADRC parameter set includes the ADRC parameters to be optimized for each of the first-order velocity active disturbance rejection controller, the second-order altitude active disturbance rejection controller, and the second-order pitch angle active disturbance rejection controller. Based on the current state vector and the current control variable, the ADRC parameters of the ADRC simulation model are optimized through a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network, including: Based on the current state vector, the current action vector is calculated through the policy network; Based on the current speed error, the current altitude error, the current pitch angle error, the current control variable, and the action vector calculated by the policy network, the current reward value is calculated through the preset reward function; Based on the current angle of attack, the current elevator deflection angle, and the current normal overload, the current cost value is calculated using the preset cost function; Based on the current state vector, the current action vector, the current reward value, and the current cost value, the parameters of the policy network are iteratively updated through the reward evaluation network and the cost evaluation network. Based on the action vectors calculated by the policy network, the current ADRC parameters of the ADRC simulation model are determined.

6. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 5, characterized in that, Also includes: Random sampling is performed on a pre-established experience pool to obtain an experience sample set; wherein, each experience sample in the experience sample set includes the state vector, action vector, reward value and cost value corresponding to the corresponding target time, as well as the state vector corresponding to the previous target time. The training steps of the reward evaluation network include: calculating the reward value of each experience sample in the experience sample set through the current reward evaluation network, calculating the reward advantage function value of the experience sample set, and then iteratively updating the parameters of the current reward evaluation network based on the reward value calculated by the current reward evaluation network and the reward advantage function value of the experience sample set to obtain the reward evaluation network. The training steps of the cost evaluation network include: calculating the cost value of each experience sample in the experience sample set through the current cost evaluation network, calculating the cost advantage function value of the experience sample set, and then iteratively updating the parameters of the current cost evaluation network based on the cost value calculated by the current cost evaluation network and the cost advantage function value of the experience sample set, thereby obtaining the cost evaluation network.

7. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 6, characterized in that, Based on the current state vector, the current action vector, the current reward value, and the current cost value, the parameters of the policy network are iteratively updated through the reward evaluation network and the cost evaluation network, including: Based on the current state vector, the current action vector, the current reward value, and the current cost value, the Lagrange operator of the preset Lagrange dominance function is iteratively updated; Based on the current reward value and the current cost value, the current Lagrange dominance function value is calculated using the preset Lagrange dominance function; Based on the current state vector, the current reward value is calculated through the reward evaluation network and the current cost value is calculated through the cost evaluation network; With the goal of maximizing the Lagrange dominance function value, the parameters of the policy network are iteratively updated based on the current reward value, the current cost value, and the current Lagrange dominance function value.

8. The ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints according to claim 7, characterized in that, The expression for the preset Lagrange dominance function is: in, express The value of the Lagrange dominance function at time t; and They represent The reward advantage function value and cost advantage function value at each moment; Indicates the first Lagrange multipliers in the next iteration.

9. A device for optimizing ADRC parameters of a variable-sweep aircraft considering deformation safety constraints, characterized in that, include: The determination module is used to determine the current deviation information and current control variables of the variable-sweep aircraft based on the desired speed, desired altitude and current flight state of the variable-sweep aircraft, through the ADRC simulation model and longitudinal motion model of the variable-sweep aircraft. The construction module is used to construct the current state vector based on the current deviation information, safety constraint information, and deformation configuration perception information of the variable sweep aircraft; wherein, the safety constraint information includes the current angle of attack, current elevator deflection angle, and current normal overload of the variable sweep aircraft, and the deformation configuration perception information includes the current sweep angle and its rate of change of the variable sweep aircraft; An optimization module is used to optimize the ADRC parameters of the ADRC simulation model based on the current state vector and the current control variable, using a preset reward function, a preset cost function, and a pre-trained policy network, reward evaluation network, and cost evaluation network; wherein, the policy network is used to calculate the corresponding action vector based on the corresponding state vector, and the action vector includes the corresponding ADRC parameter set of the ADRC simulation model.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the ADRC parameter optimization method for variable-sweep aircraft considering deformation safety constraints as described in any one of claims 1 to 8.