A multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure
Through the multi-controller fusion mechanism combining the RE-DEORL reinforcement learning algorithm and the EKF filter, the problem of insufficient robustness of traditional propulsion control strategies in multi-modal aircraft is solved, and dynamic fault-tolerant adjustment and robust performance guarantee in the event of sensor failure are realized. It is suitable for low-altitude aircraft with various propulsion configurations.
Patent Information
- Application Number
- CN202510641776.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-19
Smart Images

Figure CN120161778B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent control of aircraft, and in particular to a multi-modal propulsion control method for low-altitude aircraft that is resistant to sensor failure. Background Art
[0002] At present, the low-altitude economy, as a strategic emerging industry that is given key national support, is accelerating the emergence of new aviation platforms, especially short take-off and vertical landing aircraft, which are widely used in urban air traffic, logistics transportation, emergency response and other scenarios; compared with traditional aircraft, short take-off and vertical landing aircraft have the characteristics of multi-mode switching and mixed operation of multiple types of propellers, which put forward higher intelligent, adaptive and robust control requirements for the propulsion system.
[0003] A typical short-takeoff and vertical-drop aircraft needs to frequently undergo multi-state switching of vertical takeoff, climb, cruise, hover, transition, descent and vertical landing during the execution of its mission. The propulsion system has a complex structure, including electric propulsion, turbofan propulsion, tilting mechanism and rudder system. The control variables are highly coupled, and the dynamic response characteristics vary significantly.
[0004] Traditional propulsion control strategies are mostly single-model controllers, which are difficult to take into account multi-modal flight requirements. In complex disturbance environments, they respond slowly, have low parameter adjustment efficiency, and are prone to control performance degradation. Especially in multi-sensor systems, if there are measurement anomalies or faults, it is easy to cause state estimation distortion, which in turn affects control strategy decisions and poses a greater robustness risk.
[0005] In response to the above challenges, artificial intelligence control algorithms have gradually been introduced into the flight control field. In recent years, the multi-strategy fusion of deep reinforcement learning combined with model predictive control and fuzzy control has become a development trend. However, in practical applications, problems such as unstable strategies, uneven switching, and lack of robust processing mechanisms for sensor anomalies still exist.
[0006] To this end, the present invention proposes a multi-modal propulsion control method for low-altitude aircraft that is resistant to sensor failure, a propulsion control method that integrates control strategy output, introduces a sensor residual perception mechanism, and has fault-tolerant parameter adjustment capabilities. Summary of the Invention
[0007] The object of the present invention is to provide a multi-modal propulsion control method for low-altitude aircraft that is resistant to sensor failure. To solve the above-mentioned problems in the prior art, the present invention is achieved through the following technical solutions:
[0008] An embodiment of the present invention provides a multi-mode propulsion control method for a low-altitude aircraft that is resistant to sensor failure, specifically comprising the following steps:
[0009] S1: Construct a nonlinear model of the propulsion component;
[0010] S2: Real-time collection of flight status and sensor observation parameters, input into the flight status identification module and health parameter estimation module to identify the flight mode and the health status of each propulsion component respectively;
[0011] S3: Introduce a weighted embedding fusion network, input the state residual and health state, generate the weighted coefficients of the control strategy group, extract local disturbance features, model the multi-channel state timing dependency, and output the residual weighted vector;
[0012] S4: The identification state and health state are jointly input into the control strategy group to promote control optimization, and the control action output is calculated respectively;
[0013] S5: The control action output is weighted according to the weighted coefficients of the control strategy group to generate the final propulsion instruction, which is distributed to each propulsion subsystem of the aircraft. Each propulsion subsystem feeds back signals and sensor data to the experience pool and health parameter estimation module, and performs strategy updates and fault perception accuracy iterations.
[0014] Furthermore, the specific method for obtaining the nonlinear model is:
[0015] Define the control variable vector to describe the control process of the aircraft propulsion system;
[0016] Define the control variable vector: ;
[0017] in, and Represents the output power percentage of the two independent electric propulsion modules;
[0018] represents the deflection angle of the turbofan nozzle;
[0019] Represents the angle of the tilting mechanism;
[0020] Represents the tail nozzle opening of the turbofan or fan propeller;
[0021] Variables constitute the main control input of the system and participate in the real-time optimization process;
[0022] Define the system state vector to describe the evolution of the propulsion system under different control inputs;
[0023] Define the system state vector: ;
[0024] in, Represents the synthetic thrust of the propulsion system at the current moment;
[0025] represents the energy consumption per unit thrust;
[0026] Represents the aircraft altitude;
[0027] represents the forward speed of the aircraft;
[0028] Based on the nonlinear coupling between different propellers, the nonlinear coupling includes but is not limited to: the fast dynamic response of the electric propeller, the strong hysteresis of the turbofan system and the complex coupling of the airflow interference, the system state transfer process is achieved through the function Perform nonlinear mapping;
[0029] Furthermore, the method for obtaining the health status of the propulsion component is:
[0030] Based on the Actor network, the parameters of the Critic network are updated according to the reward and punishment values of the action. The new loss function is:
[0031]
[0032] The policy gradient update formula used by the Actor network is as follows:
[0033]
[0034] The critic network updates its parameters through the DQN algorithm. The gradient update formula is as follows:
[0035]
[0036] in, and are the learning rates of the Actor network and the Critic network respectively; and are the parameters of the Actor network and the Critic network respectively; Represents the state of the agent in the environment Next action The size of the reward obtained; represents the discount factor, which is used to measure the importance of future rewards. Represents the state The action value obtained by the action generated by the Actor network and input to the Critic network represents the time difference error. Representative Strategy About parameters In state The gradient below, Represents the strategy adopted by the Actor network;
[0037] Identify flight modes and the health status of each propulsion component based on the reward values obtained from the Actor-Critic model;
[0038] Furthermore, the method for obtaining the weighting coefficient is:
[0039] when At the maximum number of rounds, initialize a random process for action exploration , get the initial state , and calculate the initial residual ; Input to the residual perception network to obtain the residual weighted coefficient ;
[0040] Furthermore, the method for obtaining the residual weighted vector is:
[0041] Randomly initialize the current Actor network The weight parameter and the current Critic network The weight parameter ;
[0042] Initialize the target Actor network and target critic network , their respective network weight parameters are: ;
[0043] Initialize the experience replay pool Store state transfer and control residual information;
[0044] Initialize the residual perception network and build a CNN + Transformer combined network;
[0045] The network input is the residual sequence in the current state , that is, the difference between the estimated state and the actual state;
[0046] The CNN module extracts local perturbation peaks, and the Transformer module models the evolution of the residual over time and channels;
[0047] The network output is the residual weight vector ;
[0048] Furthermore, the method for propulsion control optimization is:
[0049] By formula
[0050]
[0051] Rewritten as:
[0052]
[0053] Where:
[0054]
[0055] in, is the time derivative of the state variable, is the state matrix, which describes the influence of the state variable itself on its derivative. is the change of the state variable, is the input matrix, describing the effect of the input on the state derivative, is the change in the input variable, is a matrix that reflects the influence on the state derivative, is the change in the additional variable, is the process noise, is the change in the output variable, is the output matrix, describing the influence of the state on the output, is the direct transfer matrix, which represents the direct effect of input on output. is a matrix that describes the impact on the output. To measure noise, is the new state variable derivative, is the new state vector;
[0056] Based on the design of EKF, the formula is:
[0057]
[0058] Convert to discrete form:
[0059]
[0060] Where the subscript k represents the value of the variable at step k;
[0061] definition:
[0062]
[0063] initialization:
[0064]
[0065] Status prediction:
[0066]
[0067] The forgetting factor in the filter is:
[0068]
[0069] in, is the baseline value of the forgetting factor.
[0070] and The matrix is:
[0071]
[0072] in,
[0073] Measurement prediction:
[0074]
[0075] Nonlinear terms Errors and noise introduced during linearization compared to, Negligible noise in the vicinity; Performing a first-order Taylor expansion yields
[0076]
[0077] Then, we get the batch regression formula of EKF:
[0078]
[0079] in, ;
[0080] The estimation problem in the above formula can be rewritten as
[0081]
[0082] The observation vector is , the regression matrix is , the noise is .error The covariance matrix of
[0083]
[0084] EKF update state calculated in the correction step , obtained by performing weighted least squares, linear regression:
[0085]
[0086] can be transformed into:
[0087]
[0088] The corrected step size of EKF can be obtained as follows:
[0089]
[0090] Among them, is the state vector of the k+1th step, is the discrete state transfer function, is the process noise of the k-th step, is the measurement output of the kth step, is a discrete measurement function, is the measurement noise at step k, is the Jacobian matrix of the state transfer function for state x, is the Jacobian matrix of the measurement function for the state x, is the initial state estimate, is the covariance matrix of the initial state estimate, is the prior state estimate of the k-th step, is the state transition function of the k-th step, is the posterior state estimate of the k-1th step, is the prior covariance matrix of the k-th step, is the forgetting factor of the k-th step, is the state transition function for state x in The Jacobian matrix at , is the posterior covariance matrix of the k-1th step, is the process noise covariance matrix of the k-1th step, is the forgetting factor baseline value, is the intermediate matrix for calculating the forgetting factor, is the intermediate matrix, is a matrix traces, The measurement function h is the state x in The Jacobian matrix at , is the initial covariance, is the measurement noise covariance matrix of the k-th step, is the Kalman gain matrix of the kth step, is the measurement value at step k, is the posterior state estimate at step k, is the posterior covariance matrix of the kth step, h is the nonlinear observation function, is the true state of step k, is the observation noise at step k, is the prior state estimate of the k-th step, The observation function h is The Jacobian matrix at , I is the n×n identity matrix, is the prior state estimation error, is the combined observation vector, is the combined regression matrix, is the combined noise vector, is the covariance matrix of the observation noise, is the covariance matrix of the prior state estimate, is the correction term for covariance update;
[0091] Furthermore, the method for obtaining the control action output is:
[0092] Change the decision-making process of RE-DEORL from a deterministic process to a stochastic process, and add noise based on the policy output action Achieve exploratory expansion and ultimately the actions performed by the environment The expression is:
[0093]
[0094] Among them, Set to Gaussian white noise, the mean μ is the output value of the policy network;
[0095] As the training process continues, the noise variance is continuously reduced to achieve a balance between exploration and utilization;
[0096] Furthermore, the method for obtaining the final advancement instruction is:
[0097] The thrust error does not exceed the tolerance threshold; the thruster power must not exceed the limit value; the attitude angle change rate is limited to prevent sudden maneuvers; the nozzle area, speed and voltage are constrained;
[0098] The mathematical form of the objective function is:
[0099]
[0100] Among them, J is the objective function, λ is the control penalty factor, Generate thrust The energy consumed is the actual thrust generated by the propeller, is the desired target thrust value;
[0101] Furthermore, the method for updating the strategy is:
[0102] Calculate the measurement parameters using the least squares method:
[0103]
[0104] Calculate residuals based on measured parameters:
[0105]
[0106] When all sensors are fault-free and there are no abnormal measurement points, the residual The value of is very small and is a positive integer close to 0. When some sensors have abnormal measurement points, the residual The value of is large;
[0107] because is a dimensional ill-conditioned matrix; eliminate its similar rows, that is, if , then let The i-th row element of is 0, where Representation matrix The i-th row element of the ill-conditioned matrix tolerance error is a constant greater than 0;
[0108] Using the above method, the
[0109]
[0110] The problem to be solved is transformed into the problem of solving a system of non-homogeneous linear equations; using elementary row transformation of determinant; solving
[0111]
[0112] in, is the general solution, the generalized residual For a special solution, , the general solution term represents the series combination of measurement parameters that satisfies the residual of 0, and the special solution term represents the series combination of measurement parameters that causes the residual to be non-zero. The larger the value of the special solution term element, the more abnormal the corresponding sensor measurement parameter is;
[0113] The new weight function is defined as
[0114]
[0115] The measurement prediction equation of GIRLS-EKF is:
[0116]
[0117] in, is the weight matrix, , are the measurement parameters estimated by the least squares method, for The generalized inverse matrix of , r is the residual between the observed value and the estimated value, is the element in row i of the matrix, is the tolerance error constant, is the newly defined weight function, is the residual term, is the threshold constant of the weight function, is the prior state covariance matrix, is the observation noise covariance matrix, is the input variable of step k, which is given by The weight matrix composed of weight functions;
[0118] Furthermore, the method for iterating the fault perception accuracy is:
[0119] Based on ensuring the consistency of the estimator under clean data, Should be fixed to ,in is a standard Gaussian distribution, the scale estimator Solve by iteration:
[0120]
[0121] in, is the relative standard residual, weight function:
[0122]
[0123] Selected The function is:
[0124]
[0125] The weight function is defined as:
[0126]
[0127] in, express The expectation of the function, is a standard Gaussian distribution, for function, is the variance of the scale estimator, m is the number of samples, is the weight function, is the relative standard residual, is the threshold constant, is the estimated value of the state, is an estimated value of the scale.
[0128] Beneficial effects of the present invention:
[0129] 1. This invention constructs a fusion control strategy mechanism that integrates reinforcement learning (RE-DEORL), model predictive control (MPC), and fuzzy control methods. It combines flight status and system health parameters to achieve intelligent weighted fusion of control strategy outputs. This breaks the limitation of traditional propulsion control where a single strategy is difficult to adapt to multiple flight states, effectively improving the system's generalization capability and response flexibility in multi-mode and multi-task stages. The adopted RE-DEORL reinforcement learning algorithm combines the Actor-Critic structure with the DQN concept, integrating the experience replay mechanism with the target network design, improving the stability of strategy training and sample utilization efficiency, preventing the strategy from falling into local optimality. At the same time, the dynamic disturbance expansion mechanism improves the control strategy's adaptability and convergence speed to changes in complex flight environments.
[0130] 2. A residual perception weight network based on CNN and Transformer was introduced, and the GIRLS-EKF filtering structure was combined for multi-channel sensor signal processing and health parameter estimation. It can automatically adjust the control structure in the presence of sensor anomalies, non-Gaussian noise or thruster performance degradation, thereby achieving dynamic fault-tolerant regulation and robust performance assurance of the system; the controller output weights are adaptively adjusted based on flight performance feedback and estimated health status, realizing the "self-evolution" capability of the control strategy; at the same time, the system uses a structured neural network to fuse and decouple control variables, and combines the heterogeneous response characteristics of the thruster to achieve collaborative drive and optimal power distribution among multiple propulsion components. It has good engineering integrability and platform adaptability, and is suitable for various propulsion configurations such as electric propulsion, turbofan hybrid, and tilt-rotor; it achieves the optimization of propulsion control robustness and energy efficiency of low-altitude aircraft under complex flight conditions and sensor abnormal conditions, providing a systematic solution for intelligent control architecture, with significant application value and technology promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0131] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0132] Figure 1 This is a schematic structural diagram of a multi-controller fusion mechanism for resisting sensor failures in a multi-modal propulsion control method for a low-altitude aircraft resisting sensor failures provided in Example 1 of the present invention;
[0133] Figure 2 This is an Actor-Critic network architecture and training flow chart of a multi-modal propulsion control method for a low-altitude aircraft resistant to sensor failure provided in Example 1 of the present invention;
[0134] Figure 3 This is a propulsion control optimization flow chart based on the RE-DEORL algorithm for a multi-modal propulsion control method for a low-altitude aircraft resistant to sensor failure provided in Example 1 of the present invention;
[0135] Figure 4 This is a simplified structural diagram of a method for modeling an engine fault-tolerant airborne adaptive model for a multi-modal propulsion control method for a low-altitude aircraft resistant to sensor failures provided in Example 1 of the present invention;
[0136] Figure 5 1 is a GIRLS-EKF structure flow chart of a multi-modal propulsion control method for a low-altitude aircraft resistant to sensor failure provided in Example 1 of the present invention;
[0137] Figure 6 Schematic diagram of the GIRLS principle of a multi-modal propulsion control method for a low-altitude aircraft resistant to sensor failure provided in Example 1 of the present invention;
[0138] Figure 7 This is a flowchart of the steps of a multi-modal propulsion control method for a low-altitude aircraft that is resistant to sensor failures provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0139] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention.
[0140] Example 1
[0141] like Figure 1 and Figure 7 As shown, an embodiment of the present invention provides a multi-mode propulsion control method for a low-altitude aircraft that is resistant to sensor failure, which specifically includes the following steps:
[0142] S1: Construct a nonlinear model of the propulsion components of a short take-off / vertical landing vehicle;
[0143] For multiple propulsion units in short-to-vertical (SVTOL) aircraft, such as electric propulsion, turbofan propulsion, and tilt-rotor propulsion, a unified nonlinear modeling and intelligent control architecture is constructed. A fusion propulsion control strategy based on the RE-DEORL reinforcement learning algorithm is proposed. This strategy combines a health perception mechanism, a residual weighted network, and a controller weight adjustment mechanism to dynamically coordinate the output of propulsion units.
[0144] Control variables include but are not limited to: electric propulsion power distribution ratio, nozzle area, tilt angle, nozzle deflection angle and mode switching variables;
[0145] In a specific embodiment, the method for constructing the nonlinear model is:
[0146] Define the control variable vector to describe the control process of the aircraft propulsion system;
[0147] Specifically, define the control variable vector: ;
[0148] in, and Represents the output power percentage of the two independent electric propulsion modules, which mainly affects vertical take-off and landing and low-speed attitude control;
[0149] Represents the deflection angle of the turbofan nozzle, which controls the fine adjustment of the propulsion direction, especially for attitude correction during takeoff, hovering and descent;
[0150] Represents the angle of the tilt mechanism, determines the working direction of the electric propulsion or fan, and is the key variable for achieving vertical-horizontal flight transition;
[0151] Represents the tail nozzle opening of the turbofan or fan propeller, affects the outlet airflow speed and momentum, and is an important control parameter for adjusting the balance between horizontal thrust and energy consumption;
[0152] Variables constitute the main control input of the system and participate in the real-time optimization process;
[0153] Define the system state vector to describe the evolution of the propulsion system under different control inputs;
[0154] Specifically, define the system state vector: ;
[0155] in, Represents the synthetic thrust of the propulsion system at the current moment;
[0156] Represents energy consumption per unit thrust (e.g., 100 watts per Newton), representing energy efficiency performance;
[0157] Represents the aircraft altitude, used to reflect the vertical flight state;
[0158] Represents the forward velocity of the aircraft, a measure of horizontal propulsion output;
[0159] Based on the nonlinear coupling between different propellers, the nonlinear coupling includes but is not limited to: the fast dynamic response of the electric propeller, the strong hysteresis of the turbofan system and the complex coupling of the airflow interference, the system state transfer process is achieved through the function Perform nonlinear mapping;
[0160] It should be noted that the function The system is jointly constructed by the aircraft aerodynamic model, propulsion component response model, and power conversion model, supporting iterative optimization during offline simulation and online reinforcement learning training.
[0161] Under the premise of ensuring the stability of the aircraft's mission propulsion, the unit thrust energy consumption is minimized, that is, the following nonlinear constrained optimization model is constructed:
[0162] Performance indicator: Minimize energy consumption rate ;
[0163] Constraints: Flight safety boundaries, including maximum power, motor speed limits, attitude angle limits, dynamic power balance between multiple thrusters, and thrust output that meets different flight state requirements, such as hovering thrust and horizontal cruise thrust;
[0164] The system continuously iterates and updates the policy network through reinforcement learning, and combines the outputs of other controllers to achieve fusion regulation;
[0165] S2: Real-time collection of flight status and sensor observation parameters, input into the flight status identification module and health parameter estimation module to identify the flight mode and the health status of each propulsion component respectively;
[0166] In a specific embodiment, state variables such as flight altitude, speed, mission phase, inclination angle, and thrust requirement are collected in real time. A state classification model based on a decision tree is used to determine the current flight mode, such as vertical takeoff, hovering, horizontal cruise, and transition. The output of this module will serve as a reference variable for the strategy fusion mechanism.
[0167] The control strategy group includes the following three types of control algorithm modules: the RE-DEORL master control module is responsible for master strategy generation and learning control variables; the MPC control module uses the system model to predict control variables and correct trajectories in the short-term domain; the fuzzy control module is based on empirical rules and fuzzy reasoning of state variables to adapt to sudden disturbances and uncertain environments;
[0168] The fusion controller output module takes the three controller outputs as input; it builds a weighted fusion mechanism and outputs the optimal control action; the weights change dynamically and are updated in real time by the self-learning module;
[0169] The self-learning parameter tuning module uses a lightweight neural network to build a state-weight mapping function Dynamically adjust the output weight of each controller based on the flight mission phase, system performance feedback, and customized reward and punishment mechanisms to achieve optimal matching and resource utilization of controller output in different scenarios.
[0170] The execution command generation and distribution module receives the fused control quantities, including but not limited to: propulsion power, motor control signals, and nozzle adjustment quantities; converts them into low-level control commands, and distributes them to each propulsion subsystem through the control bus;
[0171] like Figure 1 As shown in the figure, the multi-controller fusion mechanism structure that is resistant to sensor failure forms a multi-layer control system consisting of state recognition → parallel output of control strategy groups → self-learning fusion → final execution, achieving the coupled optimization of system-level responsiveness and energy efficiency;
[0172] It should be noted that in the multi-controller structure, RE-DEORL is the master control core, generating the main control strategy and providing a benchmark for system training and performance driving. The output of the policy network serves as the initial value of the control variable and also as the reference input for fuzzy control and MPC. RE-DEORL's training objective optimizes the strategy itself, guiding the fusion controller module to learn the optimal combination method under different conditions. During training, the RE-DEORL controller receives interaction data from the experience pool, compares the error between the target network and the real-time network, updates the value network, and then feeds it back to the policy network to complete step-by-step iterative optimization.
[0173] During the flight mission, the control structure automatically determines the flight conditions based on the output of the state perception module and performs the following process:
[0174] Enter the current state , for parallel processing by RE-DEORL, MPC, and fuzzy controller;
[0175] The three controllers calculate the output , , ;
[0176] Will state Input the fusion weight network and calculate the fusion weight ;
[0177] Generate the final control amount according to the fusion formula ;
[0178] Send control instructions to the actuator to complete propulsion control;
[0179] The system collects execution feedback and stores it in the experience pool for RE-DEORL training and updating;
[0180] If there is performance degradation or flight mode switching, automatically readjust the fusion weights or trigger control structure reorganization;
[0181] like Figure 2 As shown in the figure, the Actor-Critic model obtains a better algorithm structure by combining the policy gradient method with the value function method;
[0182] In the Actor-Critic model structure, the Actor network is based on the policy gradient algorithm, which selects the appropriate action from the continuous actions according to the current state;
[0183] The Critic network is based on the DQN equivalent function method, which calculates the reward and penalty values caused by the state transition after the action is executed, and evaluates whether the action is reasonable;
[0184] Actor-Critic structure; updates network parameters in a single step, avoiding the problem of low model efficiency caused by round-based updates of the policy gradient algorithm;
[0185] During the specific interaction process, the Actor network obtains the probability value of each action and then selects the behavior based on the probability;
[0186] The critic network is constantly updated to improve the reward and punishment values for each action selected in each state;
[0187] Based on the Actor network, the parameters of the Critic network are updated according to the reward and punishment values of the action. The new loss function is:
[0188]
[0189] The policy gradient update formula used by the Actor network is as follows:
[0190]
[0191] The critic network updates its parameters through the DQN algorithm. The gradient update formula is as follows:
[0192]
[0193] in, and are the learning rates of the Actor network and the Critic network respectively; and are the parameters of the Actor network and the Critic network respectively; Represents the state of the agent in the environment Next action The size of the reward obtained; represents the discount factor, which is used to measure the importance of future rewards. Represents the state The action value obtained by the action generated by the Actor network and input to the Critic network represents the time difference error. Representative Strategy About parameters In state The gradient below, Represents the strategy adopted by the Actor network;
[0194] Identify flight modes and the health status of each propulsion component based on the reward values obtained from the Actor-Critic model;
[0195] S3: Introduce a weighted embedding fusion network, input the state residual and health state, generate the weighted coefficients of the control strategy group, extract local disturbance features, model the multi-channel state timing dependency, and output the residual weighted vector;
[0196] In a specific embodiment, a residual perception weight mechanism is introduced based on the original Actor-Critic architecture to effectively enhance the system's adaptability to non-Gaussian noise and sudden abnormal signals. The idea of the DQN algorithm is also incorporated into the Actor-Critic model structure, and its structure consists of two parts: a value network and a policy network.
[0197] Reinforcement learning is usually based on the MDP model. The correlation between data leads to unstable training and difficulty in network convergence. The RE-DEORL algorithm uses the experience replay pool and target network to improve the overall performance of the algorithm.
[0198] The Actor-Critic structure is difficult to adjust robustly when faced with sensor signal disturbances, non-Gaussian distribution residuals, and sudden outliers;
[0199] The control strategy is highly sensitive to noise changes, which may lead to strategy divergence or performance degradation. The residual perception weight network is introduced to the following two core points:
[0200] State preprocessing stage: building a CNN+Transformer network structure
[0201] Input: state residual, multi-channel sensor residual stream ;
[0202] CNN module: extract local perturbation features;
[0203] Transformer module: Modeling long-term state dependencies and abnormal persistence patterns;
[0204] Output: residual weight vector , as the weighted adjustment parameter of the control variable;
[0205] The controller output weighting mechanism is introduced
[0206] Original control amount Rewritten as:
[0207]
[0208] in is the state-sensitive dynamic gain coefficient, which is used to adjust the action amplitude under abnormal disturbance;
[0209] For deep reinforcement learning neural networks, a certain amount of sample data is required when updating the weight coefficients of neurons using the gradient descent method;
[0210] If online interactive learning is used, the current data needs to be discarded after the current network update is completed, resulting in a significant reduction in data utilization. The agent needs to interact more with the environment to achieve the final convergence effect.
[0211] Experience replay technology opens up a certain size of buffer area to transfer state information Save it;
[0212] The state transition sample information enters the cache in order. If the cache is full, when a new sample enters, the oldest sample is removed from the cache.
[0213] It should be noted that the RE-DEORL algorithm is a deterministic policy gradient algorithm. After a given initial state, the interaction sequence obtained according to the policy network is fixed. The agent cannot generate different behaviors to deeply explore the environment, so the policy cannot be improved.
[0214] like Figure 3 As shown, propulsion control optimization is performed based on the RE-DEORL algorithm;
[0215] Specifically, randomly initialize the current Actor network The weight parameter and the current Critic network The weight parameter ;
[0216] Initialize the target Actor network and target critic network , their respective network weight parameters are: ;
[0217] Initialize the experience replay pool , used to store state transfer and control residual information;
[0218] Initialize the residual perception network and build a CNN + Transformer combined network to extract the disturbance features and time dependencies in the multi-channel state residual sequence;
[0219] The network input is the residual sequence in the current state , that is, the difference between the estimated state and the actual state;
[0220] The CNN module extracts local perturbation peaks, and the Transformer module models the evolution of the residual over time and channels;
[0221] The network output is the residual weight vector , used to control the dynamic adjustment of variables;
[0222] when At the maximum number of rounds, initialize a random process for action exploration , get the initial state , and calculate the initial residual ; Input to the residual perception network to obtain the residual weighted coefficient ;
[0223] Generate weighted control actions and set the state Input the current policy network and output the original action vector ; Use the residual weighting mechanism to get the final execution action Will Send to the environment for execution;
[0224] New state of environmental feedback ,award , new residual ; Transition state sample Store to the experience replay pool, from the experience replay pool Random Sampling samples as training data for the Actor network and the Critic network, let , by minimizing the loss function To update the Critic network parameters
[0225]
[0226] in, Target value, ; In calculation hour, represents the discount factor, , using the target value network and target strategy network , making The network can remain stable during training and converge more easily;
[0227] Compute the gradient of the policy network:
[0228]
[0229] Update target network and :
[0230]
[0231]
[0232] in, and are the parameters of the current policy network and the current value network respectively, and are the parameters of the target policy network and the target value network respectively, is the update coefficient, which indicates the update step size and balances the current network parameters with the target network parameters;
[0233] It should be noted that RE-DEORL, while retaining the RE-DEORL training framework, integrates the state residual self-perception adjustment mechanism. By introducing the CNN + Transformer combined perception network and action weighting structure, it significantly enhances the robustness and adaptive adjustment ability of the policy network in complex disturbance environments. It is particularly suitable for intelligent control tasks of propulsion systems in the background of multiple sensors, multiple control variables, and non-Gaussian noise.
[0234] S4: The identification state and health state are jointly input into the control strategy group to promote control optimization, and the control action output is calculated respectively;
[0235] like Figure 4 As shown, an adaptive model is constructed;
[0236] The motivation is highly nonlinear, meaning both the system equations and the measurement equations of the model are nonlinear. The EKF performs a Taylor series expansion on the original system and measurements, approximating the system as a first-order linear system before performing a linear Kalman filter estimate. This differs from the linear KF in that it uses a nonlinear model to calculate state and output values, solving the parameter estimation problem for nonlinear systems.
[0237] Since the performance degradation of the engine occurs slowly, it is considered that ; Through the formula
[0238]
[0239] Rewritten as:
[0240]
[0241] Where:
[0242]
[0243] in, is the time derivative of the state variable, is the state matrix, which describes the influence of the state variable itself on its derivative. is the change of the state variable, is the input matrix, describing the effect of the input on the state derivative, is the change in the input variable, is a matrix that reflects the influence on the state derivative, is the change in the additional variable, is the process noise, is the change in the output variable, is the output matrix, describing the influence of the state on the output, is the direct transfer matrix, which represents the direct effect of input on output. is a matrix that describes the impact on the output. To measure noise, is the new state variable derivative, is the new state vector;
[0244] Based on the design of EKF, the formula is:
[0245]
[0246] Convert to discrete form:
[0247]
[0248] Where the subscript k represents the value of the variable at step k;
[0249] definition:
[0250]
[0251] initialization:
[0252]
[0253] Status prediction:
[0254]
[0255] The forgetting factor in the filter is:
[0256]
[0257] in, is the baseline value of the forgetting factor.
[0258] and The matrix is:
[0259]
[0260] in,
[0261] Measurement prediction:
[0262]
[0263] Nonlinear terms Errors and noise introduced during linearization compared to, Negligible noise in the vicinity; Perform a first-order Taylor expansion; we get
[0264]
[0265] Then, we get the batch regression formula of EKF:
[0266]
[0267] in, ;
[0268] The estimation problem in the above formula can be rewritten as
[0269]
[0270] The observation vector is , the regression matrix is , the noise is .error The covariance matrix of
[0271]
[0272] EKF update state calculated in the correction step , obtained by performing weighted least squares, linear regression:
[0273]
[0274] can be transformed into:
[0275]
[0276] The corrected step size of EKF can be obtained as follows:
[0277]
[0278] Among them, is the state vector of the k+1th step, is the discrete state transfer function, is the process noise of the k-th step, is the measurement output of the kth step, is a discrete measurement function, is the measurement noise at step k, is the Jacobian matrix of the state transfer function for state x, is the Jacobian matrix of the measurement function for the state x, is the initial state estimate, is the covariance matrix of the initial state estimate, is the prior state estimate of the k-th step, is the state transition function of the k-th step, is the posterior state estimate of the k-1th step, is the prior covariance matrix of the k-th step, is the forgetting factor of the k-th step, is the state transition function for state x in The Jacobian matrix at , is the posterior covariance matrix of the k-1th step, is the process noise covariance matrix of the k-1th step, is the forgetting factor baseline value, is the intermediate matrix for calculating the forgetting factor, is the intermediate matrix, is a matrix traces, The measurement function h is the state x in The Jacobian matrix at , is the initial covariance, is the measurement noise covariance matrix of the k-th step, is the Kalman gain matrix of the kth step, is the measurement value at step k, is the posterior state estimate at step k, is the posterior covariance matrix of the kth step, h is the nonlinear observation function, is the true state of step k, is the observation noise at step k, is the prior state estimate of the k-th step, The observation function h is The Jacobian matrix at , I is the n×n identity matrix, is the prior state estimation error, is the combined observation vector, is the combined regression matrix, is the combined noise vector, is the covariance matrix of the observation noise, is the covariance matrix of the prior state estimate, is the correction term for covariance update;
[0279] like Figure 5 As shown, the improved iterative reweighted least squares-extended Kalman filter;
[0280] The GIRLS module calculates weights based on the weight formula. The farther the outlier is from the group, the smaller the weight assigned. When the outlier exceeds a certain limit value c, its weight will become 0.
[0281] The hybrid adaptive module updates the measurement noise covariance matrix and forgetting factor based on historical data, thereby continuously optimizing the filter parameters;
[0282] The state quantity is realized by iterative operation using the residual of the measured value and the model calculated value and the parameters calculated by the filter. estimates;
[0283] It should be noted that the engine has a large flight envelope, multiple operating modes, and a harsh operating environment, making it susceptible to interference in complex environments. In addition, the number of sensors is large and prone to failure, which can easily produce large measurement outliers. The classic weighted least squares method, namely the maximum likelihood estimator under Gaussian noise, will be affected by the loss of accuracy, which is manifested as an increase in the variance and bias of the estimated value, thus affecting the reliability and estimation accuracy, and having poor robustness.
[0284] The Huber M estimator is a method for studying abnormal observations of filters. It reconstructs the measurement noise covariance matrix by introducing weights, thereby reducing the impact of abnormal observations on the system and improving the robustness of the filter.
[0285] The Huber cost function is the most commonly used cost function in M estimation, and its expression is shown in the following formula:
[0286]
[0287] In the formula, γ is the adjustment factor, i is 1, 2, ... m, m is the dimension of the observation, is the observation residual;
[0288] The corrected measurement noise covariance matrix is:
[0289]
[0290] Where, is the weight function, .
[0291] It should be noted that by replacing the measurement noise covariance matrix in the EKF with the above formula, the robust EKF algorithm based on Huber is obtained. The measurement noise covariance matrix is reweighted, and different weights are constructed for observation residuals of different sizes, thereby reconstructing the measurement noise covariance matrix, overcoming the influence of abnormal observations on the estimation accuracy, and achieving the purpose of improving the robustness of the estimator.
[0292] According to the formula:
[0293]
[0294] It should be noted that due to the information outliers appearing in the prediction state Their influence will be through the matrix The error propagates to the observation vector on the left side of the above formula, resulting in a large estimation error. It has a certain isolation effect on small deviation outliers caused by disturbances such as noise, and improves the robustness of the filter and the estimation accuracy to a certain extent. When the sensor fails and the measured value deviates greatly from the actual value, it cannot overcome the influence of the abnormal measurement value.
[0295] The weighted least squares estimator is a maximum likelihood estimator under Gaussian noise and is considered to be the estimator that minimizes the norm of the regression residual, i.e. satisfy:
[0296]
[0297] in, is the standard deviation of the residuals, that is , the residual vector is , arg min represents the value of the variable when the following formula reaches the minimum value;
[0298] The size of the standard deviation estimate is affected by abnormal measurements, namely outliers. In order to minimize the corresponding residual square, the filter even tends to tilt towards the outliers.
[0299] like Figure 6 As shown, based on generalized iterative reweighted least squares, the solid circle on the lower left represents the circle to be estimated, the dots on the lower left represent normal observations, the dots on the upper right represent abnormal observations, the solid circle on the upper right represents the WLS fitting result, and the dotted circle represents the GIRLS fitting result;
[0300] The GIRLS method performs iterative reweighting on all observations, assigning a weight of 1 to normal observations and a weight less than 1 to abnormal observations. The farther the abnormal observation is from the normal value, the smaller the weight is, until it reaches 0.
[0301] It should be noted that when the WLS method is used for fitting, since the weight of the abnormal observation value is still 1, the fitting result will have a large deviation from the estimated value. However, the GIRLS method has a much better fitting effect than the WLS method because it isolates the abnormal observation value.
[0302] The estimate is robust if the regression minimizes the robust scale of the residuals, that is:
[0303]
[0304] in, Is a robust scale estimator that minimizes the robust scale estimator; ensuring isolation of abnormal measurements, the estimator for estimating the residual scale is defined as follows:
[0305]
[0306] Where, ,function is satisfied and Bounded function, is an even function and right Not decreasing;
[0307] Change the decision-making process of RE-DEORL from a deterministic process to a stochastic process, and add noise based on the policy output action Achieve exploratory expansion and ultimately the actions performed by the environment The expression is:
[0308]
[0309] Among them, Set to Gaussian white noise, the mean μ is the output value of the policy network;
[0310] As the training process continues, the noise variance is continuously reduced to achieve a balance between exploration and utilization;
[0311] S5: The control action output is weighted according to the weighted coefficients of the control strategy group to generate the final propulsion command, which is distributed to each propulsion subsystem of the aircraft. Each propulsion subsystem feeds back signals and sensor data to the experience pool and health parameter estimation module, and performs strategy updates and fault perception accuracy iterations.
[0312] Based on ensuring the consistency of the estimator under clean data, Should be fixed to ,in is a standard Gaussian distribution, the scale estimator Solve by iteration:
[0313]
[0314] in, is the relative standard residual, weight function:
[0315]
[0316] Selected The function is:
[0317]
[0318] The weight function is defined as:
[0319]
[0320] It should be noted that the standardized residual corresponds to an unbounded The estimator of the function, and ;in, express The expectation of the function, is a standard Gaussian distribution, for function, is the variance of the scale estimator, m is the number of samples, is the weight function, is the relative standard residual, is the threshold constant, is the estimated value of the state, is the scale estimate; the lack of robustness in this case; it is understood as using an equal weight for all residuals Perform weighting to minimize the robust scale of the residual to obtain a highly robust estimator;
[0321] The GIRLS method first calculates the measurement parameters using the least squares method:
[0322]
[0323] Calculate residuals based on measured parameters:
[0324]
[0325] When all sensors are fault-free and there are no abnormal measurement points, the residual The value of is very small and is a positive integer close to 0. When some sensors have abnormal measurement points, the residual The value of is large;
[0326] because is a dimensional ill-conditioned matrix; eliminate its similar rows, that is, if , then let The i-th row element of is 0, where Representation matrix The i-th row element of the ill-conditioned matrix tolerance error is a constant greater than 0;
[0327] Using the above method, the
[0328]
[0329] The problem to be solved is transformed into the problem of solving a system of non-homogeneous linear equations; using elementary row transformation of determinant; solving
[0330]
[0331] in, is the general solution, the generalized residual For a special solution, , the general solution term represents the series combination of measurement parameters that satisfies the residual of 0, and the special solution term represents the series combination of measurement parameters that causes the residual to be non-zero. The larger the value of the special solution term element, the more abnormal the corresponding sensor measurement parameter is;
[0332] The new weight function is defined as
[0333]
[0334] The measurement prediction equation of GIRLS-EKF is:
[0335]
[0336] in, is the weight matrix, , are the measurement parameters estimated by the least squares method, for The generalized inverse matrix of , r is the residual between the observed value and the estimated value, is the element in row i of the matrix, is the tolerance error constant, is the newly defined weight function, is the residual term, is the threshold constant of the weight function, is the prior state covariance matrix, is the observation noise covariance matrix, is the input variable of step k, which is given by The weight matrix composed of weight functions;
[0337] The minimum energy consumption control goal is to minimize the unit thrust energy consumption while outputting the desired thrust in different flight modes. ;
[0338] The control target is used to optimize the aircraft's mission control during cruise, transition, low-altitude inspection and other phases;
[0339] Control variables include: electric propulsion power , tail nozzle opening , tilt angle , nozzle deflection angle .
[0340] The constraints include: thrust error does not exceed the tolerance threshold; thruster power does not exceed the limit value; attitude angle change rate is limited to prevent sudden maneuvers; nozzle area, speed and voltage constraints;
[0341] The mathematical form of the objective function is:
[0342]
[0343] Where λ is the control penalty factor, which is used to balance the relationship between propulsion performance and energy efficiency, J is the objective function, Generate thrust The energy consumed is the actual thrust generated by the propeller, is the desired target thrust value;
[0344] It should be noted that the final problem can be reduced to an optimization problem with nonlinear dynamic constraints. During the solution process, the RE-DEORL iterative strategy is used to continuously optimize the control variable combination to approach the optimal thrust-to-energy ratio.
[0345] An embodiment of the present invention is described in detail above, but the content described is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention; the above formulas are all dimensionless and numerical calculations, and the formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field based on actual conditions and historical experience, and can be adjusted according to actual conditions; the above description is only a preferred embodiment of the present invention and is not used to limit the present invention. All equal changes and improvements made according to the scope of application of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure, characterized in that: The following steps are involved: S1: Construct a nonlinear model of the propulsion component; S2: Real-time collection of flight status and sensor observation parameters, input into the flight status identification module and health parameter estimation module to identify the flight mode and the health status of each propulsion component respectively; S3: Introducing a weighted embedding fusion network, inputting state residuals and health states, generating weighted coefficients for the control strategy group, extracting local disturbance features, modeling multi-channel state temporal dependencies, and outputting a residual weighted vector; The control strategy group includes the following three types of control algorithm modules: the RE-DEORL master control module is responsible for master strategy generation and learning control variables; the MPC control module uses the system model to predict control variables and correct trajectories in the short-term domain; the fuzzy control module is based on empirical rules and fuzzy reasoning of state variables; The fusion controller output module takes the three controller outputs as input; it builds a weighted fusion mechanism and outputs the optimal control action; the weights change dynamically and are updated in real time by the self-learning module; The self-learning parameter tuning module uses a lightweight neural network to construct a state-weight mapping function; it dynamically adjusts the output weights of each controller based on the flight mission phase, system performance feedback, and a custom reward and punishment mechanism; The execution instruction generation and issuing module receives the fused control quantity; The multi-controller fusion mechanism structure that is resistant to sensor failure forms a multi-layer control system consisting of state recognition → parallel output of control strategy groups → self-learning fusion → final execution; S4: The identification state and health state are jointly input into the control strategy group to promote control optimization, and the control action output is calculated respectively; S5: The control action output is weighted according to the weighted coefficients of the control strategy group to generate the final propulsion instruction, which is distributed to each propulsion subsystem of the aircraft. Each propulsion subsystem feeds back signals and sensor data to the experience pool and health parameter estimation module, and performs strategy updates and fault perception accuracy iterations.
2. The multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure according to claim 1, characterized in that: The specific method for obtaining the nonlinear model is: Define the control variable vector to describe the control process of the aircraft propulsion system; Define the control variable vector: ; in, and Represents the output power percentage of the two independent electric propulsion modules; represents the deflection angle of the turbofan nozzle; Represents the angle of the tilting mechanism; Represents the tail nozzle opening of the turbofan or fan propeller; Variables constitute the main control input of the system and participate in the real-time optimization process; Define the system state vector to describe the evolution of the propulsion system under different control inputs; Define the system state vector: ; in, Represents the synthetic thrust of the propulsion system at the current moment; represents the energy consumption per unit thrust; Represents the aircraft altitude; represents the forward speed of the aircraft; Based on the nonlinear coupling between different propellers, the nonlinear coupling includes but is not limited to: the fast dynamic response of the electric propeller, the strong hysteresis of the turbofan system and the complex coupling of the airflow interference, the system state transfer process is achieved through the function Perform nonlinear mapping.
3. The multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure according to claim 1, characterized in that: The method for obtaining the health status of the propulsion component is: Based on the Actor network, the parameters of the Critic network are updated according to the reward and punishment values of the action. The new loss function is: The policy gradient update formula used by the Actor network is as follows: The critic network updates its parameters through the DQN algorithm. The gradient update formula is as follows: in, and are the learning rates of the Actor network and the Critic network respectively; and These are the parameters of the Actor network and the Critic network respectively; Represents the state of the agent in the environment Next action The size of the reward obtained; represents the discount factor, which is used to measure the importance of future rewards. Represents the state The action value obtained by the action generated by the Actor network and input to the Critic network represents the time difference error. Representative Strategy About parameters In state The gradient below, Represents the strategy adopted by the Actor network; The reward value obtained based on the Actor-Critic model identifies the flight mode and the health status of each propulsion component.
4. The multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure according to claim 1, characterized in that: The method for obtaining the weighting coefficient is: when At the maximum number of rounds, initialize a random process for action exploration , get the initial state , and calculate the initial residual ; Input to the residual perception network to obtain the residual weighted coefficient .
5. The multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure according to claim 1, characterized in that: The method for obtaining the residual weighted vector is: Randomly initialize the current Actor network The weight parameter and the current Critic network The weight parameter ; Initialize the target Actor network and target critic network , their respective network weight parameters are: ; Initialize the experience replay pool Store state transfer and control residual information; Initialize the residual perception network and build a CNN+Transformer combined network; The network input is the residual sequence in the current state , that is, the difference between the estimated state and the actual state; The CNN module extracts local perturbation peaks, and the Transformer module models the evolution of the residual over time and channels; The network output is the residual weight vector .
6. The multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure according to claim 1, characterized in that: The method for obtaining the control action output is: Change the decision-making process of RE-DEORL from a deterministic process to a stochastic process, and add noise based on the policy output action Achieve exploratory expansion and ultimately the actions performed by the environment The expression is: Among them, Set to Gaussian white noise, the mean μ is the output value of the policy network; As the training process continues, the noise variance is continuously reduced to achieve a balance between exploration and utilization.
7. The multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure according to claim 1, characterized in that: The method for obtaining the final propulsion instruction is: The thrust error does not exceed the tolerance threshold; the thruster power must not exceed the limit value; the attitude angle change rate is limited to prevent sudden maneuvers; the nozzle area, speed and voltage are constrained; The mathematical form of the objective function is: Among them, J is the objective function, λ is the control penalty factor, To generate thrust The energy consumed is the actual thrust generated by the propeller, is the desired target thrust value.
Citation Information
Patent Citations
Multi-mode intelligent propulsion control method, device and equipment for short vertical aircraft and storage medium
CN120143855A