Multi-mode intelligent propulsion control method, device and equipment for short vertical aircraft and storage medium

Through the multimodal intelligent propulsion control method, combined with decision tree state classification, control strategy group and fusion device weight network, the adaptive adjustment problem of short-sag aircraft under the differentiated dynamic characteristics of multiple propulsion units is solved, and more efficient response and stability are achieved.

CN120143855AActive Publication Date: 2025-06-13XIAMEN UNIV

Patent Information

Application Number
CN202510617334.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Short-sag aircraft cannot adaptively adjust when there are differentiated power characteristics of multiple propulsion units, resulting in slow response, poor stability and low parameter adjustment efficiency.

Method used

The multimodal intelligent propulsion control method is adopted to obtain flight status data in real time, use the decision tree state classification model to generate the current flight mode, and call the control strategy group to generate control instructions with the minimum energy consumption as the optimization goal. This method combines the RE-DEORL control module, the MPC predictive control module and the fuzzy control module, and dynamically adjusts the weight of the control instructions through the fusion weight network to finally generate the final control amount.

Benefits of technology

Adaptive adjustment of short-sag aircraft under differentiated power characteristics of multiple propulsion units is realized, which improves response speed and stability and reduces parameter adjustment complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143855A_ABST
    Figure CN120143855A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode intelligent propulsion control method, device and equipment for a short vertical aircraft and a storage medium, and the method comprises the steps: firstly obtaining flight state data in real time, and classifying the flight state data according to a state classification model of a decision tree to generate a current flight mode; next, calling a control strategy group and performing parallel processing on the current flight mode by taking the lowest energy consumption as an optimization target so as to generate a plurality of control instructions, and calling a fusion device weight network to process the current flight mode so as to generate a fusion weight of each control instruction, and generating a final control quantity according to the plurality of control instructions and the fusion weight, and sending the final control quantity to a pre-constructed hybrid propulsion structure nonlinear model to control the short vertical aircraft, thereby solving the problem that the short vertical aircraft cannot be adaptively adjusted when a plurality of propulsion units have differential dynamic characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of flight control, and particularly to a multi-modal intelligent propulsion control method, device, equipment and storage medium for a short take-off and landing aircraft. Background Art

[0002] Short Take-Off and Landing Aircraft (STOL) can take off and land in limited spaces, and are flexible and efficient, making them important tools in the low-altitude economy. They are widely used in urban transportation, logistics distribution, emergency rescue and other fields.

[0003] Vertical / Short Take-Off and Landing Aircraft (referred to as short take-off and landing aircraft) often include multiple types of propulsion components (electric propulsion, turbofan, tiltrotor, etc.), and frequently experience state transitions such as "vertical take-off - climb - transition - cruise - descent - vertical landing" during missions. Traditional propulsion control methods mostly use single strategies, which have problems such as slow response, poor stability and low tuning efficiency when facing complex dynamic working conditions. Especially when there are different dynamic characteristics among multiple propulsion units, if the control system cannot perform real-time adaptive adjustment, it will seriously affect the overall performance of the aircraft.

[0004] In view of this, the present application is proposed. Summary of the Invention

[0005] The present invention discloses a multi-modal intelligent propulsion control method, device, equipment and storage medium for a short take-off and landing aircraft, which solves the problem that the short take-off and landing aircraft cannot perform adaptive adjustment when there are different dynamic characteristics among multiple propulsion units.

[0006] The first embodiment of the present invention provides a multi-modal intelligent propulsion control method for a short take-off and landing aircraft, including: Obtaining flight state data in real time, and classifying the flight state data according to the state classification model of a decision tree to generate a current flight mode; Invoking a control strategy group and taking the lowest energy consumption as the optimization goal, performing parallel processing on the current flight mode to respectively generate multiple control instructions, and invoking a fusion weight network for the current flight mode to generate a fusion weight for each control instruction, generating a final control amount according to the multiple control instructions and the fusion weight, wherein the fusion weight of each control instruction can be dynamically adjusted based on the flight mission stage, system performance feedback, and reward and punishment mechanism, and the control strategy group includes a RE-DEORL control module, an MPC predictive control module, and a fuzzy control module; Send the final control quantity to a pre-constructed non-linear model of the hybrid propulsion structure to control the short takeoff and landing aircraft, and collect feedback signals to continuously optimize the policy parameters of the RE-DEORL control module; wherein, the non-linear model of the hybrid propulsion structure includes an electric propulsion system, a turbofan propulsion system, a tilting system, and an energy scheduling system.

[0007] Preferably, the current flight modes include vertical takeoff, climb, transition, horizontal cruise, hover, descent, and vertical landing.

[0008] Preferably, the control variables of the non-linear model of the hybrid propulsion structure are:

[0009] Wherein, are the output power percentages of two independent electric propulsion modules respectively, is the turbofan nozzle deflection angle, is the tilting mechanism angle, is the tail nozzle opening of the turbofan; The state variables of the non-linear model of the hybrid propulsion structure are:

[0010] Wherein, is the combined thrust of the propulsion system at the current moment, is the energy consumption per unit thrust, is the altitude of the aircraft, is the forward speed of the aircraft.

[0011] Preferably, the RE-DEORL control module is configured to receive and process the current flight state vector, the residual weighted vector output by the residual perception network, and the historical interaction data sampled from the experience replay pool, and output the main control action vector and the residual weighted vector; The MPC predictive control module is configured to receive and process the current flight state vector, the linear time-varying model parameters, the reference trajectory, the weight matrix, and the prediction step number, and output the predictive control sequence, and use the first control quantity of the control sequence as the control instruction; The fuzzy control module is configured to receive and process the current flight state vector and the pre-constructed fuzzy rule base, and output the original control quantity after fuzzy inference and the correction signal for compensating sudden disturbances.

[0012] Preferably, the execution process of the RE-DEORL control module is: Randomly initialize the weight parameters of the current Actor network and the weight parameters of the current Critic network , and the weight parameters of the current Critic network ​ ; Initialize the target Actor network and the target Critic network , and their respective network weight parameters are as follows: , ; Initialize the experience replay pool R for storing state transitions and control residual information; Residual-aware network initialization: Construct a CNN + Transformer combined network for extracting perturbation features and temporal dependencies in the multi-channel state residual sequence; The network input is the residual sequence at the current state , where is the predicted state, and is the true state; The CNN module extracts local perturbation peaks, and the Transformer module models the evolution relationship of residuals over time and channels; The network output is the residual weighted vector for controlling variable dynamic adjustment; When entering the learning and training loop and reaching the maximum number of episodes, initialize a random process N for action exploration to obtain the initial state and calculate the initial residual ; Input it into the residual-aware network to obtain the residual weighting coefficient Generate the weighted control action: Input the state into the current policy network to output the original action vector ; Use the residual weighting mechanism to obtain the final executed action ; Send to the environment for execution; The environment feedbacks the new state , the reward , and the new residual ; Store the state transition sample in the experience replay pool, and randomly sample samples from the experience replay pool R as the training data for the Actor network and the Critic network; Let , and update the Critic network parameters by minimizing the loss function L ; Among them, is the target value; When calculating , represents the discount factor, , and the target value network and the target policy network are used to make The network can remain stable during training and converge quickly; Calculate the gradient of the policy network:

[0013] Update the target network and :

[0014]

[0015] wherein, is the update coefficient, representing the update step size, which balances between the current network parameters and the target network parameters.

[0016] Preferably, the minimum energy consumption is the optimization objective: while outputting the desired thrust, minimize the energy consumption per unit thrust; The minimum energy consumption control objective is: in different flight modes, while outputting the desired thrust, minimize the energy consumption per unit thrust; The mathematical form of the objective function is:

[0017] wherein, is the energy consumption rate at the current moment, is the combined thrust at the current moment, is the target thrust, is the control penalty factor.

[0018] The second embodiment of the present invention provides a short takeoff and landing aircraft multi-modal intelligent propulsion control device, including: A flight mode classification unit, configured to obtain flight state data in real time, and classify the flight state data according to the state classification model of the decision tree to generate the current flight mode; A final control quantity generation unit, configured to call the control policy group and take the minimum energy consumption as the optimization objective, perform parallel processing on the current flight mode to respectively generate a plurality of control instructions, and call the fusion weight network to process the current flight mode to generate the fusion weight of each control instruction, generate the final control quantity according to the plurality of control instructions and the fusion weight, wherein, the fusion weight of each control instruction can be dynamically adjusted based on the flight mission stage, system performance feedback, and reward and punishment mechanism, and the control policy group includes a RE-DEORL control module, an MPC predictive control module, and a fuzzy control module; A feedback control unit is used to send the final control quantity to a pre - constructed non - linear model of the hybrid propulsion structure to control the short - takeoff and vertical - landing aircraft, and collect feedback signals to continuously optimize the strategy parameters of the RE - DEORL control module; wherein, the non - linear model of the hybrid propulsion structure includes an electric propulsion system, a turbofan propulsion system, a tilting system, and an energy scheduling system.

[0019] The third embodiment of the present invention provides a short - takeoff and vertical - landing aircraft multi - mode intelligent propulsion control device, which is characterized by including a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement a short - takeoff and vertical - landing aircraft multi - mode intelligent propulsion control method as described in any one of the above.

[0020] The fourth embodiment of the present invention provides a computer - readable storage medium, which is characterized by storing a computer program. The computer program can be executed by the processor of the device where the computer - readable storage medium is located to implement a short - takeoff and vertical - landing aircraft multi - mode intelligent propulsion control method as described in any one of the above.

[0021] Based on the short - takeoff and vertical - landing aircraft multi - mode intelligent propulsion control method, device, equipment and storage medium provided by the present invention, first, flight state data is obtained in real - time, and the flight state data is classified according to the state classification model of the decision tree to generate the current flight mode; then, the control strategy group is called and the current flight mode is processed in parallel with the minimum energy consumption as the optimization goal to respectively generate a plurality of control instructions, and the fusion weight network is called to process the current flight mode to generate the fusion weight of each control instruction. The final control quantity is generated according to the plurality of control instructions and the fusion weight, and the final control quantity is sent to the pre - constructed non - linear model of the hybrid propulsion structure to control the short - takeoff and vertical - landing aircraft, solving the problem that the short - takeoff and vertical - landing aircraft cannot be adaptively adjusted when there are different dynamic characteristics in multiple propulsion units. Description of the Drawings

[0022] Figure 1 is a flowchart of a short - takeoff and vertical - landing aircraft multi - mode intelligent propulsion control method provided by the first embodiment of the present invention; Figure 2 is a schematic diagram of the multi - controller fusion mechanism structure of the present invention; Figure 3 is a flowchart of the Actor - Critic network architecture and training; Figure 4 is a flowchart of the propulsion control optimization based on the RE - DEORL algorithm Figure 5 is a module diagram of a short - takeoff and vertical - landing aircraft multi - mode intelligent propulsion control device provided by the second embodiment of the present invention. Detailed Embodiments

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] For a better understanding of the technical solutions of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0025] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0027] It should be understood that the term " / " used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0028] Depending on the context, the word "if" as used herein can be interpreted as "when", "while", "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detected (stated condition or event)", or "in response to detecting (stated condition or event)".

[0029] The "first" and "second" mentioned in the embodiments are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that the "first" and "second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first" and "second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0030] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0031] The present invention discloses a multi-modal intelligent propulsion control method, device, equipment and storage medium for a short takeoff and landing aircraft, which solves the problem that the short takeoff and landing aircraft cannot adaptively adjust when there are different dynamic characteristics in multiple propulsion units.

[0032] Please refer to Figure 1 and Figure 2 , a first embodiment of the present invention provides a multi-modal intelligent propulsion control method for a short takeoff and landing aircraft, which can be executed by a multi-modal intelligent propulsion control device for a short takeoff and landing aircraft (hereinafter referred to as the control device), and particularly, executed by one or more processors in the control device to at least achieve the following steps: S101, obtain flight state data in real time, and classify the flight state data according to the state classification model of the decision tree to generate the current flight mode, where the current flight mode includes vertical takeoff, climb, transition, horizontal cruise, hover, descent and vertical landing; In this embodiment, the control device may be a processor configured on the short takeoff and landing aircraft, and it can establish a communication connection with the sensor group on the short takeoff and landing aircraft.

[0033] The control device collects multi-dimensional state data of the aircraft in real time through on-board sensors, including dynamic parameters such as flight altitude, forward speed, tilt angle, current combined thrust, and mission phase requirements. These data are input into the state classification model based on the decision tree after preprocessing, and it dynamically discriminates the flight state through the feature segmentation rules constructed in the offline training stage. For example, when the flight altitude is lower than the threshold and the forward speed approaches zero, the model combines the tilt angle and thrust demand to determine that the aircraft is in the vertical takeoff or landing stage; if it is detected that the speed continues to rise and the tilt mechanism angle gradually decreases, it is determined as the transition mode from vertical takeoff to horizontal cruise; in the cruise stage, the model matches the speed stability and energy consumption rate threshold to confirm that it is currently in the horizontal cruise state. At the same time, for sudden disturbances or mission changes (such as emergency hover or transfer), the model uses the time-dependent features of the historical state sequence and combines the real-time thrust error and attitude change rate for dynamic correction. Finally, the classification result is output as the current flight mode label (including vertical takeoff, climb, transition, horizontal cruise, hover, descent and vertical landing).

[0034] S102. Invoke the control strategy group and, with the lowest energy consumption as the optimization goal, perform parallel processing on the current flight mode to respectively generate multiple control instructions, and invoke the fusion weight network to process the current flight mode to generate the fusion weight of each control instruction. Generate the final control quantity according to the multiple control instructions and the fusion weights. Among them, the fusion weight of each control instruction can be dynamically adjusted based on the flight mission phase, system performance feedback, and reward and punishment mechanism. The control strategy group includes an RE-DEORL control module, an MPC predictive control module, and a fuzzy control module. In this embodiment, the control device performs multi-dimensional collaborative control on the current flight mode through the parallel processing architecture of the control strategy group, with the lowest energy consumption as the optimization goal. The flight state vector (including real-time data such as speed, altitude, thrust, and energy consumption) and task requirements are synchronously input into the RE-DEORL control module, the MPC predictive control module, and the fuzzy control module.

[0035] Among them, the RE-DEORL control module is based on the deep reinforcement learning framework, receives the current flight state vector and the residual weighted vector dynamically generated by the residual perception network (reflecting the characteristics of environmental disturbances), and combines the historical interaction data sampled from the experience replay pool (such as state transition samples, reward feedback, and control residuals). Through the Actor-Critic network, it iteratively optimizes and generates the main control action vector (including the propulsion power distribution ratio, nozzle angle adjustment amount, etc.), and at the same time outputs the residual weighted vector to dynamically correct the gain amplitude of the control instruction.

[0036] The MPC predictive control module relies on the linear time-varying model parameters (such as the dynamic response matrix of the propulsion system, the energy consumption-thrust coupling coefficient) and the reference trajectory (the flight phase target curve), combines the weight matrix (balancing the priorities of thrust tracking and energy consumption optimization) and the prediction step (the short-time domain rolling optimization window), and online solves the optimal solution of the future multi-step control sequence, and outputs the first-step control quantity as the real-time instruction to achieve the forward-looking correction of the flight trajectory. The fuzzy control module is based on the fuzzy rule base (such as the expert experience rule of "if the thrust error is large and the energy consumption rate is high, then increase the electric propulsion power compensation"), performs fuzzy inference on the current flight state vector, and generates the original control quantity (such as the emergency attitude adjustment signal) and the correction signal (such as the nozzle deflection compensation amount when the wind field suddenly changes) that take into account the sudden disturbance compensation.

[0037] Further, the control device dynamically analyzes the current flight mode and multi-dimensional performance data through the fuser weight network to generate the fusion weights of each control instruction. This network takes the flight mission phase labels (such as vertical takeoff, cruise, etc.), real-time performance feedback (including unit thrust energy consumption, thrust tracking error, attitude stability index), and the reward signals output by the reward and punishment mechanism (such as energy consumption optimization reward, thrust deviation penalty) as inputs, and constructs a state-weight mapping relationship through a lightweight fully connected neural network. The network first encodes the flight mode into a feature vector, and performs multi-channel splicing with the performance feedback data (such as energy consumption rate , thrust error . Subsequently, it extracts high-order correlation features through two hidden layers, and finally outputs the original weight values of each control strategy , , . To meet the weight normalization constraint, the Softmax function is used to probabilize the original weights to obtain the fusion weights that satisfy . For example, during the cruise phase, if the system detects that the energy consumption rate continues to exceed the standard, the network will automatically increase the weight of the RE-DEORL control module to strengthen its energy consumption optimization ability; while in the event of a sudden wind shear, the weight of the fuzzy control module will quickly increase based on the feedback of the disturbance compensation reward signal in the reward and punishment mechanism to prioritize ensuring control robustness. Finally, the system linearly weights the main control action vector generated by RE-DEORL, the first term of the predictive control sequence of MPC, and the compensation correction signal of fuzzy control according to the fusion weights to generate the final control quantity of dynamic optimization. At the same time, the weight network parameters are continuously updated through online training: based on the historical state-action-reward samples stored in the experience replay pool, with the goal of maximizing the cumulative reward, the gradient descent method is used to optimize the network weight allocation strategy, enabling the fusion mechanism to adapt to the evolution of the flight environment and the migration of task requirements, and achieving the global optimal balance of energy consumption, stability, and response speed.

[0038] S103. Send the final control quantity to a pre-constructed non-linear model of the hybrid propulsion structure to control the short takeoff and landing aircraft, and collect feedback signals to continuously optimize the strategy parameters of the RE-DEORL control module; wherein, the non-linear model of the hybrid propulsion structure includes an electric propulsion system, a turbofan propulsion system, a tilt system, and an energy scheduling system.

[0039] Preferably, the control variables of the non-linear model of the hybrid propulsion structure are:

[0040] wherein, are respectively the output power percentages of two independent electric propulsion modules, is the deflection angle of the turbofan nozzle, is the angle of the tilting mechanism, is the opening of the turbofan's tail nozzle; The state variables of the non - linear model of the hybrid propulsion structure are:

[0041] Among them, is the combined thrust of the propulsion system at the current moment, is the energy consumption per unit thrust, is the altitude of the aircraft, is the forward speed of the aircraft.

[0042] In specific implementation, the control device sends the finally generated control quantities (including variables such as the power distribution ratio of the electric propulsion, the deflection angle of the turbofan nozzle, the angle of the tilting mechanism, and the opening of the tail nozzle) to the pre - constructed non - linear model of the hybrid propulsion structure to drive the multi - mode propulsion system of the short - takeoff - and - landing aircraft to work in coordination. This model simulates the coupled response characteristics of the electric propulsion system (distributed propeller power output), the turbofan propulsion system (horizontal thrust generation and nozzle adjustment), the tilting system (vertical - to - horizontal flight transition angle control), and the energy scheduling system (dynamic distribution of electric energy and fuel) through a high - fidelity dynamics simulation environment. Its control variables are defined as:

[0043] Among them, are the output power percentages of two independent electric propulsion modules respectively, is the deflection angle of the turbofan nozzle, is the angle of the tilting mechanism, is the opening of the turbofan's tail nozzle Among them, is the output power percentage of two independent electric propulsion modules, which directly affects the lift generation efficiency during the vertical take - off and landing phase; is the deflection angle of the turbofan nozzle, used for fine - tuning the horizontal propulsion direction and attitude correction; is the angle of the tilting mechanism, which determines the thrust direction switching rate of the electric thruster or fan; is the opening of the turbofan tail nozzle, which optimizes the horizontal thrust and energy consumption balance by adjusting the air flow cross - sectional area.

[0044] The state variables of the model feedback the dynamic performance of the aircraft in real - time:

[0045] is the combined thrust of the propulsion system at the current moment, is the energy consumption per unit thrust, is the altitude of the aircraft, is the forward speed of the aircraft Among them, is the current synthetic thrust, reflecting the collaborative output effect of multiple propulsion units; is the energy consumption per unit thrust, used to quantify the energy efficiency optimization level; and respectively characterize the flight altitude and forward speed, jointly describing the characteristics of the flight phase. The system collects the above state variables through a sensor network and inputs them into the residual perception network and the experience replay pool to form a closed-loop feedback link.

[0046] In each control cycle, the response results of the hybrid propulsion model (such as the actual thrust ), the deviation of the energy consumption rate and the attitude change rate) are compared with the preset target values to generate a residual signal . This residual signal is used by the residual perception network (CNN + Transformer structure) to extract the disturbance characteristics and output the dynamic gain coefficient , which is used to correct the main control action amplitude of the RE-DEORL control module. At the same time, the experience replay pool continuously stores the state transition samples , and updates the policy parameters of the Actor-Critic network through periodic sampling training: the Critic network optimizes the value function with the goal of minimizing the target Q-value error, and the Actor network improves the action generation quality according to the policy gradient. For example, when the model detects a high energy consumption per unit thrust during the hover phase, the feedback signal triggers the RE-DEORL module to adjust the power distribution strategy of the electric propulsion, reducing the power output ratio in the low-efficiency interval; if there is a thrust fluctuation caused by a gust disturbance, the residual perception network dynamically increases the compensation weight of the fuzzy control module and quickly corrects the tilt angle to maintain attitude stability. Through the closed-loop control and online learning mechanism, the non-linear dynamic characteristics of the hybrid propulsion model are gradually incorporated into the optimization target of the RE-DEORL policy network, realizing the adaptive iteration of the control parameters.

[0047] In a possible implementation manner of the present invention, the principle and design process of the RE-DEORL algorithm RE-DEORL (Residual Enhanced Deep Exploration and Optimization Reinforcement Learning) algorithm process. Its core is to introduce a residual-aware weight mechanism on the basis of the original Actor-Critic architecture, effectively enhancing the system's adaptability to non-Gaussian noise and sudden abnormal signals, and integrating the idea of the DQN algorithm into the Actor-Critic model structure. Its structure includes two parts: a value network and a policy network. Reinforcement learning is usually based on the MDP model, and there is a correlation between data, which easily leads to unstable training and difficulty in network convergence. Therefore, the RE-DEORL algorithm uses two methods, an experience replay pool and a target network, to improve the overall performance of the algorithm. When facing sensor signal disturbances, non-Gaussian distributed residuals, and sudden outliers, the Actor-Critic structure is difficult to robustly adjust; the control strategy is highly sensitive to noise changes, which may lead to policy divergence or performance degradation. The residual-aware weight network is introduced into the following two core points: State preprocessing stage: Construct a CNN+Transformer network structure Input: State residual (multi-channel sensor residual stream ); CNN module: Extract local perturbation features (abnormal fluctuations); Transformer module: Model long-term state dependencies and abnormal persistence patterns; Output: Residual weighted vector , as the control variable weighting adjustment parameter.

[0048] Introduce in the controller output weighting mechanism

[0049] Original control quantity Can be rewritten as:

[0050] Where is the state-sensitive dynamic gain coefficient, used to adjust the action amplitude under abnormal disturbances.

[0051] For the neural network of deep reinforcement learning, when updating the weight coefficients of neurons using the gradient descent method, a certain amount of sample data is required. If the online interactive learning method is used, the current data needs to be discarded after the current network update is completed, resulting in a significant reduction in data utilization rate. The agent needs to interact with the environment more to achieve the final convergence effect. The experience replay technique creates a buffer area of a certain size and stores the state transition information Saved. The state transition sample information enters the buffer in sequence. If the buffer is full, when a new sample enters, the oldest sample will be removed from the buffer.

[0052] As a deterministic policy gradient algorithm, the RE-DEORL algorithm has a fixed interaction sequence obtained from the policy network given an initial state. The agent cannot generate different behaviors to deeply explore the environment, so the policy cannot be improved. To change the decision-making process of RE-DEORL from deterministic to stochastic, noise N is added to the output action of the policy to achieve exploratory expansion. Finally, the action executed by the environment The expression is: . N is usually set to Gaussian white noise, and its mean is the output value of the policy network. As the training process progresses, the noise variance is continuously reduced to achieve a balance between exploration and exploitation.

[0053] Actor-Critic method: The policy gradient algorithm optimizes according to the gradient of the policy function. The parameters are slightly corrected along the gradient direction, making the optimization process relatively smooth and with small fluctuations, but the efficiency is also relatively low. In addition, the gradient method also makes the policy gradient algorithm prone to converge to a local optimal solution rather than the desired global optimal solution. Therefore, the Actor-Critic model obtains a better algorithm structure by combining the policy gradient method with the value function method, as Figure 3 shown.

[0054] In the Actor-Critic model structure, the Actor network is based on the policy gradient algorithm and can select appropriate actions from continuous actions according to the current state; the Critic network is based on value function methods such as DQN and calculates the reward and punishment value brought by the state transition after the execution of this action to evaluate whether this action is reasonable.

[0055] The Actor-Critic structure can update the network parameters in a single step, thus avoiding the problem of low model efficiency caused by the episodic update of the policy gradient algorithm. In the specific interaction process, the Actor network obtains the probability value of each action and then selects behaviors based on the probability size; the Critic network is continuously updated to improve the reward and punishment value of each action selected in each state; finally, the Actor network updates its own parameters according to the reward and punishment value of the action by the Critic network. At this time, the new loss function is:

[0056] The policy gradient update formula adopted by the Actor network is as follows:

[0057] The Critic network updates its parameters through the DQN algorithm, and the gradient update formula is as follows:

[0058] where, and are the learning rates of the Actor network and the Critic network respectively and are the parameters of the Actor network and the Critic network respectively; represents the reward value obtained by the agent when performing action in the environmental state .

[0059] Please combine with Figure 4 , the execution process of the RE-DEORL control module is as follows: Randomly initialize the weight parameters of the current Actor network and the weight parameters of the current Critic network , ; Initialize the target Actor network and the target Critic network , and their respective network weight parameters are: , , ; Initialize the experience replay pool R to store state transition and control residual information; Residual perception network initialization: Construct a CNN + Transformer combined network to extract the perturbation features and time dependence in the multi-channel state residual sequence; the network input is the residual sequence in the current state , where is the predicted state, is the real state; is the CNN module to extract the local perturbation peak, and the Transformer module models the evolution relationship of the residual over time and channels; the network output is the residual weighted vector , which is used for dynamic adjustment of control variables; When entering the learning and training loop to reach the maximum number of episodes, initialize a random process N for action exploration, obtain the initial state , and calculate the initial residual ; input it into the residual perception network to obtain the residual weighted coefficient ; Generate the weighted control action: Input the state into the current policy network, and output the original action vector ; use the residual weighting mechanism to obtain the final executed action ; Send to the environment for execution; The environment feedbacks the new state , reward , new residual ; Store the state transition sample into the experience replay pool, and randomly sample samples from the experience replay pool R as the training data for the Actor network and the Critic network; Let , and update the Critic network parameters by minimizing the loss function L ; Among them, is the target value; when calculating , represents the discount factor, , the target value network and the target policy network are used so that the network can remain stable and converge quickly during training; Calculate the gradient of the policy network:

[0060] Update the target network and :

[0061]

[0062] Among them, is the update coefficient, indicating the update step size, which balances between the current network parameters and the target network parameters.

[0063] In a possible implementation manner of the present invention, the lowest energy consumption is the optimization goal: while outputting the expected thrust, minimize the energy consumption per unit thrust; The lowest energy consumption control goal is: in different flight modes, while outputting the expected thrust, minimize the energy consumption per unit thrust; The mathematical form of the objective function is:

[0064] Among them, is the energy consumption rate at the current moment, is the combined thrust at the current moment, is the target thrust, control penalty factor.

[0065] In a specific implementation, the system constructs an energy-saving optimization model based on non-linear dynamic constraints and embeds the minimum energy consumption control target into the adjustment process of control variables under multiple flight modes. For stages such as cruise, transition, and low-altitude inspection, the system uses the energy consumption per unit thrust as the core optimization index. On the premise of ensuring that the thrust output meets the mission requirements, it dynamically adjusts the combination of control variables such as the electric propulsion power, the opening of the tail nozzle, the angle of the tilting mechanism, and the deflection angle of the nozzle. The objective function is designed as:

[0066] where the control penalty factor is jointly calibrated through offline simulation and online reinforcement learning and is used to balance the priority of thrust tracking accuracy and energy efficiency optimization. For example, it is reduced during the cruise stage to preferentially reduce energy consumption, while it is increased during an emergency climb to ensure a rapid thrust response.

[0067] To achieve the above objectives, the control device iteratively solves the constrained optimization problem through the RE-DEORL algorithm framework: First, the flight state vector (including the thrust error , the energy consumption rate , the altitude and the speed ) is input into the policy network to generate a preliminary control action; subsequently, the constraint verification module continuously checks whether the thruster power exceeds the limit (such as the maximum power of the electric propulsion module), whether the attitude angle change rate exceeds the safety threshold (to prevent instability caused by sudden pitch angle changes), and whether the nozzle area and voltage are within the physically feasible range. If the action violates the constraint, a correction mechanism is triggered - for example, when the allocated value of the electric propulsion power exceeds the upper limit, the power is scaled proportionally and the nozzle opening is compensated to maintain the thrust; if the attitude angle change rate is too high, a damping signal is injected through the fuzzy control module for smooth adjustment.

[0068] During the optimization process, the residual perception network dynamically captures the deviation characteristics between the actual thrust and the target value and outputs a residual weighted vector, which is used to adjust the gain amplitude of the main RE-DEORL control action and enhance the robustness to non-linear disturbances (such as gust interference). At the same time, the experience replay pool continuously stores state-action-reward samples, and the Actor-Critic network parameters are updated through periodic sampling training, gradually approaching the optimal control strategy. For example, during the horizontal cruise stage, the algorithm tends to reduce the opening of the turbofan nozzle to reduce air resistance, and at the same time moderately increase the proportion of the electric propulsion power to utilize the efficient range of electric energy; while during vertical landing, the aerodynamic trim is optimized by increasing the tilting angle to reduce the energy consumption of attitude adjustment.

[0069] Please refer to Figure 5 , the second embodiment of the present invention provides a multi-modal intelligent propulsion control device for a short takeoff and vertical landing aircraft, including: The flight mode classification sheet 201 is used to obtain flight status data in real time, and classify the flight status data according to the status classification model of the decision tree to generate the current flight mode; The final control quantity generation unit 202 is used to call the control strategy group and, with the lowest energy consumption as the optimization goal, perform parallel processing on the current flight mode to respectively generate a plurality of control instructions, and call the fusion weight network to process the current flight mode to generate the fusion weight of each control instruction, and generate the final control quantity according to the plurality of control instructions and the fusion weight. Among them, the fusion weight of each control instruction can be dynamically adjusted based on the flight mission stage, system performance feedback, and reward and punishment mechanism. The control strategy group includes a RE-DEORL control module, an MPC predictive control module, and a fuzzy control module; The feedback regulation unit 203 is used to send the final control quantity to a pre-constructed non-linear model of the hybrid propulsion structure to control the short takeoff and vertical landing aircraft, and collect feedback signals to continuously optimize the strategy parameters of the RE-DEORL control module; among them, the non-linear model of the hybrid propulsion structure includes an electric propulsion system, a turbofan propulsion system, a tilt system, and an energy scheduling system.

[0070] The third embodiment of the present invention provides a short takeoff and vertical landing aircraft multi-modal intelligent propulsion control device, which is characterized by including a memory and a processor. The memory stores a computer program, and the computer program can be executed by the processor to implement a short takeoff and vertical landing aircraft multi-modal intelligent propulsion control method as described in any one of the above.

[0071] The fourth embodiment of the present invention provides a computer-readable storage medium, which is characterized by storing a computer program, and the computer program can be executed by the processor of the device where the computer-readable storage medium is located to implement a short takeoff and vertical landing aircraft multi-modal intelligent propulsion control method as described in any one of the above.

[0072] Based on the short takeoff and vertical landing aircraft multi-modal intelligent propulsion control method, device, equipment and storage medium provided by the present invention, first obtain flight status data in real time, and classify the flight status data according to the status classification model of the decision tree to generate the current flight mode; then, call the control strategy group and, with the lowest energy consumption as the optimization goal, perform parallel processing on the current flight mode to respectively generate a plurality of control instructions, and call the fusion weight network to process the current flight mode to generate the fusion weight of each control instruction, and generate the final control quantity according to the plurality of control instructions and the fusion weight, and send the final control quantity to a pre-constructed non-linear model of the hybrid propulsion structure to control the short takeoff and vertical landing aircraft, solving the problem that the short takeoff and vertical landing aircraft cannot be adaptively adjusted when there are different dynamic characteristics in multiple propulsion units.

[0073] Exemplarily, the computer programs described in the third and fourth embodiments of the present invention may be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in implementing a short takeoff and vertical landing aircraft multi-modal intelligent propulsion control device. For example, the device described in the second embodiment of the present invention.

[0074] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the short takeoff and vertical landing aircraft multi-modal intelligent propulsion control method, and uses various interfaces and lines to connect the entire implementation of various parts of a short takeoff and vertical landing aircraft multi-modal intelligent propulsion control method.

[0075] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory, the processor realizes various functions of a short takeoff and vertical landing aircraft multi-modal intelligent propulsion control method. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, a text conversion function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, text message data, etc.), etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0076] Among them, if the implemented module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0077] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative effort.

[0078] As described above, the above are only the preferred specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multi-mode intelligent propulsion control method for a short-to-vertical aircraft, characterized in that: include: Acquire flight status data in real time, and classify the flight status data according to a state classification model of a decision tree to generate a current flight mode; Calling a control strategy group and taking minimum energy consumption as an optimization goal, processing the current flight mode in parallel to generate a plurality of control instructions respectively, and calling a fusion device weight network to process the current flight mode to generate a fusion weight of each control instruction, and generating a final control amount according to the plurality of control instructions and the fusion weight, wherein the fusion weight of each control instruction can be dynamically adjusted based on the flight mission stage, system performance feedback, and reward and punishment mechanism, and the control strategy group includes a RE-DEORL control module, an MPC prediction control module, and a fuzzy control module; The final control quantity is sent to a pre-built hybrid propulsion structural nonlinear model to control the short-to-vertical aircraft, and feedback signals are collected to continuously optimize the strategy parameters of the RE-DEORL control module; wherein the hybrid propulsion structural nonlinear model includes an electric propulsion system, a turbofan propulsion system, a tiltrotor system, and an energy scheduling system.

2. The multi-mode intelligent propulsion control method for short-to-vertical aircraft according to claim 1 is characterized in that: The current flight mode includes vertical take-off, climb, transition, horizontal cruise, hover, descent and vertical landing.

3. The multi-mode intelligent propulsion control method for a short-to-vertical aircraft according to claim 1 is characterized in that: The control variables of the hybrid propulsion structure nonlinear model are: in, are the output power percentages of the two independent electric propulsion modules, is the deflection angle of the turbofan nozzle, is the tilt mechanism angle, is the tail nozzle opening of the turbofan; The state variables of the nonlinear model of the hybrid propulsion structure are: in, is the synthetic thrust of the propulsion system at the current moment, is the energy consumption per unit thrust, is the aircraft altitude, is the forward speed of the aircraft.

4. The multi-mode intelligent propulsion control method for short-to-vertical aircraft according to claim 1 is characterized in that: The RE-DEORL control module is configured to receive and process the current flight state vector, the residual weighted vector output by the residual perception network, and the historical interaction data sampled in the experience replay pool, and output the main control action vector and the residual weighted vector; The MPC prediction control module is configured to receive and process the current flight state vector, the linear time-varying model parameters, the reference trajectory, the weight matrix and the prediction step number, and output a prediction control sequence, and use the first control amount of the control sequence as a control instruction; The fuzzy control module is configured to receive and process the current flight state vector and a pre-constructed fuzzy rule base, and output the original control quantity after fuzzy reasoning and a correction signal for compensating for sudden disturbances.

5. The multi-mode intelligent propulsion control method for a short-to-vertical aircraft according to claim 4 is characterized in that: The execution process of the RE-DEORL control module is: Randomly initialize the current Actor network The weight parameter , and the current critic network The weight parameter ; Initialize the target Actor network and target critic network , their respective network weight parameters are: , ; Initialize the experience replay pool R to store state transfer and control residual information; Residual perception network initialization: Construct a CNN + Transformer combined network to extract disturbance features and time dependencies in multi-channel state residual sequences; the network input is the residual sequence in the current state ,in, For the estimated state, For the real state; The CNN module extracts the local perturbation peak, and the Transformer module models the evolution of the residual over time and channels; the network output is the residual weighted vector , used to control the dynamic adjustment of variables; When the learning training cycle reaches the maximum round, a random process N is initialized for action exploration to obtain the initial state , and calculate the initial residual ; Input to the residual perception network to obtain the residual weighted coefficient ; Generate weighted control actions: transform the state Input the current policy network and output the original action vector ; Use the residual weighting mechanism to get the final execution action ;Will Send to the environment for execution; Environmental feedback new status ,award , new residual ; Transfer the state sample Store to the experience replay pool, randomly sample from the experience replay pool R Samples are used as training data for the Actor network and the Critic network; make , update the Critic network parameters by minimizing the loss function L ; in, For the goal value; in calculation hour, represents the discount factor, , using the target value network and target strategy network , so that The network can remain stable and converge quickly during training; Compute the gradient of the policy network: Update target network and : in, is the update coefficient, which indicates the update step size and balances the current network parameters with the target network parameters.

6. The multi-mode intelligent propulsion control method for a short-to-vertical aircraft according to claim 3 is characterized in that: The minimum energy consumption is the optimization goal: while outputting the desired thrust, the unit thrust energy consumption is minimized; The minimum energy consumption control target is: to minimize the unit thrust energy consumption while outputting the expected thrust in different flight modes; The mathematical form of the objective function is: in, is the energy consumption rate at the current moment, is the synthetic thrust at the current moment, Target thrust, Control the penalty factor.

7. A multi-mode intelligent propulsion control device for a short-to-vertical aircraft, characterized in that: include: A flight mode classification unit, used to obtain flight status data in real time, and classify the flight status data according to a state classification model of a decision tree to generate a current flight mode; A final control quantity generating unit is used to call the control strategy group and take the minimum energy consumption as the optimization goal, perform parallel processing on the current flight mode to generate multiple control instructions respectively, and call the fusion device weight network to process the current flight mode to generate a fusion weight of each control instruction, and generate a final control quantity according to the multiple control instructions and the fusion weight, wherein the fusion weight of each control instruction can be dynamically adjusted based on the flight mission stage, system performance feedback, and reward and punishment mechanism, and the control strategy group includes a RE-DEORL control module, an MPC prediction control module, and a fuzzy control module; A feedback regulation unit is used to send the final control amount to a pre-built hybrid propulsion structural nonlinear model to control the short-to-vertical aircraft, and collect feedback signals to continuously optimize the strategy parameters of the RE-DEORL control module; wherein the hybrid propulsion structural nonlinear model includes an electric propulsion system, a turbofan propulsion system, a tiltrotor system, and an energy scheduling system.

8. A multi-mode intelligent propulsion control device for short-to-vertical aircraft, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a multi-modal intelligent propulsion control method for a short-to-vertical aircraft as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: A computer program is stored, and the computer program can be executed by a processor of the device where the computer-readable storage medium is located to implement a multi-modal intelligent propulsion control method for a short-to-vertical aircraft as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Optimal control method for acceleration process of variable cycle engine based on computer

    CN116974194A

  • Aircraft intelligent controller training method based on deep reinforcement learning

    CN118012126A

  • Aircrew Automation System and Method

    US20170277185A1

Cited By

  • Sensor-fault-resistant multi-mode propulsion control method for low-altitude aircraft

    CN120161778A

  • A multi-mode propulsion control method for low-altitude aircraft resistant to sensor failure

    CN120161778B

  • Vertical take-off and landing aircraft self-adaptive control method and system based on distributed electric thrust array

    CN121455172A

  • Adaptive control method and system for vertical take-off and landing aircraft based on distributed electric thrust array

    CN121455172B

  • Long endurance optimization control method and system for vertical take-off and landing fixed-wing unmanned aerial vehicle

    CN121559877A