A reference speed scheduling method and device for acceleration and deceleration of an aero-engine
By using a proxy dynamics model and reinforcement learning, a reference speed sequence that satisfies multiple engineering constraints is generated, which solves the problem of speed and smoothness during the acceleration and deceleration of aero engines. This achieves fast and smooth acceleration and deceleration transitions and consistency across operating conditions, thereby improving system safety and airworthiness verifiability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF ENGINEERING THERMOPHYSICS - CHINESE ACAD OF SCI
- Filing Date
- 2026-02-06
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies struggle to simultaneously balance multiple engineering constraints (turbine inlet temperature, surge margin, speed, fuel flow) and dynamic performance (speed, smoothness, consistency across operating conditions) during the acceleration and deceleration of aero engines. This is especially true under transient conditions with high safety levels and multiple coupled constraints, where it is difficult to achieve a fast, smooth transition with adaptive capabilities.
By employing a proxy dynamics model and reinforcement learning approach, a reference speed sequence that satisfies multiple engineering constraints is generated by training the engine's transient state transition relationship. Combined with fuel action management, the safety and executability of the output command are ensured. Finally, the inner loop tracking controller uses fuel flow as the actuator to complete the acceleration and deceleration transition.
It achieves reduced acceleration and deceleration time, reduced fuel command variation variance, suppressed oscillations in fuel circuits and speed channels, improved system safety and airworthiness verifiability, and maintained performance consistency across the entire flight envelope, all while using only fuel flow as the actuator. It also features low computational overhead and ease of deployment.
Smart Images

Figure CN122151979A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aero-engine control technology, and relates to a reference speed scheduling method and device for aero-engine acceleration and deceleration based on a data-driven agent model and reinforcement learning. Background Technology
[0002] During acceleration and deceleration transitions, aero engines need to achieve a rapid and smooth response while meeting multiple engineering constraints. These engineering constraints typically include turbine inlet temperature (TNT). Upper limit, surge margin (SM) lower limit, speed (e.g., high-voltage speed) ) upper limit and fuel flow ( () Amplitude and rate of change limits.
[0003] In engineering practice, existing technologies typically employ rule-based minimum-maximum (MIN–MAX) scheduling and fixed-gain control strategies. These methods are simple to implement and offer good predictability, but they tend to be conservative in ensuring safety boundaries, which can easily lead to problems such as long transition times, jagged fuel trajectories, and inconsistent performance under different operating conditions.
[0004] To address multi-constraint coupled problems, some solutions employ model predictive control (MPC) or reference governers. MPC can explicitly consider constraints within a unified framework, but it relies on accurate dynamic models and real-time optimization solutions. Under wide-envelope conditions, it may face challenges such as model mismatch, complex parameter tuning, high real-time computational resource requirements, and significant engineering verification difficulties. Reference governers and action limiters can suppress out-of-bounds risks, but under conditions of significant variable coupling and tight constraints, they typically require sacrificing some dynamic performance.
[0005] Furthermore, with the accumulation of test / bench data and improved computing capabilities, data-driven modeling and optimization have also been applied to engine transient control. For example, historical data is used to construct surrogate models and combine them with optimization strategies to generate transitional reference trajectories. Such methods alleviate the problem of obtaining accurate mechanistic models to some extent, but there are still engineering application barriers in terms of cross-condition generalization ability, verifiability of safety constraints, robustness to sensor noise and measurement hysteresis, and achieving a trade-off between speed and smoothness using only a single fuel actuator.
[0006] In summary, existing technologies, based solely on fuel flow rate... As an actuator, its ability to simultaneously balance safety constraints, speed, and fuel smoothness still needs improvement. In particular, under transient conditions with high safety levels and multiple constraints, how to achieve a rapid and smooth transition during acceleration / deceleration and possess adaptive capability across the entire flight envelope remains a critical issue that urgently needs to be addressed. Summary of the Invention
[0007] To address the challenges of acceleration and deceleration in aircraft engines, where fuel flow alone is the primary factor... Given the limitations of actuators, it is difficult to simultaneously address multiple engineering constraints (including upper limits for turbine inlet temperature, lower limits for surge margin, upper limits for high-pressure speed, upper limits for low-pressure speed, and limits on fuel flow amplitude and rate of change) and dynamic performance (such as short transition time, smooth fuel trajectory, and good consistency across operating conditions). This invention provides a reference speed scheduling method and device for single actuators. This method achieves efficient, safe, and deployable acceleration and deceleration control under multiple constraints by integrating high-fidelity modeling, safe reinforcement learning, and online governance mechanisms.
[0008] Specifically, the method includes the following steps:
[0009] S1. Based on engine test or bench timing data, train a proxy dynamics model to learn the transient state transition relationship of the engine; S2. In the simulation environment formed by the trained surrogate dynamics model, the original control output is generated by optimizing the engine state variables and environmental parameters as inputs and the fuel flow rate as the action through reinforcement learning strategy. The optimization process is based on a reward function and satisfies engineering constraints, including upper limit of turbine inlet temperature, lower limit of surge margin, upper limit of engine speed, and limits on fuel flow amplitude and rate of change. The original control output includes the original reference speed sequence and / or the original fuel action sequence, or a combination of both. S3. Perform a treatment process on the original reference speed sequence and / or original fuel action sequence in the original control output to obtain a reference speed sequence and / or fuel action sequence. The treatment process is used to ensure that the output meets the engineering feasibility and actuator dynamic limits. S4. The reference speed sequence is input as a setpoint to the inner loop tracking controller, which uses only fuel flow as the actuator to drive the engine to complete the acceleration and deceleration transition process.
[0010] Furthermore, in step S1, the surrogate dynamics model is a recurrent neural network or an encoder-type Transformer model.
[0011] Wherein, when the surrogate dynamics model is a recurrent neural network, the recurrent neural network includes a long short-term memory network LSTM and / or a gated recurrent unit GRU, and uses a sliding time window (length of 20 to 60 steps) to model the engine's time-series state-action data; When the surrogate dynamics model is an encoder-type Transformer model, a multi-head self-attention mechanism is used to process the engine's temporal state-action data in parallel, and training is performed based on multi-step prediction loss.
[0012] Furthermore, in step S1, the surrogate dynamics model is trained to learn the transient state transition relationships of the engine, including: S11. During the training phase, the engine test or bench time series data is constructed into a state-action time series, and the future state of the recurrent neural network is predicted by a multi-step autoregressive method, or a multi-head self-attention mechanism is used to perform multi-step prediction on the encoder-type Transformer model, and the loss function is constructed by the prediction error. S12. During the training process, the recurrent neural network adopts the Scheduled Sampling strategy, gradually annealing the teacher forcing probability from 1.0 to a preset lower limit (such as 0.5) to alleviate exposure bias and suppress error accumulation in free-rolling prediction; the encoder-type Transformer model adopts the Teacher Forcing strategy to gradually reduce the teacher forcing ratio to improve the model's generalization ability. S13. Perform global norm pruning on the network gradient of the recurrent neural network or the encoder-type Transformer model, and perform regularization in combination with L2 weight decay. S14. During the inference phase, apply a hard clipping of the state increment output by the recurrent neural network or the encoder-type Transformer model, or map the state increment to a preset physical boundary using the tanh function.
[0013] Further, in step S2, the reinforcement learning policy optimization adopts any one of the following algorithms: soft actor-commenter SAC, deep deterministic policy gradient DDPG, double-delay deep deterministic policy gradient TD3, proximal policy optimization PPO, and model-based policy optimization MBPO. The reward function includes a speed tracking error term, a fuel smoothing penalty term, and constraint penalty terms corresponding to turbocharger temperature exceeding the limit, insufficient surge margin, speed exceeding the limit, and fuel change rate exceeding the limit, respectively. The weight of each term ranges from 0.1 to 10.
[0014] Further, in step S2, the state variables include high-pressure speed and / or low-pressure speed, and the environmental parameters include flight altitude H and Mach number Ma, or equivalent inlet total temperature and equivalent inlet total pressure calculated from flight altitude H and Mach number Ma; The upper limit of turbine inlet temperature and / or the lower limit of surge margin in the engineering constraints are estimated online by the surrogate dynamics model or the additional constraint estimation model. The estimation results are used to construct the reward function and serve as the basis for judgment on governance treatment.
[0015] Further, in step S3, the treatment process includes: S31. Limit the slope and / or acceleration of the reference rotational speed within a preset range; S32. The amplitude projection of the fuel action is applied to constrain it within the effective fuel flow range, and within the control cycle (10ms to 50ms) of the inner loop tracking controller, the single-step change is limited to no more than 0.5% to 3% of the full scale (FS) of the fuel flow. S33. The continuous reference trajectory is discretized into a discrete reference trajectory that can be distributed by using a piecewise interpolation method.
[0016] Furthermore, in step S4, the inner loop tracking controller is a proportional-integral-derivative PID controller or a model predictive controller (MPC), with a control cycle of 10ms to 50ms. When the original control output is the original fuel action sequence, an equivalent reference speed is generated through the setpoint filtering and limiting mechanism of the inner loop tracking controller, and closed-loop tracking is achieved.
[0017] In an improved embodiment of the above-described reference speed scheduling method for aero-engine acceleration and deceleration, the method further includes: S5. The reference speed sequence is fixed to the engine control unit in the form of a lookup table, polynomial fitting, or spline function. The input variables of the lookup table or spline function include environmental parameters and the desired steady-state target speed.
[0018] In an improved embodiment of the above-described reference speed scheduling method for aero-engine acceleration and deceleration, the method further includes: S6. Rolling predictions are made for multiple future control cycles using the trained surrogate dynamics model. If the prediction results do not meet any engineering constraints, the reference slope / acceleration is adaptively tightened, the reference trajectory is rolled back, and a safe and conservative baseline scheduling strategy is smoothly switched.
[0019] This invention also provides a reference speed scheduling device for acceleration and deceleration of an aero-engine. The device uses only fuel flow as the execution input and includes a proxy model training module, a strategy optimization module, a governance module, and a trajectory publishing module.
[0020] Specifically, the surrogate model training module is used to train a surrogate dynamics model based on engine test or bench time series data to learn the transient state transition relationships of the engine. The strategy optimization module is used to optimize and generate the original control output in a simulation environment composed of the trained agent dynamics model, with the engine's state variables and environmental parameters as inputs and fuel flow as the action, through reinforcement learning strategy. The optimization process is based on a reward function and satisfies engineering constraints, including upper limit of turbine inlet temperature, lower limit of surge margin, upper limit of engine speed, and limits on fuel flow amplitude and rate of change. The original control output includes the original reference speed sequence and / or the original fuel action sequence, or a combination of both. The treatment module is used to treat the original reference speed sequence and / or the original fuel action sequence in the original control output to obtain the reference speed sequence and / or fuel action sequence. The treatment is used to ensure that the output meets the engineering feasibility and actuator dynamic limits. The trajectory publishing module is used to input the reference speed sequence as a setpoint to the inner loop tracking controller, which uses only fuel flow as the actuator to drive the engine to complete the acceleration and deceleration transition process.
[0021] The method of this invention trains a proxy dynamics model based on engine test or bench time-series data to approximate its transient state transition relationship. In this differentiable simulation environment, a reference speed time series that satisfies the above-mentioned engineering constraints is generated through reinforcement learning strategy. Combined with the treatment of reference speed and fuel action (including slope / acceleration limiting, amplitude and rate of change projection, etc.), the safety and executability of the output command are ensured. Finally, the inner loop tracking controller completes the acceleration / deceleration transition process only with fuel flow as the actuator.
[0022] Compared with the prior art, the method of the present invention has at least the following advantages: 1. Balancing speed and smoothness: By optimizing the reinforcement learning strategy on a differentiable surrogate dynamics model and combining it with the control of reference speed and fuel action, the system achieves smooth operation even when using only fuel flow rate. Under the condition of being an actuator, it effectively shortens the rise time and settling time of the acceleration and deceleration process (for example, within the improved range of 0.5~2.0s), while significantly reducing the variation variance of fuel command and suppressing the oscillation of the oil circuit and speed channel.
[0023] 2. Controllable risk of exceeding limits: Turbine inlet temperature is explicitly embedded in both the strategy training and online execution phases. The system imposes multiple engineering constraints, including upper limits for the surge margin (SM), lower limits for the speed limit, and limits for the amplitude and rate of change of fuel flow. When the rolling forecast identifies a potential constraint violation, it can adaptively back off and smoothly switch to a safe and conservative baseline scheduling strategy, achieving zero or significant reduction of out-of-bounds events and improving system safety and airworthiness verifiability.
[0024] 3. Strong consistency and portability across operating conditions: By using environmental parameters such as flight altitude (H) and Mach number (Ma) as conditional variables, the reference trajectory maintains good performance consistency across the entire flight envelope; the surrogate models (including LSTM, GRU, or encoder-type Transformer) and reinforcement learning algorithms (including SAC, DDPG, TD3, or MBPO) are modular, supporting flexible replacement and combination, facilitating cross-platform migration and subsequent upgrades.
[0025] 4. Simple and efficient engineering implementation: It relies on only a single fuel flow actuator and can be directly integrated with existing PID or MPC inner loop controllers; the overall method supports control cycles of 10~50ms, meeting the real-time requirements of aero engines; the governance module adopts lightweight operators such as slope limiting and amplitude projection, with low computational overhead, and is easy to deploy in resource-constrained engine control units (ECUs).
[0026] 5. Good verifiability and scalability: Constraints and performance indicators are expressed in quantitative form during the training and deployment phases, facilitating verification and parameter tuning via ground bench testing; without changing the core architecture of this invention, the compressor outlet pressure can be introduced as an extension. State variables such as thrust are used to further enhance the system's robustness and applicability. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart of the reference speed scheduling method for a single actuator according to the present invention; Figure 2 This is a schematic diagram of the surrogate dynamics model; Figure 3 Flowchart for optimizing learning strategies; Figure 4 A logical flowchart for governance and processing; Figure 5 This is an architecture diagram of a reference speed scheduling device for a single actuator. Figure 6 This is a schematic diagram of the hardware structure of an electronic device; The components are as follows: 100, Data Preprocessing Module; 110, Agent Model Training Module; 120, Policy Optimization Module; 130, Governance Module; 140, Trajectory Publishing Module; 150, Inner Loop Tracking Controller; 180, Engineering Constraint Input; 201, Input Sequence; 202, Output Sequence; 211, Recursive Network Layer; 212, Readout Layer; 301, Policy Network; 302, Value Network; 303, Environment Simulation; 401, Reference Limiter; 402, Action Projector; 403, Trajectory Interpolator; 501, Processor; 502, Memory; 503, Communication Interface; 504, System Bus. Detailed Implementation
[0029] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0030] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features of the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] Furthermore, any modifications made by those skilled in the art to parameter values, module division, network structure, and algorithm replacements (such as LSTM / GRU / encoder-type Transformer; SAC / DDPG / TD3 / MBPO) without departing from the spirit of this invention fall within the protection scope of this invention. The terms "preferredly" and "furthermore" in this document refer to preferred embodiments.
[0032] It should also be noted that in the embodiments of the present invention, sometimes the subscript (such as W1) may be mistakenly written as a non-subscript form (such as W1). When the difference is not emphasized, the meaning they express is the same.
[0033] like Figure 2 The diagram shows the structure of the surrogate dynamics model, with the input sequence 201 (including high pressure and rotational speed) shown. The connection relationship between the recursive network layer 211 (LSTM / GRU) and the readout layer 212, and the output sequence 202; like Figure 3The diagram shown is an optimization flowchart of the reinforcement learning strategy. It illustrates the data flow between the policy network 301, the value network 302, and the environment simulation 303, as well as the relationship between the engineering constraint input 180 and the reward and sampling. Figure 4 The diagram is a logical block diagram of the governance process, showing the cascade relationship and input / output of the reference limiter 401, motion projector 402 and trajectory interpolator 403. This invention provides a reference speed scheduling method for a single actuator. This method achieves efficient, safe, and deployable acceleration and deceleration control under multiple constraints by integrating high-fidelity modeling, safe reinforcement learning, and online governance mechanisms.
[0034] Specifically, see Figure 1 As shown, the method includes the following steps: S1. Based on engine test or bench timing data, train a proxy dynamics model to learn the transient state transition relationship of the engine; S2. In the simulation environment formed by the trained surrogate dynamics model, the original control output is generated by optimizing the engine state variables and environmental parameters as inputs and the fuel flow rate as the action through reinforcement learning strategy. The optimization process is based on a reward function and satisfies engineering constraints, including upper limit of turbine inlet temperature, lower limit of surge margin, upper limit of engine speed, and limits on fuel flow amplitude and rate of change. The original control output includes the original reference speed sequence and / or the original fuel action sequence, or a combination of both. S3. Perform a treatment process on the original reference speed sequence and / or original fuel action sequence in the original control output to obtain a reference speed sequence and / or fuel action sequence. The treatment process is used to ensure that the output meets the engineering feasibility and actuator dynamic limits. S4. The reference speed sequence is input as a setpoint to the inner loop tracking controller, which uses only fuel flow as the actuator to drive the engine to complete the acceleration and deceleration transition process.
[0035] The core of the method of this invention is the use of a proxy dynamics model. In this differentiable environment, reinforcement learning is used to optimize and obtain the original reference rotational speed time series that satisfies the constraints. Combined with reference / action governance to ensure safe and smooth online execution, ultimately the inner loop controller relies solely on fuel flow... Complete the acceleration / deceleration transition.
[0036] In some embodiments of step S1, the collected engine test or bench raw time series data is preprocessed to filter key state parameters and performance variables of the acceleration / deceleration process, forming a parameter variable dataset, which is the time series data used to train the model.
[0037] Specifically, in the extracted parameter variable dataset, the key variables include at least time. t Environmental quantity (height) H ,Mach number Ma or imported total temperature Imported total pressure One or more of the following; the controlled variable (high-pressure speed) Selectable low-pressure speed ); Actuator quantity (fuel flow) ); optional performance indicators: thrust F .
[0038] To reduce model complexity and improve generalization ability, switching variables and weakly correlated auxiliary variables that are irrelevant to transients are removed. Variables are normalized (preferably 0-1 linear normalization or Z-score standardization), and the input sequence is constructed using a 20-60 step sliding time window. (where L is the window length, ), and construct differential features , wait.
[0039] Preferably, outliers are handled using IQR rules; the data is divided into training / validation / test sets in chronological order (e.g., a division ratio of 60% / 20% / 20%).
[0040] In some embodiments of step S1, such as Figure 2 As shown, a surrogate dynamics model containing time series can be constructed based on the temporal characteristics of aero-engine variables to approximate transient state transitions. ,in At least include and the environmental quantity mentioned above. .
[0041] Preferably, the surrogate dynamics model can be a recurrent neural network or an encoder-based Transformer model. When a recurrent neural network is used, it includes a Long Short-Term Memory (LSTM) network and / or a gated recurrent unit (GRU), and models the engine's temporal state-action data using a sliding time window (20 to 60 steps in length). When an encoder-based Transformer model is used, a multi-head self-attention mechanism is employed to process the engine's temporal state-action data in parallel, and training is performed based on a multi-step prediction loss.
[0042] Training employs a multi-step autoregressive loss and weight decay, where the weight decay parameter λ is located at 10. -6 Up to 10 -3 The learning rate is at 10 -4 Up to 10 -3 Batch size 64 to 256, early stopping criterion 5 to 15 rounds, to obtain the trained agent dynamics model, i.e., the microagent model. .
[0043] The training of the surrogate dynamics model to learn the transient state transition relationships of the engine includes: S11. During the training phase, the engine test or bench time series data is constructed into a state-action time series, and the future state of the recurrent neural network is predicted by a multi-step autoregressive method, or a multi-head self-attention mechanism is used to perform multi-step prediction on the encoder-type Transformer model, and the loss function is constructed by the prediction error. S12. During the training process, the recurrent neural network adopts the Scheduled Sampling strategy, gradually annealing the teacher forcing probability from 1.0 to a preset lower limit (such as 0.5) to alleviate exposure bias and suppress error accumulation in free-rolling prediction; the encoder-type Transformer model adopts the Teacher Forcing strategy to gradually reduce the teacher forcing ratio to improve the model's generalization ability. S13. Perform global norm pruning on the network gradient of the recurrent neural network or the encoder-type Transformer model, and perform regularization in combination with L2 weight decay. S14. During the inference phase, apply a hard clipping of the state increment output by the recurrent neural network or the encoder-type Transformer model, or map the state increment to a preset physical boundary using the tanh function.
[0044] More specifically, the process of establishing and training the surrogate dynamics model includes the following steps S101~S104: S101, Normalization and Sample Construction: S1011 aligns the data from each sensor channel according to a unified time reference and resamples them to a fixed sampling period. (Preferred range: 10~50ms). When there is slight jitter in the original channel timestamp, linear interpolation or zero-order hold is used for reconstruction; when there is a misalignment in the start and end times of multi-source records, the shortest common time period is used as the effective interval.
[0045] S1012, based on the correlation between acceleration / deceleration transients, determine the state vector and action vector: Status: At least contains and environmental quantity ( H , Ma or , One or more of the following; optional addition or thrust F ; action: .
[0046] The units of measurement for each variable were standardized: rotational speed was measured in rpm, and temperature and pressure were measured in SI units, and this standardization was maintained throughout the text.
[0047] S1013, Missing Tests and Outlier Handling: Imputation is performed on obvious missing test points, with linear interpolation preferred. Samples are removed when the missing test percentage within a window exceeds a 10% threshold. Outliers are handled using IQR or... Rules are truncated or eliminated, and the outlier threshold is calculated on the training set and used consistently during the verification / testing phase.
[0048] S1014, Denoising and Delay Alignment: Suppress high-frequency noise without altering dynamic characteristics, preferably using first-order low-pass or median filtering. When a fixed delay is known in the sensing link (e.g., response lag in the fuel-speed channel), coarse alignment can be performed using integer sampling shift; otherwise, the corresponding delay is automatically absorbed by subsequent model learning.
[0049] S1015, Normalization parameter calculation: For each continuous variable on the training set v Calculate the normalization parameters and fix them in subsequent stages: 0-1 normalization: ; Or, Z-score: ; The / or / The normalized parameters are persistently stored and applied to the validation set, test set, and online phase. The parameters must not be re-estimated on the validation / test set to avoid information leakage.
[0050] S1016, Construct a sliding window sample: Set the window length. stride (Unit: sampling points), construct time series samples: ; in, For the normalized state, for The normalized value, For multi-step horizon prediction; when Discard the sample if it exceeds the valid range.
[0051] S1017, Constructing Derived and Differential Features: To enhance short-term dynamic representation, derived features are added within the window. , ; And can add the previous limit The original values or rolling mean / variance statistics before normalization are used; derived features use the same normalization rules as the original variables.
[0052] S1018, Temporal Partitioning and Class Balance: The samples are divided into training / validation / test sets in chronological order (e.g., 60% / 20% / 20%) to avoid cross-set mixing. Preferably, the training set maintains a rough balance between acceleration and deceleration samples to improve the model's generalization ability.
[0053] S1019, Metadata Persistence and Consistency Verification: The following information will be persisted: normalized parameter set, list of selected variables, and window length. L stride s Prediction time domain K Sampling period Derived feature configuration and data segmentation boundaries. Single-step and multi-step RMSE are verified by replaying under typical working conditions to ensure that the accuracy threshold for entering the S22 / S23 training stage is met (e.g., multi-step RMSE relative threshold, out-of-bounds statistics do not deteriorate).
[0054] S102, Recurrent Neural Networks and Structure Design: S1021, Input / Output Tensor Definition: Represent the sliding window sample constructed in S101 as follows: Input sequence: , (in One-dimensional ); Target sequence: ; The "~" indicates that it has been normalized according to S101. For the selected state dimension (containing at least) Optional and environmental quantities).
[0055] S1022, Feature Embedding and Concatenation: To reduce dimensional differences and improve representational capability, linear embedding is performed on state and action respectively. , ; in, , , This forms a time-level input: ; Optionally, a derived feature of S1017 can be added ( , (etc.) enter .
[0056] S1023, Constructing a sequence encoder (LSTM main body): Using a 1~3 layer LSTM encoder, with hidden units and optional dropout between layers (0~0.3): , , ; Get the top hidden state To suppress offset drift, it is preferable to use a layer-normalized LSTM or apply LayerNorm to the hidden states.
[0057] S1024, Readout Head and Incremental Modeling: Using linear readout or a small MLP, the hidden state is mapped to a state increment, which is then added to the state from the previous step to obtain the prediction. , ; in, It is preferable to use "incremental prediction" instead of directly predicting absolute values to enhance stability.
[0058] S1025, Multi-step Autoregressive Expansion (During Training): Based on the prediction time domain K Expanding the autoregressive model: ; for : ; ; Among them, used for calculation The input sequences are trained using Scheduled Sampling (preferred): with probability... p Select real With probability Select Preferably, p in front Within one epoch, the linear annealing is reduced from 1.0 to 0.5. Furthermore, the autoregressive expansion and Scheduled Sampling here directly support the "multi-step autoregressive loss" in S103.
[0059] S1026, Physical domain constraint: To avoid non-physical output caused by numerical extrapolation, hard limiting is applied to key components during the inference phase. ,right( Similarly); maintain no cropping during the training phase. , These are the upper and lower limits after inverse normalization. For It can be handled in the same way as other limited quantities.
[0060] Optionally, hard clipping (clamp) can be enabled only during the inference phase, while the differentiable mapping described above is used during the training phase.
[0061] The domain mapping here corresponds to "avoiding non-physical boundary crossings" in the implementation method, forming a two-layer guarantee of soft constraints and hard constraints with subsequent governance.
[0062] S1027, Parameter Initialization and Numerical Stabilization: Initialization: The embedding and readout layers use Xavier / Glorot to initialize the weights, and the LSTM forget gate bias is used. Set it to 1.0~2.0 (preferably 1.0) to alleviate long-term dependence on forgetting.
[0063] Gradient clipping (preferred): Clipping based on the global norm, with a threshold of 0.5~5.0, in conjunction with the optimizer in S23.
[0064] Activation function: The readout head is linear by default; if using a two-layer MLP for readout, the hidden layer can use ReLU or Tanh functions with a width of 64~256.
[0065] S1028, Variable-length sequence and mask: When there are missing measurements or irregular segments at the end of the window, a zero-padding + mask strategy is adopted, and the loss is calculated only for the effective time steps; the corresponding mask takes effect in the loss function of S103 to avoid gradient pollution.
[0066] S103 uses multi-step autoregressive loss training. Weights, learning rate, and batch size are as follows: S1031, Multi-step Autoregressive Loss Term: The multi-step autoregressive mean squared error (optional Huber loss) is used as the main loss term, and a discount is applied to the long-term error. , ; in, The predictions obtained from autoregression according to S102: S1032, Regularization and Optional Multitasking: Total Loss is: ; in, This is the weight decay coefficient; (Optional) Used for constraint-side output (e.g.) , SM (proxy estimation), its weight .
[0067] S1033, Optimizer and Learning Rate Strategy: Adaptive first-order optimizer (preferably Adam), initial learning rate... The learning rate can be one of the following: Cosine annealing (with 5-10 epochs of linear preheating); Multi-step decay (decay by 0.1 to 0.5 times every 20 to 50 epochs); Batch size The sequence is sampled in continuous slices over time; gradient pruning is enabled (global norm threshold 0.5–5.0).
[0068] S1034, Training Stability and Regularization: Inter-layer Dropout can be set to 0–0.3; LayerNorm / LN-LSTM can be selected to optimize the network structure. Scheduled Sampling: The annealing probability p is linearly reduced from 1.0 to 0.5 (annealing cycle). ).
[0069] S1035, Early Stop and Indicator Settings: The primary monitoring indicator is the K-step RMSE of the validation set; the early stop patience value is 5-15 rounds, and the optimal checkpoint for the validation indicator is saved. Simultaneously, the following is recorded: Single-step / multi-step RMSE, MAE ( Optional ); Free rolling error and drift trend from 10 to 20 seconds; Outliers (NaN / Inf) and the number of gradient explosions (should be 0).
[0070] S104, Training Process and Model Finalization: S1041, Training and Selection: Train under a fixed random seed until early stopping is triggered or the maximum number of epochs is reached (preferably 100-300 epochs), using the K-step RMSE on the validation set as the primary metric. After early stopping is triggered, roll back to the validation-optimal weights as the final model. And save it as the lstm_model.pth file.
[0071] S1042, Normalized parameter export (input / output separation): Based on the statistical results of S1015, the following were exported: scaler_X.pkl: Input-side normalization parameters (saves min / max or mean / std for each input variable, as well as metadata such as variable name, order, and whether differencing is performed); scaler_Y.pkl: Output-side normalization parameters (preserving the same statistics and order descriptions for each output component). These two files are used consistently during the validation / testing and online phases and should not be re-estimated to prevent information leakage.
[0072] S1043, Inference Wrapper and Interface Encapsulation: An inference wrapper is built based on scaler_X.pkl / scaler_Y.pkl and lstm_model.pth, providing a unified interface (logic as follows): 1. reset(x0): Initializes the window buffer (length) L Clear hidden state; 2. preprocess(x_t, u_t): Assemble the input vector (including derived features) , wait); Transform to the normalized space using scaler_X.pkl, and form a sequence tensor according to the embedding / splicing rules of S22; 3. forward(norm_input): Load lstm_model.pth and perform forward pass through autoregression to obtain the model. K Step normalization prediction (or increment); 4. postprocess(norm_output): Inverse transform back to the physical domain using scaler_Y.pkl; By applying differentiable domain mapping or inference period clamp, we obtain ; 5. step(x_t, u_t) -> x_{t+1} is a combination of preprocess -> forward -> postprocess, returning the prediction for the next time step; 6. jacobians(x_t, u_t): A numerical or automatic derivative of the Jacobian, used for policy optimization and sensitivity analysis.
[0073] S1044, Consistency Check and Admission Threshold: Replay History under Representative Operating Conditions (SLS and Non-SLS) Sequence, verification: K-step verification shows that the RMSE is below a preset threshold (set by the project, for example, relative calibration ≤X%). Long-term free rolling without systematic drift (mean error within the allowable range); The predicted values do not exceed the physical limits (e.g., after inverse normalization). (Within the expected range). If not, return to S103 to adjust hyperparameters or S101 / S102 to adjust input and structure.
[0074] S1045, Product Packaging and Versioning: Output as an archive package: lstm_model.pth, scaler_X.pkl, scaler_Y.pkl; meta.json (Example keys: variable names and order, window) L ,horizon K Sampling period (Whether to use incremental prediction, differential feature list, training random seed, library version, weight checksum, etc.) Semantic version numbers and hash verification are used to ensure consistency and traceability with the subsequent policy optimization environment.
[0075] In some embodiments of step S2, the reinforcement learning policy optimization adopts any one of the following algorithms: soft actor-commenter SAC, deep deterministic policy gradient DDPG, double-delay deep deterministic policy gradient TD3, proximal policy optimization PPO, and model-based policy optimization MBPO. The reward function includes a speed tracking error term, a fuel smoothing penalty term, and constraint penalty terms corresponding to turbocharger temperature exceeding the limit, insufficient surge margin, speed exceeding the limit, and fuel change rate exceeding the limit, respectively. The weight of each term ranges from 0.1 to 10.
[0076] In some embodiments of step S2, the state variables include high-pressure rotational speed and / or low-pressure rotational speed, and the environmental parameters include flight altitude H and Mach number Ma, or equivalent inlet total temperature and equivalent inlet total pressure calculated from flight altitude H and Mach number Ma; The upper limit of turbine inlet temperature and / or the lower limit of surge margin in the engineering constraints are estimated online by the surrogate dynamics model or the additional constraint estimation model. The estimation results are used to construct the reward function and serve as the basis for judgment on governance treatment.
[0077] In specific implementation, The reinforcement learning strategy is executed on the constructed simulation environment to optimize and directly generate a reference rotational speed time series that meets engineering constraints. Engineering constraints include at least the following: upper limit SM Lower limit, upper limit of high-voltage rotor speed and Upper limit and / or fuel change rate limit. The reinforcement learning algorithm optimizes the soft actor-commentator (SAC) based on its temperature coefficient. The value is between 0.05 and 0.5, and the number of training steps is between 10. 5 Up to 10 6Alternatively, DDPG, TD3, or model-based MBPO can be used. The reward function must include at least: a speed tracking error term. Fuel smoothing item (for) The state vector may include at least the penalty for the aforementioned engineering constraints, and a penalty term for the aforementioned engineering constraints, with each weight ranging from 0.1 to 10. and / or as well as H , Ma (or) one or more of the following; the action variable is .
[0078] Specifically, this step is in the surrogate dynamics model The simulation is conducted in the environment described in S101~S104 above, and is used to obtain a reference rotational speed time series that meets engineering constraints. In this embodiment of the invention, the soft actor-commentator (SAC) is preferably used for strategy optimization. For example... Figure 3 As shown, the specific steps include the following: S21, Modeling Decision Problems: S211, the acceleration / deceleration process is modeled as a discrete-time Markov decision process, with the environmental step size consistent with the control sampling period (preferably 10~50ms), and a discount factor. Use a value of 0.95 to 0.999.
[0079] S212, Define the state At least including With one or more of the environmental quantities ( H , Ma or , ); optional addition ,thrust F Or the referenced export volume (e.g.) ).
[0080] S213, Define Action Two equivalent forms: Option 1 (Preferred): Action is a reference increment Updated after bandwidth limitation ; Option 2 (equivalent): Action is fuel recommendation The rest of the process remains unchanged (corresponding to claim 13).
[0081] S214, Set the round duration to cover typical acceleration / deceleration processes (preferably 5~30s), and when a serious out-of-bounds violation occurs (such as...). Exceeding limits or SM It will terminate if it fails to continue for several steps or if the preset round duration is reached.
[0082] S22, Environmental Propulsion and Constraint Injection: S221, at each time step, first perform soft limiting (slope / acceleration) on the reference to form candidates. Simulation of the inner-loop controller (a simplified model of PID / MPC) by... Receive fuel command .
[0083] S222, for and Perform projection and cropping: ,and (Preferred sampling period: 0.5%FS~3.0%FS).
[0084] S223, enter get If a constraint estimation head is configured, then synchronous results are obtained. , SM The estimated value is used for rewards and boundary crossing determination.
[0085] S23, Reward Function and Weight Setting: The reward function consists of tracking, smoothing, and constraint penalties, and includes at least: a) Speed tracking error term (penalty) Optional ); b) Fuel flow motion smoothing item (penalty) or ); c) Constraint penalty items (for) Exceeding limits SM Insufficient speed Exceeding the RPM limit as well as / (Violations will be punished); preferably, the weights of each item are 0.1 to 10.
[0086] Optionally, an endpoint shaping item (distance to the target steady state) or a round completion reward can be added to improve convergence speed.
[0087] S24, SAC strategy optimization process: S241, the policy network (Actor) and value network (Critic) adopt a multi-layer fully connected structure (2-3 layers, hidden width 64-256, activation function using ReLU / Tanh), the value network adopts a double Q-factor target network, and soft update coefficients. Take a value of 0.005 to 0.02.
[0088] S242 employs automatic temperature control to learn the temperature coefficient. Its effective working range is 0.05 to 0.5, and the target entropy is set based on the action dimension.
[0089] S243, using the experience playback buffer (capacity 10) 5 ~10 6 Small batch updates (batch size 64~256), Adam is preferred as the optimizer, and a learning rate of 10 is preferred. -4 ~10 -3 And enable global norm gradient clipping (0.5~5.0).
[0090] S244, total training interaction steps: 10 5 ~10 6 Top 10 3 ~10 4 The first step is a warm-up phase to fill the experience pool. After each environmental interaction, 1-2 gradient updates are performed, and then... Perform a soft update on the target network.
[0091] S25, Envelope-based Training and Security Assurance (Preferred): S251, to improve cross-condition generalization capability, at the beginning of each round, the environmental quantity ( H , Ma )or Random sampling is performed, targeting a steady-state range; optionally, stratified sampling is used to ensure balanced coverage of SLS / non-SLS operating conditions.
[0092] S252 performs short look-ahead checks (1-5 steps) after policy output and before environment stepping to assess whether constraints are about to be triggered. If a risk exists, the upper limit of the action or reference slope is adaptively tightened or projected, and penalties are added to the reward. This mechanism, together with the S4 governance during runtime, forms a two-layer structure of "soft guarantees during training + hard guarantees during runtime".
[0093] S26, Strategy Freeze and Reference Generation: S261 uses average round return, constraint violation rate, mean fuel variation term, and tracking error as the main monitoring indicators; when the indicators are stable within the sliding window and the violation rate is below the threshold, the strategy parameters are frozen. .
[0094] S262, Online Generation: Runs under actual control cycle Real-time output (or (This is followed by a reference limit and an inner loop controller forming a reference limit and an inner loop controller.) It drives objects and interfaces with S4 / S5.
[0095] S263, Offline Generation and Deployment: Offline rolling strategy under representative working conditions, yielding... Sequence, and by ( H , Ma )or( , ) forms a lookup table / spline as the independent variable; it is deployed in the control unit as the input of the trajectory publishing module (corresponding to claims 9 and 12). Simultaneously, it stores the strategy weights and key hyperparameters ( Weight (e.g., action boundaries, sampling period) for consistency verification.
[0096] In some embodiments of step S3 above, the treatment process includes: S31, limit the slope and / or acceleration of the reference rotational speed within a preset range; S32, the amplitude projection of the fuel action is applied to constrain it within the effective fuel flow range, and within the control cycle (10ms to 50ms) of the inner loop tracking controller, the single-step change is limited to no more than 0.5% to 3% of the full scale (FS) of the fuel flow. S33 uses a piecewise interpolation method to discretize the continuous reference trajectory into a discrete reference trajectory that can be distributed.
[0097] Specifically, to ensure the security and feasibility of online execution, for and Implement treatment and management: The slope and / or acceleration are limited within a preset range to avoid excessively steep reference angles that could lead to control saturation; Projected to interval and restrictions ,in The preferred sampling period is 0.5%FS to 3.0%FS; the reference trajectory is formed into a discrete sequence for downward distribution using piecewise splines or piecewise linear interpolation. More specifically, such as Figure 4 As shown, the specific steps include the following: S301, Parameters and Boundary Loading: S3011, Read governance parameters and engineering boundaries, including: reference slope Upper limit and / or acceleration Upper limit (units and) Consistent), upper / lower limits of fuel intensity Fuel change rate limit (Preferred to be 0.5%FS~3.0%FS / sampling period), online control of sampling period (Preferred 10~50ms).
[0098] S3012, optionally, the above parameters can be adjusted according to environmental parameters ( H , Ma or , Alternatively, the operating condition can be conditionally selected within a preset range (different thresholds are used for different operating conditions within the envelope).
[0099] S302, Reference Limiting and Smoothing (Reference Limiter 401): S3021, denote the original reference output by S2 as... Apply slope limiting to it: ; S3022, Optionally, implement acceleration limiting to suppress "mutations" (reference): ; in, For reference after acceleration control, when acceleration limiting is not enabled, let .
[0100] S3023, optionally, a first-order slope filter / suppression (such as a first-order low-pass filter) is added to further improve smoothness; the filter time constant is tuned according to the test bench.
[0101] S303, Fuel Action Safety Projection (Action Projector 402). S3031, regarding the fuel command calculated by the inner loop controller Implement amplitude projection: ; S3032, imposes limits on fuel change rate: ; in, The preferred sampling period is 0.5%FS to 3.0%FS.
[0102] S3033 can optionally output a saturation flag and a rate-of-change-limit flag, allowing the inner-loop controller to perform integral anti-saturation or feedforward compensation.
[0103] S304, Track interpolation and time alignment (track interpolator 403) S3041, the reference after treatment Its timestamp is aligned with the online sampling period This forms a discrete sequence at equal time points.
[0104] S3042, Interpolation method: Option A (preferred): Piecewise spline interpolation (maintaining first-order continuity, preferably shape-preserving splines), which can limit the slope within the segment to prevent "overshoot"; Option B: Piecewise linear interpolation, which is simple to implement and has low computational cost.
[0105] S3043 performs boundary trimming on the obtained discrete sequence to ensure that the entire sequence satisfies the slope / acceleration boundary of S42. Resampling within segments may be necessary to improve synchronization with the inner loop.
[0106] S305, Forward-looking verification and minimum guarantee switching (optional): S3051, based on a differentiable proxy model Perform a finite-step lookahead check (preferably 1-5 steps) to determine the current... and Is it possible to trigger engineering constraints (e.g.) upper limit SM (Lower limit, upper speed limit).
[0107] S3052, if a potential violation is predicted, one or a combination of the following priorities shall be executed: a) Adaptive tightening: Temporarily reduce and / or ; b) Reference rollback: Take a step up and then back a certain percentage; c) Safety switchover: Switch to a safety reference for minimum-maximum scheduling or fixed gain scheduling until the risk is eliminated.
[0108] S3053 outputs the look-ahead results (whether to tighten / roll back / switch) as a flag for use in logging and health monitoring.
[0109] S306, Output Package and Interface: S3061, Generate a reference distribution package for the trajectory publishing module: containing discrete reference sequences. Governance parameter snapshot Boundaries and flags (amplitude / rate of change saturation, look-ahead trigger, etc.).
[0110] S3062 can optionally retain a short-term buffer (preferably 3 to 20 points) to cope with communication jitter; when the new reference does not arrive in time, a degradation chain of "freeze-linear extrapolation-backup switch" is executed.
[0111] In some embodiments of step S4 above, the inner loop tracking controller is a proportional-integral-derivative PID controller or a model predictive controller (MPC), and its control cycle is 10ms to 50ms. When the original control output is the original fuel action sequence, an equivalent reference speed is generated through the setpoint filtering and limiting mechanism of the inner loop tracking controller, and closed-loop tracking is achieved.
[0112] Specifically, Online tracking is achieved by inputting an inner-loop tracking controller (PID and / or MPC), using only... To facilitate acceleration / deceleration transitions for the actuator; the optimal sampling period is 10-50 ms. The reference trajectory can be deployed in the engine control unit using lookup tables or polynomial / spline forms, and through environmental parameters (…). H , Ma or , Conditionalization enables envelope-based applications. When an engineering constraint violation is predicted to occur within a finite look-ahead time domain, a safety margin strategy (such as min-max scheduling or fixed-gain scheduling) can be switched to ensure a safety boundary.
[0113] More specifically, the reference speed sequence after treatment Input to the inner-loop tracking controller (such as PID or MPC), based solely on fuel flow. To achieve online tracking and evaluation of engine acceleration / deceleration under actuator conditions, the following sub-steps are specifically included: S41, Interface Initialization and Parameter Download: S411, Set the online control sampling period (Preferred time: 10~50ms) Synchronously load the snapshot of governance parameters output by S3. With boundary .
[0114] S412 establishes a data channel between the trajectory publishing module and the inner loop controller, and completes timestamp alignment and clock synchronization; sets a degradation chain of "freeze-linear extrapolation-safety switching" for missing input frames (consistent with S362).
[0115] S42, Reference reception and time alignment (executed by trajectory publishing module 140): S421, Receive and verify the discrete reference sequence according to Alignment is performed to an equal-time sequence; online clipping is performed when the reference slope / acceleration exceeds the boundary of S3.
[0116] S422 can optionally perform a one-time filtering (such as a first-order low-pass filter) on the reference through the setpoint filter to reduce the impact of sudden changes on the inner loop and achieve bumpless transfer.
[0117] S43, Inner Loop Controller Configuration and Operation (Executed by Inner Loop Tracking Controller 150): S431, with error For input running inner loop: PID scheme (preferred): With anti-saturation and output speed limiting; MPC scheme (optional): and (optional) Construct a finite-time-domain rolling optimization for the state / input, with constraints consistent with S3.
[0118] S432, Generate fuel command candidates The inner loop output is not sent directly; it must first be processed by the S54 action safety layer.
[0119] S44, Action Safety Layer and Execution (Execution by Actuator 160) S441, for Execute amplitude projection and rate of change limiting: , (Preferred sampling period: 0.5%FS~3.0%FS).
[0120] S442, issued after being limited in range To the fuel actuator; feed back the "amplitude saturation / rate of change limited" flag to the inner loop to trigger integral anti-saturation or feedforward compensation.
[0121] S45, Online Assessment and Health Monitoring: S451, Online calculation of evaluation quantity: Tracking error metrics (e.g., sliding window RMSE / MAE); Smoothness index ( ); Constraint monitoring ( , SM (speed limit) and violation count; Saturation event count and duration.
[0122] S452: When the violation rate or saturation rate exceeds the threshold, an alarm is triggered and the degradation strategy in S46 is implemented.
[0123] S46, Anomaly and Degradation Strategy (Safety Switching): S461, Communication Interruption: If no new reference is received for M consecutive sampling periods, execute the following steps in sequence: "Freeze the last reference → Short-time linear extrapolation → Switch to the backup reference".
[0124] S462, Outbound Risk: When it is predicted or detected that an engineering constraint will be triggered within a limited look-ahead, adaptive tightening is implemented. , If necessary, switch to minimum-maximum scheduling or fixed-gain scheduling until the risk is eliminated.
[0125] S463, Event Logging: Writes the switching reason, timestamp, and parameter snapshot to the log for later tracing.
[0126] S47, Deployment Mode and Version Management (in collaboration with device modules): S471, Online Generation and Deployment: The trajectory publishing module connects to the S3 strategy inference output in real time, and after S4 governance, it is executed according to S41~S46.
[0127] S472, Offline Table Lookup / Spline Deployment: Offline generation under representative operating conditions ,by( H , Ma )or( , The independent variable is used to form a lookup table / spline, which is then used to look up the table in the control unit and execute according to S41~S46.
[0128] S473 establishes version management and hash verification for reference versions, parameter snapshots, sampling periods, and executor boundaries to ensure consistency with training / validation.
[0129] In an improved embodiment of the above-described reference speed scheduling method for aero-engine acceleration and deceleration, the method further includes: S5. The reference speed sequence is fixed to the engine control unit in the form of a lookup table, polynomial fitting, or spline function. The input variables of the lookup table or spline function include environmental parameters and the desired steady-state target speed.
[0130] In an improved embodiment of the above-described reference speed scheduling method for aero-engine acceleration and deceleration, the method further includes: S6. Rolling predictions are made for multiple future control cycles using the trained surrogate dynamics model. If the prediction results do not meet any engineering constraints, the reference slope / acceleration is adaptively tightened, the reference trajectory is rolled back, and a safe and conservative baseline scheduling strategy is smoothly switched.
[0131] The method of this invention trains a proxy dynamics model based on engine test or bench time-series data to approximate its transient state transition relationship. In this differentiable simulation environment, a reference speed time series that satisfies the above-mentioned engineering constraints is generated through reinforcement learning strategy. Combined with the treatment of reference speed and fuel action (including slope / acceleration limiting, amplitude and rate of change projection, etc.), the safety and executability of the output command are ensured. Finally, the inner loop tracking controller completes the acceleration / deceleration transition process only with fuel flow as the actuator.
[0132] Based on the same inventive concept, this invention also provides a reference speed scheduling device for aero-engine acceleration and deceleration, as described in the following embodiments. Since the principle of the reference speed scheduling device for aero-engine acceleration and deceleration is similar to that of the reference speed scheduling method for aero-engine acceleration and deceleration, the implementation of the reference speed scheduling device for aero-engine acceleration and deceleration can refer to the implementation of the reference speed scheduling method for aero-engine acceleration and deceleration disclosed in the above embodiments, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0133] like Figure 5 The diagram shows the functional block diagram of the scheduling device. This device uses only fuel flow as the execution input and includes the coupling between the data preprocessing module 100, the agent model training module 110, the strategy optimization module 120, the governance module 130 and the trajectory publishing module 140. The structure is described below.
[0134] Specifically, the data preprocessing module 100 preprocesses the collected raw data from engine testing or bench tests to obtain engine testing or bench test timing data.
[0135] The surrogate model training module 110 is used to train a surrogate dynamics model based on engine test or bench time series data to learn the transient state transition relationship of the engine. The strategy optimization module 120 is used to optimize and generate the original control output in the simulation environment composed of the trained agent dynamics model, with the engine state variables and environmental parameters as inputs and the fuel flow rate as the action, through reinforcement learning strategy. The optimization process is based on a reward function and satisfies engineering constraints, including upper limit of turbine inlet temperature, lower limit of surge margin, upper limit of engine speed, and limits on fuel flow amplitude and rate of change. The original control output includes the original reference speed sequence and / or the original fuel action sequence, or a combination of both. The treatment module 130 is used to treat the original reference speed sequence and / or the original fuel action sequence in the original control output to obtain the reference speed sequence and / or fuel action sequence. The treatment is used to ensure that the output meets the engineering feasibility and actuator dynamic limits. The trajectory publishing module 140 is used to input the reference speed sequence as a setpoint to the inner loop tracking controller, which uses only the fuel flow rate as the actuator to drive the engine to complete the acceleration and deceleration transition process. Actuator interface and sensor interface: used for issuing commands. With collection Wait for feedback. Each module can be implemented by the processor through program instructions, or by a combination of hardware and software.
[0136] The embodiments of the present invention achieve the following technical effects: 1. Balancing speed and smoothness: By optimizing the reinforcement learning strategy on a differentiable surrogate dynamics model and combining it with the control of reference speed and fuel action, the system achieves smooth operation even when using only fuel flow rate. Under the condition of being an actuator, it effectively shortens the rise time and settling time of the acceleration and deceleration process (for example, within the improved range of 0.5~2.0s), while significantly reducing the variation variance of fuel command and suppressing the oscillation of the oil circuit and speed channel.
[0137] 2. Controllable risk of exceeding limits: Turbine inlet temperature is explicitly embedded in both the strategy training and online execution phases. The system imposes multiple engineering constraints, including upper limits for the surge margin (SM), lower limits for the speed limit, and limits for the amplitude and rate of change of fuel flow. When the rolling forecast identifies a potential constraint violation, it can adaptively back off and smoothly switch to a safe and conservative baseline scheduling strategy, achieving zero or significant reduction of out-of-bounds events and improving system safety and airworthiness verifiability.
[0138] 3. Strong consistency and portability across operating conditions: By using environmental parameters such as flight altitude (H) and Mach number (Ma) as conditional variables, the reference trajectory maintains good performance consistency across the entire flight envelope; the surrogate models (including LSTM, GRU, or encoder-type Transformer) and reinforcement learning algorithms (including SAC, DDPG, TD3, or MBPO) are modular, supporting flexible replacement and combination, facilitating cross-platform migration and subsequent upgrades.
[0139] 4. Simple and efficient engineering implementation: It relies on only a single fuel flow actuator and can be directly integrated with existing PID or MPC inner loop controllers; the overall method supports control cycles of 10~50ms, meeting the real-time requirements of aero engines; the governance module adopts lightweight operators such as slope limiting and amplitude projection, with low computational overhead, and is easy to deploy in resource-constrained engine control units (ECUs).
[0140] 5. Good verifiability and scalability: Constraints and performance indicators are expressed in quantitative form during the training and deployment phases, facilitating verification and parameter tuning via ground bench testing; without changing the core architecture of this invention, the compressor outlet pressure can be introduced as an extension. State variables such as thrust F are used to further enhance the system's robustness and applicability.
[0141] In this embodiment, a computer device is provided. Figure 6As shown, it includes a processor 501, a memory 502, a communication interface 503, and a system bus 504. The processor 501 contains a computer program stored in the memory 502 and executable on the processor. When the processor executes the computer program, it implements the aforementioned reference speed scheduling method for any of the aero-engine acceleration and deceleration methods.
[0142] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.
[0143] In this embodiment, a computer-readable storage medium is provided, which stores a computer program that executes any of the above-described reference speed scheduling methods for acceleration and deceleration of an aero-engine.
[0144] Specifically, computer-readable storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. The information can be computer-readable instructions, data structures, modules of programs, or other data, which are loaded and executed by a processor to implement the methods described above. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can store information accessible by a computing device. As defined herein, computer-readable storage media does not include transient media, such as modulated data signals and carrier waves.
[0145] Furthermore, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention can take the form of a completely or partially hardware embodiment, a completely or partially software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented in software, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any usable medium accessible to a computer or a data storage device such as a server or data center containing one or more sets of usable media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive (SSD).
[0146] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0147] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0148] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element. Furthermore, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Additionally, the character " / " in this text generally indicates an "or" relationship between the preceding and following objects, but it can also indicate an "AND / OR" relationship. Please refer to the context for specific interpretations. "At least one" refers to one or more items, while "more than" refers to two or more items. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can be represented as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0149] Furthermore, it is understood that in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0150] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0151] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of functional modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Additionally, the functional units in the various embodiments of this invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0152] If the method is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] Finally, it should be noted that the above description represents a preferred embodiment of the present invention. It should be pointed out that although preferred embodiments have been described, those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles described herein. These improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A reference speed scheduling method for acceleration and deceleration of an aero-engine, characterized in that, include: Based on engine test or bench timing data, a proxy dynamics model is trained to learn the transient state transition relationships of the engine. In the simulation environment constructed by the trained surrogate dynamics model, the original control output is generated by optimizing the engine's state variables and environmental parameters as inputs and the fuel flow rate as the action through a reinforcement learning strategy. The optimization process is based on a reward function and satisfies engineering constraints, including upper limit of turbine inlet temperature, lower limit of surge margin, upper limit of engine speed, and limits on fuel flow amplitude and rate of change. The original control output includes the original reference speed sequence and / or the original fuel action sequence, or a combination of both. The original reference speed sequence and / or original fuel action sequence in the original control output are processed to obtain a reference speed sequence and / or fuel action sequence. The processing is used to ensure that the output meets engineering feasibility and actuator dynamic limits. The reference speed sequence is input as a setpoint to the inner loop tracking controller, which uses only fuel flow as an actuator to drive the engine to complete the acceleration and deceleration transition process.
2. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, The proxy dynamics model is a recurrent neural network or an encoder-type Transformer model. When the surrogate dynamics model is a recurrent neural network, the recurrent neural network includes a long short-term memory network LSTM and / or a gated recurrent unit GRU, and a sliding time window is used to model the engine's time-series state-action data. When the surrogate dynamics model is an encoder-type Transformer model, a multi-head self-attention mechanism is used to process the engine's temporal state-action data in parallel, and training is performed based on multi-step prediction loss.
3. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 2, characterized in that, The surrogate dynamics model is trained to learn the transient state transition relationships of the engine, including: During the training phase, the engine test or bench time series data are constructed into a state-action time series, and the future state of the recurrent neural network is predicted by a multi-step autoregressive method, or a multi-head self-attention mechanism is used to perform multi-step prediction on the encoder-type Transformer model, and the loss function is constructed by the prediction error. During training, the recurrent neural network is subjected to a Scheduled Sampling strategy, which gradually anneals the teacher-forced probability from 1.0 to a preset lower limit; the encoder-type Transformer model is subjected to a Teacher Forcing strategy, which gradually reduces the teacher-forced ratio. Global norm pruning is applied to the network gradients of the recurrent neural network or the encoder-type Transformer model, and regularization is performed in combination with L2 weight decay. During the inference phase, the state increments output by the recurrent neural network or the encoder-type Transformer model are subjected to hard pruning of the domain, or the state increments are mapped to a preset physical boundary through the tanh function.
4. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, The reinforcement learning policy optimization adopts any one of the following algorithms: Soft Actor-Critic SAC, Deep Deterministic Policy Gradient (DDPG), Double Delay Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), and Model-Based Policy Optimization (MBPO). The reward function includes a speed tracking error term, a fuel smoothing penalty term, and constraint penalty terms corresponding to turbocharger temperature exceeding the limit, insufficient surge margin, speed exceeding the limit, and fuel change rate exceeding the limit, respectively. The weight of each term ranges from 0.1 to 10.
5. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, The state variables include high-pressure speed and / or low-pressure speed, and the environmental parameters include flight altitude H and Mach number Ma, or equivalent inlet total temperature and equivalent inlet total pressure calculated from flight altitude H and Mach number Ma; The upper limit of turbine inlet temperature and / or the lower limit of surge margin in the engineering constraints are estimated online by the surrogate dynamics model or the additional constraint estimation model. The estimation results are used to construct the reward function and serve as the basis for judgment on governance treatment.
6. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, The treatment process includes: Limit the slope and / or acceleration of the reference rotational speed within a preset range; The amplitude of the fuel action is projected and constrained within the effective fuel flow range. Within the control cycle of the inner loop tracking controller, the single-step change is limited to no more than 0.5% to 3% of the full scale of the fuel flow. A piecewise interpolation method is used to discretize the continuous reference trajectory into a discrete reference trajectory that can be distributed.
7. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, The inner loop tracking controller is a proportional-integral-derivative PID controller or a model predictive controller (MPC), with a control cycle of 10ms to 50ms. When the original control output is the original fuel action sequence, an equivalent reference speed is generated through the setpoint filtering and limiting mechanism of the inner loop tracking controller, and closed-loop tracking is achieved.
8. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, Also includes: The reference speed sequence is embedded into the engine control unit in the form of a lookup table, polynomial fitting, or spline function. The input variables of the lookup table or spline function include environmental parameters and the desired steady-state target speed.
9. The reference speed scheduling method for acceleration and deceleration of an aero-engine according to claim 1, characterized in that, Also includes: The trained surrogate dynamics model makes rolling predictions for multiple future control cycles. If the prediction results do not meet any engineering constraints, the reference slope / acceleration is adaptively tightened, the reference trajectory is rolled back, and a safe and conservative baseline scheduling strategy is smoothly switched.
10. A reference speed control device for acceleration and deceleration of an aircraft engine, characterized in that, The device uses only fuel flow rate as the execution input, including: The surrogate model training module is used to train a surrogate dynamics model based on engine test or bench time series data to learn the transient state transition relationships of the engine. The strategy optimization module is used to optimize and generate the original control output in a simulation environment composed of the trained agent dynamics model, with the engine's state variables and environmental parameters as inputs and fuel flow as the action, through reinforcement learning strategy. The optimization process is based on a reward function and satisfies engineering constraints, including upper limit of turbine inlet temperature, lower limit of surge margin, upper limit of engine speed, and limits on fuel flow amplitude and rate of change. The original control output includes the original reference speed sequence and / or the original fuel action sequence, or a combination of both. The treatment module is used to treat the original reference speed sequence and / or the original fuel action sequence in the original control output to obtain the reference speed sequence and / or the fuel action sequence. The treatment is used to ensure that the output meets the engineering feasibility and actuator dynamic limits. The trajectory publishing module is used to input the reference speed sequence as a setpoint to the inner loop tracking controller, which uses only fuel flow as the actuator to drive the engine to complete the acceleration and deceleration transition process.