A dual-mode variable cycle engine intelligent robust control method, device and medium
The intelligent robust controller, based on a two-layer control architecture and the DDPG algorithm, solves the robustness problem of aero-engines under model uncertainty and component degradation, and achieves stable and efficient operation of the engine under complex conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING POWER MACHINERY INST
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-29
AI Technical Summary
Existing aero-engine control algorithms lack robustness in the face of model uncertainties and component performance degradation, making it difficult to guarantee stable and efficient engine operation under complex conditions.
An intelligent robust controller with a two-layer control architecture, combining the DDPG algorithm and the DDPG robust correction module, simulates uncertainty by randomly biasing parameters, designs independent agents and reward functions, and optimizes the control strategy to adapt to model uncertainty and performance degradation.
Ensuring the engine operates safely, stably, and efficiently in the face of model uncertainties and component performance degradation improves the robustness and control accuracy of the controller.
Smart Images

Figure CN122106759A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of engine technology, and in particular to a method, device and medium for intelligent robust control of a dual-modal variable cycle engine. Background Technology
[0002] As a highly complex mechanical device, aero-engines experience a gradual decline in performance over long-term use due to the aging and wear of their components. Furthermore, manufacturing tolerances and environmental disturbances (such as dynamic coupling under high temperature and pressure) cause significant deviations between actual performance and theoretical models. Consequently, control algorithms designed based on ideal models often lack robustness in the face of the variable and uncertain operating environment of real engines, making it difficult to guarantee that the engine maintains efficient and stable operation under complex conditions. Summary of the Invention
[0003] In view of this, the purpose of this invention is to propose an intelligent robust control method for a dual-modal variable cycle engine, which can ensure the dual-modal variable cycle engine operates safely, stably, and efficiently even under adverse factors such as performance degradation or efficiency decline of engine components. Precise engine control is achieved through the DDPG algorithm, while ensuring robustness in the face of model uncertainties and performance degradation.
[0004] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows:
[0005] This invention provides an intelligent robust control method for a dual-modal variable cycle engine, comprising the following steps:
[0006] Step 1: Analyze the sources of uncertainty in the dual-modal variable cycle engine, and simulate the uncertainty of the dual-modal variable cycle engine by randomly adjusting the parameters, so as to provide a training scenario for the intelligent robust controller.
[0007] Step 2: Design a two-layer control architecture intelligent robust controller. The intelligent agent in the intelligent robust controller includes the DDPG intelligent control module and the DDPG robust correction module. The output of the DDPG robust correction module is used to achieve dynamic compensation of the output of the DDPG intelligent control module.
[0008] Step 3: Divide the dual-mode variable cycle engine into two modes, namely mode 1 and mode 2, and design independent agents for mode 1 and mode 2 respectively, and train corresponding correction strategies.
[0009] Step 4: Design action network and evaluation network structures for each agent, and adjust the learning rate and soft update rate hyperparameters of the agents to adapt to the environment of random parameter bias.
[0010] Step 5: Design and optimize the reward function for each agent. The reward function includes thrust error reward, speed error reward and safety margin reward to guide the agent to learn a strategy that ensures precise control and safety.
[0011] Step 6: Train the two agents separately to obtain two trained intelligent robust controllers;
[0012] Step 7: Verify the robustness and safety of the intelligent robust controller under random parameter bias conditions through simulation experiments.
[0013] Furthermore, step 1 specifically includes:
[0014] Step 11: Analyze the uncertainties of the dual-mode variable cycle engine and determine their sources, including modeling errors, structural errors, and parameter degradation.
[0015] Step 12: Determine parameters based on the sources of uncertainty, including fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency;
[0016] Step 13: Use the fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency as biasing parameters, with biasing coefficient k.
[0017] Step 14: Simulate random disturbances in fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency within a preset range by randomly adjusting parameters;
[0018] Step 15: The parameters after the pull are expressed as x' = k·x, where x is the original parameter and x' is the parameter after the pull.
[0019] Furthermore, step 2 specifically includes:
[0020] Step 21: Design an intelligent robust controller with a two-layer control architecture, including a first state space processing module, a second state space processing module, an intelligent agent, and an action space processing module. The intelligent agent includes a DDPG intelligent control module and a DDPG robust correction module.
[0021] Step 22: The first state space processing module and the second state space processing module normalize, integrate and differentiate the output of the dual-mode variable cycle engine and the input command composed of the target task to obtain the engine operating state, and convert it into a form suitable for processing by the DDPG algorithm.
[0022] Step 23: The DDPG intelligent control module generates basic control commands based on the engine operating state converted by the second state space processing module; the DDPG robust correction module generates control quantity compensation signals based on the engine operating state converted by the first state space processing module to dynamically compensate the basic control commands.
[0023] Step 24: The action space processing module maps the basic control commands and control quantity compensation signals to the specific physical range of the control actions, including the main fuel flow rate, the inner afterburner fuel flow rate, the outer afterburner fuel flow rate, and the inner nozzle throat area.
[0024] Furthermore, step 3 specifically includes:
[0025] Independent agents of mode 1 and mode 2 are designed based on the DDPG algorithm, and correction strategies are trained for agents of mode 1 and mode 2 respectively. Mode 1 is used to correct the main fuel flow rate, and mode 2 is used to correct the main fuel flow rate, the internal afterburner fuel flow rate, and the external afterburner fuel flow rate.
[0026] Furthermore, step 4 specifically includes:
[0027] Step 41: Design the action network structure, including the number of hidden layers, the number of neurons per layer, the learning rate, and the soft update rate. Adjust the hyperparameters of the learning rate and soft update rate to adapt to the environment of random parameter bias.
[0028] Step 42: Design the evaluation network structure, including the number of hidden layers in the state path, the number of hidden layers in the action path, the number of neurons per layer, the learning rate, and the soft update rate. Adjust the hyperparameters of learning rate and soft update rate to adapt to the environment of random parameter bias.
[0029] Step 43: The DDPG algorithm learns the mapping from state to action through the action network, evaluates the value of a specific action through the evaluation network, and then optimizes the control strategy.
[0030] Furthermore, step 5 specifically includes:
[0031] Step 51: For the agent in mode one, its reward function includes thrust error reward and safety margin reward; where:
[0032] 1) The goal of thrust tracking is to enable the actual thrust output of a dual-mode variable cycle engine with parameterized random yaw to track the thrust command. dmd The thrust relative error e is introduced into the design of the reward function. Thrust_rel Defined as:
[0033]
[0034] Define thrust error reward r Thrust The piecewise function is divided into segments starting from 5% error, defined as follows:
[0035]
[0036] 2) Safety margin reward r of a modality-one agent CBF1 as follows:
[0037]
[0038] Where a1, b1, c1, and d1 represent weighting coefficients used to balance the values of various safety terms, C represents the safety set, and NL represents the low-pressure relative speed. max NL represents the maximum relative speed at low pressure, i.e., NL = 120%, while NH represents the relative speed at high pressure. max This represents the maximum relative speed under high pressure, i.e., NH = 120%, and tT6 represents the turbine exhaust temperature. max This indicates the maximum temperature after the turbine, tT6 = 1561.15 K, and sP31 indicates the pressure after the high-pressure compressor. max This indicates the maximum pressure after the high-pressure compressor, i.e., sP31 = 5066250 Pa. , , and Indicates the scaling factor;
[0039] Step 52: For the modality 2 agent, its reward function includes thrust error reward, rotational speed error reward, and safety margin reward; where:
[0040] 1) The goal of thrust tracking is to enable the actual thrust output of a dual-mode variable cycle engine with parameterized random yaw to track the thrust command. dmd The thrust relative error e is introduced into the design of the reward function. Thrust_rel Defined as:
[0041]
[0042] Define thrust error reward r Thrust The piecewise function is divided into segments starting from 5% error, defined as follows:
[0043]
[0044] 2) Speed error reward r NH As shown in the following formula:
[0045]
[0046] 3) Safety margin reward r of modality 2 agent CBF2 as follows:
[0047]
[0048] Where a2 and b2 represent weighting coefficients used to balance the values of various safety terms, C' represents the safety set, and FAR1 represents the air-fuel ratio within the afterburner. max FAR1 represents the maximum air-fuel ratio in the internal afterburner, i.e., FAR1 = 1 / 17.5, and FAR2 represents the air-fuel ratio in the external afterburner. max This represents the maximum air-fuel ratio in the bypass afterburner, i.e., FAR2 = 1 / 17.5. and Indicates the scaling factor;
[0049] Step 53: The overall reward value reward1 for the agent in mode one is as follows:
[0050]
[0051] The overall reward value reward2 for the modality 2 agent is as follows:
[0052]
[0053] Furthermore, step 6, training the agent, specifically includes:
[0054] Step 61: Randomly set the altitude, Mach number, and target thrust command step signal for each round of training;
[0055] Step 62: Randomly reset the pull-off parameter value every n seconds to simulate a real uncertainty environment;
[0056] Step 63: The training step length is t, the training time for each round is T, and the total number of steps is T / t, where T ranges from 40 to 50 seconds and t is 0.02 seconds.
[0057] Furthermore, the pull coefficient k ranges from 0.95 to 1; for the agents of mode one and mode two, the safety constraints include: low-pressure relative speed ≤ 120%, high-pressure relative speed ≤ 120%, turbine afterburner temperature ≤ 1561.15 K, and high-pressure compressor afterburner pressure ≤ 5066250 Pa; for the agent of mode two, the safety constraints include: internal afterburner fuel-air ratio ≤ 1 / 17.5 and external afterburner fuel-air ratio ≤ 1 / 17.5.
[0058] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the intelligent robust control method for a dual-modal variable cycle engine as described above.
[0059] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described intelligent robust control method for a dual-modal variable cycle engine.
[0060] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art:
[0061] To address the uncertainties in a dual-modal variable-cycle engine model, this invention introduces parameter biasing into the simulation model of the dual-modal variable-cycle engine. Subsequently, using an engine model with stochastic parameter biasing characteristics as the controlled object, a robust control strategy for the dual-modal engine is studied. A two-layer control architecture is introduced, adding a DDPG robust correction module to the output of the DDPG intelligent control module. Similarly, a CBF control barrier function is introduced, and the update rate of the internal network of the intelligent agent is reduced, which helps to improve the robustness of the controller. Simulation experiments designed and executed on the Simulink platform verify the advantages of this two-layer architecture intelligent robust controller in terms of safety and performance. This invention can ensure that the dual-modal variable-cycle engine operates safely, stably, and efficiently even under adverse factors such as performance degradation or efficiency decline of engine components. Precise engine control is achieved through the DDPG algorithm, while ensuring robustness in the face of model uncertainties and performance degradation. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0063] Figure 1 This is an execution flowchart of an intelligent robust control method for a dual-modal variable cycle engine provided in an embodiment of the present invention.
[0064] Figure 2 This is the overall framework of the control system to which the intelligent robust controller for a dual-modal variable cycle engine with parameter uncertainty provided in the embodiments of the present invention belongs.
[0065] Figure 3 This is a graph showing the results of a pull test on an existing safety intelligent controller.
[0066] Figure 4 This is a diagram showing the effect of the biasing action in Simulink.
[0067] Figure 5 This is a diagram showing the relationship between the environment and the intelligent agent.
[0068] Figure 6 This is an action network structure diagram.
[0069] Figure 7 It is used to evaluate network structure diagrams.
[0070] Figure 8 This is a graph of the reward function of the modal-robust correction module.
[0071] Figure 9 This is the reward function curve of the modal 2 robust correction module.
[0072] Figure 10a This is a thrust tracking curve graph comparing the intelligent robust controller and the safety intelligent controller.
[0073] Figure 10b This is a pressure drop ratio tracking curve from a comparative test of the intelligent robust controller and the safety intelligent controller.
[0074] Figure 10c This is a graph showing the relative fan speeds of a fan compared to a smart robust controller and a safety smart controller.
[0075] Figure 10d This is the main fuel flow curve of the comparative test between the intelligent robust controller and the safety intelligent controller.
[0076] Figure 10e This is a curve showing the nozzle throat area of a comparative test between an intelligent robust controller and a safety intelligent controller.
[0077] Figure 10f This is a fuel flow curve graph comparing the intelligent robust controller and the safety intelligent controller.
[0078] Figure 10g This is a graph showing the relative fan speed limit compared to a smart robust controller and a safe smart controller.
[0079] Figure 10h This is a high-voltage relative speed limit curve from a comparative test of an intelligent robust controller and a safety intelligent controller.
[0080] Figure 10i This is a turbine-after temperature curve from a comparative test of the intelligent robust controller and the safety intelligent controller.
[0081] Figure 10j This is a high-pressure oil-gas ratio curve from a comparative test of the intelligent robust controller and the safety intelligent controller.
[0082] Figure 10k This is a low-pressure oil-gas ratio curve from a comparative test of the intelligent robust controller and the safety intelligent controller.
[0083] Figure 10lThis is a pressure curve of the high-pressure compressor after a comparative test between an intelligent robust controller and a safety intelligent controller.
[0084] Figure 11 This is a schematic diagram of an electronic device provided in an embodiment of the present invention.
[0085] Figure 12 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of the present invention. Detailed Implementation
[0086] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0087] Please see Figure 1 The present invention provides an intelligent robust control method for a dual-modal variable cycle engine, comprising the following steps:
[0088] Step 1: Analyze the sources of uncertainty in the dual-mode variable cycle engine, and simulate the uncertainty of the dual-mode variable cycle engine by randomly adjusting the parameters, so as to provide a training scenario for the intelligent robust controller (which can accurately reflect the uncertainty factors faced by the engine in actual operation).
[0089] In this embodiment, step 1 specifically includes:
[0090] Step 11: Analyze the uncertainties of the dual-mode variable cycle engine and determine its main sources, including modeling errors, structural errors, and parameter degradation.
[0091] As a complex aerodynamic and thermodynamic system operating under long-term high temperature, high pressure, and high-speed rotation conditions, the performance and characteristics of a dual-mode variable cycle engine are inevitably affected by perturbations and parameter fluctuations in the engine's control system during practical applications. Model uncertainties mainly originate from the following aspects: modeling errors, structural errors, and parameter degradation.
[0092] Modeling errors: The aerodynamic principles of a dual-modal variable cycle engine are exceptionally complex, involving various nonlinear and dynamic coupling effects, making it difficult to construct a completely accurate mathematical model. This results in models that can only reflect the engine's performance under certain specific operating conditions and cannot cover all possible operating scenarios.
[0093] Structural Errors: Due to inherent manufacturing process errors, each engine produced will have slight structural variations. These variations prevent the aerodynamic characteristics of each engine from perfectly matching the theoretical model. Furthermore, over time, internal engine components experience mechanical wear and aging due to prolonged operation, further exacerbating the discrepancy between the actual structure and the pre-designed model.
[0094] Parameter degradation: During long-term engine operation, critical components undergo parameter degradation due to factors such as mechanical wear, high-temperature corrosion, and airflow erosion, leading to decreased operating efficiency. Examples include fan components, compressor components, and turbine components. Parameter degradation affects the stability of the control system and the efficiency of energy conversion, causing a decrease in engine thrust output.
[0095] These uncertainties pose a severe challenge to the design and implementation of control systems, requiring control strategies to fully consider these errors and changes in order to ensure the stability and reliability of the control system under complex operating conditions.
[0096] Step 12: Determine parameters based on the sources of uncertainty, including fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency;
[0097] Step 13: Use the fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency as biasing parameters, with biasing coefficient k.
[0098] These efficiency parameters affect the engine's aerodynamic performance, energy conversion efficiency, thrust output, and fuel economy at different operating stages. A decrease in any one of these efficiency parameters can trigger a series of negative effects, potentially leading to insufficient thrust, increased fuel consumption, higher emissions, and control system instability. By adjusting these parameters, random variations caused by modeling errors, structural errors, and parameter degradation can be reflected, thus introducing uncertainties into the model (dual-modal variable cycle engine). Table 1 lists the parameter names, adjustment parameter symbols, and the range of adjustment coefficient values.
[0099] Table 1. Pulling parameters and coefficient ranges
[0100] To evaluate the robustness of existing safety intelligent controllers in the face of model uncertainties, this study set the pull coefficient k in the range of 0.95 ≤ k ≤ 1 and conducted thrust control tracking experiments. Figure 3 It can be seen that after the parameters were adjusted, both Mode 1 and Mode 2 showed large steady-state errors and failed to achieve good thrust tracking.
[0101] Step 14: Simulate random disturbances in fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency within a preset range by randomly adjusting parameters; construct a digital simulation environment covering efficiency disturbances from -5% to 0 to simulate the performance degradation and dynamic deviations of key components such as fans, compressors, and turbines, providing a training scenario for robust control algorithms;
[0102] Step 15: To improve the robustness of the algorithm, a new DDPG robust correction module is designed based on the existing safety intelligent controller. The uncertainty of the model is simulated by biasing the parameters. Assuming the biased variable is x and the biasing coefficient is k, the biased parameters are expressed as x' = k·x, where x is the original parameter and x' is the biased parameter. The biasing injection module is retained in the model design, such as... Figure 4 As shown in the figure, the random constant set is the aforementioned pull coefficient k.
[0103] Step 2: Design a two-layer control architecture intelligent robust controller. The intelligent agent in the intelligent robust controller includes the DDPG intelligent control module and the DDPG robust correction module. The output of the DDPG robust correction module is used to achieve dynamic compensation of the output of the DDPG intelligent control module.
[0104] In this embodiment, step 2 specifically includes:
[0105] Step 21: Design an intelligent robust controller with a two-layer control architecture, such as... Figure 2 As shown, it includes a first state space processing module, a second state space processing module, an intelligent agent, and an action space processing module. The intelligent agent includes a DDPG intelligent control module and a DDPG robust correction module. The pull-off parameter and other parameters are used to simulate the performance degradation and efficiency reduction that the engine may encounter in actual operation. Among them, other parameters refer to parameters that cannot be pulled off, such as the fan guide vane angle parameter. The controlled object of the control system to which the intelligent robust controller belongs is a dual-mode variable cycle engine with parameterized random pull-off. The corresponding control objectives are thrust tracking, safety protection, and mode switching smoothness of the dual-mode variable cycle engine with parameterized random pull-off.
[0106] Step 22: The first state space processing module and the second state space processing module normalize, integrate and differentiate the output of the dual-mode variable cycle engine and the input command composed of the target task to obtain the engine operating state, and convert it into a form suitable for processing by the DDPG algorithm.
[0107] Step 23: The DDPG intelligent control module generates basic control commands based on the engine operating state converted by the second state space processing module; the DDPG robust correction module generates control quantity compensation signals based on the engine operating state converted by the first state space processing module, and dynamically compensates the basic control commands; the DDPG robust correction module takes the real-time engine state as input and outputs dynamic compensation quantities for the main fuel flow, internal afterburner fuel flow, and external afterburner fuel flow, that is, it receives the real-time state information of the engine in real time, and generates control quantity compensation signals through analysis and processing of this information, and dynamically compensates (i.e. corrects) the output of the DDPG intelligent control module, thereby realizing dynamic compensation for the uncertainty of the dual-modal variable cycle engine, avoiding intelligent robust controller reconstruction and repeated training, significantly reducing trial and error costs, and accelerating network convergence; this two-layer control structure not only makes full use of the good control performance of the DDPG intelligent control module under the ideal model, but also enhances the adaptability of the control system to uncertainties in actual working conditions through the DDPG robust correction module.
[0108] Step 24: The action space processing module maps the basic control commands and control quantity compensation signals to the specific physical range of the control actions, including the main fuel flow rate, the inner afterburner fuel flow rate, the outer afterburner fuel flow rate, and the inner nozzle throat area.
[0109] Step 3: Divide the dual-mode variable cycle engine into two modes: Mode 1 and Mode 2. Design independent agents for each mode and train corresponding correction strategies. In Mode 1, the engine only needs to burn the main fuel to meet the thrust target when operating at low thrust. When the thrust gradually increases and the main fuel reaches its maximum value but still cannot meet the target, it needs to activate the internal and external afterburning fuels to burn simultaneously with the main fuel, providing greater thrust. In other words, Mode 2 involves activating both internal and external afterburning fuels, while Mode 1 involves not activating either.
[0110] In this embodiment, step 3 specifically includes: designing independent agents of mode 1 and mode 2 based on the DDPG algorithm, and training correction strategies for agents of mode 1 and mode 2 respectively;
[0111] Mode 1 is used to correct the main fuel flow rate to ensure thrust tracking accuracy; Mode 2 is used to correct the main fuel flow rate, the inner afterburner fuel flow rate, and the outer afterburner fuel flow rate to adapt to the complex requirements under afterburner conditions; when training the correction strategy, a fuel correction compensation mechanism is introduced to dynamically offset the impact of parameter bias on the control system performance, and the fuel flow rate is corrected in Mode 1 and Mode 2 respectively (that is, in Mode 1, the main fuel flow rate is corrected, and in Mode 2, the main fuel flow rate, the inner afterburner fuel flow rate, and the outer afterburner fuel flow rate are corrected).
[0112] The DDPG robust correction module takes real-time states as input, including thrust error and speed error, and outputs the correction amounts for each fuel flow rate. Independent agents for Mode 1 and Mode 2 are designed based on the DDPG algorithm, and correction strategies are trained separately for each. The observation and action values of the robust correction module for Mode 1 are shown in Table 2, and those for Mode 2 are shown in Table 3.
[0113] Table 2. Observations, action values, and symbols of the modal-intelligent robust correction module.
[0114] Table 3 Observations, Action Values, and Symbols of the Modal Two Intelligent Robust Correction Module
[0115] like Figure 2 As shown, the DDPG intelligent control module and the DDPG fuel trimming module are both designed using the DDPG algorithm (i.e., both the DDPG intelligent control and robust trimming modules use the DDPG algorithm, and the algorithm framework is the same; the innovation lies in introducing the robust trimming module to correct the normal DDPG output). The relationship between the DDPG algorithm and the environment is as follows: Figure 5 As shown, the actuator is responsible for making decisions, and the evaluator is responsible for evaluating the value of the decisions. They interact with the engine model, i.e. the environment, and store the data in the experience replay pool for learning and optimization, while introducing random perturbations into the engine input.
[0116] Step 4: Design action network and evaluation network structures for each agent, and adjust the learning rate and soft update rate hyperparameters of the agents to adapt to the environment of random parameter bias.
[0117] In this embodiment, step 4 specifically includes:
[0118] Step 41: Design the action network structure, including the number of hidden layers, the number of neurons per layer, the learning rate, and the soft update rate. See Table 4. Adjust the hyperparameters of learning rate and soft update rate to adapt to the environment of random parameter bias.
[0119] Step 42: Design the evaluation network structure, including the number of hidden layers in the state path, the number of hidden layers in the action path, the number of neurons per layer, the learning rate, and the soft update rate. See Table 4. Adjust the hyperparameters of learning rate and soft update rate to adapt to the environment of random parameter bias.
[0120] Step 43: The DDPG algorithm learns the mapping from state to action through an action network to generate the optimal control policy. The action network structure is as follows: Figure 6As shown, the evaluation network assesses the value of a specific action based on the current state, thereby optimizing the control strategy. Figure 7 As shown.
[0121] By introducing random parameter biasing, the control system correspondingly reduces the learning rate and soft update parameters of the action and evaluation networks. This adjustment effectively slows down the parameter update speed, thereby improving training stability and enabling the agent to fully utilize training data and explore the state space more deeply, ultimately converging to a local or global optimum.
[0122] Table 4 Hyperparameters of the Intelligent Robust Controller
[0123] As shown in the table above, the action network structure is designed as follows: the action network has 4 hidden layers, with 32 neurons per layer; the learning rate and soft update rate are 0.00001, and the hyperparameters are adjusted to adapt to environments with random parameter bias. The evaluation network structure is designed as follows: the evaluation network has 2 hidden layers for the state path and 3 hidden layers for the action path, with 32 neurons per layer; the learning rate and soft update rate are 0.0001. The action network was simplified to have smaller observation and action dimensions, so the depth was reduced from 6 to 4 while maintaining the width of each layer. The relatively simplified network architecture and smaller learning rate of the action and evaluation networks help maintain learning stability and improve adaptability to system parameter fluctuations in tasks with small state and action space dimensions.
[0124] Step 5: Design and optimize the reward function for each agent. The reward function includes thrust error reward, speed error reward and safety margin reward to guide the agent to learn a strategy that ensures both precision control and safety. Combined with the control barrier function (CBF), while optimizing thrust tracking accuracy, strictly ensure the safety margin of key parameters such as speed and temperature to avoid the risk of exceeding limits.
[0125] In this embodiment, step 5 specifically includes:
[0126] For the agent in mode one, thrust error is considered; for the agent in mode two, both thrust error and rotation speed error are considered. Based on the above analysis, the specific reward function design scheme is as follows:
[0127] Step 51: For the agent in mode one, its reward function includes thrust error reward and safety margin reward; where:
[0128] 1) The goal of thrust tracking is to enable the actual thrust output of a dual-mode variable cycle engine with parameterized random yaw to track the thrust command. dmdThe thrust relative error e is introduced into the design of the reward function. Thrust_rel Defined as:
[0129]
[0130] Define thrust error reward r Thrust The piecewise function is divided into segments starting from 5% error, defined as follows:
[0131]
[0132] For a modality-1 intelligent agent, the safety constraints to be met include: low-pressure relative speed ≤120%, high-pressure relative speed ≤120%, turbine after-temperature ≤1561.15K, and high-pressure compressor after-pressure ≤5066250Pa.
[0133] 2) Safety margin reward r of a modality-one agent CBF1 as follows:
[0134]
[0135] Where a1, b1, c1, and d1 represent weighting coefficients used to balance the values of various safety terms, C represents the safety set, and NL represents the low-pressure relative speed. max NL represents the maximum relative speed at low pressure, i.e., NL = 120%, while NH represents the relative speed at high pressure. max This represents the maximum relative speed under high pressure, i.e., NH = 120%, and tT6 represents the turbine exhaust temperature. max This indicates the maximum temperature after the turbine, tT6 = 1561.15 K, and sP31 indicates the pressure after the high-pressure compressor. max This indicates the maximum pressure after the high-pressure compressor, i.e., sP31 = 5066250 Pa. , , and This represents the scaling factor, which ensures adaptation to different safety boundaries, ensuring the safety of the control system while also considering performance.
[0136] Step 52: For the modality 2 agent, its reward function includes thrust error reward, rotational speed error reward, and safety margin reward; where:
[0137] 1) The goal of thrust tracking is to enable the actual thrust output of a dual-mode variable cycle engine with parameterized random yaw to track the thrust command. dmd The thrust relative error e is introduced into the design of the reward function. Thrust_rel Defined as:
[0138]
[0139] Define thrust error reward r Thrust The piecewise function is divided into segments starting from 5% error, defined as follows:
[0140]
[0141] For the modality 2 agent, the safety constraints to be met include: low-pressure relative speed ≤ 120%, high-pressure relative speed ≤ 120%, turbine after-temperature ≤ 1561.15K, and high-pressure compressor after-pressure ≤ 5066250Pa.
[0142] 2) Speed error reward r NH As shown in the following formula:
[0143]
[0144] For a modal 2 agent, the safety constraints include: the fuel-air ratio in the inner afterburner ≤ 1 / 17.5 and the fuel-air ratio in the outer afterburner ≤ 1 / 17.5.
[0145] 3) Safety margin reward r of modality 2 agent CBF2 as follows:
[0146]
[0147] Where a2 and b2 represent weighting coefficients used to balance the values of various safety terms, C' represents the safety set, and FAR1 represents the air-fuel ratio within the afterburner. max FAR1 represents the maximum air-fuel ratio in the internal afterburner, i.e., FAR1 = 1 / 17.5, and FAR2 represents the air-fuel ratio in the external afterburner. max This represents the maximum air-fuel ratio in the bypass afterburner, i.e., FAR2 = 1 / 17.5. and This represents the scaling factor, which ensures adaptation to different safety boundaries, ensuring the safety of the control system while also considering performance.
[0148] In this embodiment, the final specific parameters are set as follows: , , , , , , , , , , , .
[0149] Step 53: The overall reward value reward1 for the agent in mode one is as follows:
[0150]
[0151] The overall reward value reward2 for the modality 2 agent is as follows:
[0152]
[0153] Step 6: Train the two agents separately to obtain two trained intelligent robust controllers;
[0154] In this embodiment, step 6, training the agent, specifically includes:
[0155] Step 61: Randomly set the altitude, Mach number, and target thrust command step signal for each round of training;
[0156] Step 62: Randomly reset the pull-off parameter value every n seconds to simulate a real uncertainty environment;
[0157] Step 63: The training step length is t, the training time for each round is T, and the total number of steps is T / t, where T ranges from 40 to 50 seconds and t ranges from 0.01 to 0.02 seconds.
[0158] For example, in each training round, the altitude (0-3000) and Mach number (0-0.3) within the randomly selected area are set. Based on the randomly selected altitude and Mach number in this round, a target thrust command step signal is randomly assigned. The running step size of the environment is consistent with the running step size of the engine simulation model, which is 0.02s. The training time for each round of the agent is 40s, with a total of 2000 steps. Each round runs for a total of 80s, with the first 40 seconds being the startup time, during which the agent does not participate in control operations. In each training round, each step command is held for 5s, meaning a thrust step command is randomly assigned every 5s. Simultaneously, to expand the environmental sample, the values of each pull parameter (0.95-1) are randomly reset every 5s. The specific hyperparameters are shown in Table 5.
[0159] Table 5 Hyperparameters for Agent Training
[0160] Figure 8 and Figure 9 The overall trend of the DDPG agent's training curve shows that the reward value increases significantly with each training epoch. The reward value is low in the initial stage, gradually increasing as training progresses, reflecting the agent's continuous policy optimization during training. Finally, the reward value approaches convergence, indicating that the network parameters are updated less frequently, the training has entered a relatively stable stage, and the agent's policy is approaching stability.
[0161] Step 7: Verify the robustness and safety of the intelligent robust controller under random parameter deviation conditions through simulation experiments on the Simulink platform. Design a simulation comparison experiment to evaluate key indicators such as thrust tracking error and safety constraint violation rate, and prove the robustness and safety of the proposed method under parameter deviation conditions.
[0162] To verify the performance and safety of the proposed dual-layer control strategy's intelligent robust controller, simulation experiments were designed. The controller response was triggered by a step thrust demand signal, allowing for the acquisition of key control performance indicators. Simulation experiments were conducted at 0 altitude and Mach number 0 to test the control performance of the dual-layer architecture intelligent robust controller. During the test, the fan efficiency bias parameters, compressor efficiency bias parameters, high-pressure turbine efficiency bias parameters, and low-pressure turbine efficiency bias parameters were all set to 0.95, representing a 5% efficiency decrease. Figures 10a-10l This study compares the pull-off test results of the intelligent robust controller with those of existing safety intelligent controllers. Figures 10a to 10c For the corresponding control objectives, Figures 10d to 10f For changes in various intelligent control quantities. Figure 10g to Figure 10l For each security restriction and protection situation.
[0163] from Figures 10a-10l As can be seen, compared with the safety intelligent controller, the intelligent robust controller based on the dual-layer architecture design can achieve rapid and accurate tracking of engine thrust, speed and pressure ratio by coordinating action outputs, while always staying within the safety boundary.
[0164] Table 6 shows the test indicators of the intelligent robust controller at different heights and Mach numbers, where the fan efficiency bias parameter, compressor efficiency bias parameter, high-pressure turbine efficiency bias parameter, and low-pressure turbine efficiency bias parameter are all set to 0.95.
[0165] Table 6 Test Indicators for Intelligent Robust Controllers
[0166] After multiple sets of experimental tests, the steady-state control error was less than 1% and the overshoot was less than 5%. This verifies the effectiveness and safety of the intelligent robust controller based on a two-layer architecture designed in this invention.
[0167] like Figure 11 As shown, this embodiment of the invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the above-described intelligent robust control method for a dual-modal variable cycle engine.
[0168] like Figure 12As shown, this embodiment of the invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described intelligent robust control method for a dual-modal variable cycle engine.
[0169] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0170] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0171] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A smart robust control method for a dual-modal variable cycle engine, characterized in that, Includes the following steps: Step 1: Analyze the sources of uncertainty in the dual-modal variable cycle engine, and simulate the uncertainty of the dual-modal variable cycle engine by randomly adjusting the parameters, so as to provide a training scenario for the intelligent robust controller. Step 2: Design a two-layer control architecture intelligent robust controller. The intelligent agent in the intelligent robust controller includes the DDPG intelligent control module and the DDPG robust correction module. The output of the DDPG robust correction module is used to achieve dynamic compensation of the output of the DDPG intelligent control module. Step 3: Divide the dual-mode variable cycle engine into two modes, namely mode one and mode two, and design independent agents for mode one and mode two respectively, and train corresponding correction strategies. Step 4: Design action network and evaluation network structures for each agent, and adjust the learning rate and soft update rate hyperparameters of the agents to adapt to the environment of random parameter bias. Step 5: Design and optimize the reward function for each agent. The reward function includes thrust error reward, speed error reward and safety margin reward to guide the agent to learn a strategy that ensures precise control and safety. Step 6: Train the two agents separately to obtain two trained intelligent robust controllers; Step 7: Verify the robustness and safety of the intelligent robust controller under random parameter bias conditions through simulation experiments.
2. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, Step 1 specifically includes: Step 11: Analyze the uncertainties of the dual-mode variable cycle engine and determine their sources, including modeling errors, structural errors, and component efficiency degradation; Step 12: Determine parameters based on the sources of uncertainty, including fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency; Step 13: Use the fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency as biasing parameters, with biasing coefficient k. Step 14: Simulate random disturbances in fan efficiency, compressor efficiency, high-pressure turbine efficiency, and low-pressure turbine efficiency within a preset range by randomly adjusting parameters; Step 15: The parameters after the pull are expressed as x' = k·x, where x is the original parameter and x' is the parameter after the pull.
3. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, Step 2 specifically includes: Step 21: Design an intelligent robust controller with a two-layer control architecture, including a first state space processing module, a second state space processing module, an intelligent agent, and an action space processing module. The intelligent agent includes a DDPG intelligent control module and a DDPG robust correction module. Step 22: The first state space processing module and the second state space processing module normalize, integrate and differentiate the output of the dual-mode variable cycle engine and the input command composed of the target task to obtain the engine operating state, and convert it into a form suitable for processing by the DDPG algorithm. Step 23: The DDPG intelligent control module generates basic control commands based on the engine operating state converted by the second state space processing module; the DDPG robust correction module generates control quantity compensation signals based on the engine operating state converted by the first state space processing module to dynamically compensate the basic control commands. Step 24: The action space processing module maps the basic control commands and control quantity compensation signals to the specific physical range of the control actions, including the main fuel flow rate, the inner afterburner fuel flow rate, the outer afterburner fuel flow rate, and the inner nozzle throat area.
4. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, Step 3 specifically includes: designing independent agents of mode 1 and mode 2 based on the DDPG algorithm, and training correction strategies for agents of mode 1 and mode 2 respectively; Among them, mode one is used to correct the main fuel flow rate, and mode two is used to correct the main fuel flow rate, the internal afterburner fuel flow rate, and the external afterburner fuel flow rate.
5. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, Step 4 specifically includes: Step 41: Design the action network structure, including the number of hidden layers, the number of neurons per layer, the learning rate, and the soft update rate. Adjust the hyperparameters of the learning rate and soft update rate to adapt to the environment of random parameter bias. Step 42: Design the evaluation network structure, including the number of hidden layers in the state path, the number of hidden layers in the action path, the number of neurons per layer, the learning rate, and the soft update rate. Adjust the hyperparameters of learning rate and soft update rate to adapt to the environment of random parameter bias. Step 43: The DDPG algorithm learns the mapping from state to action through the action network, evaluates the value of a specific action through the evaluation network, and then optimizes the control strategy.
6. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, Step 5 specifically includes: Step 51: For the agent in mode one, its reward function includes thrust error reward and safety margin reward; where: 1) The goal of thrust tracking is to enable the actual thrust output of a dual-mode variable cycle engine with parameterized random yaw to track the thrust command. dmd The thrust relative error e is introduced into the design of the reward function. Thrust_rel Defined as: , Define thrust error reward r Thrust The piecewise function is divided into segments starting from 5% error, defined as follows: , 2) Safety margin reward r of a modality-one agent CBF1 as follows: , Where a1, b1, c1, and d1 represent weighting coefficients used to balance the values of various safety terms, C represents the safety set, and NL represents the low-pressure relative speed. max NL represents the maximum relative speed at low pressure, i.e., NL = 120%, while NH represents the relative speed at high pressure. max This represents the maximum relative speed under high pressure, i.e., NH = 120%, and tT6 represents the turbine exhaust temperature. max This indicates the maximum temperature after the turbine, tT6 = 1561.15 K, and sP31 indicates the pressure after the high-pressure compressor. max This indicates the maximum pressure after the high-pressure compressor, i.e., sP31 = 5066250 Pa. , , and Indicates the scaling factor; Step 52: For the modality 2 agent, its reward function includes thrust error reward, rotational speed error reward, and safety margin reward; where: 1) The goal of thrust tracking is to enable the actual thrust output of a dual-mode variable cycle engine with parameterized random yaw to track the thrust command. dmd The thrust relative error e is introduced into the design of the reward function. Thrust_rel Defined as: Define thrust error reward r Thrust The piecewise function is divided into segments starting from 5% error, defined as follows: , 2) Speed error reward r NH As shown in the following formula: , 3) Safety margin reward r of modality 2 agent CBF2 as follows: , Where a2 and b2 represent weighting coefficients used to balance the values of various safety terms, C' represents the safety set, and FAR1 represents the air-fuel ratio in the afterburner. max FAR1 represents the maximum air-fuel ratio in the internal afterburner, i.e., FAR1 = 1 / 17.5, and FAR2 represents the air-fuel ratio in the external afterburner. max This represents the maximum air-fuel ratio in the bypass afterburner, i.e., FAR2 = 1 / 17.
5. and Indicates the scaling factor; Step 53: The overall reward value reward1 for the agent in mode one is as follows: , The overall reward value reward2 for the modality 2 agent is as follows: 。 7. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, Step 6, training the agent, specifically includes: Step 61: Randomly set the altitude, Mach number, and target thrust command step signal for each round of training; Step 62: Randomly reset the pull-off parameter value every n seconds to simulate a real uncertainty environment; Step 63: The training step length is t, the training time for each round is T, and the total number of steps is T / t, where T ranges from 40 to 50 seconds and t is 0.02 seconds.
8. The intelligent robust control method for a dual-modal variable cycle engine as described in claim 1, characterized in that, The pull coefficient k ranges from 0.95 to 1; for the agents of mode one and mode two, the safety constraints include: low-pressure relative speed ≤ 120%, high-pressure relative speed ≤ 120%, turbine afterburner temperature ≤ 1561.15 K, and high-pressure compressor afterburner pressure ≤ 5066250 Pa; for the agent of mode two, the safety constraints include: internal afterburner fuel-air ratio ≤ 1 / 17.5 and external afterburner fuel-air ratio ≤ 1 / 17.
5.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a dual-modal variable cycle engine intelligent robust control method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements a smart robust control method for a dual-modal variable cycle engine as described in any one of claims 1 to 8.