Incremental control method for air-breathing engine based on reinforcement learning

Through the incremental control method based on deep reinforcement learning, multi-agent parallel training and safety restriction modules are adopted to solve the multivariate adjustment problem of aspirated engines in complex environments, and high-precision and adaptive control under model-free conditions are achieved, which improves the thrust performance and safety of the engine.

CN120251397APending Publication Date: 2025-07-04XIAMEN UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510371033.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing aspirated engine control methods rely on precise mathematical models and are difficult to flexibly cope with complex flow and variable flight conditions, which makes it difficult to achieve high-performance control, especially when multivariate adjustment and multi-condition switching.

Method used

The incremental control method based on deep reinforcement learning is adopted, and the status information is generated by collecting observation data, the deep reinforcement learning control module outputs control increments, and the constraint comparison is performed in combination with the safety restriction module. Finally, the execution signal is generated through the action space processing module, and the multi-dimensional fine adjustment of the engine is realized, and variables such as main fuel flow, connotation afterburner fuel flow are controlled in stages, and multiple agents are introduced to work in parallel with the action network and the evaluation network.

Benefits of technology

It realizes adaptive control of the engine in complex nonlinear environments without the need for precise mathematical models, improves thrust performance, fuel efficiency and safety, avoids safety hazards caused by large changes, and meets the needs of refined adjustments under multiple operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120251397A_ABST
    Figure CN120251397A_ABST
Patent Text Reader

Abstract

The invention provides an air-breathing engine incremental control method based on reinforcement learning, and relates to the technical field of air-breathing engine control. The deep reinforcement learning control module is used for outputting control increments of control variables such as the main fuel flow, the inner culvert and outer culvert thrust augmentation fuel flow and the nozzle throat area, the safety limiting module is combined for monitoring and cutting multiple monitoring parameters in real time, and finally a smooth and feasible execution signal is generated through the action space processing module. And multi-dimensional fine adjustment of the engine is realized. According to the scheme, multi-agent parallel training is adopted, switching between a low thrust mode and a high thrust mode is supported, under the condition that an accurate mathematical model is not needed, nonlinearity and coupling characteristics of an engine are flexibly coped, one-time large jump is prevented, safety and fuel efficiency are improved, and the requirements for high thrust and low fuel consumption in the flight process are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air-breathing engine control, and in particular to an incremental control method for air-breathing engines based on reinforcement learning. Background Art

[0002] In recent years, air-breathing engines have played an increasingly crucial role in aviation flight. There are numerous adjustable components inside them, and the coupling non-linearity is prominent. Traditional control methods mostly rely on accurate mathematical models. However, in the face of complex flows, different altitudes, and variable flight conditions, the models are difficult to remain accurate, resulting in the difficulty of achieving high-performance control. Deep Reinforcement Learning has the advantages of not requiring an accurate model and being able to handle high-dimensional continuous action spaces. By interacting with the environment, it adaptively learns optimal or approximately optimal strategies to adapt to the non-linear and uncertain characteristics of complex systems. Therefore, it is expected to break through the limitations of traditional model-based methods. However, in the existing technology, the control of air-breathing engines still mostly uses fixed algorithms or empirical formulas, making it difficult to flexibly meet the requirements of multi-variable regulation and multi-condition switching. Therefore, there is an urgent need for a control method that does not rely on accurate mathematical models, can take into account high thrust, low fuel consumption, and safety, and can finely adjust multi-dimensional control variables including fuel flow and nozzle throat area. Summary of the Invention

[0003] In order to overcome the defects of the prior art, the technical problem to be solved by the present invention is to propose an incremental control method for air-breathing engines based on reinforcement learning, and adopt the following technical solutions:

[0004] The incremental control method for air-breathing engines based on reinforcement learning includes the following steps:

[0005] S10: Collect the observation data of the engine and the external environment, and convert the above observation data into state information through the state space processing module;

[0006] S20: Input the above state information into the deep reinforcement learning control module, and output the control increment for the engine control variables;

[0007] S30: Constraine and compare the above control increment through the safety limit module, and if it exceeds the preset value, correct or limit it;

[0008] S40: Convert the corrected or limited above control increment into an execution signal through the action space processing module, input it into the actuator of the engine, and adjust the above control variables;

[0009] S50: Collect the new observation data after being adjusted by the above actuator, and update the above deep reinforcement learning control module using the above new observation data;

[0010] S60: Repeat steps S10 to S50, enabling the engine to achieve adaptive incremental control under different operating conditions;

[0011] Wherein, before the above-mentioned action space processing module sends the incremental control instruction output by the deep reinforcement learning control module to the engine actuator, the following operations are included:

[0012] S41: Multiply the above control increment by a preset gain coefficient to obtain an increment adjustment amount;

[0013] S42: Superimpose the above increment adjustment amount on the control increment of the previous moment, and convert it into an execution signal to be input into the actuator of the engine.

[0014] For further improvement, in step S20, the above control variables at least include the main fuel flow rate, the core afterburner fuel flow rate, the bypass afterburner fuel flow rate, the core nozzle throat area, and the bypass nozzle throat area.

[0015] For further improvement, in step S20, the above deep reinforcement learning control module is based on the deep deterministic policy gradient algorithm, adopts a multi-agent incremental control strategy, and distributes the above control variables to multiple agents for parallel training. The multiple above agents are respectively the main fuel flow rate gain agent, the core afterburner fuel flow rate gain agent, the bypass afterburner fuel flow rate gain agent, the core nozzle throat area gain agent, and the bypass nozzle throat area gain agent;

[0016] Each of the above agents outputs the increment of the corresponding control variable.

[0017] For further improvement, in step S20, it is divided into a low thrust stage and a high thrust stage according to different thrust requirements:

[0018] In the above low thrust stage, the above control variables are the main fuel flow rate, the core nozzle throat area, and the bypass nozzle throat area;

[0019] In the above high thrust stage, the above control variables are the main fuel flow rate, the core nozzle throat area, the bypass nozzle throat area, the core afterburner fuel flow rate, and the bypass afterburner fuel flow rate.

[0020] For further improvement, the above deep deterministic policy gradient algorithm includes an action network and an evaluation network:

[0021] The input layer of the above action network inputs the above state information. After passing through four first hidden layers with the activation function ReLU, the output layer with the activation function tanh generates the above control increment. Each of the above first hidden layers contains 32 neurons;

[0022] The input layer of the above evaluation network inputs the above state information and control increment. After the above state information and control increment pass through two second hidden layers respectively, they pass through a third hidden layer, and the output layer with the activation function of linear outputs the Q value used to measure the value of the above control increment. Both the above second hidden layer and the above third hidden layer contain 16 neurons.

[0023] For further improvement, in step S30, the above safety limit module collects monitoring parameters and compares them with the preset values of the monitoring parameters. The above monitoring parameters at least include the low-pressure relative rotational speed, high-pressure relative rotational speed, turbine outlet temperature, afterburner fuel-air ratio, and high-pressure compressor outlet pressure. The specific process of the above safety constraint comparison is as follows:

[0024] When it is detected that the low-pressure relative rotational speed or the high-pressure relative rotational speed approaches the upper limit, the increment of the main fuel flow is restricted.

[0025] When it is detected that the turbine outlet temperature exceeds the preset safety range, the increments of the core afterburner fuel flow and the bypass afterburner fuel flow are reduced.

[0026] When it is detected that the afterburner fuel-air ratio is higher than the set upper limit, the increment of the bypass afterburner fuel flow is reduced.

[0027] When it is detected that the high-pressure compressor outlet pressure exceeds the safety threshold, the increments of the core nozzle throat area and the bypass nozzle throat area are restricted simultaneously.

[0028] For further improvement, in step S30, the difference between each of the above monitoring parameters and the corresponding preset value of the monitoring parameter is calculated, and the minimum difference is selected as the restriction basis for the above safety constraint comparison.

[0029] For further improvement, the above state information at least includes

[0030] thrust error and its integral, which are used to control the main fuel flow, the core afterburner fuel flow, and the bypass afterburner fuel flow;

[0031] low-pressure relative rotational speed error and its differential, which are used to control the main fuel flow;

[0032] pressure ratio drop error and its differential, when the Mach number is less than 1.8, which are used to control the core nozzle throat area;

[0033] inlet flow coefficient error and its differential, when the Mach number is greater than 1.8, which are used to control and adjust the core nozzle throat area;

[0034] fan surge identification error and its differential, which are used to control the bypass nozzle throat area.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] First, the control algorithm of the present invention based on deep reinforcement learning introduces the deep deterministic policy gradient method, enabling the control system to dynamically adjust the policy according to the changes in real-time flight states and environmental parameters, adapt to the complex non-linear characteristics of the air-breathing engine within the full envelope range, and ensure that it always maintains an ideal operating state. Through the loop interaction mechanism in steps S10 to S60, state information extraction, incremental control instruction generation and execution, and continuous training and update of the observed data are carried out to form a closed-loop adaptive regulation ability, avoiding potential safety hazards and oscillations caused by large-scale changes at one time, and realizing refined regulation and adaptation under different thrust conditions.

[0037] Second, through the multi-agent incremental control strategy, the present invention distributes the main fuel flow rate, the fuel flow rate of the core afterburner, the fuel flow rate of the bypass afterburner, the throat area of the core nozzle, and the throat area of the bypass nozzle to multiple agents for parallel training, truly realizing the independent optimization and coordinated regulation of each control variable. Each agent outputs the increment of the corresponding control variable, and the action network and the evaluation network in the deep deterministic policy gradient algorithm work together to accurately learn the coupling relationship between different variables under complex flight conditions. Compared with the method of a single agent dealing with multiple variables, this approach greatly improves the control accuracy and response speed, effectively meeting the comprehensive requirements of the engine for thrust, fuel consumption, and safety in multi-condition environments.

[0038] Third, by dividing the control of the air-breathing engine into a low-thrust stage and a high-thrust stage, and implementing targeted tailoring or reduction of different control increments in the safety limit module according to monitoring parameters such as the relative low-pressure speed, relative high-pressure speed, turbine outlet temperature, fuel-air ratio in the afterburner, and pressure after the high-pressure compressor, reliable operation under high load or extreme conditions is achieved. If it is detected that multiple parameters are approaching the safety threshold simultaneously, the minimum difference can be selected as the limiting basis to prevent instantaneous jumps or overshoots, thereby significantly improving the system's ability to balance thrust performance and engine safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 Schematic diagram of the process of the incremental control method for an air-breathing engine based on reinforcement learning of the present invention;

[0041] Figure 2 Schematic diagram of the overall structure of the mode 1 controller in the present invention;

[0042] Figure 3 This is the overall structural schematic diagram of the mode two controller in the present invention;

[0043] Figure 4 This is the structural schematic diagram of the action network design in the present invention;

[0044] Figure 5 This is the structural schematic diagram of the evaluation network design in the present invention. Specific embodiments

[0045] For the convenience of those skilled in the art to understand, the embodiments will now be further described in detail in conjunction with the accompanying drawings for the structure of the present invention:

[0046] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. The terms "part", "side", "end", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0047] As Figure 1 shown, the present application provides an incremental control method for an air-breathing engine based on reinforcement learning, which specifically includes the following steps:

[0048] S10: Collect the observation data of the engine and the external environment, and convert the above observation data into state information through the state space processing module;

[0049] In step S10, the system first collects the observation data of the air-breathing engine and the external environment. The observation data refers to the real-time operation parameters of the air-breathing engine and the external environment parameters obtained by various sensors or detection devices, including but not limited to thrust, rotational speed, pressure ratio drop, or inlet duct flow coefficient, fan surge identification. The variables most relevant to control are screened from a large amount of observation data. In a specific embodiment, the observation data specifically includes thrust, low-pressure relative rotational speed, pressure ratio drop, inlet duct flow coefficient, and fan surge identification. The observation data is processed in real time through the state space processing module, and at least the thrust error and its differential, low-pressure relative rotational speed error and its differential, pressure ratio drop error and its differential, inlet duct flow coefficient error and its differential, and fan surge identification error and its differential are output as state information for subsequent signal transmission and control.

[0050] Among them, the error is the difference between the actual observation value and the target value, which is convenient for evaluating the static gap between the current state and the target state. The differential term is used to capture the instantaneous trend of the system's dynamic changes. Specifically, the thrust error and its integral are used to control the main fuel flow, the core afterburner fuel flow, and the bypass afterburner fuel flow; the low-pressure relative speed error and its differential are used to control the main fuel flow; the pressure ratio error and its differential are used to control the throat area of the core nozzle when the Mach number is less than 1.8; the inlet flow coefficient error and its differential are used to control the adjustment of the throat area of the core nozzle when the Mach number is greater than 1.8; the fan surge identification error and its differential are used to control the throat area of the bypass nozzle.

[0051] S20: Input the above state information into the deep reinforcement learning control module, and output the control increment for the engine control variables;

[0052] In step S20, the state information generated by the state space processing module is input into the deep reinforcement learning control module. The deep reinforcement learning control module is based on the principle of reinforcement learning and integrates the control core of the deep neural network structure to output the control increment for each control variable in the air-breathing engine.

[0053] Specifically, the above control variables specifically include the main fuel flow, the throat area of the core nozzle, the throat area of the bypass nozzle, the core afterburner fuel flow, and the bypass afterburner fuel flow. The deep reinforcement learning control module is based on the deep deterministic policy gradient algorithm and adopts a multi-agent incremental control strategy to distribute the above control variables to multiple agents for parallel training. In the above embodiment, the multiple agents are respectively the main fuel flow gain agent, the core afterburner fuel flow gain agent, the bypass afterburner fuel flow gain agent, the core nozzle throat area gain agent, and the bypass nozzle throat area gain agent. The control increment refers to the increment value relative to the control quantity at the previous moment. Compared with the absolute value output in traditional control (such as the nozzle opening is 70%), the output increment (such as the nozzle opening +2%) can make the control transition smoothly. Moreover, compared with the neural network, incremental learning is easier to converge than absolute value learning, reducing the risk of large deviations and ensuring the stability of the engine operation. When used for the first time, it can be combined with prior experience or a random strategy as the initial output and iteratively optimized as new observation data is continuously collected.

[0054] Preferably, it is divided into a low thrust stage and a high thrust stage according to different thrust requirements:

[0055] As Figure 2 shown, in the low thrust stage, the control variables are the main fuel flow, the throat area of the core nozzle, and the throat area of the bypass nozzle to meet the requirements of conventional cruise or medium thrust;

[0056] As Figure 3As shown in the figure, during the high-thrust stage, the control variables are the main fuel flow rate, the throat area of the core nozzle, the throat area of the bypass nozzle, the core afterburner fuel flow rate, and the bypass afterburner fuel flow rate. The afterburner is started to inject afterburner fuel to increase the thrust output.

[0057] Through the above segmented control strategy, the fuel economy and thrust performance can be taken into account, and at the same time, it can quickly switch to the high-thrust state when needed without causing system oscillation.

[0058] As Figure 4 and Figure 5 shown, the above deep deterministic policy gradient algorithm consists of an action network and an evaluation network. As Figure 4 shown, the action network takes the state information as input, passes through the first hidden layer with the ReLU activation function in four layers in sequence, each layer contains 32 neurons, and outputs the control increment through the tanh activation function; the evaluation network takes the same state information as input, and at the same time inputs the control increment. After passing through two hidden layers of 16 neurons respectively and then merging, it passes through a third hidden layer of 16 neurons to output the Q value, and finally the linear activation function gives an evaluation of the quality of the control increment. Among them, the Q value is usually the action value function, which is used to measure the value degree that can be obtained by taking a certain control increment in the current state. By comparing the Q value sizes corresponding to different actions, the deep reinforcement learning control module can continuously optimize the action network, so that the system can gradually learn the optimal or approximate optimal increment control scheme through multiple iterations and trial and errors, so as to achieve high-precision adjustment of the engine state and take into account safety and performance requirements. Through this network structure, the system can stably learn in the continuous action space, better adapt to the multi-variable coupling characteristics of the air-breathing engine, and quickly converge to an efficient control strategy.

[0059] S30: The above control increment is compared and restricted by the safety restriction module. If it exceeds the preset value, it will be corrected or restricted;

[0060] In step S30, the system hands over the control increment output from the deep reinforcement learning control module to the safety restriction module for comparison. The safety restriction module is responsible for monitoring key safety parameters, such as the low-pressure relative speed, the high-pressure relative speed, the turbine outlet temperature, the afterburner fuel-air ratio, and the high-pressure compressor outlet pressure, etc., and comparing them with the corresponding preset safety ranges. When any parameter approaches or exceeds the safety threshold, the control increment is corrected or restricted to prevent engine instability or high-temperature shock caused by excessive fuel flow or nozzle opening.

[0061] In a specific embodiment, the above-mentioned safety restriction module collects monitoring parameters and compares them with the preset values of the monitoring parameters. The above-mentioned monitoring parameters at least include the low-pressure relative rotational speed, the high-pressure relative rotational speed, the temperature after the turbine, the fuel-air ratio in the afterburner, and the pressure after the high-pressure compressor. The specific process of the above-mentioned safety constraint comparison is as follows:

[0062] When it is detected that the low-pressure relative rotational speed or the high-pressure relative rotational speed approaches the upper limit, the increment of the main fuel flow is restricted;

[0063] When it is detected that the temperature after the turbine exceeds the preset safety range, the increments of the core afterburner fuel flow and the bypass afterburner fuel flow are reduced;

[0064] When it is detected that the fuel-air ratio in the afterburner is higher than the set upper limit, the increment of the bypass afterburner fuel flow is reduced;

[0065] When it is detected that the pressure after the high-pressure compressor exceeds the safety threshold, the increments of the core nozzle throat area and the bypass nozzle throat area are restricted simultaneously.

[0066] Furthermore, the safety restriction module calculates the differences between each of the above-mentioned monitoring parameters and the corresponding preset values of the monitoring parameters, and selects the minimum difference as the restriction basis for the above-mentioned safety constraint comparison. If multiple monitoring parameters are simultaneously close to the restriction conditions, the differences between each monitoring parameter and its preset value are calculated, and the minimum difference among them is selected to determine the restriction basis for the safety constraint comparison, thereby further controlling the change rate of the increment. That is, if the remaining safety margin of any one parameter is extremely small, the control increment is preferentially tightened to prevent the engine thrust or temperature from overshooting instantaneously and ensure overall safety.

[0067] By performing safety verification before the action is executed, the stability and safety of the engine under high-load conditions or extreme environments can be significantly improved.

[0068] S40: The action space processing module converts the corrected or restricted control increment into an execution signal and inputs it into the actuator of the engine to adjust the control variable;

[0069] This step further includes:

[0070] S41: Multiply the control increment by a preset gain coefficient to obtain an increment adjustment amount;

[0071] S42: Superimpose the increment adjustment amount on the control amount of the previous moment, convert it into an execution signal, and input it into the actuator of the engine.

[0072] In a specific embodiment, such as Figure 2 and Figure 3As shown, the control increment input gain module is controlled, and then output through a delay module. The delay module is used to record the output of the previous moment and add it to the current control increment. Compared with the traditional method of directly outputting absolute control instructions, the method of increment control and delay superposition can significantly reduce the jump amplitude of the control quantity and avoid excessive impact on the engine. In addition, in practical applications, the output can also be finally trimmed in combination with the hardware upper limit to prevent risks caused by the execution of instructions beyond the feasible range and improve the safety and controllability of the entire system.

[0073] S50: Collect the new observed data adjusted by the above-mentioned actuator, and use the above-mentioned new observed data to update the above-mentioned deep reinforcement learning control module;

[0074] When the engine actuator completes the adjustment, new output characteristics will be generated. Therefore, in step S50, it is necessary to collect the observed data again and feedback it to the deep reinforcement learning control module for online or offline training and update. The advantage of doing this is that the system can improve the network parameters under the repeated trial-and-error and reward mechanisms, make the action network or evaluation network more adaptable to the current working conditions, and gradually achieve an efficient and stable control effect.

[0075] S60: Repeat steps S10 to S50 to enable the engine to achieve adaptive increment control under different operating conditions.

[0076] Through the closed-loop control method combining reinforcement learning and multiple links, this technical solution can enable the engine to obtain high-precision increment adjustment ability under complex non-linear conditions. Especially the incremental output, combined with the safety limit module and the action space processing module, can effectively prevent the engine from experiencing instantaneous over-limitation or large-amplitude oscillation; at the same time, multiple control variables are introduced to adapt to low-thrust and high-thrust modes, meeting various load requirements during flight. The entire process can be continuously iteratively trained. As the observed data accumulates, the deep reinforcement learning control module can continuously optimize the parameters of the action network and the evaluation network, enabling the system to have good self-learning and self-adaptive performance and accurately control the thrust while ensuring safety.

[0077] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An incremental control method for an air-breathing engine based on reinforcement learning, characterized in that, It includes the following steps: S10: Collect the observation data of the engine and the external environment, and convert the observation data into state information through the state space processing module; S20: Input the state information into the deep reinforcement learning control module, and output the control increment for the engine control variables; S30: Conduct constraint comparison on the control increment through the safety limit module, and correct or limit it if it exceeds the preset value; S40: Convert the corrected or limited control increment into an execution signal through the action space processing module, input it into the actuator of the engine, and adjust the control variable; S50: Collect the new observation data adjusted by the actuator, and update the deep reinforcement learning control module using the new observation data; S60: Repeat steps S10 to S50 to enable the engine to achieve adaptive incremental control under different operating conditions; Wherein, before the action space processing module sends the incremental control instruction output by the deep reinforcement learning control module to the engine actuator, it includes the following operations: S41: Multiply the control increment by a preset gain coefficient to obtain an incremental adjustment amount; S42: Superimpose the incremental adjustment amount on the control amount of the previous moment, and convert it into an execution signal, and input it into the actuator of the engine.

2. The incremental control method for an air-breathing engine based on reinforcement learning according to claim 1, wherein: In step S20, the control variables at least include the main fuel flow rate, the core afterburner fuel flow rate, the bypass afterburner fuel flow rate, the core nozzle throat area, and the bypass nozzle throat area.

3. The incremental control method for a scramjet engine based on reinforcement learning according to claim 2, wherein: In step S20, the deep reinforcement learning control module is based on the deep deterministic policy gradient algorithm, adopts a multi-agent incremental control strategy, and distributes the control variables to multiple agents for parallel training. The multiple agents are respectively the main fuel flow rate gain agent, the core afterburner fuel flow rate gain agent, the bypass afterburner fuel flow rate gain agent, the core nozzle throat area gain agent, and the bypass nozzle throat area gain agent; Each agent outputs the control increment corresponding to the control variable.

4. The incremental control method for an air-breathing engine based on reinforcement learning according to claim 3, characterized in that: In step S20, it is divided into a low thrust stage and a high thrust stage according to different thrust requirements: In the low thrust stage, the control variables are the main fuel flow rate, the core nozzle throat area, and the bypass nozzle throat area; In the high thrust stage, the control variables are the main fuel flow rate, the core nozzle throat area, the bypass nozzle throat area, the core afterburner fuel flow rate, and the bypass afterburner fuel flow rate.

5. The incremental control method for an air-breathing engine based on reinforcement learning according to claim 2, wherein: The deep deterministic policy gradient algorithm includes an action network and an evaluation network: The input layer of the action network inputs the state information. After passing through four first hidden layers with the activation function ReLU, the output layer with the activation function tanh generates the control increment. Each of the first hidden layers contains 32 neurons; The input layer of the evaluation network inputs the state information and the control increment. After the state information and the control increment pass through two second hidden layers respectively, they jointly pass through a third hidden layer, and the output layer with the activation function of linear outputs the Q value used to measure the value of the control increment. Both the second hidden layer and the third hidden layer contain 16 neurons.

6. The incremental control method for an air-breathing engine based on reinforcement learning according to claim 4, characterized in that: In step S30, the safety limit module collects monitoring parameters and compares them with the preset values of the monitoring parameters. The monitoring parameters at least include the low-pressure relative rotational speed, the high-pressure relative rotational speed, the temperature after the turbine, the fuel-air ratio in the afterburner, and the pressure after the high-pressure compressor. The specific process of the safety constraint comparison is as follows: When it is detected that the low-pressure relative rotational speed or the high-pressure relative rotational speed approaches the upper limit, the increment of the main fuel flow is restricted. When it is detected that the temperature after the turbine exceeds the preset safety range, the increments of the core afterburner fuel flow and the bypass afterburner fuel flow are reduced. When it is detected that the fuel-air ratio in the afterburner is higher than the set upper limit, the increment of the bypass afterburner fuel flow is reduced. When it is detected that the pressure after the high-pressure compressor exceeds the safety threshold, the increments of the core nozzle throat area and the bypass nozzle throat area are restricted simultaneously.

7. The incremental control method for an air-breathing engine based on reinforcement learning according to claim 1, wherein: In step S30, the differences between each of the monitoring parameters and the corresponding preset values of the monitoring parameters are calculated, and the minimum difference is selected as the restriction basis for the safety constraint comparison.

8. The incremental control method for an air-breathing engine based on reinforcement learning according to claim 1, wherein: The state information at least includes the thrust error and its integral, which are used to control the main fuel flow, the core afterburner fuel flow, and the bypass afterburner fuel flow; the low-pressure relative rotational speed error and its differential, which are used to control the main fuel flow; the pressure ratio drop error and its differential, which are used to control the core nozzle throat area when the Mach number is less than 1.8; the inlet flow coefficient error and its differential, which are used to control the adjustment of the core nozzle throat area when the Mach number is greater than 1.8; the fan surge identification error and its differential, which are used to control the bypass nozzle throat area.

Citation Information

Cited By

  • Method and device for determining pneumatic instability boundary of compression system in complete machine environment

    CN120628613A