A Method for Constructing a Fusion Model of a High-Passflow Dual-Cycle Engine Based on Reinforcement Learning
By using a high-pass flow dual-variable cycle engine fusion model based on reinforcement learning, and combining the Newton-Raphson algorithm and the TD3 algorithm, the input-output structure and reward system were designed to solve the accuracy and convergence problems of aero-engine models under full envelope and full operating conditions, achieving fast convergence and high-precision engine simulation.
Patent Information
- Application Number
- CN202411420849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-10-12
AI Technical Summary
Existing aero-engine models have difficulty guaranteeing accuracy and convergence under full envelope and full operating conditions. Traditional models can only be used for simulation at specific points, and iterative algorithms have a large computational burden.
A high-throughput dual-variable cycle engine fusion model based on reinforcement learning is adopted. Combining the Newton-Raphson algorithm and the dual-delay deep deterministic policy gradient algorithm TD3, the input-output structure and reward system are designed, the hyperparameters are tuned, and a fast-converging and high-accuracy engine dynamic performance model is established.
This achieves high accuracy and fast convergence of the engine model across the entire envelope, reduces the number of iterations, and improves the reliability and safety of the overall engine performance.
Smart Images

Figure CN119442452B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aero-engine modeling technology, and in particular to a method for constructing a high-passflow dual-variable-cycle engine fusion model based on reinforcement learning. Background Technology
[0002] As aero-engines evolve towards higher thrust-to-weight ratios, higher performance, and greater economy, improving the reliability of control and fault diagnosis systems is crucial. Aero-engine models, as the cornerstone of this research, are a current research hotspot in the aviation field, and ensuring their real-time performance and accuracy is paramount. However, to meet the power requirements of next-generation high-altitude, high-speed aircraft, engine structures are becoming increasingly complex, posing significant challenges to the accuracy and real-time performance of engine dynamic performance models. Therefore, developing real-time dynamic modeling methods for aero-engines is becoming increasingly important.
[0003] Currently widely used engine models include component-level models, piecewise linear models, linear variable parameter models, and data-driven models. Each method has its advantages and disadvantages in practical applications. Component-level models based on aerodynamic thermodynamics can ensure the accuracy requirements of the entire flight envelope, but the complex engine structure undoubtedly increases the dimensionality of the common working equations, which greatly increases the computational burden of iterative algorithms. The application of piecewise linear models to aero-engine modeling can be traced back to 1977, when Teren used this to study the shortest acceleration time of turbofan engines. Piecewise linear models have a simple structure and powerful real-time performance simulation capabilities, but research shows that their accuracy cannot be guaranteed when dealing with transient processes with significantly abrupt changes in input parameters. To address this issue, linear variable parameter models have been applied to the field of aero-engine modeling. Based on state-space equations, linear variable parameter models can effectively correlate engine inputs and outputs, with low computational cost and good dynamic performance. However, due to individual differences in aero-engines caused by manufacturing tolerances and performance degradation, neither piecewise linear models nor linear variable parameter models can cover all flight conditions. Unlike the modeling methods mentioned above, data-driven models based on large amounts of data and a reasonable network architecture can cover the entire flight envelope and operating conditions, but ensuring their accuracy remains a challenge.
[0004] In summary, considering the modeling accuracy of the entire flight envelope, engine component-level models are undoubtedly the best choice. To ensure the real-time performance of engine component-level models under complex configurations, scholars both domestically and internationally have proposed numerous improvement methods, mainly categorized into iterative algorithms, local linearization, and fusion models. However, iterative algorithms and local linearization methods suffer from some degree of accuracy loss. Therefore, research on fusion models combined with data-driven algorithms has become a hot topic. However, existing fusion modeling methods typically require engine data covering the entire flight envelope and all operating conditions to ensure accuracy, and convergence is difficult to guarantee. Summary of the Invention
[0005] To address the limitations of traditional models, which can only perform simulations at specific points and have difficulty guaranteeing convergence under the full envelope, this invention proposes a reinforcement learning-based method for constructing a high-pass flow dual-variable cycle engine fusion model. This method aims to design the fusion model architecture by combining the Newton-Raphson algorithm and the dual-delay deep deterministic policy gradient algorithm TD3, based on a selected common working equation. It also designs the input-output structure and reward system, tunes the system hyperparameters, and ultimately establishes a dynamic performance model of the engine with fast convergence, high accuracy, and good real-time performance.
[0006] This invention provides a method for constructing a high-pass flow dual-variable cycle engine fusion model based on reinforcement learning, comprising the following steps:
[0007] Step 1: Based on the aerodynamic and thermodynamic mechanism of the high-flow dual-cycle engine, select a common working equation to achieve overall engine performance matching;
[0008] Step 2: Combining the aforementioned common working equations, design a high-throughput dual-variable-cycle engine fusion model architecture based on the Newton-Raphson algorithm and the dual-delay deep deterministic policy gradient algorithm TD3;
[0009] Step 3: Design the system input / output structure and reward system based on the working mechanism of the high-flow dual-variable-cycle engine fusion model;
[0010] Step 4: Based on the residuals of the common working equations and the designed reward system, tune the hyperparameters of the high-pass flow dual-variable cycle engine fusion model.
[0011] Optionally, in one embodiment of the present invention, step 1 specifically includes:
[0012] Step 1.1: Analyze the aerodynamic and thermodynamic working mechanism of each component of the high-flow dual variable cycle engine, and establish a mathematical model of typical components of the high-flow dual variable cycle engine based on the working mechanism of existing variable cycle engine components;
[0013] Step 1.2: Analyze the working mechanism of typical components of a high-flow dual-cycle engine, including the front fan FFan, rear fan RFan, core drive fan CDFS, high-pressure compressor Comp, high-pressure turbine Hturb, intermediate-pressure turbine Mturb, and low-pressure turbine Lturb, and determine the number of common working equations.
[0014] Step 1.3: Based on the working principles of continuous engine flow, static pressure balance, and power conservation, the common working equation is selected as follows:
[0015]
[0016] in, P represents flow rate. s The values represent static pressure, η represents rotor shaft mechanical efficiency, and N represents power. The numerical subscripts indicate engine section numbers: 131 represents the third bypass outlet, 142 represents the second bypass outlet, 16 represents the rear bypass outlet, cool242 represents low-pressure turbine guide coolant, cool261 represents intermediate-pressure turbine guide coolant, 29 represents the first bypass outlet, 311 represents the heat exchanger hot-end inlet, 4 represents the combustion chamber outlet, 41 represents the high-pressure turbine guide outlet, 43 represents the high-pressure turbine outlet, 44 represents the high-pressure turbine guide outlet, 46 represents the intermediate-pressure turbine outlet, 48 represents the dual-variable combustion chamber outlet, 49 represents the low-pressure turbine guide outlet, 6 represents the low-pressure turbine outlet, 7 represents the afterburner outlet, and 9 represents the exhaust nozzle outlet. The subscripts H, M, L, and EX represent engine accessories.
[0017] Optionally, in one embodiment of the present invention, step 2 specifically includes:
[0018] Step 2.1: Compare and analyze the mechanisms of deep reinforcement learning algorithms, and select the algorithm that meets the requirements for establishing the high-pass flow dual-variable cycle engine fusion model;
[0019] Step 2.2: Combine the dual-delay deep deterministic policy gradient algorithm TD3 and the Newton-Raphson algorithm to design the high-pass flow dual variable cycle engine fusion model architecture.
[0020] Optionally, in one embodiment of the present invention, the high-pass flow dual-variable cycle engine fusion model includes two iterative systems, inner and outer loops. When the inaccurate initial iteration parameters cause the residual e of the engine's common working equation to be greater than a specified constant ε, the outer loop system based on the TD3 agent is used for iteration; when the residual e is less than the specified constant ε, the Newton-Raphson algorithm is used for inner loop iteration.
[0021] During training, since the initial iteration parameters are randomly given and the residual e of the common working equations is greater than the specified constant ε, the high-pass flow dual-variable cycle engine fusion model is trained by connecting the Actor of the outer loop system. The action A output by the Actor is... t The state S obtained by the engine flow path calculation module t S t+1 and reward r t They are stored together in the replay buffer for use in training the Critic, which calculates the expected reward Q. t The error is updated to further optimize the Actor update strategy;
[0022] During the usage phase, when the residual e is greater than the specified constant ε, the trained Actor is called to output the initial iteration parameters until the residual e is less than the specified constant ε. At this time, the high-pass flow dual variable cycle engine fusion model connects the inner loop and is iterated by the Newton-Raphson algorithm until the high-pass flow dual variable cycle engine fusion model converges.
[0023] Optionally, in one embodiment of the present invention, step 3 specifically includes:
[0024] Step 3.1: Analyze the working mechanism of the high-pass flow dual variable cycle engine fusion model. The agent based on TD3 includes two networks: Actor and Critic. In the high-pass flow dual variable cycle engine fusion model, Actor is used to generate initial iteration parameters, and Critic is used to evaluate the quality of the current initial iteration parameters and guide the Actor network to update.
[0025] Step 3.2: Design the input / output structures of the Actor and Critic networks;
[0026] Step 3.3: The agent reward system is set up to guide the agent to generate accurate iterative parameters more quickly. Therefore, the convergence and convergence speed of the high-pass flow dual-variable cyclic engine fusion model need to be comprehensively considered. The convergence during the training phase is mainly determined by the sum of the absolute values of the residuals of the common working equations. The convergence speed is expressed as K, which is the number of iterations of the required inner-loop Newton-Raphson algorithm. t The decision is based on the flag indicating that the inner loop can converge. As a boundary, when convergence is possible Reward r t The settings are as follows:
[0027]
[0028] In this context, the subscript t represents the current time.
[0029] Optionally, in one embodiment of the present invention, the quality of the initial iteration parameters of the engine depends on the engine's operating environment E. T and the current state,
[0030] The input parameters that determine the environment of a high-flow dual-cycle engine include altitude (H), Mach number (Ma), and main combustion chamber fuel flow rate (W). f Dual-variable combustion chamber fuel flow rate W fD Afterburner fuel flow rate W fA Exhaust nozzle throat area A8, exhaust nozzle outlet area A9, intermediate pressure turbine guide vane angle A MT Low-pressure turbine guide vane angle A LT Therefore, the simulation input parameters for the high-pass flow dual-variable cycle engine fusion model are as follows:
[0031] E T =[H t Ma t W f,t W fD,t W fA,t A 8,t A 9,t A MT,t A LT,t ]
[0032] Where the subscript t represents the current time;
[0033] The current state of the high-flow dual-variable-cycle engine is represented by the residual e of the common working equation. Therefore, the input to the Actor network is:
[0034] S t =[E T e t ]
[0035] Since the output of the Actor network is the initial iteration parameter I, the state S-action A pair of the Actor network input and output is:
[0036] S t =[E T e t →A t =I t
[0037] The Critic network takes the state-action pair S and A as inputs and outputs the expected reward Q. t Therefore, the input-output structure of the Critic network is as follows:
[0038] S t ~A t →Q t .
[0039] Optionally, in one embodiment of the present invention, step 4 specifically comprises:
[0040] Step 4.1: Select the Actor network and Critic network structures based on experience;
[0041] Step 4.2: Initially define the agent's hyperparameters;
[0042] Step 4.3: Tune the system hyperparameters based on the residuals and reward magnitudes of the common working equations;
[0043] Step 4.4: Obtain the high-pass flow dual-cycle engine fusion model.
[0044] The high-passflow dual-variable-cycle engine fusion model construction method based on reinforcement learning in this invention has the following characteristics:
[0045] Beneficial effects:
[0046] 1. The common working equations of the high-flow dual-variable cycle engine selected in this invention can achieve overall engine performance matching.
[0047] 2. The fusion model architecture established by this invention does not require pre-provided training data, which solves the problem of difficulty in obtaining full envelope and full-condition engine data in traditional methods.
[0048] 3. The high-pass flow dual-variable cycle engine fusion model based on deep reinforcement learning proposed in this invention has fast convergence speed, high accuracy and good real-time performance. It can overcome the limitation of traditional models that can only be simulated at specific points, reduce the number of iterations of the model under sudden input parameter conditions and ensure convergence within the full envelope. It has a positive promoting effect on promoting the application of airborne engine models and improving the overall performance reliability and safety of the engine.
[0049] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0050] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0051] Figure 1 A flowchart illustrating a high-throughput dual-cycle engine fusion model construction method based on reinforcement learning according to an embodiment of the present invention;
[0052] Figure 2 This is a schematic diagram of the high-flow dual-variable-cycle engine structure according to an embodiment of the present invention;
[0053] Figure 3 This is a schematic diagram of the high-pass flow dual-variable cycle engine fusion model architecture according to an embodiment of the present invention;
[0054] Figure 4 This is a test diagram illustrating the applicable scope of the high-pass flow dual-variable cycle engine fusion model according to an embodiment of the present invention.
[0055] Figure 5 This is a relative error diagram of the high-pass flow dual variable cycle engine fusion model under various operating conditions according to an embodiment of the present invention. Detailed Implementation
[0056] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0057] Figure 1 This is a flowchart illustrating a method for constructing a high-pass flow dual-variable cycle engine fusion model based on reinforcement learning, according to an embodiment of the present invention.
[0058] like Figure 1 As shown, the method for constructing a high-pass flow dual-variable cycle engine fusion model based on reinforcement learning includes the following steps:
[0059] Step 1: Based on the aerodynamic and thermodynamic mechanism of the high-flow dual-cycle engine, select a common working equation to achieve overall engine performance matching.
[0060] In step 1, the aero-thermodynamic working mechanism of each component of the high-flow dual variable cycle engine is first analyzed, and the model of each component of the high-flow dual variable cycle engine is established in combination with the existing modeling mechanism of each component of the variable cycle engine.
[0061] First, based on the high-flow dual-variable cycle engine, which includes seven rotating components (front fan FFan, rear fan RFan, core drive fan CDFS, high-pressure compressor Comp, high-pressure turbine Hturb, intermediate-pressure turbine Mturb, and low-pressure turbine Lturb) and three rotating shafts, the number of common working equations is determined.
[0062] Next, based on the working principles of continuous engine flow, static pressure balance, and power conservation, the following common working equation is selected:
[0063]
[0064] in, P represents flow rate. sη represents static pressure, N represents turbine component efficiency, and N represents power. The numerical subscripts indicate engine section numbers, where 131 represents the third bypass outlet, 142 represents the second bypass outlet, 16 represents the rear bypass outlet, cool242 represents low-pressure turbine guide cooling gas, cool261 represents intermediate-pressure turbine guide cooling gas, 29 represents the first bypass outlet, 311 represents the heat exchanger hot-end inlet, 4 represents the dual-variable combustion chamber outlet, 41 represents the high-pressure turbine guide outlet, 43 represents the high-pressure turbine outlet, 44 represents the high-pressure turbine guide outlet, 46 represents the intermediate-pressure turbine outlet, 48 represents the dual-variable combustion chamber outlet, 49 represents the low-pressure turbine guide outlet, 6 represents the low-pressure turbine outlet, 7 represents the afterburner outlet, and 9 represents the exhaust nozzle outlet. The subscripts H represent the high-pressure shaft, M represent the intermediate-pressure shaft, L represent the low-pressure shaft, and EX represent engine accessories.
[0065] Step 2: Combining the common working equations, design the high-pass flow dual variable cycle engine fusion model architecture based on the Newton-Raphson algorithm and the dual-delay deep deterministic policy gradient algorithm TD3.
[0066] In step 2, we first compare and analyze the mechanism of deep reinforcement learning algorithms, and select an algorithm that meets the requirements for establishing a high-pass flow dual variable cycle engine fusion model and is suitable for complex environments and high action dimensions.
[0067] Ultimately, the dual-delay deep deterministic strategy gradient algorithm TD3 was selected, and its high-pass flow dual variable cycle engine fusion model architecture was designed in combination with the Newton-Raphson algorithm.
[0068] The designed high-throughput dual-variable-cycle engine fusion model architecture includes two iterative systems: an inner loop and an outer loop. When the initial iteration parameters are inaccurate, resulting in a large residual e in the engine's common working equations, exceeding the specified constant ε, the outer loop system based on the TD3 agent is used for iteration. When the residual e is less than the specified constant ε, the Newton-Raphson algorithm is used for inner loop iteration.
[0069] During training, due to the randomized initial iteration parameters and the large residuals of the common working equations exceeding the specified constant ε, the Actor connecting the outer loop system of the fusion model is trained. The Actor outputs action A. t The state S obtained by the engine flow path calculation module t S t+1 and reward r t These are stored together in the replay buffer for use in Critic's training. Critic calculates the expected reward Q. t The error is updated to further optimize the Actor's update strategy.
[0070] During the usage phase, when the residual is large and exceeds the specified constant ε, the trained Actor is called to output the initial iteration parameters until the residual e is less than the specified constant ε. At this point, the high-pass flow dual variable cycle engine fusion model connects the inner loop and is iterated by the Newton-Raphson algorithm until the high-pass flow dual variable cycle engine fusion model converges.
[0071] Step 3: Design the system input / output structure and reward system based on the working mechanism of the high-pass flow dual variable cycle engine fusion model.
[0072] In step 3, the working mechanism of the fusion model is analyzed. It can be seen that the agent based on TD3 includes two networks: Actor and Critic. In the fusion model, Actor is mainly used to generate initial iteration parameters, and Critic is mainly used to evaluate the quality of the current initial iteration parameters and guide the Actor network to update.
[0073] The input-output structures of the Actor and Critic networks were designed, as the quality of the initial iterative parameters of the engine largely depends on the engine's operating environment E. T and the current state.
[0074] The main input parameters that determine the environment of a high-flow dual-cycle engine include altitude (H), Mach number (Ma), and main combustion chamber fuel flow rate (W). f Dual-variable combustion chamber fuel flow rate W fD Afterburner fuel flow rate W fA Exhaust nozzle throat area A8, exhaust nozzle outlet area A9, intermediate pressure turbine guide vane angle A MT Low-pressure turbine guide vane angle A LT Therefore, the environmental parameters are as follows:
[0075] E T =[H t Ma t W f,t W fD,t W fA,t A 8,t A 9,t A MT,t A LT,t ]
[0076] Where the subscript t represents the current time;
[0077] The current state of the high-flow dual-variable-cycle engine is mainly represented by the residual e of the common working equation; therefore, the input to the Actor network is:
[0078] S t =[E T e t ]
[0079] Since the output of the Actor network is the initial iteration parameter I, the state S-action A pair of the Actor network input and output is:
[0080] S t =[E T e t →A t =I t
[0081] The Critic network takes the state-action pair S and A as inputs and outputs the expected reward Q. t Therefore, the input-output structure of the Critic network is as follows:
[0082] S t ~A t →Q t
[0083] Finally, the reward system for the agent is primarily designed to guide the agent to generate accurate iterative parameters more quickly; therefore, the model's convergence and convergence speed must be considered comprehensively. The convergence during the training phase is mainly determined by the sum of the absolute values of the residuals of the common working equations. The convergence speed is mainly determined by the number of iterations K required for the inner-loop Newton-Raphson algorithm. t Decision. The flag indicating that the inner loop can converge. As a boundary, when convergence is possible Reward r t The settings are as follows:
[0084]
[0085] Step 4: Based on the residuals of the common working equations and the designed reward system, tune the hyperparameters of the high-pass flow dual-cycle engine fusion model.
[0086] In step 4, the Actor network and Critic network structures are first selected based on experience; then the agent hyperparameters are initially given; next, the system hyperparameters are tuned based on the residuals and reward size of the common working equations; finally, the high-pass flow dual-variable cycle engine fusion model is obtained.
[0087] This invention focuses on a high-pass flow dual-variable cycle engine and, combining the powerful information extraction capabilities and environmental interaction features of deep reinforcement learning, establishes a fusion model based on the dual-delay deep deterministic policy gradient (TD3) algorithm. This model exhibits good real-time performance, high accuracy, and strong convergence. It overcomes the limitation of traditional models that can only perform simulations at specific points, reduces the number of iterations under abrupt input parameters, and ensures convergence across the entire input envelope.
[0088] The embodiment of this invention is a high-flow dual-variable-cycle engine, which contains three rotating components: a high-pressure shaft, an intermediate-pressure shaft, and a low-pressure shaft. Its structural schematic diagram is shown below. Figure 2 As shown.
[0089] In one specific embodiment, step 1 is as follows:
[0090] First, the aero-thermodynamic working mechanism of each component of the high-flow dual variable cycle engine is analyzed, and the model of each component of the high-flow dual variable cycle engine is established in combination with the existing modeling mechanism of each component of the variable cycle engine.
[0091] Next, the high-flow dual-cycle engine is analyzed, including seven rotating components (front fan FFan, rear fan RFan, core drive fan CDFS, high-pressure compressor Comp, high-pressure turbine Hturb, intermediate-pressure turbine Mturb, and low-pressure turbine Lturb) and three shafts, and the number of common working equations is determined to be 10.
[0092] Finally, based on the working principles of continuous engine flow, static pressure balance, and power conservation, the following common working equation is selected:
[0093]
[0094] in, P represents flow rate. s η represents static pressure, N represents turbine component efficiency, and N represents power. The numerical subscripts indicate engine section numbers, where 131 represents the third bypass outlet, 142 represents the second bypass outlet, 16 represents the rear bypass outlet, cool242 represents low-pressure turbine guide cooling gas, cool261 represents intermediate-pressure turbine guide cooling gas, 29 represents the first bypass outlet, 311 represents the heat exchanger hot-end inlet, 4 represents the dual-variable combustion chamber outlet, 41 represents the high-pressure turbine guide outlet, 43 represents the high-pressure turbine outlet, 44 represents the high-pressure turbine guide outlet, 46 represents the intermediate-pressure turbine outlet, 48 represents the dual-variable combustion chamber outlet, 49 represents the low-pressure turbine guide outlet, 6 represents the low-pressure turbine outlet, 7 represents the afterburner outlet, and 9 represents the exhaust nozzle outlet. The subscripts H represent the high-pressure shaft, M represent the intermediate-pressure shaft, L represent the low-pressure shaft, and EX represent engine accessories.
[0095] In one specific embodiment, step 2 is as follows:
[0096] First, we compared and analyzed the mechanisms of deep reinforcement learning algorithms, selected an algorithm that meets the requirements for establishing a high-throughput dual-variable-cycle engine fusion model, is suitable for complex environments and high action dimensions, and finally selected the TD3 algorithm.
[0097] Next, combining the dual-delay deep deterministic policy gradient algorithm TD3 and the Newton-Raphson algorithm, a high-pass flow dual-variable cycle engine fusion model architecture is designed as follows: Figure 3 As shown.
[0098] The fusion model comprises two iterative systems, an inner loop and an outer loop. When inaccurate initial iterative parameters result in a large residual e in the engine's common working equation, the selection valve opens upwards, connecting the Actor network and using the outer loop system based on the TD3 agent for iteration. When the initial iterative variables given by the Actor satisfy the condition that the flow path calculation residual e is less than a specified constant ε, the selection valve opens downwards, and the Newton-Raphson algorithm is used for inner loop iteration.
[0099] During training, due to the randomized initial iteration parameters and the large residuals of the common working equations, the Actor connecting the outer loop system of the fusion model is trained. The Actor outputs action A. t The state S obtained by the engine flow path calculation module t S t+1 and reward r t These are stored together in the replay buffer for use in Critic's training. Critic calculates the expected reward Q. t The error is updated to further optimize the Actor's update strategy.
[0100] During the usage phase, when the residual is large, the trained Actor is called to output the initial iteration parameters until the residual e is less than the specified constant ε. At this point, the model connects the inner loop and is iterated by the Newton-Raphson algorithm until the high-pass flow dual variable cycle engine fusion model converges.
[0101] In one specific embodiment, step 3 is as follows:
[0102] In step 3, the working mechanism of the fusion model is analyzed. It can be seen that the agent based on TD3 includes two networks: Actor and Critic. In the fusion model, Actor is mainly used to generate initial iteration parameters, and Critic is mainly used to evaluate the quality of the current initial iteration parameters and guide the Actor network to update.
[0103] The input-output structures of the Actor and Critic networks were designed, as the quality of the initial iterative parameters of the engine largely depends on the engine's operating environment E. T and the current state.
[0104] The main input parameters that determine the environment of a high-flow dual-cycle engine include altitude (H), Mach number (Ma), and main combustion chamber fuel flow rate (W). f Dual-variable combustion chamber fuel flow rate W fD Afterburner fuel flow rate W fA Exhaust nozzle throat area A8, exhaust nozzle outlet area A9, intermediate pressure turbine guide vane angle A MT Low-pressure turbine guide vane angle ALT Therefore, the environmental parameters are as follows:
[0105] E T =[H t Ma t W f,t W fD,t W fA,t A 8,t A 9,t A MT,t A LT,t ]
[0106] In this context, the subscript t represents the current time.
[0107] The current state of the high-flow dual-variable-cycle engine is mainly represented by the residual e of the common working equation; therefore, the input to the Actor network is:
[0108] S t =[E T e t ]
[0109] Since the output of the Actor network is the initial iteration parameter I, the state S-action A pair of the Actor network input and output is:
[0110] S t =[E T e t →A t =I t
[0111] The Critic network takes the state-action pair S and A as inputs and outputs the expected reward Q. t Therefore, the input-output structure of the Critic network is as follows:
[0112] S t ~A t →Q t
[0113] Finally, the reward system for the agent is primarily designed to guide the agent to generate accurate iterative parameters more quickly; therefore, the model's convergence and convergence speed must be considered comprehensively. The convergence during the training phase is mainly determined by the sum of the absolute values of the residuals of the common working equations. The convergence speed is mainly determined by the number of iterations K required for the inner-loop Newton-Raphson algorithm. t Decision. The flag indicating that the inner loop can converge. As a boundary, when convergence is possible Reward r t The settings are as follows:
[0114]
[0115] In one specific embodiment, step 4 is as follows:
[0116] First, based on experience, the Actor network is selected to contain two hidden layers with 400 and 300 nodes respectively; the Critic network structure's state path contains two hidden layers with 400 and 300 nodes respectively, the action path contains one hidden layer with 300 nodes, and the common path contains an addition layer and an output layer.
[0117] After initially determining the agent's hyperparameters, the system hyperparameters are tuned based on the residuals of the common working equations and the reward magnitude. The final selected fusion model parameters are shown in Table 1.
[0118] Table 1 Hyperparameters of the fusion model
[0119]
[0120] The high-pass flow dual-variable cycle engine fusion model was finally obtained through training.
[0121] Simulation results Figure 4 This indicates that the established fusion model can cover almost 70%-100% of the maximum speed within the full envelope, i.e., from idle speed to maximum speed, demonstrating its ability to operate at various typical operating points for safe engine operation. Figure 5 As can be seen, the relative errors of the key parameters of the fusion model in the slow speed state, subsonic cruise state and maximum afterburner state are all no more than 1.5%, indicating high model accuracy. When the real-time performance of the fusion model was tested on a host computer with a CPU i7-14700 and a main frequency of 2.10GHz, the average single-step simulation time was no more than 0.157ms. When the real-time performance was tested on a P2020 development board with a main frequency of 1GHz, the average single-step simulation time of the model was no more than 2.949ms.
[0122] The high-pass flow dual-variable cycle engine fusion model constructed by the reinforcement learning-based method proposed in this embodiment of the invention has fast convergence speed, high accuracy and good real-time performance. It can overcome the limitation of traditional models that can only be simulated at specific points, reduce the number of iterations of the model under abrupt input parameters and ensure convergence within the entire envelope.
[0123] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0124] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0125] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
Claims
1. A method for constructing a fusion model of a high-passflow dual-variable-cycle engine based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Based on the aerodynamic and thermodynamic mechanism of the high-flow dual-cycle engine, select a common working equation to achieve overall engine performance matching; Step 2: Combining the aforementioned common working equations, design a high-throughput dual-variable-cycle engine fusion model architecture based on the Newton-Raphson algorithm and the dual-delay deep deterministic policy gradient algorithm TD3; Step 3: Design the system input / output structure and reward system based on the working mechanism of the high-flow dual-variable-cycle engine fusion model; Step 4: Based on the residuals of the common working equations and the designed reward system, tune the hyperparameters of the high-pass flow dual-variable cycle engine fusion model; The high-throughput dual-cycle engine fusion model includes two iterative systems, an inner loop and an outer loop. When the inaccurate initial iteration parameters cause the residual e of the engine's common working equation to be greater than the specified constant ε, the outer loop system based on the TD3 agent is used for iteration; when the residual e is less than the specified constant ε, the Newton-Raphson algorithm is used for inner loop iteration. During training, since the initial iteration parameters are randomly given and the residual e of the common working equation is greater than the specified constant ε, the high-pass flow dual-variable cycle engine fusion model is trained by connecting the Actor of the outer loop system. The action A output by the Actor is... t The state S obtained by the engine flow path calculation module t S t+1 and reward r t They are stored together in the replay buffer for use in training the Critic, which calculates the expected reward Q. t The error is updated to further optimize the Actor update strategy; During the usage phase, when the residual e is greater than the specified constant ε, the trained Actor is called to output the initial iteration parameters until the residual e is less than the specified constant ε. At this time, the high-pass flow dual variable cycle engine fusion model connects the inner loop and is iterated by the Newton-Raphson algorithm until the high-pass flow dual variable cycle engine fusion model converges.
2. The method according to claim 1, characterized in that, Step 1 specifically includes: Step 1.1: Analyze the aerodynamic and thermodynamic working mechanism of each component of the high-flow dual variable cycle engine, and establish a mathematical model of typical components of the high-flow dual variable cycle engine based on the working mechanism of existing variable cycle engine components; Step 1.2: Analyze the working mechanism of typical components of a high-flow dual-cycle engine, including the front fan FFan, rear fan RFan, core drive fan CDFS, high-pressure compressor Comp, high-pressure turbine Hturb, intermediate-pressure turbine Mturb, and low-pressure turbine Lturb, and determine the number of common working equations. Step 1.3: Based on the working principles of continuous engine flow, static pressure balance, and power conservation, the common working equation is selected as follows: in, P represents flow rate. s The values represent static pressure, η represents rotor shaft mechanical efficiency, and N represents power. The numerical subscripts indicate engine section numbers: 131 represents the third bypass outlet, 142 represents the second bypass outlet, 16 represents the rear bypass outlet, cool242 represents low-pressure turbine guide coolant, cool261 represents intermediate-pressure turbine guide coolant, 29 represents the first bypass outlet, cool311 represents the heat exchanger hot-end inlet, 4 represents the combustion chamber outlet, 41 represents the high-pressure turbine guide outlet, 43 represents the high-pressure turbine outlet, 44 represents the intermediate-pressure turbine guide outlet, 46 represents the intermediate-pressure turbine outlet, 48 represents the dual-variable combustion chamber outlet, 49 represents the low-pressure turbine guide outlet, 6 represents the low-pressure turbine outlet, 7 represents the afterburner outlet, and 9 represents the exhaust nozzle outlet. The subscripts H, M, L, and EX represent engine accessories.
3. The method according to claim 1, characterized in that, Step 2 specifically includes: Step 2.1: Compare and analyze the mechanisms of deep reinforcement learning algorithms, and select the algorithm that meets the requirements for establishing the high-pass flow dual-variable cycle engine fusion model; Step 2.2: Combine the dual-delay deep deterministic policy gradient algorithm TD3 and the Newton-Raphson algorithm to design the high-pass flow dual variable cycle engine fusion model architecture.
4. The method according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1: Analyze the working mechanism of the high-pass flow dual variable cycle engine fusion model. The agent based on TD3 includes two networks: Actor and Critic. In the high-pass flow dual variable cycle engine fusion model, Actor is used to generate initial iteration parameters, and Critic is used to evaluate the quality of the current initial iteration parameters and guide the Actor network to update. Step 3.2: Design the input / output structures of the Actor and Critic networks; Step 3.3: The agent reward system is set up to guide the agent to generate accurate iterative parameters more quickly. Therefore, the convergence and convergence speed of the high-pass flow dual-variable cyclic engine fusion model need to be comprehensively considered. The convergence during the training phase is determined by the sum of the absolute values of the residuals of the common working equations. The convergence speed is expressed as K, which is the number of iterations of the required inner-loop Newton-Raphson algorithm. t The decision is based on the flag indicating that the inner loop can converge. As a boundary, when convergence is possible Reward r t The settings are as follows: In this context, the subscript t represents the current time.
5. The method according to claim 4, characterized in that, The quality of the engine's initial iteration parameters depends on the engine's operating environment E. T and the current state, The input parameters that determine the environment of a high-flow dual-cycle engine include altitude (H), Mach number (Ma), and main combustion chamber fuel flow rate (W). f Dual-variable combustion chamber fuel flow rate W fD Afterburner fuel flow rate W fA Exhaust nozzle throat area A8, exhaust nozzle outlet area A9, intermediate pressure turbine guide vane angle A MT Low-pressure turbine guide vane angle A LT Therefore, the simulation input parameters for the high-pass flow dual-variable cycle engine fusion model are as follows: HAVE BEEN T =[H t Ma t W f,t W fD,t W fA,t A 8,t A 9,t A MT,t A LT,t ] Where the subscript t represents the current time; The current state of the high-flow dual-variable-cycle engine is represented by the residual e of the common working equation. Therefore, the input to the Actor network is: S t =[E T yes t ] Since the output of the Actor network is the initial iteration parameter I, the state S-action A pair of the Actor network input and output is: S t =[E T yes t ]→A t =I t The Critic network takes the state-action pair S and A as inputs and outputs the expected reward Q. t Therefore, the input-output structure of the Critic network is as follows: S t ~A t →Q t 。 6. The method according to claim 5, characterized in that, Step 4 is as follows: Step 4.1: Select the Actor network and Critic network structures based on experience; Step 4.2: Initially define the agent's hyperparameters; Step 4.3: Tune the system hyperparameters based on the residuals and reward magnitudes of the common working equations; Step 4.4: Obtain the high-pass flow dual-cycle engine fusion model.
Citation Information
Patent Citations
Aero-engine direct thrust control method based on reinforcement learning
CN115840354A
Reactive power optimization method based on double-delay depth deterministic strategy gradient
CN116468159A