Engine multi-parameter self-adjusting model modeling method, system, medium and equipment

Through a self-adjusting module based on reinforcement learning, parameters such as fuel flow and nozzle area are corrected, which solves the problem of model accuracy being affected by input errors and component degradation in traditional modeling methods, and achieves higher engine state parameter estimation accuracy.

CN119442877BActive Publication Date: 2025-09-05XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411509003.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-09-05
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

In traditional component-level modeling methods, the estimation accuracy of the engine model's state parameters is affected by the measurement error of the engine input parameters and component degradation, resulting in larger modeling errors.

Method used

A reinforcement learning-based engine multi-parameter self-adjustment model modeling method is adopted. The self-adjustment module calculates the parameter adjustment coefficient according to the residual between the component-level model output parameters and the sensor measurement values, corrects parameters such as fuel flow rate and nozzle area, and constructs online strategies and value networks for training to improve model accuracy.

Benefits of technology

The impact of engine input measurement uncertainty and component degradation on model accuracy is reduced, and the estimation accuracy of engine state parameters is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119442877B_ABST
    Figure CN119442877B_ABST
Patent Text Reader

Abstract

A reinforcement learning-based engine multi-parameter self-adjusting model modeling method, system, medium and equipment, in which the engine input parameters and gas path parameters are collected; the action is adjusted according to the engine input parameters and the parameters of the self-adjusting module a t , calculate the gas path parameters of the engine component-level model; calculate the residual signal between the component-level model after parameter adjustment and the engine gas path parameters, and record the residual signal of the engine gas path parameters as the state vector s t ; Design a reward function to calculate the reward value obtained after each parameter adjustment r t ;Store experience data pairs ( s t , a t , r t , s t+1 ) to build an experience pool; during the self-adjustment module application phase, the trained policy network is used to calculate the corresponding parameter adjustment actions based on the state vector, participate in the calculation of the component-level model, and predict the engine gas path parameters. This invention improves the modeling accuracy of the engine model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aero-engine modeling, and in particular to a reinforcement learning-based engine multi-parameter self-adjusting modeling method, system, medium and equipment. Background Art

[0002] Performance prediction, fault-tolerant control, and health management are key areas of aircraft engine research and development, and are also key technologies for ensuring safe and efficient engine operation. Building an accurate and reliable engine model is particularly critical. The accuracy of engine state parameter estimation using traditional component-level modeling methods is affected by the metering errors of engine input parameters and component degradation. 1) Fuel flow rate and nozzle area are important input parameters for engine component-level models, and their metering errors can affect the accuracy of the engine model's output parameters. 2) As the engine's service life increases, the performance of the actual engine's gas path components degrades, increasing the modeling error between the model and the actual engine, which in turn affects the accuracy of the engine state parameter estimation.

[0003] The above information disclosed in this Background section is only for enhancement of understanding of the background of the invention and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0004] The present invention provides a reinforcement learning-based engine multi-parameter self-adjusting model modeling method, system, medium and equipment, which improves the estimation accuracy of engine state parameters by adaptively adjusting engine input parameters and component performance parameters.

[0005] The modeling method of the engine multi-parameter self-tuning model based on reinforcement learning includes:

[0006] Step 100: Collect the engine input parameters I E (t) and the engine's gas path parameters O E (t);

[0007] Step 200: construct an engine component level model with a self-adjusting module based on the engine, and E (t) and the parameter adjustment action a of the self-adjusting module t , calculate the gas path parameters of the engine component level model, where the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients to adjust the input parameters and component characteristic parameters of the engine component level model. The engine component level model formula is expressed as , where O M (t) is the gas path parameter output by the engine model, f M is the function representing the engine model;

[0008] Step 300: Calculate the gas path parameters of the engine component level model after parameter adjustment. M (t) and the engine's gas path parameters O E (t), the residual signal is recorded as the state vector s t ;

[0009] Step 400: Design a reward function to calculate the reward value r obtained after each parameter adjustment. t ;

[0010] Step 500, store the empirical data pair (s t ,a t ,r t ,s t+1 ), build an experience pool;

[0011] Step 600: construct an online policy network μ(s t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network;

[0012] Step 700: construct an online value network q(s) in the self-adjusting module. t ,a t ; w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network;

[0013] Step 800: In the self-adjustment module training phase, the online policy network μ(s t ;θ) and the online value network q(s t ,a t ;w) conduct training;

[0014] Step 900, in the self-adjustment module application phase, the trained online strategy network μ(s t ;θ) Calculate the corresponding parameter adjustment action based on the state vector, participate in the calculation of the engine component level model, and estimate the engine state parameters.

[0015] In the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 100, the engine input parameter I E (t)={T t0 (t), P t0 (t), W f (t), A8(t)}, where T t0(t) is the total intake air temperature, P t0 (t) is the total intake pressure, W f (t) is the fuel flow rate, A8(t) is the nozzle area, and the gas path parameters of the engine are E (t)={N L,E (t),N H,E (t), P t25,E (t), P t31,E (t), T t6,E (t), P t6,E (t)}, where N L,E (t) is the measured value of the low-pressure rotor speed, N H,E (t) is the measured value of the high-pressure rotor speed, P t25,E (t) is the measured value of the pressure at the outlet of the low-pressure compressor, P t31,E (t) is the measured value of the pressure at the high pressure compressor outlet, T t6,E (t) is the measured value of the temperature at the outlet of the low-pressure turbine, P t6,E (t) is the measured value of the pressure at the outlet of the low-pressure turbine.

[0016] In the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, the gas path parameters O of the engine component level model M (t)={N L,M (t), N H,M (t), P t25,M (t), P t31,M (t), T t6,M (t), P t6,M (t)},N L,M (t) is the predicted value of low-pressure rotor speed calculated by the model, N H,M (t) is the predicted value of high-pressure rotor speed calculated by the model, P t25,M (t) is the predicted value of the pressure at the outlet of the low-pressure compressor calculated by the model, P t31,M (t) is the predicted value of the high pressure compressor outlet pressure calculated by the model, T t6,M (t) is the predicted value of the temperature at the outlet of the low-pressure turbine calculated by the model, P t6,M (t) is the predicted value of the pressure at the low-pressure turbine outlet calculated by the model.

[0017] In the engine multi-parameter self-adjustment model modeling method based on reinforcement learning, the vector k(t) composed of the model parameter adjustment coefficients of the engine component level model includes the fuel flow adjustment coefficient , nozzle area adjustment coefficient , Fan flow adjustment coefficient , Fan efficiency adjustment factor , High-pressure compressor flow adjustment coefficient , High-pressure compressor efficiency adjustment coefficient , High-pressure turbine flow adjustment coefficient , High-pressure turbine efficiency adjustment factor , low-pressure turbine flow adjustment coefficient and low-pressure turbine efficiency adjustment factor .

[0018] In the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 300, the state vector s t It is composed of the residual between the calculated value of the gas path parameter of the engine component level model and the measured value of the actual engine gas path parameter, which can be expressed as .

[0019] In the reinforcement learning-based engine multi-parameter self-tuning model modeling method, in step 400, the formula of the reward function is: ,in,

[0020] is the reward value corresponding to the calculated value of the i-th gas path parameter at time t, and the calculation formula is ;

[0021] in, is the absolute value of the relative error of the i-th gas path parameter, and the calculation formula is: , is a constant to ensure the stability of numerical calculations.

[0022] In the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 500, the empirical data pair (s t ,a t ,r t ,s t+1 ) is represented by the state vector s at time t t , parameter adjustment action a at time t t , reward value r at time t t , execute parameter adjustment action a t The state vector s formed at time t+1 is then entered t+1 constitute.

[0023] In the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 600, the online strategy network μ(s t ;θ) includes input layer, intermediate layer and output layer. The input layer receives the state vector s t , the output layer output parameter adjustment action a t .

[0024] In the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 700, the online value network q(s t ,a t ;w) includes input layer, intermediate layer and output layer, the input layer receives the state vector s t and parameter adjustment action a t , output layer output (s t ,a t ) corresponds to the value Q t .

[0025] In the reinforcement learning-based engine multi-parameter self-tuning model modeling method, in step 800, the self-tuning module training phase includes the following steps:

[0026] Step 801: Use the deep deterministic policy gradient algorithm DDPG to build an online policy network μ(s t ;θ), online value network q(s t ,a t ;w), target strategy network μ T (s t ;θ T ), target value network q T (s t ,a t ;w T ),in,

[0027] Online strategy network μ(s t ;θ) and the number of neurons in the input layer and the state vector s t The dimension is the same as that of the action vector a. t The dimensions are the same,

[0028] Online value network q(s t ,a t ;w) The number of neurons in the input layer is the state vector s t The dimension of the action vector a t The sum of the dimensions of , the number of neurons in the output layer is 1,

[0029] Target policy network μ T (s t ;θ T ) and the online policy network μ(s t ;θ) has the same network structure,

[0030] Target value network q T (s t ,a t ;w T ) and the online value network q(s t ,a t;w) has the same network structure,

[0031] And initialize the online strategy network parameters θ, online value network parameters w, target strategy network parameters θ T , target value network parameter w T , the learning rate α of the online policy network parameters, the learning rate β of the online value network parameters, the discount factor and update coefficients ;

[0032] Step 802: Observe the state vector s corresponding to the residual signal of the engine gas path parameter at time t t , calculate the online policy network μ(s t ;θ t ) corresponding parameter adjustment action a t ;

[0033] Step 803: Let the engine component level model perform parameter adjustment action a t , so that the parameter adjustment coefficients are involved in the calculation of the engine component level model;

[0034] Step 804: Calculate the reward value r at time t t , observe the predicted value of the gas path parameters at time t+1, and calculate the next corresponding state vector s t+1 , the empirical data (s t ,a t ,r t ,s t+1 ) is stored in the experience pool;

[0035] Step 805: collect N experience data pairs (s i ,a i ,r i ,s i+1 ), calculate the target value y i

[0036] ;

[0037] Step 806: Minimize the loss function by gradient descent , update the online value network parameter w, the update formula is ,in, is the online value network parameter of the k+1th iteration, is the online value network parameter of the kth iteration, The loss function is The gradient at

[0038] Step 807: Maximize the scoring function by gradient ascent method , update the online policy network parameters θ, the update formula is ,in, is the online policy network parameter of the k+1th iteration, is the online policy network parameter of the kth iteration, The loss function is The gradient at

[0039] Step 808: Update the parameters of the target value network and the target strategy network. ;

[0040] Step 809: Set the network in the self-tuning module to perform 1000 rounds of training, repeat steps 802 to 808, and the training ends when the number of training rounds reaches the set value.

[0041] In the reinforcement learning-based engine multi-parameter self-tuning model modeling method, in step 900, the self-tuning module application stage includes the following steps:

[0042] Step 901: Observe the engine gas path parameter residual signal corresponding to the state vector s at time t t , calculate the target policy network μ T (s t ;θ T ) corresponding parameter adjustment action a t ;

[0043] Step 902: Let the engine component level model perform parameter adjustment action a t , so that the parameter adjustment coefficient participates in the calculation of the engine component level model and predicts the engine gas path parameters,

[0044] Step 903: repeat steps 901 to 902 until the parameter prediction process is completed.

[0045] A reinforcement learning-based engine multi-parameter self-tuning model modeling system includes:

[0046] The acquisition unit is used to: acquire the input parameters I of the engine E (t) and the engine's gas path parameters O E (t);

[0047] The model building unit is used to: build an engine component level model with a self-tuning module based on the engine, according to the input parameters I of the engine E (t) and the parameter adjustment action a of the self-adjusting module t , calculate the gas path parameters of the engine component level model, where the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients to adjust the input parameters and component characteristic parameters of the engine component level model. The engine component level model formula is expressed as , where O M (t) is the gas path parameter output by the engine model, f M is the function representing the engine model;

[0048] The residual unit is used to calculate the gas path parameters of the engine component level model after parameter adjustment and the gas path parameters of the engine. E (t), the residual signal is recorded as the state vector s t ;

[0049] The reinforcement learning machine includes a self-adjustment module, wherein the reinforcement learning machine designs a reward function for calculating the reward value r obtained after each parameter adjustment. t , store the empirical data pairs (s t ,a t ,r t ,s t+1 ), build the experience pool; build the online strategy network μ(s) in the self-adjusting module t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network, and the online value network q(s t ,a t ; w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network; in the self-adjusting module training phase, the experience data pairs in the experience pool are used to adjust the online policy network μ(s t ;θ) and the online value network q(s t ,a t ;w) conduct training;

[0050] The adjustment coefficient prediction unit uses the trained online strategy network μ(s t ;θ) calculates the corresponding parameter adjustment action based on the state vector and participates in the calculation of the engine component level model.

[0051] A computer storage medium includes computer instructions, which, when executed on a computer, cause the computer to execute the method described above.

[0052] An electronic device, characterized in that the electronic device comprises:

[0053] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein:

[0054] When the processor executes the program, the method described is implemented.

[0055] Compared with the existing technology, the present invention has the following advantages: the present invention constructs a self-adjustment module, calculates the parameter adjustment coefficient according to the residual between the component-level model output parameter and the sensor measurement value, and corrects the fuel flow, nozzle area and gas path component performance parameters of the component-level model, thereby improving the modeling accuracy of the engine model and reducing the susceptibility of the accuracy to the uncertainty of engine input measurement and component degradation. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Various other advantages and benefits of the present invention will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The accompanying drawings are intended only to illustrate preferred embodiments and are not to be construed as limiting the present invention. It should be understood that the drawings described below are merely examples of the present invention, and that those skilled in the art will be able to derive other drawings from these drawings without inventive effort. Throughout the drawings, identical reference numerals are used to denote identical components.

[0057] In the attached figure:

[0058] Figure 1 This is an architecture diagram of a reinforcement learning-based engine multi-parameter self-tuning model modeling method provided by one embodiment of the present disclosure;

[0059] Figure 2 This is a schematic diagram of a reward value change curve during the training process of the self-adjusting module provided by one embodiment of the present disclosure;

[0060] FIG3( a ) shows an embodiment of the present disclosure, wherein the engine low-pressure rotor speed N is used and the non-used parameter self-adjusting module is used. L Schematic diagram of the comparison;

[0061] FIG3( b ) shows an embodiment of the present disclosure, wherein the engine high pressure rotor speed N is used and the non-used parameter self-adjusting module is used. H Schematic diagram of the comparison;

[0062] FIG3( c ) shows an embodiment of the present disclosure, wherein the engine compressor pressure P is used and not used. t31 Schematic diagram of the comparison;

[0063] FIG3( d ) shows an embodiment of the present disclosure, wherein the engine low-pressure turbine after-temperature T is adopted and not adopted the parameter self-adjusting module. t6 Schematic comparison diagram.

[0064] The present invention will be further explained below with reference to the accompanying drawings and embodiments. DETAILED DESCRIPTION

[0065] Specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although specific embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0066] It should be noted that certain words are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. This specification and claims do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of the components as the criterion for distinction. As mentioned throughout the specification and claims, "including" or "comprising" is an open term, so it should be interpreted as "including but not limited to". The subsequent description of the specification is a preferred embodiment of the present invention, but the description is based on the general principles of the specification and is not intended to limit the scope of the invention. The scope of protection of the present invention shall be as defined in the attached claims.

[0067] To facilitate understanding of the embodiments of the present invention, further explanation will be given below using specific embodiments as examples in conjunction with the accompanying drawings, and the accompanying drawings do not constitute a limitation on the embodiments of the present invention.

[0068] like Figures 1 to 3(d) As shown, the engine multi-parameter self-tuning model modeling method based on reinforcement learning includes the following steps:

[0069] Step 100: Collect the engine input parameters I E (t) and the engine's gas path parameters O E (t);

[0070] Step 200: construct an engine component level model with a self-adjusting module based on the engine, and E (t) and the parameter adjustment action a of the self-adjusting module t , calculate the gas path parameters of the engine component level model, where the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients to adjust the input parameters and component characteristic parameters of the engine component level model. The engine component level model formula is expressed as , where O M (t) is the gas path parameter output by the engine model, f M is the function representing the engine model;

[0071] Step 300: Calculate the gas path parameters of the engine component level model after parameter adjustment. M(t) and the engine's gas path parameters O E (t), the residual signal is recorded as the state vector s t ;

[0072] Step 400: Design a reward function to calculate the reward value r obtained after each parameter adjustment. t ;

[0073] Step 500, store the empirical data pair (s t ,a t ,r t ,s t+1 ), build an experience pool;

[0074] Step 600: construct an online policy network μ(s t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network;

[0075] Step 700: construct an online value network q(s) in the self-adjusting module. t ,a t ; w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network;

[0076] Step 800: In the self-adjustment module training phase, the online policy network μ(s t ;θ) and the online value network q(s t ,a t ;w) conduct training;

[0077] Step 900, in the self-adjustment module application phase, the trained online strategy network μ(s t ;θ) Calculate the corresponding parameter adjustment action based on the state vector, participate in the calculation of the engine component level model, and estimate the engine state parameters.

[0078] In a preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 100, the engine input parameter I E (t)={T t0 (t), P t0 (t), W f (t), A8(t)}, where T t0 (t) is the total intake air temperature, P t0 (t) is the total intake pressure, W f(t) is the fuel flow rate, A8(t) is the nozzle area, and the gas path parameters of the engine are E (t)={N L,E (t), N H,E (t), P t25,E (t), P t31,E (t), T t6,E (t), P t6,E (t)}, where N L,E (t) is the measured value of the low-pressure rotor speed, N H,E (t) is the measured value of the high-pressure rotor speed, P t25,E (t) is the measured value of the pressure at the outlet of the low-pressure compressor, P t31,E (t) is the measured value of the pressure at the high pressure compressor outlet, T t6,E (t) is the measured value of the temperature at the outlet of the low-pressure turbine, P t6,E (t) is the measured value of the pressure at the outlet of the low-pressure turbine.

[0079] In the preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, the gas path parameters O of the engine component level model are M (t)={N L,M (t), N H,M (t), P t25,M (t), P t31,M (t), T t6,M (t),P t6,M (t)},N L,M (t) is the predicted value of low-pressure rotor speed calculated by the model, N H,M (t) is the predicted value of high-pressure rotor speed calculated by the model, P t25,M (t) is the predicted value of the pressure at the outlet of the low-pressure compressor calculated by the model, P t31,M (t) is the predicted value of the high pressure compressor outlet pressure calculated by the model, T t6,M (t) is the predicted value of the temperature at the outlet of the low-pressure turbine calculated by the model, P t6,M (t) is the predicted value of the pressure at the low-pressure turbine outlet calculated by the model.

[0080] In the preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, the vector k(t) composed of the model parameter adjustment coefficients of the engine component level model includes the fuel flow adjustment coefficient , nozzle area adjustment coefficient , Fan flow adjustment coefficient , Fan efficiency adjustment factor , High-pressure compressor flow adjustment coefficient , High-pressure compressor efficiency adjustment coefficient , High-pressure turbine flow adjustment coefficient , High-pressure turbine efficiency adjustment factor , low-pressure turbine flow adjustment coefficient and low-pressure turbine efficiency adjustment factor .

[0081] In the preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 300, the state vector s t It is composed of the residual between the calculated value of the gas path parameter of the engine component level model and the measured value of the actual engine gas path parameter, which can be expressed as .

[0082] In a preferred embodiment of the reinforcement learning-based engine multi-parameter self-tuning model modeling method, in step 400, the formula of the reward function is: ,in,

[0083] is the reward value corresponding to the calculated value of the i-th gas path parameter at time t, and the calculation formula is ;

[0084] in, is the absolute value of the relative error of the i-th gas path parameter, and the calculation formula is: , is a constant to ensure the stability of numerical calculations.

[0085] In the preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 500, the empirical data pair (s t ,a t ,r t ,s t+1 ) is represented by the state vector s at time t t , parameter adjustment action a at time t t , reward value r at time t t , execute parameter adjustment action a t The state vector s formed at time t+1 is then entered t+1 constitute.

[0086] In a preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 600, the online strategy network μ(s t ;θ) includes input layer, intermediate layer and output layer. The input layer receives the state vector s t , the output layer output parameter adjustment action a t .

[0087] In the preferred embodiment of the engine multi-parameter self-adjusting model modeling method based on reinforcement learning, in step 700, the online value network q(s t ,a t ;w) includes input layer, intermediate layer and output layer, the input layer receives the state vector s t and parameter adjustment action a t , output layer output (s t ,a t ) corresponds to the value Q t .

[0088] In a preferred embodiment of the engine multi-parameter self-tuning model modeling method based on reinforcement learning, in step 800, the self-tuning module training phase includes the following steps:

[0089] Step 801: Use the deep deterministic policy gradient algorithm DDPG to build an online policy network μ(s t ;θ), online value network q(s t ,a t ;w), target strategy network μ T (s t ;θ T ), target value network q T (s t ,a t ;w T ),in,

[0090] Online strategy network μ(s t ;θ) and the number of neurons in the input layer and the state vector s t The dimension is the same as that of the action vector a. t The dimensions are the same,

[0091] Online value network q(s t ,a t ;w) The number of neurons in the input layer is the state vector s t The dimension of the action vector a t The sum of the dimensions of , the number of neurons in the output layer is 1,

[0092] Target policy network μ T (s t ;θ T ) and the online policy network μ(s t ;θ) has the same network structure,

[0093] Target value network q T (s t ,a t ;w T ) and the online value network q(s t,a t ;w) has the same network structure,

[0094] And initialize the online strategy network parameters θ, online value network parameters w, target strategy network parameters θ T , target value network parameter w T , the learning rate α of the online policy network parameters, the learning rate β of the online value network parameters, the discount factor and update coefficients ;

[0095] Step 802: Observe the state vector s corresponding to the residual signal of the engine gas path parameter at time t t , calculate the online policy network μ(s t ;θ t ) corresponding parameter adjustment action a t ;

[0096] Step 803: Let the engine component level model perform parameter adjustment action a t , so that the parameter adjustment coefficients are involved in the calculation of the engine component level model;

[0097] Step 804: Calculate the reward value r at time t t , observe the predicted value of the gas path parameters at time t+1, and calculate the next corresponding state vector s t+1 , the empirical data (s t ,a t ,r t ,s t+1 ) is stored in the experience pool;

[0098] Step 805: collect N experience data pairs (s i ,a i ,r i ,s i+1 ), calculate the target value y i

[0099] ;

[0100] Step 806: Minimize the loss function by gradient descent , update the online value network parameter w, the update formula is ,in, is the online value network parameter of the k+1th iteration, is the online value network parameter of the kth iteration, The loss function is The gradient at

[0101] Step 807: Maximize the scoring function by gradient ascent method , update the online policy network parameters θ, the update formula is ,in, is the online policy network parameter of the k+1th iteration, is the online policy network parameter of the kth iteration, The loss function is The gradient at

[0102] Step 808: Update the parameters of the target value network and the target strategy network. ;

[0103] Step 809: Set the network in the self-tuning module to perform 1000 rounds of training, repeat steps 802 to 808, and the training ends when the number of training rounds reaches the set value.

[0104] In a preferred embodiment of the engine multi-parameter self-tuning model modeling method based on reinforcement learning, in step 900, the self-tuning module application stage includes the following steps:

[0105] Step 901: Observe the engine gas path parameter residual signal corresponding to the state vector s at time t t , calculate the target policy network μ T (s t ;θ T ) corresponding parameter adjustment action a t ;

[0106] Step 902: Let the engine component level model perform parameter adjustment action a t , so that the parameter adjustment coefficient participates in the calculation of the engine component level model and predicts the engine gas path parameters,

[0107] Step 903: repeat steps 901 to 902 until the parameter prediction process is completed.

[0108] In one embodiment, Figure 1 As shown, the method includes the following steps:

[0109] Step 100, collecting engine input parameters and gas path parameters;

[0110] Step 200, adjusting action a according to the input parameters of the engine and the parameters of the self-adjusting module t , calculate the gas path parameters of the engine component level model;

[0111] Step 300: Calculate the residual signal between the component-level model after parameter adjustment and the engine gas path parameters, and record the residual signal of the engine gas path parameters as the state vector s t ;

[0112] Step 400: Design a reward function to calculate the reward value r obtained after each parameter adjustment. t ;

[0113] Step 500, store the empirical data pair (s t ,a t ,r t ,s t+1 ), build an experience pool;

[0114] Step 600: In the self-adjusting module, construct an online strategy network μ(s t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network;

[0115] Step 700: In the self-adjusting module, construct an online value network q(s t ,a t ;w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network;

[0116] Step 800 , in the self-adjusting module training phase, the policy network and the value network are trained using the experience data pairs in the experience pool;

[0117] Step 900, in the self-tuning module application phase, uses the trained policy network to calculate the corresponding parameter adjustment action according to the state vector, participates in the calculation of the component-level model, and estimates the engine state parameters.

[0118] In step 100, the engine input parameter I E (t)={T t0 (t), P t0 (t), W f (t), A8(t)}, where T t0 (t) is the total intake air temperature, P t0 (t) is the total intake pressure, W f (t) is the fuel flow rate, A8(t) is the nozzle area. The engine gas path parameters O E (t)={N L,E (t), N H,E (t), P t25,E (t), P t31,E (t), T t6,E (t), P t6,E (t)}, where N L,E (t) is the measured value of the low-pressure rotor speed, N H,E(t) is the measured value of the high-pressure rotor speed, P t25,E (t) is the measured value of the pressure at the outlet of the low-pressure compressor, P t31,E (t) is the measured value of the pressure at the high pressure compressor outlet, T t6,E (t) is the measured value of the temperature at the outlet of the low-pressure turbine, P t6,E (t) is the measured value of the pressure at the outlet of the low-pressure turbine.

[0119] In step 200, the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients, which is used to adjust the input parameters and component characteristic parameters of the component level model. The engine component level model is expressed as .

[0120] Among them, the output of the engine model O M (t)={N L,M (t), N H,M (t), P t25,M (t), P t31,M (t), T t6,M (t), P t6,M (t)},N L,M (t) is the predicted value of low-pressure rotor speed calculated by the model, N H,M (t) is the predicted value of high-pressure rotor speed calculated by the model, P t25,M (t) is the predicted value of the pressure at the outlet of the low-pressure compressor calculated by the model, P t31,M (t) is the predicted value of the high pressure compressor outlet pressure calculated by the model, T t6,M (t) is the predicted value of the temperature at the outlet of the low-pressure turbine calculated by the model, P t6,M (t) is the predicted value of the pressure at the low-pressure turbine outlet calculated by the model.

[0121] The adjustment coefficient vector k(t) includes the fuel flow adjustment coefficient , nozzle area adjustment coefficient , Fan flow adjustment coefficient , Fan efficiency adjustment factor , High-pressure compressor flow adjustment coefficient , High-pressure compressor efficiency adjustment coefficient , High-pressure turbine flow adjustment coefficient , High-pressure turbine efficiency adjustment factor , low-pressure turbine flow adjustment coefficient and low-pressure turbine efficiency adjustment factor .

[0122] In step 300, the state vector s tIt is composed of the residual between the calculated value of the engine model gas path parameter and the measured value of the actual engine gas path parameter, which can be expressed as , further, can be written as .

[0123] In step 400, the formula of the reward function is: ,in,

[0124] is the reward value corresponding to the calculated value of the i-th gas path parameter at time t, and the calculation formula is ;

[0125] in, is the absolute value of the relative error of the i-th gas path parameter, and the calculation formula is: , is a constant to ensure the stability of numerical calculations.

[0126] In step 500, the empirical data pair (s t ,a t ,r t ,s t+1 ) is represented by the state vector s at time t t , parameter adjustment action a at time t t , reward value r at time t t , perform action a t The state vector s formed at time t+1 is then entered t+1 constitute.

[0127] In step 600, the policy network μ(s t ;θ), including input layer, intermediate layer and output layer, the input layer receives the state vector s t , the output layer output parameter adjustment action a t .

[0128] In step 700, the value network q(s t ,a t ;w), including input layer, intermediate layer and output layer, the input layer receives the state vector s t and parameter adjustment action a t , output layer output (s t ,a t ) corresponds to the value Q t .

[0129] In step 800, the self-adjusting module training phase includes the following steps:

[0130] Step 801: Use the Deep Deterministic Policy Gradient (DDPG) algorithm to build an online policy network μ(s t ;θ), online value network q(st ,a t ;w), target policy network μ T (s t ;θ T ), target value network q T (s t ,a t ;w T ),in,

[0131] Online strategy network μ(s t ;θ) and the number of neurons in the input layer and the state vector s t The dimension is the same as that of the action vector a. t The dimensions are the same,

[0132] Online value network q(s t ,a t ;w) The number of neurons in the input layer is the state vector s t The dimension of the action vector a t The sum of the dimensions of , the number of neurons in the output layer is 1,

[0133] Target policy network μ T (s t ;θ T ) and the online policy network μ(s t ;θ) has the same network structure,

[0134] Target value network q T (s t ,a t ;w T ) and the online value network q(s t ,a t ;w) has the same network structure,

[0135] And initialize the online strategy network parameters θ, online value network parameters w, target strategy network parameters θ T , target value network parameter w T , the learning rate α of the online policy network parameters, the learning rate β of the online value network parameters, the discount factor and update coefficients ;

[0136] Step 802: Observe the state vector s corresponding to the residual signal of the engine gas path parameter at time t t , calculate the policy network μ(s t ;θ t ) corresponding action a t ;

[0137] Step 803: Let the component-level model perform action a t, so that the parameter adjustment coefficient participates in the calculation of the component-level model;

[0138] Step 804: Calculate the reward value r at time t t , observe the predicted value of the gas path parameters at time t+1, and calculate the next corresponding state vector s t+1 , the empirical data (s t ,a t ,r t ,s t+1 ) is stored in the experience pool;

[0139] Step 805: collect N experience data pairs (s i ,a i ,r i ,s i+1 ), calculate the target value y i

[0140] ;

[0141] Step 806: Minimize the loss function by gradient descent , update the online value network parameter w, the update formula is ,in, is the online value network parameter of the k+1th iteration, is the online value network parameter of the kth iteration, The loss function is The gradient at

[0142] Step 807: Maximize the scoring function by gradient ascent method , update the online policy network parameters θ, the update formula is ,in, is the online policy network parameter of the k+1th iteration, is the online policy network parameter of the kth iteration, The loss function is The gradient at

[0143] Step 808: Update the parameters of the target value network and the target strategy network. ;

[0144] Step 809: Set the network in the self-tuning module to perform 1000 rounds of training, repeat steps 802 to 808, and the training ends when the number of training rounds reaches the set value.

[0145] In step 900, the self-adjusting module application phase includes the following steps:

[0146] Step 901: Observe the engine gas path parameter residual signal corresponding to the state vector s at time tt , calculate the target policy network μ T (s t ;θ T ) corresponding action a t ;

[0147] Step 902: Let the component-level model perform action a t , so that the parameter adjustment coefficient is involved in the calculation of the component-level model to predict the engine gas path parameters.

[0148] Step 903: repeat steps 901 to 902 until the parameter prediction process is completed.

[0149] Figure 2 This is a curve showing the change in reward value during the training process of the self-adjusting module provided by one embodiment of the present disclosure.

[0150] FIG3( a ) shows an embodiment of the present disclosure, wherein the engine low-pressure rotor speed N is used and the non-used parameter self-adjusting module is used. L A comparative case.

[0151] FIG3( b ) shows an embodiment of the present disclosure, wherein the engine high pressure rotor speed N is used and the non-used parameter self-adjusting module is used. H A comparative case.

[0152] FIG3( c ) shows an embodiment of the present disclosure, wherein the engine compressor pressure P is used and not used. t31 A comparative case.

[0153] FIG3( d ) shows an embodiment of the present disclosure, wherein the engine low-pressure turbine after-temperature T is adopted and not adopted the parameter self-adjusting module. t6 A comparative case.

[0154] It can be seen that the predicted values ​​of the gas path parameters of the component-level model using the self-adjusting module are closer to the engine measurement values. The strategy network constructed using reinforcement learning corrects the fuel flow, nozzle area and gas path component performance parameters of the component-level model, thereby improving the modeling accuracy of the engine model.

[0155] In one embodiment, the steps include: collecting engine input parameters and gas path parameters; adjusting action a according to the engine input parameters and the parameters of the self-adjusting module; t , calculate the gas path parameters of the engine component level model; calculate the residual signal between the component level model after parameter adjustment and the engine gas path parameters, and record the residual signal of the engine gas path parameters as the state vector s t ; Design a reward function to calculate the reward value r obtained after each parameter adjustment t ;Store experience data pairs (s t,a t ,r t ,s t+1 ), build the experience pool; in the self-adjustment module, build the online strategy network μ(s t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t ;Build online value network q(s t ,a t ;w), used to calculate the state vector s t and parameter adjustment action a t During the self-adjustment module training phase, the policy network and value network are trained using empirical data pairs from the experience replay pool. During the self-adjustment module application phase, the trained policy network is used to calculate corresponding parameter adjustment actions based on the state vector, participate in component-level model calculations, and predict engine gas path parameters. This invention can reduce the impact of engine input parameter measurement errors and component degradation on model accuracy, thereby improving the modeling accuracy of the engine model.

[0156] In one embodiment, the engine is an aircraft engine.

[0157] A reinforcement learning-based engine multi-parameter self-tuning model modeling system includes:

[0158] Acquisition unit, collects the input parameters of the engine I E (t) and the engine's gas path parameters O E (t);

[0159] The model building unit is used to: build an engine component level model with a self-tuning module based on the engine, according to the input parameters I of the engine E (t) and the parameter adjustment action a of the self-adjusting module t , calculate the gas path parameters of the engine component level model, where the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients to adjust the input parameters and component characteristic parameters of the engine component level model. The engine component level model formula is expressed as , where O M (t) is the gas path parameter output by the engine model, f M is the function representing the engine model;

[0160] Residual unit, calculates the gas path parameters of the engine component level model after parameter adjustment and the gas path parameters of the engine E (t), the residual signal is recorded as the state vector s t ;

[0161] The reinforcement learning machine includes a self-adjustment module, wherein the reinforcement learning machine designs a reward function for calculating the reward value r obtained after each parameter adjustment. t , store the empirical data pairs (s t ,a t ,r t ,s t+1 ), build the experience pool; build the online strategy network μ(s) in the self-adjusting module t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network, and the online value network q(s t ,a t ; w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network; in the self-adjusting module training phase, the experience data pairs in the experience pool are used to adjust the online policy network μ(s t ;θ) and the online value network q(s t ,a t ;w) conduct training;

[0162] The adjustment coefficient prediction unit uses the trained online strategy network μ(s t ;θ) calculates the corresponding parameter adjustment action based on the state vector and participates in the calculation of the engine component level model.

[0163] A computer storage medium includes computer instructions, which, when executed on a computer, cause the computer to execute the method described above.

[0164] An electronic device, characterized in that the electronic device comprises:

[0165] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein:

[0166] When the processor executes the program, the method described is implemented.

[0167] Although the embodiments of the present invention have been described above with reference to the accompanying drawings, the present invention is not limited to the above-mentioned specific embodiments and application fields. The above-mentioned specific embodiments are merely illustrative and instructive, and are not restrictive. A person skilled in the art, guided by this specification and without departing from the scope of protection of the claims of the present invention, may also devise various forms, all of which fall within the scope of protection of the present invention.

Claims

1. A method for modeling an engine multi-parameter self-tuning model based on reinforcement learning, characterized in that: The steps include: Step 100: Collect the engine input parameters I E (t) and the engine's gas path parameters O E (t); Step 200: construct an engine component level model with a self-adjusting module based on the engine, and E (t) and the parameter adjustment action a of the self-adjusting module t , calculate the gas path parameters of the engine component level model, where the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients to adjust the input parameters and component characteristic parameters of the engine component level model. The engine component level model formula is expressed as , where O M (t) is the gas path parameter output by the engine model, f M is the function representing the engine model; Step 300: Calculate the gas path parameters of the engine component level model after parameter adjustment. M (t) and the engine's gas path parameters O E (t), the residual signal is recorded as the state vector s t ; Step 400: Design a reward function to calculate the reward value r obtained after each parameter adjustment. t ; Step 500, store the empirical data pair (s t ,a t ,r t ,s t+1 ), build an experience pool; Step 600: construct an online policy network μ(s t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network; Step 700: construct an online value network q(s) in the self-adjusting module. t ,a t ; w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network; Step 800: In the self-adjustment module training phase, the online policy network μ(s t ;θ) and the online value network q(s t ,a t ;w) conduct training; Step 900, in the self-adjustment module application phase, the trained online strategy network μ(s t ;θ) Calculate the corresponding parameter adjustment action based on the state vector, participate in the calculation of the engine component level model, and estimate the engine state parameters.

2. The engine multi-parameter self-tuning model modeling method based on reinforcement learning according to claim 1, characterized in that: In step 100, the engine input parameter I E (t)={T t0 (t), P t0 (t), W f (t), A8(t)}, where T t0 (t) is the total intake air temperature, P t0 (t) is the total intake pressure, W f (t) is the fuel flow rate, A8(t) is the nozzle area, and the gas path parameters of the engine are E (t)={N L,E (t), N H,E (t), P t25,E (t), P t31,E (t), T t6,E (t), P t6,E (t)}, where N L,E (t) is the measured value of the low-pressure rotor speed, N H,E (t) is the measured value of the high-pressure rotor speed, P t25,E (t) is the measured value of the pressure at the outlet of the low-pressure compressor, P t31,E (t) is the measured value of the pressure at the high pressure compressor outlet, T t6,E (t) is the measured value of the temperature at the outlet of the low-pressure turbine, P t6,E (t) is the measured value of the pressure at the outlet of the low-pressure turbine.

3. The engine multi-parameter self-tuning model modeling method based on reinforcement learning according to claim 1, characterized in that: Gas path parameters O of engine component level model M (t)={N L,M (t), N H,M (t), P t25,M (t), P t31,M (t), T t6,M (t), P t6,M (t)},N L,M (t) is the predicted value of low-pressure rotor speed calculated by the model, N H,M (t) is the predicted value of high-pressure rotor speed calculated by the model, P t25,M (t) is the predicted value of the pressure at the outlet of the low-pressure compressor calculated by the model, P t31,M (t) is the predicted value of the high pressure compressor outlet pressure calculated by the model, T t6,M (t) is the predicted value of the temperature at the outlet of the low-pressure turbine calculated by the model, P t6,M (t) is the predicted value of the pressure at the low-pressure turbine outlet calculated by the model.

4. The engine multi-parameter self-tuning model modeling method based on reinforcement learning according to claim 1, characterized in that: The vector k(t) of the model parameter adjustment coefficients of the engine component level model includes the fuel flow adjustment coefficient , nozzle area adjustment coefficient , Fan flow adjustment coefficient , Fan efficiency adjustment factor , High-pressure compressor flow adjustment coefficient , High-pressure compressor efficiency adjustment coefficient , High-pressure turbine flow adjustment coefficient , High-pressure turbine efficiency adjustment factor , low-pressure turbine flow adjustment coefficient and low-pressure turbine efficiency adjustment factor .

5. The engine multi-parameter self-adjusting model modeling method based on reinforcement learning according to claim 1, characterized in that: In step 300, the state vector s t It is composed of the residual between the calculated value of the gas path parameter of the engine component level model and the measured value of the actual engine gas path parameter, which can be expressed as .

6. The engine multi-parameter self-tuning model modeling method based on reinforcement learning according to claim 1, characterized in that: In step 800, the self-adjusting module training phase includes the following steps: Step 801: Use the deep deterministic policy gradient algorithm DDPG to build an online policy network μ(s t ;θ), online value network q(s t ,a t ;w), target strategy network μ T (s t ;θ T ), target value network q T (s t ,a t ;w T ),in, Online strategy network μ(s t ;θ) includes input layer, intermediate layer and output layer. The input layer receives the state vector s t , the output layer output parameter adjustment action a t , the number of neurons in the input layer and the state vector s t The dimension is the same as that of the action vector a. t The dimensions are the same, Online value network q(s t ,a t ;w) includes input layer, intermediate layer and output layer, the input layer receives the state vector s t and parameter adjustment action a t , output layer output (s t ,a t ) corresponds to the value Q t , the number of neurons in the input layer is the state vector s t The dimension of the action vector a t The sum of the dimensions of , the number of neurons in the output layer is 1, Target policy network μ T (s t ;θ T ) and the online policy network μ(s t ;θ) has the same network structure, Target value network q T (s t ,a t ;w T ) and the online value network q(s t ,a t ;w) has the same network structure, And initialize the online strategy network parameters θ, online value network parameters w, target strategy network parameters θ T , target value network parameter w T , the learning rate α of the online policy network parameters, the learning rate β of the online value network parameters, the discount factor and update coefficients ; Step 802: Observe the state vector s corresponding to the residual signal of the engine gas path parameter at time t t , calculate the online policy network μ(s t ;θ t ) corresponding parameter adjustment action a t ; Step 803: Let the engine component level model perform parameter adjustment action a t , so that the parameter adjustment coefficients are involved in the calculation of the engine component level model; Step 804: Calculate the reward value r at time t t , observe the predicted value of the gas path parameters at time t+1, and calculate the next corresponding state vector s t+1 , the empirical data (s t ,a t ,r t ,s t+1 ) is stored in the experience pool; Step 805: collect N experience data pairs (s i ,a i ,r i ,s i+1 ), calculate the target value y i ; Step 806: Minimize the loss function by gradient descent , update the online value network parameter w, the update formula is ,in, is the online value network parameter of the k+1th iteration, is the online value network parameter of the kth iteration, The loss function is The gradient at Step 807: Maximize the scoring function by gradient ascent method , update the online policy network parameters θ, the update formula is ,in, is the online policy network parameter of the k+1th iteration, is the online policy network parameter of the kth iteration, The loss function is The gradient at Step 808: Update the parameters of the target value network and the target strategy network. ; Step 809: Set the network in the self-tuning module to perform 1000 rounds of training, repeat steps 802 to 808, and the training ends when the number of training rounds reaches the set value.

7. The engine multi-parameter self-adjusting model modeling method based on reinforcement learning according to claim 1, characterized in that: In step 900, the self-adjusting module application phase includes the following steps: Step 901: Observe the engine gas path parameter residual signal corresponding to the state vector s at time t t , calculate the target policy network μ T (s t ;θ T ) corresponding parameter adjustment action a t ; Step 902: Let the engine component level model perform parameter adjustment action a t , so that the parameter adjustment coefficient participates in the calculation of the engine component level model and predicts the engine gas path parameters, Step 903: repeat steps 901 to 902 until the parameter prediction process is completed.

8. An engine multi-parameter self-adjusting model modeling system based on reinforcement learning, characterized in that: It includes: The acquisition unit is used to: acquire the input parameters I of the engine E (t) and the engine's gas path parameters O E (t); The model building unit is used to: build an engine component level model with a self-tuning module based on the engine, according to the input parameters I of the engine E (t) and the parameter adjustment action a of the self-adjusting module t , calculate the gas path parameters of the engine component level model, where the parameter adjustment action a of the self-adjusting module t The vector k(t) is composed of model parameter adjustment coefficients to adjust the input parameters and component characteristic parameters of the engine component level model. The engine component level model formula is expressed as , where O M (t) is the gas path parameter output by the engine model, f M is the function representing the engine model; The residual unit is used to calculate the gas path parameters of the engine component level model after parameter adjustment and the gas path parameters of the engine. E (t), the residual signal is recorded as the state vector s t ; The reinforcement learning machine includes a self-adjustment module, wherein the reinforcement learning machine designs a reward function for calculating the reward value r obtained after each parameter adjustment. t , store the empirical data pairs (s t ,a t ,r t ,s t+1 ), build the experience pool; build the online strategy network μ(s) in the self-adjusting module t ;θ), used to calculate the state vector s t The corresponding parameter adjustment action a t , θ is the network parameter of the value network, and the online value network q(s t ,a t ; w), used to calculate the state vector s t and parameter adjustment action a t The corresponding value, w is the network parameter of the value network; in the self-adjustment module training phase, the experience data pairs in the experience pool are used to adjust the online policy network μ(s t ;θ) and the online value network q(s t ,a t ;w) conduct training; The adjustment coefficient prediction unit uses the trained online strategy network μ(s t ;θ) calculates the corresponding parameter adjustment action based on the state vector and participates in the calculation of the engine component level model.

9. A computer storage medium, characterized in that The storage medium includes computer instructions, which, when executed on a computer, enable the computer to perform the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Gas turbine engine fuel flow correction method based on network model

    CN111523276A

  • Robot path planning method and device based on HER-SAC algorithm

    CN117873070A