An Intelligent Direct Thrust Control Method for Turbofan Engines Based on Reinforcement Learning

By designing the direct thrust intelligent controller of the turbofan engine based on reinforcement learning, the problems of low control accuracy and imbalance of thrust output of the existing indirect thrust control technology are solved, and the control effect with excellent dynamic performance in the full-enclosed line is achieved.

CN114527654BActive Publication Date: 2025-07-01NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210088552.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-07-01
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The existing indirect thrust control technology of turbofan engines has problems such as low control accuracy, large individual differences, complex nonlinear change relationships and imbalance in thrust output, which is difficult to meet the control needs of high-performance aero engines.

Method used

A turbofan engine direct thrust intelligent controller is designed based on reinforcement learning, and the fully connected neural network is replaced by LSTM RNN. It combines the PPO algorithm to train the full-inclusive direct thrust controller, design a new controller form and reward system, and add a points link to improve control accuracy and stability.

Benefits of technology

It realizes a controller with excellent dynamic performance in the all-inclusive line, improves the accuracy and stability of direct thrust control, protects key safety parameters, and enhances the real-time and adaptability of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114527654B_ABST
    Figure CN114527654B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent direct thrust control method for a turbofan engine based on reinforcement learning, comprising the following steps: Step 1), select the structures and parameters of the policy and evaluation networks, and design the form of the direct thrust controller considering the protection of key safety parameters and the form of the reward of the reinforcement learning environment; Step 2), based on the continuous policy gradient reinforcement learning algorithm, use the component-level model to build an environment for exploration, and train the intelligent agent policy network and evaluation network through the experience obtained from the exploration; Step 3), test the control performance of the intelligent agent within the full envelope range, and optimize the network structure and parameters. The present invention solves the problems of poor dynamic performance, high conservatism, and inaccurate thrust control in the indirect thrust control of turbofan engines. Through the reward designed by the present invention, the intelligent agent is motivated to search for the direct thrust controller with the optimal dynamic performance within the full envelope range, and it is ensured that the key safety parameters of the engine do not exceed the limit during the control process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of aeroengine control, and particularly relates to a direct thrust intelligent control method for a turbofan engine based on reinforcement learning. Background Art

[0002] With the development of aeroengines towards high thrust-to-weight ratio / power-to-weight ratio, high performance and high economy, the traditional indirect thrust control method of aeroengines can no longer meet the requirements of advanced aeroengine control.

[0003] The traditional indirect thrust control method usually controls the thrust output level of the engine by controlling the pressure ratio and rotational speed of the engine. However, this control method has many drawbacks. First, the off-line designed controller has a large conservatism, which limits the performance level of the engine. Second, the individual differences between engines will cause deviations in thrust control. Third, the change relationship between the throttle lever angle and thrust in indirect thrust control is non-linear, which increases the difficulty and response time for pilots to control the aircraft and adds risks to the combat process. Fourth, twin-engine civil airliners will yaw due to unbalanced engine thrust output, and additional rudder control is required to correct it, which increases the wind resistance of the airliner and reduces the economy of the airliner. Based on the above four points, indirect thrust control severely restricts the development of high-performance aeroengines. Therefore, the research on direct thrust control technology is necessary.

[0004] The intelligent engine (IEC) in the United States, as one of the top four advanced control concepts, one of the core research contents is direct thrust control. The aviation technology development seminar initiated by NASA also proposed the need to reduce the workload of pilots and improve the autonomy of aircraft engine operation. The research on direct thrust control is receiving extensive attention worldwide.

[0005] However, there are still many key technical problems in direct thrust control technology that have not been broken through. First, in terms of the control architecture, there are roughly three existing direct thrust control architectures: the double-loop architecture, the multivariable control architecture, and the control architecture based on optimization algorithms. The double-loop architecture usually only nests an outer thrust feedback control on the basis of the PI control architecture, which increases the complexity of the system and is difficult to obtain satisfactory dynamic characteristics. In the multivariable control architecture based on linear algorithms, when the number of control variables is more than the number of controlled variables, there will be a problem of difficult solution of overdetermined equations. Therefore, it is necessary to add additional controlled variables, which is redundant for thrust control and will affect the effect of thrust control. For the control architecture based on optimization algorithms, although better control effects can be obtained, the control accuracy of this architecture depends on the modeling accuracy, and the online optimization calculation amount is large, which is difficult to meet the real-time requirements of control. At present, it is very difficult to have an architecture that can be applied to engineering practice. There is an urgent need to develop a direct thrust control method suitable for multivariable aeroengines and with excellent dynamic quality. Summary of the Invention

[0006] Object of the Invention: In order to overcome the deficiencies of the existing indirect thrust control technology of turbofan engines in thrust control, the present invention proposes a design method for a direct thrust intelligent controller of a turbofan engine. Considering the strong nonlinearity and strong time series characteristics of the turbofan engine, by designing a new controller and reward form, using LSTM RNN to replace the fully connected neural network in the reinforcement learning strategy and evaluation network, and training an all-envelope direct thrust controller based on the PPO algorithm. By adopting the above method, a controller with excellent dynamic performance within the all-envelope can be obtained, and safe and effective direct thrust control can be realized.

[0007] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A design method for a direct thrust intelligent controller of a turbofan engine, comprising the following steps:

[0009] Step 1), select the structure and parameters of the strategy and evaluation network, and design the form of the direct thrust controller considering the protection of key safety parameters and the reward form of the reinforcement learning environment;

[0010] Step 2), train the all-envelope intelligent direct thrust controller: based on the continuous policy gradient reinforcement learning algorithm, use the component-level model to build an environment for exploration, and train the intelligent agent policy network and evaluation network through the experience obtained from the exploration;

[0011] Step 3), test the control performance of the intelligent agent within the all-envelope range, and optimize the network structure and parameters.

[0012] Further, the specific steps in the above Step 1) are as follows:

[0013] Step 1.1), select the LSTM RNN network as the policy network and evaluation network of reinforcement learning, and define the input and output of the policy network; select the atmospheric environment parameters and the measurable state parameters of the engine to form the engine state s at time t t , select the current and the previous n moments' states of the engine to form the engine state group X at time t t = [s t-2 , s t-1 , s t , as the input of the policy network; the output of the policy network is two vectors with a dimension of 4, representing the mean u t of the action a t given by the policy network under the current state X t and the variance σ t , and the action a t under the current state X t = [a u,t a σ,t ; by defining the input and output of the policy network, obtain the input-output relationship formed by X t = [s t-2 , s t-1 , s t → Y t = [a u,t a σ,t ; the input of the evaluation network is the same as that of the policy network, and the output is a scalar, representing the value v(s t |a t ) of taking the action a t under the current engine state X t , and obtain the input-output relationship formed by X t = [s t-2 , s t-1 , s t → Y t = [v(s t |a t )];

[0014] Step 1.2), define the form of the controller, add an integral link to the controller, and obtain the controller expressions for the fuel flow rate W fb and the area A8 of the tail nozzle as

[0015]

[0016]

[0017] where, g1 and g2 are the inertia coefficients of the fuel flow rate and the change of the tail nozzle area, and e F is the error F ref - F of the thrust control, and W fb,tDenotes the normalized initial state of fuel at time t, A 8,t Denotes the normalized initial state of the nozzle area at time t, e F Denotes the thrust error, ΔW fb Denotes the fuel flow rate increment Denotes the fuel integration parameter Denotes the nozzle integration parameter;

[0018] Step 1.3), design the reward form of the reinforcement learning environment. The reward consists of the control precision reward r e,t 、the control stability reward r s,t and the critical parameter overrun penalty r l,t These three parts;

[0019] The control precision reward r e,t is

[0020]

[0021] F ref Denotes the thrust command, and F denotes the normalized command of the thrust;

[0022] The control stability reward r s,t is a constant

[0023] r s,t = 0.1

[0024] The critical safety parameter overrun penalty r l,t is

[0025]

[0026] Denote the fan surge penalty, the compressor surge penalty, the high-pressure compressor outlet total pressure overrun penalty, the low-pressure turbine outlet total temperature overrun penalty, and the low-pressure rotor speed overrun penalty respectively;

[0027] The total reward at time t can be written as

[0028] r t = r e,t + r s,t + r l,t .

[0029] Furthermore, the specific steps of the full envelope intelligent direct thrust controller training method in the said step 2) are as follows:

[0030] Step 2.1), set the learning rate, the maximum number of episodes, the episode length, the policy update frequency, the batch data dimension, and the number of policy updates according to experience;

[0031] Step 2.2), in the same episode, randomly select the operating point, initial state, and thrust command, and explore in the set environment until the episode time limit is reached; record the state, action, and reward information during the exploration process to generate an experience pool for the agent to update and use;

[0032] Step 2.3), repeat the steps of Step 2.2). When the number of update episodes is reached, update the policy and value networks of the agent.

[0033] Furthermore, the specific steps for generating the experience pool in Step 2.2) are as follows:

[0034] Step 2.2.1), randomly select the operating point (H, Ma) and initial state within the envelope, find the feasible thrust range at this point, and randomly select a thrust value within this range as the thrust command F ref ;

[0035] Step 2.2.2), at time t, generate the engine state group X t = [s t-2 , s t-1 , s t according to the engine state, and use it as the input of the policy network to obtain the action a t under the current state X t = [a u,t a σ,t ; According to the mean u t and variance σ t of the action, sample the current action under the normal distribution, obtain the engine control quantity through the controller form defined in Step 1), use the control quantity as the input of the component-level model, after dynamic calculation, generate the engine state group X t+1 at time t + 1 through the component-level model state at time t + 1, and calculate the reward r t+1 at time t + 1 according to the reward form determined in Step 1); Record the engine state groups X t and X t+1 , the action a t of the policy network, and the reward r t+1 during one-step dynamic process, and regard them as a set of data to add to the experience pool for the agent to train and use;

[0036] Step 2.2.3), continuously sample at the same operating point according to the method described in Step 2.2.2 until the episode time length limit is reached.

[0037] Furthermore, the specific steps for updating the policy network and value network in Step 2.3) are as follows:

[0038] Step 2.3.1), for a set of data in the experience pool in Step 2.2), calculate the return Gt = r t + r t+1 + … + r T ; The subscript T represents the time series at the last moment in a round. With the help of Markov recursion, this process can be simplified to G t = r t+1 + γG t+1 , where γ is the discount factor, indicating the importance of future moments for the current moment; Take the engine state group X t+1 at time t + 1 as the input of the evaluation network to obtain the value V t+1 , and use V t+1 to replace G t+1 , to get G t = r t+1 + γV t+1 ;

[0039] Step 2.3.2), by comparing the return G t and the value V t , calculate the advantage of the action (a t |X t ) at time t Compare the probabilities π θ (a t |X t ) of the current action selected by the new policy and the old policy and to obtain the probability ratio By setting the clipping parameter ε and restricting it between [1 - ε, 1 + ε], the policy update amplitude is restricted;

[0040] Step 2.3.3), based on the PPO algorithm, calculate the loss function and update the agent's policy network and value network.

[0041] Further, the specific steps for optimizing the agent network structure and parameters in step 3) are as follows:

[0042] Step 3.1), randomly generate l groups of 1×3 arrays c i = [W fb,i , A 8,i , F i (i = 1,..., l), and form where W fb,i represents the normalized initial state of fuel in the i-th group of experiments, A 8,i represents the normalized initial state of the nozzle area in the i-th group of experiments, F i represents the normalized command of thrust in the i-th group of experiments, and W fb,i ∈ [0, 1], A 8,i ∈ [0, 1], F i ∈ [0, 1];

[0043] Step 3.2), at the same operating point (H, Ma), respectively obtain the maximum thrust F max and the minimum thrust F min , the maximum fuel flow rate W fb,max and the minimum fuel flow rate W fb,min , the maximum nozzle area A 8,max and the minimum nozzle area, and according to the in Step 3.1), establish a number of test cases, and perform anti-normalization to obtain the initial fuel, the initial nozzle area, and the thrust command value;

[0044] W fb,ini,i = W fb,i ·(W fb,max - W fb,min ) + W fb,min

[0045] A 8,ini,i = A 8,i ·(A 8,max - A 8,min ) + A 8,min

[0046] F ref,i = F i ·(F max - F min ) + F min

[0047] Step 3.3), based on the obtained policy network and the test cases generated in Step 3.2), control the thrust;

[0048] Step 3.4), calculate the reward value r t at each moment during the control process according to the reward expression in Step 1) (t = 1,..., T), and record the sum of the reward values at each moment as the score of this test case at this operating point

[0049] Step 3.5), within the full envelope, divide the altitude and Mach number into n × m grid points, perform tests at these grid points respectively, and record the score of each point, the average score of the full envelope, and the variance;

[0050] Step 3.6), evaluate the controller performance according to the scores, and adjust the network structure parameters and training parameters.

[0051] Further, the specific steps of the thrust control based on the policy network in Step 3.3) are as follows:

[0052] Step 3.3.1), based on the component-level model, generate the engine state group X t at time t, and take Xt As the input of the policy network, the current state X is obtained. t The action a t = [a u,t a σ,t ;

[0053] In step 3.3.2), during the actual control process of the agent, the random sampling of the action distribution is discarded, and the mean value a u,t is directly used as the action and substituted into the controller expression described in step 1) to calculate the control quantity of the engine, which is used as the input of the engine to complete the control of the engine;

[0054] In step 3.3.3), the above process is repeated until the control period ends.

[0055] Beneficial effects: A design method for a direct thrust intelligent controller of a turbofan engine provided by the present invention has the following technical effects compared with the prior art by adopting the above technical solutions:

[0056] (1) Aiming at the direct thrust control problem of the turbofan engine, the present invention designs a new controller form, making the change of the control quantity smoother, and the added integral link effectively eliminates the control error existing in the control using neural networks.

[0057] (2) The present invention proposes a new direct thrust control reward system for the PPO algorithm for the turbofan engine, effectively improving the dynamic characteristics of the control system within the full envelope range and protecting the key safety parameters of the system from exceeding the limit during the dynamic process.

[0058] (3) Aiming at the strong time series characteristics of the turbofan engine, the present invention uses LSTM RNN to replace the fully connected neural network in the PPO algorithm, effectively improving the convergence speed and stability of the policy network and the evaluation network, and improving the performance of the controller within the full envelope range. Description of the Drawings

[0059] Figure 1 is the design flow of the direct thrust controller of the turbofan engine based on reinforcement learning.

[0060] Figure 2 is the structural schematic diagram of the turbofan engine.

[0061] Figure 3 is the direct thrust intelligent control architecture diagram of the present invention.

[0062] Figure 4 is the direct thrust control diagram of Case 1.

[0063] Figure 5 is the direct thrust control diagram of Case 2.

[0064] Figure 6 It is the full envelope simulation flight path.

[0065] Figure 7 It is the full envelope thrust control effect. Detailed implementation manners

[0066] The present invention relates to a design method for a direct thrust intelligent controller of a turbofan engine, comprising the following steps:

[0067] Step 1), select the strategy and evaluation network structure and parameters, and design the form of the direct thrust controller considering the protection of key safety parameters and the form of the reward of the reinforcement learning environment

[0068] Step 1.1), select the LSTM RNN network as the strategy network and evaluation network of reinforcement learning, and define the input and output of the strategy network; select the atmospheric environment parameters and the measurable state parameters of the engine to form the state s of the engine at time t t , select the current state and the previous n states of the engine to form the engine state group X at time t t =[s t-2 , s t-1 , s t as the input of the strategy network; the output of the strategy network is two vectors with a dimension of 4, respectively representing the mean u t of the action a t given by the strategy network under the current state X t and the variance σ t , the action a t under the current state X t =[a u,t a σ,t ; by defining the input and output of the strategy network, the input-output relationship formed by X t =[s t-2 , s t-1 , s t →Y t =[a u,t a σ,t is obtained; the input of the evaluation network is the same as the input of the strategy network, and the output is a scalar, representing the value v(s t |a t ) of taking the action a t under the current engine state X t , and the input-output relationship formed by X t =[s t-2 , s t-1 , s t →Y t =[v(s t |a t )] is obtained;

[0069] Step 1.2), define the form of the controller, add an integral link to the controller, and obtain the fuel flow rate W fb The controller expressions for the fuel flow rate W and the nozzle area A8 are

[0070]

[0071]

[0072] where g1 and g2 are the inertia coefficients of the changes in the fuel flow rate and the nozzle area, and e F is the error of thrust control F ref -F, W fb,t represents the normalized initial state of the fuel at time t, A 8,t represents the normalized initial state of the nozzle area at time t, e F represents the thrust error, ΔW fb represents the fuel flow rate increment, represents the fuel integral parameter, represents the nozzle integral parameter;

[0073] Step 1.3), design the reward form of the reinforcement learning environment. The reward consists of the control accuracy reward r e,t , the control stability reward r s,t and the critical parameter overrun penalty r l,t which consists of three parts;

[0074] The control accuracy reward r e,t is

[0075]

[0076] F ref represents the thrust command, and F represents the normalized thrust command;

[0077] The control stability reward r s,t is a constant

[0078] r s,t = 0.1

[0079] The critical safety parameter overrun penalty r l,t is

[0080]

[0081] respectively represent the fan surge penalty, the compressor surge penalty, the overrun penalty of the total pressure at the outlet of the high-pressure compressor, the overrun penalty of the total temperature at the outlet of the low-pressure turbine, and the overrun penalty of the low-pressure rotor speed;

[0082] The total reward at time t can be written as

[0083] r t = r e,t + r s,t + r l,t 。

[0084] Step 2), train the full envelope intelligent direct thrust controller: Based on the continuous policy gradient reinforcement learning algorithm, use the component-level model to build an environment for exploration, and train the agent policy network and evaluation network with the experience obtained from the exploration;

[0085] Step 2.1), set the learning rate, maximum number of episodes, episode length, policy update frequency, batch data dimension, and number of policy updates according to experience;

[0086] Step 2.2), in the same episode, randomly select the operating point, initial state, and thrust command, and conduct exploration in the set environment until the episode time limit is reached; record the state, action, and reward information during the exploration process to generate an experience pool for agent update;

[0087] Step 2.2.1), randomly select the operating point (H, Ma) and initial state within the envelope, find the feasible thrust range at this point, and randomly select a thrust value within this range as the thrust command F ref ;

[0088] Step 2.2.2), generate the engine state group X t = [s t-2 , s t-1 , s t at time t according to the engine state, and use it as the input of the policy network to obtain the action a t under the current state X t = [a u,t a σ,t ; According to the mean u t and variance σ t of the action, sample the current action under the normal distribution, obtain the engine control quantity through the controller form defined in Step 1, use the control quantity as the input of the component-level model, and generate the engine state group X t+1 at time t + 1 through dynamic calculation after the component-level model state at time t + 1. Calculate the reward r t+1 at time t + 1 according to the reward form determined in Step 1; Record the engine state groups X t and X t+1 , the action a t of the policy network, and the reward r t+1 during one-step dynamic process, and regard them as a set of data to be added to the experience pool for agent training;

[0089] Step 2.2.3), according to the method described in Step 2.2.2), continuously sample at the same operating point until the round time length upper limit is reached.

[0090] Step 2.3), repeat the steps of Step 2.2). When the update round number is reached, update the agent's policy and value network.

[0091] Step 2.3.1), for a set of data in the experience pool in Step 2.2), obtain the return G t = r t + r t+1 +…+ r T ; the subscript T represents the time series at the last moment in a round. With the help of Markov recursion, this process can be simplified to G t = r t+1 + γG t+1 , where γ is the discount factor, indicating the importance of future moments for the current moment; take the engine state group X t+1 at time t + 1 as the input of the evaluation network to obtain the value V t+1 , and use V t+1 to replace G t+1 , to get G t = r t+1 + γV t+1 ;

[0092] Step 2.3.2), by comparing the return G t and the value V t , calculate the advantage of the action (a t |X t ) at time t Compare the probabilities π θ (a t |X t ) of the new policy and the old policy for choosing the current action and to obtain the probability ratio By setting the clipping parameter ε and restricting it between [1 - ε, 1 + ε], the policy update amplitude is restricted;

[0093] Step 2.3.3), based on the PPO algorithm, calculate the loss function and update the agent's policy network and value network.

[0094] Step 3), test the control performance of the agent within the full envelope range and optimize the network structure and parameters;

[0095] Step 3.1), randomly generate l groups of 1×3 arrays c i = [W fb,i , A 8,i , F i (i = 1,..., l), and form Where W fb,i represents the normalized initial state of the fuel in the i-th group of tests, and A 8,i represents the normalized initial state of the nozzle area in the i-th group of tests, and F i represents the normalized command of the thrust in the i-th group of tests, and W fb,i ∈ [0, 1], A 8,i ∈ [0, 1], F i ∈ [0, 1];

[0096] Step 3.2), at the same operating point (H, Ma), respectively obtain the maximum thrust F max and the minimum thrust F min , the maximum fuel flow rate W fb,max and the minimum fuel flow rate W fb,min , the maximum nozzle area A 8,max and the minimum nozzle area, and establish a number of test cases according to in Step 3.1), and perform denormalization to obtain the initial fuel, the initial nozzle area, and the thrust command value;

[0097] W fb,ini,i = W fb,i · (W fb,max - W fb,min ) + W fb,min

[0098] A 8,ini,i = A 8,i · (A 8,max - A 8,min ) + A 8,min

[0099] F ref,i = F i · (F max - F min ) + F min

[0100] Step 3.3), based on the obtained policy network and the test cases generated in Step 3.2), control the thrust;

[0101] Step 3.3.1), based on the component-level model, generate the engine state group X t at time t, and use X t as the input of the policy network to obtain the action a t at the current state X t = [a u,t a σ,t ;

[0102] Step 3.3.2), in the actual control process of the agent, discard the random sampling of the action distribution and directly use the mean value au,t As an action, substitute it into the controller expression described in step 1), calculate the control quantity of the engine, and use it as the input of the engine to complete the control of the engine;

[0103] Step 3.3.3), repeat the above process until the control period ends.

[0104] Step 3.4), calculate the reward value r at each moment during the control process according to the reward expression in step 1) t (t = 1,..., T), and record the sum of the reward values at each moment as the score of this test case at this operating point

[0105] Step 3.5), within the full envelope range, divide n×m grid points according to altitude and Mach number, conduct tests at these grid points respectively, and record the score of each point, the average score of the full envelope, and the variance;

[0106] Step 3.6), evaluate the controller performance according to the score, and adjust the network structure parameters and training parameters.

[0107] Embodiment

[0108] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0109] According to the following embodiments, the present invention can be better understood. However, those skilled in the art can easily understand that the specific material ratios, process conditions and their results described in the embodiments are only used to illustrate the present invention, and should not and will not limit the present invention described in detail in the claims.

[0110] The object described in the embodiment is a certain type of two-spool military turbofan engine, which contains two rotor components, a high-pressure shaft and a low-pressure shaft. Its structural schematic diagram is as Figure 2 shown. In this embodiment, the control quantities of the turbofan engine are the fuel flow rate W fb and the area A8 of the tail nozzle, and the controlled quantity is the thrust F.

[0111] Figure 3 is the control architecture diagram of the direct thrust intelligent control method for turbofan engines based on reinforcement learning described in the present invention. The part within the box in the figure is the controller designed and optimized in this paper. When designing the control law offline, the policy and value networks are updated synchronously through the engine state and reward information. When using the control law online, the value network stops updating, and only the action mean output by the policy network is used for thrust feedback control. Since the thrust is not measurable, the thrust used in the present invention is estimated by a state estimator, which will not be elaborated here as it is not the content of the present invention.

[0112] Step A: Select the strategy and evaluation network structure and parameters, and design a controller and reward form that consider the direct control of thrust and the limit protection of key parameters;

[0113] First, it is necessary to determine the network structures of the reinforcement learning policy network and evaluation network. Since the engine is a strongly nonlinear object and both the input and output of the engine are time series, the LSTM RNN network is selected as the reinforcement learning policy network and evaluation network. Select the atmospheric environment parameters altitude H, Mach number Ma, and the engine measurable state parameters low-pressure shaft speed n L , high-pressure shaft speed n H , turbine pressure ratio EPR, high-pressure compressor outlet atmospheric pressure P3, low-pressure turbine outlet temperature T6, and thrust error e F to form the engine state x at time t t = [H t , Ma t , n L,t , n H,t , EPR t , P 3,t , T 6,t , e F , and select the current and previous two moments of the engine state to form the engine state group X at time t t = [s t-2 , s t-1 , s t as the input of the LSTM RNN neural network; the output of the policy network is two vectors with a dimension of 4, respectively representing the mean u t and variance σ t of the action a t given by the policy network under the current state X t . The four dimensions in the two vectors respectively represent four control quantities, namely the fuel flow increment ΔW fb , the afterburner nozzle area ΔA8, the fuel integration parameter the nozzle integration parameter written as and Denote Y t = a t = [a u,t a σ,t , and obtain the network input-output relationship formed by X t = [s t-2 , s t-1 , s t → Y t = [a u,t a σ,t . The input of the evaluation network is the same as that of the policy network, and the output of the policy network is a scalar, representing the current engine taking action a under the state X t ​t The value v(s t |a t ), to obtain the network input-output relationship formed by X t =[s t-2 , s t-1 , s t → Y t =[v(s t |a t )].

[0114] Next, it is necessary to design the form of the controller. To reduce the steady-state error of the control, an integral link is added to the controller. The expressions of the control variables, the fuel flow rate W fb and the area A8 of the tail nozzle are

[0115]

[0116]

[0117] where g1 and g2 are the inertia coefficients of the fuel flow rate and the change in the area of the tail nozzle, and e F is the error of the thrust control F ref -F.

[0118] Finally, to achieve the tracking control of multiple control variables for the thrust within the full envelope, considering the limit protection of the key safety parameters in the dynamic process, the form of the reinforcement learning environment reward is designed. The total reward consists of three parts: the accuracy reward, the stability reward, and the penalty for exceeding the limit of key parameters. The accuracy reward r e,t is

[0119]

[0120] The stability reward r s,t is a constant

[0121] r s,t =0.1

[0122] The penalty for exceeding the limit of key parameters r l,t is

[0123]

[0124] The expressions of the parameters in the formula are shown in Table 1

[0125] Table 1 Expression of the safety reward for key parameters

[0126]

[0127] The total reward at time t can be written as r t =r e,t +rs,t +r l,t 。

[0128] Step B: Train the full envelope intelligent direct thrust controller;

[0129] Select the network structure parameters and training parameters. The selected parameters are shown in Table 2.

[0130] Table 2 Network Structure and Training Parameters

[0131]

[0132] Next, it is necessary to generate the empirical data in the data pool for use in policy update.

[0133] Randomly select the operating points (H, Ma) and initial states within the envelope, and randomly select the thrust command F between the maximum and minimum thrusts ref . At time t, generate the engine state group X t =[s t-2 , s t-1 , s t and use it as the input of the policy network to obtain the action a t at the current state X t =[a u,t a σ,t . According to the mean u t and variance σ t of the action, sample under the normal distribution to obtain the current fuel flow rate ΔW fb , the area of the tail nozzle ΔA8, the fuel integration parameter and the nozzle integration parameter Obtain the engine control quantity through the controller form defined in Step A. Use the control quantity as the input of the component-level model. After dynamic calculation, obtain the model output at the next moment to get the engine state group X t+1 at time t + 1. Calculate the reward r t+1 at time t + 1 according to the reward form in Step A). Record the engine state groups X t and X t+1 , the action a t of the policy network, and the reward r t+1 in one-step dynamic process to form the data group at time t and add it to the data pool for training.

[0134] Continuously sample according to the above method until the upper limit of the episode time length is reached, and complete the sampling of this episode. Change the operating point and continue sampling until the set update episodes are reached. At this time, it is necessary to update the policy and value networks.

[0135] For a set of data in the data pool, take the discount factor γ = 0.99 to calculate the state return G at time t t = r t+1 + γG t+1 . Among them, the state return G at time t+1 t+1 can be replaced by the value V obtained through the value evaluation network, resulting in t+1 G

[0136] = r t + γV t+1 . t+1 .

[0137] By comparing the state return G obtained through exploration at time t t and the value V obtained through the value network t , the advantage of taking the action (a t |X t ) at time t can be obtained

[0138]

[0139] Through importance sampling, compare the probability π θ (a t |X t ) of the new policy choosing the current action and the probability of the old policy choosing the current action to obtain the probability ratio

[0140]

[0141] By the clipping parameter ε, limit the policy update amplitude. In the present invention, take ε = 0.2, and limit the probability ratio r t (θ) between [1 - ε, 1 + ε] to obtain the following CLIP loss function

[0142]

[0143] Finally, by calculating the value evaluation network loss function and the cross-entropy S[π(θ)](X t ), obtain the loss function at time t

[0144]

[0145] Similarly, calculate the loss function for all data in the data pool, and train in batches according to the batch size. Calculate the average loss function for the divided batches, and calculate the network update gradient according to the loss function to update the policy network and the value network.

[0146] Step C: Design an evaluation system for the full envelope thrust control of the engine, and optimize the network structure and training parameters based on the evaluation indicators;

[0147] Select l = 100, and randomly generate one hundred groups of 1×3 random arrays c i =[W fb,i , A 8,i , F i (i = 1,..., 100) to form where W fb,i represents the normalized initial state of fuel in the i-th group of experiments, A 8,i represents the normalized initial state of the nozzle area in the i-th group of experiments, F i represents the normalized command of thrust in the i-th group of experiments, and W fb,i ∈[0, 1], A 8,i ∈[0, 1], F i ∈[0, 1].

[0148] At the same operating point (H, Ma), obtain the maximum thrust F max and the minimum thrust F min , the maximum fuel flow rate W fb,max and the minimum fuel flow rate W fb,min , the maximum nozzle area A 8,max and the minimum nozzle area respectively. According to c i =[W fb,i , A 8,i , r i (i = 1,..., 100) for denormalization to obtain the initial fuel, initial nozzle area and thrust command values, and establish one hundred test cases. The denormalization process can be expressed as

[0149] W fb,ini,i =W fb,i ·(W fb,max -W fb,min )+W fb,min

[0150] A 8,ini,i =A 8,i ·(A 8,max -A 8,min )+A 8,min

[0151] F ref,i =F i ·(F max -F min )+F min

[0152] Based on the designed controller, control is carried out on these test samples. In each control cycle, the state x of the engine at the current moment is recorded t =[[H t , Ma t , n L,t , n H,t , EPR t , P 3,t , T 6,t , e F , generating the state group X of the engine at time t t , and taking X t as the input of the policy network, obtaining the output a of the policy network under the current state X t = [a t a u,t a σ,t . When actually using the controller, the random sampling of the action distribution is discarded, and the action mean is directly used as the fuel flow increment ΔW fb , the area of the tail nozzle ΔA8, the fuel integral parameter the nozzle integral parameter are substituted into the controller expression in step A to obtain the control quantities of the engine: the fuel flow W fb,t and the area of the tail nozzle A 8,t , and taking them as the input of the engine to complete the control of the engine.

[0153] During the control process, according to the reward form designed in step A, calculate the reward value r at each moment in the control process t (t = 1,..., T), and record the total reward as the score of a single test sample at this operating point After completing all test samples at this operating point, calculate the average score at this operating point as the score of this operating point.

[0154] Within the full envelope, 32 points are divided by altitude intervals and 11 points are divided by Mach number intervals, forming 32×11 = 352 operating points within the full envelope, and tests are carried out at these operating points. According to the average reward value of each operating point calculate the overall score of direct thrust control within the entire envelope and the score variance Based on this, adjust the controller, reward, network structure and training parameters to obtain the optimal control effect.

[0155] Randomly select points within the full envelope to test the designed direct thrust controller and test the control effect. Figure 4The simulation results of Test Case 1 are shown. This test was conducted at an altitude of 4092 m with a Mach number of 0.6. In this test, the adjustment time was 0.02 s, the overshoot was 0.78%, the steady-state error was 0.14%, the dynamic quality of thrust control was good, and the key safety parameters of the system were all within the limits. Figure 5 The simulation results of Test Case 2 are shown. This test was conducted at an altitude of 804 m with a Mach number of 0.4. In this test, the adjustment time was 0.16 s, the overshoot was 4.54%, the steady-state error was 0.73%, the dynamic quality of thrust control was good, and the key safety parameters of the system were all within the limits.

[0156] Next, a thrust control test was conducted in the full envelope. Figure 6 It is a flight path diagram. The state of the aircraft from startup to landing was simulated according to the flight path, but the time of each stage was shortened to verify the quality of the control system. Figure 7 It is the simulation result. From the simulation result, it can be seen that during the simulated flight, although the change speed of the thrust command was very fast, under the action of the controller of the present invention, the thrust always followed the command, and the maximum thrust error was only 0.1%, and all the key safety parameters of the system were within the safe range. It can be seen that the present invention is effective.

[0157] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A direct thrust intelligent control method for a turbofan engine based on reinforcement learning, characterized in that: It includes the following steps: Step 1), select the policy and evaluation network structure and parameters, and design the form of the direct thrust controller considering the protection of key safety parameters and the form of the reward of the reinforcement learning environment; Step 2), train the full envelope intelligent direct thrust controller: based on the continuous policy gradient reinforcement learning algorithm, use the component-level model to build the environment for exploration, and train the intelligent agent policy network and evaluation network through the experience obtained from the exploration; Step 3), test the control performance of the intelligent agent within the full envelope range, and optimize the network structure and parameters; The specific steps in the above Step 1) are as follows: Step 1.1), select the LSTMRNN network as the policy network and evaluation network for reinforcement learning, and define the input and output of the policy network; select the atmospheric environment parameters and the measurable state parameters of the engine to form the engine state s at time t t , select the current and the previous n states of the engine to form the engine state group X at time t t = [s t-2 , s t-1 , s t , as the input of the policy network; the output of the policy network is two vectors with a dimension of 4, respectively representing the mean u t of the action a t given by the policy network under the current state X t and the variance σ t , and the action a t under the current state X t = [a u,t a σ,t ; by defining the input and output of the policy network, the input-output relationship formed by X t = [s t-2 , s t-1 , s t → Y t = [a u,t a σ,t is obtained; the input of the evaluation network is the same as that of the policy network, and the output is a scalar, representing the value v(s t | a t ) of taking the action a t under the current engine state X t , and the input-output relationship formed by X t = [s t-2 , s t-1 , s t → Y t = [v(s t | a t )] is obtained; Step 1.2), define the form of the controller, add an integral link to the controller, and obtain the controller expressions for the fuel flow rate W fb and the nozzle area A8 as wherein, g1 and g2 are inertia coefficients of the fuel flow rate and the nozzle area change, W fb,t represents the normalized initial state of the fuel at time t, A 8,t represents the normalized initial state of the nozzle area at time t, e F represents the thrust error, ΔW fb represents the fuel flow rate increment, represents the fuel integration parameter, represents the nozzle integration parameter; Step 1.3), design the reward form of the reinforcement learning environment. The reward consists of three parts: the control accuracy reward r e,t , the control stability reward r s,t , and the penalty r l,t for exceeding the limit of key parameters; Control precision reward r e,t For F ref represents the thrust command, and F represents the normalized command of the thrust; Control stability reward r s,t Is a constant r s,t =0.1 Critical safety parameter overrun penalty r l,t For respectively represent the fan surge penalty, the compressor surge penalty, the overlimit penalty of the total pressure at the outlet of the high-pressure compressor, the overlimit penalty of the total temperature at the outlet of the low-pressure turbine, and the overlimit penalty of the low-pressure rotor speed; The total reward at time t can be written as r t =r e,t +r s,t +r l,t ; The specific steps of the training method of the full envelope intelligent direct thrust controller in the above Step 2) are as follows: Step 2.1), set the learning rate, maximum number of episodes, episode length, policy update frequency, batch data dimension, and number of policy updates according to experience; Step 2.2), in the same episode, randomly select the operating point, initial state, and thrust command, and conduct exploration in the set environment until the episode time limit is reached; record the state, action, and reward information during the exploration, and generate an experience pool for the intelligent agent to update and use; Step 2.3), repeat the steps of Step 2.2), and when the number of updated episodes is reached, update the policy and value network of the intelligent agent; The specific steps of generating the experience pool in the above Step 2.2) are as follows: Step 2.2.1), randomly select an operating point (H, Ma) and an initial state within the wire wrapping, find the feasible thrust range at this point, and randomly select a thrust value within this range as the thrust command F ref ; Step 2.2.2), generate the engine state group X according to the engine state at time t t =[s t-2 , s t-1 , s t , and use it as the input of the policy network to obtain the action a t at the current state X t =[a u,t a σ,t ; According to the mean u t and variance σ t of the action, sample the current action under the normal distribution, obtain the engine control quantity through the controller form defined in step 1), use the control quantity as the input of the component-level model, after dynamic calculation, generate the engine state group X t+1 at time t + 1 through the component-level model state at time t + 1, and calculate the reward r t+1 at time t + 1 according to the reward form determined in step 1); Record the engine state groups X t and X t+1 , the action a t of the policy network, and the reward r t+1 in one-step dynamic process, and regard them as a set of data to be added to the experience pool for the agent to train with; Step 2.2.3), continuously sample at the same operating point according to the method described in Step 2.2.2) until the episode time length limit is reached; The specific steps of updating the policy network and value network in the above Step 2.3) are as follows: Step 2.3.1), in Step 2.2), for a set of data in the experience pool, calculate the return G t = r t + r t+1 + … + r T ; The subscript T represents the time series at the last moment in a round; With the help of Markov recursion, this process can be simplified to G t = r t+1 + γG t+1 , where γ is the discount factor, indicating the importance of future moments for the current moment; Take the engine state group X t+1 at time t + 1 as the input of the evaluation network to obtain the value V t+1 , and use V t+1 to replace G t+1 , to get G t = r t+1 + γV t+1 ; Step 2.3.2), calculate the advantage of the action (a t |X t ) at time t by comparing the return Gt and the value Vt Compare the probabilities π θ (a t |X t ) of the current action selected by the new policy and the old policy and obtain the probability ratio Limit the policy update amplitude by setting a clipping parameter ε and restricting it between [1 - ε, 1 + ε]; Step 2.3.3), based on the PPO algorithm, calculate the loss function, and update the intelligent agent policy network and value network; The specific steps of optimizing the intelligent agent network structure and parameters in the above Step 3) are as follows: Step 3.1), randomly generate l groups of 1×3 arrays c i =[[W fb,i ,[[A 8,i ,[[F i , i = 1, ..., l, and form where W fb,i represents the normalized initial state of fuel in the i-th group of experiments, A 8,i represents the normalized initial state of the nozzle area in the i-th group of experiments, F i represents the normalized command of thrust in the i-th group of experiments, and W fb,i ∈[0, 1], A 8,i ∈[0, 1], F i ∈[0, 1]; Step 3.2), at the same operating point (H, Ma), respectively obtain the maximum thrust F max and the minimum thrust F min , the maximum fuel flow rate W fb,max and the minimum fuel flow rate W fb,min , the maximum nozzle area A 8,max and the minimum nozzle area. According to in Step 3.1), establish a number of test cases, and perform anti-normalization to obtain the initial fuel, the initial nozzle area, and the thrust command value; W fb,ini,i = W fb,i · (W fb,max - W fb,min ) + W fb,min A 8,ini,i = A 8,i ·(A 8,max - A 8,min ) + A 8,min F ref,i = F i · (F max - F min ) + F min Step 3.3), based on the obtained policy network and the test cases generated in Step 3.2), control the thrust; Step 3.4), calculate the reward value r at each moment during the control process according to the reward expression in Step 1) t , where t = 1,..., T, and denote the sum of the reward values at each moment as the score of this test case at this operating point Step 3.5), within the full envelope range, divide into n×m grid points according to altitude and Mach number, conduct tests on these grid points respectively, and record the score of each point, the average score of the full envelope, and the variance; Step 3.6), evaluate the controller performance according to the score, and adjust the network structure parameters and training parameters.

2. The intelligent direct thrust control method for a turbofan engine based on reinforcement learning according to claim 1, wherein: The specific steps of thrust control based on the policy network in the above Step 3.3) are as follows: Step 3.3.1), generate the engine state group X at time t based on the component-level model t , and use X t as the input of the policy network to obtain the action a t under the current state X t = [a u,t a σ,t ; Step 3.3.2), during the actual control process of the agent, the random sampling of the action distribution is discarded, and the mean value a u,t is directly used as the action and substituted into the controller expression described in Step 1) to calculate the control quantity of the engine, which is used as the input of the engine to complete the control of the engine; Step 3.3.3), repeat the above process until the control cycle ends.

Citation Information

Patent Citations

  • Turbofan engine thrust prediction method and controller based on neural network

    CN110579962A

  • Variable-cycle engine modeling method and variable-cycle engine component-level model

    CN111914365A