Actuating mechanism fault-tolerant control method based on deep reinforcement learning
Through the actuator fault-tolerant control method based on deep reinforcement learning, the problem that aircraft engine actuators are difficult to maintain performance and safety when facing failures is solved, automatic fault diagnosis and fault-tolerant control are achieved, the shortcomings of traditional methods are overcome, and the efficient and reliable operation of the engine is ensured.
Patent Information
- Application Number
- CN202311540167.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
When the aircraft engine actuator faces system failures, uncertainties or abnormal conditions, it is difficult to maintain the performance and safety of the aircraft, and traditional fault-tolerant controls are difficult to effectively deal with stuck faults.
The actuator fault-tolerant control method based on deep reinforcement learning is adopted. By establishing an engine mathematical model and a deep reinforcement learning agent simulation training environment, the agent is trained using the deep deterministic policy gradient (DDPG) algorithm to achieve automatic fault diagnosis and fault-tolerant control.
When different types of faults occur in the actuator, the input of the actuator is adjusted to ensure the efficient and reliable operation of the aircraft engine to the greatest extent, and the active fault tolerance control is quickly and smoothly, overcoming the shortcomings of traditional methods that are difficult to deal with stuck faults.
Smart Images

Figure CN120020653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft engine control, and in particular to an actuator fault-tolerant control method based on deep reinforcement learning. Background Technology
[0002] In the field of aviation, especially in the field of aircraft engines, the reliability of actuators is of vital importance. The goal of fault-tolerant control is to ensure that these actuators can maintain the performance and safety of the aircraft in the face of system failures, uncertainties or abnormal situations. Because any failure of the engine during flight may lead to serious consequences, the fault-tolerant control of aircraft engine actuators has extremely high requirements.
[0003] The application of fault-tolerant control on actuators can enhance the safety, stability and robustness of the system. However, during the flight, some parameters of the engine itself will also change. Due to the limited ground intervention capability of the aircraft, these internal and external changes will bring great challenges to the fault diagnosis and fault-tolerant control of the actuator. This requires the controller to have strong robustness and adaptability to ensure the fault tolerance of the entire control loop.
[0004] With the development of artificial intelligence technology, researchers have gradually expanded the methods of active fault-tolerant control and used intelligent learning methods to solve fault-tolerant control problems. Fault-tolerant control based on artificial intelligence technology belongs to the category of active fault-tolerant control. It has received widespread attention due to its good adaptability and robustness. Deep reinforcement learning, as a machine learning technology for autonomous learning and decision-making, provides new possibilities for improving the fault-tolerant control of aircraft engine actuators. It allows the system to automatically learn how to adjust the actuators in the face of system failures or abnormal conditions to minimize the adverse effects on flight performance. SUMMARY OF THE INVENTION
[0005] In order to solve the above problems, the present invention proposes a fault-tolerant control method for actuators based on deep reinforcement learning, so that when an actuator fails, the controller automatically identifies the fault and performs fault-tolerant control to ensure flight safety.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A fault-tolerant control method for an actuator based on deep reinforcement learning, comprising the following steps:
[0008] (1) Establish a mathematical model of the engine including the actuator;
[0009] (2) Establish a deep reinforcement learning agent simulation training environment for engine actuator fault-tolerant control tasks;
[0010] (3) Train the agent using the Deep Deterministic Policy Gradient (DDPG) algorithm;
[0011] (4) Arrange the trained agent on the microcontroller to perform fault-tolerant control on the faulty actuator.
[0012] For the engine mathematical model with actuators established in step (1), based on the control structure of the actuators, establish a model including various faults; use the component-level modeling method to establish the variable cycle engine mathematical model.
[0013] Furthermore, for the model including various faults established based on the control structure of the actuators, taking the fuel metering device as an example, the actuator faults include bias fault, jamming fault, and performance degradation, which are specifically expressed as follows:
[0014] Actuator bias fault:
[0015] W f0 = W f + ΔW f
[0016] where ΔW f is a constant. When ΔW f = 0, the actuator is in the normal working state.
[0017] Actuator jamming fault:
[0018]
[0019] where a is a constant, and the value range of a is obtained from the working range envelope of the actuator.
[0020] Actuator performance degradation:
[0021] W f0 = W f + t * k
[0022] where t is time and k is the degradation coefficient.
[0023] Furthermore, for the method of using component-level modeling to establish the variable cycle engine mathematical model, establish the thermodynamic models of each engine component according to the modular idea, and then perform the overall calculation of each part of the engine.
[0024] The establishment of the deep reinforcement learning agent simulation training environment for the engine actuator fault-tolerant control task in step (2) includes:
[0025] Step (2.1), given the controlled variables of the engine controller as the high-pressure rotor speed n H and the turbine pressure ratio π t , select the control variable as the fuel flow rate W f, the throat area A of the tail nozzle 8 , the outer flow area A of the CDFS 125 ;
[0026] Step (2.2), select the state variable s t as the rotational speed n H , the rotational speed error Δn H , the pressure ratio π t , the pressure ratio error Δπ t , and the action variable a t as the fuel flow rate W f , the throat area A of the tail nozzle 8 and the outer flow area A of the CDFS 125 , the state variable s at time t t and the action variable a t are expressed as follows:
[0027] s t = [n H , Δn H , π t , Δπ t
[0028] a t = [W f , A 8 , A 125
[0029] Step (2.3), design the action network and evaluation network of the agent:
[0030] The action network π of the agent θ consists of an input layer, a fully connected layer 1, a Relu activation function layer, a fully connected layer 2, a Relu activation function layer, an output layer, and a Tanh activation function layer in sequence. The input parameter of the action network is the state variable s t , and the output parameter is the action variable a t ;
[0031] The evaluation network Q of the agent ω consists of an input layer, a fully connected layer 1, a Relu activation function layer, a fully connected layer 2, a Relu activation function layer, an output layer, and a Tanh activation function layer in sequence. The input parameters of the evaluation network are the state variable s t and the action variable a t , and the output parameter is the expectation of the reward that the current state variable s t and the action variable a t can obtain;
[0032] Step (2.4), design the reward function r according to the state variable and the action variable, including the dense reward r 1 , and the sparse reward r 2 And the reward r for the change amplitude of the control quantity 3 , specifically as follows:
[0033] r = r 1 + r 2 + r 3
[0034] Where:
[0035]
[0036]
[0037]
[0038] Based on the deep reinforcement learning agent simulation training environment built in step (2), the agent is trained using the Deep Deterministic Policy Gradient (DDPG) algorithm, specifically including the following steps:
[0039] Step (3.1), load the actuator model and the engine component-level model;
[0040] Step (3.2), establish the target action network π θ - and the target evaluation network Q ω - , whose structure is the same as that of the action network π θ and the evaluation network Q ω ;
[0041] Step (3.3), initialize the weight parameters θ and ω of the action network π θ and the evaluation network Q ω with random parameters, and then copy the weight parameters θ and ω to the target network parameters θ - and ω - ;
[0042] Step (3.4), initialize the experience replay pool R, and set the number of training episodes E, the simulation time T, the simulation sampling step size ΔT, the discount factor γ, and the Soft-Update update coefficient σ;
[0043] Step (3.5), start the episode loop;
[0044] Step (3.6), randomly initialize the engine model, the control target, and observe the initial state variable s 0 ;
[0045] Step (3.7), start the simulation loop;
[0046] Step (3.8), inject actuator faults at random time steps;
[0047] Step (3.9), the action network outputs the action variable a according to the state variable s t ; t ;
[0048] Step (3.10), observe the action variable a t Change the input value of the control quantity, the engine model runs for one step length, and calculate the reward r t , and the environmental state variable is s at this time t+1 ;
[0049] Step (3.11), store the current information frame e t (s t , a t , r t , s t+1 ) and put it into the experience replay pool R;
[0050] Step (3.12), randomly sample a small batch from the replay pool R, and update each network parameter according to the following formula;
[0051]
[0052]
[0053] ω - =τω+(1 - τ)τω -
[0054] θ - =τθ+(1 - τ)τθ -
[0055] Step (3.13), execute steps (3.7) - (3.12) until the simulation loop ends;
[0056] Step (3.14), execute steps (3.5) - (3.13) until the episode loop ends;
[0057] Step (4), deploy the trained agent on the single-chip microcomputer to perform fault-tolerant control on the faulty actuator. After training, save the weight parameters of the agent's action network, and deploy the forward channel of the action network to the single-chip microcomputer. The single-chip microcomputer receives the engine state variable s transmitted from the upper computer t as the input of the action network, and outputs the action variable a t and send it to the upper computer actuator to achieve fault-tolerant control of the actuator.
[0058] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0059] (1) The actuator fault-tolerant control method based on deep reinforcement learning proposed in the present invention is a general data-based active fault-tolerant control method. Its active fault-tolerant control framework can not only automatically realize fault diagnosis, but also realize active fault-tolerant control. When different types of faults occur in the actuator, by adjusting the actuator input, the efficient and reliable operation of the aircraft engine is guaranteed to the greatest extent, and fast and smooth active fault-tolerant control is realized;
[0060] (2) The method of the present invention can overcome the disadvantage of the traditional fault-tolerant control that is difficult to handle stuck faults. Through multivariable control, the control effect of the faulty actuator is distributed to the fault-free actuator, and the closed-loop control of the target quantity is continued, solving the coupling problem between the control loops;
[0061] (3) The method of the present invention has good fault-tolerant control effect for different faults of different actuators, strong portability and wide versatility. Brief Description of the Figures
[0062] Figure 1 is the technical principle diagram of the present invention;
[0063] Figure 2 is a graph showing the changes in the reward function during the agent training process in an embodiment of the present invention;
[0064] Figure 3(a) shows the high pressure rotor speed n in the embodiment of the present invention H Control effect diagram;
[0065] Figure 3(b) shows the pressure drop ratio π in the embodiment of the present invention t Control effect diagram. Specific implementation method
[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention
[0067] Compare with Figure 1 , a fault-tolerant control method for actuators based on deep reinforcement learning, comprising the following steps:
[0068] (1) Establish a mathematical model of the engine including the actuator;
[0069] Based on the control structure of the actuator, a model containing multiple faults is established. Taking the fuel metering device as an example, the actuator faults include bias fault, stuck fault, and performance degradation, which are specifically expressed as follows:
[0070] Actuator offset fault:
[0071] W f0 = W f + ΔW f
[0072] where ΔW f is a constant. When ΔW f = 0, the actuator is in a normal working state.
[0073] Actuator jamming fault:
[0074]
[0075] where a is a constant, and the value range of a is obtained from the working range envelope of the actuator.
[0076] Actuator performance degradation:
[0077] W f0 = W f + t * k
[0078] where t is time and k is the degradation coefficient.
[0079] According to the modular idea, the thermodynamic models of each component of the variable cycle engine are established by using the component-level modeling method, and then the overall engine calculations are carried out.
[0080] (2) Establish a deep reinforcement learning agent simulation training environment for the fault-tolerant control task of the engine actuator;
[0081] Step (2.1), given the controlled variables of the engine controller as the high-pressure rotor speed n H and the turbine pressure ratio π t , select the control variables as the fuel flow rate W f , the throat area A of the nozzle 8 , and the outer flow area A of the CDFS 125 ;
[0082] Step (2.2), select the state variable s t as the speed n H , the speed error Δn H , the pressure ratio π t , the pressure ratio error Δπ t , the action variable a t as the fuel flow rate W f , the throat area A of the nozzle 8 and the outer flow area A of the CDFS 125 , the state variable s t and the action variable a t at time t are expressed as follows:
[0083] s t = [n H , Δn H , π t , Δπ t
[0084] a t = [W f , A 8 , A 125
[0085] Step (2.3), design the action network and evaluation network of the agent:
[0086] The action network π of the agent θ is successively composed of an input layer, a fully connected layer 1, a Relu activation function layer, a fully connected layer 2, a Relu activation function layer, an output layer, and a Tanh activation function layer. The input parameter of the action network is the state variable s t , and the output parameter is the action variable a t ;
[0087] The evaluation network Q of the agent ω is successively composed of an input layer, a fully connected layer 1, a Relu activation function layer, a fully connected layer 2, a Relu activation function layer, an output layer, and a Tanh activation function layer. The input parameters of the evaluation network are the state variable s t and the action variable a t , and the output parameter is the expectation of the reward that the current state variable s t and the action variable a t can obtain;
[0088] Step (2.4), design the reward function r according to the state variable and the action variable, including the dense reward r 1 , the sparse reward r 2 and the control variable change amplitude reward r 3 , specifically as follows:
[0089] r = r 1 + r 2 + r 3
[0090] Among them:
[0091]
[0092]
[0093]
[0094] (3) Use the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent;
[0095] Step (3.1), load the actuator model and the engine component-level model;
[0096] Step (3.2), establish the target action network π θ - and the target evaluation network Q ω - , whose structures are the same as those of the action network π θ and the evaluation network Q ω ;
[0097] Step (3.3), initialize the weight parameters θ and ω of the action network π θ and the evaluation network Q ω with random parameters, and then copy the weight parameters θ and ω to the target network parameters θ - and ω - ;
[0098] Step (3.4), initialize the experience replay pool R, and set the number of training episodes E, the simulation time T, the simulation sampling step size ΔT, the discount factor γ, and the Soft-Update update coefficient σ;
[0099] Step (3.5), start the episode loop;
[0100] Step (3.6), randomly initialize the engine model and the control target, and observe the initial state variable s 0 ;
[0101] Step (3.7), start the simulation loop;
[0102] Step (3.8), inject an actuator fault at a random time step;
[0103] Step (3.9), the action network outputs the action variable a according to the state variable s t ; t ;
[0104] Step (3.10), observe the action variable a t to change the control amount, input the value to run the engine model for one step, and calculate the reward r t , and at this time the environmental state variable is s t+1 ;
[0105] Step (3.11), store the current information frame e t (s t , a t , r t , s t+1 ) into the experience replay pool R;
[0106] Step (3.12), randomly sample a small batch from the replay pool R, and update each network parameter according to the following formula;
[0107]
[0108]
[0109] ω - =τω+(1 - τ)τω -
[0110] θ - =τθ+(1 - τ)τθ -
[0111] Step (3.13), execute Step (3.7) - Step (3.12) until the simulation loop ends;
[0112] Step (3.14), execute Step (3.5) - Step (3.13) until the episode loop ends;
[0113] Appendix Figure 2 shows the change of the reward function during the process of training the above agent using the DDPG algorithm. It can be seen that after 2300 episodes of training, the reward value has converged to a relatively high level.
[0114] (4) Arrange the trained agent on the microcontroller to perform fault-tolerant control on the faulty actuator. After training is completed, save the weight parameters of the agent's action network, and deploy the forward channel of the action network to the microcontroller. The microcontroller receives the engine state variable s t transmitted from the host computer as the input of the action network, and outputs the action variable a t and sends it to the host computer actuator to achieve fault-tolerant control of the actuator.
[0115] To verify the effectiveness of the actuator fault-tolerant control method proposed in the present invention, we conducted a detailed verification based on the actuator and engine model of this embodiment. The verification results are shown in Fig. 3(a) and Fig. 3(b). During the verification process, 0 - 4 seconds represent the normal operating state of the engine actuator. At 4 seconds, we introduced an actuator bias fault, which lasted until the end of the simulation. At the same time, we changed the control target pressure ratio π tt . It can be clearly seen from the verification results that when the actuator encounters a fault, our controller can automatically identify the fault and effectively control the engine to keep it in the target state. Even in the case of actuator faults, when changing the target pressure ratio π tt, the controller can still maintain excellent control effect. Both dynamic and steady-state errors are kept in a small range, ensuring the stable operation of the engine. The verification results show that the actuator fault-tolerant control method based on deep reinforcement learning proposed by us shows excellent robustness and adaptability in the face of actuator failures. This provides strong support for the stability and reliability of the engine system and proves the practicality and effectiveness of the invention in practical applications.
[0116] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field within the technical scope disclosed by the present invention can make equivalent replacements or changes based on the technical solution and inventive concept of the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A fault-tolerant control method for actuators based on deep reinforcement learning, characterized in that The following steps are involved: (1) Establish a mathematical model of the engine including the actuator; (2) Establish a deep reinforcement learning agent simulation training environment for engine actuator fault-tolerant control tasks; (3) Use the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent; (4) The trained intelligent agent is placed on a single-chip microcomputer to perform fault-tolerant control on the faulty actuator.
2. The method for fault-tolerant control of an actuator based on deep reinforcement learning according to claim 1, characterized in that The step (2) establishes a deep reinforcement learning agent simulation training environment for the engine actuator fault-tolerant control task, selects appropriate variable parameters and designs an appropriate reward function, including: Step (2.1), given that the controlled variable of the engine controller is the high-pressure rotor speed n H and turbine pressure ratio π t , select the control quantity as fuel flow W f , tail nozzle throat area A8, CDFS outer duct area A 125 ; Step (2.2), select the state variable s t is the speed n H , speed error Δn H , pressure drop ratio π t , pressure drop ratio error Δπ t , action variable a t is the fuel flow rate W f , tail nozzle throat area A8 and CDFS outer area A 125 , state variable s at time t t and action variable a t It is expressed as follows: s t =[n H ,Δn H ,p t ,Dp t ] a t =[W f ,A8,A 125 ] Step (2.3), design the action network and evaluation network of the intelligent agent: The agent's action network π θ It consists of input layer, fully connected layer 1, Relu activation function layer, fully connected layer 2, Relu activation function layer, output layer, and Tanh activation function layer. The input parameter of the action network is the state variable s t , the output parameter is the action variable a t ; The agent's evaluation network Q ω It consists of input layer, fully connected layer 1, Relu activation function layer, fully connected layer 2, Relu activation function layer, output layer, and Tanh activation function layer. The input parameter of the evaluation network is the state variable s t and action variable a t , the output parameter is the current state variable s t and action variable a t Expectation of being rewarded; Step (2.4), design the reward function r according to the state variables and action variables, including dense reward r1, sparse reward r2 and control amount change range reward r3, as follows: r=r1+r2+r3 in:
3. The method for fault-tolerant control of actuators based on deep reinforcement learning according to claim 1 is characterized in that In step (4), the trained agent is arranged on the single-chip microcomputer to perform fault-tolerant control on the faulty actuator. After the training is completed, the weight parameters of the agent action network are saved, and the forward channel of the action network is deployed to the single-chip microcomputer. The single-chip microcomputer receives the engine state variable s transmitted from the host computer. t As the input of the action network, the output action variable a t Send it to the upper computer actuator to realize fault-tolerant control of the actuator.