Fault prediction model and fault prediction method based on reinforcement learning
By using Actor-Critic reinforcement learning network structure and features extracted by the autoencoder in fault prediction, the problems of large fluctuations in predicting value and low accuracy caused by underutilization of timing correlation are solved, and more stable and accurate fault prediction is achieved.
Patent Information
- Application Number
- CN202510317618.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art does not fully utilize sensor monitoring data and timing correlation of predicted values at the time point in fault prediction, resulting in large fluctuations in predicted values and low prediction accuracy.
The Actor-Critic reinforcement learning network structure is adopted, and the fault time change is predicted through the Actor network output, combined with the low-dimensional key features extracted by the autoencoder and the multi-dimensional fault time features, the state variables of reinforcement learning are constructed to realize the training of the fault prediction model.
By making full use of the timing correlation and historical fault time information of the monitoring data, the fluctuations in the predicted value are significantly reduced and the accuracy and stability of fault prediction are improved.
Smart Images

Figure CN120216956A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to fault prediction and health management, and more specifically, relates to a fault prediction model and a fault prediction method based on reinforcement learning. Background Art
[0002] Fault prediction and health management technology uses a large amount of device monitoring data and prediction algorithms to evaluate the health status of devices in real time, control potential risks, and improve the safety and reliability of device operation. Fault prediction is a key component of PHM. Accurately predicting the time point when a device fails is crucial for formulating a reasonable maintenance plan and avoiding accident risks.
[0003] Generally speaking, the common paradigm of fault prediction methods is to calculate the time point when a future device fails by predicting the remaining useful life value of the current device, which can usually be divided into model-based methods and data-driven methods. Among them, model-based methods usually require sufficient prior knowledge of mechanical systems to construct an empirical model of the degradation process. Although the method based on an accurate degradation model is effective, it is difficult to establish an accurate physical model for complex machines in practical industrial applications. Data-driven methods aim to establish a complex mapping relationship between device monitoring data and degradation information. In recent years, due to the rapid development of deep learning (DL) and sensor technology, a large number of deep learning models have been used for data-driven fault prediction, but the application of these methods still faces some challenges. First, traditional DL methods shuffle the training set data and then randomly sample to train the model, which ignores the temporal sequence relationship of the monitoring data of the same device. Second, these methods do not consider the continuity of the predicted values during the device degradation process. The predicted values at adjacent time points are independent of each other, resulting in large fluctuations in the predicted values, which may lead to misjudgment of the device state and redundant maintenance. Using reinforcement learning (RL) methods can transform the device degradation process into a step-by-step interaction process between the agent and the environment in the model. On the one hand, it can make full use of the temporal correlation of the monitoring data, and on the other hand, it makes the prediction results more stable and smooth. At present, there are few studies on using RL methods for device fault prediction. In the existing research, when constructing the reinforcement learning state, the single-dimensional fault time information in the model is easily submerged. The reward function that only considers the prediction error at the current time step cannot capture the change of the prediction error, and the training method lacks guidance for the agent to learn the device degradation information, resulting in poor stability and accuracy of the prediction results. In addition, the (Twin Delayed Deep Deterministic policy gradient, TD3) algorithm used in the existing research is very sensitive to model parameters, and it is difficult to debug and train the model in practical applications. Summary of the Invention
[0004] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a fault prediction model and a fault prediction method based on reinforcement learning, which solve the problems of large fluctuations in prediction values and low prediction accuracy caused by the insufficient utilization of sensor monitoring data and the temporal correlation of predicted values of fault time points in fault prediction.
[0005] To achieve the above object, according to one aspect of the present invention, there is provided a fault prediction model based on reinforcement learning. The fault prediction model adopts an Actor-Critic reinforcement learning network structure. The input of the Actor network in the fault prediction model is the state variable s at the current time step t , and the output is the predicted change in fault time at the current time step. The s t = {e t , r t}, s t is the state variable at time step t, e t is the low-dimensional key feature at time step t, r t is the multi-dimensional fault time feature, is the predicted fault time at time step t - 1, is the predicted fault time at time step t - 2, is the predicted fault time at time step t - Dim;
[0006] After inputting the state variable corresponding to the current time step into the Actor network, the predicted fault time and the action reward value are calculated using the output predicted change in fault time. The state variable, the predicted change in fault time, and the corresponding action reward value at the current time step are stored as experience in the experience pool, and the experience is sampled from the experience pool to update the fault prediction model, thereby realizing the training of the fault prediction model;
[0007] Further preferably, the low-dimensional key feature e t is obtained by extraction using an autoencoder. The autoencoder includes an encoder and a decoder with symmetric structures, both of which are convolutional neural networks. The encoder is used to extract features from the preprocessed monitoring data to obtain the low-dimensional key feature, and the decoder reconstructs the signal using the output of the encoder. The autoencoder is trained in an unsupervised manner such that the reconstructed signal is close to the input preprocessed monitoring data, thereby obtaining the trained autoencoder. The trained autoencoder is used for the extraction of the low-dimensional key feature e t .
[0008] Further preferably, the loss function of the autoencoder is as follows:
[0009]
[0010] where n is the number of time steps, and x i and are the original signal and the reconstructed signal at the i-th time step respectively, is the mean squared error loss function.
[0011] Further preferably, the Actor-Critic reinforcement learning network structure includes an Actor network, a Critic network, and a Critic Target network.
[0012] Further preferably, the calculation relationship of the reward value of the Actor network is as follows:
[0013] reward t = reward1 + reward2
[0014]
[0015] where f t-1 and f t respectively represent the absolute values of the prediction errors at the (t - 1)-th time step and the t-th time step. The prediction error is the error between the predicted fault time and the actual fault time. Δt is the time interval, and k1, k2, and k3 are the coefficients of the corresponding reward function terms respectively. reward1 measures the prediction error magnitude of the model at each time step, reward2 measures the change in the prediction errors between adjacent time steps, and β is the threshold determining the sign of the reward value.
[0016] Further preferably, the condition for terminating the training of the fault prediction model is:
[0017]
[0018] K’ is the updated early stopping threshold, step is the update step size, k is the current error threshold, t k is the interaction round counter, used to record the number of rounds that have been interacted under the error threshold k, interaction max is the preset maximum interaction round threshold, n k is the number of interaction rounds completed under the error threshold k, and n is the number of interaction rounds when traversing the entire dataset.
[0019] Further preferably, the loss function of the fault prediction model is as follows:
[0020] The loss function of the Actor network is as follows:
[0021] The loss function of the Critic network is as follows:
[0022]
[0023] The loss function of the entropy regularization term coefficient α in the fault prediction model is as follows:
[0024]
[0025] where S t , a t, reward t are the state, action, and reward values at time step t, respectively, and S t+1 , a t+1 are the state and action at time step t + 1, respectively, R is the experience pool, ∈ t is the noise, is the Gaussian distribution, α is the entropy regularization term coefficient, γ is the discount factor, and f θ (∈ t ; S t ) is the sampled action under the condition of state S t , that is, the noise ∈ t , and π θ is the policy network with parameter θ, is the Critic network with parameter ω j , is the Critic Targte network with parameter , and a t+1 ~π θ (·|S t+1 ) is the action sampled according to the policy π θ at time step t + 1, and Q ω (S t , a t ) represents the output value of the critic network when the state is S t and the action is a t . is the target entropy size, and V is the target value when updating the critic network.
[0026] Further preferably, the calculation relation for predicting the fault time is as follows:
[0027]
[0028] where is the predicted fault time at time step t, is the predicted fault time at time step t - 1, and a t is the change amount of the predicted fault time at time step t.
[0029] According to another aspect of the present invention, there is provided a method for fault prediction using the above-mentioned fault prediction model, characterized in that the monitoring data of mechanical equipment is preprocessed, and an autoencoder is used to extract the low-dimensional key features of the preprocessed data to construct the state variable of the current time step. The state variable of the current time step is input into the above-mentioned fault prediction model to output the predicted fault time change amount, and the predicted fault time of the current time step is calculated using the predicted fault time change amount, thereby realizing fault prediction.
[0030] Further preferably, the preprocessing is to cluster the monitoring data of the entire life cycle of the mechanical equipment according to the working conditions, and then normalize the clustered data.
[0031] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the beneficial effects are as follows:
[0032] 1. The present invention transforms the degradation process of the equipment into an interaction process of reinforcement learning. At each time step, the Actor network outputs the predicted fault time change amount. The predicted fault time value of the previous time step minus this output gives the predicted fault time value of the current moment. There is an explicit constraint relationship between the predicted fault time values of adjacent moments, solving the problems of large fluctuations in the predicted values and low prediction accuracy caused by the insufficient utilization of the sensor monitoring data and the temporal correlation of the predicted fault time values.
[0033] 2. The present invention uses the SAC algorithm to encourage the agent to explore a wider action space by maximizing the action entropy, which is beneficial to improving the prediction accuracy of the model. And the SAC algorithm is not very sensitive to hyperparameters, and the model is easier to debug and train.
[0034] 3. The present invention constructs a multi-dimensional fault time feature containing the predicted values of multiple historical time steps, uses an autoencoder to extract key features from the sensor monitoring data, and combines the multi-dimensional fault time feature with the key features to obtain the state of reinforcement learning, providing richer historical fault time information.
[0035] 4. The reward function of the present invention includes an extended term of error gradient reward, which helps the model more effectively capture the changing trend of the prediction error, and then optimizes the prediction performance.
[0036] 5. The present invention proposes a training method for the reinforcement learning model of progressive early stopping. As the model is trained, the interaction stop condition is gradually shrunk, and the agent is gradually guided to learn the degradation trend information of the equipment. This training method can more precisely control the model performance and improve the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a method for predicting mechanical equipment faults based on reinforcement learning constructed according to the preferred embodiment of the present invention;
[0038] Figure 2 It is a schematic structural diagram of a fault prediction model constructed according to the preferred embodiment of the present invention;
[0039] Figure 3 It is a schematic diagram of the so-called fault time feature construction constructed according to the preferred embodiment of the present invention;
[0040] Figure 4 It is a schematic flow diagram of a progressive early stopping training method constructed according to the preferred embodiment of the present invention;
[0041] Figure 5 It is a visualization diagram of the true value and predicted value on a specific example constructed according to the preferred embodiment of the present invention; among them, (a) is the visualization result on the FD001 sub-dataset, (b) is the visualization result on the FD002 sub-dataset, (c) is the visualization result on the FD003 sub-dataset, and (d) is the visualization result on the FD004 sub-dataset;
[0042] Figure 6 It is a visualization diagram of the full life cycle true value and predicted value on an engine unit example close to complete failure according to the present invention; among them, (a) is the visualization result of the 34th engine on the FD001 sub-dataset, (b) is the visualization result of the 100th engine on the FD002 sub-dataset, (c) is the visualization result of the 79th engine on the FD003 sub-dataset, and (d) is the visualization result of the 2nd engine on the FD004 sub-dataset. Specific Embodiments
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0044] Figure 1 It is the flow of a mechanical equipment fault prediction method based on reinforcement learning constructed according to the preferred example of the present invention. This method includes the following steps:
[0045] S1. Use multiple sensors to collect the full life cycle monitoring data of mechanical equipment; according to the research purpose, the full life cycle monitoring data of multiple mechanical equipment can be collected as the training data set and the test data set; the monitoring data of the training set is used for model training, and the test data is used to evaluate the model performance;
[0046] Preprocess the collected monitoring data. The preprocessing process includes sensor screening, clustering of monitoring data according to working conditions, and normalization of monitoring data. Assume that N sensors are used for data collection, and the collected data contains monitoring information at T moments. The original monitoring data is X all =[[X1 X2 … X N . Screen out the sensor data with an obvious degradation trend during the equipment degradation process, and eliminate the sensor data that remains unchanged or has a small correlation degree during the degradation process. Finally, retain N F sensor monitoring data where N F is less than N. The operation of the equipment under different working conditions will also cause changes in sensor data. To eliminate the influence of working conditions on monitoring data, cluster the monitoring data according to working conditions. Assume that it is clustered into k categories, representing k working conditions respectively. Since the information monitored by different sensors is different, there will be a large difference in the absolute value of the data. Normalize the data of N F sensors according to working conditions. Among them, and respectively represent the minimum and maximum values of the monitoring data of the jth sensor under the kth working condition, and the normalized monitoring data is calculated.
[0047] S2. Construct a fault prediction model
[0048] (1) Autoencoder
[0049] The autoencoder consists of an encoder and a decoder with symmetric structures, both of which are one-dimensional convolutional neural networks. The encoder is used to extract key features and reduce the dimension of high-dimensional data within the window, and the decoder uses the output of the encoder to reconstruct the signal. Train the autoencoder on the entire training dataset in an unsupervised manner.
[0050] Train the autoencoder using monitoring data. The autoencoder is an unsupervised neural network that attempts to make the input and output as similar as possible. A typical autoencoder consists of two parts: an encoder and a decoder. The original input signal x passes through the encoder to obtain the low-dimensional feature h (i.e., encoding), and h is used as the input of the decoder, and the decoder outputs the reconstructed signal This process can be expressed as:
[0051] h = f encoder (x)
[0052]
[0053] Here, f encoder and f decoder respectively represent the mapping functions of the encoder and the decoder, usually parameterized neural networks. During the training process of the autoencoder, the loss function is usually defined as the reconstruction error:
[0054]
[0055] After training is completed, use the encoder to extract the key features e from the original high-dimensional data t . In this way, on the one hand, the data dimension can be effectively reduced, and on the other hand, the key information in the equipment degradation process can be retained.
[0056] The trained encoder inputs the sensor data within the time window and outputs the low-dimensional key feature e t . Use the predicted values of the first three time steps to construct the multi-dimensional fault time point feature r t , use the low-dimensional key feature e t and the multi-dimensional fault time point feature r t to perform feature concatenation to construct the state variable s of RL t . Compared with the state variable s in the prior art t which only uses the fault time point of the previous time step. On the one hand, the multi-dimensional fault time point feature r t can increase the weight of the fault time point information in the state and effectively avoid the drowning of the fault time point information. On the other hand, using the historical predicted values of multiple time steps can help the agent better capture the dynamic features of equipment degradation, including the equipment degradation speed, trend, etc., which can provide better support for the agent's decision-making, reduce prediction errors and uncertainties.
[0057] (2) SAC reinforcement learning algorithm
[0058] In reinforcement learning, the action of the agent at the current time step affects the state of the next time step. Therefore, when using reinforcement learning to solve the fault prediction problem, the state must contain the predicted value information of the historical time steps. The present invention proposes to construct a multi-dimensional fault time feature, and use the predicted value of the previous time step and the historical fault time predicted value j = t - 2, t - 3..., to construct the fault time feature r t ,
[0059] where Dim represents the dimension size of the constructed multi-dimensional fault time state r t . On the one hand, the multi-dimensional fault time feature r t can increase the weight of the fault time information in the state and effectively avoid the drowning of the fault time information. On the other hand, using the historical predicted values of multiple time steps can help the agent better capture the dynamic features of equipment degradation, including the equipment degradation speed, trend, etc., which can provide better support for the agent's decision-making, reduce prediction errors and uncertainties.
[0060] Finally, the state variable s of reinforcement learningt Composed of the low-dimensional key feature e t and the multi-dimensional fault time feature r t through feature concatenation, s t ={e t , r t}.
[0061] The Actor network is composed of multiple layers of fully connected networks. The input of the Actor network is the state variable of RL, and the output action is a one-dimensional value, representing the change amount a t of the predicted value. This action acts on the predicted value of the previous time step to obtain the predicted value at the current moment. Among them, the RL algorithm based on SAC used in this method adopts the reparameterization technique when the action a t takes values, where f represents the Actor network, and the two numerical values of its output action represent the mean and variance respectively. A Gaussian distribution of the action is constructed according to the mean and variance, and the specific numerical value of the action is sampled from the distribution. Compared with the existing technology that uses a deterministic policy network such as DQN or TD3, the SAC algorithm used in the present invention can provide more randomness for the action. At the same time, this can also reduce the sensitivity of the model to hyperparameters and make the model easier to debug.
[0062] In the MGP-SAC model proposed by the present invention, the action space is a continuous interval A, a t varies within the interval A. Among them the design of the action space having negative numerical values is convenient for the agent to adjust the action in a timely manner in the subsequent decision-making process after making a wrong decision, so as to reduce the prediction error. At the same time, the asymmetric design of the action interval is reasonable, and the degradation process of the device is generally irreversible. The agent executes the action a t acting on the predicted value of the previous moment to obtain the predicted value at the current time step The fault time feature at the (t + 1)-th time step is r t+1 ,
[0063] According to the error between the predicted value and the true value and the change situation of the prediction error between adjacent time steps, a reward value reward t is given to the agent. Then the time window slides one time step along the time axis, and the encoder extracts the low-dimensional key feature e t+1 of the sensor data within the new time window. The state s t+1 at the next time step = {e t+1 , r t+1}, determine whether the current round of interaction ends according to whether the error between the RUL predicted value and the true value exceeds the given threshold k. If it ends, done t = True, otherwise done t = False. Finally, store the interaction experience {s t , a t , reward t , s t+1 , done t} into the experience pool.
[0064] The reward value consists of two parts: the error between the predicted value and the true value at the current time step, and the difference between the predicted errors at the current time step and the previous time step. Let the reward value at the t-th time step be reward t , reward t 's calculation expression is as follows:
[0065] reward t = reward1 + reward2
[0066]
[0067] In the formula, f t-1 , f t respectively represent the absolute values of the prediction errors at the (t - 1)-th time step and the t-th time step. There is a one-unit time step interval between them, and Δt = 1. k1, k2, and k3 respectively represent the coefficients of the corresponding reward function terms. This design can inject richer gradient information into the model, enabling the model to capture the data change trend, rather than just the error value at a single time step. The reward function considering the error gradient can better reflect the performance of the model at different time points, improving the robustness and generalization ability of the prediction. The agent can better adjust its own strategy and perform more accurate actions for different states.
[0068] Calculate the reward value reward of the current interaction according to the prediction error and the change amount of the error at the current time step t , and store the interaction experience at the current time step into the experience pool. The time window slides along the time axis, interacts over the entire life cycle of the device, obtains the predicted values at each time step, and stores the interaction experience. After a round of interaction ends, extract the experience from the experience pool to update the Actor network, Critic network, Critic Target network, and the entropy regularization term coefficient α in the RL model. After the update is completed, use the updated Actor network for the next round of interaction. After completing multiple rounds of interaction and updating the network parameters of the RL model, obtain the final Actor network for device fault prediction.
[0069] The Critic network is also a multi-layer fully connected network. The input of the Critic network is the state s in RL t and the action a t . The network structure of the Critic Target is exactly the same as that of the Critic network. The input of the Critic Target network is the state s at the next time step t+1 and the action a t+1 . The coefficient α of the entropy regularization term is also a quantity with gradient that can be updated. When the optimal action is uncertain, the value of the entropy regularization term coefficient α should be relatively large. In the state where the optimal action is relatively certain, the entropy regularization term coefficient α should be relatively small. After one round of interaction on the training dataset, experiences are sampled from the experience pool to update the parameters of the Actor network, the Critic network, and the Critic network. In the next round of interaction, the agent uses the updated Actor network to interact with the constructed environment
[0070] After each round of interaction, S4 samples experiences from the experience pool and updates the network parameters of the model. Specifically, it includes the updates of the Actor network, the Critic network, the Critic Target network, and the entropy regularization term coefficient α. In the SAC algorithm used in the present invention, two action-value functions (parameters are ω1 and ω2 respectively) and a policy function π (parameter is θ) are modeled. Based on the idea of Double DQN, SAC uses two Critic networks, and alleviates the problem of overestimation of Q-values by selecting the network with a smaller output value Q of the Critic network. The loss function of the action-value function Q of the Critic network is
[0071]
[0072] In the formula, R is the data collected by the policy in the past time steps. To make the training more stable, two target Q networks are used corresponding to the two Q networks respectively. For the environment with a continuous action space, the policy output of the SAC algorithm uses the reparameterization trick
[0073] The loss function of the Actor network for rewriting the policy is
[0074]
[0075] In the SAC algorithm, the entropy regularization term coefficient α is very important. In a certain state where the optimal action is uncertain, the value of the entropy should be relatively large; in the case where the optimal action is relatively certain, the value of the entropy should be relatively small. During the training process, the regularization term coefficient α is also set as a variable with gradient that can be updated. Set the target value of the entropy to The loss function of α
[0076] In the current reinforcement learning technology, the interaction termination conditions used are generally fixed, that is, when the effect of the model or the number of actions performed by the agent reaches the threshold, the current round of interaction ends.
[0077] Based on this, this paper proposes a training method of progressive early stopping. At the initial stage of model training, a relatively large early stopping error threshold k is set, allowing a large error between the predicted value and the true value, enhancing the exploration of the model in the initial stage of training, and avoiding it falling into local optima. During the model training process, when the agent can traverse the entire dataset and the prediction error does not exceed the early stopping error threshold k, the early stopping threshold is reduced by a step size step. If the model cannot complete the interaction on the entire dataset under the current error threshold within a continuously preset number of interaction rounds, it is considered that the model performance cannot be further improved, and the training process terminates. Assume the current error threshold k, the interaction round counter t k is used to record the number of rounds that have been interacted under the error threshold k, and the preset maximum interaction round threshold interaction max . When each round of interaction terminates, there are the following situations.
[0078] 1) The model completes the interaction on the entire dataset:
[0079] k′ = k - step, t k′ = 0.
[0080] 2) The model fails to complete the interaction on the entire dataset and t k < interaction max :
[0081] t k = t k + 1.
[0082] 3) The model fails to complete the interaction on the entire dataset and t k ≥ interaction max , the model training terminates. Compared with the existing technology, the design of the dynamically changing stop threshold in the present invention can both avoid the problem that when the early stopping threshold is set too small, it is difficult for the model to complete the interaction on the entire dataset, resulting in insufficient model training. And it can dynamically capture the change of the model prediction accuracy and continuously shrink the early stopping threshold, which is beneficial to improving the prediction accuracy of the model.
[0083] Taking an aero-engine, a mechanical device, as a specific object, the method for predicting mechanical device faults based on reinforcement learning of the present invention is further described in detail with specific examples. The specific steps are as follows:
[0084] (1) The C-MAPSS dataset is used for training and testing. The C-MAPSS dataset consists of four sub-datasets (FD001, FD002, FD003, and FD004), containing monitoring data collected by 21 sensors, with different failure modes and operating conditions. Sensor data that is constant throughout the engine life cycle or only exhibits weak changes is excluded. Finally, sensor data numbered 2, 3, 4, 7, 8, 9, 11, 12, 13, 14, 15, 17, 20, and 21 is retained. Then, the retained sensor data is clustered by operating conditions and normalized for offline training of the model. The test set only contains sensor data before the occurrence of a failure and is used for online fault prediction. The present invention comprehensively evaluates all four sub-datasets.
[0085] Table 1 C-MAPSS Dataset
[0086]
[0087] (2) In the stage of training the autoencoder, first, time window processing is performed on the monitoring data. The selected time window dimension is size×14, where size represents the sampling data of consecutive size time instants, and 14 represents the number of sensors selected. The size values of the FD001, FD002, FD003, and FD004 sub-datasets are set to 30, 20, 30, and 15 respectively. The specific network structure of the encoder is shown in Table 2.
[0088] (3) Use the trained encoder to extract the low-dimensional key features e of the sensor data within the time window of each time step t , and construct the multi-dimensional fault time feature r using the predicted values of the first three time steps t . Use the low-dimensional key feature e t and the multi-dimensional fault time feature r t to perform feature concatenation to construct the state s of each time step t . During the training process, the state s t is input into the Actor network to obtain the action value a of each time step t , and a t acts on the predicted value of the previous time instant to obtain the predicted value of the current time instant, and then the reward value is calculated according to the reward function. After each interaction ends, the interaction experience is stored in the experience pool, and it is judged whether this round of interaction ends according to whether the early stop threshold is triggered. After each round of interaction ends, samples are taken from the experience pool to update the network parameters of the model. Specifically, the structural parameters of the Actor network, Critic network, and Critic Target network used in the model are shown in Table 3. The hyperparameters used in the model are shown in Table 4.
[0089] Table 2 Encoder Network Structure
[0090]
[0091] (4) Visualize the prediction results. Taking the test data of FD001, FD002, FD003, and FD004 as an example, using the fault prediction model based on reinforcement learning, the prediction results are as Figure 5 and 6 shown. It can be intuitively seen from Figure 5 and 6 that most of the predicted values are very close to the label values. For the engine test units approaching failure, this method has very good prediction results. One engine unit close to complete degradation is selected from each sub-test set to demonstrate the prediction performance during the entire degradation process. The predicted values remain unchanged or decrease slowly in the early stage of equipment operation. As the number of working cycles increases, the predicted values show an obvious downward trend. The method proposed in the present invention can obtain good prediction effects during the entire equipment degradation process and generally maintains a monotonically smooth degradation trend, which is in good agreement with the actual degradation process of the equipment.
[0092] Table 3 Structural parameters of the Actor network, Critic network, and Critic Target network
[0093]
[0094] Table 4 Important hyperparameters in the model
[0095]
[0096]
[0097] To further verify the advantages of the model designed in this patent compared with the prior art, three baseline prediction models were selected for comparison, including representative deep learning methods: Deep Convolutional Neural Network (DCNN), Long Short-Term Memory (LSTM), and the existing reinforcement learning model DQN (Deep Q-learning), etc. The Root Mean Square Error (RMSE) and Score Function were used to analyze the experimental results, and their calculation methods are specifically described as follows:
[0098]
[0099] Among them, and y iThey represent the predicted value and the true value respectively. As shown in Table 5, Table 5 summarizes all the estimation results based on the root mean square error and the score.
[0100] It can be seen by comparison that both the RMSE index and the Score index of the proposed method on the four test sets are better than those of other methods, which further illustrates the superiority of the proposed method compared with the existing technologies.
[0101] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0102] Table 5 Estimation Results Based on Root Mean Square Error and Score
[0103]
Claims
1. A fault prediction model based on reinforcement learning, characterized in that: The fault prediction model adopts an Actor-Critic reinforcement learning network structure. The input of the Actor network in the fault prediction model is the state variable s of the current time step. t , the output is the predicted failure time change at the current time step, the s t ={e t , r t }, s t is the state variable at time step t, e t is the low-dimensional key feature at time step t, r t is the multidimensional failure time feature, is the predicted failure time at time step t-1, is the predicted failure time at time step t-2, is the predicted failure time at time step t-Dim; After the state variables corresponding to the current time step are input into the Actor network, the predicted fault time and action reward value are calculated using the output predicted fault time change, the state variables of the current time step, the predicted fault time change and the corresponding action reward value are stored as experience in the experience pool, and experience is extracted from the experience pool to update the fault prediction model, thereby realizing the training of the fault prediction model.
2. A mechanical equipment fault prediction method based on reinforcement learning as described in claim 1 or 2, characterized in that: The low-dimensional key feature e t It is obtained by extracting it using an autoencoder, which includes an encoder and a decoder with a symmetrical structure, both of which are convolutional neural networks. The encoder is used to extract features from the preprocessed monitoring data to obtain low-dimensional key features, and the decoder reconstructs the signal using the output of the encoder. The autoencoder is trained in an unsupervised manner so that the reconstructed signal is close to the input preprocessed monitoring data, thereby obtaining the trained autoencoder, which is used for low-dimensional key features. t Extraction.
3. A mechanical equipment fault prediction method based on reinforcement learning as claimed in claim 2, characterized in that: The loss function of the autoencoder is as follows: Where n is the number of time steps, x i and are the original signal and the reconstructed signal at the i-th time step, is the mean square error loss function.
4. A mechanical equipment fault prediction method based on reinforcement learning as claimed in claim 1, characterized in that: The Actor-Critic reinforcement learning network structure includes an Actor network, a Critic network and a Critic Target network.
5. A mechanical equipment fault prediction method based on reinforcement learning as claimed in claim 1 or 2, characterized in that: The calculation formula of the reward value of the Actor network is as follows: reward t =reward1+reward2 Among them, f t-1 、f t They represent the absolute values of the prediction errors at the t-1th time step and the tth time step respectively. The prediction error is the error between the predicted failure time and the actual failure time. Δt is the time interval. k1, k2 and k3 are the coefficients of the corresponding reward function terms respectively. reward1 is a measure of the prediction error of the model at each time step. reward2 is a measure of the change in the prediction errors of adjacent time steps. β is the threshold that determines the positive or negative reward value.
6. A mechanical equipment fault prediction method based on reinforcement learning as claimed in claim 1 or 2, characterized in that: The loss function of the fault prediction model is as follows: The loss function of the Actor network is as follows: The loss function of the network is as follows: The loss function of the entropy regularization term coefficient α in the fault prediction model is as follows: Among them, E is the expected value of calculation, S t ,a t, reward t are the state, action and reward value of time step t, S t+1 ,a t+1 are the state and action at time step t+1 respectively, R is the experience pool, ∈ t It is noise. is a Gaussian distribution, α is the entropy regularization term coefficient, γ is the discount factor, and f θ (∈ t ; S t ) is state S t That is, noise∈ t Sampling action under condition, π θ is a policy network with parameters θ, The parameter is ω j Critic network, The parameter is Critic Target network, j is the jth Critic or Critic Target network, According to the strategy π θ The action sampled at time step t+1, Q β (S t ,a t ) indicates that in state S t , action is a t When is the output value of the critic network, is the target entropy size, and V is the target value when updating the critic network.
7. A mechanical equipment fault prediction method based on reinforcement learning as claimed in claim 1 or 2, characterized in that: The conditions for terminating the fault prediction model training are: K' is the updated early stopping threshold, step is the update step length k is the current error threshold, t k is the interaction round counter, which is used to record the number of rounds of interaction under the error threshold k. max is the preset maximum interaction round threshold, n k is the number of interaction rounds completed under the error threshold k, and n is the number of interaction rounds when traversing the entire data set.
8. A mechanical equipment fault prediction method based on reinforcement learning as claimed in claim 1, characterized in that: The calculation relationship of the predicted failure time is as follows: in, is the predicted failure time at time step t, is the predicted failure time at time step t-1, a t is the change in predicted failure time at time step t.
9. A method for fault prediction using the fault prediction model according to any one of claims 1 to 8, characterized in that: The monitoring data of the mechanical equipment is preprocessed, and the low-dimensional key features of the preprocessed data are extracted using an autoencoder to construct the state variables of the current time step. The state variables of the current time step are input into the fault prediction model described in any one of claims 1-8 to output the predicted fault time change, and the predicted fault time change is used to calculate the predicted fault time of the current time step, thereby realizing fault prediction.
10. The fault prediction method according to claim 9, characterized in that: The preprocessing is to cluster the collected monitoring data of the entire life cycle of the mechanical equipment according to the working conditions, and then normalize the clustered data.