Equipment prediction maintenance framework based on probability residual life

Through the maintenance framework based on Bayesian neural network probability prediction and reinforcement learning, the uncertainty problem of equipment life prediction in the prior art is solved, and more accurate maintenance decisions and equipment reliability are achieved.

CN120197498APending Publication Date: 2025-06-24BEIHANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510336307.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing equipment maintenance strategies cannot accurately reflect the probability distribution characteristics of equipment life, resulting in blindness and uncertainty in maintenance decisions, increasing maintenance costs and potential safety hazards.

Method used

The Bayesian neural network-based method is used to predict the probability of the remaining life of the device, and a predictive maintenance framework is constructed in combination with reinforcement learning, and maintenance decisions are dynamically optimized through real-time data.

Benefits of technology

The probability distribution prediction of the remaining life of the equipment is realized, providing a more accurate basis for maintenance decision-making, reducing maintenance costs and downtime, and improving the overall reliability of the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197498A_ABST
    Figure CN120197498A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment prediction maintenance framework based on probability residual life, and belongs to the technical field of fault prediction and health management (PHM). According to the framework, through combining a Bayesian neural network and a reinforcement learning technology, probability prediction and dynamic maintenance decision optimization of the residual life of equipment are realized. The method specifically comprises the following steps: acquiring sensor data in equipment operation, and preprocessing to generate a training data set; constructing a Bayesian neural network (BNN), utilizing variation reasoning to approximate posteriori distribution, and outputting probability distribution of residual life through Monte Carlo sampling; based on a probability prediction result, constructing a reinforcement learning environment model, and defining a state space containing residual life distribution, a spare part state and a maintenance action and a reward function; a Double DQN algorithm is adopted to optimize a maintenance strategy, and intelligent decision-making of the optimal maintenance time and the optimal ordering time of equipment is dynamically realized through an epsilon-greedy algorithm, so that the maintenance cost and the fault risk are minimized. According to the method, probability residual life distribution and reinforcement learning decision are innovatively combined, the problem that a traditional point estimation model ignores uncertainty is solved, an intelligent maintenance framework is constructed through a reinforcement learning method, and dynamic optimization of maintenance decision is achieved. The NASA aero-engine data set verification shows that compared with a traditional method, the optimization effects of prediction errors, uncertainty quantification, the maintenance cost rate and the like are remarkable, and the reliability and economical efficiency of equipment are effectively balanced.
Need to check novelty before this filing date? Find Prior Art

Description

(1) Technical Field

[0001] The present invention provides a device predictive maintenance framework based on probabilistic remaining useful life, which is used to predict the remaining life distribution of a device and formulate an optimized maintenance strategy, belonging to the technical field of prognostics and health management (PHM). (2) Background Art

[0002] In modern industry, the reliable operation of devices is crucial for production efficiency, product quality, and safety. Unexpected failures of devices may lead to production interruptions, economic losses, and even endanger the safety of personnel. Therefore, how to effectively predict the remaining useful life (RUL) of devices and formulate reasonable maintenance strategies has become a research hotspot in the industrial and academic fields. Traditional device maintenance strategies mainly include corrective maintenance and preventive maintenance at regular intervals. Corrective maintenance means that the device is repaired after a failure occurs. Although it is simple and direct, it may lead to unexpected downtime, resulting in high downtime costs and potential safety hazards. Preventive maintenance is based on statistical data and experience, and the device is subjected to preventive inspections and replacements at preset time intervals. However, to avoid failures, premature and frequent maintenance is usually adopted, which increases unnecessary maintenance costs. Moreover, due to the complexity and randomness of the device deterioration process, it cannot accurately reflect the health status of the device, and may miss key failure omens, leading to blindness in maintenance decisions.

[0003] With the rapid development of sensor technology, device maintenance strategies are gradually shifting towards intelligence and digitization. Predictive maintenance (PdM), as an emerging maintenance strategy, uses real-time monitoring data, historical operation data, and advanced analysis techniques to evaluate and predict the health status of a device, so as to perform maintenance at the optimal time point. Its core lies in accurately predicting the RUL of the device and formulating an intelligent maintenance strategy, thereby realizing the optimized scheduling of maintenance activities, minimizing maintenance costs and downtime to the greatest extent, and improving the overall reliability of the device. Current predictive maintenance methods mainly rely on point-estimation prediction models of RUL, which output a definite remaining life value for making maintenance decisions. However, in the actual operation process of a device, it will be affected by various uncertain factors, such as changes in environmental conditions, load fluctuations, and material batch differences. These factors lead to significant uncertainty in RUL prediction. The RUL prediction model based on point estimation only provides a definite remaining life value and cannot comprehensively reflect the probabilistic distribution characteristics of the device life. At the same time, the introduction of uncertainty also increases the decision-making difficulty. Traditional decision-making methods are difficult to match complex maintenance scenarios. This limitation makes it difficult for maintenance decisions to accurately evaluate potential risks, reducing the reliability and optimization effect of maintenance strategies.

[0004] Therefore, this paper proposes a device predictive maintenance framework based on probabilistic remaining useful life, predicts the remaining life distribution of the device based on Bayesian neural network, and constructs a predictive maintenance framework in combination with reinforcement learning to dynamically optimize maintenance decisions according to real-time data, which can provide a strategic basis for device health management. (III) Summary of the Invention

[0005] The present invention is a device predictive maintenance framework based on probabilistic remaining useful life, which conducts probabilistic prediction of the device RUL based on Bayesian neural network and stochastic process, and makes maintenance decisions on real-time data through reinforcement learning. The method flow of the present invention is as Figure 1 shown, specifically including the following:

[0006] Step 1: Collect sensor data generated during the operation of the device, and reverse-infer the remaining service life label corresponding to each moment according to the data from operation to failure. Clean, denoise, normalize and extract features from the sensor data to form a data set for model training and prediction;

[0007] Step 2: Use the time-series sensor data and life label data [X, Y] as training samples to construct a Bayesian neural network (BNN);

[0008] Step 3: Adopt a variational inference method to approximately infer the actual posterior distribution, construct a variational distribution q φ (ω), and solve the Bayesian neural network prediction model;

[0009] Step 4: Considering the random uncertainty in prediction, set the output to a normal distribution, adjust the network hyperparameters and train the neural network model, and use the Monte Carlo sampling method to output the predicted probabilistic remaining useful life result;

[0010] Step 5: Based on the probabilistic remaining useful life prediction result, construct an environment model of the device predictive maintenance framework based on reinforcement learning, mainly including state, action and reward.

[0011] Step 6: Considering that the RUL degradation is a long-term time-series process, adopt the Double DQN (Double Deep Q-Network) algorithm to make maintenance decision judgments, and update the model by updating the target network.

[0012] Step 7: Select an appropriate maintenance action a t according to the ε-greedy algorithm, and calculate the corresponding cost and reward function to realize the intelligent maintenance decision of the degraded device.

[0013] Among them, in step 2, the long short-term memory neural network (LSTM) is selected as the time series prediction model. The uncertainty quantification model in the network mainly consists of the prior distribution p(ω) in the parameter space and the likelihood function of Bayesian regression and is usually characterized by the Gaussian distribution l(y (i) ∣f ω (x (i) ))). Among them, the model parameter ω is independent of the feature sample X. According to Bayes' theorem, the posterior distribution of the model parameter is:

[0014]

[0015] The Bayesian neural network prediction model can be expressed as:

[0016] p(y∣r,X,Y)=∫l(y∣f π (r))p(ω∣X,Y)dω (2)

[0017] In step 3, construct the probability distribution family q φ (ω) as the variational distribution family, and minimize the Kullback-Leibler (KL) divergence between the probability distribution q φ (ω) and p(ω∣X,Y) with respect to the parameter φ to find the optimal variational distribution. The KL divergence is defined as:

[0018]

[0019] Among them, an estimator of the KL divergence can be expressed as:

[0020]

[0021] In the formula, can be calculated by the following formula:

[0022]

[0023] The can be minimized using a stochastic optimizer * to obtain the optimal φ because φ is an estimator of KL(q * (ω)‖p(ω∣X,Y)), and φ φ is also optimal for KL(q (ω)‖p(ω∣X,Y)). Use

[0024] to approximate p(ω∣X,Y), and then complete the solution of the prediction model. In step 4, construct the damage function of the Bayesian neural network based on the variational distribution

[0025]

[0026] In the formula, λ is the model decay coefficient.

[0027] The remaining useful life prediction model is sampled N times by the Monte Carlo method, and the sampling results can be expressed as:

[0028]

[0029] The predicted mean value of RUL under N samplings can be expressed as:

[0030]

[0031] The variance of the predicted RUL under N samplings can be expressed as:

[0032]

[0033] That is, the distribution characteristics of the predicted RUL are y ~ N(μ, σ 2 ).

[0034] In step 5, for the current moment t, the state S t consists of the following elements:

[0035] S t = [RUL(t), t, P t , O t , T space (10)

[0036] where RUL(t) is the remaining useful life prediction distribution at the current moment, following the distribution (μ, σ 2 ); P t is the spare part status function, 1 indicates that the spare part has arrived, and 0 indicates that it has not arrived; O t is the order status function, 1 indicates that an order for the spare part has been placed, and 0 indicates that no order has been placed; T space is the arrival time of the spare part, set to ∞ if no order has been placed.

[0037] At any moment, the available action a t is defined as follows:

[0038] a t = 0: Perform preventive maintenance;

[0039] a t = 1: No operation;

[0040] a t = 2: Order spare parts.

[0041] For a given state-action pair (S t , a t ), the transferred state S t+1 is obtained, and the immediate reward r t is calculated. The reward is calculated based on the negative or reciprocal of the actual cost to minimize the total cost. At the same time, the reward function is adaptively increased or decreased for certain expected or unexpected events. Specifically:

[0042] (1) When a t = 0 and the spare part has arrived, the reward function is:

[0043] r t = -[C p + C l × max(0, T fail - t) + C i × max(0, t - T space )] (11)

[0044] Where C p is the preventive maintenance cost, C l is the waste cost, C i is the storage cost, T fail is the actual failure time, and T space is the spare part arrival time.

[0045] (2) When a t = 1 and it finally leads to a failure (t ≥ T fail ), the reward function is:

[0046] If the spare part has arrived, the corrective maintenance cost C c and the storage cost C i are incurred:

[0047] r t = -[C c + C i × max(0, t - T space )] (12)

[0048] If the spare part is not in time for arrival, the corrective maintenance cost and the downtime and out-of-stock cost are incurred:

[0049] r t = -[C c + C os × max(0, T space - t)] (13)

[0050] When a t = 1 and no failure occurs, a small negative reward is given to encourage timely decision-making.

[0051] (3) When a t = 2, if an order is not placed, a reward is generated, and if an order has been placed and is ordered again, a penalty is generated:

[0052]

[0053] In step 6, the Q-value function in Double DQN is based on the Bellman optimality equation, representing the long-term reward obtained by taking a certain action in a given state. The Q-value function is iteratively updated by training a neural network to approximate the optimal policy, that is:

[0054] Q(s t , a t ) = Q(s t , a t ) + α[r t + γmaxQ(s t+1 , a) - Q(s t , a t )] (15)

[0055] Among them, Q(s t , a t ) is the Q-value for (state s t , action a t ), α is the learning rate, r t is the reward at time t, γ is the discount factor, and maxQ(s t+1 , a) is the maximum Q-value estimate among all actions in the next state s t+1 .

[0056] Double DQN is an improvement over traditional DQN. By introducing a target network, it solves the problem of overestimation of Q-values in traditional DQN due to using the same network to calculate Q-values and select the maximum action. Specifically as follows:

[0057] y t = r t + γQ(s t+1 , a t+1 ; θ - ) (16)

[0058] Among them, y t is the target Q-value, used to update the network parameters online, θ - are the target network parameters, and a t+1 is the action that selects the maximum Q-value in the current state.

[0059] For the experience replay sample (s t , a t , r t , s t+1 , d t ), where dt Indicates whether the decision ends. Placing the experiences in a replay pool and randomly sampling from it to start training can break the correlation between samples. The loss function for training is chosen as SmoothL1 Loss, which is specifically defined as:

[0060]

[0061] In step 7, the exploration - exploitation balance in reinforcement learning is achieved through the ε - greedy algorithm, and the specific implementation is as follows: For the given state space S and action A, the goal of the algorithm is to find an optimal policy π starting from the initial state s0 * , which selects the optimal action a t at each state s t , and maximizes the long - term return starting from state s t :

[0062]

[0063] In the Double DQN algorithm, the basic process of the ε - greedy algorithm is as follows:

[0064] Exploration: With probability ε, randomly select an action a ∈ A, which may not be the current optimal action;

[0065] Exploitation: With probability 1 - ε, select the action with the highest current Q - value

[0066] Meanwhile, as the training progresses, continuously decay the value of ε:

[0067] ε new = ε old × ε f (19)

[0068] where ε f is the decay coefficient, indicating that as the training progresses, the agent increasingly relies on the learned policy rather than exploring new methods.

[0069] Through the probability remaining useful life prediction distribution of the input device, the device operating state is continuously updated to obtain the real - time optimal maintenance time and optimal ordering time. Combining with the actual failure time of the device, calculate the total maintenance cost. Compared with the traditional regular maintenance and corrective maintenance methods, it reflects the effectiveness of intelligent maintenance decision - making.

[0070] The present invention is a device predictive maintenance framework based on probability remaining useful life, and its advantages and effects are as follows:

[0071] 1. The present invention uses a method based on Bayesian neural network to consider the influence of uncertainty, obtains the probability RUL prediction distribution, and provides sample input for maintenance decision - making;

[0072] 2. The present invention adopts a method based on reinforcement learning to construct a predictive maintenance framework. By defining the device state, maintenance actions, and reward function, it realizes intelligent decision-making on the optimal maintenance time and optimal ordering time of the device, reduces maintenance costs, and improves the operation efficiency and reliability of the device. (IV) Description of the Drawings

[0073] Figure 1 It is the flowchart of the method described in the present invention;

[0074] Figure 2 It is the RUL prediction result of the M2 method;

[0075] Figure 3 It is a comparison diagram between the regular maintenance method and the predictive maintenance method.

[0076] The labels and symbols in the figure are explained as follows:

[0077] RUL represents the remaining useful life. (V) Specific Embodiments

[0078] The specific process of implementing the present invention uses the aero-engine dataset C-MAPSS provided by the NASA Ames Prognostics Center of Excellence. This dataset is a commonly used dataset for carrying out predictive maintenance. Select the operating data of 100 engines in the "FD001" sub-dataset. This dataset has collected 24 groups of sensor data (3 groups of operation data and 21 groups of sensor detection data), and recorded the complete multivariate time series data from operation to failure.

[0079] Five indicators are selected to measure the effects of RUL prediction and maintenance decision-making respectively. Among them, for RUL prediction, the root mean square error (RMSE) is used to evaluate the mean prediction performance, and the prediction interval average bandwidth (PINAW) and prediction interval coverage probability (PICP) are used to evaluate the interval prediction performance; for maintenance decision-making, the failure rate (FR) and average cost rate (ACR) are used as indicators. The calculation formulas of the six indicators are as follows:

[0080] (1) Root Mean Square Error (RMSE)

[0081]

[0082] (2) Coefficient of Determination (R 2 )

[0083]

[0084] (3) Prediction Interval Coverage Probability (PICP)

[0085]

[0086] (4) Prediction Interval Average Width (PINAW)

[0087]

[0088] (5) Failure Rate (FR)

[0089]

[0090] (6) Average Cost Rate (ACR)

[0091]

[0092] Wherein, RUL i is the true RUL, RUL max is the maximum value of the true RUL, RUL min is the minimum value of the true RUL, L k is the predicted RUL, U i is the predicted upper confidence bound, L i is the predicted lower confidence bound, C i is the maintenance cost of this group of engines, t i is the maintenance time of this group of engines.

[0093] In the experiment, FD001 (T1-T80) is used as the training set, and FD001 (T81-T100) is used as the validation set. The sensor data is used as the input feature, and each input feature is normalized by the maximum-minimum method. Since the degradation characteristics of the engine are not obvious in the initial stage of operation, the segmented RUL label is selected to replace the linear RUL label. For this dataset, the maximum value of RUL is set to 130 time cycles, and the RUL label is calculated as follows:

[0094]

[0095] Wherein, L is the maximum cycle life of the sample.

[0096] The present invention uses the long short-term memory neural network (LSTM NN) commonly used in time series prediction problems as the benchmark model to verify the effectiveness of RUL prediction, and the confidence level of interval prediction is taken as 95%. The traditional LSTM NN (M1) that does not consider uncertainty is compared with the proposed LSTM-BNN (M2), and the same dataset and parameter settings are used in both groups of experiments. Table 1 shows the parameter settings of the prediction model.

[0097] Table 1 Prediction model parameters

[0098]

[0099]

[0100] Figure 2 The prediction results of the data of Engine #83 in the test set based on the M2 method are shown. It can be seen that the prediction results of the proposed method are in good agreement with the actual RUL. At the same time, the comparison of the results of four indicators of RUL prediction by the two methods is given in Table 2. It can be seen that compared with the M1 method, the M2 method has a smaller RMSE and a larger R 2 and has the ability of interval prediction. The results of PICP and PINAW also meet the requirements of the confidence interval, which can provide data input for the subsequent predictive maintenance framework.

[0101] Table 2 RUL prediction results

[0102]

[0103] The obtained RUL prediction distribution is input into the reinforcement learning algorithm, and the Adam optimizer is used for optimization. The target network is updated through soft update, and the coefficient τ = 0.01 is used to achieve the smooth update of the target grid, thereby indirectly adjusting the learning rate. Comparing the traditional regular maintenance strategy (M3) with the predictive maintenance strategy (M4) proposed in this paper, the regular maintenance strategy is based on the historical mean time to failure, and its cost rate is expressed as:

[0104]

[0105] where t ture is the actual failure time of this group of engines, is the mean time to failure calculated based on historical data.

[0106] The parameters in the predictive maintenance framework are shown in Table 3:

[0107] Table 3 Predictive maintenance framework parameters

[0108]

[0109]

[0110] Both groups of experiments use the RUL dataset predicted by LSTM - BNN to predict the optimal maintenance time and the optimal order time, and at the same time calculate the failure incidence rate and the average cost rate using the real RUL data. Some decision results of the predictive maintenance framework are shown in Table 4, and the comparison with the regular maintenance strategy is as follows Figure 3As shown, it can be seen that through the predictive maintenance method, the actual probability of failure during operation has decreased significantly, indicating that this method can effectively prevent the occurrence of failures and reduce the probability of accidents. At the same time, the average cost rate of maintenance has also decreased significantly, reflecting the accuracy of the prediction, which can better balance the relationship between reliability and economy and reduce maintenance costs on the premise of ensuring reliability.

[0111] Table 4 Predictive Maintenance Decision Results (Partial)

[0112]

Claims

1. A predictive maintenance framework for equipment based on probabilistic remaining life, characterized in that: The following steps are involved: Step 1: Collect sensor data generated during the operation of the equipment, and infer the corresponding remaining service life label at each moment based on the data from operation to failure. Clean, denoise, normalize and extract features of the sensor data to form a data set for model training and prediction; Step 2: Use the time series sensor data and life label data [X, Y] as training samples to build a Bayesian Neural Network (BNN); Step 3: Use variational inference to approximate the actual posterior distribution and construct the variational distribution q φ (ω), solving the Bayesian neural network prediction model; Step 4: Considering the random uncertainty in the prediction, set the output to normal distribution, adjust the network hyperparameters and train the neural network model, and use the Monte Carlo sampling method to output the predicted probabilistic remaining life results; Step 5: Based on the probabilistic remaining life prediction results, build an equipment predictive maintenance framework environment model based on reinforcement learning, which mainly includes states, actions and rewards. Step 6: Considering that RUL degradation is a long-term temporal process, the Double DQN (Double Deep Q-Network) algorithm is used to make maintenance decision judgments, and the model is updated by updating the target network. Step 7: Select appropriate maintenance action a according to the ε-greedy algorithm t , and calculate the corresponding cost and reward function to realize intelligent maintenance decision-making for degraded equipment.

2. The equipment predictive maintenance framework based on probabilistic remaining life according to claim 1 is characterized by: In step 2, the construction of the Bayesian neural network includes the following sub-steps: (1) Using long short-term memory neural network (LSTM) as the time series prediction model; (2) Quantify model uncertainty through the prior distribution of parameter space and the likelihood function of Bayesian regression; (3) Based on the variational distribution family, the Kullback-Leibler (KL) divergence is minimized to solve the optimal variational distribution of the model parameters and construct the loss function of the Bayesian neural network; (4) The remaining life prediction model is sampled multiple times using the Monte Carlo method, and the mean and variance of the predicted RUL are output.

3. The equipment predictive maintenance framework based on probabilistic remaining life according to claim 1 is characterized by: In step 3, the elements of the state space include: the predicted distribution of remaining life at the current moment, the spare parts arrival status, the ordering status and the spare parts arrival time; the elements of the action space include performing preventive maintenance, no operation or ordering spare parts; the reward function is dynamically calculated based on the maintenance cost, downtime cost and spare parts related costs.

4. The equipment predictive maintenance framework based on probabilistic remaining life according to claim 1 is characterized by: In step 4, when the Double DQN algorithm is used, the Q value function is updated through the target network and the online network separation mechanism, the network parameters are optimized by combining the experience replay pool and the Smooth L1Loss function, and the exploration rate ε is attenuated as the training progresses, thereby gradually improving the strategy stability.

5. The equipment predictive maintenance framework based on probabilistic remaining life according to claim 1 is characterized by: In step 5, the calculation of the total maintenance cost includes preventive maintenance cost, corrective maintenance cost, storage cost and out-of-stock cost, and the effectiveness of the framework is verified by comparing the failure rate (FR) and average cost rate (ACR) of the predictive maintenance strategy with that of the traditional regular maintenance strategy.

6. The equipment predictive maintenance framework based on probabilistic remaining life according to claim 1 is characterized by: The implementation of the framework uses an aviation turbofan engine dataset, outputs probabilistic remaining life distribution through segmented RUL label processing, time series data normalization and LSTM-BNN model training, and combines reinforcement learning algorithm to achieve dynamic optimization of maintenance decisions. The maintenance decision framework supports the coordinated optimization of spare parts ordering and maintenance execution, dynamically adjusts maintenance actions according to spare parts arrival time and equipment degradation status, avoids excessive maintenance or maintenance delays, and reduces maintenance costs while ensuring reliable operation.

Citation Information

Cited By

  • Method and system for monitoring reliability of photovoltaic converter in plateau special environment

    CN120908559A

  • Aero-engine model Bayesian optimization method for quantizing uncertainty

    CN121031378A

  • Flight simulation equipment management and inventory optimization method and system

    CN121480328A

  • A method and system for managing and optimizing inventory of flight simulation equipment

    CN121480328B