A method for predicting uncertainty of remaining useful life based on multi-attention mechanism
By introducing multi-attention mechanisms and time convolution networks in the residual service life prediction of gas turbines, the problems of insufficient flexibility and error accumulation of existing models are solved, high-precision and efficient prediction are achieved, and uncertain quantification performance is provided.
Patent Information
- Application Number
- CN202210636893.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-07
AI Technical Summary
Existing deep learning models have problems of insufficient flexibility and error accumulation in predicting the remaining service life of gas turbines, and TCN has been studied less in this field.
A method of prediction of residual service life uncertainty based on multi-attention mechanism is proposed. The characteristic dimension and time step dimension are weighted through the self-attention mechanism, combined with a time convolution network with shared parameters, improve prediction efficiency, and use a non-parametric probability prediction framework to provide an estimated confidence interval for the remaining service life.
It realizes high-precision residual service life prediction, can adaptively extract feature information, improve prediction efficiency, and effectively demonstrate the performance of uncertainty quantization.
Smart Images

Figure CN115204463B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remaining useful life prediction, and in particular to a remaining useful life uncertainty prediction method based on a multi-attention mechanism. Background Art
[0002] Predicting the remaining useful life of an asset allows engineers to make more reasonable maintenance plans, reduce downtime, prevent catastrophic failures, and reduce costs and improve efficiency. Therefore, accurately predicting the remaining useful life is of great significance.
[0003] With the development of deep learning, in terms of predicting the remaining useful life, traditional deep learning models such as RNN do not have the ability to process in parallel. The prediction of the subsequent time must wait for the completion of the previous step, which reduces the flexibility of the model and causes the error to accumulate step by step; CNN is not suitable for modeling time series problems due to the limited size of its convolution kernel and cannot effectively capture long-term dependency information. In order to obtain a prediction method that can adaptively extract features and output remaining useful life estimates, the machine learning model should have the ability to learn useful information and end-to-end trainable parameters from training data. TCN uses a combination of residual networks and dilated convolutions to enhance network memory, making it more effective in long-term series prediction tasks. However, there is currently little research on TCNs in prediction. Therefore, simplifying the process of quantifying arbitrary uncertainties in remaining useful life prediction and exploring TCN networks in remaining useful life prediction have very important application value. Summary of the invention
[0004] The purpose of the present invention is to propose a method for predicting the uncertainty of the remaining useful life of a gas turbine based on a multi-attention mechanism to solve the problem of predicting the remaining useful life of a gas turbine.
[0005] The technical solution to achieve the purpose of the present invention is a method for predicting uncertainty of remaining useful life based on a multi-attention mechanism, comprising the following steps:
[0006] Step 1, obtain two-dimensional data of a gas turbine engine that gradually degrades from a good state to a scrapped state over time to construct a training set, and perform preprocessing, the preprocessed data is three-dimensional data including the degradation process, the time lag order, and the number of sensors;
[0007] Step 2: construct a TCN network based on a multi-attention mechanism, use the self-attention mechanism to weight the feature dimensions of the preprocessed three-dimensional data, and obtain the sequence data of the sensor with weighted feature information; use a temporal convolutional network with shared parameters to learn the sequence data with weighted feature information, and obtain the sequence data of the sensor with time and space correlation information; use the self-attention mechanism to weight the sequence data with time and space correlation information in the time dimension, and obtain the sequence data of the sensor with weighted time information; use the fully connected layer to predict the sequence data with weighted time information, and obtain the predicted remaining service life;
[0008] Step 3: Use the zero-start training method and set the quantile loss as the loss of model training to train the TCN network based on the multi-attention mechanism, and use the grid search technology to obtain the optimal parameters of the network;
[0009] Step 4: pre-process the gas turbine engine sensor data for which the remaining useful life is to be predicted, input the trained model, and complete the remaining useful life prediction.
[0010] Further, in step 1, two-dimensional data of the gas turbine engine gradually degrading from a good state to being scrapped over time is obtained and preprocessed. The specific steps are as follows:
[0011] Step 1.1: Draw a trend graph of the data of each sensor of all engines over time. Based on the trend of the graph, discard the engines that have no effect on sensor degradation, and finally obtain the data of the engines that affect sensor degradation. These data are recorded from the brand new state until the sensor is completely scrapped;
[0012] Step 1.2: Perform Z-value normalization on the engine data that affects sensor degradation;
[0013] Step 1.3: Assume that the remaining service life decreases linearly to zero. Select a time node for the engine data after Z value standardization. Before this time, the degradation of any engine is not obvious. Cut it off from here. That is, the remaining service life of the sensor is greater than the node value, and it is set to the node value. Finally, the remaining service life gradually decreases after this node.
[0014] Step 1.4: Add the time-lagged data of the sensor to the truncated data and discard the entries with missing historical data. The data dimension changes from two-dimensional (n, s) to three-dimensional (np*(t n -1), t n , s), where n is the total number of records of all sensors in the training set, t nis the lag order, s is the number of sensors recording the engine working status, and p is the number of engines. Finally, three-dimensional data including the degradation process, time lag order, and number of sensors are obtained.
[0015] Furthermore, in the TCN network based on the multi-attention mechanism, the self-attention mechanism is used to weight the feature dimensions of the preprocessed three-dimensional data to obtain the sequence data of the sensor with weighted feature information. The specific steps are as follows:
[0016] Step 2.1: For the sensor measurements x collected at time step t t ={x 1,t , x 2,t , …, x S,t}, calculate the importance weight, and its calculation formula is as follows:
[0017]
[0018] Where s is the sensor number, t is the time step, and x t is the measured value of the sensor, h w1 is the hidden vector to be learned during training, α s,t is the feature dimension weight of the sth sensor at the tth time step. Therefore, the feature dimension importance weight vector of the sensor is
[0019] Step 2.2: Based on the importance weight, calculate the average importance weight of the sth sensor feature dimension The calculation formula is as follows:
[0020]
[0021] Among them, t is the time step, α s,t is the feature dimension weight of the sth sensor at the tth time step, t n is the lag order;
[0022] Step 2.3: Calculate the sensor data weighted by the feature dimension according to the average value of the feature dimension importance weight and the sensor measurement value. The calculation formula is as follows:
[0023]
[0024] Among them, x s,t is the sensor measurement value of the sth sensor at the tth time step, is the average value of the importance weight of the sth sensor.
[0025] Furthermore, in the TCN network based on the multi-attention mechanism, a temporal convolutional network with shared parameters is used to learn the sequence data obtained in step 2 to obtain the sequence data of the sensor with temporal and spatial correlation information. is the output of the temporal convolutional network at the tth time step of the sth sensor, where:
[0026] The temporal convolutional network consists of two residual blocks, each of which consists of two sparse causal convolutional layers and one convolutional layer. In each residual block, the residual connection input and the output after the convolutional layer are used; each causal convolutional layer is connected to a gated activation layer and a batch normalization layer, where the gated activation layer is defined as:
[0027]
[0028] Where w represents the convolution parameter, o represents the output of the dilated causal convolution layer, * is the convolution operation, ⊙ is the element-wise product, tanh is the inverse sine activation function, and sigmoid is the activation function that maps variables to 0-1. is the output of the gated activation.
[0029] Furthermore, in the TCN network based on the multi-attention mechanism, the self-attention mechanism is used to weight the sequence data obtained in step 3 in the time dimension to obtain the sequence data of the sensor with weighted time information. The specific steps are:
[0030] Step 4.1: Apply the softmax function to the output of the temporal convolutional network The calculation formula is as follows:
[0031]
[0032] in is the output of the temporal convolutional network of the sth sensor, is a randomly generated hidden vector, and the time dimension importance weight vector at the tth time step is λ t =(λ 1,t , ..., λ S,t );
[0033] Step 4.2: Based on the importance weight, calculate the average value of the importance weight of the time dimension The calculation formula is as follows:
[0034]
[0035] Among them, λ t is the time dimension importance weight vector of the tth time step, t is the time step, t n is the lag order;
[0036] Step 4.3: Calculate the weighted sensor according to the average value of the importance weight of the time dimension and the sensor measurement value. The calculation formula is as follows:
[0037]
[0038] in, is the output of the temporal convolutional network at the tth time step of the sth sensor, is the average value of the importance weight of the time dimension.
[0039] Furthermore, in step 3, the zero-start training method is used, and the quantile loss is set as the loss of model training. The TCN network based on the multi-attention mechanism is trained, and the grid search technology is used to obtain the optimal parameters of the network. The specific steps are:
[0040] The loss function of the model was set to 0.1, 0.5, and 0.9 quantile losses, the batch size of the training network was 512, the time lag order was set to 40, and in the grid search, a variable learning rate was used for training. The initial learning rate was 0.001, and the number of network training rounds was 40 to obtain the optimal parameters.
[0041] Further, in step 4, the gas turbine engine sensor data for which the remaining service life is to be predicted is preprocessed and input into the trained model to complete the remaining service life prediction. The specific steps are as follows:
[0042] The optimal weights are loaded into the trained model, and the preprocessed data is predicted on the server. The forward inference process does not perform loss calculation and return loss. The network structure is the same as that during training, and what is returned is the predicted remaining service life of the gas turbine engine.
[0043] A remaining useful life uncertainty prediction system based on a multi-attention mechanism realizes remaining useful life prediction based on a probabilistic remaining useful life prediction framework based on deep learning.
[0044] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the remaining useful life prediction is realized based on a remaining useful life uncertainty prediction framework based on a multi-attention mechanism.
[0045] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements remaining useful life prediction based on a remaining useful life uncertainty prediction framework based on a multi-attention mechanism.
[0046] Compared with the prior art, the present invention has the following significant advantages: 1) The self-attention mechanism is used to weight the data in the feature dimension and the time step dimension respectively, so that information can be adaptively extracted from different features and time steps. 2) TCN with shared parameters is used to apply to the sequence data of all sensors to improve the prediction efficiency. 3) The non-parametric probabilistic remaining useful life prediction framework can provide relevant remaining useful life estimation confidence intervals, better demonstrate the performance of uncertainty quantification, and the non-parametric method shows that the uncertainty decreases with the increase of the cycle. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a remaining useful life uncertainty prediction framework diagram based on a multi-attention mechanism of the present invention.
[0048] Figure 2 This is a framework diagram of the non-parametric multi-attention temporal convolutional network of the present invention.
[0049] Figure 3 This is a comparison chart of the RMSE and Score values of the NPMSA-TCN of the present invention and other methods.
[0050] Figure 4 This is a diagram showing the remaining useful life prediction results of the four engine units in FD003 of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] The present invention provides a remaining useful life uncertainty prediction framework based on a multi-attention mechanism. First, the two-dimensional data of a gas turbine engine gradually degrades from a good state to being scrapped over time is obtained, and then the characteristic dimensions of the sensor data are weighted through a self-attention mechanism. Then, a temporal convolutional network with shared parameters is used to learn the temporal and spatial correlation information of the sensor data. The time dimension is weighted again through a self-attention mechanism, and the sensor data is predicted using a fully connected layer. Then, a zero-start training method is used, and the quantile loss is set as the loss for model training to train a TCN model based on a multi-attention mechanism. Finally, the framework can output a high-precision interval remaining useful life estimate. Figure 1 As shown, the present invention provides a method for predicting uncertainty of remaining useful life based on a multi-attention mechanism, which specifically includes the following steps:
[0053] Step 1: Obtain the two-dimensional data of the gas turbine engine gradually degraded from a good state to scrapped over time to construct a training set, and perform preprocessing. The preprocessed data is three-dimensional data including the degradation process, time lag order, and number of sensors. The specific steps are:
[0054] Step 1.1: Draw a trend graph of the data of each sensor of all engines in the training set over time. Based on the trend of the graph, discard the engines that have no effect on sensor degradation. Finally, get the data of engines that affect sensor degradation. These data are recorded from the brand new state until the sensor is completely scrapped.
[0055] Step 1.2: Perform Z-value normalization on the discarded sensor data.
[0056] Step 1.3: Assume that the remaining service life decreases linearly to zero. Select a suitable time node for the engine data that will affect sensor degradation. Before this time, the degradation of any engine is not obvious. Truncation is performed from here, that is, the remaining service life of the sensor that is greater than the node value is set to the node value. Finally, the data of the remaining service life gradually decreases after this node is obtained.
[0057] Step 1.4: Add the time-lagged data of the sensor to the truncated data and discard the entries with missing historical data. The data dimension changes from two-dimensional (n, s) to three-dimensional (np*(t n -1), t n , s), where n is the total number of records of all sensors in the training set, t n is the lag order, s is the number of sensors recording the engine working status, and p is the number of engines. Finally, the preprocessed data is obtained.
[0058] Step 2: Use the self-attention mechanism to weight the feature dimensions of the preprocessed three-dimensional data to obtain the sensor sequence data with weighted feature information. The specific steps are:
[0059] Step 2.1: For the sensor measurements x collected at time step t t ={x 1,t , x 2,t , …, x S,t}, calculate the importance weight, and its calculation formula is as follows:
[0060]
[0061] Where s is the sensor number, t is the time step, and x t is the measured value of the sensor, h w1 is the hidden vector to be learned during training, α s,tis the weight of the feature dimension of the sth sensor at the tth time step, so the feature dimension importance weight vector of the sensor is
[0062] Step 2.2: Based on the importance weight, calculate the average importance weight of the sth sensor feature dimension The calculation formula is as follows:
[0063]
[0064] Among them, t is the time step, α s,t is the weight of the feature dimension of the sth sensor at the tth time step, t n is the lag order.
[0065] Step 2.3: Calculate the sensor data weighted by the feature dimension according to the average value of the feature dimension importance weight and the sensor measurement value. The calculation formula is as follows:
[0066]
[0067] Among them, x s,t is the sensor measurement value of the sth sensor at the tth time step, is the average value of the importance weight of the sth sensor.
[0068] Step 3: Use a temporal convolutional network with shared parameters to learn the sequence data obtained in step 2 to obtain the sequence data of the sensor with temporal and spatial correlation information. is the output of the temporal convolutional network at the tth time step of the sth sensor). The steps for building the temporal convolutional network model are:
[0069] The temporal convolutional network consists of two residual blocks, each of which consists of two sparse causal convolutional layers and one convolutional layer. In each residual block, the residual connection input and the output after the convolutional layer are used; each causal convolutional layer is connected to a gated activation layer and a batch normalization layer, where the gated activation layer is defined as:
[0070]
[0071] Where w represents the convolution parameter, o represents the output of the dilated causal convolution layer, * is the convolution operation, ⊙ is the element-wise product, tanh is the inverse sine activation function, and sigmoid is the activation function that maps variables to 0-1. is the output of the gated activation.
[0072] Step 4: Use the self-attention mechanism to weight the sequence data obtained in step 3 in the time dimension to obtain the sequence data of the sensor with weighted time information. The specific steps are:
[0073] Step 4.1: Apply the softmax function to the output of the temporal convolutional network The calculation formula is as follows:
[0074]
[0075] in is the output of the temporal convolutional network of the sth sensor, is a randomly generated hidden vector, and the time dimension importance weight vector at the tth time step is λ t =(λ 1,t , ..., λ S,t ).
[0076] Step 4.2: Based on the importance weight, calculate the average value of the importance weight of the time dimension The calculation formula is as follows:
[0077]
[0078] Among them, λ t is the time dimension importance weight vector of the tth time step, t is the time step, t n is the lag order.
[0079] Step 4.3: Calculate the weighted sensor according to the average value of the importance weight of the time dimension and the sensor measurement value. The calculation formula is as follows:
[0080]
[0081] in, is the output of the temporal convolutional network at the tth time step of the sth sensor, is the average value of the importance weight of the time dimension.
[0082] Step 5: Use the fully connected layer to predict the sensor data and obtain the predicted remaining service life.
[0083] Step 6: Use the zero-start training method and set the quantile loss as the loss of model training to train the TCN network based on the multi-attention mechanism constructed in steps 2-5, and use the grid search technology to obtain the optimal parameters of the network. The specific steps are:
[0084] The loss function of the model was set to 0.1, 0.5, and 0.9 quantile losses. The batch size of the training network was 512, the time lag order was set to 40, and in the grid search, a variable learning rate was used for training. The initial learning rate was 0.001, and the number of network training rounds was 40. The optimal parameters were obtained.
[0085] Step 7: Input the preprocessed 3D data into the trained model to complete the remaining service life prediction. The specific steps are:
[0086] The optimal weights are loaded into the trained model, and the preprocessed data is predicted on the server. The forward inference process does not perform loss calculation and return loss. The network structure is the same as that during training, and what is returned is the predicted remaining service life of the gas turbine engine.
[0087] Example
[0088] In order to verify the effectiveness of the scheme of the present invention, the following experiment was carried out.
[0089] In this embodiment, a public turbine engine degradation dataset provided by NASA is used to evaluate the performance of this framework. C-MAPSS is a tool that simulates the entire degradation process of large commercial turbine engines under different operating conditions and failure modes. It contains many customizable input parameters to simulate different degradation processes. C-MAPSS generates four sub-datasets, recorded as FD001, FD002, FD003 and FD004. Each sub-dataset contains 26 features, of which 21 measurements are time series data collected by the suit sensors. A non-parametric remaining useful life prediction method based on a multi-step self-attention temporal convolutional neural network specifically includes the following steps:
[0090] In an embodiment, a public turbine engine degradation dataset provided by NASA is used to evaluate the performance of the present framework.
[0091] Table 1 Parameters of this framework
[0092]
[0093] The two-dimensional data of a gas turbine engine gradually degraded from a good state to scrapped over time are obtained, divided into training set and test set, and preprocessed to obtain three-dimensional data including degradation process, time lag order, and number of sensors. The detailed information of the four sub-datasets is listed in Table 2.
[0094] Table 2 Details of the four sub-datasets in the C-MAPSS dataset
[0095]
[0096] Data feature selection: 7 measurements with unchanged values were removed. The remaining set contains 14 sensor measurements S new =(2, 3, 4, 7, 8, 9, 11, 12, 13, 14, 15, 17, 20, 21) for experiments.
[0097] Normalization based on operating conditions: For each sensor measurement S new Normalize:
[0098]
[0099] in and Respectively represent u th Normalized and raw measured values of the sensors in the engine at each operating condition th . and Respectively represent s th Mean and standard deviation of the sensors under each operating condition.
[0100] In this part, a common sliding time window technique is applied to generate sequential training and test samples. The sliding time window means that the input samples of the network are measured from the first 1~l, the second 2~(l+1), the third 3~(l+2) timestamps, etc. for each engine in the queue. The size of the time window is equal to 30. Finally, the dimension of each training sample is x∈R 1×30×14 , respectively represent the input channel, window size and number of features. In addition, the remaining useful life labels of the training samples are generated by a piecewise linear function with a constant remaining useful life in the early stage. In this embodiment, the constant remaining useful life is set to 125.
[0101] This model applies the root mean square error (RMSE) and score function. The two indicators are as follows:
[0102]
[0103]
[0104] Where Δ i Indicates the actual remaining useful life y i and predict remaining useful life The difference between * is the number of test samples. If the remaining useful life is underestimated, α is 1 / 13, and if it is overestimated, α is 1 / 10. Therefore, the scoring function is asymmetric and penalizes overestimation of the remaining useful life.
[0105] The performance of the probabilistic remaining useful life prediction is evaluated by the quantile loss at a predefined quantile level q, denoted as QLq (For example, QL 0.1 ). The loss function of the model is set to the sum of 0.1, 0.5, and 0.9 quantile losses. The Adam optimizer is used to train the prediction model. The initial learning rate is set to 0.001, the batch size is 256, and the epoch is 60. In addition, learning rate annealing is used during training, and the prediction results of each experiment are the average of the last 20 rounds.
[0106] 1) Impact of different time window sizes: The size of the time window significantly affects the results of the remaining useful life prediction. Figure 3 The prediction results of FD003 with different time window sizes are shown. The results of this framework shown in the following experiments are based on these time window sizes.
[0107] 2) Impact of multi-step self-attention mechanism: Two attention mechanisms are used for both time steps and features. Four experiments are implemented: 1) Basic TCN; 2) TCN with time step self-attention only; 3) TCN with feature self-attention only; 4) TCN with multi-step self-attention; The experimental results are shown in Table 3.
[0108] Table 3 Performance comparison between basic TCN, TCN with time-step self-attention, TCN with feature self-attention, and TCN with multi-step self-attention
[0109]
[0110]
[0111] 3) Performance comparison with other methods: All experiments were performed five times, and the mean and standard deviation (STD) values of RMSE and Score were used as the results.
[0112] Table 4 Comparison of prediction performance
[0113]
[0114] 4) Non-parametric uncertainty prediction: The probability prediction results of the four sub-datasets are shown in Table 5.
[0115] Table 5 Comparison of prediction performance between parametric and non-parametric methods
[0116]
[0117] A nonparametric method based on quantile regression (quantile ranks Q = {0.05, 0.5, 0.95}) and a parametric method based on the Gaussian assumption are used in this section, with a trade-off parameter λ = 1.
[0118] exist Figure 4The 90% confidence interval remaining useful life predictions for four randomly selected units from FD003 are shown in Figure . Figure 4 The estimated remaining useful life predicted by this framework is shown compared with the actual remaining useful life.
[0119] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0120] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for predicting uncertainty of remaining useful life based on multi-attention mechanism, It is characterized in that The steps include: Step 1, obtain two-dimensional data of a gas turbine engine that gradually degrades from a good state to a scrapped state over time to construct a training set, and perform preprocessing, the preprocessed data is three-dimensional data including the degradation process, the time lag order, and the number of sensors; Step 2: construct a TCN network based on a multi-attention mechanism, use the self-attention mechanism to weight the feature dimensions of the preprocessed three-dimensional data, and obtain the sequence data of the sensor with weighted feature information; use a temporal convolutional network with shared parameters to learn the sequence data with weighted feature information, and obtain the sequence data of the sensor with time and space correlation information; use the self-attention mechanism to weight the sequence data with time and space correlation information in the time dimension, and obtain the sequence data of the sensor with weighted time information; use the fully connected layer to predict the sequence data with weighted time information, and obtain the predicted remaining service life; Step 3: Use the zero-start training method and set the quantile loss as the loss of model training to train the TCN network based on the multi-attention mechanism, and use the grid search technology to obtain the optimal parameters of the network; Step 4, preprocessing the gas turbine engine sensor data for which the remaining service life is to be predicted, inputting the trained model, and completing the remaining service life prediction; Step 1: Obtain two-dimensional data of the gas turbine engine gradually degrading from a good state to scrapping over time, and perform preprocessing. The specific steps are as follows: Step 1.1: Draw a trend graph of the data of each sensor of all engines over time. Based on the trend of the graph, discard the engines that have no effect on sensor degradation, and finally obtain the data of the engines that affect sensor degradation. These data are recorded from the brand new state until the sensor is completely scrapped; Step 1.2: Perform Z-value normalization on the engine data that affects sensor degradation; Step 1.3: Assume that the remaining service life decreases linearly to zero. Select a time node for the engine data after Z value standardization. Before this time, the degradation of any engine is not obvious. Cut it off from here. That is, the remaining service life of the sensor is greater than the node value, and it is set to the node value. Finally, the remaining service life gradually decreases after this node. Step 1.4: Add the time-lagged data of the sensor to the truncated data and discard the entries with missing historical data. The data dimension changes from two-dimensional (n, s) to three-dimensional (np*(t n -1), t n , s), where n is the total number of records of all sensors, t n is the lag order, s is the number of sensors recording the engine working status, and p is the number of engines. Finally, three-dimensional data including the degradation process, time lag order, and number of sensors are obtained. In the TCN network based on the multi-attention mechanism, the self-attention mechanism is used to weight the feature dimensions of the preprocessed three-dimensional data to obtain the sequence data of the sensor with weighted feature information. The specific steps are as follows: Step 2.1: For the sensor measurements x collected at time step t t ={x 1,t ,x 2,t ,...,x S,t }, calculate the importance weight, and its calculation formula is as follows: Where s is the sensor number, t is the time step, and x t is the measured value of the sensor, h w1 is the hidden vector to be learned during training, α s,t is the weight of the feature dimension of the sth sensor at the tth time step. Therefore, the feature dimension importance weight vector of the sensor is Step 2.2: Based on the importance weight, calculate the average importance weight of the sth sensor feature dimension The calculation formula is as follows: Among them, t is the time step, α s,t is the weight of the feature dimension of the sth sensor at the tth time step, t n is the lag order; Step 2.3: Calculate the sensor data weighted by the feature dimension according to the average value of the feature dimension importance weight and the sensor measurement value. The calculation formula is as follows: Among them, x s,t is the sensor measurement value of the sth sensor at the tth time step, is the average value of the importance weight of the sth sensor.
2. According to claim 1, a method for predicting uncertainty of remaining useful life based on a multi-attention mechanism, It is characterized in that In the TCN network based on the multi-attention mechanism, a temporal convolutional network with shared parameters is used to learn the sequence data obtained in step 2 to obtain the sequence data of the sensor with temporal and spatial correlation information. is the output of the temporal convolutional network at the tth time step of the sth sensor, where: The temporal convolutional network consists of two residual blocks, each of which consists of two sparse causal convolutional layers and one convolutional layer. In each residual block, the residual connection input and the output after the convolutional layer are used; each causal convolutional layer is connected to a gated activation layer and a batch normalization layer, where the gated activation layer is defined as: Where w represents the convolution parameter, o represents the output of the dilated causal convolution layer, * is the convolution operation, ⊙ is the element-wise product, tanh is the inverse sine activation function, and sigmoid is the activation function that maps variables to 0-1. is the output of the gated activation.
3. According to claim 2, a method for predicting uncertainty of remaining useful life based on a multi-attention mechanism, It is characterized in that In the TCN network based on the multi-attention mechanism, the self-attention mechanism is used to weight the sequence data obtained in step 3 in the time dimension to obtain the sequence data of the sensor with weighted time information. The specific steps are: Step 4.1: Apply the softmax function to the output of the temporal convolutional network The calculation formula is as follows: in is the output of the temporal convolutional network of the sth sensor, is a randomly generated hidden vector, and the time dimension importance weight vector at the tth time step is λ t =(λ 1,t ,...,λ S,t ); Step 4.2: Based on the importance weight, calculate the average value of the importance weight of the time dimension The calculation formula is as follows: Among them, λ t is the time dimension importance weight vector of the tth time step, t is the time step, t n is the lag order; Step 4.3: Calculate the weighted sensor according to the average value of the importance weight of the time dimension and the sensor measurement value. The calculation formula is as follows: in, is the output of the temporal convolutional network at the tth time step of the sth sensor, is the average value of the importance weight of the time dimension.
4. According to claim 1, a method for predicting uncertainty of remaining useful life based on a multi-attention mechanism, It is characterized in that Step 3: Use the zero-start training method and set the quantile loss as the loss of model training to train the TCN network based on the multi-attention mechanism and use the grid search technology to obtain the optimal parameters of the network. The specific steps are as follows: The loss function of the model was set to 0.1, 0.5, and 0.9 quantile losses, the batch size of the training network was 512, the time lag order was set to 40, and in the grid search, a variable learning rate was used for training. The initial learning rate was 0.001, and the number of network training rounds was 40 to obtain the optimal parameters.
5. According to claim 1, a framework for predicting uncertainty of remaining useful life based on a multi-attention mechanism, It is characterized in that Step 4, preprocessing the gas turbine engine sensor data for which the remaining service life is to be predicted, the specific steps are as follows: The optimal weights are loaded into the trained model, and the preprocessed data is predicted on the server. The forward inference process does not perform loss calculation and return loss. The network structure is the same as that during training, and what is returned is the predicted remaining service life of the gas turbine engine.
6. A Remaining Useful Life Uncertainty Prediction System Based on Multi-Attention Mechanism, It is characterized in that The remaining useful life uncertainty prediction method based on the multi-attention mechanism described in any one of claims 1-5 realizes accurate prediction of the remaining useful life.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, an accurate prediction of the remaining useful life is achieved based on the remaining useful life uncertainty prediction method based on a multi-attention mechanism as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, an accurate prediction of the remaining useful life is achieved based on the remaining useful life uncertainty prediction method based on a multi-attention mechanism as described in any one of claims 1-5.
Citation Information
Patent Citations
Method for predicting residual life of equipment in industrial process
CN113486578A
Prediction model establishment method and prediction method for residual life of engineering mechanical part
CN114169091A