A method for predicting the remaining service life of water pump equipment based on improved Transformer model

Through the improved Transformer model, the important attention screening mechanism is used to solve the problems of low prediction accuracy of the remaining service life of the water pump equipment and model redundancy, achieving higher prediction accuracy and model simplification, while ensuring the timeliness and accuracy of predictions.

CN115688556BActive Publication Date: 2025-05-23ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211094317.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-05-23
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

When the prior art predicts the remaining service life of the water pump equipment, the prediction accuracy is not high and the model is redundant, which affects the accuracy of the prediction.

Method used

The improved Transformer model is adopted to filter out important attention by introducing a metric based on the attention mechanism, thereby predicting the remaining service life of the water pump equipment. While ensuring prediction accuracy, the model reduces the parameters of the model.

Benefits of technology

The accuracy of the remaining service life prediction of the water pump equipment is improved, the complexity and parameter volume of the model are reduced, and the asymmetric loss function is used to ensure that the predicted value is less than the true value, and the equipment that is about to be damaged is replaced in time to avoid serious losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115688556B_ABST
    Figure CN115688556B_ABST
Patent Text Reader

Abstract

A method for predicting the remaining service life of a water pump equipment based on an improved Transformer model comprises the following steps: step S1 collects data and constructs and preprocesses a data set, step S2 divides the data set into a training set and a validation set in proportion, step S3 constructs an input vector for the data set, and performs encoding of the input vector in step S4, first-layer improved multi-head self-attention mechanism processing in step S5, and second-layer multi-head attention mechanism processing in step S6 to obtain an output vector, step S7 calculates the loss value of the output vector through an asymmetric loss function, step S8 repeats steps S3 to S7 to train the model for n rounds to obtain an optimal model, and step S9 collects real-time data and inputs it into the model to predict the remaining service life of the water pump equipment; the method improves the multi-head attention mechanism to screen out important attention, improves the information extraction capability of the model and reduces model parameters, so as to achieve better prediction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for predicting the remaining service life of a water pump equipment based on an improved Transformer model. Background Art

[0002] In the context of rapid urbanization in my country, the demand for residential and industrial water has surged, which has put forward higher requirements for water supply companies to deliver stable water supply. Therefore, the reliability of water pump equipment operation is very important. Accurately predicting the remaining service life of water pump equipment can not only ensure the stability of water supply, but also reduce the maintenance cost of water pump equipment and replace the damaged water pump equipment in time.

[0003] At present, the prediction methods for the remaining service life of equipment can be divided into methods based on physical models, statistical models and artificial intelligence. Among them, artificial intelligence methods can accurately predict the remaining service life of equipment due to their excellent information extraction and information representation capabilities. In the field of artificial intelligence, recurrent neural networks and their variants are widely used, but the training speed of recurrent neural network models is slow. As the length of the input sequence increases, the prediction ability will decrease significantly, affecting the prediction accuracy of the model. Summary of the invention

[0004] In order to overcome the problems of low prediction accuracy and redundant prediction models in the existing methods for predicting the remaining service life of water pump equipment, the present invention proposes to predict the remaining service life of water pump equipment based on an improved Transformer model. This model makes improvements based on the attention mechanism and proposes a metric to screen out important attention. The remaining service life is predicted using the screened attention, which reduces the parameters of the model while ensuring the prediction accuracy of the model.

[0005] The technical solution of the present invention is as follows: a method for predicting the remaining service life of a water pump equipment based on an improved Transformer model, comprising the following steps:

[0006] S1. Collect sensor measurement features at different times and the remaining service life of the pump equipment to form a data set S, and preprocess the data set S to obtain the data set S 1 , the dataset S is represented as {S i |S i =(t, x 1 , x 2 , ..., x k , r), 1≤i≤n}, where S i represents the i-th data, t represents the acquisition time, x 1 , x 2 , ..., x krepresents the measurement features of the 1st to the kth sensors, r represents the remaining service life, k represents the number of sensor measurement features, and n is the size of the data set S;

[0007] S2, the data set S 1 Divide into training set T and validation set V in proportion;

[0008] S3, for data set S 1 Use the sliding window to construct the input vector, the input vector is the vector F input and vector F output ;

[0009] S4, vector F input and vector F output Encode them separately to get the corresponding vectors and vector

[0010] S5. Vector and vector After the first layer of multi-head self-attention mechanism processing, it is input into the residual network and fully connected layer of the Transformer model to obtain the corresponding output vector and vector

[0011] S6. Vector and vector At the same time, the second layer of multi-head attention mechanism is processed to obtain the vector Then the vector Input to the residual network and fully connected layer of the Transformer model to obtain the final output result F out ;

[0012] S7. Calculate the output result F through the asymmetric loss function out The loss value is output as F out When the input vector comes from the training set T, back propagation updates the model parameters, F out Record the loss value when the input vector comes from the validation set V;

[0013] S8, repeatedly execute S3 to S7 to train the Transformer model, record the loss value of each sliding window in the validation set V and calculate the average value as the loss value of this round of training, train the model for n rounds, and save the Transformer model parameters with the smallest loss value as the optimal Transformer model parameters;

[0014] S9. Obtain the detection value of the sensor of the water pump equipment, obtain a vector with a dimension of (l-1)×(m+2), and input it into the Transformer model to obtain the remaining service life r.

[0015] Preferably, in step S1, the pre-processing step includes:

[0016] S1.1. Remove invalid sensor features, calculate the standard deviation of each sensor feature, remove sensor features whose data standard deviation is less than the preset standard deviation, and obtain the data set S 1 Expressed as Where m≤k;

[0017] S1.2. For dataset S 1 Perform standardization, including standardization of time feature t and data feature x 1 , x 2 , ..., x m , r standardization, time feature t standardization is to convert the time feature t into a feature vector with a value distribution in the interval [-0.5, 0.5], and the dimension is converted from n×1 to n×a, where a is the length of the converted time feature t, data feature x 1 , x 2 , ..., x m , r is standardized by the formula Get a standardized vector with a mean of 0, a standard deviation of 1, and a dimension of n×(m+1). Finally, add the vector of time feature t and data feature x 1 , x 2 , ..., x m , the vectors of r are concatenated by matrices to obtain a dataset S with a dimension of n×(a+m+1) 1 , where μ 1 , μ 2 , ...μ m , μ r and σ 1 , σ 2 , ...σ m , σ r The datasets S are 1 Data feature x 1 , x 2 ,...x m , the mean and standard deviation of r.

[0018] Preferably, the method of constructing the input vector by sliding window in step S3 is:

[0019] The dimension of the feature vector selected by the sliding window is l×(a+m+1), the length of the sliding window is l, the step size is 1, and the first l-1 data from the current sliding window are taken as the vector F input , vector F input The dimension is (l-1)×(a+m+1), then The data is a vector Foutput , vector F output The dimension is

[0020] Preferably, in step S4, the encoding step includes:

[0021] S4.1. Vector F input Data features x 1 , x 2 , ..., x m , r is value encoded, and the vector F is converted into input The dimension of is increased from (l-1)×(m+1) to (l-1)×512, and the value matrix A is obtained. v ;

[0022] S4.2. Vector F input Data features x 1 , x 2 , ..., x m , r is used for position encoding, through the formula or The calculated position matrix A is (l-1)×512 in dimension p , where 0≤j<256;

[0023] S4.3. Vector F input The time feature t is temporally encoded, and the dimension of the time feature t is increased from (l-1)×a to (l-1)×512 through the fully connected layer to obtain the time matrix A t ;

[0024] S4.4. The value matrix A v 、Position matrix A p and the time matrix A t Add to get vector vector The dimension is (l-1)×512;

[0025] S4.5. Vector F output The last row of x 1 , x 2 , ..., x m , r feature is set to 0 occlusion, and the vector F is calculated output The value matrix A v 、Position matrix A p , time matrix A t And add them together to get the vector The dimension is

[0026] Preferably, in step S5, the vector and vector After the first layer of multi-head self-attention mechanism processing, it is input into the residual network and fully connected layer of the Transformer model to obtain the corresponding output vector and vector The steps are as follows:

[0027] S5.1. Vector To process the multi-head self-attention mechanism, first transform the vector Assign values ​​to Q, K, and V respectively, and then divide Q, K, and V into 8 self-attentions to get Q t , K t 、V t , where 1≤t≤8, vector The dimension of Q is adjusted from (l-1)×512 to 8×(l-1)×64. t , K t 、V t Use the metric formula: Calculate, where q i ∈Q t , k j ∈K t , 1≤i, j≤l-1, the calculated M(q i , K t ) in descending order and select the first (l-1) / 2 values, round down the number of values, and t q i ,q i Statement Q t For the i-th row of data in , use the attention formula: Calculate the attention value for V t By row averaging, after row averaging, V t The q i The position index of the position index is replaced by the calculated attention value to obtain F t , put 8 F t By concatenating the matrices, we get a vector F with a dimension of (l-1)×512. att ;

[0028] S5.2. Vector F att Input to the residual network and fully connected layer of the Transformer model to obtain the vector

[0029] S5.3. Processing vectors using multi-head self-attention mechanism Where M(q i , K t) are sorted in descending order and the top (l+1) / 4 values ​​are selected, and the processed vectors are input into the residual network and fully connected layer of the Transformer model to obtain the vector

[0030] Preferably, the step S6 includes the following steps:

[0031] S6.1. Vector Assign to Q, vector Assign to, vectors Q, K and V are split into 8 attention blocks and then calculate the attention separately, and concatenate the 8 attention block matrices into vector

[0032] S6.2. Vector Input to the residual network and fully connected layer of the Transformer model to obtain the vector F out , vector F out The dimension is

[0033] Preferably, the loss value calculation process in step S7 is:

[0034] F out The last row of data is recorded as the predicted value F pred , vector F output The last row of data is recorded as the true value F true , the loss value is calculated by the asymmetric loss function Loss.

[0035] Preferably, the asymmetric loss function Loss is:

[0036] in

[0037] The beneficial effects of the present invention are as follows: by collecting the running time and sensor information of the water pump equipment to predict the remaining service life of the water pump, the improved multi-head attention mechanism is used to screen out important attention, which not only ensures the ability to extract important information, but also reduces the parameters of the model and achieves better prediction results. At the same time, an asymmetric loss function is added to the prediction model to make the remaining service life prediction of the model as small as possible from the true value, so as to timely replace the water pump equipment with insufficient service life and avoid serious losses caused by the predicted value being too biased. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of the method for predicting the remaining service life of a water pump device according to the present invention.

[0039] Figure 2 To improve the Transformer model training flow chart

[0040] Figure 3 Schematic diagram of sliding windows to improve Transformer model training.

[0041] Figure 4 Schematic diagram of the prediction model of the method of the present invention. DETAILED DESCRIPTION

[0042] The present invention will be further described below in conjunction with the accompanying drawings.

[0043] Reference Figure 1 to Figure 4 , a prediction of the remaining service life of a water pump equipment based on an improved Transformer model, comprising the following steps:

[0044] Step S1, preprocessing the data set S. The data set S is obtained by collecting the vibration signal of the water pump. The rated power of the water pump is 20KW. The vibration signals of 6 different parts are measured by sensors. Each measurement interval is 15 minutes. All data from the beginning of use to the failure of the water pump are measured, with a total of 4605 data. Table 1 shows part of the measured data, t is time, x is 1 , x 2 , x 3 , x 4 , x 5 , x 6 is the sensor measurement characteristic, and r is the remaining service life.

[0045] Table 1. Sensor measurement data

[0046]

[0047] S1.1, remove invalid sensor features, by calculating x 1 , x 2 , x 3 , x 4 , x 5 , x 6 The standard deviation of x 3 The standard deviation of the feature column is less than 1×10 -6 , x 3 The feature is deleted from the dataset S, and the dataset S is obtained. 1 .

[0048] S1.2. For dataset S 1 After standardization, the dimension is 4605×11. 1The standardization is divided into two parts. One part is to standardize the time feature t, and convert the time feature t into a feature vector with a value in the interval [-0.5, 0.5]. The calculation method is: calculate the current minute divided by 59, calculate the hour divided by 23, calculate the day of the week divided by 6, calculate the day of the month divided by 30, calculate the day of the year divided by 365, and finally subtract 0.5 from the calculation result to get the feature vector. Taking the time feature in this embodiment as an example, the time feature vector of 2020-03-01 11:15 is [-0.24576, -0.02174, 0.50000, -0.50000, -0.33562]. In this embodiment, a is 5, and the dimension of the time feature vector is 4605×5. The other part is x 1 , x 2 , x 4 , x 5 , x 6 , r feature standardization, through the formula Calculated, where μ 1 , μ 2 , μ 4 , μ 5 , μ 6 , μ r The datasets S are 1 Medium 1 , x 2 , x 4 , x 5 , x 6 , the average value of r feature column, σ 1 , σ 2 , σ 4 , σ 5 , σ 6 , σ r The datasets S are 1 Medium 1 , x 2 , x 4 , x 5 , x 6 , the standard deviation of the feature column r. In this embodiment, m is 5, and the standardized data with a mean of 0, a standard deviation of 1 and a dimension of 4605×6 is calculated. The vector of the time feature t and the data feature x are concatenated by the matrix 1 , x 2 , x 4 , x 5 , x 6 , the vector of r gets the data set S with a dimension of 4605×11 1 .

[0049] Step S2: Set the data set S 1The data is divided into a training set T and a validation set V in a ratio of 4:1. In this embodiment, the training set T is 3684 data items, and the validation set V is 921 data items.

[0050] Step S3: Set the data set S 1 Use the sliding window to construct the input vector, the input vector is the vector F input , vector F output In this embodiment, the length of the sliding window l is set to 25, the step size is 1, the selected data dimension is 25×11, including the time feature t dimension of 25×5 and the data feature x 1 , x 2 , x 4 , x 5 , x 6 , r dimension is 25×6. Take the first 24 data as F input The dimension is 24×11, and the last 6 data are taken as F output The dimension is 13 × 11. The training set T obtains 3660 sets of input vectors, and the validation set V obtains 897 sets of input vectors.

[0051] Step S4: vector F input and vector F output Encode to get vector and vector The details are as follows:

[0052] S4.1. Vector F input Data features x 1 , x 2 , x 4 , x 5 , x 6 , r is used for value encoding. Through a one-dimensional convolution operation with a convolution kernel of 3, the dimension of 24×6 is increased to 24×512, and the value matrix A is obtained. v .

[0053] S4.2. Vector F input Data features x 1 , x 2 , x 4 , x 5 , x 6 , r is used for position encoding, and the calculation formula is: Where 0≤j<256, the dimension is increased from 24×6 to 24×512 to obtain the position matrix A p .

[0054] S4.3. Vector F input The time feature t is temporally encoded, and the time feature dimension 24×5 is increased to 24×512 through the fully connected layer to obtain the time matrix At .

[0055] S4.4. The value matrix A v 、Position matrix A p , time matrix A t Add together to get a vector of dimension 24×512

[0056] S4.5. Vector F output The last row of x 1 , x 2 , x 4 , x 5 , x 6 , r feature is set to 0 occlusion, and the vector F is calculated output The value matrix A v 、Position matrix A p , time matrix A t , and then add them together to get a dimension of 13×512

[0057] Step S5: vector and vector After the first layer of improved multi-head self-attention mechanism processing, the respective vectors F are obtained att , and then input them into the residual network and the fully connected layer to get the corresponding output vector and vector The details are as follows:

[0058] S5.1. Vector Perform an improved multi-head self-attention mechanism. First, vector Assign values ​​to Q, K, and V respectively, and then divide Q, K, and V into 8 self-attentions to obtain Q t , K t 、V t , where 1≤t≤8, the dimension changes from 24×512 to 8×24×64. t , K t 、V t Substituting into the measurement formula: where q i ∈Q t , k j ∈K t , 1≤i, j≤24. i , K t ) in descending order and select the top 12 values. t q i ,q i Statement Q t For the i-th row of data in , use the attention formula: Calculate the attention value for V t By row averaging, after row averaging, V t The q i The position index of the position index is replaced by the calculated attention value to obtain F t , put 8 F t By concatenating the matrices, we get a vector F with a dimension of (l-1)×512. att ;

[0059]

[0060] S5.1. Update the F att The residual network and the fully connected layer obtain a vector of size 24×512

[0061] S5.1. Process the vector using the improved multi-head self-attention mechanism in steps (5.1) and (5.2) Among them, for M(q i , K t ) values ​​in descending order and select the top 6 values, finally obtaining a vector with a dimension of 13×512

[0062] Step S6: vector and vector After the second layer of multi-head attention mechanism processing, the respective vectors are obtained Then input it into the residual network and the fully connected layer to get the final output result F out .

[0063] S6.1. Vector Assign to Q, vector Assign values ​​to K and V, then split into 8 attention blocks, calculate the attention separately, and finally concatenate the 8 attention block matrices into a vector

[0064] S6.1. Vector Input to the residual network and the fully connected layer to get the vector F out , with dimensions of 13×6.

[0065]

[0066] Step S7: Calculate the vector F by using an asymmetric loss function out Loss value, vector F out When the input vector comes from the training set T, the model parameters are updated through back propagation, and the vector F outThe loss value is recorded when the input vector of F comes from the validation set V. In this embodiment, the loss value is calculated by: intercepting F out The last row of data is recorded as the predicted value F pred =[0.3323, 0.5051, -0.9062, 0.0625, -0.1143, 0.0922], intercept vector F output The last row of data is recorded as the true value F true =[1.2921, -1.5065, -0.0267, 0.0177, -0.2252, -1.1172], the loss is calculated by the asymmetric loss function, where the asymmetric loss function is:

[0067]

[0068] Step S8, repeating steps S3 to S7 to train the model, and averaging the loss values ​​of 897 sets of sliding window records in the validation set V to obtain the loss value of this round of training. In this embodiment, the improved Transformer model is trained for 6 rounds, and the improved Transformer model parameters with the smallest loss value are saved as the optimal improved Transformer model parameters.

[0069] Step S9: Load the trained improved Transformer model to predict the remaining service life of the water pump equipment. Input the vector with a dimension of 24×7 into the improved Transformer model to obtain a prediction vector with a dimension of 1×6, which contains the data feature x at the next moment. 1 , x 2 , x 4 , x 5 , x 6 , the predicted value of r, extracting the feature r in the prediction vector to obtain the remaining service life. In this embodiment, the final predicted remaining service life is 695.19.

[0070] Table 2 shows the comparison between the improved Transformer model, the normal Transformer model and the LSTM model. The mean square error (MSE) and mean absolute error (MAE) of the improved Transformer model are both the lowest, and the prediction results are more accurate.

[0071] Table 2. Comparison of the improved Transformer model with the Transformer model and the LSTM model

[0072]

[0073] Those skilled in the art should recognize that the above contents are only used to illustrate the present invention and are not used to limit the present invention. As long as they are within the spirit of the present invention, changes and modifications to the above examples will fall within the scope of the claims of the present invention.

Claims

1. A method for predicting the remaining service life of water pump equipment based on an improved Transformer model. It is characterized in that The following steps are involved: S1. Collect sensor measurement features at different times and the remaining service life of the pump equipment to form a data set S, and preprocess the data set S to obtain the data set S 1 , the dataset S is represented as {S i |S i =(t,x 1 ,x 2 ,…,x k ,r),1≤i≤n}, where S i represents the i-th data, t represents the acquisition time, x 1 ,x 2 ,…,x k represents the measurement features of the 1st to the kth sensors, r represents the remaining service life, k represents the number of sensor measurement features, and n is the size of the data set S; S2, the data set S 1 Divide into training set T and validation set V in proportion; S3, for data set S 1 Use the sliding window to construct the input vector, the input vector is the vector F input and vector F output ; S4, vector F input and vector F output Encode them separately to get the corresponding vectors and vector S5. Vector and vector After the first layer of multi-head self-attention mechanism processing, it is input into the residual network and fully connected layer of the Transformer model to obtain the corresponding output vector and vector S6. Simultaneously process the vector and the vector through the second-layer multi-head attention mechanism to obtain the vector Then, input the vector into the residual network and the fully connected layer of the Transformer model to obtain the final output result F out ; S7. Calculate the output result F through the asymmetric loss function out The loss value is output as F out When the input vector comes from the training set T, back propagation updates the model parameters and the output result F out Record the loss value when the input vector comes from the validation set V; S8, repeatedly execute S3 to S7 to train the Transformer model, record the loss value of each sliding window in the validation set V and calculate the average value as the loss value of this round of training, train the model for n rounds, and save the Transformer model parameters with the smallest loss value as the optimal Transformer model parameters; S9, obtain the detection value of the sensor of the water pump equipment, obtain a vector with a dimension of (l-1)×(m+2), and input it into the Transformer model to obtain the remaining service life r; In step S5, the vector and vector After the first layer of multi-head self-attention mechanism processing, it is input into the residual network and fully connected layer of the Transformer model to obtain the corresponding output vector and vector The steps are as follows: S5.

1. Vector To process the multi-head self-attention mechanism, first transform the vector Assign values ​​to Q, K, and V respectively, and then divide Q, K, and V into 8 self-attentions to get Q t , K t 、V t , where 1≤t≤8, vector The dimension of Q is adjusted from (l-1)×512 to 8×(l-1)×64. t , K t 、V t Use the metric formula: Calculate, where q i ∈Q t , k j ∈K t , 1≤i, j≤l-1, the calculated M(q i ,K t ) in descending order and select the first (l-1) / 2 values, round down the number of values, and t q i ,q i Statement Q t For the i-th row of data in , use the attention formula: Calculate the attention value for V t By row averaging, after row averaging, V t The q i The position index of the position index is replaced by the calculated attention value to obtain F t , put 8 F t By concatenating the matrices, we get a vector F with a dimension of (l-1)×512. att ; S5.

2. Vector F att Input to the residual network and fully connected layer of the Transformer model to obtain the vector S5.

3. Processing vectors using multi-head self-attention mechanism Where M(q i ,K t ) are sorted in descending order and the top (l+1) / 4 values ​​are selected, and the processed vectors are input into the residual network and fully connected layer of the Transformer model to obtain the vector The step S6 includes the following steps: S6.

1. Vector is assigned to Q, and vector is assigned to. After splitting vectors Q, K, and V into 8 attention blocks, attention is calculated separately, and the 8 attention block matrices are concatenated into a vector S6.

2. Vector Input to the residual network and fully connected layer of the Transformer model to obtain the vector F out , vector F out The dimension is 2. A method for predicting the remaining useful life of a water pump equipment based on an improved Transformer model as claimed in claim 1, It is characterized in that In step S1, the pre-processing step includes: S1.

1. Remove invalid sensor features, calculate the standard deviation of each sensor feature, remove sensor features whose data standard deviation is less than the preset standard deviation, and obtain the data set S 1 Expressed as Where m≤k; S1.

2. For dataset S 1 Perform standardization, including standardization of time feature t and data feature x 1 ,x 2 ,…,x m ,r standardization, time feature t standardization is to convert the time feature t into a feature vector with a value distribution in the interval [-0.5, 0.5], and the dimension is converted from n×1 to n×a, where a is the length of the converted time feature t, and the data feature x 1 ,x 2 ,…,x m ,r normalization is done by the formula Get a standardized vector with a mean of 0, a standard deviation of 1, and a dimension of n×(m+1). Finally, add the vector of time feature t and data feature x 1 ,x 2 ,…,x m , r's vectors are concatenated by matrices to obtain a dataset S with a dimension of n×(a+m+1) 1 , where μ 1 ,μ 2 ,…μ m ,μ r and σ 1 ,σ 2 ,…σ m ,σ r The datasets S are 1 Data feature x 1 ,x 2 ,…x m , the mean and standard deviation of r.

3. A method for predicting the remaining useful life of a water pump equipment based on an improved Transformer model as claimed in claim 2, It is characterized in that The method of constructing the input vector by sliding window in step S3 is: The dimension of the feature vector selected by the sliding window is l×(a+m+1), the length of the sliding window is l, the step size is 1, and the first l-1 data from the current sliding window are taken as the vector F input , vector F input The dimension is (l-1)×(a+m+1), then The data is a vector F output , vector F output The dimension is 4. A method for predicting the remaining useful life of a water pump equipment based on an improved Transformer model as claimed in claim 1, It is characterized in that In step S4, the encoding step includes: S4.

1. Vector F input Data features x 1 ,x 2 ,…,x m ,r is value encoded, and the vector F is transformed into input The dimension of is increased from (l-1)×(m+1) to (l-1)×512, and the value matrix A is obtained. v ; S4.

2. Vector F input Data features x 1 ,x 2 ,…,x m ,r position encoding, through the formula or The calculated position matrix A is (l-1)×512 in dimension p , where 0≤j<256; S4.

3. Vector F input The time feature t is temporally encoded, and the dimension of the time feature t is increased from (l-1)×a to (l-1)×512 through the fully connected layer to obtain the time matrix A t ; S4.

4. The value matrix A v 、Position matrix A p and the time matrix A t Add to get vector vector The dimension is (l-1)×512; S4.

5. Vector F output The last row of x 1 ,x 2 ,…,x m ,r feature is set to 0 occlusion, calculate vector F output The value matrix A v 、Position matrix A p , time matrix A t And add them together to get the vector The dimension is 5. A method for predicting the remaining useful life of a water pump equipment based on an improved Transformer model as claimed in claim 1, It is characterized in that The loss value calculation process in step S7 is: F out The last row of data is recorded as the predicted value F pred , vector F output The last row of data is recorded as the true value F true , the loss value is calculated by the asymmetric loss function Loss.

6. A method for predicting the remaining useful life of a water pump equipment based on an improved Transformer model as claimed in claim 5, It is characterized in that The asymmetric loss function Loss is: in Y j ∈F true .

Citation Information

Patent Citations

  • Aero-engine residual life prediction method based on full attention deep network and dynamic ensemble learning

    CN114297918A

  • A system and method for training machine-learning algorithms for processing biology-related data, a microscope and a trained machine learning algorithm

    US20220246244A1