Industrial equipment RUL prediction method based on deconvolution and Transformer
By setting up multiple sensors on industrial equipment to collect data, and using deconvolution and Transformer models for preprocessing and feature fusion, the robustness and interpretability problems of existing methods when processing complex data are solved, and high accuracy RUL prediction of industrial equipment under complex operating conditions is achieved.
Patent Information
- Application Number
- CN202510259282.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
In the IoT industrial scenario, existing data-driven methods are difficult to improve the robustness and interpretability of prediction results under complex operating conditions of industrial equipment when processing high-dimensional, nonlinear and unbalanced data.
The RUL prediction method based on deconvolution and Transformer is adopted. By setting multiple sensors around industrial equipment, multiple sensors are collected, time series data of multiple sensors are carried out, normalized preprocessing and deconvolution data mapping is performed, and the feature fusion method of time domain nonlinear mutual information is combined, information from different data channels is integrated, and finally input into the Transformer model for training.
It realizes accurate prediction of the remaining service life of industrial equipment under complex working conditions, improves the robustness and interpretability of the prediction results, and has good feasibility and superiority.
Smart Images

Figure CN120105910A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of Internet of Things, and specifically is a Transformer-based industrial equipment remaining useful life (RUL) prediction method. Background Art
[0002] In modern industrial production, equipment reliability and maintenance costs are crucial factors. With the development of Industry 4.0, intelligent manufacturing and predictive maintenance have become key technologies to improve production efficiency and reduce maintenance costs. Remaining Useful Lifetime (RUL) prediction is the core content of predictive maintenance. By predicting the time that the equipment can continue to operate normally in the current state, the maintenance plan can be optimized to avoid sudden failures and unnecessary downtime.
[0003] Traditional equipment health management methods mainly include physical model-based methods and statistical methods. The physical model method relies on the working principle and structural characteristics of the equipment, but usually requires an accurate equipment model and a large amount of experimental data, and has poor applicability. The statistical method predicts RUL by analyzing historical data and equipment operation data, but may not perform well when dealing with complex nonlinear relationships. In recent years, with the rapid development of artificial intelligence and machine learning technologies, data-driven RUL prediction methods have gradually become a research hotspot. These methods can more accurately capture the operating status and trends of equipment through a large amount of sensor data and machine learning algorithms, thereby improving the accuracy of RUL prediction. However, existing data-driven methods still face certain challenges in dealing with high-dimensional, nonlinear and unbalanced data, especially under complex working conditions of industrial equipment. The robustness and interpretability of the prediction results still need to be further improved. Summary of the invention
[0004] The purpose of the present invention is to solve the problem of predicting the remaining service life of industrial equipment in a future period of time based on the data of multiple sensors (such as temperature, air pressure, etc.) and the estimated remaining service life at each time point during a period of normal operation in the industrial scenario of the Internet of Things, and to provide a RUL prediction method based on deconvolution and Transformer.
[0005] A RUL prediction method based on deconvolution and Transformer includes the following steps:
[0006] (1) Different types of sensors (turbine temperature, compressor outlet pressure, fan speed, fuel flow, etc.) are installed around the industrial equipment involved in the application scenario (taking aerospace engines as an example) and its various components. Sensor data is usually collected at fixed time steps. The collected data is arranged according to the time step to produce a multivariate sensor time series data set. The remaining useful life (RUL) of each time slice is obtained through the simulation of the physical degradation model and used as the label of each data set. At this point, the necessary data sets are basically constructed, and the main features of each data set include the sensor data of the engine (such as temperature, pressure, speed, etc.). The label of each data set contains the remaining useful life (RUL) information, that is, the estimated remaining useful life of the engine corresponding to each sample point at that time point. The data set usually also includes the measurement values of multiple sensors;
[0007] (2) After arranging the data obtained in step (1), normalization preprocessing is performed, and then the deconvolution method is applied to the data to expand the information source of the data in the time dimension and channel (different sensor data in the same time slice), that is, the original distribution of the data is mapped to a new high-dimensional distribution. Then, the feature fusion method based on time-domain nonlinear mutual information (for dimensionality reduction and fusion information) is used to integrate different distributions from different data channels, and finally a reasonable input data distribution is formed after denormalization;
[0008] (3) The data input finally obtained in step (2) is distributed to the Transformer model to prepare for input embedding and position encoding. Then, the most critical encoding layer and output layer are entered, and the computer is allowed to perform neural network training for several rounds based on the data and model, and finally the remaining service life prediction model for this industrial equipment under this working condition is obtained, and the RUL prediction value is output. The encoding layer uses an improved attention mechanism of the gated linear unit (GLU) containing the SiLU function. SiLU has been proven to be superior to the multi-layer perceptron in various natural language processing tasks.
[0009] Furthermore, the multivariate sensor time series data set in step (1) can be expressed by the following formula:
[0010] RUL=f(X t ,t) (1)
[0011]
[0012] Where t is the timestamp, usually the time point of measurement or the sample number, X trepresents the input features at time t, usually signals from multiple sensors (temperature, pressure, speed, etc.), RUL is the remaining useful life (or health status), which is the predicted output of the physical degradation model and serves as the label for each data. i (t) represents the measurement value of the i-th sensor at time t, where n is the number of features.
[0013] Furthermore, the specific implementation process in step (2) is as follows:
[0014] 2.1 Implement the preprocessing described in step (2), specifically, for each (X t ,t) is normalized, that is, the data is converted into a distribution with a mean of 0 and a standard deviation of 1, in order to eliminate the different magnitudes of its features, and obtain (X Norm ,t);
[0015] 2.2 is the same preparation work. In step 2.1, we get X Norm Based on this, we can transpose the matrix and get Then, put X Norm As input in the time dimension, X Norm As input of the channel dimension, enter the next step.
[0016] 2.3 Set the size S of the convolution kernel, that is, the convolution kernel is equivalent to a matrix of size S*S. The specific parameters w and b are set manually, and finally a convolution kernel K∈S*S is obtained, and then proceed to the next step.
[0017] 2.4 Implement the deconvolution data mapping described in step (2) and transform the X in step 2.2 Norm and Use the convolution kernel K in 2.3 to perform transposed convolution at the same time. Transposed convolution is equivalent to applying the effect of the convolution kernel in reverse to obtain a higher resolution output data distribution Y T and Y C .
[0018] 2.5 Implement the feature fusion based on time domain nonlinear mutual information in step (2) to obtain the final data distribution Y′ after fusion. T and Y C , all use the same feature fusion operation: first, calculate the statistics L (mean μ and standard deviation σ) in the distribution in units of sliding windows to represent the data distribution obtained by deconvolution; then, calculate the remaining useful life of the target variable and each L k The mutual information of the statistic is used to generate Y T and Y CThe weight of each tensor in is calculated and feature fusion is performed. Finally, the results of the time dimension T and the channel dimension C are summarized, that is, the distributions of the two are added. The calculation method of mutual information is time-domain weighting and nonlinear enhancement for time-series sensitive scenarios (equipment degradation prediction). The former gives higher weight to recent data through the time attenuation factor γ, focusing on the key stage at the end of equipment degradation; the latter enhances the ability to capture nonlinear correlations through power transformation. Through the collaborative design of time-domain weighting and nonlinear enhancement, the key features of dynamic evolution in time series data can be captured more accurately. At this point, the operation described in step (2) is completed, and the final data distribution Y′ is obtained after denormalization.
[0019] The expression of the deconvolution method in step (2) can be expressed by the following formula:
[0020] m D,i,j,k =μ(x i,j )=w D,j,k *x i,j +b D,j,k (3)
[0021] Among them, the mass function m D,i,j,k It is generated by a fuzzy membership function and is used to describe the relationship between data at a specific location and certain specific conditions (such as time points, channels, events). i,j ) is a function that maps input data to the range [0,1] and is used to describe an input x i,j The degree of membership in a certain set. D corresponds to the input data (X Norm ,t), i and j are the subscripts of the time dimension and channel dimension respectively, and the parameter w D,j,k and b D,j,k represents the slope and intercept of the membership function, and k corresponds to the event dimension, that is, the convolution kernel size S*S in step 2.3.
[0022] The feature fusion method based on temporal nonlinear mutual information (TN_MI) in step (2) can be expressed by the following formula:
[0023]
[0024] Among them, formula (4) and formula (5) are for Y k (kth generated variable) and each of its channels c, calculate its sliding window statistics ( Represents the tensor Y k The mean and standard deviation of channel c in the time window [t-W+1,t]) form a new variable L k,c, which is used to calculate mutual information in the subsequent step. In formula (4) and formula (5), W is the sliding window size (number of time steps), T is the total number of time steps, t is the time step of a single window, b represents the number of batches, and i is the accumulated variable in the summation calculation. Formula (6) is the key to this step and is used to calculate Y k The time-domain nonlinear mutual information with the target variable RUL, where is the joint probability density of sensor window statistics and remaining service life, p(rul t ) is the marginal probability density of the sensor window statistics and the remaining life. In this way, the correlation between the two variables can be measured based on the result of the logarithmic term. If there is no correlation at all, the logarithm is 0; γ is the time decay coefficient, which controls the decay speed of the weight over time (γ>0, by adjusting γ to control the weight decay curve to dynamically focus on the critical stage of the final degradation period); α is the nonlinear sensitivity parameter, and the size of α is adjusted (greater than 1, strengthening the correlation; less than 1, weakening the correlation) to achieve the purpose of enhancing the adaptability to complex degradation patterns. Finally, formulas (7) and (8) perform weight calculation and feature fusion based on mutual information to obtain a comprehensive feature representation Y′ for each.
[0025] Furthermore, the specific implementation process in step (3) is as follows:
[0026] 3.1 To implement the neural network model training of step (3), preliminary preprocessing is required. Specifically, input embedding and position encoding are performed on the input Y′. The input Y′ is first converted into a matrix of appropriate size for the model through the embedding layer, and then the position encoding is added to the matrix. The position encoding provides the position information of each feature element in each Y′ in the sequence to make up for the defect that the self-attention mechanism is not aware of the order.
[0027] 3.2 After completing the preprocessing required in step (3), to build a formal training model, you need to build an encoder layer, the most critical step of which is the multi-head self-attention layer. Specifically, attention then generates new representations by calculating the relationship between each input and all other inputs. In each encoder layer, attention calculation is performed by calculating the query, key, and value. The query, key, and value are obtained from the output of the previous layer through linear transformation.
[0028] 3.3 After the output of step 3.2, apply residual connection and normalization in each sub-layer to ensure stable model training. Specifically, the residual connection means that the input of each sub-layer (such as the multi-head self-attention layer and the feedforward network layer) is added to the output of the layer, and finally passes through the normalization layer.
[0029] 3.4 After the multi-head self-attention layer in step 3.2, the encoder also includes a feed-forward neural network, which includes two linear layers and a Swish activation function.
[0030] 3.5 Apply the residual connection and normalization in step 3.3 again to process the output of step 3.4.
[0031] 3.6 Use steps 3.2 to 3.5 as the encoder, and then repeatedly stack N encoder layers to enhance the feature abstraction ability layer by layer, and finally output a context-related representation.
[0032] 3.7 Complete the neural network model training in step (3). The last step is to apply the model constructed in 3.6, load the training data sets in steps (1) and (2) on the corresponding computer software, and initialize the weights of the Transformer encoder (the values of various preset parameters). Finally, batch training is performed, and the following operations are repeated in each round of training: obtain input data in batches; forward propagation calculates output; calculates loss (parameters that can reflect model performance, that is, the error with real data); back propagation updates parameters (weights). Until the training converges (the loss no longer changes significantly), all operations in step (3) are completed.
[0033] Furthermore, the formula of the feedforward neural network in step 3.4 is as follows:
[0034] Swish β (x) = x*σ(β*x) (9)
[0035]
[0036] Among them, Swish in formula (9) is an activation function, which is used as a substitute for traditional activation functions (such as ReLU) in deep learning, and σ(x) is the Sigmoid function, which is defined as:
[0037]
[0038] β is a trainable parameter and is usually optimized during the training process. 1 and x 2 are the two parts of the input vector, usually obtained by splitting the input in half, x 1 Processed by Swish activation function, x 2 Directly with the activated x 1 Multiplication; W 1 ,W 2 is the learned weight matrix, b 1 and b 2 is the bias term. Formula (10) is partially expressed by x2 It is used as a gating mechanism to weight the results after activation. Only those input information with higher weights will flow through the activation function, which allows the model to selectively control the flow of information during calculation.
[0039] Furthermore, the update expression of the gated encoder layer in step 3.2 is as follows: l =
[0040] MSA(Linear(Z l ))+Z l (12)
[0041] Z l+1 =SwiGLU(Linear(Y l ))+Y l (13)
[0042] Where l represents the number of layers of the current encoder (in Transformer, multiple encoders are stacked to allow the model to fully learn), and l+1 represents the next layer. In the above formula (13), if l+1 is the current layer, then Z l The output of the previous layer is used as the input of the current layer. After linear transformation and MSA processing, it is finally added to the output of SiLU to obtain the output Z of the current layer. l +1 The Linear() in the above formulas represents linear transformation processing, and the MSA() function represents multi-head attention.
[0043] Beneficial effects of the present invention:
[0044] In view of the insufficient data distribution information in the problem scenario and the sequential nature of the time series data itself, the present invention proposes a deep learning model method based on an improved Transformer based on deconvolution, which achieves the goal of predicting the remaining service life of industrial equipment under complex working conditions. The method of the present invention is simple to calculate, the results are effective, and it is feasible and superior in specific industrial production scenarios. In addition, the present invention has good reusability in similar application scenarios, and the practical value of the invention is relatively strong. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of the execution flow of the present invention;
[0046] Figure 2 It is a schematic diagram of the structure of the model of the present invention;
[0047] Figure 3 Schematic diagram of the encoder layer structure of the present invention. DETAILED DESCRIPTION
[0048] In order to enable relevant personnel to more clearly understand the technical content of the present invention, the technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods. The present invention discloses a RUL prediction method for industrial equipment based on deconvolution and Transformer. Sensors can be used to collect various indicators of industrial equipment to predict the remaining service life in the future to guide the development of maintenance work. In view of the insufficient data distribution information in the problem scenario and the sequential nature of the time series data itself, the present invention proposes a deep learning model method based on an improved Transformer based on deconvolution, which achieves the goal of predicting the remaining service life of industrial equipment under complex working conditions. The method of the present invention is simple to calculate, the results are effective, and it is feasible and superior in specific industrial production scenarios. In addition, the present invention has good reusability in similar application scenarios, and the practical value of the invention is relatively strong.
[0049] In order to solve the above problems, the present invention adopts Transformer to learn the complex features in the multi-sensor time series of industrial equipment. The main components of Transformer can be divided into encoder and decoder, and both parts are stacked by multiple identical layers. In order to reduce the heavy degree of model training, the present invention only uses the encoder part. The input of the encoder is a sequence (such as time series, text, etc.), and its task is to convert the input sequence into a fixed-length representation, usually a context-related vector sequence. Each layer of the encoder consists of two sublayers: a multi-head self-attention mechanism layer and a feedforward neural network layer. Each sublayer has a residual connection and layer normalization, which helps the propagation of gradients and prevents problems in the training of deep networks. Its unique self-attention mechanism enables it to effectively capture long-distance dependencies in sequence data, which is of great significance for RUL prediction. Specifically, there are the following points: The Transformer model can simultaneously consider all time steps in the sequence data through the self-attention mechanism, thereby effectively capturing long-distance dependencies. This is very helpful for predicting the long-term operating trends and potential failures of equipment; the Transformer model has powerful representation capabilities and can handle high-dimensional and nonlinear sensor data. Through the multi-head self-attention mechanism and feedforward neural network, the Transformer is able to learn complex data features and patterns; the self-attention mechanism of the Transformer model provides an intuitive way to explain the decision-making process of the model. By analyzing the self-attention weights, the time steps and features that have the greatest impact on the prediction results can be identified, thereby improving the interpretability and robustness of the model. In summary, the Transformer-based RUL prediction method provides a new technical path for the health management of industrial equipment, which can significantly improve the performance and efficiency of traditional methods in practical applications while improving prediction accuracy.
[0050] Deconvolution itself is a classic method for signal restoration and image reconstruction, and is often used to solve the problem of reverse recovery of signals or images blurred by convolution operations. The core of the deconvolution operation is to expand low-resolution feature maps to higher resolution forms by learning weights. Unlike direct interpolation (such as nearest neighbor interpolation or bilinear interpolation), deconvolution can use patterns and distributions in training data to learn complex mapping relationships, thereby generating more detailed and structured high-resolution outputs. The multivariate time series data in RUL prediction, that is, the indicators obtained by multiple different sensors in equal time, can be used to benchmark signal processing, use deconvolution to expand the information of the data, and improve the quality of the input data. Therefore, deconvolution technology is very suitable for studying the problem of chaotic or insufficient RUL data information of industrial equipment. Deconvolution technology is an operation for generating high-resolution feature maps from low-resolution feature maps. First, deconvolution receives a set of input feature maps, which are usually low-resolution data from previous convolution or pooling operations, containing key information that needs to be expanded. Then define the convolution kernel, also known as the filter, which is a small weight matrix whose size and shape are set according to specific task requirements, such as 3×3 or 5×5. This convolution kernel learns the mapping relationship between input and output features through training. In order to expand the spatial resolution of the feature map, zero values are inserted between pixels of the input feature map when the stride is set to be greater than 1. For example, if the stride is 2, a zero value is inserted between every two pixels, making the feature map twice the original size. Next, the convolution kernel moves on the expanded feature map like a sliding window, covering a specific range of pixels each time. Each sliding operation of the convolution kernel multiplies the pixel value at the corresponding position, and these results are added to generate a new pixel value, which is filled in the output feature map. Finally, the deconvolution operation generates a high-resolution feature map. Compared with the direct interpolation method, deconvolution can generate more detailed and realistic high-resolution results based on the input data by training the convolution kernel weights. However, the deconvolution operation itself has limitations: details may be lost or noise may be introduced in the process. At this time, feature fusion can significantly improve the model's ability to represent complex patterns by integrating feature information from different levels or sources. The feature tensor after deconvolution usually corresponds to a single scale and is difficult to adapt to changes in actual scenes. Therefore, the multi-scale features output by different deconvolution layers are integrated through feature fusion, so that the model can capture the overall structure of large targets and the local features of small targets at the same time. In addition to deconvolution, the present invention also improves the attention mechanism in Transformer, that is, a unique gated linear unit (Gated Linear Unit) is added. The design purpose of GLU is to selectively transmit information through a gating mechanism, thereby improving the expressiveness and performance of the model. It is widely used in natural language processing (NLP) and computer vision (CV) tasks.Therefore, this method applies deconvolution modeling and proposes a RUL prediction method based on deconvolution and Transformer.
[0051] A method for predicting RUL of industrial equipment based on deconvolution and Transformer, the implementation example includes the following steps:
[0052] Step 1. Set different types of sensors (turbine temperature, compressor outlet pressure, fan speed, fuel flow, etc.) around the industrial equipment involved in the application scenario (taking aerospace engines as an example) and its various components. Sensor data is usually collected at fixed time steps. The collected data is arranged according to the time step to produce a multivariate sensor time series data set. The remaining useful life (RUL) of each time slice is obtained by simulating the physical degradation model and used as the label of each data set. At this point, the necessary data sets are basically constructed, and the main features of each data set include the sensor data of the engine (such as temperature, pressure, speed, etc.). The label of each data set contains the remaining useful life (RUL) information, that is, the estimated remaining useful life of the engine corresponding to each sample point at that time point.
[0053] Step 2: Convert the data in the dataset into a matrix. The mathematical expression is as follows:
[0054] RUL=f(X t ,t) (1)
[0055]
[0056] Where t is the timestamp, usually the time point of measurement or the sample number, X t represents the input features at time t, usually signals from multiple sensors (temperature, pressure, speed, etc.), RUL is the remaining useful life (or health status), which is the predicted output of the physical model and serves as the label for each data. i (t) represents the measurement value of the i-th sensor at time t, where n is the number of features.
[0057] Step 3: After arranging the data obtained in step 1, normalize and preprocess them to obtain X Norm , set the size of the convolution kernel S, that is, the convolution kernel is equivalent to a matrix of size S*S, and then manually set the specific parameters w, b, and finally get a convolution kernel K∈S*S. Then apply the deconvolution method to it, expand it in the time dimension and channel (different sensor data in the same time slice) dimension to expand the information source of the data, that is, map the original distribution of the data to a new high-dimensional distribution, and perform the deconvolution operation according to the following formula:
[0058] m D,i,j,k =μ(xi,j )=w D,j,k *x i,j +b D,j,k (3)
[0059] Among them, the mass function m D,i,j,k It is generated by a fuzzy membership function and is used to describe the relationship between data at a specific location and certain specific conditions (such as time points, channels, events). i,j ) is a function that maps input data to the range [0,1] and is used to describe an input x i,j The degree of membership in a certain set. D corresponds to the input data (X Norm ,t), i,j are the subscripts of the time dimension and channel dimension respectively, parameters w and b represent the slope and intercept of the membership function, k corresponds to the event dimension, that is, the convolution kernel size K;
[0060] Step 4: Perform the same deconvolution operation on the two-dimensional distributions in step 3. T and m C , all perform similar expectation calculation operations, which are performed by weighted summation of all possible time points and states according to the following formula;
[0061]
[0062] Among them, formula (4) and formula (5) are for Y k (kth generated variable) and each of its channels c, calculate its sliding window statistics ( Represents the tensor Y k The mean and standard deviation of channel c in the time window [t-W+1,t]) form a new variable L k,c , which is used to calculate mutual information in the subsequent step. In formula (4) and formula (5), W is the sliding window size (number of time steps), T is the total number of time steps, t is the time step of a single window, b represents the number of batches, and i is the accumulated variable in the summation calculation. Formula (6) is the key to this step and is used to calculate Y k The time-domain nonlinear mutual information with the target variable RUL, where is the joint probability density of sensor window statistics and remaining service life, p(rul t) is the marginal probability density of the sensor window statistics and the remaining life. In this way, the correlation between the two variables can be measured based on the result of the logarithmic term. If there is no correlation at all, the logarithm is 0; γ is the time decay coefficient, which controls the decay rate of the weight over time (γ>0, by adjusting γ to control the weight decay curve to dynamically focus on the key stage of the final degradation stage); α is the nonlinear sensitivity parameter, and the size of a is adjusted (greater than 1, strengthening the correlation; less than 1, weakening the correlation) to achieve the purpose of enhancing the adaptability to complex degradation patterns. Finally, formula (7) and formula (8) perform weight calculation and feature fusion based on mutual information to obtain a comprehensive feature representation Y′ for each. The above two steps can be seen in the attached Figure 2 Schematic diagram of the model structure;
[0063] Step 5: For neural network model training, preliminary preprocessing is required. Specifically, for the input Y′, input embedding and position encoding are performed. The input sequence is first converted into a matrix of appropriate size for the model through the embedding layer, and then the position encoding is added to the embedding vector.
[0064] Step 6. To build a formal training model, you need to stack multiple encoder layers. First, build a multi-head self-attention layer. Specifically, attention generates new representations by calculating the relationship between each input and all other inputs. In each encoder layer, attention calculation is performed by calculating the query, key, and value. The query, key, and value are obtained from the output of the previous layer through linear transformation;
[0065] Step 7: Apply residual connection and normalization in each sub-layer to ensure stable model training. Specifically, the residual connection means that the input of each sub-layer (such as the multi-head self-attention layer and the feedforward network layer) is added to the output of the layer and then passed through the normalization layer;
[0066] Step 8. After the multi-head self-attention layer in step 3.2, the encoder also includes a feed-forward neural network, which includes two linear layers and a Swish activation function.
[0067] Step 9: Apply the residual connection and normalization in step 7 again to process the output of step 8, use the encoder described in steps 3.2 to 3.5, and then repeatedly stack N encoder layers to enhance the abstract ability of features layer by layer, and finally output a context-related representation, as shown in the attached figure. Figure 3 The encoder layer structure diagram is shown and can be expressed by the following formula:
[0068] Y l =MSA(Linear(Z l ))+Z l (9)
[0069] Zl+1 =SwiGLU(Linear(Y l ))+Y l (10)
[0070] Where l represents the number of layers of the current encoder (in Transformer, multiple encoders are stacked to allow the model to fully learn), and l+1 represents the next layer. In the above formula, if l+1 is the current layer, then Z l The output of the previous layer is used as the input of the current layer. After linear transformation and MSA processing, it is finally added to the output of SiLU to obtain the output Z of the current layer. l The Linear() in the above formulas represents the linear transformation process, and the MSA() function represents multi-head attention, which can be expressed as:
[0071]
[0072] Among them, Q is the query matrix, K is the key matrix, and V is the value matrix. is the scaled dot product, d k is the dimension of the key, used for scaling. The softmax operation converts a set of values into a probability distribution for calculating attention weights;
[0073] Step 10. On the corresponding computer software, load the training data set in step 1 and step 2, and initialize the weights of the Transformer encoder (the values of various preset parameters). Finally, perform a large number of batch training, and repeat the following operations in each round of training: obtain input data in batches; forward propagate to calculate output; calculate loss (parameters that can reflect model performance, that is, the error with real data); back propagate to update parameters (weights). End all operations until the training converges (the loss no longer changes significantly). The above is as attached. Figure 1 The execution flow diagram is shown in the figure.
[0074] After completing the above steps, the method tested the revenue results on the aircraft engine data set with the number of vehicles in the system under four different working conditions, and compared it with the existing advanced prediction model. The overall error results are shown in Table 1. The results show that the method generally exhibits good RUL prediction effect on the data set, verifying the accuracy and robustness of the present invention.
[0075] Table 1
[0076] method FD001 FD002 FD003 FD004 average BiLSTM (2018) 13.65 23.18 13.74 24.86 18.85 DCNN (2020) 12.61 22.36 12.64 23.31 17.73 MHT(2024) 11.92 13.7 10.63 17.73 13.5 DAA(2022) 12.25 17.08 13.39 19.86 15.65 BiGRU-AS(2021) 13.68 20.81 15.53 27.31 19.33 MSDCNN-LSTM (2023) 12.96 18.7 11.78 21.57 16.25 DSAN(2022) 13.4 22.06 15.12 21.03 17.9 IMDSSN(2023) 12.14 17.4 12.35 19.78 15.42 This method 12.08 18.09 12.56 25.12 16.96
[0077] It should be known that the parts not elaborated in detail in this specification belong to the prior art. Relevant technical personnel should understand that the above embodiments are only to help readers understand the principles and implementation methods of the present invention, and the scope of protection of the present invention is not limited to such embodiments. All equivalent replacements made on the basis of the present invention are within the protection scope of the rights of the present invention.
Claims
1. A method for predicting RUL of industrial equipment based on deconvolution and Transformer, characterized in that: The steps include: Step 1: Set up different types of sensors around the industrial equipment and its components in the application scenario, and arrange the collected data according to time steps to produce a multivariate sensor time series data set; Step 2: After normalizing the multivariate sensor time series data set, apply deconvolution and feature fusion method based on time domain nonlinear mutual information, and then form input data distribution Y′ after denormalization; Step 3: Input the input data distribution Y′ into the Transformer model, pass through the encoding layer and the output layer, and finally obtain the remaining useful life prediction model for industrial equipment under this working condition, and output the RUL prediction value.
2. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 1 is characterized in that: The step 1 is specifically implemented as follows: different types of sensors are set around the industrial equipment involved in the application scenario and its various components, the collected data are arranged according to time steps to produce a multivariate sensor time series data set, and the remaining service life RUL of each time slice is obtained by simulating the physical degradation model, which is used as the label of each data set, and the characteristics of each data set include the sensor data of the engine; The multivariate sensor time series data set is expressed by the following formula: RUL=f(X t ,t) (1) Where t is the timestamp, X t represents the input features at time t, which are signals from multiple sensors, RUL is the remaining useful life, which is the predicted output of the physical degradation model and serves as the label of each data; x i (t) represents the measurement value of the i-th sensor at time t, where n is the number of features.
3. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 2 is characterized in that: The specific implementation process of step 2 is as follows: 2.1 For each (X t ,t) is normalized to get (X Norm ,t); 2.2 After obtaining X Norm Based on this, we can transpose the matrix and get Then, put X Norm As input in the time dimension, X Norm As input of channel dimension; 2.3 Set the size S and parameters w,b of the convolution kernel to obtain a convolution kernel K∈S*S; 2.4 Implementing deconvolution data mapping: X Norm and Use convolution kernel K to perform transposed convolution at the same time to obtain the output data distribution Y T and Y C ; 2.5 Realizing feature fusion based on time-domain nonlinear mutual information: for Y T and Y C , all use the same feature fusion operation: First, the statistic L in the distribution is calculated in units of sliding windows, which represents the data distribution obtained by deconvolution. Then, the remaining useful life of the target variable and the relationship between each L are calculated. k The mutual information of the statistic is used to generate Y T and Y C The weight of each tensor in is calculated and feature fusion is performed; finally, the results of the time dimension T and the channel dimension C are summarized, that is, the distribution of the two is added together; the calculation method of mutual information is time-domain weighted and nonlinearly enhanced for time-sensitive scenarios. The former gives higher weight to recent data through the time attenuation factor γ, focusing on the key stage at the end of device degradation; the latter enhances the ability to capture nonlinear correlations through power transformation; After denormalization, the final data distribution Y′ is obtained.
4. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 3 is characterized in that: The expression of the deconvolution is as follows: m D,i,j,k =μ(x i,j )=w D,j,k *x i,j +b D,j,k (3) Among them, the mass function m D,i,j,k It is generated by the fuzzy membership function, which describes the relationship between the location data and the conditions; the fuzzy membership function μ(x i,j ) is a function that maps input data to the range [0,1] and is used to describe an input x i,j The degree of membership in the set; D corresponds to the input data (X Norm ,t), i and j are the subscripts of the time dimension and channel dimension respectively, and the parameter W D,j,k and b D,j,k represents the slope and intercept of the membership function, and k corresponds to the event dimension, that is, the convolution kernel size S*S in step 2.
3.
5. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 4 is characterized in that: The feature fusion method based on time domain nonlinear mutual information is expressed by the following formula: Among them, formula (4) and formula (5) are for the kth generated variable Y k And each channel c, calculate its sliding window statistics to form a new variable L k,c , Represents the tensor Y k The mean and standard deviation of channel c in the time window [t-W+1,t], where W in formula (4) and formula (5) is the sliding window size, T is the total number of time steps, t is the time step of a single window, b represents the number of batches, and i is the accumulated variable in the summation calculation; Formula (6) calculates Y k The time-domain nonlinear mutual information with the target variable RUL, where is the joint probability density of sensor window statistics and remaining service life, p(rul t ) is the marginal probability density of the sensor window statistic and the remaining life. In this way, the correlation between the two variables is measured based on the result of the logarithmic term. If they are completely unrelated, the logarithm is 0; γ is the time decay coefficient, which controls the decay rate of the weight over time; α is the nonlinear sensitivity parameter. Formulas (7) and (8) perform weight calculation and feature fusion based on mutual information to obtain a comprehensive feature representation Y′ for each.
6. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 5, characterized in that: The specific implementation process of step 3 is as follows: 3.1 For input Y′, input embedding and position encoding are performed. The input Y′ is first converted into a matrix through the embedding layer, and then the position encoding is added to the matrix; the position encoding provides the position information of each feature element in Y′ in the sequence; 3.2 Build an encoder layer, where multi-head self-attention generates new representations by calculating the relationship between each input and all other inputs; in each encoder layer, attention calculation is performed, where the query, key, and value are obtained from the output of the previous layer through linear transformation; 3.3 The output of step 3.2, apply residual connection and normalization in each sub-layer. The residual connection is to add the input of each sub-layer to the output of the layer, and finally pass through the normalization layer; 3.4 After the multi-head self-attention layer, the encoder also includes a feed-forward neural network, which includes two linear layers and a Swish activation function; 3.5 Apply the residual connection and normalization in step 3.3 again to process the output of step 3.4; 3.6 Use steps 3.2 to 3.5 as the encoder, and then repeatedly stack N encoder layers to finally output a context-dependent representation; 3.7 Perform back-propagation training through loss function.
7. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 6, characterized in that: The formula of the feedforward neural network in step 3.4 is as follows: Swish β (x)=x*σ(β*x) (9) Where, Swish in formula (9) is the activation function, and σ(x) is the Sigmoid function; β is a trainable parameter, which is usually optimized during training. x1 and x2 are two parts of the input vector, which are obtained by dividing the input into two halves. x1 is processed by the Swish activation function, and x2 is directly multiplied by the activated x1. W1 and W2 are learned weight matrices, and b1 and b2 are bias terms.
8. The industrial equipment RUL prediction method based on deconvolution and Transformer according to claim 7, characterized in that: The update expression of the encoder layer in step 3.2 is as follows: Y l =MSA(Linear(Z l ))+Z l (11) Z l+1 =SwiGLU(Linear(Y l ))+Y l (12) Where l represents the number of layers of the current encoder, and l+1 represents the next layer; in the above formula (12), if l+1 is the current layer, then Z l The output of the previous layer is used as the input of the current layer. After linear transformation and MSA processing, it is finally added to the output of SiLU to obtain the output Z of the current layer. l+1 ; The Linear() in the above formulas represents linear transformation processing, and the MSA() function represents multi-head attention.