Multivariable bearing residual life prediction method and system
By using a bidirectional time convolution capsule network in the residual life prediction of rolling bearings combined with a multi-scale attention mechanism and capsule network, the shortcomings in the prediction accuracy and stability of the existing models are solved, and more accurate and efficient life prediction is achieved.
Patent Information
- Application Number
- CN202510110517.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing deep learning models have the disadvantages of inability to parallel computing, cumulative error, gradient problems and large memory usage in the remaining life prediction of rolling bearings. Moreover, traditional CNNs are not suitable for time series modeling. The time convolutional network ignores the utilization of future time features, limiting the improvement of its prediction performance.
Bidirectional time convolution capsule network (Bi-TCC) is used to combine multi-scale attention mechanisms and capsule networks to process time series data through forward and reverse branches, capture historical and future time characteristics, and enhance the understanding of long sequences through multi-scale attention mechanisms and self-attention mechanisms.
It improves the multivariable vibration signal feature capture capability of rolling bearings under different working conditions, improves the accuracy and stability of residual life prediction, and reduces gradient explosion and overfitting problems.
Smart Images

Figure CN120030302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent prediction technology, and more specifically to a multivariable bearing remaining life prediction method and system. Background Art
[0002] At present, with the increasing complexity of industrial equipment, the health management and remaining useful life (RUL) prediction of rolling bearings are crucial to ensuring system reliability, reducing maintenance costs, and optimizing enterprise operation strategies. Traditional RUL prediction methods based on physical or statistical models are difficult to accurately depict the equipment degradation process when facing actual complex working conditions. Data-driven methods, especially deep learning technology, have been widely used in the field of mechanical health monitoring because they can reveal the potential correlation between data and health status.
[0003] However, existing deep learning models such as RNN and its variants have the disadvantages of being unable to parallelize, accumulating errors, gradient problems, and large memory usage; although traditional CNNs can effectively extract features in processing high-dimensional data, they are generally considered unsuitable for time series modeling because their convolution kernel size limits the ability to capture long-term information. In addition, although temporal convolutional networks (TCNs) overcome some of the above challenges by adopting dilated causal convolutions and residual connections, and perform well in modeling long-term dependencies, they often ignore the use of future time features, which limits the further improvement of their prediction performance.
[0004] Therefore, developing a method that can effectively capture both historical and future time characteristics is an urgent problem that technicians in this field need to solve. Summary of the invention
[0005] In view of this, the present invention provides a multivariable bearing remaining life prediction method and system to overcome the above-mentioned defects.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] A multivariable bearing remaining life prediction method, the specific steps are:
[0008] Collect multiple operating status data of rotating machinery;
[0009] fusing a plurality of the operating status data to obtain two-dimensional time series data;
[0010] The two-dimensional time series data is input into a trained life prediction model, and the remaining service life of the rotating machinery is output.
[0011] Furthermore, the steps of acquiring the two-dimensional time series data are:
[0012] Normalizing the collected multiple operating status data;
[0013] The normalized running status data is spliced by channel to generate the two-dimensional time series data.
[0014] Furthermore, the life prediction model is constructed based on a bidirectional causal convolutional capsule network, including a forward branch and a reverse branch connected in parallel, wherein the forward branch includes a TCN module, a multi-scale attention mechanism and a capsule network connected in sequence; and the reverse branch includes the TCN module, the self-attention mechanism and the capsule network connected in sequence.
[0015] Further, the TCN modules of the forward branch and the reverse branch each include three stacked TCN blocks of different configurations.
[0016] Furthermore, the TCN block includes a first branch and a second branch of a residual connection, the first branch includes a plurality of consecutive dilated causal convolutional layers, each of the dilated causal convolutional layers is followed by a layer normalization and a Dropout layer; the second branch includes a one-dimensional convolutional layer.
[0017] Furthermore, the TCN block also includes a GeLU activation function, and the first branch and the second branch of the residual connection are connected to the GeLU activation function.
[0018] Furthermore, the working principle of the multi-scale attention mechanism is:
[0019] Extract features of different scales from input data;
[0020] Merge the extracted features of different scales to obtain merged features;
[0021] Calculating attention weights based on the combined features;
[0022] The input data is adjusted according to the attention weight to obtain adjusted data.
[0023] Furthermore, the output expression of the capsule network is:
[0024]
[0025] in,
[0026] u ji =W ij u i ;
[0027] In the formula, s j For advanced capsules; W ij is the weight matrix; ui For low-grade capsules; c ij The coupling coefficient determined iteratively for dynamic routing; Prediction vector for low-level capsules.
[0028] Furthermore, the life prediction model is evaluated by the mean absolute error, the root mean square error and the score function, wherein the expression of the score function is:
[0029]
[0030] in,
[0031] In the formula, ω 1 and ω 2 are all prediction weights; m is the number of time steps; n is the total number of time steps; A t is the accuracy scoring function; t is the difference between the predicted RUL and the actual RUL at time t.
[0032] A multivariable bearing remaining life prediction system, comprising:
[0033] A data acquisition module, used for collecting multiple operating status data of the rotating machinery;
[0034] A data processing module, used for fusing the plurality of operating status data to obtain two-dimensional time series data;
[0035] The life prediction module is used to input the two-dimensional time series data into a trained life prediction model and output the remaining service life of the rotating machinery.
[0036] It can be seen from the above technical solutions that the present invention discloses a multivariable bearing remaining life prediction method and system, which has the following beneficial effects compared with the prior art:
[0037] 1. By introducing the bidirectional temporal convolution capsule network and integrating the attention mechanism, the multivariate vibration signal characteristics of rolling bearings under different working conditions can be captured more accurately, thus improving the RUL prediction accuracy;
[0038] 2. The multi-scale attention mechanism combined with the self-attention mechanism enhances the ability to capture key complex information in long sequences, enabling the model to better understand and process subtle changes in time series data;
[0039] 3. The use of layer normalization and GeLU activation function not only helps to capture complex feature relationships, but also effectively prevents the gradient explosion phenomenon, alleviates the overfitting problem, and improves the stability and generalization ability of the model training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0041] Figure 1 A schematic diagram of the method provided by the present invention;
[0042] Figure 2 A schematic diagram of the fusion of multi-sensor data provided by the present invention;
[0043] Figure 3 A schematic diagram of the structure of the Bi-TCC network provided by the present invention;
[0044] Figure 4 A schematic diagram of a TCN block structure provided by the present invention;
[0045] Figure 5 A schematic diagram of the dilated causal convolution structure provided by the present invention;
[0046] Figure 6 A schematic diagram of the multi-scale attention mechanism structure provided by the present invention;
[0047] Figure 7 A schematic diagram of the main structure of the capsule network provided by the present invention;
[0048] Figure 8 A schematic diagram of the self-attention mechanism structure provided by the present invention;
[0049] FIG9(a) is a schematic diagram of a vibration signal in the horizontal direction when the A1_1 bearing in the PHM2012 data set provided by the present invention runs to a fault; FIG9(b) is a schematic diagram of a vibration signal in the vertical direction when the A2_1 bearing in the PHM2012 data set provided by the present invention runs to a fault;
[0050] FIG. 10( a ) is a schematic diagram of the prediction results of the A2_4 bearing in the PHM2012 data set provided by the present invention; FIG. 10( b ) is a schematic diagram of the prediction results of the A2_6 bearing in the PHM2012 data set provided by the present invention;
[0051] FIG11( a ) is a schematic diagram of a vibration signal in the horizontal direction when the A1_1 bearing in the XJTU-SY data set provided by the present invention runs to a fault; FIG11( b ) is a schematic diagram of a vibration signal in the vertical direction when the A2_1 bearing in the XJTU-SY data set provided by the present invention runs to a fault;
[0052] FIG. 12( a ) is a schematic diagram of the prediction results of the A1_4 bearing in the XJTU-SY data set provided by the present invention; FIG. 12( b ) is a schematic diagram of the prediction results of the A2_4 bearing in the XJTU-SY data set provided by the present invention;
[0053] FIG. 13( a ) is a schematic diagram of the results of the ablation method for the A1_1 bearing provided by the present invention; FIG. 13( b ) is a schematic diagram of the results of the ablation method for the B1_4 bearing provided by the present invention. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] On the one hand, an embodiment of the present invention discloses a multivariable bearing remaining life prediction method, such as Figure 1 As shown, the specific steps are:
[0056] Step 1: Collect multiple operating status data of the rotating machinery;
[0057] Step 2: Fusing multiple operation status data to obtain two-dimensional time series data;
[0058] Step 3: Input the two-dimensional time series data into the trained life prediction model and output the remaining service life of the rotating machinery.
[0059] In one embodiment, the steps of acquiring two-dimensional time series data are:
[0060] Normalizing the collected multiple operation status data;
[0061] The normalized running status data is spliced by channel to generate two-dimensional time series data.
[0062] Furthermore, if Figure 2 As shown in the figure, by integrating data from multiple sensors, more comprehensive and accurate bearing status information can be provided. By fusing sensor data, the network can not only analyze the temporal trend of a single sensor, but also reveal the relationship between multiple sensors. In addition, if a sensor fails or data is lost, the data of other sensors can be used as a backup to ensure the continuous operation of the system. The design of multi-sensor fusion provides the model with rich multi-dimensional information, enabling it to more comprehensively grasp the operating status of the bearing, specifically:
[0063] First, performance degradation data is collected from C sensors of M bearings, sampled at time intervals t. The data sample collected by the Cth sensor at time t is represented by X c,t ∈R H×1 , where H is the length of each sample. In order to reduce the difference in data distribution between different bearings, the raw data of each sensor is normalized to obtain the normalized data As shown in formula (1):
[0064]
[0065] In the formula, μ c and σ c are the mean and standard deviation of the Cth sensor data respectively.
[0066] The normalized data is spliced by channel to form multi-channel fusion data X m,t , as shown in formula (2):
[0067] X m,t ={x 1,t ,x 2,t ,...,x K,t} (2);
[0068] In the formula, X m,t ∈R H×K is the tth sample of the mth bearing.
[0069] In one embodiment, a life prediction model is constructed based on a bidirectional causal convolutional capsule network, including a forward branch and a reverse branch connected in parallel, wherein the forward branch includes a TCN module, a multi-scale attention mechanism, and a capsule network connected in sequence; and the reverse branch includes a TCN module, a self-attention mechanism, and a capsule network connected in sequence.
[0070] In one embodiment, the TCN modules of the forward branch and the reverse branch each include three stacked TCN blocks of different configurations.
[0071] In one embodiment, the TCN block includes a first branch and a second branch of a residual connection, the first branch includes a plurality of consecutive dilated causal convolutional layers, each of which is followed by a layer normalization and a Dropout layer; the second branch includes a one-dimensional convolutional layer.
[0072] Furthermore, the Bidirectional Causal Convolutional Capsule Network (Bi-TCC) combines the Temporal Convolutional Network (TCN), multi-scale attention mechanism, self-attention mechanism and Capsule Network to process serialized input data, such as Figure 3As shown in the figure. The input is a two-dimensional tensor containing multi-channel time series data from C sensors, and each channel contains a sequence of length H. Bi-TCC predicts the remaining useful life of time series through two parallel channels, processing forward and reverse time series inputs respectively. The forward time series input is processed through three stacked TCN blocks, a multi-head attention mechanism, and a one-dimensional capsule network to learn the feature information of the time series. The reverse time series input is processed through three stacked TCN blocks, a self-attention mechanism, and a one-dimensional capsule network, using the reverse time series to help the model capture backward dependencies, thereby learning the relationship between current values and future values. This bidirectional structure enables the network to more comprehensively understand the forward and backward dependencies of the time series, improving the accuracy of the remaining useful life prediction.
[0073] Furthermore, the TCN block is specifically designed to process time series data, such as Figure 4 As shown, the TCN block extracts features through two consecutive dilated causal convolutional layers, each followed by layer normalization (LN) and Dropout layers to stabilize the training process and reduce overfitting. To ensure that the dimensions of the input and output remain consistent, the network uses an optional 1x1 convolutional layer to implement residual connections. Finally, the input and convolutional layer outputs are added and processed through the GeLU activation function, which enhances the network's learning and generalization capabilities for time series data. It not only helps capture long-term dependencies in time series, but also effectively alleviates the gradient vanishing problem in deep networks through residual connections, allowing the network to better learn complex time series features.
[0074] Layer normalization reduces internal covariate shift by normalizing the input of each layer, thereby speeding up the training process of the model and improving its stability. Since layer normalization does not depend on batch data, it is very suitable for parallel computing and is robust to changes in the scale and distribution of input data. In addition, layer normalization helps improve gradient flow, reduces dependence on initialization methods, and improves the overall performance of the network when processing complex data. In layer normalization, for the input vector x, the mean μ and variance σ 2 As shown in formula (3):
[0075]
[0076] Where D is the dimension of vector x.
[0077] The normalization process is shown in formula (4):
[0078]
[0079] Here, ò is a small constant to avoid division by zero.
[0080] The input is scaled and translated as shown in formula (5):
[0081]
[0082] Where γ and β are learnable parameters used to restore the representation ability of the network.
[0083] In one embodiment, the TCN block further includes a GeLU activation function, and the first branch and the second branch of the residual connection are connected to the GeLU activation function.
[0084] Furthermore, the GeLU activation function is defined by the approximate error function (ERF), where x is the input value, and the GeLU activation function maps x to a nonlinear value in the context of a normal distribution, as shown in formula (6):
[0085]
[0086] When x is small, GeLU is close to 0; when x is large, GeLU approaches a linear mapping of x. This feature makes GeLU particularly effective when processing data with normal distribution characteristics. GeLU is continuous and smooth over the entire real number domain, which helps the gradient propagate more smoothly during training, thereby improving training efficiency and accelerating network convergence. GeLU can adaptively adjust the activation value according to the characteristics of the input, thereby helping the network to automatically learn important features. In addition, GeLU's randomness and adaptability to the input distribution help alleviate overfitting, thereby improving the performance of the model in complex tasks. In particular, the output of GeLU is close to a Gaussian distribution when the input is close to 0. This property helps to enhance the generalization ability of neural networks, enabling them to better adapt to different data distributions.
[0087] Among them, dilated causal convolution is a special convolution operation that combines the characteristics of causal convolution and dilated convolution. It is widely used in the processing of time series data, such as Figure 5 As shown in Figure 2. In time series analysis, causal convolution avoids information leakage by ensuring that the model only uses current and past information when making predictions, without relying on future data. However, one limitation of causal convolution is that its receptive field is small, that is, the input area covered by the convolution kernel is limited. In order to expand the receptive field, it is usually necessary to increase the number of convolution layers or increase the size of the convolution kernel, which will increase the complexity of the model.
[0088] The causal convolution is shown in formula (7):
[0089]
[0090] In the formula, y t is the output, wi is the weight of the convolution kernel, x t-i is the input signal, t is the current time point, k is the size of the convolution kernel, and i is the index of the current element in the convolution kernel.
[0091] To solve this problem, dilated convolution was introduced. Dilated convolution expands the receptive field by inserting gaps between convolution kernel elements without increasing the number of parameters or computational burden. By skipping part of the input data, dilated convolution enables the convolution kernel to act on an input area larger than its own length, thereby effectively expanding the receptive field of the model. Dilated convolution introduces the dilation rate d based on formula (7), as expressed in formula (8):
[0092]
[0093] In the formula, y t is the output, w i is the weight of the convolution kernel, x t-i is the input signal, t is the current time point, k is the size of the convolution kernel, i is the index of the current element in the convolution kernel; d represents the interval between elements in the convolution operation.
[0094] In the dilated causal convolution structure, the dilation rate of the convolution kernel usually increases exponentially with the increase of network depth, which significantly expands the receptive field while avoiding a significant increase in computational complexity. In this way, the dilated causal convolution can capture data dependencies in longer time series and ensure that the output depends only on the current and previous inputs and is not affected by future information, which is particularly critical for tasks such as time series prediction. In deep networks, the dilated causal convolution, with its unique structure, alleviates the problem of gradient vanishing or exploding in traditional convolutions. In addition, the dilated causal convolution allows the entire sequence to be processed in parallel, greatly improving computational efficiency. More importantly, the dilated causal convolution can expand the receptive field without adding additional parameters, which makes the model more efficient and effectively reduces the risk of overfitting.
[0095] In one embodiment, the working principle of the multi-scale attention mechanism is:
[0096] Extract features of different scales from input data;
[0097] Merge the extracted features of different scales to obtain merged features;
[0098] Calculate attention weights based on the merged features;
[0099] The input data is adjusted according to the attention weight to obtain adjusted data.
[0100] Furthermore, the Multi-Scale Attention Mechanism is Figure 6 As shown, the input data extracts features of different scales through multiple convolutional layers. In this embodiment, three different sizes of convolution kernels are used, corresponding to feature extraction of different scales. A convolution layer with a convolution kernel size of 1 is used to capture subtle local features, a convolution layer with a convolution kernel size of 3 captures medium-scale features, and a convolution layer with a convolution kernel size of 5 is used to extract features of a wider scale. In order to enhance the nonlinear expression ability and stability of the network, batch normalization (BatchNormalization) and ReLU activation function are introduced after the convolution layer. Subsequently, a fusion convolution layer is used to merge these features of different scales, and batch normalization and Sigmoid activation function are combined to calculate the attention weight. The network differentially weights the features of each scale to highlight the key parts of the input data. By multiplying the attention weight with the input, the importance of each position in the feature map is adjusted. Ultimately, the output result can enhance the relevant features according to the feature importance of each position, thereby more effectively capturing and utilizing the key information in the input data.
[0101] In one embodiment, a capsule network (CapsuleNetwork), such as Figure 7 As shown in the figure, the basic unit of the capsule network is the capsule, and the output of each capsule is a vector rather than a scalar. The length of this vector represents the probability of the entity's existence, while its direction encodes the specific attributes of the entity. In order to transmit information, the capsule network adopts a mechanism called dynamic routing. In this process, the path of information transmission between the low-level capsule and the high-level capsule is determined in an iterative manner. The weight of the transmission is determined by the "coupling coefficient", which is optimized through learning to ensure that the high-level capsule can be correctly activated. The design of the capsule network enables it to have equivariance and can recognize objects in different positions and postures. This feature is achieved through the direction of the capsule vector, because the direction of the vector encodes the posture information of the object.
[0102] uth i capsules are multiplied by an independent weight matrix W ij To predict each high-level capsule, as shown in formula (9). High-level capsule s j are all the prediction vectors of the lower-level capsules The weighted sum of advanced capsules j The final output vector v is obtained by the nonlinear Squash activation function j , as shown in formulas (10) and (11):
[0103] u ji =W ij ui (9);
[0104]
[0105] Each prediction vector is distributed to all high-level capsules by dynamic routing, where c ij is the coupling coefficient determined by dynamic routing iteration, which is determined by the softmax function, and the sum of the coupling coefficients is 1, as shown in formula (12). Its purpose is to enable the input neuron to autonomously select the best path to transmit to the next layer of capsules. ij The prediction vector is updated by dynamic routing Should be coupled to advanced capsules j The logarithmic prior probability of is shown in formula (13). The final capsule completes the forward propagation between the two capsule layers.
[0106]
[0107] b ij =b ij +v j u ji (13);
[0108] Because capsule networks can effectively encode pose information, they outperform traditional convolutional neural networks (CNNs) when recognizing objects with different poses. Capsule networks are also better at targeting the spatial relationship between features, which enables them to provide higher accuracy than CNNs when handling complex images and object recognition tasks. In addition, capsule networks are more robust to input data because capsules can capture the spatial properties of features, making the network more adaptable when dealing with transformations.
[0109] In one embodiment, the self-attention mechanism is a deep learning mechanism for processing sequence data, which enables the model to directly establish connections between different positions, thereby effectively capturing long-distance dependencies, such as Figure 8 As shown in the figure. The self-attention mechanism does not depend on the temporal order of the sequence during the calculation process, so it can process the entire sequence in parallel, which significantly improves the calculation efficiency. Thanks to the advantages of parallel computing, the self-attention mechanism avoids the common gradient vanishing or gradient exploding problems of RNN and LSTM, making the training of deep networks more stable. In addition, the self-attention mechanism improves the adaptability of the model to the input data by dynamically calculating weights, allowing the model to flexibly focus on different parts of the sequence when processing the sequence, thereby better capturing long-distance dependencies and complex patterns.
[0110] When X represents the input sequence, through the weight matrix W q , W k, W v The query (Q), key (K), and value (V) matrices are generated by linear transformation of , as shown in formula (14).
[0111]
[0112] Compute the dot product of the query matrix and the key matrix, and then divide by As shown in formula (15).
[0113]
[0114] The attention scores are normalized using the softmax function to obtain the attention weight of each element. The attention weights are applied to the value matrix and the weighted values are summed to obtain the output of the self-attention mechanism, as shown in formula (16).
[0115]
[0116] In the formula, is the dimension of the key vector, and V is the value matrix.
[0117] By simultaneously inputting the forward time series data and its reverse time series data, the network can capture the forward and reverse dependencies in the data, thereby more comprehensively understanding the changes in the time series. The parallel bidirectional approach not only enhances the network's expressiveness and generalization capabilities, but also reduces overfitting of the time series order. Taking into account the impact of the reverse time series, the network can better handle nonlinearity and complexity, while taking advantage of time symmetry and enhancing its robustness. In application scenarios such as predictive maintenance, this bidirectional learning strategy can detect potential signs of failure in advance and improve the accuracy of predictions. Finally, the network concatenates the outputs of the forward and reverse channels, flattens the concatenated results, and uses the ReLU activation function through the fully connected layer to predict the remaining service life. The network parameters are shown in Table 1.
[0118] Table 1 Network parameter configuration
[0119]
[0120] According to the above structure, the data processing steps include:
[0121] Forward: The forward pass first extracts features through three TCN blocks, each of which uses dilated causal convolution to capture temporal features at different scales while maintaining the causality of the data. The first TCN block uses 64 filters, a kernel size of 3, and a dilation rate of 1, aiming to capture local temporal features and maintain the causality of the time series. The second TCN block uses 32 filters, a kernel size of 3, and a dilation rate of 2 to further deepen the model's understanding of the time series. The third TCN block uses 16 filters, a kernel size of 3, and a dilation rate of 4, with increasing dilation rates to capture longer-term dependencies. The network then introduces a multi-scale attention mechanism, which processes time series data at different scales through 16 attention filters to enhance the ability to identify key features. The capsule network and digital capsule layer then further refine the features and help the network learn the spatial features of the data. The capsule layer contains 10 capsules, each with a dimension of 16, and performs 3 iterations of dynamic routing to effectively learn spatial features. Finally, the digital capsule layer further processes the output of the capsule layer to extract high-level features related to specific tasks, thereby improving the representation and prediction capabilities of the network.
[0122] Reverse: First, the input time series data is reversed so that the network can capture the feature information in the time series from back to front. Next, the features of the reverse sequence are extracted through the TCN block, and each TCN block uses a different configuration of dilated causal convolution to capture time series features of different scales. Each TCN block uses 32 filters, a convolution kernel size of 5, and dilation rates of 1, 2, and 4, respectively. Unlike the decreasing filter configuration used in the forward sequence, the configuration used in the reverse sequence can more fully extract its features while still maintaining the causal relationship of the sequence. Subsequently, the output of TCN is processed by the self-attention mechanism, which enables the network to focus on the key parts in the reverse sequence. Finally, the capsule layer and the digital capsule layer are used to refine and learn the data space features of the reverse sequence, providing rich feature representations for the reverse sequence. The parameter settings of these capsule layers and digital capsule layers are consistent with the forward input part to ensure the consistency and effectiveness of the network in the forward and reverse sequences.
[0123] Finally, the features of the reverse sequence are fused with the features of the forward sequence and processed through a fully connected layer to output the prediction result. The activation function uses ReLU. The entire network is constructed as an end-to-end time series analysis framework, which aims to simultaneously utilize the forward and reverse data of the time series to improve the network's understanding and prediction performance of time series data.
[0124] In one embodiment, the life prediction model is constructed based on the prediction framework of Bi-TCC, and the construction steps are:
[0125] Signal acquisition: different sensors are used to collect life data of rotating machinery;
[0126] Data fusion: Integrate data from multiple sensors and convert multiple one-dimensional time series data into two-dimensional time series data.
[0127] Life prediction: The training samples are input into the proposed Bi-TCC, and the time series features are learned and the service life of the rotating machinery is predicted.
[0128] Predictive analysis: predict the service life of bearings under different faults and conduct visual analysis based on usage.
[0129] Among them, the training samples are constructed by processing the data according to the above normalization steps. The specific steps are: all samples of the mth bearing are Where N is the total sampling time. The label of each sample is represented by y m,t , all sample labels are represented as {y m,t} t=1 . Construct the data set of the mth bearing For M bearings, construct the entire data set D = {D 1 ,D 2 ,...,D M}.
[0130] In order to verify the effectiveness and superiority of the proposed Bi-TCC method in the prediction of the remaining useful life (RUL) of rolling bearings, the IEEE PHM2012 dataset and the XJTU-SY dataset are used for experimental verification. The experimental environment is configured as follows: the CPU is i9-9900K, the GPU is NVIDIA RTX 4070Ti (video memory 12GB), the running memory is 32GB, and the experimental framework is based on Python's TensorFlow 2.5.0 and Keras.
[0131] During the model training process, the back-propagation algorithm was used and the Adam optimizer was used for gradient optimization. At the same time, in order to reduce overfitting, an L2 regularization term was added to the causal convolution layer. The learning rate was set to 0.001, and the weights were fine-tuned according to each parameter update to adjust the gradient of the loss function. In addition, the first decay rate was set to 0.9 and the second decay rate was set to 0.999 to calculate the exponentially weighted average of the gradient and its square to adapt to different learning rate changes. During the training process, the batch size was set to 128 and the number of traversals was 400.
[0132] In one embodiment, the life prediction model is evaluated by using mean absolute error, root mean square error, and score function to iteratively optimize the life prediction model.
[0133] In the online prediction process, the test data is input into the trained prediction network model for real-time prediction, and the prediction performance of the model is verified by setting evaluation indicators. The evaluation indicators include mean absolute error (MAE), root mean square error (RMSE), and a designed score function called "Score" to comprehensively evaluate the effect of the proposed method. In order to facilitate comparison with other related studies, MAE and RMSE are used as the main evaluation criteria, and their expressions are shown in formulas (17) and (18).
[0134]
[0135] In the formula, er t is the difference between the predicted RUL and the actual RUL at time t, that is n is the total life of the bearing, and lower MAE and RMSE represent better prediction results. In the model, the final score is calculated by a specific formula that takes into account the performance of the model at different stages of the entire time series, as shown in formula (20). Formula (19) is the accuracy score function, ω 1 and ω 2 Represent the prediction weights of the early stage and the late stage respectively. In this embodiment, set ω 1 :ω 2 The ratio is 2:3, which means that the prediction of the second half of the equipment life is more critical and is therefore given a higher weight. m is defined as half of the total number of time steps n, that is, The score is calculated in the range of 0 to 1. The higher the score, the better the prediction performance of the model. This scoring mechanism helps to identify and optimize the performance of the method at different time stages.
[0136]
[0137]
[0138] In order to verify the superiority of the proposed method, six advanced prediction methods were selected for comparison. These methods include: TCN-RSA, TCN-SA, threshold CNN, and Bi-TCN. TCN-RSA introduces a self-attention mechanism in traditional TCN, so that it can simultaneously learn the time-frequency information and spatial information of the vibration signal. TCN-SA combines the soft threshold attention mechanism with the prediction model of TCN, in which the soft threshold attention mechanism significantly improves the robustness of the model. Threshold CNN replaces the dilated causal convolution in TCN with standard convolution and enhances the prediction performance through the soft threshold attention mechanism. Bi-TCN processes time series data through two temporal convolutional networks (TCNs), one TCN is used to encode past covariates and the other is used to encode future covariates.
[0139] Experiment 1
[0140] The experimental data comes from the PRONOSTIA platform, which is used to collect data sets from normal operation to failure of bearings under different working conditions. The PRONOSTIA platform consists of three main parts: the rotating part, the degradation generating part (applying radial force on the bearing) and the measuring part, including NIDAQ acquisition card, pressure regulator, test bearing, coupling, reducer, AC motor and cylinder pressure measurement device, force sensor, accelerometer, temperature sensor, torque sensor, speed sensor, etc. The data of the platform covers the degradation process of ball bearings throughout their life until complete failure, and each degraded bearing almost contains various types of defects such as inner and outer rings, balls and cages. The rotating part is equipped with a 250W motor, which transmits rotational motion through a gearbox, and the rated speed of the motor is 2830 rpm. The load part consists of a pneumatic jack, which can apply a dynamic load of up to 4000N to the bearing. The test data mainly includes two types: vibration and temperature. The vibration data is collected by two micro accelerometers at 90° to each other with a sampling frequency of 25.6kHz; the temperature data is collected by a resistance temperature detector with a sampling frequency of 0.1Hz. The two accelerometers are placed in the horizontal and vertical directions of the bearing to capture the vibration signals in these two directions. Figure 9(a)-Figure 9(b) The vibration signals of two different bearings, A1_1 and A2_1, in the horizontal and vertical directions are shown. To ensure the safety of the experiment, when the amplitude of the vibration data exceeds 20g (1g = 9.8m / s 2 ), the experiment will automatically stop. The data collected in 0.1 seconds is divided into a sample, that is, the input data is x∈R 2560×2 .
[0141] In order to accelerate the degradation process of the bearing, the test bench applies radial load and controls it using a load regulator. In this experiment, the bearing degradation data under two working conditions in the PHM2012 dataset were selected. Each working condition contains data of 7 bearings, namely A1_1 to A1_7 and A2_1 to A2_7. Under each working condition, the degradation data of 6 bearings are randomly selected for training, and the degradation data of the remaining 1 bearing is used for testing. In addition, Table 2 lists the basic characteristics of the bearings, and Table 3 provides a detailed description of the dataset. Considering that bearings in different degradation states may have the same remaining useful life (RUL) value, in order to unify the processing, the experiment normalizes the actual RUL value of the bearing and uses the life percentage as the output label. The label is shown in formula (21):
[0142]
[0143] In the formula, is the output label, yt is the real RUL of the bearing at time t, and y is the RUL of the total time. In the figure, 100% represents a healthy bearing and 0% represents a failed bearing.
[0144] Table 2 Experimental bearing characteristics
[0145]
[0146] Table 3 Experimental data description
[0147]
[0148]
[0149] The prediction analysis of 14 bearing lifespans in the PHM 2012 data set was carried out, and the performance results of different models were obtained, as shown in Table 4. TCN-RSA performed best on the A2_4 bearing, but was inferior on the A2_7 bearing, with an overall average score of 0.69, which was a medium level. TCN-SA performed best on the A2_4 bearing, but the score on the A2_2 bearing was only 0.44, with an overall average score of 0.75, showing good stability. Threshold CNN performed well on the A1_4 bearing, but performed poorly on the A2_1 bearing, with an average score of 0.70, and a relatively balanced performance. Bi-TCN performed poorly on the A2_2 bearing, but achieved good results on the A1_6 bearing. When the data complexity increased, the performance of TCN-RSA and Threshold CNN on some data sets fluctuated greatly. The low score of Threshold CNN on most bearings shows that the performance of traditional convolutional networks in predicting the remaining service life is still lacking. Bi-TCC performs well on multiple data sets, especially on the two bearing data A1_1 and A2_6, where its MAE is significantly better than other methods, with the highest overall average score of 0.82, showing excellent prediction ability. Bi-TCC shows good stability on complex data and can maintain relatively stable performance under different test conditions. In summary, the effectiveness of the bidirectional temporal convolutional capsule network structure in such prediction tasks, Bi-TCC has demonstrated strong prediction ability with its consistent excellent performance on multiple data sets, especially in processing multi-sensor time series data and capturing long-term dependencies.
[0150] Figures 10(a) and 10(b) show the remaining useful life (RUL) prediction curves of bearings A2_4 and A2_6, respectively. The results show that the proposed Bi-TCC method can accurately predict the overall health status of the bearings. In order to present the prediction effect more clearly, the life prediction curves of bearing A2_4 in the early stage and bearing A2_6 in the later stage are enlarged respectively. Although the proposed method may have local oscillation in some cases, it can generally predict the state of the bearing in the later stage of its life more accurately.
[0151] Table 4 Performance comparison of different methods on the PHM2012 dataset
[0152]
[0153]
[0154] Experiment 2
[0155] In order to further verify the effectiveness and superiority of the proposed TCN-RSA model, the XJTU-SY dataset provided by Xi'an Jiaotong University was used. This dataset was collected in an accelerated degradation experiment and records the entire process of rolling bearings from normal operation to failure. The bearing test bench used in the experiment consists of an AC induction motor, a motor speed controller, a support shaft, two heavy-duty roller bearings, a hydraulic oil lubricator, and a bearing seat. In the experiment, a total of 15 LDK UER204 bearings were tested under three different working conditions. The radial force was generated by a hydraulic loading system and applied to the bearing housing, while the speed was set and kept stable by the speed controller of the AC induction motor. The specific specifications of the tested bearings are shown in Table 5. The failure types of the bearings are diverse, including inner raceway wear, cage fracture, outer raceway wear, and outer raceway fracture.
[0156] The test bench simulated different working conditions by applying different radial forces on the bearing seat of the test bearing. To verify the performance of the proposed method, the LDK UER204 rolling bearing data under two working conditions were selected for testing. Each working condition contains the full life cycle data of 5 bearings, namely B1_1 to B1_5 and B2_1 to B2_5, as shown in Table 6. Figure 11(a) and Figure 11(b) show the vibration signals of two different bearings, B1_1 and B2_1, in the horizontal and vertical directions, respectively. In the experiment, the multi-sensor fusion data of any 4 bearings were used as training sets, and the data of the remaining bearing were used as test sets. Two PCB 352C33 single-axis accelerometers were installed in the horizontal and vertical positions of the bearing seat to collect vibration signals. The sampling rate was set to 25.6kHz, and the vibration signal of 1.28 seconds per minute (i.e., 32,768 data points) was recorded, and the profile was described as x∈R 32768×2The label processing method used in this case is the same as that in Experiment 1.
[0157] Table 5 Experimental bearing characteristics
[0158]
[0159] Table 6 Description of experimental datasets
[0160]
[0161] When predicting the remaining useful life (RUL) of the XJTU-SY dataset, the proposed Bi-TCC uses the same parameters as those listed in Experiment 1, and uses RMSE, SMAPE, and score as performance evaluation indicators. The results are shown in Table 7. By performing life prediction analysis on 10 bearings in the XJTU-SY dataset, Bi-TCC shows low MAE and RMSE values on multiple bearings and obtains high scores, reflecting its excellent performance in prediction accuracy and reliability. The performance of Bi-TCC on each test bearing is relatively stable, with an average score of 0.79 and a small fluctuation, indicating that the method can maintain good robustness when facing different data sets and has a wide range of application potential. However, considering that different types of complex faults may occur in the bearing during the degradation process, the prediction method does not achieve the best RUL prediction results on every bearing. As a bidirectional network, Bi-TCC can effectively utilize the previous and next information in the time series data and enhance the ability of feature extraction and understanding of the degradation process. These characteristics make Bi-TCC very suitable for processing complex multi-sensor time series data. In most cases, the prediction effect of Bi-TCC is better than other methods. Figure 12(a) and Figure 12(b) are the RUL prediction results of bearings A1-4 and A2-4 respectively. The RUL prediction curve at the end of the full life is very close to the true value, and the predicted RUL tracks the true RUL well.
[0162] Table 7 Performance comparison of different methods on the XJTU-SY dataset
[0163]
[0164] On the two benchmark datasets PHM2012 and XJTU-SY, ablation experiments are conducted on the improved TCN Block proposed by Bi-TCC and the different attention mechanisms of its internal bidirectional channels.
[0165] Method 1: Directly use the unimproved TCN Block, which mainly uses layer normalization BatchNormalization and activation function ReLU.
[0166] Method 2: The attention mechanism in the forward and reverse channels of Bi-TCC all uses the multi-scale attention mechanism (MSA).
[0167] Method 3: The attention mechanism in the forward and reverse channels of Bi-TCC all uses the self-attention mechanism (SA).
[0168] Method 4: All attention mechanisms in the Bi-TCC bidirectional channel are eliminated.
[0169] According to the prediction results in Tables 8 and 9, the comprehensive evaluation results of Method 1 are only better than those of Method 4, indicating that the improved TCN Block can better capture complex feature relationships and effectively prevent the gradient explosion problem by using layer normalization and GeLU activation function. This improvement enhances the expression of positive activation features while suppressing negative activation features. In contrast, Method 4 performed the worst, highlighting the advantages of the attention mechanism in processing long sequence time series data. It can effectively capture long-distance dependencies, thereby avoiding the information loss problem in traditional methods. The attention mechanism also improves the efficiency and accuracy of the method in processing complex data by selectively focusing, dynamically assigning weights, and enhancing interpretability.
[0170] In addition, the comprehensive evaluation results of Method 2 and Method 3 are better than those of Method 1 and Method 4, which further verifies the significant role of the attention mechanism in improving the performance of the method. However, when using a single attention mechanism, the feature information of the reverse time series input may be partially coupled and confused with the forward time series features, resulting in overfitting of the model's prediction results in the later stages of the lifespan, which in turn affects the overall evaluation score and causes large fluctuations in the actual RUL prediction.
[0171] Figures 13(a) and 13(b) show the comparative radar charts of two benchmark datasets (bearings A1_1 and B1_4). It can be clearly seen that the prediction results of the proposed Bi-TCC show the lowest prediction error and the highest score, showing its effectiveness in this task. Although Method 2 performs well, it still fails to surpass Bi-TCC. Especially in complex data scenarios, it proves the key role of the attention mechanism in improving the performance of the method.
[0172] Table 8 Ablation experiment results of PHM2012 dataset
[0173]
[0174]
[0175] Table 9 Ablation experiment results of XJTU-SY dataset
[0176]
[0177] On the other hand, this embodiment discloses a multivariable bearing remaining life prediction system, including:
[0178] A data acquisition module, used for collecting multiple operating status data of the rotating machinery;
[0179] A data processing module is used to fuse multiple operating status data to obtain two-dimensional time series data;
[0180] The life prediction module is used to input two-dimensional time series data into the trained life prediction model and output the remaining service life of the rotating machinery.
[0181] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0182] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multivariable bearing remaining life prediction method, characterized in that: The specific steps are: Collect multiple operating status data of rotating machinery; fusing a plurality of the operating status data to obtain two-dimensional time series data; The two-dimensional time series data is input into a trained life prediction model, and the remaining service life of the rotating machinery is output.
2. A multivariable bearing remaining life prediction method according to claim 1, characterized in that: The steps for obtaining the two-dimensional time series data are as follows: Normalizing the collected multiple operating status data; The normalized running status data is spliced by channel to generate the two-dimensional time series data.
3. A multivariable bearing remaining life prediction method according to claim 1, characterized in that: The life prediction model is constructed based on a bidirectional causal convolutional capsule network, including a forward branch and a reverse branch connected in parallel, wherein the forward branch includes a TCN module, a multi-scale attention mechanism and a capsule network connected in sequence; and the reverse branch includes the TCN module, the self-attention mechanism and the capsule network connected in sequence.
4. A multivariable bearing remaining life prediction method according to claim 3, characterized in that: The TCN modules of the forward branch and the reverse branch each include three stacked TCN blocks of different configurations.
5. A multivariable bearing remaining life prediction method according to claim 4, characterized in that: The TCN block includes a first branch and a second branch of a residual connection, wherein the first branch includes a plurality of consecutive dilated causal convolutional layers, each of which is followed by a layer normalization and a Dropout layer; and the second branch includes a one-dimensional convolutional layer.
6. A multivariable bearing remaining life prediction method according to claim 4, characterized in that: The TCN block also includes a GeLU activation function, and the first branch and the second branch of the residual connection are connected to the GeLU activation function.
7. A multivariable bearing remaining life prediction method according to claim 3, characterized in that: The working principle of the multi-scale attention mechanism is: Extract features of different scales from input data; Merge the extracted features of different scales to obtain merged features; Calculating attention weights based on the combined features; The input data is adjusted according to the attention weight to obtain adjusted data.
8. A multivariable bearing remaining life prediction method according to claim 3, characterized in that: The output expression of the capsule network is: in, u j|i =W ij u i ; In the formula, s j For advanced capsules; W ij is the weight matrix; u i For low-grade capsules; c ij The coupling coefficient determined iteratively for dynamic routing; Prediction vector for low-level capsules.
9. A multivariable bearing remaining life prediction method according to claim 1, characterized in that: The life prediction model is evaluated by the mean absolute error, the root mean square error and the score function, wherein the expression of the score function is: in, Where ω1 and ω2 are prediction weights; m is the number of time steps; n is the total number of time steps; A t is the accuracy scoring function; t is the difference between the predicted RUL and the actual RUL at time t.
10. A multivariable bearing remaining life prediction system, characterized in that: include: A data acquisition module, used for collecting multiple operating status data of the rotating machinery; A data processing module, used for fusing the plurality of operating status data to obtain two-dimensional time series data; The life prediction module is used to input the two-dimensional time series data into a trained life prediction model and output the remaining service life of the rotating machinery.
Citation Information
Cited By
Electric drive transmission system comprehensive life prediction method based on AI
CN120373147A
AI-based comprehensive life prediction method for electric drive transmission systems
CN120373147B
Self-powered temperature difference vibration sensor energy management method based on capsule network
CN120951255A