Rolling bearing fault detection method based on improved Transform-BiGRU

By improving the Transformer-BiGRU network, the problems of insufficient feature extraction and gradient vanishing in the processing of complex sequence data are solved, and high-precision rolling bearing fault detection is achieved.

CN120369328APending Publication Date: 2025-07-25CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760068.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When processing complex sequence data, existing models lack feature extraction capabilities, making it difficult to capture the global dependence and local timing dynamic characteristics of vibration signals at the same time. In addition, deep networks are prone to gradient disappearance problems, and classification accuracy needs to be improved.

Method used

Build an improved Transformer-BiGRU network, and improve the Transformer encoder and BiGRU network, and use technical means such as reshaping layer, position encoding layer, multi-head attention layer, LayerNorm, and combine residual connection and multi-head attention layer feature fusion to improve feature extraction capabilities and gradient transmission.

Benefits of technology

The gradient vanishing problem is effectively solved, the feature extraction accuracy and classification accuracy are improved, and the test accuracy of 99.99% is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120369328A_ABST
    Figure CN120369328A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of rolling bearings, in particular to a rolling bearing fault detection method based on an improved Transform-BiGRU, and the method comprises the steps: obtaining the vibration data of a bearing; the method comprises the following steps of: constructing an improved Transform-BiGRU network, wherein the improved Transform-BiGRU network comprises an improved Transform encoder and an improved BiGRU network, and constructing the improved Transform-BiGRU network, wherein the improved Transform-BiGRU network comprises the improved Transform encoder and the improved BiGRU network; the improved Transform encoder comprises a remodeling layer, a position encoding layer, a multi-head attention layer, a first LayerNorm, a feed-forward layer and a second LayerNorm, wherein an output feature vector of the position encoding layer is in residual connection with the second LayerNorm. The method solves the problems that when an existing model processes complex sequence data, the feature extraction capacity is insufficient, and a single model is difficult to capture the global dependency relationship and the local time sequence dynamic characteristics of vibration signals at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rolling bearings, and in particular to a rolling bearing fault detection method based on an improved Transformer-BiGRU. Background Art

[0002] Once a rolling bearing fails, it will not only interfere with the normal operation of the equipment, but may also cause equipment damage, economic losses, and even endanger personal safety.

[0003] Existing methods, such as a bearing fault diagnosis method combining PCA and SVM, reduce the dimensionality of the high-dimensional features of the signal through PCA, and then use SVM for fault mode recognition; there is also a method of sparse coding using a discriminative dictionary.

[0004] There is also research on using autoencoders to extract features from multi-sensor data and input them into a deep belief network (DBN). Empirical mode decomposition (EMD) and Hilbert Huang transform (HHT) are also used to calculate the instantaneous frequency peak, and fault classification is performed through an ANN with multiple hidden layers.

[0005] However, when dealing with complex sequence data, the above models have the following problems:

[0006] 1. Insufficient feature extraction ability;

[0007] 2. The limitation of a single model architecture is high. It is difficult for a single model to simultaneously capture the global dependence and local temporal dynamic characteristics of vibration signals, and the deep network is prone to the problem of gradient disappearance;

[0008] 3. The classification accuracy needs to be improved: in some complex scenarios, the classification accuracy of the model still needs to be further optimized. Summary of the Invention

[0009] Aiming at the deficiencies of the existing methods, the present invention solves the problems that the processing accuracy and processing efficiency of the existing models need to be further improved.

[0010] The technical solution adopted by the present invention is: the rolling bearing fault detection method based on the improved Transformer-BiGRU includes the following steps:

[0011] Step 1: Obtain bearing vibration data;

[0012] As a preferred embodiment of the present invention, the bearing vibration data includes: the CWRU dataset.

[0013] As a preferred embodiment of the present invention, the bearing vibration data is cut into overlapping samples with a preset overlap rate.

[0014] As a preferred embodiment of the present invention, the bearing vibration data is decomposed by VMD to obtain the highest frequency component, the second highest frequency component, the medium-low frequency component, and the lowest frequency component.

[0015] Step 2: Construct an improved Transformer-BiGRU network, which includes an improved Transformer encoder and an improved BiGRU network; the improved Transformer encoder includes: a reshaping layer, a position encoding layer, a multi-head attention layer, a first LayerNorm, a feed-forward layer, and a second LayerNorm, and the output feature vector of the position encoding layer is connected with the second LayerNorm by residual connection.

[0016] As a preferred embodiment of the present invention, the number of improved Transformer encoders is 3.

[0017] As a preferred embodiment of the present invention, the improved BiGRU network includes: a first BiGRU layer, a second BiGRU layer, a third BiGRU layer, a second multi-head attention layer, a normalization layer, a fully connected layer, and an output layer; the time-step features and average pooling features output by the second multi-head attention layer are fused by addition feature fusion.

[0018] As a preferred embodiment of the present invention, the number of the second multi-head attention layers is 4.

[0019] As a preferred embodiment of the present invention, the cross-entropy loss function and the Adam optimizer are used during the training of the improved Transformer-BiGRU network.

[0020] As a preferred embodiment of the present invention, a rolling bearing fault detection system based on the improved Transformer-BiGRU includes: a memory for storing instructions executable by a processor; a processor for executing the instructions to implement a rolling bearing fault detection method based on the improved Transformer-BiGRU.

[0021] As a preferred embodiment of the present invention, a computer-readable medium storing computer program code, and the computer program code implements a rolling bearing fault detection method based on the improved Transformer-BiGRU when executed by a processor.

[0022] The beneficial effects of the present invention:

[0023] 1. The present invention improves the residual design of the Transformer encoder, effectively solves the problem of gradient disappearance existing in the existing Transformer encoder, and maintains the transmission of feature information;

[0024] 2. The present invention improves the BiGRU network, effectively captures the correlations at different positions in the sequence by using the multi-head self-attention mechanism layer, and improves the accuracy of feature extraction;

[0025] 3. Additively fuse the time-step features and average pooling features of the multi-head attention layer of the improved BiGRU network to achieve an organic combination of feature extraction and sequence modeling. Description of the Drawings

[0026] Figure 1 is the flowchart of the rolling bearing fault detection method based on the improved Transformer-BiGRU of the present invention;

[0027] Figure 2 is the curve graph of the loss and accuracy of the training and validation sets of the present invention;

[0028] Figure 3 is the feature distribution diagram before and after t-SNE dimensionality reduction of the present invention;

[0029] Figure 4 is the confusion matrix of the classification results of the CWRU data set of the present invention. Detailed Embodiments

[0030] The present invention will be further described below in conjunction with the drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.

[0031] As Figure 1 shown, the rolling bearing fault detection method based on the improved Transformer-BiGRU includes the following steps:

[0032] Step 1: Obtain bearing vibration data;

[0033] The bearing vibration data can be self-collected or obtained through a public data set. For example, a data set containing various bearing vibration information can be downloaded from the CWRU bearing data website. The data set is a mat file, and the training set, validation set, and test set are randomly divided in a ratio of 7:2:1 for model training, validation, and testing;

[0034] Data preprocessing: First, organize the original signal data into a table, with each column representing a bearing state. Each column of the original data has approximately 120,000 points. Cut the data into multiple overlapping samples (each sample contains 1024 points), and the overlapping rate of adjacent samples is 50%; after converting the data from the table format to the tensor format usable by PyTorch;

[0035] Each sample is processed by VMD (Variational Mode Decomposition), and the complex signal is decomposed into 4 simple components, including the highest frequency component of IMF1 (2 - 4 kHz), the second highest frequency component of IMF2 (1 - 2 kHz), the medium and low frequency component of IMF3 (500 - 1000 Hz), and the lowest frequency component of IMF4 (0 - 500 Hz).

[0036] Step 2: Construct an improved Transformer - BiGRU network, and the improved Transformer - BiGRU network includes an improved Transformer encoder and an improved BiGRU network;

[0037] The improved Transformer encoder includes: a reshaping layer, a position encoding layer, a multi - head attention layer, a first LayerNorm, a feed - forward layer, and a second LayerNorm. The output feature vector of the position encoding layer is connected with the second LayerNorm through a residual connection.

[0038] Preferably, the number of Transformer encoders is 3.

[0039] Each sub - layer of the existing single Transformer encoder uses a residual connection (Residual Connection) and layer normalization (LayerNorm), that is, 2 times of residuals are performed, while the single Transformer encoder of the present invention only performs one residual. This design of the residual connection can effectively solve the problem of gradient disappearance in deep networks and maintain the transmission of feature information.

[0040] LayerOutput = LayerNorm(x + FFN(LayerNorm(x + MultiHeadAttn(x))));

[0041] Among them, MultiHeadAttn is the multi - head attention layer, which contains 8 attention heads; FFN is the feed - forward network, and the hidden dimension is extended to 256;

[0042] The output feature vector of the improved Transformer encoder is used as the input of the improved BiGRU network.

[0043] Preferably, the improved BiGRU network is used as the decoder. The improved BiGRU network includes: a first BiGRU layer, a second BiGRU layer, a third BiGRU layer, a second multi - head attention layer, a normalization layer, a fully - connected layer, and an output layer; the time - step features and average - pooling features output by the second multi - head attention layer are fused by addition feature fusion;

[0044] The number of attention heads of the second multi - head attention layer is 4.

[0045] In the existing BiGRU network, the normalization layer is directly connected after the third BiGRU layer. In the present invention, a multi-head attention layer and a feature fusion layer are incorporated between the third BiGRU layer and the normalization layer; the multi-head attention layer is used for feature extraction to effectively capture the correlations at different positions in the sequence and improve the accuracy of feature extraction; then, innovatively, the last time step feature and the average pooling feature are combined, and in an additive fusion manner, using a similar residual connection idea, the last time step feature and the average pooling feature are respectively taken for additive feature fusion.

[0046] The specific process is as follows:

[0047] First, the input data [32, 4, 1024] is reshaped into [32, 64, 64]. Among them, there are 32 samples, each sample has 4 channels, and each channel has 1024 points. It is reshaped to better suit the Transformer encoder; after the input passes through the multi-head self-attention and the encoder, a residual connection is made with the original input, and the output shape is still: [32, 64, 64];

[0048] The number of multi-head self-attention is 8, which improves the parallel processing ability and enhances the model's ability to capture different attention information.

[0049] The feature vector output by the Transformer encoder is input into the improved BiGRU network. The BiGRU layer of the improved BiGRU network consists of two independent GRUs: one processes the sequence in the forward direction of time, and the other processes it in the reverse direction;

[0050] Forward GRU: Starting from the first element of the sequence, it processes to the last element in turn.

[0051] →h_t = GRU_forward(x_t, →h_{t - 1});

[0052] Reverse GRU: Starting from the last element of the sequence, it processes to the first element in turn.

[0053] ←h_t = GRU_backward(x_t, ←h_{t + 1});

[0054] Combined output: At each time step, the output of the BiGRU layer is the concatenation of the hidden states of the forward and reverse GRUs, and the output at each time step is the splicing of the outputs of the two-direction GRUs.

[0055] h_t = [→h_t, ←h_t];

[0056] The number of BiGRU layers is 3, and the hidden layer sizes are set to: [64, 128, 192]. The number of neurons in each layer increases to extract more complex sequence features and temporal dependencies. The first BiGRU layer: input [32, 64, 64] → output [32, 64, 128] (since it is bidirectional, the output dimension doubles); the second BiGRU layer: input [32, 64, 128] → output [32, 64, 256]; the third BiGRU layer: input [32, 64, 256] → output [32, 64, 384].

[0057] After the output [32, 64, 384] of the third BiGRU layer, a multi-head self-attention layer with 4 heads is added for feature extraction. The output of additive feature fusion is [32, 384], which combines the final state of the sequence and the overall average information, and finally classifies through a fully connected classifier.

[0058] The output dimension of the fully connected classifier: 10, applicable to 10-class classification tasks.

[0059] During model training, the cross-entropy loss function and Adam optimizer are adopted. The learning rate is set to 0.0002, and the weight decay is 1e-5. Backpropagation is used to update the weights.

[0060] Learning rate scheduler: ReduceLROnPlateau, patience 5, decay factor 0.5; number of training epochs: 40 epochs. Load the best model parameters recorded during training, and finally complete bearing fault classification. Combine the learning rate scheduler and L2 regularization to prevent model overfitting. Adopt the early stopping strategy and learning rate scheduler to dynamically adjust training parameters to ensure model convergence.

[0061] To verify the effectiveness of the method proposed in the present invention, experiments were carried out on the CWRU dataset. According to the experimental records, the test accuracy is stable at 99.99%.

[0062] Table 1 shows the comparison of measurement accuracies between the method of the present invention and existing conventional methods;

[0063] Table 1 Comparison experiment results on the CWRU dataset

[0064]

[0065] Table 2 Model ablation experiment results

[0066]

[0067] As Figure 2 It can be seen that the loss and accuracy of the method of the present invention converge rapidly with the increase of the number of iterations, indicating the effectiveness of the method of the present invention;

[0068] Figure 3After performing t-SNE dimensionality reduction on the VMD feature decomposition, it can be seen that the sample points corresponding to different modes (IMFs) after VMD decomposition form clear groupings on the 2D plane, indicating that the VMD decomposition is effective and each mode is distinguishable.

[0069] Figure 4 This is the confusion matrix of the classification results of the CWRU dataset of the present invention. It can be seen that the confusion matrix shows diagonal elements, indicating clear classification between classes;

[0070] Table 2 presents the ablation experiment of the model of the present invention, and it can be seen that the improvement effect of the method of the present invention is significant. Taking the ideal embodiments of the present invention as the inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A rolling bearing fault detection method based on an improved Transformer-BiGRU, characterized in that It includes the following steps: Step 1, obtain bearing vibration data; Step 2, construct an improved Transformer-BiGRU network, and the improved Transformer-BiGRU network includes an improved Transformer encoder and an improved BiGRU network; The improved Transformer encoder includes: a reshaping layer, a position encoding layer, a multi-head attention layer, a first LayerNorm, a feed-forward layer, a second LayerNorm, and the output feature vector of the position encoding layer is connected in residual with the second LayerNorm.

2. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 1, characterized in that The improved BiGRU network includes: a first BiGRU layer, a second BiGRU layer, a third BiGRU layer, a second multi-head attention layer, a normalization layer, a fully connected layer, and an output layer; the time-step features and average pooling features output by the second multi-head attention layer are fused by addition.

3. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 1, wherein The bearing vibration data includes: the CWRU dataset.

4. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 1, wherein The bearing vibration data is cut into overlapping samples with a preset overlap rate.

5. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 1, wherein The bearing vibration data is decomposed by VMD to obtain the highest-frequency component, the second-highest-frequency component, the medium-low-frequency component, and the lowest-frequency component.

6. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 1, characterized in that, The number of improved Transformer encoders is 3.

7. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 2, wherein The number of the second multi-head attention layers is 4.

8. The rolling bearing fault detection method based on the improved Transformer-BiGRU according to claim 1, wherein When training the improved Transformer-BiGRU network, a cross-entropy loss function and an Adam optimizer are adopted.

9. Rolling bearing fault detection system based on improved Transformer-BiGRU, characterized in that It includes: A memory for storing instructions executable by a processor; A processor for executing instructions to implement the rolling bearing fault detection method based on the improved Transformer-BiGRU as described in any one of claims 1-8.

10. A computer-readable medium storing computer program code, characterized in that, The computer program code implements the rolling bearing fault detection method based on the improved Transformer-BiGRU as described in any one of claims 1-8 when executed by the processor.