CNN Rolling Bearing Fault Diagnosis Method Based on Improved GAF and SA
Through the improved GAF and SA modules, the one-dimensional vibration data is converted into two-dimensional images, and combined with CNN for rolling bearing fault diagnosis, the problem of degraded diagnostic performance in high-noise environments is solved, and fault classification with high accuracy is achieved.
Patent Information
- Application Number
- CN202210659742.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-13
AI Technical Summary
Traditional rolling bearing fault diagnosis technology has degraded diagnostic performance in high-noise environments, and relies on manual feature extraction and shallow learning, which lacks large-scale data processing capabilities.
The improved Gram Angle Field (GAF) is used to convert one-dimensional vibration data into two-dimensional images, and combined with the convolutional neural network (CNN) and self-attention (SA) modules, features are extracted through multi-layer convolution and pooling units, and SA attention mechanism is added to perform fault classification.
In high noise environment, the accuracy of rolling bearing fault diagnosis is significantly improved, the classification performance of the model is optimized, and automatic feature extraction and noise resistance are achieved.
Smart Images

Figure CN115221916B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of bearing fault diagnosis, and particularly to a CNN rolling bearing fault diagnosis method based on improved GAF and SA. Background Technique
[0002] As an important component of rotating equipment, rolling bearings often operate in high-load and strong-noise environments. Without necessary supervision and maintenance, it may cause bearing failure and equipment shutdown, resulting in significant economic losses. Therefore, it is very important to conduct fault diagnosis on rolling bearings at an early stage. Traditional fault diagnosis techniques mainly rely on signal processing methods, such as Fast Fourier Transform (FFT), Short-time Fourier Transform (STFT), wavelet transform, etc.; in recent years, with the development of machine learning, common signal processing methods are used to extract features and then fed into specific classifier models for classification, such as Support Vector Machine (SVM), artificial neural network, etc. However, such methods often rely on manual means to extract features and expert experience, and are often shallow learning, lacking the ability to process large-scale data. In order to get rid of the influence of manual extraction means on the results, some deep learning algorithms have begun to play a role in rolling bearing fault diagnosis. Such as convolutional neural network, autoencoder model, etc., but in the high-noise environment of actual bearing operation, the diagnostic performance also decreases to varying degrees. Summary of the Invention
[0003] This application provides a CNN rolling bearing fault diagnosis method based on improved GAF and SA, and its technical purpose is to improve the bearing fault diagnosis performance and the model diagnosis rate in a high-noise environment.
[0004] The above technical purpose of this application is achieved through the following technical solutions:
[0005] A CNN rolling bearing fault diagnosis method based on improved GAF and SA includes:
[0006] S1: Collect the vibration data of the rolling bearing, preprocess the vibration data to obtain a one-dimensional vibration data set;
[0007] S2: Convert the one-dimensional vibration data set into a two-dimensional image through improved GAF, and divide the two-dimensional image into a training set, a validation set, and a test set;
[0008] S3: Input both the training set and the validation set into the GAF-CNN model, set the initial parameters of the GAF-CNN model, train the GAF-CNN model with the training set. When the GAF-CNN model converges iteratively, complete the network training, save the network parameters, and obtain the trained GAF-CNN model. Input the validation set into the trained GAF-CNN model to verify the accuracy. If the accuracy of the validation set reaches the preset standard, go to step S4; otherwise, repeat step S3 until the accuracy of the validation set reaches the preset standard.
[0009] S4: Input the test set into the trained GAF-CNN model to obtain the test set accuracy and complete the fault diagnosis.
[0010] Among them, the GAF-CNN model includes a convolutional neural network and an SA attention module. The convolutional neural network includes four cascaded convolutional pooling modules and one fully connected layer. The SA attention module is arranged between the last convolutional pooling module and the fully connected layer.
[0011] Furthermore, the principle of improving the GAF (Gram Angular Field) method is as follows:
[0012] The Gram Angular Field is a method of encoding one-dimensional signals into images according to specific rules. Its special feature is that we represent the perturbation signal sequence in the polar coordinate system instead of the typical Cartesian coordinate system. By the polar coordinate method, the time dependence of the sequence is maintained, and the original time series can be restored through the polar coordinate mapping diagram.
[0013] The improved GAF is expressed as:
[0014]
[0015]
[0016] where x i represents the i-th value of the time series data. For the matrix element G (i,j||i-j=k ), when k = 0, the main diagonal G i,j is composed of the original values of the scaled time series.
[0017] During the actual operation of the rolling bearing, the amount of recorded data is quite large, and the amount of data increases significantly after being transformed by the above-mentioned Gram Angular Field method, which brings certain difficulties to model training. In order to reduce the amount of data, the Piecewise Aggregate Approximation method is used to reduce the dimension of the data. By transforming the data s with length n into data s' with length m and approximating the sequence segment with the mean value of the data elements included in this segment, the sequence length is shortened, the amount of data is greatly reduced, and the training time is reduced.
[0018] Further, in step S2, before converting the one-dimensional vibration data set into a two-dimensional image, the piecewise aggregate approximation method is used to reduce the dimension of the one-dimensional vibration data set first, and then the reduced one-dimensional vibration data set is converted into a two-dimensional image.
[0019] Further, the ratio of the training set, the validation set, and the test set is 60%, 15%, and 25%; training is performed using the Adam optimizer with a learning rate of 0.001 and a mini-batch size of 64; the loss function is the cross-entropy loss function.
[0020] The beneficial effects of this application are as follows: The CNN rolling bearing fault diagnosis method based on improved GAF and SA proposed in this application first reduces the dimension of one-dimensional time-series vibration data, and then encodes the reduced data into a two-dimensional image. Utilize the ability of the convolutional neural network's multi-layer "convolution + pooling" units to adaptively extract image features, and add the SA attention mechanism to achieve bearing fault classification. The effectiveness and anti-noise performance of GAF and the SA attention mechanism are verified through experiments. The advantage of this application is that it encodes the one-dimensional vibration signal after dimension reduction into an image using improved GAF, fully retaining and mining the fault information in the bearing data under a noisy environment. Use the convolutional neural network's multi-layer "convolution + pooling" units to automatically extract image features, and introduce the SA attention mechanism to adaptively weight the features, assigning greater weights to important features, which can further optimize the classification performance of the model and achieve effective diagnosis of rolling bearing faults under a noisy environment. Description of the Drawings
[0021] Figure 1 It is a schematic flowchart of an embodiment of this application;
[0022] Figure 2 It is a schematic diagram of the improved GAF two-dimensional diagram of an embodiment of this application;
[0023] Figure 3 It is a schematic diagram of the GAF-CNN network structure of an embodiment of this application;
[0024] Figure 4 It is a schematic diagram of the SA attention module structure of an embodiment of this application;
[0025] Figure 5 It is a schematic diagram of the training set loss function and accuracy rate during the training process of the GAF-CNN network of an embodiment of this application;
[0026] Figure 6 It is a schematic diagram of the test set confusion matrix during the training process of the GAF-CNN network of an embodiment of this application. Detailed Implementation Modes
[0027] The technical solution of the present application will be described in detail below in conjunction with the accompanying drawings.
[0028] This embodiment adopts a CNN rolling bearing fault diagnosis method based on improved GAF and SA. Refer to Figure 1 , including the following steps:
[0029] S1: Use the vibration data of rolling bearings from Case Western Reserve University, and perform preprocessing such as outlier processing and missing value filling on the data to obtain a one-dimensional vibration data set.
[0030] Specifically, it includes the following steps:
[0031] Step S11: The experimental data set uses the bearing fault data set of Case Western Reserve University. Take the data of the deep groove ball bearing 6205-2RS JEM SKF at the drive end as an example, and the sampling frequency is 12KHz. There are four loads in the experiment: 0hp, 1hp, 2hp, 3hp, corresponding to 1797 (r / min), 1772 (r / min), 1750 (r / min), 1730 (r / min) respectively. The data under the condition of no load 1797 (r / min) is selected to train the model.
[0032] Step S12: The fault positions of rolling bearings are divided into three types: rolling element fault (Ball), outer race fault (OuterRace), and inner race fault (InnerRace). Each fault position contains three fault sizes: 0.007inch, 0.014inch, 0.021inch. Therefore, the bearing faults can be divided into 9 fault states and 1 normal state. In data processing, after piecewise aggregate approximation, 170 points are taken as a sequence for each state, and 1000 samples are constructed for each type of signal and randomly shuffled.
[0033] S2: Use the improved GAF (Gram Angular Field method) to transform the one-dimensional vibration data set into a two-dimensional image, and then divide it into a training set, a validation set, and a test set.
[0034] Specifically, it includes:
[0035] Step S21: Reduce the dimension of the one-dimensional vibration data set by the piecewise aggregate approximation method. For a vibration data S=(x1, x2, x3,..., x n-1 , x n ) with a length of n, it is represented by a data S1=(y1, y2, y3,..., y m-1 , y m ) with a length of m, where m < n and n is divisible by m. The compression ratio of the entire time series is N k = n / m. In the present invention, the compression ratio is taken as 3, and the transformed sequence S1 can be expressed as:
[0036]
[0037] Step S22: Encode each sequence into a 170*170*3 two-dimensional image by the GASF method, and divide the corresponding training set, validation set, and test set. The images corresponding to the fault types are encoded as Figure 2 shown.
[0038] S3: Input the samples of the training set and the validation set into the GAF-CNN model, set the initial parameters, and perform model training. When the iteration converges, complete the network training, save the network parameters, and put the validation set into the model to calculate the accuracy.
[0039] As Figure 3 shown, the GAF-CNN network of this embodiment is composed of a convolutional neural network (CNN) and an SA (Shuffle Attention) attention module. The convolutional neural network includes multiple convolutional pooling modules and a fully connected layer; the SA attention module is set between the last convolutional pooling module and the fully connected layer.
[0040] A convolutional neural network usually consists of an input layer, a hidden layer, and an output layer. Among them, the hidden layer is usually composed of a convolutional layer, a pooling layer, and a fully connected layer. The hidden layer is usually composed of several groups of convolutional layers and pooling layers connected alternately plus a fully connected layer. The convolutional layer extracts features from the input data through a convolutional kernel. Each convolutional kernel corresponds to a weight coefficient and a bias, similar to the neurons of a feedforward neural network. The specific mathematical model of the convolutional layer is as follows:
[0041]
[0042] Among them, n_in represents the number of input matrices; X k represents the kth input matrix; W k represents the kth sub-convolutional kernel matrix of the convolutional kernel; s(i, j) represents the value of the corresponding element of the output matrix corresponding to the convolutional kernel W.
[0043] In the selection of the activation function of the convolutional network, compared with the sigmoid function and the tanh function, the ReLU function can better solve the problem of gradient explosion or gradient disappearance, and can also accelerate the convergence speed. The calculation formula of the ReLU activation function is as follows:
[0044] f(x) = max(0, x);
[0045] The pooling layer is usually located after the convolutional layer. The feature map output by the convolutional layer is sent to the pooling layer, and downsampling is performed through windows of different sizes for feature selection and filtering. Common pooling methods include max pooling, average pooling, etc. The mathematical model of the pooling layer is expressed as:
[0046]
[0047] Among them, down(.) represents the subsampling layer function; the sum of each different n×n sub-block of the input image is calculated, so that the output image is n times smaller in both spatial dimensions.
[0048] The fully connected layer is usually located after the last convolutional pooling layer. Each neuron in the fully connected layer is connected to all the activated neurons in the previous layer, and a non-linear combination of the functions extracted in the convolutional layer and pooling layer is performed to obtain the output.
[0049] The output layer often uses a logical function or a normalization function to output the classification label. In this application, the Softmax function is used, and its definition is as follows:
[0050]
[0051] Among them, the SA attention module integrates channel attention and spatial attention into one module by splitting the input feature map into multiple groups and then using Shuffle units for each group, and then aggregates all the features to perform information transfer in the channel dimension; the fully connected layer is used to integrate the features in the channel dimension and output the prediction value. The structure of the SA attention module is as Figure 4 shown.
[0052] The principle of the SA attention module in this embodiment is as follows:
[0053] First, feature grouping is performed. For a given feature map X∈R C×H×W , where C, H, and W represent the number of channels, spatial height, and spatial width respectively. X is divided into G groups along the channel dimension X = [X1,..., X G , and each sub-feature X i will obtain a corresponding importance coefficient during the training process. At the beginning of each attention unit, the input of X i is divided into two branches X k1 , X k2 along the channel dimension. One branch utilizes the mutual relationship between channels to output the channel attention map, and the other branch utilizes the spatial relationship of the features to generate the spatial attention map, giving higher weights to the effective features through these two branches.
[0054] For channel attention, first, global average pooling is used to generate channel statistics, S∈R C / 2G , embedding global information, and shrinking X k1 along the spatial dimension H×W. Finally, the output mathematical model combined with the sigmoid activation function is as follows:
[0055]
[0056] For spatial attention, the Group Normalization method is used to obtain spatial statistics, and then for X k2 feature enhancement is performed. The mathematical model is as follows:
[0057] X k2 ′ = σ(f(GN(X k2 ))).X k2 ;
[0058] Finally, the results of the two branches are concatenated to keep the number of channels unchanged.
[0059] During training, the proportions of the training set, test set, and validation set in the experimental dataset are 60%, 15%, and 25% respectively. Adam optimizer is used for training, the learning rate is 0.001, the mini - batch size is 64, and the loss function is the cross - entropy loss function. Iteration stops when the loss function values of the training set and the validation set are both small and tend to be stable.
[0060] As Figure 5 shown, it is the change of the loss function of the training set of the GAF - CNN network with the number of iterations. It can be seen that at the beginning of training, the loss function is large and the accuracy is low. As training progresses, the loss function decreases rapidly and the accuracy increases rapidly. When the number of iterations reaches 45 rounds, the loss function is basically stable at around 0.01 and the accuracy is stable at around 0.99.
[0061] S4: Input the samples of the test set into the trained GAF - CNN network to obtain the fault diagnosis results of the test set.
[0062] To verify the performance of the GAF - CNN model proposed in this application, three models commonly used for fault diagnosis are used for comparison, namely 1D - CNN, SVM, and RF. The same test data is input into these three network models of 1D - CNN, SVM, and RF respectively, and the diagnostic results are compared. Figure 6 It is the confusion matrix of the test set.
[0063] Among them, 1D - CNN directly inputs a 1*170 sequence into the corresponding 1D - CNN model. The number of network layers is the same as that of GAF - CNN, only changing from two - dimensional to one - dimensional in structure.
[0064] Table 1 shows the accuracies of the GAF - CNN model and the three comparison models. It can be seen that after adding the attention module, the accuracy on the test set is improved to 99.56%, which is 5% higher than that of using the ordinary convolutional neural network, verifying the effectiveness of the method proposed in this application.
[0065] Table 1
[0066] Model Training set accuracy Test set accuracy GAF-CNN 0.9899 0.9956 1D-CNN 0.9797 0.9488 SVM 0.9389 0.9156 RF 0.9611 0.9378
[0067] In this embodiment, the fault diagnosis accuracy rates of SVM and RF are relatively low, both below 94%. The main reason is that the learning ability of such models strongly depends on the extracted features, and it is difficult for manually extracted features to uncover more comprehensive and deeper information of the signals. This also highlights the advantages of automatically extracting features based on deep learning. From the comparison between the GAF-CNN and 1D-CNN models, it can be seen that the classification results of the time series after using the improved GAF transformation are better than those directly classified by the one-dimensional convolutional neural network, indicating that the encoding method using the Gram angle field can more effectively extract the features of the original data and can reflect the characteristics of different faults through the relationships between data points.
[0068] Those of ordinary skill in the art can understand that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An improved GAF and SA-based CNN rolling bearing fault diagnosis method, characterized in that, Including: S1: Collect vibration data of the rolling bearing, preprocess the vibration data to obtain a one-dimensional vibration data set; S2: Transform the one-dimensional vibration data set into a two-dimensional image by improved GAF, and divide the two-dimensional image into a training set, a validation set and a test set; S3: Input both the training set and the validation set into the GAF-CNN model, set the initial parameters of the GAF-CNN model, train the GAF-CNN model with the training set, and when the GAF-CNN model converges iteratively, complete the network training, save the network parameters, and obtain the trained GAF-CNN model; Input the validation set into the trained GAF-CNN model to verify the accuracy. If the accuracy of the validation set reaches the preset standard, go to step S4, otherwise repeat step S3 until the accuracy of the validation set reaches the preset standard; S4: Input the test set into the trained GAF-CNN model to obtain the test set accuracy and complete the fault diagnosis; Wherein, the GAF-CNN model includes a convolutional neural network and an SA attention module, and the convolutional neural network includes four cascaded convolutional pooling modules and one fully connected layer; The SA attention module is arranged between the last convolutional pooling module and the fully connected layer; The improved GAF is expressed as: where x i represents the i-th value of the time series data. For the matrix element G (i,j||i-j=k) , when k = 0, the main diagonal G i,j is composed of the original values of the scaled time series.
2. The fault diagnosis method according to claim 1, wherein, In step S2, before transforming the one-dimensional vibration data set into a two-dimensional image, the piecewise aggregate approximation method is used to first reduce the dimension of the one-dimensional vibration data set, and then transform the dimension-reduced one-dimensional vibration data set into a two-dimensional image.
3. The fault diagnosis method according to claim 1, characterized in that, The ratio of the training set, the validation set and the test set is 60%, 15% and 25%; Training is carried out through the Adam optimizer, with a learning rate of 0.001 and a mini-batch size of 64; The loss function is the cross-entropy loss function.
Citation Information
Patent Citations
A rolling bearing fault identification method under variable working conditions based on ATT-CNN
CN109902399A
MWDCNN-based bearing fault diagnosis method under variable working condition
CN111964908A