Bearing Fault Diagnosis Method, System, Device and Storage Medium

By combining the interpretability technology of CNN and Grad-CAM and the Swin Transformer model, the bearing vibration signal is converted into two-dimensional images for fault diagnosis, solving the problem of insufficient classification interpretability in the prior art, and achieving high accuracy and interpretable bearing fault diagnosis.

CN116610993BActive Publication Date: 2025-07-04HUIZHI NEW ENERGY (SHANDONG) INFORMATION TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310493855.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-07-04
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

The existing bearing fault diagnosis methods have shortcomings in classification interpretability and accuracy, and traditional deep learning models such as RNN are limited in parallel computing and long-term dependency problems. The Transformer model has not been widely used in the field of fault diagnosis.

Method used

Convolutional neural network (CNN) combined with interpretability technology Grad-CAM is used for feature visualization, and the latest Transformer model such as Swin Transformer is introduced. The vibration signal is converted into two-dimensional images through sliding window sampling, Gram Angle Field (GAF) and Wavelet Transformer (WT) to construct a deep learning model for troubleshooting.

Benefits of technology

It improves the accuracy and interpretability of bearing fault diagnosis, reduces equipment failure rate and downtime, and improves equipment utilization and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116610993B_ABST
    Figure CN116610993B_ABST
Patent Text Reader

Abstract

The present invention discloses a bearing fault diagnosis method, system, device and storage medium. The method includes: obtaining a data set, where the data set is a two-dimensional image with known bearing fault type labels, and the two-dimensional image is obtained by converting the bearing vibration signal; the data set is divided into a training set and a test set; respectively training a first deep learning model and a second deep learning model based on the training set to obtain a trained first deep learning model and a trained second deep learning model; based on the test set, testing the trained first deep learning model and the trained second deep learning model, and screening out the deep learning model with high classification accuracy as the final deep learning model for output; obtaining the vibration signal of the bearing to be diagnosed, converting the vibration signal of the bearing to be diagnosed into a two-dimensional image, and inputting the two-dimensional image of the bearing to be diagnosed into the final deep learning model to output the bearing fault diagnosis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of mechanical fault diagnosis and interpretability, and particularly to a bearing fault diagnosis method, system, device, and storage medium. Background Technique

[0002] The statements in this part only mention the background technique related to the present invention and do not necessarily constitute the prior art.

[0003] In recent years, rotating machinery has been widely used in modern industries, including industrial fields such as automobiles, motors, and aviation. However, due to the harsh working environment, rotating machinery is prone to failure. According to statistics, about 40%-50% of rotating machinery failures are caused by rolling bearing damage. Rolling bearing failures may cause huge economic losses and even endanger the safety of operators in severe cases. Therefore, accurate and timely diagnosis of bearing faults is of great significance for improving the reliability and operating safety of rotating machinery.

[0004] Fault diagnosis methods generally include model-based methods and data-driven methods. The former requires a large amount of prior knowledge and it is difficult to accurately establish a diagnostic model under complex conditions. The latter uses the data collected by sensors for fault diagnosis. It does not rely on prior knowledge and provides accurate fault diagnosis results by analyzing mechanical signals. Therefore, with the development of intelligent algorithms, data-driven methods have been widely used in fault diagnosis. The most common intelligent fault diagnosis methods currently are developed on machine learning (ML) methods, such as support vector machines (SVM), K-nearest neighbor (KNN) algorithms, self-organizing map (SOM) networks, etc. Most rolling bearing intelligent fault diagnosis methods are based on the processing and analysis of vibration signals. The use of ML requires artificial extraction of features from the acquired original signals and does not fully utilize the advantages of intelligent algorithms. In recent years, deep learning (DL), as an important branch of ML, has developed rapidly and has achieved great success in many fields such as natural language processing (NLP) and computer vision (CV) with its powerful data processing ability and learning ability. DL models have also been widely used in the field of bearing fault diagnosis, such as deep belief networks (DBN), convolutional neural networks (CNN), autoencoders (AE), recurrent neural networks (RNN), generative adversarial networks (GAN), etc.

[0005] In recent years, aiming at the problem of difficult identification of bearing fault types, the convolutional neural network (CNN), as one of the most effective algorithms in the DL framework, has achieved remarkable results in multiple detection and identification problems. It automatically extracts relevant features according to the learning objectives and does not rely on expert experience, and has been increasingly widely used in the field of bearing fault diagnosis. Although with the continuous innovation of intelligent algorithms, great progress has been made in the technology of bearing fault diagnosis, there are still some problems that need further research. For example, most CNN-based fault diagnosis methods only classify the fault types and do not consider the interpretability of the classification. The interpretability technology of CNN is widely used in the field of computer vision, but there is little related research in the field of mechanical fault diagnosis. Therefore, the present invention needs to use a visualization method to explain the deep fault features, establish the connection between the key activation regions of the neural network and the target categories, and solve the problem of the black-box operation of the neural network. Secondly, a new DL model Transformer has been recently proposed, which only contains the multi-head self-attention (MSA) mechanism and basic fully connected layers. This model has achieved excellent results in the NLP field. The Transformer model has a self-attention mechanism, which can effectively obtain global information, and the multi-head attention mechanism can map the information to multiple spaces, making the model more expressive, and it breaks through the limitation that the recurrent neural network (RNN) model cannot perform parallel computing. At the same time, the improved model of Transformer is widely used in the CV field. In 2021, Dosovitskiy et al. proposed a Vision Transformer model for image classification tasks and achieved amazing results on larger datasets. Liu et al. proposed a Swin Transformer model for applications such as image classification, object detection, and semantic segmentation. Compared with Vision Transformer, it is more suitable for downstream tasks. However, although the Transformer model has achieved excellent results in both the CV and NLP fields, it has not been widely applied in the field of fault diagnosis.

[0006] Chinese invention patent (Application No.: 202211138667.9, Patent Name: A Rolling Bearing Fault Diagnosis Method). First, the obtained bearing vibration signal is subjected to time-frequency conversion to obtain a time-frequency signal, and then the time-frequency signal is input into the constructed fault diagnosis model for feature extraction and fault identification. The fault diagnosis model of this invention can achieve good learning effects with fewer samples and has better performance than other models under some complex working conditions. However, this invention only classifies the faults and does not consider the interpretability of the neural network.

[0007] Chinese invention patent (Application No.: 202010662863.0, Patent Title: A Fault Diagnosis Method for Mechanical Equipment Based on Deep Learning). First, data is collected from mechanical equipment and preprocessed. Then, the dataset is divided (into a training set, a validation set, and a test set) by means of cross-validation. Finally, a fault diagnosis model based on CNN and bidirectional long short-term memory network (BD-LSTM) is established, and the dataset is input into the fault diagnosis model to extract hidden features and identify the fault type. This invention solves the problem of uncertainty caused by environmental interference. However, the variant BD-LSTM of RNN does not completely solve the long-term dependence problem, and the inherent sequential property of RNN hinders the parallelization among training samples. Summary of the Invention

[0008] To solve the deficiencies of the prior art, the present invention provides a bearing fault diagnosis method, system, device, and storage medium; the fault diagnosis of bearings is realized through CNN, and the interpretability technology of CNN is applied to the field of fault diagnosis, and the attribution of the fault diagnosis result is understood through visualization. Secondly, the latest Transformer model with the attention mechanism as the core is applied to the field of bearing fault diagnosis, so as to more accurately identify the fault type. Timely and accurate fault diagnosis can greatly reduce the failure rate and downtime of equipment, improve the equipment utilization rate, thus bringing greater economic benefits to enterprises, and at the same time eliminating the inestimable safety hazards caused by sudden failures.

[0009] In the first aspect, the present invention provides a bearing fault diagnosis method;

[0010] The bearing fault diagnosis method includes:

[0011] Obtain a dataset, where the dataset is a two-dimensional image with known bearing fault type labels, and the two-dimensional image is obtained by converting the bearing vibration signal; the dataset is divided into a training set and a test set;

[0012] Based on the training set, train the first deep learning model and the second deep learning model respectively to obtain the trained first deep learning model and the trained second deep learning model; based on the test set, test the trained first deep learning model and the trained second deep learning model, and select the deep learning model with high classification accuracy as the final deep learning model for output;

[0013] Obtain the vibration signal of the bearing to be diagnosed, convert the vibration signal of the bearing to be diagnosed into a two-dimensional image, and input the two-dimensional image of the bearing to be diagnosed into the final deep learning model to output the bearing fault diagnosis result.

[0014] In the second aspect, the present invention provides a bearing fault diagnosis system;

[0015] Bearing fault diagnosis system, comprising:

[0016] An acquisition module, configured to: acquire a data set, the data set being a two-dimensional image with known bearing fault type labels, the two-dimensional image being obtained by converting a bearing vibration signal; the data set is divided into a training set and a test set;

[0017] A training module, configured to: respectively train a first deep learning model and a second deep learning model based on the training set to obtain a trained first deep learning model and a trained second deep learning model; based on the test set, test the trained first deep learning model and the trained second deep learning model, and screen out the deep learning model with high classification accuracy as the final deep learning model for output;

[0018] A diagnosis module, configured to: acquire the vibration signal of the bearing to be diagnosed, convert the vibration signal of the bearing to be diagnosed into a two-dimensional image, input the two-dimensional image of the bearing to be diagnosed into the final deep learning model, and output the bearing fault diagnosis result.

[0019] In a third aspect, the present invention further provides an electronic device, comprising:

[0020] A memory for non-temporarily storing computer-readable instructions; and

[0021] A processor for running the computer-readable instructions,

[0022] wherein, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.

[0023] In a fourth aspect, the present invention further provides a storage medium for non-temporarily storing computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.

[0024] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, the computer program being used to implement the method described in the first aspect above when running on one or more processors.

[0025] Compared with the prior art, the beneficial effects of the present invention are:

[0026] The present invention first obtains the vibration signal of a bearing, performs sliding window sampling on the signal to obtain more samples for training a more accurate model; then converts the vibration signal after sliding window sampling into different types of two-dimensional images through the Gramian Angular Field (GAF) and Wavelet Transform (WT) methods; randomly divides the two-dimensional images into a training set and a test set; constructs a Convolutional Neural Network (CNN) and a Transformer model, inputs the training set into the model for feature extraction, and the test set is used to verify the performance of the model; conducts experiments through transfer learning, verifies the performance of the constructed model, and compares it with other methods; uses the interpretability technique Gradient-weighted Class Activation Mapping (Grad-CAM) of the CNN model for visualization to explain the attribution of the fault diagnosis results. The method of the present invention has a higher fault diagnosis accuracy rate, which is superior to many existing deep learning models. Description of the Drawings

[0027] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0028] Figure 1 is the flowchart of the method of the present invention;

[0029] Figures 2(a)-2(j) is the GAF image;

[0030] Figures 3(a)-3(j) is the WT image;

[0031] Figure 4 is the process diagram of bearing fault diagnosis based on CNN;

[0032] Figure 5 is the process diagram of bearing fault diagnosis based on Transformer;

[0033] Figures 6(a) and 6(b) are the bearing fault diagnosis results based on CNN;

[0034] Figures 7(a) and 7(b) are the confusion matrices of the bearing fault diagnosis results based on CNN;

[0035] Figures 8(a) and 8(b) are the bearing fault diagnosis results based on Transformer;

[0036] Figures 9(a) and 9(b) are the confusion matrices of the bearing fault diagnosis results based on Transformer;

[0037] Figures 10(a)-10(d) is the visualization result diagram of the GAF image;

[0038] Figures 11(a)-11(d) is the visualization result diagram of the WT image;

[0039] Figure 12 It is a process diagram of sliding window sampling. Specific implementation manner

[0040] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0041] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0042] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0043] Embodiment 1

[0044] This embodiment provides a bearing fault diagnosis method;

[0045] As Figure 1 shown, the bearing fault diagnosis method includes:

[0046] S101: Obtain a data set, where the data set is a two-dimensional image with known bearing fault type labels, and the two-dimensional image is obtained by converting the bearing vibration signal; the data set is divided into a training set and a test set;

[0047] S102: Train the first deep learning model and the second deep learning model respectively based on the training set to obtain the trained first deep learning model and the trained second deep learning model;

[0048] Based on the test set, test the trained first deep learning model and the trained second deep learning model, and select the deep learning model with high classification accuracy as the final deep learning model for output;

[0049] S103: Obtain the vibration signal of the bearing to be diagnosed, convert the vibration signal of the bearing to be diagnosed into a two-dimensional image, input the two-dimensional image of the bearing to be diagnosed into the final deep learning model, and output the bearing fault diagnosis result.

[0050] Further, S101: Obtain a dataset. Select the Case Western Reserve University (CWRU) bearing dataset. Each fault state in the original vibration signals collected from the bearings contains a different number of time - process measurement values. Reshape the samples by means of sliding - window sampling to ensure that each sample has a consistent number of time - process measurement values, achieving the purpose of increasing the number of samples to train a more accurate model.

[0051] The Case Western Reserve University (CWRU) bearing dataset is selected. The fault points on the bearings of this dataset are all artificially processed, and their diameters are divided into 5 types: 7 mils, 14 mils, 21 mils, 28 mils, and 40 mils (1 mil = 0.001 inches). SKF bearings are used for faults with diameters of 7, 14, and 21 mils, and NTN equivalent bearings are used for faults with diameters of 28 mils and 40 mils. The bearing vibration signals are collected using accelerometers, and the accelerometers are fixed to the housing with magnetic bases. In this invention, the SKF bearing drive - end dataset and the normal dataset at a sampling frequency of 12 kHz are selected for experiments, and 10 different fault types under 1 load condition are studied. This invention uses the sliding - window sampling method for data augmentation. Every 1200 data points are used as a sample, and the size of the sliding window is set to 600, so as to obtain more sample numbers for training a more accurate model. The process of sliding - window sampling is as Figure 12 shown. The detailed information of the dataset is listed in Table 1.

[0052] Table 1: Detailed information of the dataset

[0053]

[0054] Different fault states in the original vibration signals collected from the bearings contain different numbers of time - process measurement values. To increase the number of samples to train a more accurate model, this invention reshapes the samples by means of sliding - window sampling. Every 1200 data points are used as a sample, and the size of the sliding window is set to 600. The process of sliding - window sampling is as Figure 12 shown.

[0055] Further, the two - dimensional image is obtained by transforming the bearing vibration signal, including:

[0056] Transform the bearing vibration signal through the Gramian angular field algorithm to obtain a two - dimensional image of the bearing vibration signal; transform the bearing vibration signal through the wavelet transform algorithm to obtain a two - dimensional image of the bearing vibration signal.

[0057] Use the Gramian angular field (GAF) and wavelet transform (WT) methods to transform the reshaped samples into different types of two - dimensional images.

[0058] Exemplarily, the bearing vibration signal is transformed through the Gramian Angular Field (GAF) algorithm to obtain a two-dimensional image of the bearing vibration signal, which specifically includes:

[0059] S101-a1: Normalize the original bearing vibration signal X = {x1, x2,..., x n} and scale X to the range [0, 1] using Equation (1).

[0060]

[0061] where xi i is the i-th value in X, and is the value of xi i after scaling.

[0062] S101-a2: Transform the scaled timestamp and amplitude into the radius r and cosine of the angle in polar coordinates to form a new time series The formula is as follows:

[0063]

[0064] where t i is the timestamp and N is the regularization constant factor in polar coordinates.

[0065] S101-a3: Calculate the correlation between any two data points through trigonometric functions. Equation (3) can transform the time series into an image, i.e., GAF.

[0066]

[0067] where I is the unit row vector, and is the transposed vector of . The GAF image is as shown in Figures 2(a)-2(j) .

[0068] Exemplarily, the bearing vibration signal is transformed through the wavelet transform algorithm to obtain a two-dimensional image of the bearing vibration signal, which specifically includes:

[0069]

[0070]

[0071] where a is the scale factor; τ is the translation factor; ψ a,τ (t) is the wavelet basis function. The WT image is as shown in Figures 3(a)-3(j) .

[0072] Wavelet transform (WT) is a new transform analysis method. It inherits and develops the idea of localization of the short-time Fourier transform, and at the same time overcomes the disadvantages such as the window size not changing with frequency. It can provide a "time-frequency" window that changes with frequency and is an ideal tool for time-frequency analysis and processing of signals. It can automatically adapt to the requirements of time-frequency signal analysis, thus focusing on any details of the signal and solving the difficult problems of Fourier transform.

[0073] The original vibration signal is converted into different types of two-dimensional images using the GAF and WT methods. The GAF diagram completely describes the instantaneous characteristics of the signal, such as Figures 2(a)-2(j) shown. The WT diagram describes the overall effect of the signal, such as Figures 3(a)-3(j) shown.

[0074] Furthermore, the dataset is divided into a training set and a test set, which are randomly divided into a training set and a test set at a ratio of 8:2 for the two-dimensional images. The training set is used to train the constructed deep learning model, and the test set is used to verify the performance of the model.

[0075] In the present invention, the 2000 reconstructed samples are randomly divided into a training set and a test set at a ratio of 8:2, that is, 1600 training samples and 400 test samples. In the present invention, the neural network model is trained using the training set, and the test set is used to evaluate the performance of the model.

[0076] Furthermore, S102: The first deep learning model and the second deep learning model are respectively trained based on the training set to obtain the trained first deep learning model and the trained second deep learning model. Among them, the first deep learning model includes: the Conv1 layer of the Resnet34 network, the Conv2_x of the Resnet34 network, the Conv3_x layer of the Resnet34 network, the Conv4_x layer of the Resnet34 network, the Conv5_x layer of the Resnet34 network, an interpretable layer, an average pooling layer, and a fully connected layer connected in sequence.

[0077] The present invention selects the Resnet34 network for the fault classification task for the following reasons. First, the Resnet34 network contains a residual structure, which can simplify the training of the network and reduce the training parameters of the network. Second, the Resnet34 network uses a Batch Normalization layer, which can solve the problem of gradient disappearance or gradient explosion. Third, most industrial monitoring data has local correlation. Since the CNN filter can learn local correlation information, it is suitable for industrial fault detection. Finally, the industrial environment is often polluted by noise. CNN can extract translation-invariant features, enhance robustness, and reduce the adverse effects of noise. The constructed Resnet model is asFigure 4 as shown

[0078] The 2D image is input into the Resnet34 network model in the CNN. The first convolutional operation of this network uses 64 convolutional kernels with a size of (7, 7), and the moving step size is 2. After each convolutional operation, it has to go through a batch normalization layer and an activation layer in sequence. The operations of the remaining convolutional layers are similar to those of the first layer. Each curved arrow represents a residual structure, and the number next to it represents how many identical residual structures are connected together. Figure 4 The first residual structures corresponding to Conv3_x, Conv4_x, and Conv5_x in [] are dotted-line residual structures. This residual structure needs to reduce the height and width of the feature matrix to half of the original, and adjust the number of channels. After a series of convolutional operations, it passes through an average pooling layer, a fully connected layer, and the Softmax function in sequence to identify the type and location of the fault. The present invention conducts experiments on the CWRU bearing dataset. The fault categories of the bearing can be divided into four categories: inner race fault (IR), outer race fault (OR), ball fault (BR), and normal (Normal). The entire fault diagnosis process of the CNN is as Figure 4 as shown

[0079] The detailed steps of the fault diagnosis method based on CNN are summarized in the following algorithm:

[0080] Input: GAF or WT image

[0081] Step 1: Feature extraction.

[0082] The feature extractor can extract features from the original data. The following are the operations of the core layer of feature extraction:

[0083] Convolutional layer:

[0084] Batch normalization layer:

[0085] ReLU activation layer: R r = ReLU(H r ) = max(H r , 0)

[0086] Max pooling layer:

[0087] Step 2: Fault classification.

[0088] The fault classifier identifies the fault type based on the extracted features. The following are the operations of the core layer of fault classification:

[0089] Average pooling layer:

[0090] Fully connected layer: Z = [Z 1 , Z 2 ,..., Z r ,..., Z R , G = ReLU(W g Z + b g )

[0091] Softmax function:

[0092] Output: Target classes (Normal, BR007,....., OR021).

[0093] In the above algorithm, q = 1, 2, 3,..., Q, where Q is the number of predefined filters; p = 1, 2, 3,..., P, where P is the number of channels; W p,q and b q represent the weights and biases of the filters respectively; c is the mini - batch of feature C; u and σ 2 are the mean and variance of the mini - batch respectively; ε is the minimum value close to 0; S represents the range of the pooling function; W g and b g represent the weights and biases of the fully connected layer; O j represents the estimated probability of class j; θ (j) is the parameter of the Softmax function; K is the number of target classes.

[0094] Furthermore, the interpretation layer refers to Grad - CAM (Gradient - weighted Class Activation Mapping);

[0095] The class discriminative localization map Grad - CAM L c is calculated as follows:

[0096]

[0097] where A k refers to the k - th feature map of the CNN layer; is obtained by combining the gradient of the c - class score with the Y of the feature map A. is obtained by globally average pooling the gradient of the c - class score with the y k of the feature map A c ; f(·) is the ReLU activation function, which emphasizes the features that have a positive impact on the class of interest.

[0098] To better understand CNN, a visual interpretation is carried out for it to better make decisions about the model. Gradient-weighted Class Activation Mapping (Grad-CAM) is proposed. Compared with the previous work CAM, Grad-CAM can visualize CNNs with any structure without modifying the network structure or retraining. Therefore, to explain the fault classification results, the present invention uses Grad-CAM on the last convolutional layer of the Resnet34 network to locate and highlight the distinguishing regions. The interpretable technology of CNN has rarely been applied in the field of fault diagnosis. To fill this gap, the present invention applies it to fault diagnosis. The visualization results are as Figures 10(a)-10(d) and Figures 11(a)-11(d) shown.

[0099] It should be understood that to explain the fault diagnosis results, Gradient-weighted Class Activation Mapping (Grad-CAM) is used. Grad-CAM is a class-discriminative localization method that assigns scores to each class using the filter gradients based on backpropagation and the convolutional activation values, providing a way to see which specific parts of the image affect the model's decision-making.

[0100] To explain the fault classification results, the present invention uses Grad-CAM on the last convolutional layer of the CNN to locate and highlight the distinguishing regions. One GAF image is selected from each of the four major types of bearing fault samples for visualization. The visualization results are as Figures 10(a)-10(d) shown.

[0101] The experimental results show that the vibration signal of the normal bearing is weak and there is no obvious impact peak. The attention regions of the network model are scattered on most segments of the input signal, indicating that the contribution degrees of most signal sequences to the output are basically the same. For bearings with outer race, inner race, and rolling element faults, the parts with larger network activation degrees are basically concentrated in the parts with stronger vibration signals. The information in this part has a higher weight for the classification results of the network, indicating that this position contains more fault feature information. One sample is selected from each of the four major types of bearing fault WT images for visualization. The visualization results are as Figures 11(a)-11(d) shown. The experimental results show that the key regions focused on by the network are the parts with larger amplitudes. Therefore, there is a basic similarity between the classification and recognition of samples by the convolutional neural network in the field of bearing fault diagnosis and the human cognitive law, which can prove that the regions focused on by the method proposed in the present invention are correct when performing fault classification, and solve the problem of the black box operation of the neural network.

[0102] Further, in S102: The first deep learning model and the second deep learning model are respectively trained based on the training set to obtain the trained first deep learning model and the trained second deep learning model, wherein the second deep learning model is a Swin Transformer network.

[0103] The Transformer model with the attention mechanism as the core has been successfully applied in the field of natural language processing and almost replaced the RNN. In addition, the improved models of Transformer, Vision Transformer and Swin Transformer, have been successfully applied in the field of computer vision and achieved unprecedented results. However, the latest Transformer models have not been applied in the field of fault diagnosis, and the Swin Transformer model is more suitable for downstream tasks than the Vision Transformer model. The Swin Transformer model has shown excellent performance in the field of computer vision, such as image classification, semantic segmentation, object detection, etc. Therefore, the present invention proposes a method based on the combination of two-dimensional images and the Swin Transformer model for bearing fault diagnosis.

[0104] Swin Transformer has a hierarchical structure similar to CNN. The overall architecture of Swin Transformer based on fault diagnosis is as Figure 5 shown. The Swin Transformer Block module is the core of the Swin Transformer model, which is composed of a window-based multi-head self-attention mechanism (Window based Multi-head Self-Attention, W-MSA) module and a shifted window-based multi-head self-attention mechanism (Shifted Window based Multi-head Self-Attention, SW-MSA) module. Layer normalization (LayerNorm, LN) is applied before each multi-head self-attention mechanism (Multi-head Self-Attention, MSA) module and each multi-layer perceptron (Multilayer perceptron, MLP), and a residual connection is applied after each module.

[0105] The calculation of two consecutive Swin Transformer Blocks can be expressed as:

[0106]

[0107]

[0108]

[0109]

[0110] Among them, and x l represent the output features of the W-MSA module and the MLP module in block l.

[0111] The self-attention calculation with relative position bias is as follows:

[0112]

[0113] Among them, are the query, key, and value matrices; d is the dimension of the query / key; M 2 is the number of patches in the window.

[0114] The overall network architecture of the Swin Transformer model based on the moving window adopts a hierarchical design, including a total of 4 Stages. Each Stage reduces the resolution of the input feature map, similar to the CNN operation. For an input image with a size of 224×224×3, first, like the Vision Transformer, the image is divided into patch blocks. Here, the patch size used in Swin Transformer is 4×4, different from the size of 16×16 used in Vision Transformer. After Patch Partition, the size of the image becomes 56×56×48. The main function of the Linear Embedding layer is to change the dimension of the vector to the value preset in the present invention, that is, to meet the value that can be input by the Transformer. The size of the number of channels C in the Swin Transformer model used in the present invention is 96, and the network output value obtained after Stage1 is 56x56x96. In Stage2, similar to the pooling operation of the convolutional neural network, Patch Merging is used to reduce the resolution and adjust the number of channels, so as to realize the hierarchical design of the model. Here, the downsampling is 2 each time, elements are selected at every other point in the row and column directions, and then spliced together and unfolded, and the size of the number of channels becomes 2C. After the Stage2 operation, the size of the network output becomes 28×28×192. The same applies to Stage3 and Stage4. Finally, the size of the network output becomes 7×7×768. After flattening, the sequence length becomes 49×768. Then, it passes through the Avgpool layer, Flatten layer, Linear layer, and Softmax layer in turn. The probability value of the input image belonging to each category is obtained through the Softmax layer, so as to obtain the fault type of the sample. The process of the fault diagnosis task based on Swin Transformer is as Figure 5 shown.

[0115] Further, after the S102 and before the S103, it also includes: using the method of transfer learning for experiments, and an ideal model can be quickly trained even when the dataset is small. At the same time, the performance of the model of the present invention is compared with that of other excellent models.

[0116] The present invention conducts experiments using the method of transfer learning, and can quickly train an ideal model even when the dataset is small. The present invention adjusts the parameters based on the pre-trained Resnet34 network according to different classification and prediction targets. The size of the time-series image is reconstructed to (224, 224) by the Resize method. After operations of multiple stacked convolutional layers, pooling layers, and batch normalization layers, the size of the output feature map changes from (224, 224) to (7, 7, 512). Then the output feature map passes through the average pooling layer, fully connected layer, and Softmax function to identify the category of the time-series image. During the training process, the learning rate of the Adam optimization algorithm is set to 0.0001, and the number of training iterations is set to 50. The bearing fault diagnosis results based on CNN are shown in Figures 6(a) and 6(b).

[0117] The bearing fault diagnosis results based on CNN are shown in Figures 6(a) and 6(b). Figure 6(a) is about the GAF image, and Figure 6(b) is about the WT image. The experimental results show that the fault diagnosis accuracy of the method of the present invention reaches 96% and 100% on the GAF image and WT image of the test set respectively. The features contained in the GAF image are relatively complex, and the features extracted from the WT image are more intuitive, so the neural network has a higher accuracy on the WT image. To more clearly understand the details of fault classification, the present invention draws the confusion matrices as shown in Figures 7(a) and 7(b). Figure 7(a) is the confusion matrix about the GAF image, and Figure 7(b) is the confusion matrix about the WT image. The horizontal axis represents the true label, and the vertical axis represents the predicted label. For the GAF image, there are relatively more misclassified samples in the BR007 and BR021 classes, but these two classes of samples belong to the same major class, so it is somewhat difficult for the neural network to classify. There are 2 and 1 misclassified samples in the IR021 and OR014@6-1 respectively. For the WT image, all samples are classified into the correct classes.

[0118] The present invention adjusts the parameters in the Swin Transformer to make the model suitable for the fault diagnosis task of the present invention. The size of the input image is set to 224×224, the number of classes is 10, the learning rate of the Adam optimization algorithm is set to 0.0001, and the number of iterations is 50. The present invention conducts transfer learning on the pre-trained Swin Transformer model, and can train an ideal model even when the training set is small. During the training process, except for the head layer, the parameters of other layers are all frozen. The bearing fault diagnosis results based on Swin Transformer are shown in Figures 8(a) and 8(b).

[0119] The experimental results of bearing fault classification based on Swin Transformer are shown in Figures 8(a) and 8(b). Figure 8(a) is about GAF images, and Figure 8(b) is about WT images. The experimental results show that the method of the present invention achieves fault diagnosis accuracies of 98.5% and 100% on the test set of GAF and WT images respectively. After reconstruction, the CWRU dataset only contains 2000 samples, and the model of the present invention has achieved good performance on small samples. To more clearly understand the details of fault classification, the present invention draws confusion matrices as shown in Figures 9(a) and 9(b). Figure 9(a) is the confusion matrix for GAF images, and Figure 9(b) is the confusion matrix for WT images. The horizontal axis represents the true labels, and the vertical axis represents the predicted labels. For GAF images, only a few samples of the BR021 class are misclassified, and the remaining images are correctly classified into the corresponding classes. For WT images, all samples are classified into the correct classes.

[0120] The experimental results show that the method of the present invention can accurately classify the bearing fault types based on the bearing vibration signals and labels without relying on any prior knowledge. Because the multi-head attention mechanism in the Swin Transformer model can effectively obtain global information, the Swin Transformer model has a higher accuracy on GAF images compared to the CNN model. To further illustrate the effectiveness of the method of the present invention, it is compared with other methods using the same dataset, and the comparison results are shown in Table 2. To avoid the contingency of experimental results, the present invention conducts 5 experiments, and the experimental results in Table 2 are the averages of the 5 experimental results. It can be seen from Table 2 that the method of the present invention outperforms many excellent deep learning models.

[0121] Table 2: Comparison of the method of the present invention with other deep learning models.

[0122] Method Number of classes Accuracy LS-SVM 4 89.50% BPNN 10 81.35% LFGRU 10 97.32% ConvRNN 10 94.74% ANN 2 92% 2DViT 10 94% CNN(GAF) 10 96% CNN(WT) 10 100% Transformer(GAF) 10 98.5% Transformer(WT) 10 100%

[0123] Embodiment 2

[0124] This embodiment provides a bearing fault diagnosis model training system;

[0125] A bearing fault diagnosis system, comprising:

[0126] An acquisition module, configured to: acquire a dataset, the dataset being two-dimensional images with known bearing fault type labels, and the two-dimensional images being obtained by converting bearing vibration signals; the dataset is divided into a training set and a test set;

[0127] A training module, which is configured to: train a first deep learning model and a second deep learning model respectively based on a training set to obtain a trained first deep learning model and a trained second deep learning model; based on a test set, test the trained first deep learning model and the trained second deep learning model, and select a deep learning model with high classification accuracy as the final deep learning model for output;

[0128] A diagnosis module, which is configured to: obtain the vibration signal of the bearing to be diagnosed, convert the vibration signal of the bearing to be diagnosed into a two-dimensional image, input the two-dimensional image of the bearing to be diagnosed into the final deep learning model, and output the bearing fault diagnosis result.

[0129] It should be noted here that the above-mentioned acquisition module, training module and diagnosis module correspond to steps S101 to S103 in Embodiment 1. The examples and application scenarios implemented by the above-mentioned modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above-mentioned modules can be executed in a computer system such as a set of computer executable instructions as part of the system.

[0130] Embodiment 3

[0131] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above-mentioned one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in Embodiment 1 above.

[0132] Embodiment 4

[0133] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in Embodiment 1 is completed.

[0134] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A bearing fault diagnosis method, characterized in that Including: Obtain a data set, where the data set is a two-dimensional image with known bearing fault type labels, and the two-dimensional image is obtained by converting the bearing vibration signal; The data set is divided into a training set and a test set. The two-dimensional image is obtained by converting the bearing vibration signal, including: Convert the bearing vibration signal through the Gram angular field algorithm to obtain a two-dimensional image of the bearing vibration signal; convert the bearing vibration signal through the wavelet transform algorithm to obtain a two-dimensional image of the bearing vibration signal; Train the first deep learning model and the second deep learning model respectively based on the training set to obtain the trained first deep learning model and the trained second deep learning model. Among them, the first deep learning model includes: the Conv1 layer of the Resnet34 network, the Conv2_x of the Resnet34 network, the Conv3_x layer of the Resnet34 network, the Conv4_x layer of the Resnet34 network, the Conv5_x layer of the Resnet34 network, an interpretable layer, an average pooling layer, and a fully connected layer connected in sequence. The interpretable layer refers to the Gradient-weighted Class Activation Mapping (Grad-CAM); Class discrimination and localization map The calculation is as follows: (6) Among them, refers to the k-th feature map of the CNN layer; is obtained by combining the gradient of the c-class scores with Y of the feature map A; is obtained by performing global average pooling on the gradient of the c-class scores and the feature map of ; f(•) is the ReLU activation function, which is used to emphasize the features that have a positive impact on the category of interest. The second deep learning model is the Swin Transformer network; based on the test set, the trained first deep learning model and the trained second deep learning model are tested, and the deep learning model with high classification accuracy is selected and output as the final deep learning model. Obtain the vibration signal of the bearing to be diagnosed, convert the vibration signal of the bearing to be diagnosed into a two-dimensional image, input the two-dimensional image of the bearing to be diagnosed into the final deep learning model, and output the bearing fault diagnosis result.

2. The bearing fault diagnosis method according to claim 1, characterized in that The conversion of the bearing vibration signal through the Gram angular field algorithm to obtain a two-dimensional image of the bearing vibration signal specifically includes: S101-a1: Normalize the original bearing vibration signal , and perform normalization processing. Use equation (1) to scale it to the range of [0, 1]; (1) Among them, is the i-th value in is the scaled value; S101-a2: The scaled timestamp and amplitude are converted into the radius in polar coordinates and the cosine of the angle , forming a new time series , and the formula is as follows: (2) Among them, is the timestamp, is the regularization constant factor in polar coordinates; S101-a3: Calculate the correlation between any two data points through trigonometric functions; Equation (3) converts the time series into an image: (3) Among them, is a unit row vector, is the transposed vector of.

3. The bearing fault diagnosis method according to claim 1, characterized in that The conversion of the bearing vibration signal through the wavelet transform algorithm to obtain a two-dimensional image of the bearing vibration signal specifically includes: (4) (5) where a is a scaling factor; is a translation factor; is a wavelet basis function.

4. Bearing fault diagnosis system, characterized in that, Including: An acquisition module configured to: obtain a data set, where the data set is a two-dimensional image with known bearing fault type labels, and the two-dimensional image is obtained by converting the bearing vibration signal; the data set is divided into a training set and a test set. The two-dimensional image is obtained by converting the bearing vibration signal, including: Convert the bearing vibration signal through the Gram angular field algorithm to obtain a two-dimensional image of the bearing vibration signal; convert the bearing vibration signal through the wavelet transform algorithm to obtain a two-dimensional image of the bearing vibration signal; A training module configured to: train the first deep learning model and the second deep learning model respectively based on the training set to obtain the trained first deep learning model and the trained second deep learning model. Among them, the first deep learning model includes: the Conv1 layer of the Resnet34 network, the Conv2_x of the Resnet34 network, the Conv3_x layer of the Resnet34 network, the Conv4_x layer of the Resnet34 network, the Conv5_x layer of the Resnet34 network, an interpretable layer, an average pooling layer, and a fully connected layer connected in sequence. The interpretable layer refers to the Gradient-weighted Class Activation Mapping (Grad-CAM); Class discrimination and positioning map The calculation is as follows: (6) Among them, refers to the k-th feature map of the CNN layer; is obtained by combining the gradient of the c-class scores with Y of the feature map A; is obtained by performing global average pooling on the gradient of the c-class scores and the feature map of ; f(•) is the ReLU activation function, which is used to emphasize the features that have a positive impact on the category of interest. The second deep learning model is the Swin Transformer network; based on the test set, the trained first deep learning model and the trained second deep learning model are tested, and the deep learning model with high classification accuracy is selected as the final deep learning model output; A diagnostic module, configured to: acquire a vibration signal of a bearing to be diagnosed, convert the vibration signal of the bearing to be diagnosed into a two-dimensional image, input the two-dimensional image of the bearing to be diagnosed into a final deep learning model, and output a bearing fault diagnosis result.

5. An electronic device, characterized by comprising: a memory for non-temporarily storing computer-readable instructions; and a processor for running the computer-readable instructions, wherein, when the computer-readable instructions are run by the processor, the method according to any one of claims 1-3 above is executed.

6. A storage medium, characterized in that, Non-temporarily store computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method according to any one of claims 1-3 are executed.

Citation Information

Patent Citations

  • A Deep Learning-Based Fault Diagnosis Method for Mechanical Equipment

    CN111813084B

  • Rolling bearing fault diagnosis method

    CN115705396A