Bearing fault diagnosis method based on structural similarity generative adversarial network under small sample
The wavelet threshold denoising and structural similarity generation adversarial network generates high-quality auxiliary samples, which solves the problems of sample scarcity and overfitting in bearing fault diagnosis, and achieves high-precision fault recognition effect. It is suitable for fault monitoring and early warning systems in industrial scenarios.
Patent Information
- Application Number
- CN202510535664.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
The existing bearing fault diagnosis methods are insufficient in the condition of scarce sample number, the quality of the generated samples is not high, and the training process is easy to overfit, making it difficult to effectively apply in actual industrial environments.
Wavelet threshold denoising processing is used to convert one-dimensional vibration signals into two-dimensional GADF images, combine structural similarity generation adversarial network (SSGAN) to generate high-quality auxiliary samples, and filter through structural similarity index (SSIM) to build a deep convolutional neural network (DCNN) for training, and optimize the training set composition.
High-precision fault identification is achieved under small sample conditions, with a diagnostic accuracy of more than 96%, which improves the generalization ability of the model and the quality of the generated samples, and solves the diagnostic bottleneck in the small sample environment.
Smart Images

Figure CN120448813A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault diagnosis, and in particular to a small-sample bearing fault diagnosis method that combines a generative adversarial network and a structural similarity index, and belongs to the interdisciplinary technical field of intelligent operation and maintenance of mechanical equipment and artificial intelligence. Background Art
[0002] Existing rolling element bearing fault diagnosis methods rely on large amounts of labeled fault data, using deep learning models to extract and classify fault features. However, in real industrial applications, due to the low failure probability of high-reliability equipment, fault samples are often scarce, resulting in insufficient training and reduced diagnostic performance. While existing methods have attempted to mitigate the small sample size issue through signal enhancement or transfer learning, these methods suffer from issues such as low sample quality and model overfitting. Therefore, a high-quality data augmentation and screening mechanism is urgently needed to improve fault identification capabilities under small sample size conditions. Summary of the Invention
[0003] The purpose of this invention is to address existing bearing fault diagnosis technologies, which suffer from insufficient diagnostic accuracy, low sample quality, and overfitting during training, especially in real-world industrial environments where bearing failure rates are relatively low. Acquiring large amounts of real-world fault data is a significant challenge, and existing methods such as traditional data augmentation, meta-learning, and transfer learning are limited by high complexity and insufficient sample generation efficiency.
[0004] To this end, the present invention proposes a bearing fault diagnosis method based on structural similarity generative adversarial network, which aims to make full use of the image generation capability of generative adversarial network. At the same time, a structural similarity evaluation mechanism is introduced to screen auxiliary samples, thereby improving the quality of generated samples, optimizing the composition of training sets, and improving the diagnostic performance and generalization ability of the model in a small sample environment.
[0005] The method comprises the following steps:
[0006] (1) The one-dimensional vibration signal of the bearing during operation is collected and denoised using a wavelet threshold denoising method. Specifically, the method includes: performing multi-scale wavelet decomposition on the signal, applying a soft threshold function to compress the high-frequency detail coefficients, and obtaining a denoised smooth signal through wavelet reconstruction.
[0007] (2) Slice the denoised vibration signal to obtain multiple fixed-length signal segments, and use GADF transformation to convert the one-dimensional signal into a two-dimensional image for each segment, and divide it into training set and test set according to the proportion;
[0008] (3) The small sample training set is used as the input of the SSGAN model, and the real samples are trained to generate auxiliary image samples. The generated samples are compared with the real images to calculate the structural similarity index (SSIM), and images with similarity below the set threshold are eliminated, and only high-quality samples are retained for subsequent training.
[0009] (4) The screened auxiliary training samples are combined with the original sample set as the input data of the DCNN model to perform supervised learning training on various bearing states;
[0010] (5) Apply the trained model to the test set and output the corresponding fault status label to achieve automatic fault identification in a small sample environment.
[0011] The wavelet threshold denoising process adopts Daubechies wavelet basis to perform three-layer or four-layer wavelet decomposition, and the detail coefficient threshold processing method is soft threshold compression.
[0012] The GADF in step (1) converts a one-dimensional time series into a two-dimensional image through polar coordinate transformation and trigonometric function mapping. Each pixel value is composed of the cosine of the angle difference:
[0013]
[0014] Where x i is the normalized time series signal point; is the polar coordinate of the angle cosine.
[0015] The SSIM threshold is set between 0.85 and 0.95 to filter out auxiliary samples that have large structural differences from the real image. The formula for calculating the structural similarity index is:
[0016] S SSIM (x, y) = [l(x, y)] α *[c(x,y)] β *[s(x,y)] γ
[0017]
[0018]
[0019] Where l(x, y), c(x, y), s(x, y) are brightness similarity, contrast similarity, and structure similarity, respectively; x, y are the pixel values of the two images; * is the convolution operation; μ x , μ y are the means of x and y respectively; σ x ,σ y are the variances of x and y respectively; σ xyis the covariance of x and y; C1, C2, and C3 are constants. α, β, and γ are usually set to 1.
[0020] In the generative adversarial network, the generator uses the ReLU activation function (the output layer uses the tanh function), and the discriminator uses the LeakyReLU activation function (the output layer uses the Sigmoid function).
[0021] The generator and discriminator are both convolutional neural network structures, constructed using Conv2D two-dimensional convolution, and do not contain fully connected layers. Global mean pooling is used instead of full connection.
[0022] The DCNN network includes 4 convolutional layers, 4 pooling layers, 2 fully connected layers and 1 Softmax output layer.
[0023] The main innovative features of the present invention include:
[0024] (1) GADF is used for image conversion of bearing vibration signals to maintain the structural information of the time series and enhance the expressiveness of image features. Compared with traditional spectrograms or grayscale images, GADF retains more "geometric relationships" between time series features, greatly improving the perception efficiency of deep network models, and enhancing the discriminability and diversity of generated samples, especially showing stronger generalization ability under small sample conditions.
[0025] (2) Combine SSIM for auxiliary sample screening: Based on the generative adversarial network, a structural similarity evaluation mechanism is introduced. The generated images are evaluated at the brightness, contrast and structure levels through SSIM, and images that are significantly different from the real samples are automatically eliminated, thereby improving the reliability and representativeness of the auxiliary training data.
[0026] (3) Network structure optimization design: Compared with traditional generative adversarial networks, the proposed SSGAN has optimized and innovated its network structure. By removing the fully connected layer, adopting global pooling, and using activation functions and regularization methods that are more suitable for image tasks, the model's generation capability and training stability are effectively improved, ultimately significantly improving the performance of the fault diagnosis model in a small sample environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 : Schematic diagram of converting one-dimensional vibration signal into GADF diagram.
[0028] Figure 2 : Schematic diagram of SSGAN network flow.
[0029] Figure 3 : Fault diagnosis flow chart based on the present invention.
[0030] Figure 4: t-SNE visualization clustering effect diagram after feature dimensionality reduction. DETAILED DESCRIPTION
[0031] During the signal processing phase, accelerometers were used to collect vibration signals from industrial equipment during operation, sampling at a frequency of 12 kHz. The resulting one-dimensional time-domain signal served as the raw data input. To eliminate high-frequency interference and background noise and improve subsequent image representation, the raw signal was first subjected to wavelet threshold denoising. Using the Daubechies-4 (db4) wavelet basis, the signal was subjected to a 3-5-layer discrete wavelet transform to obtain approximate coefficients and multiple layers of detail coefficients. A soft threshold function was applied to each layer of detail coefficients, and the compressed detail coefficients and the original approximate coefficients were reconstructed into a denoised time series signal.
[0032] The denoised time series signals are segmented into multiple windows to generate a large number of signal segments for image conversion. All segmented time series signals are normalized using the Z-score, mapping them to the [-1, 1] interval to eliminate the influence of differences in the original vibration amplitude. Each normalized sequence is then mapped into an angle sequence, and a GADF matrix is constructed to convert the intercepted sample data into a two-dimensional image.
[0033] Auxiliary sample generation and structural similarity screening are performed through SSGAN. First, real GADF image samples are input into SSGAN for training to generate a large number of auxiliary image samples. The SSIM is calculated for each generated image and its corresponding real image to evaluate their similarity in brightness, contrast and structure. Auxiliary images that exceed the SSIM threshold are screened out, and only high-quality auxiliary samples are retained for model training.
[0034] The generator structure parameters are shown in the following table:
[0035]
[0036] The discriminator structure parameters are shown in the following table:
[0037]
[0038] Among them, convolution operations with a step size greater than 1 are used to replace traditional spatial pooling layers to achieve feature downsampling without losing image resolution, effectively reducing information loss; the entire generator architecture does not contain a fully connected structure, which is replaced by a global mean pooling operation to reduce the risk of overfitting and improve network training stability; the choice of activation function takes into account both training stability and nonlinear expression capabilities; all convolutional layers include BatchNorm regularization by default, which effectively enhances the network generalization ability and training speed; the discriminator outputs a probability in the interval [0,1], which is used to determine whether the image is real.
[0039] Here, a 100-dimensional random vector is used as the generator input, and a real GADF image is used as the discriminator input. Auxiliary images are generated using a GAN. The Adam optimizer is used with an initial learning rate of 0.0002, a batch size of 64, and 100 iterations. The discriminator and generator are trained alternately to maintain adversarial balance. A structural similarity index is then calculated between the generated images and their corresponding real images. Only images with a SSIM of ≥ 0.9 are retained for training, and images with significant quality differences are discarded.
[0040] Then, a deep convolutional neural network (DCNN) classifier was constructed, which includes 4 convolutional layers: containing 3×3 convolution kernels, the number of channels increases layer by layer (32→64→128→256), and all use ReLU activation; 4 pooling layers: 2×2 maximum pooling; 2 fully connected layers: 256→64 dimensions; the output layer uses a Softmax multi-classifier to output the bearing status category.
[0041] During training, only 25 real GADF images were used for each bearing fault condition. SSGAN was used to generate 1,000 high-quality auxiliary images, which were then combined with the real images to form a training set. The DCNN model then fed these images for recognition. The recognition accuracy exceeded 96% on 300 test images.
[0042] This paper proposes a bearing fault diagnosis method for small sample sizes. By integrating key technologies such as wavelet threshold denoising, time series signal image conversion, structural similarity generative adversarial network (SSGAN) sample enhancement, and deep convolutional neural network (DCNN) classification, it establishes an end-to-end, well-structured fault identification process. This method overcomes the diagnostic bottleneck in small sample sizes and offers significant advantages in generated data quality, diagnostic accuracy, and algorithm stability, promising promising industrial applications.
Claims
1. A bearing fault diagnosis method based on structural similarity generative adversarial network under small sample conditions, characterized by The following steps are involved: (1) The one-dimensional vibration signal of the bearing during operation is collected and denoised using a wavelet threshold denoising method. Specifically, the method includes: performing multi-scale wavelet decomposition on the signal, applying a soft threshold function to compress the high-frequency detail coefficients, and obtaining a denoised smooth signal through wavelet reconstruction. (2) Slice the denoised vibration signal to obtain multiple fixed-length signal segments, and use GADF transformation to convert the one-dimensional signal into a two-dimensional image for each segment, and divide it into training set and test set according to the proportion; (3) The small sample training set is used as the input of the SSGAN model, and the real samples are trained to generate auxiliary image samples. The generated samples are compared with the real images to calculate the structural similarity index (SSIM), and images with similarity below the set threshold are eliminated, and only high-quality samples are retained for subsequent training. (4) The screened auxiliary training samples are combined with the original sample set as the input data of the DCNN model to perform supervised learning training on various bearing states; (5) Apply the trained model to the test set and output the corresponding fault status label to achieve automatic fault identification in a small sample environment.
2. The method according to claim 1, wherein: The wavelet threshold denoising process adopts Daubechies wavelet basis to perform three-layer or four-layer wavelet decomposition, and the detail coefficient threshold processing method is soft threshold compression.
3. The method according to claim 1, wherein: The GADF in step (1) converts a one-dimensional time series into a two-dimensional image through polar coordinate transformation and trigonometric function mapping; each pixel value is composed of the cosine of the angle difference: Where x i is the normalized time series signal point; is the polar coordinate of the angle cosine.
4. The method according to claim 1, wherein: The SSIM threshold is set between 0.85 and 0.95 to filter out auxiliary samples with large structural differences from the real image; The formula for calculating the structural similarity index is: S SSIM (x,y)=[l(x,y)] α *[c(x,y)] β *[s(x,y)] γ Where l(x, y), c(x, y), s(x, y) are brightness similarity, contrast similarity, and structure similarity, respectively; x, y are the pixel values of the two images; * is the convolution operation; μ x , μ y are the means of x and y respectively; σ x , σ y are the variances of x and y respectively; σ xy is the covariance of x and y; C1, C2, C3 are constants; α, β, γ are usually set to 1.
5. The method according to claim 1, wherein: In the generative adversarial network, the generator uses the ReLU activation function (the output layer uses the tanh function), and the discriminator uses the LeakyReLU activation function (the output layer uses the Sigmoid function).
6. The method according to claim 1, wherein: The generator and discriminator are both convolutional neural network structures, constructed using Conv2D two-dimensional convolution, and do not contain fully connected layers. Global mean pooling is used instead of full connection.
7. The method according to claim 1, wherein: The DCNN network includes 4 convolutional layers, 4 pooling layers, 2 fully connected layers and 1 Softmax output layer.