Small sample bearing fault diagnosis method based on deep residual network and bidirectional GRU
By combining a deep residual network and bidirectional GRU approach with Gramian Angular Field and short-time Fourier transform to generate multi-channel input tensors, the problem of insufficient small-sample learning and feature extraction in bearing fault diagnosis is solved, and efficient fault identification and diagnosis in complex industrial environments is achieved.
Patent Information
- Application Number
- CN202511116178.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for bearing fault diagnosis suffer from insufficient small-sample learning ability, limitations in diagnostic performance due to single feature extraction methods, limited ability to capture temporal features, and inflexible model structure, resulting in insufficient applicability and accuracy of the models in complex industrial environments.
A method based on deep residual networks and bidirectional GRUs is adopted. Multi-channel input tensors are generated through Gramian Angular Field and short-time Fourier transform. Combined with ResNet and BiGRU models, temporal and spatial features are captured. End-to-end training is carried out using transfer learning and cross-entropy loss function to achieve fault diagnosis under small sample conditions.
It effectively improves the accuracy and robustness of bearing fault diagnosis, can capture the dynamic changes and temporal correlations of complex signals under small sample conditions, adapts to complex industrial environments, and improves the model's generalization ability and fault identification ability.
Smart Images

Figure CN120974113A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bearing fault diagnosis, and particularly relates to a small sample bearing fault diagnosis method based on a deep residual network and a bidirectional GRU. BACKGROUND
[0002] With the development of industrial automation and intelligent manufacturing, the health monitoring and fault diagnosis technology of mechanical equipment becomes particularly important. Bearings, as the key vulnerable parts in mechanical equipment, their running state is directly related to the safety and production efficiency of the equipment. Through the analysis of bearing vibration signals, early identification and classification of bearing faults can be achieved, thereby effectively avoiding accidents and reducing maintenance costs.
[0003] In recent years, with the rapid development of deep learning technology, fault diagnosis methods based on convolutional neural networks (CNN) and recurrent neural networks (RNN) have become a research hotspot. These methods can automatically extract features, avoiding the drawbacks of traditional methods that rely on manual feature design, significantly improving the accuracy and efficiency of diagnosis. Time-frequency analysis techniques (such as short-time Fourier transform STFT, SFFT) and signal image conversion methods (such as Gramian Angular Field, GAF) are widely used in feature extraction of vibration signals, which can reveal the time-varying frequency structure of the signal from different angles and improve the expression ability of fault features.
[0004] Although the existing technology has made some progress in the field of bearing fault diagnosis, there are still the following main deficiencies:
[0005] Insufficient small sample learning ability, existing deep learning models usually rely on a large amount of labeled data for training, while high-quality fault data in industrial sites is difficult to collect in large quantities, especially for some rare fault types, leading to model overfitting and poor generalization ability, affecting the actual application effect;
[0006] Single feature extraction method limits diagnostic performance, traditional methods often use only single features in time or frequency domain, or simple time-frequency analysis, without fully integrating multi-angle and multi-scale signal information, resulting in insufficient expression of fault features and affecting classification accuracy;
[0007] Limited time series feature capture ability, most CNN-based diagnosis methods focus on spatial feature extraction and have weak dependence on time series modeling of vibration signals, which cannot effectively capture the dynamic changes and time series correlation of signals, reducing the recognition ability of complex fault patterns;
[0008] The model structure is not flexible enough, and the existing part of the model structure is single, lacking of collaborative modeling of time series and spatial features, which is difficult to adapt to the characteristics of signal diversification and non-stationarity, limiting the applicability of the model in complex industrial environments. SUMMARY
[0009] The application aims to provide a small sample bearing fault diagnosis method based on a deep residual network and a bidirectional GRU, so as to solve the above-mentioned problems of the prior art.
[0010] To achieve the above-mentioned purpose, the application adopts the following technical solutions:
[0011] A small sample bearing fault diagnosis method based on a deep residual network and a bidirectional GRU comprises the following steps:
[0012] S1, collecting a bearing original vibration signal and performing preprocessing;
[0013] S2, converting the preprocessed signal into two-dimensional feature representations of two different modalities respectively, and splicing the two two-dimensional feature representations according to the channel dimension to form a multi-channel input tensor;
[0014] Among them, the first modality is used to represent the global time sequence correlation of the signal, and the second modality is used to represent the time-frequency dynamic characteristics of the signal;
[0015] S3, inputting the multi-channel input tensor into a ResNet-BiGRU model, extracting spatial features through an improved deep residual network (ResNet) and outputting a feature map, and converting the feature map into sequence data to input a bidirectional gated recurrent unit (BiGRU) to capture the time sequence dependence relationship, and generating a feature vector based on the final time step feature output by the bidirectional gated recurrent unit;
[0016] S4, outputting a bearing fault category probability distribution according to the feature vector, and determining a diagnosis result.
[0017] Further, the preprocessing of the bearing original vibration signal adopts normalization processing:
[0018]
[0019] In formula (1), x' represents the vibration signal after normalization processing, the value range is [0, 1]; i = 1, 2, …, N, represents a sampling point, wherein N represents the signal length; x = {x1, x2, …, xN} represents the bearing original vibration signal; max(x) and min(x) respectively represent the maximum value and the minimum value of all sampling points of the vibration signal. N
[0020] Further, step S2 comprises:
[0021] S21, mapping the vibration signal after normalization processing into a polar coordinate form, and generating a two-dimensional angle amplitude image of the first modality through Gramian Angular Field (GAF) conversion:
[0022] G i,j =cos(φ i +φ j ) (2);
[0023] φ i =arccos(x’ i ) (3);
[0024] In formula (2)-(3): j is same as i, representing a sampling point; G i,j represents a GAF image with a size of N*N, wherein each element represents a cosine value of the polar angle sum of the signal at the sampling points i and j;
[0025] S22, applying a short-time Fourier transform to the normalized vibration signal to generate a time-frequency amplitude map of the second mode:
[0026]
[0027] In formula (4): S(t,f) represents an amplitude value at time t and frequency f; x' m represents the mth sampling point of the normalized signal; w represents a window function, such as a Hanning window; and M represents a window length; is an imaginary unit;
[0028] S23, the calculated time-frequency amplitude map is normalized and spliced with the two-dimensional angle amplitude image according to the channel dimension to form a double-channel input tensor.
[0029] Further, the improved residual network: the first layer convolution kernel channel number is adapted to the dimension of the multi-channel input tensor; the network parameters are initialized by loading the ImageNet pre-training weight through transfer learning.
[0030] Further, the improved residual network extracts spatial features in the input feature map through multiple residual blocks, and the output feature map is:
[0031]
[0032] In formula (5): F represents the output feature map; B represents the batch size; C represents the channel number, that is, the feature dimension; and H and W represent the height and width of the feature map, respectively.
[0033] Further, the conversion of the feature map into sequence data is specifically represented by the following formula (6):
[0034]
[0035] In formula (6): F seq represents the sequence form of the feature map, and each feature map position corresponds to a time step of the sequence.
[0036] Further, the final time step feature output by the bidirectional gated recurrent unit is composed of the forward hidden state and the reverse hidden state spliced together:
[0037]
[0038] In formula (7), h t represents the final time step feature; represents the input feature at the tth time step; GRUCell represents the recursive calculation of the gated recurrent unit; represents the forward hidden state; represents the reverse hidden state.
[0039] Further, step S4 comprises:
[0040] S41, taking the bidirectional hidden state at the last time step of the sequence data as the sample feature representation, calculating the bearing fault class probability through a fully connected layer and a Softmax classification layer:
[0041] y = Softmax (Wh L +b) (8);
[0042] In formula (8), y represents the predicted probability distribution of the model; represents the classification layer weight matrix; h L represents the bidirectional hidden state at the last time step of the sequence data; represents the bias; C out represents the probability of the fault class.
[0043] S42, selecting the fault class corresponding to the maximum probability as the diagnosis result.
[0044] Further, the training and optimization of the ResNet-BiGRU model comprises:
[0045] The Adam optimization algorithm is used to iteratively update the model parameters, and the initial learning rate is set to 1x10 -4 and the cosine annealing scheduling is adopted, and the batch size is not more than 16 to adapt to small sample training, and the learning rate and the number of training rounds are adjusted in real time according to the results;
[0046] The loss function adopts the cross-entropy function:
[0047]
[0048] In formula (9), L represents the cross-entropy loss function; y c represents the true label; is the one-hot encoding of the true label.
[0049] Compared with the prior art, the application has the following advantages:
[0050] 1、The application converts one-dimensional vibration signals into two-dimensional images through Gramian Angular Field, retains the global angle relationship of time series by polar coordinate mapping, forms a diagonal explicit Gram matrix, and captures long-term time sequence dependence characteristics; meanwhile, short-time fast Fourier transform is adopted to generate a time-frequency graph, and a sliding window Fourier transform is used to capture transient frequency characteristics and non-stationary signal characteristics; the GAF image and the SFFT time-frequency graph are spliced in the channel dimension to construct a space-time-frequency domain joint feature, the problem of time domain and frequency domain feature separation is solved, and complementary feature expression is formed.
[0051] 2、The application improves the input layer convolution kernel structure of ResNet, supports double-channel image input, solves the deep network gradient disappearance problem through residual jump connection, maintains high representation ability under small sample conditions, and the feature reuse mechanism of ResNet significantly improves the feature utilization rate.
[0052] 3、The application converts the spatial feature map output by ResNet into a sequence form, inputs a bidirectional gated recurrent unit (BiGRU), cooperates the forward-backward hidden state at the same time, captures the history and future dependence relationship of the fault signal, and overcomes the time sequence modeling limitation of traditional CNN and unidirectional RNN.
[0053] 4、The application initializes ResNet by using ImageNet pre-training weight, alleviates the overfitting risk of small sample training through parameter migration, adopts end-to-end training, constructs a complete process from signal preprocessing to classification output, adopts a joint optimization objective function, and avoids the error accumulation caused by the separation of feature extraction and classifier in traditional methods. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 It is a step flow diagram of the fault diagnosis method of the application.
[0055] Figure 2 It is a system structure diagram of the diagnosis method of the application. DETAILED DESCRIPTION
[0056] A preferred embodiment of the application will be described in detail below with reference to the accompanying drawings.
[0057] As shown in Figure 1 and 2 , a small sample bearing fault diagnosis method based on a deep residual network and a bidirectional GRU specifically includes the following steps:
[0058] S1, collect the original vibration signal of the bearing and perform preprocessing.
[0059] Specifically, the original vibration signal of the bearing can be collected by an acceleration sensor. In order to eliminate the influence of amplitude difference on subsequent processing, the pre-processing of the original vibration signal of the bearing adopts normalization processing:
[0060]
[0061] In formula (1), x' represents the normalized vibration signal, the value range of which is [0, 1]; i = 1, 2, …, N, representing the sampling point, wherein N represents the signal length; x = {x1, x2, …, xN} represents the original vibration signal of the bearing; max(x) and min(x) respectively represent the maximum value and the minimum value in all sampling points of the vibration signal. N
[0062] S2, the pre-processed signal is converted into two different modal two-dimensional feature representations respectively, and the two two-dimensional feature representations are spliced according to the channel dimension to form a multi-channel input tensor.
[0063] In this step, the first modal is used to represent the global time sequence correlation of the signal, and the second modal is used to represent the time-frequency dynamic characteristics of the signal. Specifically, step S2 includes:
[0064] S21, the normalized vibration signal is mapped into a polar coordinate form, and a two-dimensional angle amplitude image of the first modal is generated by Gramian Angular Field (GAF) conversion; the global angle relationship of the time sequence is retained by polar coordinate mapping, a diagonal explicit Gram matrix is formed to capture long-term time sequence dependence features; at the same time, by converting the time domain signal into a two-dimensional angle amplitude image, the global angle information of the signal is retained, and the spatial representation of the time sequence is strengthened:
[0065] G i,j = cos (φ i + φ j ) (2);
[0066] φ i = arccos (x' i ) (3);
[0067] In formula (2)-(3), j is the same as i, representing the sampling point; G i,j represents a GAF image with a size of N×N, wherein each element represents the cosine value of the polar angle of the signal at sampling points i and j;
[0068] S22, short-time Fourier transform is applied to the normalized vibration signal to generate a time-frequency amplitude image of the second modal, and the local time-frequency features of the signal are extracted to reflect the change of the frequency component of the signal with time:
[0069]
[0070] In formula (4), S(t, f) represents the amplitude value of frequency f at time t; x' m represents the mth sampling point of the normalized signal; w represents a window function, such as a Hanning window, which captures transient frequency characteristics and non-stationary signal characteristics through a sliding window Fourier transform, and is particularly sensitive to weak impact responses of early faults; and M represents the window length; is a unit imaginary number;
[0071] S23, the time-frequency amplitude map calculated after normalization is spliced with the two-dimensional angle amplitude image according to the channel dimension to form a double-channel input tensor.
[0072] The size of the double-channel input tensor is [2, N, N][2, N, N], indicating that the sizes of the two channel images are both N x N. Specifically, the GAF image focuses on time sequence dependent modeling, and the SFFT image reflects frequency distribution characteristics, and the combination of the two can comprehensively depict the dynamic characteristics of the fault signal, effectively distinguish different types of bearing faults, and enhance the sensitivity to non-stationary signals and early faults. At the same time, SFFT has good response ability to short-time frequency disturbances in the signal, and GAF can identify the detailed structure of periodic changes, and the combination of the two can effectively improve the detection accuracy of the model under complex working conditions. The preferred embodiment realizes the complementary fusion of time sequence characteristics and frequency domain characteristics and the collaborative learning of time domain and frequency domain information through the double-channel fusion mode, and provides more rich representation information for the model.
[0073] The fused double-channel input tensor can unify image input to adapt to a deep neural network: the double-channel image structure is naturally compatible with the CNN-ResNet model, supports end-to-end training, effectively avoids the problem of disconnection between traditional feature extraction and classification, and improves the automatic learning ability. At the same time, it can adapt to small sample training and improve the robustness of the model: the image-based time-frequency features are more stable in the spatial structure, which helps the deep network to extract more discriminative features under small sample conditions and reduces the risk of overfitting.
[0074] S3, inputting the multi-channel input tensor into the ResNet-BiGRU model, extracting spatial features through the improved deep residual network (ResNet) and outputting a feature map, and converting the feature map into sequence data to input into a bidirectional gated recurrent unit (BiGRU) to capture time sequence dependency, and generating a feature vector based on the final time step feature output by the bidirectional gated recurrent unit.
[0075] The improved residual network in the preferred embodiment includes:
[0076] The first layer convolution kernel channel number adapts the dimension of the multi-channel input tensor, preserves the residual connection structure, and solves the gradient vanishing problem of the deep network.
[0077] The network parameters are initialized by loading the ImageNet pre-training weights through transfer learning, which significantly improves the classification accuracy and generalization ability under small samples. The deep residual structure enhances the network expression ability while avoiding overfitting. The transfer learning strategy can significantly reduce the demand for training data volume, improve the generalization ability under small samples, and support end-to-end diagnosis process, avoiding error accumulation caused by separation of feature extraction and classifier in traditional methods, and improving the accuracy of automatic feature extraction and fault identification of the model.
[0078] By constructing a ResNet model supporting multi-channel input and fine-tuning parameters through transfer learning, automatic identification and classification of bearing fault types under small sample data are realized. The final output diagnosis result is, for example, normal, inner ring fault, outer ring fault, rolling element fault, etc. Specifically, the improved residual network extracts spatial features in the input feature map through multiple residual blocks, and the output feature map is:
[0079]
[0080] In formula (5), F represents the output feature map, B represents the batch size, C represents the channel number, i.e., the feature dimension, and H and W represent the height and width of the feature map, respectively.
[0081] Further, the conversion of the feature map into sequence data is specifically represented by the following formula (6):
[0082]
[0083] In formula (6), F seq represents the sequence form of the feature map, and each feature map position corresponds to a time step of the sequence.
[0084] The final time step feature output by the BiGRU in the preferred embodiment is composed of the forward hidden state and the backward hidden state. By converting the spatial feature map output by the ResNet into a sequence form and then inputting it into the BiGRU unit for time series modeling, the history and future time dependence of the fault signal can be captured simultaneously, the modeling capability is stronger than that of the traditional RNN or unidirectional GRU, and the dynamic feature extraction for long sequences is more robust, suitable for processing non-stationary signals, and makes up for the problem of insufficient time series modeling capability of the traditional CNN structure:
[0085]
[0086] In formula (7), h t represents the final time step feature. represents the input feature of the t-th time step; GRUCell represents the recursive calculation of the Gated Recurrent Unit (GRU); represents the forward hidden state; represents the reverse hidden state.
[0087] S4, output the bearing fault class probability distribution according to the feature vector, and determine the diagnosis result.
[0088] Specifically, step S4 includes:
[0089] S41, taking the bidirectional hidden state of the last time step of the sequence data as a sample feature representation, calculating the bearing fault class probability through a fully connected layer and a Softmax classification layer:
[0090] y = Softmax (Wh L +b) (8);
[0091] In formula (8): represents the predicted probability distribution of the model; represents the weight matrix of the classification layer; h L represents the bidirectional hidden state of the last time step of the sequence data; represents the bias; C out represents the probability of the fault class.
[0092] S42, selecting the fault class corresponding to the maximum probability as the diagnosis result.
[0093] In specific operation, in order to further improve the accuracy of model fault recognition, the following training and optimization are adopted:
[0094] The Adam optimization algorithm is used to iteratively update the model parameters, and the initial learning rate is set to 1x10 -4 and cosine annealing scheduling is adopted, and the batch size is not more than 16 to adapt to small sample training, and the learning rate and the number of training rounds are adjusted in real time according to the results;
[0095] The loss function adopts a cross-entropy function:
[0096]
[0097] In formula (9), L represents the cross-entropy loss function; y c represents the true label; is the one-hot encoding of the true label.
[0098] The above described embodiments are merely intended to describe the preferred embodiments of the present application, and are not intended to limit the scope of the present application, and various modifications and improvements made by those skilled in the art to the technical solutions of the present application without departing from the design spirit of the present application shall fall within the protection scope of the present application as defined by the claims.
Claims
1. A small-sample bearing fault diagnosis method based on deep residual networks and bidirectional GRU, characterized in that, Includes the following steps: S1. Acquire the original vibration signal of the bearing and perform preprocessing; S2. The preprocessed signal is converted into two different modal two-dimensional feature representations, and the two two-dimensional feature representations are concatenated according to the channel dimension to form a multi-channel input tensor. The first mode is used to characterize the global temporal correlation of the signal, and the second mode is used to characterize the time-frequency dynamic characteristics of the signal. S3. Input the multi-channel input tensor into the ResNet-BiGRU model, extract spatial features and output feature maps through the improved deep residual network (ResNet), and convert the feature maps into sequence data and input them into the bidirectional gated recurrent unit (BiGRU) to capture temporal dependencies. At the same time, generate feature vectors based on the final time step features output by the bidirectional gated recurrent unit. S4. Output the bearing fault category probability distribution based on the feature vector and determine the diagnosis result.
2. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 1, characterized in that, The preprocessing of the original vibration signal of the bearing employs normalization: In formula (1): x′ represents the normalized vibration signal with a range of [0,1]; i=1,2,…,N represents the sampling points, where N represents the signal length; x={x1,x2,…,x N } represents the original vibration signal of the bearing; max(x) and min(x) represent the maximum and minimum values among all sampling points of the vibration signal, respectively.
3. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 2, characterized in that, Step S2 includes: S21. Map the normalized vibration signal to polar coordinates and generate a two-dimensional angular amplitude image of the first mode using Gramian Angular Field (GAF) transformation: G i,j =cos(φ i +φ j ) (2); f i =arccos(x' i ) (3); In formulas (2)-(3): j is the same as i, representing the sampling point; G i,j This represents a GAF image of size N×N, where each element represents the cosine of the sum of the polar angles of the signal at sampling points i and j; S22. Apply a short-time Fourier transform to the normalized vibration signal to generate the time-frequency amplitude diagram of the second mode: In formula (4): S(t,f) represents the amplitude value at time t with frequency f; x' m This represents the m-th sampling point of the signal after normalization; w represents the window function, such as the Hanning window; M represents the window length. The imaginary unit; S23. After normalizing the calculated time-frequency amplitude image, it is stitched together with the two-dimensional angle amplitude image according to the channel dimension to form a dual-channel input tensor.
4. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 1, characterized in that, The improved residual network: the number of channels in the first layer convolutional kernel is adapted to the dimension of the multi-channel input tensor; the network parameters are initialized by loading ImageNet pre-trained weights through transfer learning.
5. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRUs according to claim 4, characterized in that, The improved residual network extracts spatial features from the input feature map through multiple layers of residual blocks, and the output feature map is: In formula (5): F represents the output feature map; B represents the batch size; C represents the number of channels, i.e. the feature dimension; H and W represent the height and width of the feature map, respectively.
6. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 5, characterized in that, The conversion of feature maps into sequence data is specifically represented by the following formula (6): In formula (6): F seq It represents the sequence form of the feature map, with each feature map position corresponding to a time step in the sequence.
7. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 1, characterized in that, The final time step feature output by the bidirectional gated loop unit is composed of a forward hidden state and a reverse hidden state: In formula (7): h t Indicates the final time step features; represents the input feature at time step t; GRUCell represents the recursive computation of the gated recurrent unit; Indicates a positive hidden state; This indicates a reverse hidden state.
8. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 7, characterized in that, Step S4 includes: S41. Take the bidirectional hidden state of the last time step of the sequence data as the sample feature representation, and calculate the bearing failure category probability through a fully connected layer and a Softmax classification layer: y=Softmax(Wh L +b) (8)? In formula (8): This represents the predicted probability distribution of the model; h represents the classification layer weight matrix; L This represents the bidirectional hidden state of the sequence data at the last time step; Indicates bias; C out Indicates the probability of a fault category; S42. Select the fault category with the highest probability as the diagnostic result.
9. The method for small-sample bearing fault diagnosis based on deep residual networks and bidirectional GRU according to claim 1, characterized in that, The training and optimization of the ResNet-BiGRU model includes: The Adam optimization algorithm is used to iteratively update the model parameters, with an initial learning rate of 1×10⁻⁶. -4 Cosine annealing scheduling is used, with a batch size not exceeding 16, to adapt to small sample training, and the learning rate and number of training rounds are adjusted in real time according to the results; The loss function used is cross-entropy: In formula (9): L represents the cross-entropy loss function; y c Indicates the true label; One-hot encoding for the real label.
Citation Information
Patent Citations
Rolling bearing fault identification method based on GAF-CNN-BiGRU network
CN112179654A
Rolling bearing fault diagnosis method based on small samples and GAF-DCGAN
CN114266339A
Rolling bearing fault diagnosis method based on GAF-DRSN
CN114595730A
Bearing fault diagnosis method based on weight adaptive feature fusion
CN115753101A
Bearing fault diagnosis method based on lightweight neural network and dimension expansion
CN115761398A
Cited By
A small sample rolling bearing fault diagnosis method based on federated learning
CN122551021A