Bearing fault diagnosis method based on wavelet transform and mixed attention mechanism
The convolutional neural network constructed through wavelet transformation and mixed attention mechanism solves the problems of low recognition rate and poor generalization in variable noise environments, and realizes accurate identification and high-precision diagnosis under complex conditions.
Patent Information
- Application Number
- CN202510341336.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing bearing fault diagnosis methods have low recognition rate, poor general use, low diagnostic accuracy, and difficult to adapt to complex industrial environments.
The convolutional neural network is constructed using wavelet transform and mixed attention mechanism. By learning convolution, the weight can be adjusted, combined with the residual structure, BiLSTM and self-attention mechanism, multi-scale features and long-distance dependencies are extracted, and the model parameters are optimized to adapt to the variable noise conditions.
In a variable noise environment, the recognition rate and generality of bearing fault diagnosis are improved, and more accurate fault diagnosis results are output, which improves the adaptability and accuracy of the model.
Smart Images

Figure CN120277463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bearing fault diagnosis, and particularly to a bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism. Background Art
[0002] In modern industrial production, rotating machinery has become the core equipment, and bearings, as its key components, are crucial for the stable operation and safety of the equipment. However, in actual operation, bearings are often affected by various internal and external factors, such as wear, high temperature and pressure, super-large load, and external impact. These factors can lead to the degradation or even failure of bearing performance. Once a bearing fails, it will not only reduce product quality and equipment operation efficiency, but may also trigger major safety accidents, causing huge economic losses and casualties. Therefore, ensuring the normal operation of rolling bearings is not only a technical requirement, but also a responsibility for equipment reliability and production safety.
[0003] Bearing fault diagnosis technology has emerged as the times require. Its core is to collect relevant data of mechanical equipment and use expert experience or intelligent diagnostic algorithms to monitor the equipment status in real time, so as to detect faults in time and take measures to reduce the risk of accidents.
[0004] In recent years, deep learning technology has been widely used in the field of bearing fault diagnosis, and has significantly improved the accuracy and efficiency of diagnosis. However, the successful application of deep learning is usually based on an assumption that the training data set and the test data set must come from the same noise conditions. But in the actual industrial environment, this assumption is often difficult to meet. On the one hand, the working conditions of rotating machinery change frequently, resulting in unstable data distribution; on the other hand, it is necessary to obtain enough vibration data and their corresponding fault labels to train models adapted to each working condition. However, the types of noise are also diverse, which makes it difficult for models under a single noise condition to adapt to the changing noise environment. In order to cover all possible fault situations under different noise conditions, collecting a large amount of fault data and separately constructing deep learning models for each noise condition not only requires huge economic investment, but also requires a large amount of human resources, and the implementation difficulty is extremely high.
[0005] Facing these challenges, researchers are committed to developing methods that can accurately diagnose bearing faults under complex noise and unbalanced data conditions. In recent years, although there are many bearing fault diagnosis methods based on deep learning, there are still some problems that cannot be accurately identified under complex conditions. Although traditional model methods can achieve good recognition rates under a single noise condition, they are often difficult to maintain a high recognition rate under changing noise conditions. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism to solve the problems of low recognition rate of faulty bearings in a variable noise environment, poor versatility of traditional bearing fault diagnosis methods, and low diagnostic accuracy of bearing faults.
[0007] To achieve the above technical objectives, the present invention provides the following technical solutions:
[0008] A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism specifically includes the following steps:
[0009] S1. Use a sensor to obtain the bearing vibration signal to obtain various types of original fault data;
[0010] S2. Perform wavelet transform, add Gaussian noise of different intensities, and block sampling on the obtained original fault data in sequence; and generate category labels and position labels for the sampled fault data;
[0011] S3. Based on the hybrid attention mechanism, residual structure, and learnable convolution, construct a fault diagnosis model based on the convolutional neural network CNN as the basic model;
[0012] S4. Divide the sampled fault data into a training set and a test set, and use them as the input data of the model for training, and monitor the training results in real time for parameter optimization until the optimal model is trained; and visualize the fault diagnosis results output by the model;
[0013] S5. Conduct a comparative experiment with the existing model under various noise conditions to test the performance of the constructed fault diagnosis model.
[0014] By means of the above technical solutions, the present invention has at least the following beneficial effects:
[0015] 1. The method proposed by the present invention denoises the noisy signal through wavelet transform, and automatically adjusts the weight by using learnable convolution. When the noise intensity or noise frequency changes, the weight is automatically adjusted to enhance the ability to extract effective features under noise interference, and the parameters are continuously optimized;
[0016] 2. Design a residual parallel module composed of three residual blocks with different structures, extract the key features that can best reflect bearing faults from three perspectives of capturing multi-scale features, enhancing important features, and capturing long-distance dependence relationships, and combine the advantages of BiLSTM and self-attention mechanism to capture global features to avoid falling into local optima; at the same time, the original information is retained through multiple residual structures to avoid feature loss during the extraction process;
[0017] 3. The method proposed by the present invention solves the problems of low recognition rate of faulty bearings in a variable noise environment, poor versatility of traditional bearing fault diagnosis methods, and relatively low diagnostic accuracy of bearing faults. It can output more accurate bearing fault diagnosis results, realizes the advantage of accurately identifying bearing faults by the model under variable noise conditions, and achieves the purpose of improving the versatility and accuracy rate of the bearing fault diagnosis model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is the overall flowchart of a bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism proposed by the present invention;
[0020] Figure 2 is the structural diagram of the fault diagnosis model designed by the method proposed by the present invention;
[0021] Figure 3 is the model framework diagram of the first residual block in the residual block parallel module;
[0022] Figure 4 is the model framework diagram of the second residual block in the residual block parallel module;
[0023] Figure 5 is the model framework diagram of the third residual block in the residual block parallel module;
[0024] Figure 6 is the line graph comparing the performance of the method proposed by the present invention with other models;
[0025] Figure 7 is the accuracy rate and loss graph of the model proposed by the present invention under the noise condition of 0.02;
[0026] Figure 8 is the accuracy rate and loss graph of the model proposed by the present invention under the noise condition of 0.04;
[0027] Figure 9 is the accuracy rate and loss graph of the model proposed by the present invention under the noise condition of 0.06;
[0028] Figure 10 is the accuracy rate and loss graph of the model proposed by the present invention under the noise condition of 0.08;
[0029] Figure 11 is the accuracy rate and loss graph of the model proposed by the present invention under the noise condition of 0.1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein involves and encompasses any and all possible combinations of one or more of the associated listed items.
[0032] As Figures 1-11 shown, a specific embodiment of the present invention is shown; as Figure 1 and 2 shown, a bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism proposed by the present invention specifically includes:
[0033] S1. Use a sensor to obtain the bearing vibration signal to obtain various types of original fault data;
[0034] As a preferred embodiment, step S1 specifically includes:
[0035] S11. Install the sensor for collecting the bearing vibration signal on the bearing to ensure that the sensor is correctly connected and can work normally;
[0036] S12. Collect the original fault data Z at different positions of the bearing, including the bearing vibration signal data of the inner ring, outer ring, and rolling elements under the condition of zero horsepower of the motor power; in this embodiment, the original fault data is denoted as
[0037] Z = [z1, z2... z N , where N is the data length; for any one item z n , it represents the nth bearing diagnosis signal data, and since it is time series data, it can also be further expressed as z n (t) to characterize its time correlation.
[0038] S2. Perform wavelet transform, add Gaussian noise of different intensities, and block sampling on the obtained original fault data in sequence; and generate class labels and position labels for the sampled fault data;
[0039] As a preferred embodiment, step S2 specifically includes:
[0040] S21. Perform wavelet transform on the original fault data, decompose it into frequency components of different scales to extract local features; after decomposing it into frequency components of different scales, obtain the approximation coefficient C A and the detail coefficient C D . Connect C A and C D into a vector X transformed = [C A , C D as the fault data after wavelet transform;
[0041] In this embodiment, db4 wavelet transform is selected; the approximation coefficient and the detail coefficient represent the low-frequency component and the high-frequency component of the input data respectively; the data after wavelet transform is decomposed into different frequencies on the one hand and can be denoised to a certain extent on the other hand; while extracting the local features of the signal, it is beneficial to adjust the weights of the subsequent learnable convolution;
[0042] S22. Add Gaussian noise of different intensities to the data after wavelet transform, and then perform block sampling. In this embodiment, the window size is set to 1024, so each sampling block contains 1024 samples, and the fault data after wavelet transform is divided into multiple sample data blocks of a fixed size, denoted as X sampled ;
[0043] S23. Generate a class label for each data in the sample data block through one-hot encoding, and then generate the corresponding position label;
[0044] In this embodiment, the class label is a one-hot encoded label for bearing fault classification or classification results. This label indicates which class the sample belongs to; the position label is used to record the class number of each sample in the dataset. Different from the class label, the position label directly represents which class the signal belongs to and is stored in integer form. These two labels are actually different representations of the same information.
[0045] The finally obtained class label matrix Y and position label matrix LabelPositional are used for the training and evaluation of the subsequent model.
[0046] S3. Construct a fault diagnosis model based on a convolutional neural network CNN as the basic model using a hybrid attention mechanism, a residual structure, and a learnable convolution;
[0047] As a preferred implementation manner, as Figure 2 shown, step S3 specifically includes:
[0048] S31. Replace the ordinary convolution in the original CNN model with a learnable convolution having the same structure and calculation logic to adaptively adjust the weights. The data after wavelet transform and block sampling first undergoes feature extraction through the learnable convolution, and then dimensionality reduction and compression are performed through a max pooling layer to obtain the primary feature X.
[0049] It should be noted here that one of the innovations of this application lies in replacing the ordinary convolution in the traditional CNN model with a learnable convolution; the convolution kernel weights of the learnable convolution layer are automatically adjusted through the backpropagation algorithm during the training process. Specifically, when the noise frequency changes, the spectral characteristics of the input signal will also change accordingly. The learnable convolution layer can adapt to this change by adjusting the frequency response characteristics of the convolution kernel, that is: the adjustment of the convolution kernel weights will change the sensitivity of the convolution operation in different frequency bands; if the noise frequency is mainly concentrated in a certain specific frequency band, the convolution kernel weights will be adjusted to have a stronger inhibitory effect in this frequency band, and at the same time better extract the useful features of the signal in other frequency bands. In this way, even if the noise frequency changes, the model can maintain the effective extraction of signal features by adjusting the convolution kernel weights, thereby improving the recognition performance in a variable noise environment, especially being able to extract effective features from the periodic pulses emitted from bearing vibration signal data under multi-noise conditions.
[0050] In the fault diagnosis model designed by the method proposed in the present invention, the learnable convolutions used everywhere are the same, with the same calculation logic and the same structure. However, each time it is called, its behavior will change due to parameters. Although the implementation of the learnable convolution is the same, its use at different positions may exhibit different behaviors due to different parameters. This flexibility enables it to adapt to different convolution operation requirements, such as at the input layer of the model, in the residual block, or when adjusting the size of the feature map. By customizing the convolution kernel and bias terms, the learnable convolution provides a more flexible convolution operation ability than the built-in convolution.
[0051] S32. Input the obtained primary feature X into the residual block parallel module; the residual block parallel module contains three residual blocks with different structures, and each residual block independently processes the input, respectively used to capture multi-scale features, enhance important features, and capture long-distance dependence relationships. The output features of the three residual blocks are denoted as X1, X2, and X3 respectively; the three output features are concatenated together along the channel dimension to form the enhanced feature X. enh ;
[0052] In this embodiment, the three residual blocks have different focuses on the input data. Figures 3-5 Shows the internal structure of each residual block:
[0053] First is the first residual block, which is used to capture multi-scale features, such asFigure 3 As shown in the figure, the specific process is as follows:
[0054] The first residual block is divided into three branches. The first branch performs a one-dimensional convolution operation to extract local important features; the second branch performs a k-dimensional convolution operation to extract deeper features; the third branch first performs a k-dimensional convolution operation to extract the maximum scale features, and then performs data smoothing through a one-dimensional convolution. The features extracted by the three branches are concatenated in the channel dimension, and after adjusting the dimension through a one-dimensional convolution, they are residual with the input primary feature X. Finally, through an activation function, the output feature X1 of the first residual block is obtained.
[0055] In this embodiment, the first residual block combines the Inception module and the residual connection. By performing parallel convolution operations to extract multi-scale features, compared with directly stacking multiple convolutional layers, the Inception structure can reduce the number of parameters while maintaining a high feature extraction ability. By connecting the outputs of multiple branches, the richness of feature representation is increased. In addition, using the residual connection structure enhances gradient propagation and model stability, which can not only alleviate the gradient vanishing problem, but also ensure the stability and convergence of the deep network, and can better retain the original feature information to avoid its loss during the model extraction process. The residual connection structure is adopted in many places in this application for the same consideration as here, and will not be elaborated elsewhere.
[0056] The second residual block is used to enhance key features, such as Figure 4 As shown in the figure, the specific process is as follows:
[0057] Within the second residual block, the primary feature X first passes through a learnable convolution, and the formula is expressed as:
[0058] X’ = LearnableConv1D(X);
[0059] where LearnableConv1D() is the learnable convolution;
[0060] Then it is input into a convolutional block attention module CBAM to weight the features. First, the channel attention mechanism is used to enhance the attention to important features to obtain the channel-weighted feature X c-att ;
[0061] The channel attention mechanism is used to assign a weight to each channel, indicating the contribution degree of the channel to the final output. The specific steps are as follows:
[0062] First, global average pooling and global max pooling operations are performed to obtain the global statistical features of each channel. The input feature is X'∈R N×L×C , N is the number of samples, L is the length of the input signal, and C is the number of channels of the feature map. First, perform global average pooling:
[0063]
[0064] where \(x'\) i is the \(i\)-th input signal of \(X'\), and \(A\) avg ∈ ℝ N×C is the average pooling result for each channel.
[0065] Then perform global max pooling:
[0066]
[0067] where \(A\) max ∈ ℝ N×C is the max pooling result for each channel.
[0068] Next, calculate the weights for each channel by passing the pooled results through a series of fully connected layers; process the results of average pooling and max pooling separately:
[0069] \(A'\) avg = Dense1(\(A\) avg );
[0070] \(A'\) max = Dense1(\(A\) max );
[0071] Then merge these two results through the second fully connected layer to obtain the final channel attention weight \(A\) c
[0072] \(A\) c = σ(Dense2(\(A'\) avg +\(A'\) max ));
[0073] where σ is the Sigmoid activation function, Dense1 and Dense2 are the first and second fully connected layers, and the output \(A\) c ∈ ℝ N×C represents the attention weight for each channel;
[0074] Finally, multiply the channel attention weight \(A\) C by the feature \(X'\) to obtain the channel-weighted feature \(X\) c-att = \(A\) C \(X'\);
[0075] Then perform weighted processing in the spatial dimension through the spatial attention mechanism to obtain the final weighted feature \(X\) s-att ;
[0076] The spatial attention mechanism enhances the features in the spatial domain by assigning a weight to each time step (each position in the feature map); the spatial attention mechanism takes the channel-weighted feature as \(X\)s-att For the input, the spatial attention mechanism performs pooling operations on the features of each channel to obtain the global average pooling and maximum pooling features at each position:
[0077] Average pooling:
[0078]
[0079] where S avg ∈R N×L is the average pooling result for each channel;
[0080] Maximum pooling:
[0081]
[0082] where S max ∈R N×L is the maximum pooling result for each channel;
[0083] Concatenate the pooled results and calculate the spatial attention weights through a convolutional layer:
[0084] S space = Conv1D(S avg , S max );
[0085] where S space ∈R N×L×1 represents the spatial attention weights;
[0086] Finally, multiply the channel-weighted feature X c-att by the spatial attention weights S space to obtain the final weighted feature X s-att = S space X c-att ;
[0087] The final weighted feature X s-att then passes through an SE module to adaptively adjust the weights of each channel, enhance the feature expression of important channels, and obtain X SE ; The attention mechanism SE module compresses each channel through global average pooling, then performs excitation through a fully connected layer, and then applies it to the feature map, so as to be able to adaptively adjust the weights of each channel and strengthen the expression of important channels; The formula is expressed as:
[0088] X SE = SEBlock(X s-att );
[0089] The feature X SEAfter the second learnable convolution, it passes through a Dropout layer and a one-dimensional convolution in sequence, and then makes a residual connection with the output primary feature X to obtain the output feature X2 of the second residual block;
[0090] In this embodiment, the second residual block combines learnable convolution, the CBAM module, and the SE module. The CBAM module mechanism is used to enhance the model's ability to capture important features in the bearing vibration signal, and at the same time, the SE module is used to weight the channels to improve the feature expression ability; the Dropout layer randomly discards the output to suppress the risk of overfitting, and the 1x1 one-dimensional convolution ensures that the output and input channel numbers are the same during residual connection; the combined use of the Dropout layer and the one-dimensional convolution elsewhere in the text also has the same effect, so it will not be elaborated too much in this embodiment;
[0091] The third residual block is used to capture long-range dependencies, such as Figure 5 shown, and it specifically includes:
[0092] In the third residual block, the primary feature X first passes through a learnable convolution layer to extract features, obtaining the feature X l , and then the feature X l enters the Transformer module, calculates the correlation between features through the multi-head self-attention mechanism, and then further extracts the feature X through the feed-forward network l’ ; X l’ is sequentially smoothed through the second learnable convolution layer, the Dropout layer, and the one-dimensional convolution, and finally makes a residual connection with the primary feature X to obtain the output feature X3 of the third residual block. Since the fault features may be far apart in the time series, this application designs the third residual block for the time series data of the bearing vibration signal, which can effectively capture the long-range dependencies in the signal; through the Transformer module, the model can better understand the correlation between different time parts of the signal, thereby improving the accuracy of fault diagnosis.
[0093] In addition, in this application, two learnable convolution layers are used in both the second and third residual blocks, which can smooth the input features and extract edge information. High-frequency noise will be suppressed by the convolution operation, and it has a smoothing effect when extracting features. Low-frequency features are extracted by larger convolution kernels, playing a role in denoising and feature enhancement; however, there are the following considerations for using two learnable convolutions in this application:
[0094] Since the bearing vibration signal usually contains a large amount of noise and weak fault features, before feature extraction, a learnable convolution layer is first used to extract features and reduce the dimension of the signal, reducing redundant information and retaining the main features; while the second learnable convolution layer performs multi-scale fusion on the extracted features to capture complex fault mode features.
[0095] Up to step S32, in this application, the designed residual parallel module extracts and enhances key fault features from three different perspectives, which helps to perform accurate fault diagnosis under multi-noise conditions.
[0096] S33. The enhanced feature X obtained by the residual block parallel module enh Then, it is sequentially processed through an average pooling layer (used to reduce the model complexity and fix the output feature length), a fully connected layer (introducing non-linearity to enhance the model's expressive ability), and a Dropout layer (suppressing overfitting) to obtain the feature X'. enh ;
[0097] S34. The feature X' obtained after the data processing in step S33 enh Then, it passes through a residual connection structure composed of a BiLSTM module and a self-attention mechanism module to capture the global feature X all ; The BiLSTM module is used to further capture the context information in the time series, and the self-attention mechanism is used to strengthen the learning of important features;
[0098] As a preferred embodiment, in this embodiment, as Figure 2 shown, step S34 specifically includes:
[0099] S341. The residual connection structure includes a BiLSTM module, a Dropout layer, and a self-attention mechanism module connected in sequence; first, the BiLSTM module captures the long-term dependence in two directions for the input X' enh to obtain a feature sequence H containing the forward and backward LSTM outputs
[0100] The forward LSTM processes the input sequence from left to right. The weights and gating structures of the backward LSTM are the same as those of the forward LSTM, but in the opposite direction; the feature at each time step can be expressed as:
[0101]
[0102] where are the features output by the forward LSTM and the backward LSTM at the current time step t, respectively;
[0103] S342. H is input into the self-attention mechanism module after passing through a Dropout layer, and the information at other positions in the reference feature sequence is referred to to capture the global dependence, obtaining the output H of the self-attention mechanism module att ; Specifically:
[0104] For the input sequence representation H = [h1, h2... h T , first calculate the query, key, and value:
[0105] Q = HW Q 、K = HW K 、V = HW V ;
[0106] Wherein, W Q W K W V is the weight matrix obtained through training.
[0107] Calculate the attention weights:
[0108] Wherein, is the dot product of the query and the key, d k is the dimension of the key, used for scaling, and softmax normalizes the weights.
[0109] Finally, the representation after weighting by the self-attention mechanism:
[0110] H att = Attention(Q, K, V);
[0111] S343. Then, the input X' enh is subjected to a residual connection with the output H att of the self-attention mechanism module to obtain the global feature X all = H att + X' enh ; Thus, enhanced key features and global information dependencies are obtained.
[0112] S35. The global feature X all after being processed by the residual connection structure enters the last two fully connected layers. The first fully connected layer maps the feature X all to a higher-dimensional space, increasing the non-linear expression ability of the model and helping the model learn more complex patterns from rich features. The second fully connected layer then maps the feature to the class space to obtain the final output feature X output , that is, the classification result is obtained.
[0113] S4. Divide the sampled fault data into a training set and a test set, and use them as the input data of the model for training, and monitor the training results in real time for parameter optimization until the optimal model is trained; and visualize the fault diagnosis results output by the model;
[0114] As a preferred implementation method, step S4 specifically includes:
[0115] S41. During the model training process, continuously observe the losses and accuracies of the training set and the test set. If there are large differences and fluctuations, terminate the model training in a timely manner, modify the parameters, and retrain; prevent the model from being in an overfitting state;
[0116] S42. Perform a visualization operation on the model results and output the confusion matrix.
[0117] In this embodiment, the confusion matrix is used to evaluate the performance of the fault diagnosis model, which shows the relationship between the model prediction results (obtained in step S3) and the actual labels (set in step S2). Each row of the confusion matrix represents the actual category, and each column represents the predicted category. By analyzing the confusion matrix, it is possible to understand in which categories the model performs well and in which categories there are problems.
[0118] S5. Conduct a comparative experiment with the existing model under various noise conditions to test the performance of the constructed fault diagnosis model.
[0119] So far, this embodiment shows the specific process of the method proposed by the present invention. Through the method designed by the present invention, a fault diagnosis model for bearing faults is designed. Therefore, to verify the performance of the fault diagnosis model constructed by the method proposed by the present invention, the following experimental examples are provided in this embodiment:
[0120] This method is experimented on the publicly available bearing fault dataset provided by Case Western Reserve University. The Case Western Reserve University dataset includes a total of 9 ".mat" data files for the inner ring, outer ring, and rolling elements, which are fault data, and 1 ".mat" file for normal data. Each file contains the vibration acceleration signal of the bearing collected by an acceleration sensor placed above the bearing housing at the motor drive end under the condition of a sampling frequency of 12 kHz. All faulty bearings are damaged by electrical discharge machining at a single point. The bearing types collected in the publicly available bearing fault dataset of Case Western Reserve University are 3 faulty bearings and 1 normal bearing, and the fault positions of the faulty bearings are the outer ring of the bearing, the rolling element of the bearing, and the inner ring of the bearing. All data files are sampled, with the signal length of each sampling being 1024. The sampled dataset is divided into a training set and a test set, with the training set accounting for 75% and the test set accounting for 25%. In the experiment, the publicly available bearing fault dataset of Case Western Reserve University is used as the training set and the test set by the present invention.
[0121] Figure 6 And Table 1 below shows that this model has a higher fault recognition rate compared to other basic models, especially under variable noise conditions, where the advantage is more obvious.
[0122] Table 1 Comparison results of each model under variable noise conditions
[0123]
[0124] The present invention sets the number of training times to 140, the batch size to 32, and uses the Adam algorithm with an initial learning rate of 0.00004. AsFigures 7-11 As shown, under different intensities of noise, the training and test accuracy of the model of the present invention can be higher than 95%, the training loss of the model of the present invention can be lower than 0.1, and the test loss can be maintained at about 0.1. Even with the continuous increase of the noise intensity, the fault diagnosis model constructed by the present invention has not lost its effectiveness. Therefore, the bearing fault method proposed by the present invention is reliable.
[0125] In summary, the method proposed by the present invention can accurately diagnose bearing faults, enabling the model to have the advantage of accurately identifying bearing faults under variable noise conditions, improving the versatility and accuracy of the bearing fault diagnosis model, and realizing the identification of bearing faults under variable noise conditions.
[0126] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0127] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices.
[0128] The above embodiments have introduced the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism, characterized in that, Specifically, it includes the following steps: S1. Use sensors to obtain bearing vibration signals and obtain various types of original fault data; S2. Perform wavelet transform, add Gaussian noise with different intensities, and block sampling on the obtained original fault data in sequence; and generate class labels and position labels for the sampled fault data; S3. Build a fault diagnosis model based on a convolutional neural network CNN as the basic model using a hybrid attention mechanism, residual structure, and learnable convolution; S4. Divide the sampled fault data into a training set and a test set, and use them as input data for the model for training, and monitor the training results in real time for parameter optimization until the optimal model is trained; and visualize the fault diagnosis results output by the model; S5. Conduct comparative experiments with existing models under various noise conditions to test the performance of the constructed fault diagnosis model.
2. The bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 1, wherein Step S1 specifically includes: S11. Install the sensor for collecting bearing vibration signals on the bearing to ensure that the sensor is correctly connected and can work normally; S12. Collect the original fault data Y at different positions of the bearing, including the bearing vibration signal data of the inner ring, outer ring, and rolling elements under the condition of zero horsepower of the motor.
3. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 1, characterized in that, Step S2 specifically includes: S21. Perform wavelet transform on the original fault data, decompose it into frequency components of different scales to extract local features; after decomposing it into frequency components of different scales, obtain the approximation coefficient C A and the detail coefficient C D . Connect C A and C D into a vector X transformed = [C A , C D , which serves as the fault data after wavelet transform; S22. After wavelet transform, Gaussian noise with different intensities is added to the data, and then block sampling is performed. By means of a sliding window, the fault data after wavelet transform is divided into multiple sample data blocks of a fixed size, denoted as X sampled ; S23. Generate class labels for each data in the sample data block through one-hot encoding, and then generate corresponding position labels; the class labels and position labels are used for subsequent model training and evaluation.
4. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 1, characterized in that Step S3 specifically includes: S31. Replace the ordinary convolution in the original CNN model with learnable convolution with the same structure and calculation logic to adaptively adjust the weights. The data after wavelet transform and block sampling first undergoes learnable convolution for feature extraction, and then passes through a max pooling layer for dimensionality reduction and compression to obtain the primary feature X; S32. Input the obtained primary feature X into the residual block parallel module. The residual block parallel module contains three residual blocks with different structures. Each residual block independently processes the input, which are respectively used to capture multi-scale features, enhance important features, and capture long-range dependencies. The output features of the three residual blocks are denoted as X1, X2, and X3 respectively. Concatenate the three output features along the channel dimension to form the enhanced feature X enh ; S33. The enhanced feature X obtained through the residual block parallel module enh Then, it is successively processed by an average pooling layer, a fully connected layer, and a Dropout layer to further process the data, obtaining the feature X' enh ; Feature X' obtained after data processing in step S33 enh Then, a residual connection structure composed of a BiLSTM module and a self-attention mechanism module is used to capture the global feature X all ; S35. The global feature X processed by the residual connection structure all enters the last two fully connected layers. The first fully connected layer maps the feature X all to a higher-dimensional space, increasing the non-linear expression ability of the model and helping the model learn more complex patterns from rich features. The second fully connected layer then maps the feature to the class space to obtain the final output feature X output , that is, the classification result is obtained.
5. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 4, characterized in that In step S32, the first residual block is used to capture multi-scale features. The specific process is as follows: The first residual block is divided into three branches. The first branch performs a one-dimensional convolution operation to extract local important features; the second branch performs a k-dimensional convolution operation to extract deeper features; the third branch first performs a k-dimensional convolution operation to extract the maximum scale features, and then undergoes a one-dimensional convolution for data smoothing; the features extracted by the three branches are concatenated in the channel dimension, and after adjusting the dimensions through a one-dimensional convolution, they are residual with the input primary feature X, and finally pass through an activation function to obtain the output feature X1 of the first residual block.
6. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 4, characterized in that, In step S32, the second residual block is used to enhance key features. The specific process is as follows: Within the second residual block, the primary feature X first passes through a layer of learnable convolution and then is input into a convolutional attention mechanism module CBAM to weight the feature. First, the channel attention mechanism is used to enhance the attention to important features, obtaining the channel-weighted feature X c-att , and then through the spatial attention mechanism for weighting in the spatial dimension, obtaining the final weighted feature X s-att ; The final weighted feature X s-att then passes through an SE module to adaptively adjust the weights of each channel, enhancing the feature expression of important channels, obtaining X SE ; Feature X output by the SE module SE After passing through the second learnable convolution, it sequentially passes through a Dropout layer and a one-dimensional convolution, and then performs a residual connection with the output primary feature X to obtain the output feature X2 of the second residual block.
7. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 4, characterized in that, In step S32, the third residual block is used to capture long-distance dependencies. The specific process is as follows: In the third residual block, the primary feature X first passes through a learnable convolutional layer to extract features, obtaining the feature X l , and then the feature X l enters the Transformer module, calculates the correlations between features through the multi-head self-attention mechanism, and then further extracts the feature X through the feed-forward network l, ; X l, successively passes through the second learnable convolutional layer, the Dropout layer, and the one-dimensional convolution for smoothing, and finally is connected residually with the primary feature X to obtain the output feature X3 of the third residual block.
8. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 4, characterized in that, Step S34 specifically includes: S341. The residual connection structure includes a BiLSTM module, a Dropout layer, and a self-attention mechanism module connected in sequence. First, in the BiLSTM module, X' input from two directions enh captures long-term dependencies to obtain a feature sequence H containing the forward and backward LSTM outputs; S342 and H are input into the self-attention mechanism module after passing through a Dropout layer, referring to the information at other positions of the reference feature sequence to capture global dependencies, and obtaining the output H of the self-attention mechanism module att ; S343. Then, input X' enh is subjected to residual connection with the output H of the self-attention mechanism module att to obtain the global feature X all .
9. A bearing fault diagnosis method based on wavelet transform and hybrid attention mechanism according to claim 1, characterized in that, Step S4 specifically includes: S41. During the model training process, continuously observe the losses and accuracies of the training set and the test set. If there are large differences and fluctuations, terminate the model training in time, modify the parameters, and retrain; prevent the model from overfitting; S42. Perform a visualization operation on the model results and output a confusion matrix.
Citation Information
Patent Citations
Fault diagnosis method based on fusion of residual learning and attention mechanism
CN115640531A
Rolling bearing fault diagnosis method based on WPD and AFRB-LWUNet under small sample
CN116399588A
Bearing fault diagnosis method based on wavelet transform and depth residual attention mechanism
CN116718377A
Bearing fault diagnosis method based on dual-channel multi-scale attention feature fusion
CN118918424A
Rolling bearing fault diagnosis method based on autoregression data generation
CN118936887A
Cited By
Marine bearing dark damage diagnosis method based on convolution of mixed attention and Inception fusion graph
CN122346767A