Bearing fault classification method based on bidirectional time convolution and Transform

By combining VMD and FFT with the method based on bidirectional time convolution and Transformer, feature extraction is performed, and feature fusion is performed using the multi-head self-attention mechanism and cross-attention mechanism, the problem of insufficient fusion of time-frequency domain features in the existing technology is solved, and a high-precision and high-rootability bearing fault classification is achieved.

CN120356483APending Publication Date: 2025-07-22NORTHEASTERN UNIV CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510437583.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate time-frequency domain features, and is not robust when facing audio bearing failure data sets in mixed working conditions, resulting in low accuracy and efficiency of bearing failure classification.

Method used

Using a method based on bidirectional time convolution and Transformer, bearing fault audio is collected through microphone arrays, feature extraction is used using VMD and FFT to form a feature matrix containing time domain and frequency domain information, and fault classification is performed by combining bidirectional time convolution and Transformer parallel structure, and feature fusion is performed by using multi-head self-attention mechanism and cross-attention mechanism.

Benefits of technology

It improves the accuracy and robustness of bearing fault classification, and can accurately identify fault types under ideal experimental and mixed conditions, adapt to dynamic changes in the industrial environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356483A_ABST
    Figure CN120356483A_ABST
Patent Text Reader

Abstract

The invention discloses a bearing fault classification method based on bidirectional time convolution and Transform, and the method comprises the steps: collecting a plurality of bearing fault audios through a microphone array, and carrying out the preprocessing of the class labels; performing decomposition of a VMD method on the audio data to obtain a plurality of IMFs with different center frequencies and bandwidths; fFT processing is carried out on the audio data, and frequency spectrum information, namely frequency distribution, of the signals is obtained; performing feature fusion on the plurality of IMFs and the frequency spectrum information in a stacking manner to form a feature matrix containing time domain and frequency domain information; inputting the feature matrix into a constructed fault classification model containing a bidirectional time convolution and Transform parallel structure, and obtaining the probability distribution of each type of fault; training model parameters by minimizing a loss function and optimizing an evaluation index until a set threshold value is met, and storing the model parameters; meanwhile, the algorithm precision is measured through the optimal evaluation index, and the system performance is evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent fault diagnosis, and relates to a bearing fault classification method based on bidirectional temporal convolution and Transformer. Background Art

[0002] In the wave of industrial intelligent transformation, intelligent manufacturing has put forward higher requirements for the reliability and efficiency of mechanical equipment. With the rapid development of artificial intelligence technology, intelligent diagnosis has gradually penetrated into the maintenance of industrial equipment, significantly improving the accuracy and efficiency of fault diagnosis by using cutting-edge information and technology. Bearing fault classification, as a key link among them, is particularly important. In the past, the identification of bearing faults mainly relied on the experience of professional technicians. However, in the face of complex and diverse audio or vibration data, traditional methods often fell short. Bearing fault classification has started to develop towards intelligence, experiencing an evolution from traditional signal processing methods (such as Fourier transform, short-time Fourier transform, wavelet transform) to machine learning and then to deep learning. Traditional methods have limitations in dealing with non-stationary and non-linear signals, and the parameter selection is complex, making it difficult to adapt to the dynamic changes in the industrial environment. Machine learning methods combine artificial feature extraction and classification models, improving the level of intelligence, but still relying on artificial features and difficult to meet the requirements of real-time and high accuracy. Fortunately, the rise of deep learning technology provides an effective solution to this problem. Deep learning can autonomously learn the fault patterns and corresponding features in the signal, not only significantly improving the accuracy and stability of fault classification, but also becoming increasingly efficient as the bearing fault data increases and industrial requirements continue to upgrade.

[0003] The emergence of deep learning technology has provided new solutions for bearing fault diagnosis. Models such as CNN and Transformer have been widely used in bearing fault diagnosis due to their excellent performance. As one of the pioneers in the field of deep learning, CNN can efficiently extract local and global features in one-dimensional fault signals through convolutional layers, pooling layers, and fully connected layers, such as time-domain, frequency-domain features, and spectral structure information. These features are crucial for accurately classifying bearing faults. The convolutional layer extracts local features, the pooling layer reduces computational complexity, and the fully connected layer completes the classification task. Through this hierarchical feature learning, CNN significantly improves the accuracy and robustness of bearing fault classification. Subsequently, AlexNet made significant progress in bearing fault diagnosis by deepening the network structure and accelerating training with GPUs. GoogLeNet introduced the Inception module, which processes information through multi-scale convolutional kernels, improving the computational efficiency and accuracy of the network. ResNet solved the problem of gradient disappearance in deep networks by introducing residual connections, making it possible to train deeper networks and thus improving the accuracy and robustness of fault diagnosis. EfficientNet optimized the network structure through neural architecture search and compound scaling strategies, significantly reducing the computational cost while ensuring high accuracy. TCN effectively processes time series data and captures long-range dependency information by introducing one-dimensional convolution instead of RNN and has achieved good results in sequence modeling tasks. Different from CNN, Transformer initially shone in the field of natural language processing and was later successfully introduced into the field of fault diagnosis. The core part of Transformer is its self-attention mechanism, which can capture long-range dependencies between different positions in the input sequence, avoiding information loss when CNN processes long time series data. By calculating the weighted relationships between positions, the self-attention mechanism enables the model to focus on important features, thus improving its performance in complex tasks. QCNN improves the model's ability to extract global features by introducing an attention mechanism into WDCNN. When the relatively classic ViT model derived from Transformer is used for fault diagnosis, one-dimensional data is first transformed into image data through methods such as CWT or STFT, and then the image data is split into multiple patches, each patch representing a part of the image. ViT uses the self-attention mechanism to model these patches, capture local and global features, and finally fuse the features through the Transformer layer to achieve fault classification. In this way, ViT can effectively improve the accuracy and robustness of fault diagnosis.

[0004] However, the above method cannot well fuse time-frequency domain features, nor can it take into account the extraction of global and local features. It lacks robustness when dealing with the audio bearing fault dataset under mixed working conditions. The main purpose of the present invention is to design and verify a method that can accurately and efficiently classify ideal experimental data and the audio bearing fault dataset under mixed working conditions. Summary of the Invention

[0005] To solve the above technical problems, the object of the present invention is to provide a bearing fault classification method based on bidirectional temporal convolution and Transformer.

[0006] A bearing fault classification method based on bidirectional temporal convolution and Transformer of the present invention includes:

[0007] Step 1: Use a microphone array to collect various bearing fault audio simulated by a simulation bearing fault test bench, and preprocess its category labels;

[0008] Step 2: Decompose the audio data based on the VMD method to obtain four IMFs with different center frequencies and bandwidths;

[0009] Step 3: Perform FFT processing on the audio data to obtain the spectral information of the signal, that is, the frequency distribution characteristics;

[0010] Step 4: Fuse the multiple IMFs obtained by the transformation and the spectral information in a stacked manner to form a feature matrix containing time domain and frequency domain information;

[0011] Step 5: Input the feature matrix into the constructed fault classification model with a parallel structure of bidirectional temporal convolution and Transformer to obtain the probability distribution of various types of faults;

[0012] Step 6: By minimizing the loss function and optimizing the evaluation index, train the model parameters until the set threshold is met, and save the model parameters; at the same time, measure the algorithm accuracy with the optimal evaluation index and evaluate the system performance.

[0013] A bearing fault classification method based on bidirectional temporal convolution and Transformer of the present invention has the following beneficial effects:

[0014] (1) A bearing fault classification method based on bidirectional temporal convolution and Transformer is proposed, which has high accuracy and good robustness.

[0015] (2) Through VMD and FFT, multi-scale and multi-band feature extraction is performed on the one-dimensional vibration signal, and the data is more deeply analyzed and represented in the time domain and frequency domain, enabling the model to learn richer features from these two dimensions (time domain and frequency domain).

[0016] (3) Based on temporal convolution and combined with the design of a bidirectional structure, it can capture the dynamic features of the signal in both the forward and backward directions simultaneously, improving the accuracy and stability of the model in temporal feature extraction.

[0017] (4) Introducing Transformer, it captures long-range dependence relationships through the multi-head self-attention mechanism, significantly enhancing the model's ability to model global features, thereby improving the accuracy of fault classification.

[0018] (5) Dynamically weighting and fusing local and global features through the cross-attention mechanism, adaptively learning key features, and improving the complementarity of feature representation and diagnostic performance. Brief Description of the Drawings

[0019] Figure 1 is a flowchart of a bearing fault classification method based on bidirectional temporal convolution and Transformer of the present invention;

[0020] Figure 2 is a multi-modal time-domain feature map obtained after VMD;

[0021] Figure 3 is a spectrogram obtained after FFT of audio data; (a) Spectrogram of normal audio data; (b) Spectrogram of pitting fault audio data; (c) Spectrogram of broken tooth fault audio data; (d) Spectrogram of wear fault audio data;

[0022] Figure 4 is a visualization diagram of the feature matrix obtained after feature fusion;

[0023] Figure 5 is a structural diagram of a fault classification model with a parallel structure of bidirectional temporal convolution and Transformer;

[0024] Figure 6 is a confusion matrix diagram on the test sets of different datasets, (a) Confusion matrix diagram of dataset 1; (b) Confusion matrix diagram of dataset 2;

[0025] Figure 7 is an example waveform diagram of data under various working conditions of the rolling element;

[0026] Figure 8 is a confusion matrix diagram under three mixed working condition tasks, (a) Confusion matrix diagram of task 1; (b) Confusion matrix diagram of task 2; (3) Confusion matrix diagram of task 3. Detailed Embodiments

[0027] As Figure 1 shown, a bearing fault classification method based on bidirectional temporal convolution and Transformer of the present invention includes:

[0028] Step 1: Use a microphone array to collect various bearing fault audio simulated by a simulated bearing fault test bench and preprocess its class labels. Specifically:

[0029] Step 1.1: Use a microphone array to collect data from the simulated bearing fault test bench and perform analog-to-digital conversion on the collected data. The device model is not limited, the sampling frequency is set to 25.6 kHz, and the bearing speed is set to 800 r / min.

[0030] Step 1.2: Divide the audio data under each working condition according to different working condition series, and organize the multiple fault audio data under the same working condition into a csv file. Assign class labels according to the categories to be classified (such as: inner ring wear, ball missing). Each column in the csv file corresponds to different fault category data, including the fault category label and the collected signal. The label field clearly indicates the fault type corresponding to each piece of data for subsequent model training.

[0031] Step 1.3: Normalize the data so that the values of all features are within the range of (-1, 1).

[0032] Step 1.4: Divide the time series data into multiple samples according to a set length, such as 2048 or 4096 sample points.

[0033] Step 2: Decompose the audio data using the VMD (Variational Mode Decomposition) method to obtain multiple IMFs (Intrinsic Mode Functions) with different center frequencies and bandwidths, making it easier for the model to identify fault-related frequency components. Specifically:

[0034] Step 2.1: Perform VMD decomposition on the fault audio, which is an algorithm that solves the variational problem through an optimization process. Decompose the preprocessed audio signal f(t) into K IMFs, and minimize the bandwidth of each IMF. Each IMF is represented as u k (t), that is, the k-th IMF of the signal at time t, and its center frequency is represented as ω k , δ(t) is the Dirac function used for signal convolution operation. The calculation formula for the variational problem is:

[0035]

[0036] At the same time, there are constraints:

[0037]

[0038] Step 2.2: To simplify the optimization process of the variational problem in 2.1, introduce the quadratic penalty factor α and the Lagrange multiplier operator λ(t) to transform the constrained variational problem into an unconstrained variational problem:

[0039]

[0040] Step 2.3: Utilize the IMF components and their frequency information optimized through Step 2.2 to further update each IMF, and determine the final IMF components (u1, u2... u k ) and the corresponding central frequencies (ω1, ω2... ω k ) through iterative updating. Iteratively update the IMF component u k (t):

[0041]

[0042] Wherein, is the Fourier transform of the k-th modal component after the (n + 1)-th iteration, is the Fourier transform of the preprocessed audio signal f(t).

[0043] Iteratively update the central frequency ω k :

[0044]

[0045] Wherein, is the value of the central frequency of the k-th mode after the (n + 1)-th iteration, is the Fourier transform of the k-th modal component.

[0046] Step 2.4: When the iterative process converges, output the u k (t) and central frequency ω k of multiple IMFs, as shown in Figure 2 . The final multiple modal features are in the format of B is the batch size, L is the sequence length, and C1 is equal to K.

[0047] Step 3: Perform FFT (Fast Fourier Transform) processing on the audio data to obtain the spectral information of the signal, i.e., the frequency distribution characteristics, specifically:

[0048] Step 3.1: Convert the preprocessed audio signal f(t) from analog to digital as f[n], and then perform FFT processing to obtain its frequency distribution. The mathematical expression of FFT is based on DFT and is defined as follows:

[0049]

[0050] Wherein, X[h] is the h-th sample in the frequency domain, f[m] is the m-th sample in the time domain, and M is the total number of samples;

[0051] Step 3.2: Take the modulus of a series of complex numbers obtained after FFT processing to obtain the frequency distribution characteristics of the signal, i.e., spectral information:

[0052]

[0053] Among them,

[0054] Figure 3 is the spectrogram obtained after the audio data undergoes FFT; (a) spectrogram of normal audio data; (b) spectrogram of audio data with pitting fault; (c) spectrogram of audio data with broken tooth fault; (d) spectrogram of audio data with wear fault.

[0055] Step 4: Stack and fuse the features of multiple IMFs (intrinsic mode functions) obtained by transformation and the spectral information in a stacked manner to form a feature matrix containing time-domain and frequency-domain information, specifically:

[0056] Stack the spectral information obtained in Step 3 and the IMFs obtained in Step 2 according to the feature dimension to form a feature matrix f ∈ R B×L×D , where B represents the batch size, L represents the sequence length, and D = C1 + C2 represents the sum of the spectral feature dimension and the feature dimension of the IMFs, enabling the model to utilize the features of both the time domain and the frequency domain simultaneously. The visualization diagram of the feature matrix is as shown in Figure 4 shown.

[0057] Step 5: Input the feature matrix into the constructed fault classification model with a parallel structure of bidirectional temporal convolution and Transformer to obtain the probability distribution of each type of fault. The fault classification model is as shown in Figure 5 shown. The fault classification model with a parallel structure of bidirectional temporal convolution and Transformer is specifically:

[0058] Extract local temporal features through the bidirectional temporal convolution module, and extract global features through the Transformer module; fuse the features of the two parts through the cross attention layer and output them to the classifier, and the classifier calculates the probability distribution of each fault category to complete fault classification.

[0059] The bidirectional temporal convolution module includes a forward temporal convolution part and a backward temporal convolution part, and the basic components all include: dilated causal convolution, Chomp1d layer, Relu activation function, and Dropout layer;

[0060] The operation process of the forward temporal convolution part is as follows:

[0061] Dilated Causal Convolution: In the forward-time convolutional part, dilated causal convolution is adopted to extract features of time series data. The output of the dilated convolution at time step t is determined by the sum of the products of the values of the feature matrix f at different time steps and the elements of the convolutional kernel g. The index s of the convolutional kernel starts from 0 until L g -1 to ensure that the output meets the causality requirement. The output tensor formula is:

[0062]

[0063] where, * d,c denotes the dilated convolution operation, d is the dilation rate, c is the causality, L g is the convolutional kernel length, and g(s) is the value of the convolutional kernel at position s;

[0064] Chomp1d Layer: After the dilated causal convolution, the Chomp1d layer is applied to trim the tail padding that may be generated due to the dilated convolution, ensuring that the length of the output sequence is the same as that of the input sequence. If the output tensor is represented as (f * d,c g)(t) ∈ R B×L×C , where C is the number of channels, then the formula is:

[0065] (f * d,c g)(t)' = (f * d,c g)(t)[:, :L - chomp_size, :] (9)

[0066] where, (f * d,c g)(t)' represents the tensor after being processed by the Chomp1d layer, chomp_size is the number of elements at the end of the sequence to be removed, which is usually equal to the padding added due to the dilated convolution;[:, :L - chomp_size, :] means retaining all elements of the first dimension B of the output tensor, retaining all elements of the third dimension C of the output tensor, and on the second dimension L of the output tensor, retaining the elements from the beginning to the position of L - chomp_size;

[0067] ReLU Activation Function: The ReLU activation function is applied to introduce non-linearity:

[0068] (f * d,c g)(t)” = ReLU[(f * d,c g)(t)'] = max[0, (f * d,c g)(t)'] (10)

[0069] where, (f * d,c g)(t)” represents the tensor after being processed by the ReLU function;

[0070] Dropout layer: Add a Dropout layer after the ReLU activation function to randomly discard some neurons in the network to prevent the model from overfitting to the training data:

[0071] X forward = (f * d,c g)(t)” ε Bernoulli(p) (11)

[0072] where, ⊙ represents element-wise multiplication, and Bernoulli(p) is a Bernoulli random variable, where its parameter p is the probability of retaining neurons;

[0073] The operation process of the backward-time convolution part is as follows:

[0074] Define the backward-time convolution part in the bidirectional time convolution module. Similar to the forward-time convolution, the backward-time convolution first flips the input sequence, then applies the same structure as the forward-time convolution to extract features, and finally flips the output again to restore its original time order to obtain

[0075] After concatenating the output of the forward-time convolution part and the output of the backward-time convolution part in the feature dimension, perform Dropout to reduce overfitting to obtain the final local feature representation

[0076] Define that the transformer module includes the following basic components: input embedding and positional encoding, multi-head attention mechanism MHA, residual connection, LN layer, and FFN layer. The operations of the transformer module are as follows:

[0077] Input embedding and positional encoding: Convert the feature matrix f ∈ R B×L×D containing time-domain and frequency-domain information into an embedding vector z0 ∈ R B×L×D and add positional encoding to retain the order information of time steps. The definition of positional encoding is:

[0078]

[0079] where is the index of the t time step, and i is the index of the feature dimension;

[0080] The calculation result of the input embedding is:

[0081] z0 = f + PE (13)

[0082] MHA layer: In the multi-head attention mechanism, the input z0 ∈ R B×L×D is linearly transformed into three parts, namely query key value dk is the feature dimension of the query and the key, d v is the feature dimension of the value. The calculation formula for a single attention head is:

[0083]

[0084] The multi-head attention mechanism forms the final output by calculating the results of multiple attention heads in parallel, concatenating these results, and then mapping them back to the feature dimension through the output weight matrix;

[0085] LN and residual connection: After each multi-head attention mechanism, residual connection and layer normalization LN are used to improve training stability and model performance. The residual connection is expressed as:

[0086] z' = z0 + MHA(z0) (15)

[0087] where z' ∈ R B×L×D ;

[0088] The application of LN is expressed as:

[0089]

[0090] where μ ∈ R and σ ∈ R are the mean and variance of the feature dimension, ε is used to prevent division by zero, γ ∈ R and β ∈ R are learnable scaling parameters and offset parameters;

[0091] FFN: The output of the multi-head attention mechanism is further passed through a two-layer feed-forward network for feature extraction and non-linear transformation; After FFN, residual connection and layer normalization are performed again to obtain

[0092] The above modules are stacked 4 layers by stacking multiple encoders to obtain the final global feature representation X transformer ∈ R B ×L×D 。

[0093] The two parts of features are fused through the cross attention layer and output to the classifier. The classifier calculates the probability distribution of each fault category to complete fault classification. Specifically:

[0094] The Cross attention module is used to model the correlation of features from two feature extraction modules and achieve information fusion between different features through the cross attention mechanism; The local features of the output of the bidirectional temporal convolution module are used as the query sequence Q, and the global feature X of the output of the transformer module is used transformer ∈ R B×L×DSimultaneously serve as the key sequence K and the value sequence V, and then obtain Q through linear mapping emb ,K emb ,V emb ∈R B×L×D , calculate the attention weights using the projected query sequence and the key sequence and weight the value sequence:

[0095]

[0096] where, X ∈ R B×L×D is the fused feature representation. The local features are used as the query, which can actively search for relevant information from the global features, strengthen the focus of fine-grained spatio-temporal features, and at the same time supplement the global context information. The global features provide the overall temporal dependence and are suitable as the target for the query;

[0097] The fused feature X ∈ R B×L×D First, swap the positions of the feature dimension D and the time dimension L, and then use the pooling operation to compress the time dimension L to 1 to get X ∈ R B×D×1 , flatten the pooled feature tensor into a two-dimensional form flat_tensor ∈ R B×D , input it into the fully connected layer to map the features to the target classification space and finally output outputs ∈ R B×C , where C is the number of classification categories; the output of each sample is the scores of C categories, which are used for subsequent classification tasks.

[0098] Step 6: By minimizing the loss function and optimizing the evaluation metrics, train the model parameters until the set threshold is met, and save the model parameters; at the same time, measure the algorithm accuracy with the optimal evaluation metrics and evaluate the system performance, specifically:

[0099] Step 6.1: Perform fault classification on the collected dataset. The sampling frequency and feature extraction method of the dataset samples are as in Steps 2, 3, and 4. The feature matrix containing time-domain and frequency-domain information obtained after preprocessing is used as the input of the fault classification model;

[0100] Step 6.2: Use the predicted probability distribution and the true class distribution as the loss function, and use accuracy, precision, recall, and F1-score as the optimal evaluation metrics. The expression of the loss function is:

[0101]

[0102] where, y is the probability distribution of the true label, is the probability distribution predicted by the model, and C is the number of classification categories;

[0103] Step 6.2: By minimizing the predicted probability distribution and the true class distribution until the number of training times reaches the set threshold or the value of the loss function reaches the set range, it is considered that the model parameters have been trained, and the model parameters are saved; at the same time, the optimal evaluation index is selected to measure the accuracy of the algorithm and evaluate the performance of the system.

[0104] The present invention will be further described below with reference to examples.

[0105] Example 1:

[0106] In order to verify the performance and robustness of the proposed bearing fault classification method based on bidirectional temporal convolution and Transformer under different operating conditions, the experiment takes self-collected data as an example. The focus is on comparing the performance of the method of the present invention with that of existing mainstream methods in terms of accuracy, robustness, etc. The data statistics table after preprocessing the data set is shown in Table 1.

[0107] Table 1 is a statistical table of the parameters and quantities of different data sets

[0108]

[0109] Since the bearing speed of Data Set 2 is relatively fast, a segment length of 2048 is selected in the preprocessing stage, and each sample contains about 2.4 revolutions. The Adam optimizer and the cross-entropy loss function are used during the training process, the learning rate is set to 0.001, and 32 samples are iterated each time during the training process. The data set is randomly allocated, with 70% used as the training set, 20% used as the validation set, and 10% used as the test set. In order to comprehensively evaluate whether the model is effective, six parameter indicators are taken for comparison:

[0110] Acc: Accuracy, that is, the proportion of all correctly classified samples in the total samples.

[0111]

[0112] Pre: Precision, that is, among the samples predicted as positive by the model, the proportion of those that are truly positive.

[0113]

[0114] Rec: Recall rate, that is, among the samples that are truly positive, the proportion predicted as positive by the model.

[0115]

[0116] F1-Score: F1 score, which comprehensively considers the harmonic mean of precision and recall rate and is used to measure the overall classification performance of the model, especially in the problem of class imbalance.

[0117]

[0118] Among them, TP is True Positive, which refers to the number of samples predicted as positive samples and actually being positive samples; TN is True Negative, which refers to the number of samples predicted as negative samples and actually being negative samples; FP is False Positive, which refers to the number of samples predicted as positive samples but actually being negative samples; FN is False Negative, which refers to the number of samples predicted as negative samples but actually being positive samples. Table 2 shows the comparison results of the weighted average values of the four major parameters of this model for each category when using different data sets. Figure 6 It is the confusion matrix diagram on the test sets of different data sets. (a) The confusion matrix diagram of data set 1; (b) The confusion matrix diagram of data set 2. Table 3 shows the comparison results when using different algorithms. It can be seen that compared with other classical algorithms under ideal experimental data, our algorithm has better performance.

[0119] Table 2 Comparison results of the weighted average values of the four major parameters of this model when using different data sets

[0120]

[0121] Table 3 Comparison results of the weighted average values of the four major parameters when using different algorithms

[0122]

[0123] Example 2:

[0124] In this example, bearing fault audio under different working conditions is selected, and the bearing fault categories of the output data are obtained.

[0125] First, the bearing signal data is standardized to ensure that the model can process data in different ranges. Based on data set 1, fault audio data with sampling rates of 12.4KHz and 41.4KHz and bearing speeds of 600r / min and 1000r / min is continuously supplemented and collected. There are a total of nine working conditions, and the time-domain waveform examples of the rolling elements under the nine working conditions are as Figure 7 shown. The mixed-condition experiment is carried out for training and evaluation on the data set. The training process uses the training method defined in step 6, and four major parameters are selected for model performance verification.

[0126] To evaluate the mixed-condition performance of the model, three groups of mixed-condition data in Table 4 are constructed, corresponding to three mixed-condition tasks respectively.

[0127] Table 4 Statistical table of the parameters and quantities of each mixed-condition task

[0128]

[0129] Among them, for Task 1, the rotational speed of the fixed bearing is set, and the robustness of the test model to data with various different sampling rates is tested; for Task 2, the sampling rate is fixed, and the robustness of the test model to data with various different bearing rotational speeds is tested; for Task 3, the comprehensive performance of the test model in the face of mixed data is tested. The proportion of the data set is adjusted to 60% for the training set, 15% for the validation set, and 25% for the test set. The experimental results are statistically analyzed, and the four major parameters of the method of the present invention are compared with those of other comparative models under the experimental conditions of the same mixed working conditions. The results are shown in Table 5.

[0130] Table 5 Comparison results of the weighted average values of the four major parameters when using different algorithms under the mixed working condition tasks

[0131]

[0132] In addition, in order to more intuitively display the performance of the model, the confusion matrices under three groups of mixed working condition tasks are statistically analyzed. Figure 8 It is the confusion matrix diagram under three kinds of mixed working condition tasks. (a) The confusion matrix diagram of Task 1; (b) The confusion matrix diagram of Task 2; (3) The confusion matrix diagram of Task 3. The confusion matrix can clearly reflect the prediction results of each category and the distribution of classification errors.

[0133] Through the above comparison, it can be found that the present method is superior to the comparative method under most mixed working conditions. Especially under the complex working conditions including nine different kinds of data in Task 3, it has stronger robustness and stability. It is verified that the bearing fault classification method based on bidirectional temporal convolution and Transformer of the present invention has excellent robustness both in the normal environment and the mixed working condition environment, proving that it is applicable to the fault diagnosis task in the actual industrial scenario.

[0134] The above are only the preferred embodiments of the present invention, and are not intended to limit the idea of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A bearing fault classification method based on bidirectional temporal convolution and Transformer, characterized in that, Including: Step 1: Use a microphone array to collect various bearing fault audio simulated by a simulated bearing fault test bench, and preprocess their class labels; Step 2: Decompose the audio data based on the VMD method to obtain multiple IMFs with different center frequencies and bandwidths; Step 3: Perform FFT processing on the audio data to obtain the spectral information of the signal, that is, the frequency distribution characteristics; Step 4: Feature fusion is performed on the multiple IMFs obtained by the transformation and the spectral information in a stacked manner to form a feature matrix containing time-domain and frequency-domain information; Step 5: Input the feature matrix into the constructed fault classification model with a parallel structure of bidirectional temporal convolution and Transformer to obtain the probability distribution of various fault categories; Step 6: By minimizing the loss function and optimizing the evaluation index, train the model parameters until the set threshold is met, and save the model parameters; at the same time, measure the algorithm accuracy with the optimal evaluation index and evaluate the system performance.

2. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 1, wherein The specific content of Step 1 is as follows: Step 1.1: Use a microphone array to collect data from the simulated bearing fault test bench; Step 1.2: Assign class labels according to the categories to be classified, and integrate the fault audio data in a csv file, with each column corresponding to different fault category data; Step 1.3: Perform normalization processing on the data so that the values of all features are within the range of (-1, 1); Step 1.4: Divide the time series data into multiple samples according to the number of sample points with a set length.

3. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 1, characterized in that, The specific content of Step 2 is as follows: Step 2.1: Perform VMD decomposition on the faulty audio. It is an algorithm that solves the variational problem through an optimization process. The preprocessed audio signal f(t) is decomposed into K IMFs, and the bandwidth of each IMF is minimized. Each IMF is represented as u k (t), that is, the k-th IMF of the signal at time t, and its center frequency is represented as ω k , δ(t) is the Dirac function, which is used for the convolution operation of the signal. The calculation formula of the variational problem is: There are also constraint conditions: Step 2.2: In order to simplify the optimization process of the variational problem in 2.1, introduce the quadratic penalty factor α and the Lagrange multiplier operator λ(t), and transform the constrained variational problem into an unconstrained variational problem: Step 2.3: Using the IMF components and their frequency information optimized through Step 2.2, further update each IMF, and determine the final IMF components (u1, u2... u k ) and the corresponding central frequencies (ω1, ω2... ω k ) through iterative updating. Iteratively update the IMF component u k (t): Among them, is the Fourier transform of the k-th modal component after the (n + 1)-th iteration, is the Fourier transform of the preprocessed audio signal f(t); Iteratively update the center frequency ω k : wherein, is the value of the center frequency of the k-th mode after the (n + 1)-th iteration, is the Fourier transform of the k-th modal component; Step 2.4: When the iteration process converges, output the u of multiple IMFs k (t) and the central frequency ω k , and the final multiple modal feature format is B is the batch size, L is the sequence length, and C1 is equal to K.

4. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 1, characterized in that The specific content of Step 3 is as follows: Step 3.1: Convert the preprocessed audio signal f(t) from analog to digital as f[n], and then perform FFT processing to obtain its frequency distribution. The mathematical expression of FFT is based on DFT and is defined as follows: Among them, X[h] is the h-th sample in the frequency domain, f[m] is the m-th sample in the time domain, and M is the total number of samples; Step 3.2: Take the modulus of a series of complex numbers obtained after FFT processing to obtain the frequency distribution characteristics of the signal, that is, the spectral information: Among them, 5. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 1, characterized in that, The specific content of Step 4 is as follows: Stack the spectral information obtained in step 3 and the IMFs obtained in step 2 according to the feature dimension to form a feature matrix \(f\in\mathbb{R}\) B×L×D , where \(B\) represents the batch size, \(L\) represents the sequence length, and \(D = C1 + C2\) represents the sum of the spectral feature dimension and the feature dimension of the IMFs, enabling the model to utilize the features in both the time domain and the frequency domain simultaneously.

6. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 1, wherein The specific fault classification model with a parallel structure of bidirectional temporal convolution and Transformer in Step 5 is as follows: Extract local temporal features through the bidirectional temporal convolution module, and extract global features through the Transformer module; fuse the features of the two parts through the cross attention layer and output them to the classifier, and the classifier calculates the probability distribution of each fault category to complete fault classification.

7. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 6, characterized in that, The bidirectional temporal convolution module includes a forward temporal convolution part and a backward temporal convolution part, and the basic components all include: dilated causal convolution, Chomp1d layer, Relu activation function, and Dropout layer; The operation process of the forward temporal convolution part is as follows: Dilated causal convolution: In the forward temporal convolution part, dilated causal convolution is adopted to extract the features of time series data. The output of the dilated convolution at time step t is determined by the sum of the products of the values of the feature matrix f at different time steps and the elements of the convolution kernel g. The index s of the convolution kernel starts from 0 and goes up to L g - 1 to ensure that the output meets the causality requirement. The output tensor formula is: Among them, * d,c represents the dilated convolution operation, d is the dilation rate, c is the causality, and L g is the length of the convolution kernel, and g(s) is the value of the convolution kernel at position s; Chomp1d layer: After dilated causal convolution, the Chomp1d layer is applied to crop the tail padding that may be generated due to dilated convolution, ensuring that the length of the output sequence is the same as that of the input sequence. If the output tensor is represented as (f * d,c g)(t) ∈ R B ×L×C , where C is the number of channels, the formula is: (f * d,c g)(t)' = (f * d,c g)(t)[:, :L - chomp_size, :] (9) where (f * d,c g)(t)' represents the tensor after being processed by the Chomp1d layer, chomp_size is the number of elements at the end of the sequence to be removed, which is usually equal to the padding number added due to dilated convolution;[:, :L - chomp_size, :] means to keep all elements of the first dimension B of the output tensor, keep all elements of the third dimension C of the output tensor, and on the second dimension L of the output tensor, keep the elements from the beginning to the position of L - chomp_size. Relu activation function: Apply the ReLU activation function to introduce nonlinearity: (f * d,c g)(t)” = ReLU[(f * d,c g)(t)'] = max[0, (f * d,c g)(t)'] (10) where, (f * d,c g)(t)” represents the tensor processed by the ReLU function; Dropout layer: Add a Dropout layer after the ReLU activation function to randomly discard some neurons in the network to prevent the model from overfitting to the training data: X forward = (f * d,c g)(t)” ε Bernoulli(p) (11) Among them, ⊙ represents element-wise multiplication, and Bernoulli(p) is a Bernoulli random variable, where the parameter p is the probability of retaining neurons; The operation process of the backward time convolution part is as follows: Define the backward temporal convolution part in the bidirectional temporal convolution module. Similar to the forward temporal convolution, the backward temporal convolution first flips the input sequence, then applies the same structure as the forward temporal convolution to extract features, and finally flips the output again to restore its original temporal order, obtaining After concatenating the outputs of the forward-time convolutional part and the backward-time convolutional part in the feature dimension, Dropout is performed to reduce overfitting, resulting in the final local feature representation 8. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 6, wherein Define the transformer module to include the following basic components: input embedding and position encoding, multi-head attention mechanism MHA, residual connection, LN layer, and FFN layer. The operation of the transformer module is as follows: Input Embedding and Positional Encoding: Convert the feature matrix \(f\in\mathbb{R}\) containing time-domain and frequency-domain information B×L×D into an embedding vector \(z_0\in\mathbb{R}\) B×L×D and add positional encoding to preserve the sequential information of time steps. The definition of positional encoding is as follows: Among them, is the index of the t time step, and i is the index of the feature dimension; The calculation result of the input embedding is: z0 = f + PE (13) MHA layer: In the multi-head attention mechanism, the input z0 ∈ R B×L×D is linearly transformed into three parts, namely the query key value d k is the feature dimension of the query and the key, and d v is the feature dimension of the value. The calculation formula for a single attention head is: The multi-head attention mechanism forms the final output by parallelly calculating the results of multiple attention heads, concatenating these results, and then mapping them back to the feature dimension through the output weight matrix; LN and residual connection: After each multi-head attention mechanism, use residual connection and layer normalization LN to improve training stability and model performance. The residual connection is expressed as: z' = z0 + MHA(z0) (15) where z' ∈ R B×L×D ; The application of LN is expressed as: where, μ ∈ R and σ ∈ R are the mean and variance of the feature dimension, ε is used to prevent division by zero, γ ∈ R and β ∈ R are learnable scaling and offset parameters; FFN: The output of the multi-head attention mechanism is further passed through a two-layer feed-forward network for feature extraction and non-linear transformation; residual connection and layer normalization are performed again after FFN to obtain The above module is stacked 4 layers by stacking multiple layers of encoders to obtain the final global feature representation X transformer ∈R B×L×D .

9. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 6, wherein, Fuse the features of the two parts through the cross attention layer and output them to the classifier. The classifier calculates the probability distribution of each fault category to complete fault classification. Specifically: The Cross attention module is used to model the correlation of features from two feature extraction modules and achieve information fusion between different features through the cross-attention mechanism; the local features output by the bidirectional temporal convolutional module are used as the query sequence Q, and the global features X output by the transformer module transformer ∈R B×L×D are used as the key sequence K and the value sequence V at the same time, and then Q emb , K emb , V emb ∈R B×L×D are obtained through linear mapping. The attention weights are calculated using the projected query sequence and key sequence, and the value sequence is weighted: where X ∈ R B×L×D is the fused feature representation. The local features are used as queries, which can actively search for relevant information from the global features, strengthen the emphasis on fine-grained spatio-temporal features, and at the same time supplement the global context information. The global features provide the overall temporal dependence and are suitable as the target for queries; The fused feature X ∈ R B×L×D First, swap the positions of the feature dimension D and the time dimension L, and then use the pooling operation to compress the time dimension L to 1, resulting in X ∈ R B×D×1 , flatten the pooled feature tensor into a two-dimensional form flat_tensor ∈ R B×D , input it into the fully connected layer to map the feature to the target classification space and finally output outputs ∈ R B×C , where C is the number of classification categories; the output of each sample is the scores of C categories, which are used for subsequent classification tasks.

10. The bearing fault classification method based on bidirectional temporal convolution and Transformer according to claim 1, characterized in that, The specific content of step 6 is: Step 6.1: Perform fault classification on the collected data set. The sampling frequency and feature extraction method of the data set samples are as in steps 2, 3, and 4. The feature matrix containing time-domain and frequency-domain information obtained after preprocessing is used as the input of the fault classification model; Step 6.2: Use the predicted probability distribution and the true class distribution as the loss function, and use accuracy, precision, recall, and F1 score as the optimal evaluation metrics. The expression of the loss function is: where y is the probability distribution of the true labels, is the probability distribution predicted by the model, and C is the number of classification categories; Step 6.3: Minimize the predicted probability distribution and the true class distribution until the number of training times reaches the set threshold or the value of the loss function reaches the set range, then it is considered that the model parameters have been trained, and save the model parameters; at the same time, select the optimal evaluation metric to measure the accuracy of the algorithm and evaluate the performance of the system.

Citation Information

Cited By

  • Bearing fault diagnosis method based on variational mode decomposition and time sequence block cross attention fusion

    CN121881113A