Few-sample bearing fault diagnosis method based on converter and mahalanobis distance
Through the method of combining the converter and Mahayana distance, the problem of data acquisition difficulties in bearing fault diagnosis is solved, efficient and accurate diagnosis under the condition of few samples is achieved, and the performance of bearing fault diagnosis is improved.
Patent Information
- Application Number
- CN202510443512.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
AI Technical Summary
When data acquisition is difficult, existing bearing fault diagnosis methods have problems such as insufficient diagnostic accuracy and poor robustness. Especially traditional machine learning methods are time-consuming and labor-intensive, deep learning models are overly dependent on local features and are susceptible to noise under small samples.
The method based on the transformer and Marshall distance is adopted, and the bearing fault diagnosis is performed through multi-scale feature extraction, global converter and local Marshall distance calculation, combined with weighted fusion and nearest neighbor strategy, and the multi-scale large core feature extraction module is used to share weights, reducing the amount of parameters, and improving training speed and diagnostic accuracy.
Under the condition of few samples, the accuracy and stability of bearing fault diagnosis is significantly improved, and the calculation complexity is reduced, making it suitable for industrial scenarios where data acquisition is difficult.
Smart Images

Figure CN120372392A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bearing fault diagnosis, and particularly relates to a few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance. Background Art
[0002] In industrial fields such as manufacturing, aerospace, and power electronics, electric motors are key components of mechanical equipment. As an important part of electric motors, the fault diagnosis of bearings is crucial for ensuring the normal operation of equipment. According to statistics, approximately 40% of electric motor faults are attributed to bearing faults. Traditional bearing fault diagnosis methods mainly rely on vibration signal analysis, and these methods show good performance when there is sufficient data. However, in practical applications, it is often difficult and costly to obtain a large amount of fault data. Therefore, few-shot learning methods have important application value in bearing fault diagnosis.
[0003] Traditional machine learning methods such as diagnostic methods combined with principal component analysis (PCA), support vector machine (SVM), particle swarm optimization-support vector machine (PSO-SVM), etc. require manual feature extraction and adjustment of a large number of hyperparameters, which is time-consuming and laborious. Although the convolutional neural network (CNN) in deep learning overcomes the deficiencies of traditional methods to a certain extent, it still faces problems such as difficulty in data acquisition and over-reliance on local features in bearing fault diagnosis. Although few-shot learning, as a meta-learning method, has been introduced, existing models such as those based on siamese neural network (SNN) and its improved models, prototype network, matching network, etc. have problems such as complex methods, multi-stage diagnosis, smooth distance metrics being vulnerable to noise when the data volume is small, and high costs of advanced metrics when the number of samples is small. Therefore, there is an urgent need for an efficient and accurate bearing fault diagnosis technology applicable to few samples. Summary of the Invention
[0004] To solve the deficiencies of the prior art and achieve the purpose of improving the accuracy and robustness of bearing fault diagnosis when the training data is limited, the present invention adopts the following technical solutions:
[0005] A few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance, comprising the following steps:
[0006] Step 101: Obtain bearing vibration signals to construct a support data set and a query data set, convert the data set into a spectrogram for feature extraction, and extract multi-scale features;
[0007] Step 102: Pass the extracted features through a global transducer Transformer to obtain the global similarity between the support data set and the query data set;
[0008] Step 103: Calculate the local Mahalanobis distance for the extracted features to measure the local similarity between the support dataset and the query dataset;
[0009] Step 104: Perform weighted fusion on the global similarity and the local similarity to generate the total similarity, and classify the query dataset through the nearest neighbor strategy to determine the bearing fault category.
[0010] Further, the feature extraction in step 101 adopts a multi-scale large kernel feature extraction module. The multi-scale large kernel feature extraction module includes multiple residual large kernel attention modules. The residual large kernel attention module extracts local features through the depthwise separable convolution it contains, captures long-range dependencies through dilated depthwise separable convolution, and mixes information between different channels through pointwise convolution to extract multi-scale features. And the input feature map and the output feature map are added through a residual connection to retain the input information and prevent gradient vanishing. Among them, the feature extraction of the support dataset and the query dataset shares weights. The shared weights refer to that the residual large kernel attention module uses the same convolutional kernel weights at different scales. By sharing weights at different scales, the number of model parameters can be significantly reduced. In the training and inference processes of the weight sharing mode, the convolutional kernel weights in the same residual large kernel attention module can be reused multiple times, reducing repeated calculations. This not only speeds up the training speed of the model but also reduces the computational complexity during inference. At the same time, since all residual large kernel attention modules share the same weights, the features extracted by them at different scales have consistency.
[0011] Further, the multi-scale large kernel feature extraction module includes a convolutional layer module, a residual large kernel attention module, and a feedforward convolutional neural network FFCN module;
[0012] The convolutional layer module extracts local features from the input spectrogram through the convolutional layer, reduces the dimension through the max-pooling layer to reduce the amount of calculation. Then, layer normalization is performed to stabilize the training process and prevent gradient explosion, and activation processing is introduced to introduce non-linearity to enhance the expression ability of features;
[0013] The residual large kernel attention module obtains the feature map output by the convolutional layer module, extracts multi-scale features and performs layer normalization;
[0014] The feedforward convolutional neural network module performs non-linear fusion through the convolutional layer and the activation function to generate the support dataset feature vector and the query dataset feature vector.
[0015] Further, the calculation formula of the residual large kernel attention module is as follows:
[0016]
[0017] Among them, ResLKA represents the residual large kernel attention module, X represents the input feature map, DW represents the depthwise separable convolution, DDW represents the dilated depthwise separable convolution, and PW represents the pointwise convolution. represents element-wise multiplication;
[0018] The layer normalization enhances the expressive ability of features and unifies the feature scale, avoiding certain features from dominating the training process due to their large scales, and normalizes all features of each sample. The formula is as follows:
[0019]
[0020] where x represents the input feature, μ represents the mean of the input feature, σ represents the standard deviation of the input feature, and γ and β represent learnable parameters used to adjust the normalized feature scale and offset. The role of the layer normalization module is to stabilize the training process, improve the convergence speed of the model, and enhance the expressive ability of features. By normalizing the input of each layer, layer normalization can reduce the problems of gradient explosion or gradient vanishing, making the model more stable during the training process;
[0021] The role of the residual connection is to directly add the input feature map to the output feature map after being processed by the residual large kernel attention module and layer normalization, thereby retaining the original input information and preventing gradient vanishing. The specific operation is as follows:
[0022] F out = F norm + X
[0023] where X represents the input feature map, and F norm represents the feature map after being processed by the residual large kernel attention module and layer normalization, and F out represents the output feature map after the residual connection.
[0024] Furthermore, the global Transformer in step 102 includes an encoder and a decoder. The extracted features are respectively subjected to a dimensionality reduction operation by the global average pooling module to make the model training process more stable and reduce the risk of overfitting. The encoder finds the feature correlations between each category of samples in the support dataset, and the decoder projects based on the output of the encoder to generate learnable weights for calculating the global similarity between the query dataset samples and the support dataset samples.
[0025] Furthermore, the encoder linearly transforms the support dataset features to obtain (Q s , K s , V s ), and inputs them into the scaled dot-product attention module to calculate the correlations between the support dataset samples using the dot-product technique. The formula is as follows:
[0026]
[0027] Among them, Q s represents the query vector of the encoder, K s represents the key vector of the encoder, V s represents the value vector of the encoder, and c is the dimension of K s used to scale the dot product result to prevent gradient explosion. This attention mechanism helps to synthesize information in a concentrated and selective manner in the support dataset; the output information of the scaled dot product attention mechanism is represented as follows:
[0028]
[0029] Among them, LayerNorm represents layer normalization, which is used to stabilize the training process. The output ε of the encoder is processed by a feed-forward network, and the formula is as follows:
[0030]
[0031] Among them, FFN represents the network, which is used for non-linear transformation of features, represents the processed output information.
[0032] Furthermore, the decoder obtains the query dataset features through linear transformation to get (Q q , K q , V q ), and inputs them into the scaled dot product attention module to calculate the correlation, and the formula is as follows:
[0033]
[0034] Among them, Q q represents the query vector of the decoder, K q represents the key vector of the decoder, V q represents the value vector of the decoder, and c is the dimension of K q ; then, the intermediate output of the decoder is generated through layer normalization and residual connection:
[0035]
[0036] Then, a linear projection is performed on the intermediate output D of the decoder to generate the query vector :
[0037]
[0038] Among them, W Q' represents the learnable weight matrix, and then, the output of the encoder Perform a linear projection to generate W K and W V :
[0039]
[0040] where W K and W V represent learnable weight matrices, and then calculate the similarity between the query dataset samples and the support dataset samples through dot product:
[0041]
[0042] Finally, the decoder generates the final output through the cross-attention mechanism module The formula is as follows:
[0043]
[0044] The final output of the decoder is the global similarity matrix between the query dataset samples and the support dataset samples. This matrix represents the correlation between the query dataset samples and the support dataset samples and is used for subsequent classification tasks. The formula is as follows:
[0045]
[0046] where z global is the global similarity matrix, representing the global similarity between the query dataset and the support dataset.
[0047] Furthermore, in step 103, obtain the features extracted from the support dataset. For each category, combine the pixel values at the same position into local feature vectors, calculate the local covariance through the local feature vectors, and calculate the Mahalanobis distance between the sample features in the query dataset and the local covariance to generate a local similarity matrix.
[0048] Furthermore, the features extracted from the support dataset For the features of category k Generate a set of local feature vectors where d = C × h × ω, n s represents the total number of samples in the support set, h represents the height of the feature map, ω represents the width of the feature map, c represents the number of channels of the feature map, C represents the number of samples in each category, and d represents the total number of local feature vectors;
[0049] For category k, calculate the local covariance matrix :
[0050]
[0051] Among them, represents the local feature vector of a single category k, and μ k represents the mean of the local feature vectors of category k;
[0052] For each sample F in the query dataset q , calculate its Mahalanobis distance from the local covariance matrix of category k:
[0053]
[0054] Combine the calculation results of the Mahalanobis distances of each sample in the query dataset from all categories k into a local similarity matrix z local :
[0055]
[0056] where z local is the local similarity matrix, representing the local similarity between the query dataset and the support dataset.
[0057] Furthermore, in step 104, through a learnable weight matrix ω, the global similarity matrix z global output by the global transformer is transformed to obtain a new global similarity matrix:
[0058] M global = z global × ω T
[0059] The local similarity matrix z local obtained by calculating the local Mahalanobis distance is subjected to one-dimensional convolution calculation to obtain a new local
[0060] similarity matrix:
[0061] M local = conv1d(z local )
[0062] The new global similarity matrix M global and the new local similarity matrix M local are weighted and fused to generate a total similarity matrix:
[0063] M total = μ × M global + (1 - μ) × M local
[0064] where μ ∈ (0, 1) is a hyperparameter used to balance the weights of global and local similarity information;
[0065] Classify using the nearest neighbor strategy; for each sample in the query dataset, find the sample in the support dataset with the highest similarity to it and use its label as the classification result Y s i 。
[0066] The advantages and beneficial effects of the present invention are as follows:
[0067] The few-shot bearing fault diagnosis method based on transformers and Mahalanobis distance of the present invention can efficiently extract features and fuse global and local information under few-shot conditions by combining multi-scale feature extraction and global-local information fusion, significantly improving the accuracy and stability of diagnosis, enhancing the performance of bearing fault diagnosis, and being applicable to industrial scenarios where data acquisition is difficult. Brief Description of the Drawings
[0068] Figure 1 is a flowchart of executing the bearing fault diagnosis method in an embodiment of the present invention.
[0069] Figure 2 is a schematic structural diagram of the multi-scale large kernel feature extraction module in an embodiment of the present invention.
[0070] Figure 3 is a schematic internal structure diagram of the convolutional layer in an embodiment of the present invention.
[0071] Figure 4 is a schematic internal structure diagram of the residual large kernel attention module in an embodiment of the present invention.
[0072] Figure 5 is a schematic internal structure diagram of the non-linear combination module FFCN in an embodiment of the present invention.
[0073] Figure 6 is a schematic internal structure diagram of the global Transformer module in an embodiment of the present invention.
[0074] Figure 7 is a schematic internal structure diagram of the local Mahalanobis distance metric module in an embodiment of the present invention.
[0075] Figure 8 is a schematic internal structure diagram of the integration and verification module in an embodiment of the present invention. Detailed Embodiments
[0076] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustration and explanation of the present invention, and are not intended to limit the present invention.
[0077] Such as Figure 1As shown, a few-shot bearing fault diagnosis method based on Transformer and Mahalanobis distance metric learning first performs data preprocessing. In the embodiments of the present invention, the dataset is divided into a support dataset and a query dataset. The support dataset is the model training dataset, and the query dataset is the model validation dataset. Preprocessing operations need to be performed before inputting the dataset into the model. The purpose is to convert the input vibration signal into a spectrogram through short-time Fourier transform (STFT) to more effectively separate features. The preprocessing process includes: First, the vibration signal obtained from the sensor, usually time series data; then, use short-time Fourier transform to convert the time series data into a spectrogram; the spectrogram is a two-dimensional image, where the horizontal axis represents time, the vertical axis represents frequency, and the color or grayscale represents the amplitude. For example but not limited to, a spectrogram can be generated using a sliding window of 2048 points and a step size of 80 points. The bearing fault diagnosis method specifically includes the following steps:
[0078] Step 101: Use a multi-scale large kernel feature extraction (MLKFE) module to extract multi-scale features from the spectrogram.
[0079] Step 102: Input the extracted features into the global Transformer module, process the global information through the encoder and decoder, and obtain the correlation between the query set and the support set.
[0080] Step 103: Input the extracted features into the local Mahalanobis distance module, calculate the local covariance matrix of each category in the support set, and use the Mahalanobis distance to measure the correlation between the samples in the query set and the support set.
[0081] Step 104: Perform weighted fusion on the outputs of the global Transformer module 102 and the local Mahalanobis distance module 103, and classify through the nearest neighbor strategy to determine the fault category to which the samples in the query set belong.
[0082] Step 101 of the method is a multi-scale large kernel feature extraction (MLKFE) module, and its internal structure is as Figure 2As shown. This method uses two multi-scale large kernel feature extraction modules to process the support dataset and the query dataset in parallel. Weights need to be shared between the two multi-scale large kernel feature extraction modules. In the multi-scale large kernel feature extraction module, multiple residual large kernel attention modules are used to extract features of different scales. Weight sharing means that these residual large kernel attention modules use the same convolutional kernel weights at different scales, rather than training a separate set of weights for each scale. By sharing weights at different scales, this method can significantly reduce the number of model parameters. In the weight sharing mode of this method, during the training and inference processes, the convolutional kernel weights in the same residual large kernel attention module can be reused multiple times, reducing redundant calculations. This not only speeds up the model training speed but also reduces the computational complexity during inference. At the same time, since all residual large kernel attention modules share the same weights, the features they extract at different scales are consistent.
[0083] Next, the specific processing steps of the MLKFR module 101 are described in detail.
[0084] The specific processing process of the MLKFE module 101 is as follows: The spectrogram after data preprocessing is first input to the convolutional layer module 201, and its internal structure is as Figure 3 shown. The convolutional layer module first extracts local features through the convolutional layer, then reduces the spatial dimension of the feature map through the max pooling layer to reduce the computational amount, then uses layer normalization to stabilize the training process and prevent gradient explosion, and finally introduces non-linearity through the activation function to enhance the expression ability of the features. This structure can be repeated multiple times to extract deeper features. In this method, the convolutional layer module 201 usually serves as the pre-module of the MLKFE module 101 to provide an initial feature representation for the residual large kernel attention module 202 for subsequent further extraction of multi-scale features.
[0085] The feature map after feature extraction by the convolutional layer module 201 is then input to k residual large kernel attention modules 202, and its internal structure is as Figure 4 shown. In this module, local features are first extracted through depthwise separable convolution, then dilated depthwise separable convolution is used to capture long-range dependencies, then pointwise convolution is used to mix information between different channels, and finally the input feature map and the output feature map are added through a residual connection to retain the original input information and prevent gradient disappearance. The specific formula is:
[0086]
[0087] where, denotes element-wise multiplication, X is the input feature map, ResLKA is the residual large kernel attention module, DW is the depthwise separable convolution, DDW is the dilated depthwise separable convolution, and PW is the pointwise convolution.
[0088] Next, layer normalization is performed to enhance the expressive ability of features and unify the feature scale, avoiding certain features from dominating the training process due to their large scales. Normalization is performed on all features of each sample, and the specific formula is as follows:
[0089]
[0090] where \(x\) is the input feature, \(\mu\) is the mean of the input features, \(\sigma\) is the standard deviation of the input features, and \(\gamma\) and \(\beta\) are learnable parameters used to adjust the scale and offset of the normalized features. The role of the layer normalization module is to stabilize the training process, improve the convergence speed of the model, and enhance the expressive ability of features. By normalizing the input of each layer, layer normalization can reduce the problems of gradient explosion or gradient vanishing, making the model more stable during training.
[0091] The role of the residual connection is to directly add the input feature map to the output feature map after passing through the residual large kernel attention module and layer normalization, thereby retaining the original input information and preventing gradient vanishing. The specific operation is expressed as:
[0092] F out = F norm + X
[0093] where \(X\) is the input feature map, \(F\) norm is the feature map after passing through the residual large kernel attention module and layer normalization, and \(F\) out is the output feature map after the residual connection.
[0094] Subsequently, the feature map after the residual connection is input into the FFCN module 203, and its internal structure is as Figure 5 shown. Since the residual large kernel attention module extracts rich local and global features through convolutional operations of different scales, but these features may be relatively scattered in the spatial and channel dimensions, the FFCN module effectively fuses them through a series of convolutional layers and non-linear activation functions. The specific process is as follows: First, features are extracted through a convolutional layer with a kernel size of 3×3, then non-linearity is introduced through an activation function (such as ReLU or GELU, but not limited to this), and finally, optionally, the training process is stabilized through a normalization layer. The FFCN module further performs non-linear combination and fusion on the multi-scale features extracted by the residual large kernel attention module.
[0095] After being processed by the MLKFE module, the preprocessed spectrogram in this method is converted into the query dataset feature vector \(F\) q and the support dataset feature vector \(F\) s , and the two feature vectors are both input into the global Transformer module 102 and the local Mahalanobis distance module 103 to extract local and global features.
[0096] Next, the specific steps of the global Transformer module 102 will be described in detail.
[0097] The internal structure of the global Transformer module 102 is as Figure 6 shown. This module is used to process global information and capture the correlation between the query dataset and the support dataset. The specific process is as follows: Feature vectors F q and feature vector F s After entering the module, they first undergo dimensionality reduction operations through the global average pooling module (GAP) respectively. Global average pooling makes the model training process more stable and reduces the risk of overfitting by reducing the feature dimensions. Then, the encoder module 301 finds the feature correlations between each class of samples in the support dataset. Specifically, from the support dataset features obtain (Q s , K s , V s ) through linear transformation and input them into the scaled dot-product attention module 303. This module uses the dot-product technique to calculate the correlation between support samples, and the formula is:
[0098]
[0099] where c is the dimension of K s and is used to scale the dot-product result to prevent gradient explosion. This attention mechanism helps to synthesize information in a concentrated and selective manner in the support dataset. The output information of the scaled dot-product attention mechanism can be expressed as:
[0100]
[0101] where LayerNorm is layer normalization, which is used to stabilize the training process. The output of the encoder can be further processed by the feed-forward network (FFN), and the formula is:
[0102]
[0103] where FFN is the feed-forward network, which is used to perform non-linear transformation on the features.
[0104] Next, the decoder module 302 calculates the global similarity between the query dataset samples and the support dataset samples. Specifically, from the query dataset obtain (Q q , K q , V q ) through linear transformation and input them into the scaled dot-product attention module 303 to calculate the correlation, and the formula is:
[0105]
[0106] Among them, c is the dimension of K q . Then, the intermediate output D of the decoder is generated through layer normalization and residual connection:
[0107]
[0108] Then, a linear projection is performed on the intermediate output D of the decoder to generate a query vector :
[0109]
[0110] Among them, W Q' is a learnable weight matrix. Then, a linear projection is performed on the output of the encoder to generate W K and W V :
[0111]
[0112] Among them, W K and W V are learnable weight matrices. Then, the similarity between the query sample and the support sample is calculated through dot product:
[0113]
[0114] Finally, the decoder generates the final output through the cross-attention mechanism module, and its formula is:
[0115]
[0116] The final output of the decoder is the global similarity matrix between the query sample and the support sample. This matrix represents the correlation between the query sample and the support sample and can be used for subsequent classification tasks. The formula is:
[0117]
[0118] Among them, z global is the global similarity matrix, representing the global similarity between the query dataset and the support dataset.
[0119] The internal structure of the local Mahalanobis distance module 103 described in this method is as Figure 7 shown. The module 103 is used to process local information and capture the correlation between the query dataset and the support dataset.
[0120] Next, the specific steps of the local Mahalanobis distance module 103 will be introduced in detail.
[0121] First, generate local feature vectors by combining the features extracted from the support dataset as input. For each class k, combine the pixel values at the same position into local feature vectors. Specifically, for the features of class k generate a set of local feature vectors where d = C × h × ω, and n s is the total number of samples in the support set, h is the height of the feature map, ω is the width of the feature map, c is the number of channels of the feature map, C is the number of samples per class, and d is the total number of local feature vectors.
[0122] Then, calculate the local covariance matrix. For class k, calculate the local covariance matrix :
[0123]
[0124] where μ k is the mean of the local feature vectors of class k.
[0125] Next, calculate the Mahalanobis distance. For each sample in the query dataset calculate its Mahalanobis distance from the local covariance matrix of class k :
[0126]
[0127] Finally, combine the Mahalanobis distance calculation results of each sample in the query dataset with all classes k into a local similarity matrix z local :
[0128]
[0129] where z local is the local similarity matrix, representing the local similarity between the query dataset and the support dataset.
[0130] The internal structure of the integration and verification module 104 described in this method is as shown in Figure 8 The main function of the module is to fuse the global and local similarity information captured by the global Transformer module and the local Mahalanobis distance module, query the similarity score of the score, and generate the final classification result.
[0131] Next, the specific steps of the integration and verification module 104 will be introduced in detail.
[0132] First, calculate the global similarity matrix M global , the global Transformer module outputs the global similarity matrix z global , and through a learnable weight matrix ω, zglobal Convert to the global metric M global :
[0133] M global = z global × ω T
[0134] Then, calculate the local similarity matrix M local , by performing one-dimensional convolution on z local to calculate the local metric M local :
[0135] M local = conv1d(z local )
[0136] Next, perform weighted fusion on the global similarity matrix M global and the local similarity matrix M local to generate the total similarity matrix M total :
[0137] M total = μ × M global + (1 - μ) × M local
[0138] where μ ∈ (0, 1) is a hyperparameter used to balance the weights of global and local similarity information.
[0139] Finally, use the nearest neighbor strategy for classification. For each sample in the query set find the support set sample with the highest similarity to it and use its label as the classification result
[0140] The present invention also proposes a few-shot bearing fault diagnosis system based on Transformer and Mahalanobis distance metric learning, which solves the problems of large data requirements, low feature extraction efficiency, and insufficient few-shot diagnosis accuracy of traditional deep learning models, and includes a multi-scale large kernel feature extraction (MLKFE) module, a global Transformer module, a local Mahalanobis distance module, and an integration and verification module.
[0141] The multi-scale large kernel feature extraction (MLKFE) module efficiently extracts multi-scale features, combines with the global Transformer module to capture the correlation between the query set and the support set, and the local Mahalanobis distance module ensures the non-singularity of the covariance matrix and enhances robustness. Finally, the integration and verification module performs weighted fusion and nearest neighbor classification to achieve end-to-end diagnosis. Compared with the prior art, the present invention improves the diagnosis accuracy and stability under few-shot conditions, reduces data dependence and computational complexity, and provides reliable support for the safe operation of equipment.
[0142] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A few-shot bearing fault diagnosis method based on a transformer and Mahalanobis distance, characterized in that It includes the following steps: Step 101: Obtain bearing vibration signals to construct a support dataset and a query dataset, convert the datasets into spectrograms for feature extraction, and extract multi-scale features; Step 102: Pass the extracted features through a global transformer to obtain the global similarity between the support dataset and the query dataset; Step 103: Pass the extracted features through calculating the local Mahalanobis distance to measure the local similarity between the support dataset and the query dataset; Step 104: Perform weighted fusion on the global similarity and the local similarity to generate the total similarity, and classify the query dataset through the nearest neighbor strategy to determine the bearing fault category.
2. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 1, wherein: The feature extraction in Step 101 adopts a multi-scale large kernel feature extraction module. The multi-scale large kernel feature extraction module includes multiple residual large kernel attention modules. The residual large kernel attention module extracts local features through the depthwise separable convolution it contains, captures long-range dependencies through dilated depthwise separable convolution, and mixes information between different channels through pointwise convolution to extract multi-scale features, and adds the input feature map and the output feature map through residual connection. Among them, the feature extraction of the support dataset and the query dataset shares weights. The shared weight means that the residual large kernel attention module uses the same convolution kernel weights at different scales.
3. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 2, characterized in that: The multi-scale large kernel feature extraction module includes a convolutional layer module, a residual large kernel attention module, and a feedforward convolutional neural network module; The convolutional layer module extracts local features from the input spectrogram through the convolutional layer, reduces the dimension through the max pooling layer, and then performs layer normalization and activation processing; The residual large kernel attention module obtains the feature map output by the convolutional layer module, extracts multi-scale features and performs layer normalization; The feedforward convolutional neural network module performs non-linear fusion through the convolutional layer and the activation function to generate the support dataset feature vector and the query dataset feature vector.
4. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 3, characterized in that: The calculation formula of the residual large kernel attention module is as follows: Among them, ResLKA represents the residual large kernel attention module, X represents the input feature map, DW represents depthwise separable convolution, DDW represents dilated depthwise separable convolution, and PW represents pointwise convolution. denotes element-wise multiplication; The formula of the layer normalization is as follows: Among them, x represents the input feature, μ represents the mean of the input feature, σ represents the standard deviation of the input feature, γ and β represent learnable parameters used to adjust the scale and offset of the normalized feature. The role of the layer normalization module is to stabilize the training process, improve the convergence speed of the model, and enhance the expression ability of the features. By normalizing the input of each layer, the layer normalization can reduce the problems of gradient explosion or gradient disappearance, making the model more stable during the training process; The role of the residual connection is to directly add the input feature map to the output feature map after being processed by the residual large kernel attention module and layer normalization, so as to retain the original input information and prevent gradient disappearance. The specific operation is as follows: F out = F norm + X Among them, X represents the input feature map, and F norm represents the feature map after passing through the residual large kernel attention module and layer normalization, and F out represents the output feature map after residual connection.
5. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 1, wherein: The global transformer in Step 102 includes an encoder and a decoder. The extracted features are respectively subjected to a dimensionality reduction operation through the global average pooling module. The encoder finds the feature correlation between each category sample in the support dataset, and the decoder projects based on the output of the encoder to generate learnable weights for calculating the global similarity between the query dataset samples and the support dataset samples.
6. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 5, characterized in that: The encoder will support dataset features obtained by linear transformation to (Q s , K s , V s ), and input to the scaled dot-product attention module to calculate the correlation between support dataset samples using the dot-product technique. The formula is as follows: Among them, Q s represents the query vector of the encoder, K s represents the key vector of the encoder, V s represents the value vector of the encoder, c is the dimension of K s used to scale the dot product result, and the output information of the scaled dot product attention mechanism is expressed as follows: Among them, LayerNorm represents layer normalization, and the output ε of the encoder is processed by a feed-forward network, and the formula is as follows: Among them, FFN represents a network used for non-linear transformation of features, indicating the processed output information.
7. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 5, wherein: The decoder will query the dataset features and obtain (Q q , K q , V q ) through linear transformation, and input them into the scaled dot-product attention module to calculate the correlation. The formula is as follows: Among them, Q q represents the query vector of the decoder, K q represents the key vector of the decoder, V q represents the value vector of the decoder, and c is the dimension of K q ; then, the intermediate output of the decoder is generated through layer normalization and residual connection: Then, perform a linear projection on the intermediate output D of the decoder to generate a query vector Among them, W Q' represents a learnable weight matrix, and then, the output of the encoder is linearly projected to generate W K and W V : Among them, W K and W V represent learnable weight matrices, and then calculate the similarity between the query dataset samples and the support dataset samples through dot product: Finally, the decoder generates the final output through the cross-attention mechanism module The formula is as follows: The final output of the decoder is the global similarity matrix between the query dataset samples and the support dataset samples, which represents the correlation between the query dataset samples and the support dataset samples and is used for subsequent classification tasks. The formula is as follows: Among them, z global is the global similarity matrix, representing the global similarity between the query dataset and the support dataset.
8. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 1, wherein: In step 103, the features extracted from the support dataset are obtained. For each category, the pixel values at the same position are combined into local feature vectors, the local covariance is calculated through the local feature vectors, and for the sample features in the query dataset, the Mahalanobis distance between them and the local covariance is calculated to generate a local similarity matrix.
9. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 8, characterized in that: Features extracted from the support set For the features of class k Generate a set of local feature vectors where d = C × h × ω, n s represents the total number of samples in the support set, h represents the height of the feature map, and ω represents the width of the feature map c represents the number of channels of the feature map, C represents the number of samples in each category, and d represents the total number of local feature vectors; For class k, calculate the local covariance matrix Among them, represents the local feature vector of a single category k, and μ k represents the mean of the local feature vectors of category k; For each sample in the query dataset compute its Mahalanobis distance from the local covariance matrix of class k as follows: Combine the Mahalanobis distance calculation results of each sample in the query dataset with all classes k to form a local similarity matrix z local : Among them, z local is the local similarity matrix, representing the local similarity between the query data set and the support data set.
10. The few-shot bearing fault diagnosis method based on a transducer and Mahalanobis distance according to claim 1, wherein: In the step 104, a new global similarity matrix is obtained by transforming the global similarity matrix z output by the global transformer through a learnable weight matrix ω global as follows: M global = z global × ω T The local similarity matrix z obtained by calculating the local Mahalanobis distance local Perform one-dimensional convolution calculation to obtain a new local similarity matrix: M local = conv1d(z local ) The new global similarity matrix M global and the new local similarity matrix M local are weighted and fused to generate the total similarity matrix: M total = μ × M global + (1 - μ) × M local Among them, μ ∈ (0, 1) is a hyperparameter used to balance the weights of global and local similarity information; The nearest neighbor strategy is used for classification; for each sample in the query dataset, the sample in the support dataset with the highest similarity to it is found and its label is used as the classification result.