Mechanical fault diagnosis method based on MK-ACFormer model

By constructing a MK-ACFormer model combined with a multi-scale nuclear attention convolutional neural network (MK-ACNN) and the Transformer module, the problem of poor mechanical fault diagnosis effect in noise environments is solved, and the fault diagnosis effect with high accuracy and strong robustness is achieved.

CN120046040APending Publication Date: 2025-05-27GUANGDONG OCEAN UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510141130.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing mechanical fault diagnosis method based on the Transformer model is poor in diagnosis in noisy environments, generalization ability needs to be improved, and there is a lack of effective utilization of multi-scale convolutional channel information and effective fusion of feature information.

Method used

A multi-scale nuclear attention convolutional neural network (MK-ACNN) combined with the Transformer module is used to build an MK-ACFormer model for mechanical fault diagnosis. The model extracts and fuses multi-scale features of vibration signals through feature screening, multi-scale feature extraction and feature fusion modules, improving the accuracy and robustness of diagnosis.

Benefits of technology

It significantly improves the fault diagnosis capability of the model in a noisy environment, enhances generalization and robustness, and achieves high-accuracy fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046040A_ABST
    Figure CN120046040A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical fault diagnosis method based on an MK-ACFormer model. The method comprises the following steps: acquiring a vibration signal of a to-be-detected mechanical key component; inputting the vibration signal of the to-be-detected mechanical key component into a fault diagnosis model to obtain a fault diagnosis result; wherein the fault diagnosis model is trained through a training set, the training set is mechanical key component vibration signals under different health conditions, and the fault diagnosis model is constructed based on a multi-scale kernel attention convolutional neural network and a Transform module. According to the method, the fault diagnosis capability of the model in a noise environment can be remarkably improved, the robustness is very high, the fault diagnosis accuracy is very high on different data sets, and the generalization is very high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault diagnosis, and particularly relates to a mechanical fault diagnosis method based on the MK-ACFormer model. Background Art

[0002] Mechanical equipment is widely used in various industries. As key components of mechanical equipment, bearings and gears are prone to damage during long-term operation, leading to equipment failures, serious economic losses and personnel accidents. Researching advanced fault diagnosis technologies and accurately diagnosing faults in equipment can greatly improve economic and social benefits.

[0003] Intelligent fault diagnosis methods based on shallow learning identify faults by combining signal processing technology and pattern recognition. Signal processing methods such as Fourier transform, wavelet transform, and empirical mode analysis are used to extract fault features from vibration signals, and then shallow machine learning models such as k-nearest neighbor, artificial neural network, and support vector machine algorithms are used for pattern recognition.

[0004] Deep learning methods have received extensive attention from scholars because they can directly extract features from raw data and can perform end-to-end diagnosis. Bearing fault diagnosis methods based on convolutional neural networks such as AlexNet, ResNet, GoogLeNet, and ShuffleNet have emerged continuously. The self-attention network based on Transformer was first proposed in 2017 and has since been widely used in fields such as natural language and machine vision due to its many advantages such as parallel computing, strong ability to capture long-distance dependencies, and strong global feature learning ability. At the same time, research on fault diagnosis based on Transformer has also become a hot topic among scholars.

[0005] However, the current fault diagnosis methods based on the Transformer model still have the following deficiencies: (1) Lack of effective utilization of multi-scale convolutional channel information. (2) Simple splicing of feature information extracted by different methods, unable to effectively perform feature fusion. This will affect the accuracy, generalization, and robustness of the model's fault diagnosis. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes a mechanical fault diagnosis method based on the MK-ACFormer model, which is used to solve the problem that the existing Transformer model has poor diagnostic effect on mechanical faults in a noisy environment and its generalization ability needs to be improved.

[0007] The present invention provides a mechanical fault diagnosis method based on the MK-ACFormer model, including:

[0008] Obtain the vibration signal of the key components of the machine to be detected;

[0009] Input the vibration signal of the key components of the machine to be detected into the fault diagnosis model to obtain the fault diagnosis result; wherein, the fault diagnosis model is trained by a training set, the training set is the vibration signals of the key components of machines in different health conditions, and the fault diagnosis model is constructed based on a multi-scale kernel attention convolutional neural network fused with a Transformer module.

[0010] Optionally, obtaining the training set includes:

[0011] Obtain the vibration signals of different original health conditions;

[0012] Use a conversion function to perform standardization processing on the vibration signals of different original health conditions to obtain the training set.

[0013] Optionally, the fault diagnosis model includes: a feature screening module, a feature extraction module, and a feature fusion module;

[0014] The feature screening module is used to screen the features of the vibration signal;

[0015] The feature extraction module is used to perform multi-scale extraction on the screened signal features;

[0016] The feature fusion and classification module is used to perform feature fusion and classification on the multi-scale features.

[0017] Optionally, screening the features of the vibration signal includes:

[0018] Input the vibration signal of the key components of the machine to be detected into a wide convolutional layer, perform batch normalization, pass through a non-linear activation, perform max pooling processing to reduce the dimension, extract the features of the vibration signal, and obtain effective vibration signal feature information.

[0019] Optionally, the feature extraction module includes: an MK-ACNN sub-module and a Transformer Encoder sub-module;

[0020] The MK-ACNN sub-module includes: two improved CNN blocks, wherein the improved CNN block is stacked in sequence by a convolutional layer, a Relu layer, a maxpool_1d layer, a Relu layer, an SE layer, and a Relu layer.

[0021] Optionally, the SE layer includes: a compression sub-layer and an excitation sub-layer;

[0022] The compression sub-layer is used to perform global average pooling on the feature map obtained by convolution to obtain global information, wherein the global information is used for channel weight learning;

[0023] The excitation sub-layer is used to learn channel weights and enhance important channels.

[0024] Optionally, the method for obtaining the global information is:

[0025]

[0026] where z c is the c-th element of the global information z, U C is the c-th channel of the feature map U, H and W are the height and width of the input feature map, and F sq is a compression operation, and U c (i, j) is the value at the position (i, j) on the c-th channel of the input feature map;

[0027] The method for processing the channel dimension is:

[0028] s = F ex (z, W) = σ(W 2 δ(W 1 z))

[0029] where S is the channel attention weight, δ is the ReLU function, σ is the sigmoid function, W is the weight, and F ex is an excitation operation, W 1 is the weight matrix of the first fully connected layer, and W 2 is the weight matrix of the second fully connected layer.

[0030] Optionally, the feature fusion of multi-scale features includes:

[0031] Obtain the channels corresponding to the multi-scale features, adaptively adjust the channels using an effective channel attention block, and obtain the adjusted channels;

[0032] Calculate the adjusted channel weights and perform feature fusion.

[0033] Optionally, the method for adaptively adjusting the channels using an effective channel attention block is:

[0034]

[0035] where C is the number of channels of the input feature, b and γ are constants, k is the size of the 1D convolution kernel, and ψ() is a function of the number of channels C;

[0036] The method for calculating the adjusted channel weights is:

[0037] ω = σ(C1D k (y))

[0038] where y is the input feature, and C1Dk The 1D convolution of size k, σ is the Sigmoid activation function, and ω is the obtained channel weight.

[0039] Compared with the prior art, the present invention has the following advantages and technical effects:

[0040] The present invention designs a multi-scale convolutional kernel attention convolutional module to extract multi-local receptive field features of vibration signals, extract the correlation between channels, and reasonably allocate channel weights. Secondly, an effective channel attention module is used to fuse the features extracted by different-scale convolutions and the features extracted by combining convolutions with Transformers. By adaptively adjusting the feature channels and allocating different weights, information redundancy is reduced. Experiments show that compared with the recent fault diagnosis methods that combine Transformers with CNNs and improved CNN-based methods, the present invention can significantly improve the fault diagnosis ability of the model in a noisy environment and has strong robustness. On different data sets, it has a high fault diagnosis accuracy rate and strong generalization ability. Description of the Drawings

[0041] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0042] Figure 1 is the flowchart of the mechanical fault diagnosis method based on the MK-ACFormer model in the embodiment of the present invention;

[0043] Figure 2 is the schematic structural diagram of the MK-ACFormer in the embodiment of the present invention;

[0044] Figure 3 is the schematic structural diagram of the squeeze-and-excitation module in the embodiment of the present invention;

[0045] Figure 4 is the composition diagram of the Transformer encoder in the embodiment of the present invention;

[0046] Figure 5 is the schematic structural diagram of the channel attention network in the embodiment of the present invention;

[0047] Figure 6 is the schematic diagram of the improvement process of the CNN module in the embodiment of the present invention;

[0048] Figure 7 is the comparison diagram of the average training accuracy of each method provided in the embodiment of the present invention;

[0049] Figure 8This is a comparison chart of the average verification accuracy of each method provided by the embodiments of the present invention;

[0050] Figure 9 This is a confusion matrix diagram of the diagnosis accuracy of the method of the present invention under noise interference provided by the embodiments of the present invention;

[0051] Figure 10 This is a confusion matrix diagram of the diagnosis accuracy of LiConvFormer under noise interference provided by the embodiments of the present invention;

[0052] Figure 11 This is a distribution diagram of the accuracy of each method under different noise interferences provided by the embodiments of the present invention;

[0053] Figure 12 This is a t-SNE visualization clustering scatter plot provided by the embodiments of the present invention;

[0054] Figure 13 This is a comparison chart of the accuracy of various methods under different noise intensities provided by the embodiments of the present invention;

[0055] Figure 14 This is a comparison chart of the accuracy of different scales under different noise intensities provided by the embodiments of the present invention;

[0056] Figure 15 This is a comparison chart of the accuracy of different kernel sizes under different noise intensities provided by the embodiments of the present invention. Detailed implementation manners

[0057] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0058] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0059] The present invention provides a mechanical fault diagnosis method based on the MK-ACFormer model, as Figure 1 shown, which specifically includes:

[0060] Obtain the vibration signals of the key components of the mechanical equipment to be detected, where the key components of the mechanical equipment include bearings and gears, etc.

[0061] Input the vibration signal of the key components of the machine to be detected into the fault diagnosis model to obtain the fault diagnosis result. Among them, the fault diagnosis model is trained by a training set, which is the vibration signals of the key components of the machine in different health conditions. The fault diagnosis model is constructed based on the multi-scale kernel attention convolutional neural network fused with the Transformer module.

[0062] Specifically, S1: Obtain the vibration signal and perform standardization processing on the signal.

[0063] S2: Perform overlapping sampling segmentation on the standardized vibration data and divide it into a training set, a validation set, and a test set according to a ratio.

[0064] S3: Use the training set and the validation set to train the Transformer model based on the multi-scale kernel attention convolutional neural network to obtain the trained model. The Transformer model based on the multi-scale kernel attention convolutional neural network includes: a feature screening part with a wide kernel convolutional module for screening the vibration signal features; a multi-scale feature extraction layer part designed with MK-ACNN and Transformer Encoder modules for extracting multi-scale features; a feature fusion and classification part designed with effective channel attention for feature fusion.

[0065] S4: Input the test set into the trained model to obtain the fault diagnosis result.

[0066] Furthermore, obtaining the training set includes:

[0067] Obtain the vibration signals of different original health conditions.

[0068] Use the conversion function to perform standardization processing on the vibration signals of different original health conditions to obtain the training set.

[0069] Specifically, perform standardization processing on the collected fault vibration signals. The standardization processing processes multiple single-channel vibration signals separately to make the data conform to the standard normal distribution, that is, the mean is 0 and the standard deviation is 1. Its conversion function is:

[0070]

[0071] Among them, y is the standardized data, μ is the mean of all sample data, S is the standard deviation of all sample data, and X is the sample data.

[0072] Furthermore, as Figure 2 shown, the fault diagnosis model includes: a feature screening module, a feature extraction module, and a feature fusion module.

[0073] The feature screening module is used to screen the vibration signal features.

[0074] A feature extraction module for multi-scale extraction of the screened signal features;

[0075] A feature fusion and classification module for feature fusion and classification of the multi-scale features.

[0076] Specifically, for the feature screening part, the one-dimensional original bearing vibration signal is input into the wide convolution layer, and then batch normalization (BN) is used to accelerate the training speed. After passing through the ReLU layer for non-linear activation, the max pooling layer is used for dimensionality reduction and reducing the parameter calculation amount. The calculation formula is:

[0077] Y j = σ(P(σ(BN(k j ×x i +b j ))))

[0078] Where Y j is the j-th output, k j is the weight associated with the j-th output. b j is the bias term, BN is batch normalization, σ is the ReLU function, P is max pooling, and x i is the i-th element in the input vector;

[0079] Furthermore, the vibration signal feature screening includes:

[0080] Input the vibration signal of the mechanical key component to be detected into the wide convolution layer, perform batch normalization, pass through non-linear activation, perform max pooling processing, reduce the dimension, extract the vibration signal features, and obtain effective vibration signal feature information.

[0081] Specifically, the meaning of the effective vibration signal feature information is as follows: Feature information is the key index that can effectively characterize the equipment state extracted from the original vibration signal through mathematical, statistical, or machine learning methods.

[0082] The feature information is obtained by its own learning, and this information is composed of data.

[0083] Essentially, the feature information is digital data extracted from the original vibration signal, which can describe the characteristics of the signal and provide basic data for subsequent analysis and models.

[0084] Furthermore, the feature extraction module includes: a multi-scale kernel attention convolutional neural network (MK-ACNN) sub-module and a Transformer encoder sub-module;

[0085] The MK-ACNN sub-module includes: two improved CNN blocks, where the improved CNN block is composed of a convolutional layer, a Relu layer, a maxpool_1d layer, a Relu layer, an SE layer, and a Relu layer stacked in sequence.

[0086] Specifically, the multi-scale feature extraction part inputs the data output from the previous layer processing into the multi-scale extraction layer, which is composed of MK-ACNN and Transformer Encoder. MK-ACNN uses MK-ACNN blocks with five different convolutional kernels of 1×5, 1×7, 1×9, 1×11, and 1×13. The MK-ACNN block is constructed by stacking two improved CNN blocks, as Figure 6 shown. The improved CNN block is composed of a convolutional layer, Relu, maxpool_1d, Relu, SE module, and Relu stacked in sequence; finally, a Transformer Encoder is added after the MK-ACNN block with a 1×9 convolutional kernel (as Figure 4 shown).

[0087] Furthermore, as Figure 3 shown, the SE layer includes: a squeeze sub-layer and an excitation sub-layer;

[0088] The squeeze sub-layer is used to perform global average pooling on the feature map obtained by convolution to obtain global information;

[0089] The excitation sub-layer is used to process the channel dimension.

[0090] The channel attention SE module (SE module) consists of two main operations: squeeze and excitation;

[0091] Squeeze: For the feature map U obtained by convolution, global average pooling is performed to obtain global information z. Expressed as the following formula:

[0092]

[0093] where z c is the c-th element of the global information z, U C is the c-th channel of the feature map U, H and W are the height and width of the input feature map, F sq is the squeeze operation, and U c (i,j) is the value at the position (i,j) on the channel c of the input feature map.

[0094] Excitation: It consists of two fully connected layers, Relu and sigmoid. The first fully connected layer reduces the channel dimension, uses Relu for non-linear activation, then the second fully connected layer restores the channel dimension, and finally sigmoid processing is performed. Expressed as the following formula:

[0095] s = F ex (z, W) = σ(W 2 δ(W 1 z))

[0096] where s is the channel attention weight, δ is the ReLU function, σ is the sigmoid function, W is the weight, and F ex is the excitation operation, W 1 is the weight matrix of the first fully connected layer, and W 2 is the weight matrix of the second fully connected layer.

[0097] The original input feature map is scaled by the channel feature weights output after compression and excitation. It is expressed by the following formula:

[0098]

[0099] where F scale (u c , s c ) refers to the multiplication of channels between the scalar s c and the feature map u c .

[0100] In this step, as an additional embodiment:

[0101] When the Transformer Encoder uses single-head attention, the method of using multiple attention heads to calculate the results in parallel is:

[0102] MultiHead(Q, K, V) = Concat(head 1 ,..., head h )W O

[0103] MultiHead(.) represents the output function of the multi-head attention mechanism, and Concat(·) is the concatenation operation, where each attention head:

[0104]

[0105] where are the linear projection weights of the query, key, and value respectively, and W O is the weight matrix of the linear transformation with dimensions and

[0106]

[0107] Among them, Q, K, and V respectively represent the matrices of query, key, and value. T is the transpose operation, and d k is the dimension of the key vector, and softmax is the activation function;

[0108] Furthermore, the feature fusion of multi-scale features includes:

[0109] Obtain the channels corresponding to the multi-scale features, use the effective channel attention block to adaptively adjust the channels, and obtain the adjusted channels. Among them, the effective channel attention block is as Figure 5 shown;

[0110] Calculate the weights of the adjusted channels and perform feature fusion.

[0111] Specifically, in the feature fusion and classification part, during the feature fusion process, a 1D convolution with an adaptive size in the effective channel attention block is used to adaptively adjust the fused feature channels, and different weights are assigned to each channel. Finally, the GAP layer is used to reduce the feature parameters after adaptive fusion, preventing model overfitting to a certain extent, and then connecting to the fully connected layer and outputting to the Softmax layer for classification.

[0112] Furthermore, the method for adaptively adjusting the channels using the effective channel attention block is:

[0113]

[0114] Among them, C is the number of channels of the input features, b and γ are constants, k is the size of the 1D convolution kernel, and ψ() is a function of channel C;

[0115] The method for calculating the weights of the adjusted channels is:

[0116] ω = σ(C1D k (y))

[0117] Among them, y is the input feature, C1D k is the 1D convolution of size k, σ is the Sigmoid activation function, and ω is the obtained channel weight.

[0118] The following elaborates on this embodiment in conjunction with the accompanying drawings:

[0119] This embodiment selects the bearing dataset of the University of Ottawa under accelerated conditions, with 5 groups of bearing vibration signals in different health conditions. The partitions of the training set, validation set, and test set are 600, 300, and 300 respectively, the sliding window is 1024, and at the same time, to avoid data leakage, each sample does not overlap.

[0120] To verify the noise resistance performance of the model, in this embodiment, the following two different types of noise are added to the test set to simulate the differences between the training set data and the test set data in the actual industrial environment.

[0121] Add random Gaussian distribution noise: s i = s i + ε;

[0122] Random scaling: s i = σ + s i ;

[0123] Where s i represents the i-th sampling point of the sample; ε ~ N(0, λ) represents the added Gaussian noise; σ ~ N(1, λ) represents the scaling factor; λ is the variance, and the larger the λ value, the more significant the difference between the training set and the test set. To better simulate the uncertainty of the noise, the random addition probability of these two types of noise for each sample is 50%. In this embodiment, four novel CNN end-to-end fault diagnosis methods combined with Transformer, namely LiConvFormer, CLFormer, ConvformerNSE, and MCSwin-T, and four classic CNN methods, MobileNet, MobileNet-V2, ResNet18, and MK-ResCNN, are selected for comparison with the proposed method. To reduce randomness, each experiment is repeated five times. Each training is performed for 100 iterations with a batch size of 32, and the learning rate adopts an adaptive decay mode. The initial learning rate of each method is set using the grid search method according to the accuracy of the validation set.

[0124] Table 1

[0125]

[0126] As can be seen from Table 1, although the proposed method is slightly lower than some methods when λ = 0 on this dataset, it performs the best after adding noise. Under the interference of three different levels of noise, the average accuracy exceeds 90%. Moreover, when λ = 0.6, the accuracy is nearly 8% higher than that of the sub-optimal method LiConvFormer.

[0127] During the training process, as Figure 7 shown, the accuracies of various methods on the training set and the validation set are recorded, Figure 8 which is the average accuracy of various methods on the training set under 5 repeated experiments. Figure 9It is the average accuracy of various methods on the validation set under 5 repeated experiments. During the early iteration process, the accuracy fluctuation of each method on the training set is lower than that on the validation set. After 80 epochs, the accuracy of the training set and the validation set gradually stabilizes at a relatively high level. Compared with other methods, the accuracy of CLFormer on the validation set is the lowest, only reaching about 0.9.

[0128] From Figure 9 and Figure 10 's confusion matrix, it can be seen the classification situation of the proposed method and LiConvFormer on the test set when λ = 0.6. Compared with the method proposed in this embodiment, it can be seen that under the interference of noise, LiConvFormer misdiagnoses some fault categories 2 (inner raceway fault) as 3 (outer raceway fault) and 4 (rolling element fault), resulting in a poor overall diagnosis effect.

[0129] In addition, on the dataset from the Aero Engine Research Institute of Xi'an Jiaotong University, non-overlapping partitions are made for two vibration signals in different directions using a sliding window of size 1024. The training set, validation set, and test set are 500, 300, and 400 respectively. Simulate complex industrial environment test conditions, adding random Gaussian distribution noise and random scaling.

[0130] Table 2

[0131]

[0132] It can be seen from Table 2 that although the proposed method is slightly lower than other methods when λ = 0. However, after adding noise interference, the accuracy of the proposed method is significantly higher than other methods. When λ = 0.2, the average accuracy of 5 repeated experiments is improved by 2.41% compared with the sub-optimal LiConvFormer model. As the added noise increases, the performance gap between the proposed method and other methods becomes more significant. When λ = 0.6, the average accuracy is increased by about 12% compared with the LiConvFormer model.

[0133] Figure 11 It can be seen that under noise interference, the proposed method is not only the best in terms of accuracy, but also the difference between each test is small, the probability distribution is relatively concentrated, and the performance is relatively stable. Convformer-NSE shows a large range of fluctuations under noise interference. When λ = 0.2, the lowest value is the lowest among all methods, below 80%.

[0134] Figure 12Shows the visualization pictures of the features of different layer outputs after t-SNE processing. From the t-SNE visualization, it can be seen that after the dimensionality reduction of the original vibration signal, the features of various categories are mixed together. After the first-layer wide kernel convolution, the four types of bearing faults, namely rolling element fault, inner race fault, outer race fault, and mixed fault, initially show aggregation. After passing through the MK-ACNN blocks with different convolution kernels, the clustering effects are significantly different. Block3 basically completes the feature extraction of the four categories of bearing faults, and the clustering effect is the most obvious. Other categories also show a trend of beginning to separate. Block2 also completes the clustering of the features of the inner race fault of the bearing. After passing through the Transformer Encoder, various categories are basically clustered, but there are still a small number of discrete samples. After passing through the classification layer, all samples are clustered, and various categories are separated, with compactness within the category. Overall, the model first completes the classification of bearing faults, and then completes the classification of gear faults and normal states.

[0135] To verify the functions of each part module of the model, ablation experiments are carried out. The Transformer Encoder module of the proposed model is removed and named MK-ACFormer1. The SE module of the proposed model is removed and named MK-ACFormer2. The effective channel attention module of the proposed model is removed and named MK-ACFormer3. Figure 13 Shows that under the condition of no noise interference, the accuracies of each method are relatively close. Under the interference of high noise, the accuracy of the model without the Transformer Encoder module drops significantly, which has the greatest impact on the model. Followed by the model without the SE module, and finally the model without the effective channel attention module. This also shows that these three modules can capture the vibration signal features more accurately and have a certain anti-noise effect.

[0136] To further explore the performance of the model and verify the influence of different multi-scale convolution kernels and different wide kernel convolutions on the model. In this embodiment, the MK-ACNN Blocks with 3 scales and 7 scales are respectively selected for comparison, and the results are as Figure 14 shown. The 3-scale convolution kernel cannot fully extract the signal features and affects the model performance. While selecting the 7-scale convolution kernel for feature extraction extracts more redundant information, which cannot improve the model performance but instead increases the computational amount and slows down the calculation speed.

[0137] In this embodiment, wide convolution kernels of 64 and 256 are selected for comparison, and the results are as Figure 15As shown. The size of the convolutional kernel directly determines its receptive field for the input data, that is, the range of local information that the convolutional kernel can capture. Generally speaking, the larger the convolutional kernel, the larger the receptive field and the more feature information can be captured, but it is likely to cause problems such as large computational volume and overfitting. For a convolutional kernel of 64, the extracted feature information is not sufficient, and the accuracy drops in the presence of high noise. And there is no obvious improvement in the model with a convolutional kernel of 256, and the larger convolutional kernel leads to a large computational volume.

[0138] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A mechanical fault diagnosis method based on the MK-ACFormer model, characterized in that: include: Obtain vibration signals of key mechanical components to be tested; The vibration signal of the key mechanical component to be detected is input into the fault diagnosis model to obtain the fault diagnosis result; wherein, the fault diagnosis model is trained by a training set, and the training set is the vibration signals of the key mechanical components in different health conditions. The fault diagnosis model is constructed based on a multi-scale kernel attention convolutional neural network fused with a Transformer module.

2. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 1 is characterized in that: Acquiring the training set includes: Obtain original vibration signals of different health conditions; The original vibration signals of different health conditions are standardized by using a conversion function to obtain the training set.

3. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 1 is characterized in that: The fault diagnosis model includes: a feature screening module, a feature extraction module and a feature fusion module; The feature screening module is used to screen the vibration signal features; The feature extraction module is used to perform multi-scale extraction on the filtered signal features; The feature fusion and classification module is used to perform feature fusion and classification on multi-scale features.

4. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 3 is characterized in that: Vibration signal feature screening includes: The vibration signal of the key mechanical component to be detected is input into a wide convolutional layer for batch normalization. After nonlinear activation, maximum pooling processing is performed to reduce the dimension, extract the vibration signal features, and obtain effective vibration signal feature information.

5. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 3 is characterized in that: The feature extraction module includes: an MK-ACNN submodule and a Transformer Encoder submodule; The MK-ACNN submodule includes: two improved CNN blocks, wherein the improved CNN block is formed by stacking a convolutional layer, a Relu layer, a maxpool_1d layer, a Relu layer, a SE layer, and a Relu layer in sequence.

6. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 5 is characterized in that: The SE layer includes: a compression sublayer and an excitation sublayer; The compression sublayer is used to perform global average pooling on the feature map obtained by convolution to obtain global information, wherein the global information is used for channel weight learning; The excitation sublayer is used to learn channel weights and enhance important channels.

7. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 6 is characterized in that: The method for obtaining the global information is: Among them, z c is the cth element of the global information z, U C is the cth channel of feature map U, H and W are the height and width of the input feature map, F sq For compression operation, U c (i, j) is the value of the input feature map at position (i, j) on channel c; The method for handling the channel dimension is: s=F ex (z,W)=σ(W2δ(W1z)) Among them, s is the channel attention weight, δ is the ReLU function, σ is the sigmoid function, W is the weight, F ex For the excitation operation, W1 is the weight matrix of the first fully connected layer, and W2 is the weight matrix of the second fully connected layer.

8. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 3 is characterized in that: Feature fusion of multi-scale features includes: Obtain a channel corresponding to the multi-scale feature, and adaptively adjust the channel using an effective channel attention block to obtain an adjusted channel; Calculate the adjusted channel weights and perform feature fusion.

9. The mechanical fault diagnosis method based on the MK-ACFormer model according to claim 8, characterized in that: The method of adaptively adjusting the channel using the effective channel attention block is: Where C is the number of channels of the input feature, b and γ are constants, k is the size of the 1D convolution kernel, and ψ() is a function of the number of channels C; The method to calculate the adjusted channel weights is: ω=σ(C1Dk(y)) Among them, y is the input feature, C1D k is a 1D convolution of size k, σ is the Sigmoid activation function, and ω is the obtained channel weight.

Citation Information

Patent Citations

  • Multi-scale feature fusion gearbox fault diagnosis method based on self-attention mechanism

    CN116010900A

  • Circuit breaker fault diagnosis method and system based on multi-signal fusion characteristics

    CN117192346A

  • Fault diagnosis method based on combination of VMD decomposition and multi-scale convolutional network

    CN118194160A

  • Bearing fault diagnosis method under strong noise background based on wavelet domain self-attention mechanism

    CN118443309A

  • Cross-domain mechanical fault diagnosis method based on multi-channel feature fusion of CBAM and use thereof

    US20240142342A1