Bearing fault diagnosis method and system, storage medium and computer
By extracting the time-frequency characteristics of bearing vibration signals and combining them with deep learning technology, the problem that traditional methods are difficult to extract early signal characteristics of bearing failures is solved, achieving higher diagnostic accuracy and robustness.
Patent Information
- Application Number
- CN202510052298.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional bearing fault diagnosis methods are difficult to effectively extract the early signal characteristics of axle box bearing failures, and when dealing with complex failure modes, it is impossible to fully capture the complex dependencies between the features, resulting in unsatisfactory diagnostic performance.
A bearing fault diagnosis method is proposed. By obtaining bearing vibration signals, the time-frequency features are extracted using variational modal decomposition and fast Fourier transform, the frequency domain features are extracted in combination with adaptive multi-scale dynamic convolutional neural network and ECA attention mechanism, and the time-domain features are captured through the Transformer model, and finally the feature fusion and classification are performed through the bidirectional cross attention mechanism.
It effectively suppresses noise interference, highlights the key characteristics of fault signals, improves the effect of signal separation and feature extraction, enhances the selectivity of feature channels, realizes the deep fusion of time-frequency characteristics, and significantly improves the accuracy and robustness of bearing fault diagnosis.
Smart Images

Figure CN119989142A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis, and in particular to a bearing fault diagnosis method, system, storage medium and computer. Background Art
[0002] Axlebox bearings are one of the most important rotating components in rail vehicles. The health of their operating status is closely related to the performance and service life of the equipment. Bearing failures not only lead to reduced equipment performance, but may also cause safety accidents. Therefore, bearing fault diagnosis has become an important research direction in the industry.
[0003] Traditional bearing fault diagnosis methods usually rely on manual feature extraction of time domain or frequency domain analysis, such as fast Fourier transform (FFT), empirical mode decomposition (EMD), wavelet transform (WT) and other signal processing methods. Although these methods can extract some effective features, the feature extraction process relies on expert experience and the feature expression ability is limited. In recent years, deep learning technology has been widely used in the field of fault diagnosis due to its powerful automatic feature extraction ability. Convolutional neural network (CNN) performs well in extracting spatial features, while Transformer has advantages in capturing temporal dependencies. However, in the early stage of axle box bearing failure, due to the weak fault signal and often accompanied by a large amount of noise interference, traditional time domain or frequency domain analysis methods are difficult to effectively extract fault features; when dealing with complex bearing fault modes, there are certain limitations in performance, and often cannot fully capture the complex dependencies between features, resulting in unsatisfactory diagnostic performance; affecting the results of fault diagnosis. Summary of the invention
[0004] Based on this, the purpose of the present invention is to provide a bearing fault diagnosis method, system, storage medium and computer to solve the technical problems existing in the prior art.
[0005] The present invention provides a bearing fault diagnosis method, comprising:
[0006] Acquire a bearing vibration signal, and extract time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform;
[0007] The time-frequency features are used as input of an adaptive multi-scale dynamic convolutional neural network to perform a convolution operation, and an ECA attention mechanism is used to process the convolution output results to extract frequency domain features from the time-frequency features;
[0008] Capturing the global temporal dependency in the time-frequency features based on the Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features;
[0009] The extracted frequency domain features are fused with the time domain features through a bidirectional cross-attention mechanism, the fused features are passed to the fully connected layer for classification, the fault category of the bearing is output, and the classification result is converted into a probability distribution, and the category with the highest probability is selected as the diagnosis result.
[0010] Preferably, the step of extracting the time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform comprises:
[0011] Decomposing the bearing vibration signal into a plurality of eigenmode functions by means of the variational mode decomposition;
[0012] Converting the bearing vibration signal into a frequency spectrum amplitude according to a fast Fourier transform;
[0013] A plurality of the intrinsic mode functions and the spectrum amplitudes are stacked vertically to obtain a time-frequency feature matrix.
[0014] Preferably, the bearing vibration signal decomposition expression is:
[0015]
[0016] Where x(n) is the bearing vibration signal, u m (t) is the mth eigenmode function, M is the number of modes;
[0017] The expression of the time-frequency feature matrix is:
[0018] X(t,f)=[u1(t),u2(t),…,u M (t),|X(f)|]
[0019] Where X(t,f) is the time-frequency feature matrix, u1(t),u2(t),…,u M (t) is the M intrinsic mode functions, |X(f)| is the spectrum amplitude of the fast Fourier transform.
[0020] Preferably, the step of performing a convolution operation on the time-frequency features as input of an adaptive multi-scale dynamic convolutional neural network and processing the convolution output results using an ECA attention mechanism to extract frequency domain features from the time-frequency features includes:
[0021] Obtain a multi-scale convolution kernel, perform a first convolution process on the convolution kernel and the time-frequency feature matrix, and perform a second convolution process on the time-frequency feature matrix based on an adaptive weight generation network to obtain a weight matrix;
[0022] Obtaining a convolution output result of an adaptive multi-scale dynamic convolutional neural network according to the first convolution output and the weight matrix, and obtaining a multi-scale fused convolution output;
[0023] The multi-scale fused convolution output is used as the input of the ECA attention mechanism and the ECA operation is performed to obtain the channel weighting coefficient of the ECA attention mechanism input, and the output result of the ECA attention mechanism is calculated according to the channel weighting coefficient and the multi-scale fused convolution output;
[0024] The output results of the ECA attention mechanism are batch normalized and nonlinearly transformed, and the results after the transformation are sampled through a maximum pooling layer to obtain the frequency domain features in the time-frequency features.
[0025] Preferably, the expression of the first convolution process is:
[0026] Conv3(x)=Conv1d(X(t,f),g3)
[0027] Conv5(x)=Conv1d(X(t,f),g5)
[0028] Conv7(x)=Conv1d(X(t,f),g7)
[0029] Where g3, g5, and g7 are the sizes of the one-dimensional convolution kernels, Conv3(x), Conv5(x), and Conv7(x) are the convolution outputs corresponding to the sizes of the one-dimensional convolution kernels of 3, 5, and 7, respectively. Conv1d represents one-dimensional convolution, and X(t, f) is the time-frequency feature matrix.
[0030] The expression of the second convolution process is:
[0031] W3(t,f)=Softmax(Conv3(x),g3)
[0032] W5(t,f)=Softmax(Conv5(x),g5)
[0033] W7(t,f)=Softmax(Conv7(x),g7)
[0034] Wherein, W3(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 3, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 3; W5(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 5, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 5; W7(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 7, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 7, and Softmax is the normalized exponential activation function;
[0035] The expression of the multi-scale fused convolution output is:
[0036] y3(x)=W3(t,f)·Conv3(x)
[0037] y5(x)=W5(t,f)·Conv5(x)
[0038] y7(x)=W7(t,f)·Conv7(x)
[0039] Where y3(x), y5(x), and y7(x) are the multi-scale fused convolution outputs corresponding to the one-dimensional convolution kernel sizes of 3, 5, and 7, respectively;
[0040] The expression of the channel weighting coefficient is:
[0041] p3=Sigmoid(Conv1d(GlobalAvgPool(y3(x))))
[0042] p5=Sigmoid(Conv1d(GlobalAvgPool(y5(x))))
[0043] p7=Sigmoid(Conv1d(GlobalAvgPool(y7(x))))
[0044] In the formula, p3, p5, p5 are the channel weight coefficients corresponding to the one-dimensional convolution kernel size of 3, 5, and 7 respectively, Sigmoid is the S-type activation function, Conv1d represents one-dimensional convolution, and GlobalAvgPool is the pooling operation for each channel of the ECA attention mechanism input;
[0045] The output result of the ECA attention mechanism is expressed as:
[0046] y(x)′=y3(x)′+y5(x)′+y7(x)′
[0047] Where y′ is the output result of the ECA attention mechanism;
[0048] y3(x)′=y3(x)·p3
[0049] y5(x)′=y5(x)·p5
[0050] y7(x)′=y7(x)·p7
[0051] Where y3(x)′, y5(x)′, and y7(x)′ are the outputs of the ECA attention mechanism when the size of the one-dimensional convolution kernel is 3, 5, and 7, respectively;
[0052] The expressions of the batch normalization and nonlinear transformation operations are:
[0053] yfinal =ReLU(BatchNorm(y(x)′))
[0054] In the formula, y final It is the result of batch normalization and nonlinear transformation operations, ReLU is the rectified linear activation function, and BatchNorm is the batch normalization operation;
[0055] The expression of the maximum pooling layer sampling process is:
[0056] F(t,f)=MaxPool(y final )
[0057] Where MaxPool is the maximum pooling layer operation.
[0058] Preferably, the step of capturing the global temporal dependency in the time-frequency features based on the Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features includes:
[0059] Map the time-frequency feature matrix based on the multi-head attention mechanism in the Transformer model to generate a query vector, a key vector, and a value vector corresponding to each head;
[0060] Calculate the corresponding attention weight according to the query vector and the key vector corresponding to each head, and obtain the output corresponding to each head according to the attention weight and the value vector;
[0061] The output of each head is concatenated, and the concatenated output is transformed by a linear transformation matrix to obtain the time domain features of the long-distance dependence in the time-frequency features.
[0062] Preferably, the step of fusing the extracted frequency domain features with the time domain features through a bidirectional cross attention mechanism comprises:
[0063] Calculating first cross-attention weights of the time domain features and the frequency domain features, and performing weighted summation of the first cross-attention weights and the frequency domain values in the frequency domain features to obtain a first attention output of the time domain to the frequency domain;
[0064] Calculating the second cross-attention weights of the frequency domain features and the time domain features, and performing weighted summation on the second cross-attention weights and the time domain values in the time domain features to obtain a second attention output of the frequency domain to the time domain;
[0065] The first attention output and the second attention output are fused to obtain comprehensive features of frequency domain features and time domain features.
[0066] The present invention also provides a bearing fault diagnosis system, comprising:
[0067] An extraction module, used to obtain a bearing vibration signal, and extract time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform;
[0068] A convolution module, used to perform a convolution operation on the time-frequency features as input of an adaptive multi-scale dynamic convolutional neural network, and process the convolution output results using an ECA attention mechanism to extract frequency domain features from the time-frequency features;
[0069] A capture module, used to capture the global temporal dependency in the time-frequency features based on a Transformer model, so as to extract the time domain features of the long-distance dependency in the time-frequency features;
[0070] The diagnosis module is used to fuse the extracted frequency domain features with the time domain features through a bidirectional cross-attention mechanism, pass the fused features into the fully connected layer for classification, output the fault category of the bearing, and convert the classification results into probability distribution, and select the category with the highest probability as the diagnosis result.
[0071] The present invention also provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned bearing fault diagnosis method is implemented.
[0072] The present invention also proposes a computer, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the above-mentioned bearing fault diagnosis method is implemented when the processor executes the computer program.
[0073] Compared with the prior art, the present invention has the following beneficial effects: the bearing fault diagnosis method provided by the present application obtains the bearing vibration signal, and extracts the time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform, which can effectively suppress noise interference, highlight the key features of the fault signal, and significantly improve the effect of signal separation and feature extraction; the frequency domain features in the signal are extracted by an adaptive multi-scale dynamic convolution module, and combined with the capture of time domain features by Transformer, the effective fusion of time-frequency features is achieved, and the accuracy and robustness of bearing fault diagnosis are improved; the ECA attention mechanism effectively enhances the selectivity of feature channels by introducing dynamic weight distribution between channels, so that the model can adaptively focus on the most representative feature channels when processing different types of faults, thereby improving the accuracy of feature extraction; the bidirectional cross attention mechanism realizes the deep fusion of time domain and frequency domain information through bidirectional information interaction between time-frequency features, and can more comprehensively capture the characteristics of complex signals, which is particularly suitable for bearing signal analysis with multiple fault modes, greatly improving the fault diagnosis capability of the model, and suitable for large-scale promotion.
[0074] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 This is a flow chart of a bearing fault diagnosis method in Embodiment 1 of the present invention;
[0076] Figure 2 This is a structural block diagram of a computer in Embodiment 4 of the present invention.
[0077] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0078] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0080] Embodiment 1
[0081] See also Figure 1 , which shows a bearing fault diagnosis method in a first embodiment of the present invention. The bearing fault diagnosis method specifically includes steps S10 to S40:
[0082] S10, acquiring a bearing vibration signal, and extracting time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform;
[0083] Optionally, the step of extracting the time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform includes:
[0084] Decomposing the bearing vibration signal into a plurality of eigenmode functions by means of the variational mode decomposition;
[0085] Converting the bearing vibration signal into a frequency spectrum amplitude according to a fast Fourier transform;
[0086] A plurality of the intrinsic mode functions and the spectrum amplitudes are stacked vertically to obtain a time-frequency feature matrix.
[0087] The expression for the decomposition of the bearing vibration signal is:
[0088]
[0089] Where x(n) is the bearing vibration signal, u m (t) is the mth eigenmode function, M is the number of modes;
[0090] The expression of the time-frequency feature matrix is:
[0091] X(t,f)=[u1(t),u2(t),…,u M (t),|X(f)|]
[0092] Where X(t,f) is the time-frequency feature matrix, u1(t),u2(t),…,u M (t) is the M intrinsic mode functions, |X(f)| is the spectrum amplitude of the fast Fourier transform.
[0093] In this embodiment, the time-frequency feature matrix is composed of the IMF components obtained by VMD decomposition and the FFT spectrum. The matrix contains information in the time and frequency dimensions, and is used to represent the time-frequency characteristics of the bearing vibration signal; the bearing vibration signal is decomposed into several intrinsic mode functions (IMFs) by variational mode decomposition (VMD), representing the time domain characteristics of the collected bearing vibration signal x(n); the bearing vibration signal is converted into a spectrum amplitude |X(f)| by fast Fourier transform (FFT), and finally, the time domain IMFs obtained by VMD decomposition and the spectrum amplitude after FFT transformation are vertically stacked to construct a time-frequency feature matrix. The goal of time-frequency feature fusion is to combine time domain and frequency domain information, so as to more comprehensively describe the characteristics of the fault signal, enrich the characteristics of the signal, and improve the accuracy of fault identification.
[0094] S20, performing a convolution operation on the time-frequency feature as an input of an adaptive multi-scale dynamic convolutional neural network, and processing the convolution output result using an ECA attention mechanism to extract frequency domain features from the time-frequency feature;
[0095] Optionally, the step of performing a convolution operation on the time-frequency feature as an input of an adaptive multi-scale dynamic convolutional neural network, and processing the convolution output result using an ECA attention mechanism to extract frequency domain features from the time-frequency feature includes:
[0096] Obtain a multi-scale convolution kernel, perform a first convolution process on the convolution kernel and the time-frequency feature matrix, and perform a second convolution process on the time-frequency feature matrix based on an adaptive weight generation network to obtain a weight matrix;
[0097] Obtaining a convolution output result of an adaptive multi-scale dynamic convolutional neural network according to the first convolution output and the weight matrix to obtain a multi-scale fused convolution output;
[0098] The multi-scale fused convolution output is used as the input of the ECA attention mechanism and the ECA operation is performed to obtain the channel weighting coefficient of the ECA attention mechanism input, and the output result of the ECA attention mechanism is calculated according to the channel weighting coefficient and the multi-scale fused convolution output;
[0099] The output results of the ECA attention mechanism are batch normalized and nonlinearly transformed, and the results after the transformation are sampled through a maximum pooling layer to obtain the frequency domain features in the time-frequency features.
[0100] The expression of the first convolution process is:
[0101] Conv3(x)=Conv1d(X(t,f),g3)
[0102] Conv5(x)=Conv1d(X(t,f),g5)
[0103] Conv7(x)=Conv1d(X(t,f),g7)
[0104] Where g3, g5, and g7 are the sizes of the one-dimensional convolution kernels, Conv3(x), Conv5(x), and Conv7(x) are the convolution outputs corresponding to the sizes of the one-dimensional convolution kernels of 3, 5, and 7, respectively. Conv1d represents one-dimensional convolution, and X(t, f) is the time-frequency feature matrix.
[0105] The expression of the second convolution process is:
[0106] W3(t,f)=Softmax(Conv3(x),g3)
[0107] W5(t,f)=Softmax(Conv5(x),g5)
[0108] W7(t,f)=Softmax(Conv7(x),g7)
[0109] Wherein, W3(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 3, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 3; W5(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 5, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 5; W7(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 7, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 7, and Softmax is the normalized exponential activation function;
[0110] The expression of the multi-scale fused convolution output is:
[0111] y3(x)=W3(t,f)·Conv3(x)
[0112] y5(x)=W5(t,f)·Conv5(x)
[0113] y7(x)=W7(t,f)·Conv7(x)
[0114] Where y3(x), y5(x), and y7(x) are the multi-scale fused convolution outputs corresponding to the one-dimensional convolution kernel sizes of 3, 5, and 7, respectively;
[0115] The expression of the channel weighting coefficient is:
[0116] p3=Sgimoid(Conv1d(GlobalAvgPool(y3(x))))
[0117] p5=Sigmoid(Conv1d(GlobalAvgPool(y5(x))))
[0118] p7=Sigmoid(Conv1d(GlobalAvgPool(y7(x))))
[0119] In the formula, po3, p5, p5 are the channel weight coefficients corresponding to the one-dimensional convolution kernel size of 3, 5, and 7 respectively, Sigmoid is the S-type activation function, Conv1d represents one-dimensional convolution, and GlobalAvgPool is the pooling operation for each channel of the ECA attention mechanism input;
[0120] The output result of the ECA attention mechanism is expressed as:
[0121] y(x)′=y3(x)′+y5(x)′+y7(x)′
[0122] Where y′ is the output result of the ECA attention mechanism;
[0123] y3(x)′=y3(x)·p3
[0124] y5(x)′=y5(x)·p5
[0125] y7(x)′=y7(x)·p7
[0126] Where y3(x)′, y5(x)′, and y7(x)′ are the outputs of the ECA attention mechanism when the size of the one-dimensional convolution kernel is 3, 5, and 7, respectively;
[0127] The expressions of the batch normalization and nonlinear transformation operations are:
[0128] y final =ReLU(BatchNorm(y(x)′))
[0129] In the formula, y final It is the result of batch normalization and nonlinear transformation operations, ReLU is the rectified linear activation function, and BatchNorm is the batch normalization operation;
[0130] The expression of the maximum pooling layer sampling process is:
[0131] F(t,f)=MaxPool(y final )
[0132] Where MaxPool is the maximum pooling layer operation.
[0133] In this embodiment, in order to extract frequency domain features of different scales from the time-frequency feature matrix, a multi-scale one-dimensional convolution kernel is used. The size of the convolution kernel can be 3, 5 and 7, respectively represented as g3, g5, and g7; they are used to capture features of different scales. Each convolution kernel can capture signal patterns in different frequency ranges, thereby improving the network's sensitivity to frequency domain features. In order to dynamically adjust the impact of different convolution kernels, an adaptive weight generation network is used. The network obtains a weight matrix by performing 1×1 convolution and Softmax activation function calculation on the input time-frequency feature matrix, which is used to weight the output of different convolution kernels; by weighting the outputs of the three convolutions and then applying the ECA attention mechanism, the effect of feature representation can be further improved. ECA performs more fine-grained channel-level attention weighting on the weighted features, which can enable the model to pay more attention to important feature channels, especially in complex tasks. This strategy can enhance the expressive power of the model and improve performance. Perform ECA operation on the weighted convolution outputs y3(x), y5(x), y7(x), that is, perform global average pooling, one-dimensional convolution, and channel attention calculation; feature y final Downsampling is performed through the maximum pooling layer to reduce the dimension of the features and retain the most significant time-frequency feature information: the final F(t,f) is the frequency domain feature extracted from the collected bearing vibration signal x(n) through the adaptive multi-scale dynamic convolutional neural network combined with the channel attention mechanism ECA, which provides high-quality frequency domain feature representation for the subsequent process, thereby improving the accuracy and robustness of bearing fault diagnosis.
[0134] Furthermore, ECA (Efficient Channel Attention) aims to enhance the expressiveness of important channel features by assigning weights to features of different channels. The ECA channel attention mechanism focuses on the weighted coefficients of each channel, rather than the weighted coefficients of the convolution kernel. Each convolution output generates a weighted coefficient using a small convolution kernel after calculating global pooling. This weighted coefficient is used to weight the channel of each convolution output. Ultimately, the weighted outputs of all convolutions are summed to obtain the final result. The purpose of the ECA attention mechanism is to enhance the model's attention to specific channels in a simple and efficient way, avoiding complex channel-to-channel interactions, and significantly improving the expressiveness of channel features without introducing excessive computational costs.
[0135] S30, capturing the global temporal dependency in the time-frequency features based on a Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features;
[0136] Optionally, the step of capturing the global temporal dependency in the time-frequency features based on the Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features includes:
[0137] Map the time-frequency feature matrix based on the multi-head attention mechanism in the Transformer model to generate a query vector, a key vector, and a value vector corresponding to each head;
[0138] Calculate the corresponding attention weight according to the query vector and the key vector corresponding to each head, and obtain the output corresponding to each head according to the attention weight and the value vector;
[0139] The output of each head is concatenated, and the concatenated output is transformed by a linear transformation matrix to obtain the time domain features of the long-distance dependence in the time-frequency features.
[0140] In this embodiment, the Transformer model uses four parallel attention heads to extract time domain features; specifically, each head has an independent weight matrix The input time-frequency feature matrix X(t,f) is mapped through linear transformation to generate multiple different query vectors Q r , key K r Vector and value vector V r ;
[0141]
[0142] Where i∈{1,2,3,4} represents different heads.
[0143] The weight matrix is optimized during training using the back-propagation algorithm. The initial values are random and initialized using a uniform distribution.
[0144] The expression of attention weight is:
[0145]
[0146] Where i∈{1,2,3,4} represents different heads. is the attention score of different heads, Softmax is the normalized exponential activation function, Softmax will query Q r and key K r The similarity of is normalized so that the weight of each position can reflect the importance of the position to the current query. is the scaling factor, where D is the query Q of each head r and key K r The dimension D is equal to The denominator 4 is the number of heads, and the numerator j is the feature dimension of the input time-frequency feature matrix X(t,f), which ensures that the size of the attention weight remains in an appropriate range.
[0147] The output expression corresponding to each head is:
[0148]
[0149] This operation extracts the relationship between each time step in the input time-frequency feature matrix X(t,f), reflecting the time domain characteristics. In the multi-head attention mechanism, the outputs of all heads will be spliced together to obtain a new matrix Z(t,f);
[0150]
[0151] The transformed time domain feature expression is:
[0152] T(t,f)=Z(t,f)·W o
[0153] Where W o A trainable parameter, called the linear transformation matrix, can help the model adjust the output of each attention head to improve task performance. The final Transformer output T(t,f) has the same shape as the input time-frequency feature matrix X(t,f), but after being processed by the multi-head attention mechanism, T(t,f) contains rich time-domain dependency information.
[0154] S40, the extracted frequency domain features are fused with the time domain features through a bidirectional cross-attention mechanism, the fused features are passed to the fully connected layer for classification, the fault category of the bearing is output, and the classification result is converted into a probability distribution, and the category with the highest probability is selected as the diagnosis result.
[0155] Optionally, the step of fusing the extracted frequency domain features with the time domain features through a bidirectional cross attention mechanism includes:
[0156] Calculating first cross-attention weights of the time domain features and the frequency domain features, and performing weighted summation of the first cross-attention weights and the frequency domain values in the frequency domain features to obtain a first attention output of the time domain to the frequency domain;
[0157] Calculating the second cross-attention weights of the frequency domain features and the time domain features, and performing weighted summation on the second cross-attention weights and the time domain values in the time domain features to obtain a second attention output of the frequency domain to the time domain;
[0158] The first attention output and the second attention output are fused to obtain comprehensive features of frequency domain features and time domain features.
[0159] After the previous frequency domain and time domain feature extraction, the frequency domain feature F(t,f) and the time domain feature T(t,f) are obtained respectively, and the two types of information are fused through the bidirectional cross attention mechanism;
[0160] In the forward cross-attention mechanism (from time-domain query to frequency-domain key), Q t The query generated for the time domain feature, K f The key generated for the frequency domain feature, V f The value generated for the frequency domain feature. t With the frequency domain key K f The dot product of gets the expression of the first attention weight as:
[0161]
[0162] The expression of the first attention output is:
[0163] F1=A qkt ·V f
[0164] In the reverse criss-cross attention mechanism (from frequency domain query to time domain key), Q f For the query generated by frequency domain features, K t The key generated for the time domain feature, V t The value generated for the time domain feature. f With the time domain key K t The dot product of gets the expression of the second attention weight as:
[0165]
[0166] The expression of the second attention output is:
[0167] F2=A qkf ·V t
[0168] Finally, the forward cross-attention output F1 and the reverse cross-attention output F2 are weighted averaged by a learnable weighting coefficient to obtain the final feature output:
[0169] F=βF1+(1-β)F2
[0170] In the formula, β is a learnable parameter. Learnable parameters refer to parameters that are automatically adjusted by gradient descent during model training. The initial values of these parameters are generally random. During the training process, the model gradually updates these parameters based on the feedback of the data, making the model's prediction results more accurate. β controls the contribution ratio of forward and reverse cross attention.
[0171] The goal of this bidirectional cross-attention mechanism is to maximize the complementary information of time domain and frequency domain features by modeling the bidirectional dependency between time and frequency information, thereby improving the accuracy and robustness of fault pattern recognition. Through this bidirectional cross-attention mechanism, the model can capture the interaction and deep dependency between time domain and frequency domain features, and achieve accurate classification and fault detection of complex signals.
[0172] Furthermore, the output feature F is adaptively averaged and all positions of each channel are averaged to obtain the pooled feature F p , and then the pooled feature tensor F p Flatten to F l Then the output F is obtained through the fully connected layer fc . Finally, F fc Perform Softmax operation to obtain the category probability of bearing fault.
[0173] The expression of the bearing fault category probability is:
[0174]
[0175] Where P b,c is the probability that bearing sample b belongs to bearing fault category c, (F fc,b,c ) is the score of the b-th bearing sample on the bearing fault category c after the output of the fully connected layer. The indexation is to increase the influence of larger scores and reduce the influence of smaller scores. c′ exp(F fc,b,c′) is the sum of all bearing fault categories c′, and the sum of the indexed scores of all bearing fault categories is calculated. This operation is to ensure that the sum of the output probability distribution is 1.
[0176] Y = argmax(P b )
[0177] In the formula, Y is the category with the highest probability of the final prediction, that is, for the b-th bearing sample, find P b,c The maximum bearing failure category is c.
[0178] In summary, the bearing fault diagnosis method provided in the present application obtains the bearing vibration signal, and extracts the time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform, which can effectively suppress noise interference, highlight the key features of the fault signal, and significantly improve the effect of signal separation and feature extraction; the frequency domain features in the signal are extracted by an adaptive multi-scale dynamic convolution module, and combined with the capture of time domain features by Transformer, the effective fusion of time-frequency features is realized, and the accuracy and robustness of bearing fault diagnosis are improved; the ECA attention mechanism effectively enhances the selectivity of feature channels by introducing dynamic weight distribution between channels, so that the model can adaptively focus on the most representative feature channels when dealing with different types of faults, thereby improving the accuracy of feature extraction; the bidirectional cross attention mechanism realizes the deep fusion of time domain and frequency domain information through bidirectional information interaction between time-frequency features, and can more comprehensively capture the characteristics of complex signals, which is particularly suitable for bearing signal analysis with multiple fault modes, greatly improving the fault diagnosis capability of the model.
[0179] Embodiment 2
[0180] This embodiment provides a bearing fault diagnosis system, including:
[0181] An extraction module, used to obtain a bearing vibration signal, and extract time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform;
[0182] A convolution module, used to perform a convolution operation on the time-frequency features as input of an adaptive multi-scale dynamic convolutional neural network, and process the convolution output results using an ECA attention mechanism to extract frequency domain features from the time-frequency features;
[0183] A capture module, used to capture the global temporal dependency in the time-frequency features based on a Transformer model, so as to extract the time domain features of the long-distance dependency in the time-frequency features;
[0184] The diagnosis module is used to fuse the extracted frequency domain features with the time domain features through a bidirectional cross-attention mechanism, pass the fused features into the fully connected layer for classification, output the fault category of the bearing, and convert the classification results into probability distribution, and select the category with the highest probability as the diagnosis result.
[0185] Wherein, the step of extracting the time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform comprises:
[0186] Decomposing the bearing vibration signal into a plurality of eigenmode functions by means of the variational mode decomposition;
[0187] Converting the bearing vibration signal into a frequency spectrum amplitude according to a fast Fourier transform;
[0188] A plurality of the intrinsic mode functions and the spectrum amplitudes are stacked vertically to obtain a time-frequency feature matrix.
[0189] Preferably, the bearing vibration signal decomposition expression is:
[0190]
[0191] Where x(n) is the bearing vibration signal, u m (t) is the mth eigenmode function, M is the number of modes;
[0192] The expression of the time-frequency feature matrix is:
[0193] X(t,f)=[u1(t),u2(t),…,u M (t),|X(f)|]
[0194] Where X(t,f) is the time-frequency feature matrix, u1(t),u2(t),…,u M (t) is the M intrinsic mode functions, |X(f)| is the spectrum amplitude of the fast Fourier transform.
[0195] Preferably, the step of performing a convolution operation on the time-frequency features as input of an adaptive multi-scale dynamic convolutional neural network and processing the convolution output results using an ECA attention mechanism to extract frequency domain features from the time-frequency features includes:
[0196] Obtain a multi-scale convolution kernel, perform a first convolution process on the convolution kernel and the time-frequency feature matrix, and perform a second convolution process on the time-frequency feature matrix based on an adaptive weight generation network to obtain a weight matrix;
[0197] Obtaining a convolution output result of an adaptive multi-scale dynamic convolutional neural network according to the first convolution output and the weight matrix to obtain a multi-scale fused convolution output;
[0198] The multi-scale fused convolution output is used as the input of the ECA attention mechanism and the ECA operation is performed to obtain the channel weighting coefficient of the ECA attention mechanism input, and the output result of the ECA attention mechanism is calculated according to the channel weighting coefficient and the multi-scale fused convolution output;
[0199] The output results of the ECA attention mechanism are batch normalized and nonlinearly transformed, and the results after the transformation are sampled through a maximum pooling layer to obtain the frequency domain features in the time-frequency features.
[0200] Preferably, the expression of the first convolution process is:
[0201] Conv3(x)=Conv1d(X(t,f),g3)
[0202] Conv5(x)=Conv1d(X(t,f),g5)
[0203] Conv7(x)=Conv1d(X(t,f),g7)
[0204] Where g3, g5, and g7 are the sizes of the one-dimensional convolution kernels, Conv3(x), Conv5(x), and Conv7(x) are the convolution outputs corresponding to the sizes of the one-dimensional convolution kernels of 3, 5, and 7, respectively. Conv1d represents one-dimensional convolution, and X(t, f) is the time-frequency feature matrix.
[0205] The expression of the second convolution process is:
[0206] W3(t,f)=Softmax(Conv3(x),g1)
[0207] W5(t,f)=Softmax(Conv5(x),g5)
[0208] W7(t,f)=Softmax(Conv7(x),g7)
[0209] Wherein, w3(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 3, indicating the importance of the convolution with a one-dimensional convolution kernel size of 3; W5(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 5, indicating the importance of the convolution with a one-dimensional convolution kernel size of 5; W7(t,f) is the adaptive weight generated by the input X(t,f) when the one-dimensional convolution kernel is 7, indicating the importance of the convolution with a one-dimensional convolution kernel size of 7, and Softmax is the normalized exponential activation function;
[0210] The expression of the multi-scale fused convolution output is:
[0211] y3(x)=W3(t,f)·Conv3(x)
[0212] y5(x)=W5(t,f)·Conv5(x)
[0213] y7(x)=W7(t,f)·Conv7(x)
[0214] Where y3(x), y5(x), and y7(x) are the multi-scale fused convolution outputs corresponding to the one-dimensional convolution kernel sizes of 3, 5, and 7, respectively;
[0215] The expression of the channel weighting coefficient is:
[0216] p3=Sigmoid(Conv1d(GlobalAvgPool(y3(x))))
[0217] p5=Sigmoid(Conv1d(GlobalAvgPool(y5(x))))
[0218] p7=Sigmoid(Conv1d(GlobalAvgPool(y7(x))))
[0219] In the formula, p3, p5, p5 are the channel weight coefficients corresponding to the one-dimensional convolution kernel size of 3, 5, and 7 respectively, sigmoid is the S-type activation function, Conv1d represents one-dimensional convolution, and GlobalAvgPool is the pooling operation for each channel input by the ECA attention mechanism;
[0220] The output result of the ECA attention mechanism is expressed as:
[0221] y(x)′=y3(x)′+y5(x)′+y7(x)′
[0222] Where y′ is the output result of the ECA attention mechanism;
[0223] y3(x)′=y3(x)·p3
[0224] y5(x)′=y5(x)·p5
[0225] y7(x)′=y7(x)·p7
[0226] Where y3(x)′, y5(x)′, and y7(x)′ are the outputs of the ECA attention mechanism when the size of the one-dimensional convolution kernel is 3, 5, and 7, respectively;
[0227] The expressions of the batch normalization and nonlinear transformation operations are:
[0228] y final =ReLU(BatchNorm(y(x)′))
[0229] In the formula, y final It is the result of batch normalization and nonlinear transformation operations, ReLU is the rectified linear activation function, and BatchNorm is the batch normalization operation;
[0230] The expression of the maximum pooling layer sampling process is:
[0231] F(t,f)=MaxPool(y final )
[0232] Where MaxPool is the maximum pooling layer operation.
[0233] Preferably, the step of capturing the global temporal dependency in the time-frequency features based on the Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features includes:
[0234] Map the time-frequency feature matrix based on the multi-head attention mechanism in the Transformer model to generate a query vector, a key vector, and a value vector corresponding to each head;
[0235] Calculate the corresponding attention weight according to the query vector and the key vector corresponding to each head, and obtain the output corresponding to each head according to the attention weight and the value vector;
[0236] The output of each head is concatenated, and the concatenated output is transformed by a linear transformation matrix to obtain the time domain features of the long-distance dependence in the time-frequency features.
[0237] Preferably, the step of fusing the extracted frequency domain features with the time domain features through a bidirectional cross attention mechanism comprises:
[0238] Calculating first cross-attention weights of the time domain features and the frequency domain features, and performing weighted summation of the first cross-attention weights and the frequency domain values in the frequency domain features to obtain a first attention output of the time domain to the frequency domain;
[0239] Calculating the second cross-attention weights of the frequency domain features and the time domain features, and performing weighted summation on the second cross-attention weights and the time domain values in the time domain features to obtain a second attention output of the frequency domain to the time domain;
[0240] The first attention output and the second attention output are fused to obtain comprehensive features of frequency domain features and time domain features.
[0241] Embodiment 3
[0242] This embodiment provides a storage medium on which a computer program is stored. When the program is executed by a processor, the bearing fault diagnosis method as described above is implemented.
[0243] Embodiment 4
[0244] The present invention also provides a computer, see Figure 2 , shown is a computer in an embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored in the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, the above-mentioned bearing fault diagnosis method is implemented.
[0245] The memory 10 includes at least one type of storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 10 may be an internal storage unit of a computer, such as a hard disk of the computer. In other embodiments, the memory 10 may also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. Further, the memory 10 may also include both an internal storage unit of a computer and an external storage device. The memory 10 may be used not only to store application software and various types of data installed in the computer, but also to temporarily store data that has been output or is to be output.
[0246] Among them, in some embodiments, the processor 20 can be an electronic control unit (Electronic Control Unit, abbreviated as ECU, also known as a vehicle computer), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run the program code stored in the memory 10 or process data, such as executing access restriction programs, etc.
[0247] It should be pointed out that Figure 2 The structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0248] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0249] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0250] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0251] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0252] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A bearing fault diagnosis method, characterized in that: include: Acquire a bearing vibration signal, and extract time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform; The time-frequency features are used as input of an adaptive multi-scale dynamic convolutional neural network to perform a convolution operation, and an ECA attention mechanism is used to process the convolution output results to extract frequency domain features from the time-frequency features; Capturing the global temporal dependency in the time-frequency features based on the Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features; The extracted frequency domain features are fused with the time domain features through a bidirectional cross-attention mechanism, the fused features are passed to the fully connected layer for classification, the fault category of the bearing is output, and the classification result is converted into a probability distribution, and the category with the highest probability is selected as the diagnosis result.
2. The bearing fault diagnosis method according to claim 1, characterized in that: The step of extracting the time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform comprises: Decomposing the bearing vibration signal into a plurality of eigenmode functions by means of the variational mode decomposition; Converting the bearing vibration signal into a frequency spectrum amplitude according to a fast Fourier transform; A plurality of the intrinsic mode functions and the spectrum amplitudes are stacked vertically to obtain a time-frequency feature matrix.
3. The bearing fault diagnosis method according to claim 2, characterized in that: The expression for the decomposition of the bearing vibration signal is: Where x(n) is the bearing vibration signal, u m (t) is the mth eigenmode function, M is the number of modes; The expression of the time-frequency feature matrix is: X(t,f)=[u1(t),u2(t),...,u M (t),|X(f)|] Where X(t, f) is the time-frequency feature matrix, u1(t), u2(t), ..., u M (t) is the M intrinsic mode functions, |X(f)| is the spectrum amplitude of the fast Fourier transform.
4. The bearing fault diagnosis method according to claim 3, characterized in that: The step of performing a convolution operation on the time-frequency feature as the input of the adaptive multi-scale dynamic convolutional neural network and processing the convolution output result using the ECA attention mechanism to extract the frequency domain feature from the time-frequency feature includes: Obtain a multi-scale convolution kernel, perform a first convolution process on the convolution kernel and the time-frequency feature matrix, and perform a second convolution process on the time-frequency feature matrix based on an adaptive weight generation network to obtain a weight matrix; Obtaining a convolution output result of an adaptive multi-scale dynamic convolutional neural network according to the first convolution output and the weight matrix to obtain a multi-scale fused convolution output; The multi-scale fused convolution output is used as the input of the ECA attention mechanism and the ECA operation is performed to obtain the channel weighting coefficient of the ECA attention mechanism input, and the output result of the ECA attention mechanism is calculated according to the channel weighting coefficient and the multi-scale fused convolution output; The output results of the ECA attention mechanism are batch normalized and nonlinearly transformed, and the results after the transformation are sampled through a maximum pooling layer to obtain the frequency domain features in the time-frequency features.
5. The bearing fault diagnosis method according to claim 4, characterized in that: The expression of the first convolution process is: Conv3(x)=Conv1d(X(t,f),g3) Conv5(x)=Conv1d(X(t,f),g5) Conv7(x)=Conv1d(X(t,f),g7) Where g3, g5, and g7 are the sizes of the one-dimensional convolution kernels, Conv3(x), Conv5(x), and Conv7(x) are the convolution outputs corresponding to the sizes of the one-dimensional convolution kernels of 3, 5, and 7, respectively. Conv1d represents one-dimensional convolution, and X(t, f) is the time-frequency feature matrix. The expression of the second convolution process is: W3(t, f)=Softmax(Conv3(x), g3) W5(t, f)=Softmax(Conv5(x), g5) W7(t, f)=Softmax(Conv7(x), g7) Wherein, W3(t, f) is the adaptive weight generated by the input X(t, f) when the one-dimensional convolution kernel is 3, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 3; W5(t, f) is the adaptive weight generated by the input X(t, f) when the one-dimensional convolution kernel is 5, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 5; W7(t, f) is the adaptive weight generated by the input X(t, f) when the one-dimensional convolution kernel is 7, which indicates the importance of the convolution with a one-dimensional convolution kernel size of 7, and Softmax is the normalized exponential activation function; The expression of the multi-scale fused convolution output is: y3(x)=W3(t,f)·Conv3(x) y5(x)=W5(t,f)·Conv5(x) y7(x)=W7(t,f)·Conv7(x) Where y3(x), y5(x), and y7(x) are the multi-scale fused convolution outputs corresponding to the one-dimensional convolution kernel sizes of 3, 5, and 7, respectively; The expression of the channel weighting coefficient is: p3=Sigmoid(Conv1d(GlobalAvgPool(y3(x)))) p5=Sigmoid(Conv1d(GlobalAvgPool(y5(x)))) p7=Sigmoid(Conv1d(GlobalAvgPool(y7(x)))) In the formula, p3, p5, p5 are the channel weight coefficients corresponding to the one-dimensional convolution kernel size of 3, 5, and 7 respectively, Sigmoid is the S-type activation function, Conv1d represents one-dimensional convolution, and GlobalAvgPool is the pooling operation for each channel of the ECA attention mechanism input; The output result of the ECA attention mechanism is expressed as: y(x)′=y3(x)′+y5(x)′+y7(x)′ Where y′ is the output result of the ECA attention mechanism; y3(x)′=y3(x)·p3 y5(x)′=y5(x)·p5 y7(x)′=y7(x)·p7 Where y3(x)′, y5(x)′, and y7(x)′ are the outputs of the ECA attention mechanism when the size of the one-dimensional convolution kernel is 3, 5, and 7, respectively; The expressions of the batch normalization and nonlinear transformation operations are: y final =ReLU(BatchNorm(y(x)′)) In the formula, y final It is the result of batch normalization and nonlinear transformation operations, ReLU is the rectified linear activation function, and BatchNorm is the batch normalization operation; The expression of the maximum pooling layer sampling process is: F(t,f)=MaxPool(y final ) Where MaxPool is the maximum pooling layer operation.
6. The bearing fault diagnosis method according to claim 2, characterized in that: The step of capturing the global temporal dependency in the time-frequency features based on the Transformer model to extract the time domain features of the long-distance dependency in the time-frequency features includes: Map the time-frequency feature matrix based on the multi-head attention mechanism in the Transformer model to generate a query vector, a key vector, and a value vector corresponding to each head; Calculate the corresponding attention weight according to the query vector and the key vector corresponding to each head, and obtain the output corresponding to each head according to the attention weight and the value vector; The output of each head is concatenated, and the concatenated output is transformed by a linear transformation matrix to obtain the time domain features of the long-distance dependence in the time-frequency features.
7. The bearing fault diagnosis method according to claim 6, characterized in that: The step of fusing the extracted frequency domain features with the time domain features through a bidirectional cross attention mechanism includes: Calculating first cross-attention weights of the time domain features and the frequency domain features, and performing weighted summation of the first cross-attention weights and the frequency domain values in the frequency domain features to obtain a first attention output of the time domain to the frequency domain; Calculating the second cross-attention weights of the frequency domain features and the time domain features, and performing weighted summation on the second cross-attention weights and the time domain values in the time domain features to obtain a second attention output of the frequency domain to the time domain; The first attention output and the second attention output are fused to obtain comprehensive features of frequency domain features and time domain features.
8. A bearing fault diagnosis system, characterized in that: include; An extraction module, used to obtain a bearing vibration signal, and extract time-frequency features in the bearing vibration signal according to variational mode decomposition and fast Fourier transform; A convolution module, used to perform a convolution operation on the time-frequency features as input of an adaptive multi-scale dynamic convolutional neural network, and process the convolution output results using an ECA attention mechanism to extract frequency domain features from the time-frequency features; A capture module, used to capture the global temporal dependency in the time-frequency features based on a Transformer model, so as to extract the time domain features of the long-distance dependency in the time-frequency features; The diagnosis module is used to fuse the extracted frequency domain features with the time domain features through a bidirectional cross-attention mechanism, pass the fused features into the fully connected layer for classification, output the fault category of the bearing, and convert the classification results into probability distribution, and select the category with the highest probability as the diagnosis result.
9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the bearing fault diagnosis method as described in any one of claims 1 to 7 is implemented.
10. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the bearing fault diagnosis method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Blockage detection method, system and equipment of flow measuring device and medium
CN120293267A
A method, system, device and medium for detecting blockage of a flow measuring device
CN120293267B
Fault early warning method and system for wind generating set
CN120576044A
Training method of coal mine pressure prediction model and coal mine pressure prediction method
CN120744850A
Rolling bearing fault diagnosis method based on time-frequency fusion and double-branch deep network
CN120992200A