Rolling bearing fault diagnosis method based on multi-scale residual attention network and adaptive Transform encoder

By combining the multi-scale residual attention network and the adaptive Transformer encoder, the problem of insufficient multi-scale and temporal feature extraction in rolling bearing fault diagnosis is solved, and efficient and accurate fault diagnosis effects are achieved.

CN120654144APending Publication Date: 2025-09-16CHINA THREE GORGES UNIV
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510731678.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In rolling bearing fault diagnosis, existing technologies lack the ability to extract multi-scale fault features. Moreover, when CNN and Transformer are used alone, it is difficult to fully utilize the various feature information of vibration signals, resulting in limited diagnostic efficiency and accuracy.

Method used

A multi-scale residual attention network is combined with an adaptive Transformer encoder to extract local features. The feature expression is optimized through residual connection and channel attention mechanism. At the same time, an adaptive Transformer encoder is introduced to capture temporal features, and adaptive position encoding is used to improve the ability to capture long-term dependencies.

Benefits of technology

It improves the accuracy and efficiency of rolling bearing fault diagnosis, enhances the ability to extract multi-scale and time series features, reduces computational costs, and improves the adaptability and diagnostic performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654144A_ABST
    Figure CN120654144A_ABST
Patent Text Reader

Abstract

The invention discloses a rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transform encoder. The rolling bearing fault diagnosis method comprises the following steps: acquiring original vibration data in the running process of a rolling bearing; segmenting the collected original vibration data into samples with specified lengths, and dividing the samples into a training data set and a test data set; inputting the training data set into a multi-scale residual attention network to perform preliminary multi-scale feature extraction; inputting the feature information extracted by the multi-scale residual attention network into an adaptive Transform encoder to obtain time sequence features; finally obtained feature information is subjected to GAP processing and then is input into a Softmax layer for fault diagnosis; the forward propagation calculation and the back propagation calculation are repeatedly executed to optimize model parameters until the diagnosis accuracy and loss of the training data set reach a stable level; and inputting the test data set into the trained model for fault diagnosis, and determining the health condition of the rolling bearing. According to the method, the adaptability and the diagnosis accuracy in time sequence dependence scenes such as rolling bearing fault diagnosis are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rolling bearing fault diagnosis, and in particular to a rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder. Background Art

[0002] With the rapid development of modern industry, rotating machinery is widely used in manufacturing, transportation, aerospace, and other fields. As a core component of rotating machinery, rolling bearings play a vital role in ensuring the safety and efficiency of mechanical operation. However, long-term high-load operation under complex operating conditions can lead to bearing damage, resulting in severe economic losses and even safety accidents. Therefore, improving the accuracy and efficiency of rolling bearing fault diagnosis has become a research hotspot in the industrial field.

[0003] In recent years, deep learning has attracted widespread attention in various fields due to its superior feature mining capabilities, opening up new research directions in fault diagnosis. Compared to traditional methods, deep learning can automatically extract key fault features from vibration signals, reducing information loss that can occur during manual processing. Consequently, a large number of deep learning models have been applied in the field of fault diagnosis. Convolutional Neural Networks (CNNs), a typical deep learning model, can automatically extract local features from vibration signals through multi-layer convolution, thereby improving the expressiveness of fault features. However, CNNs primarily rely on a fixed-size receptive field, which limits their ability to extract multi-scale fault features. To address this deficiency, researchers have proposed multi-scale convolutional neural networks (MSCNNs) to enhance the model's ability to perceive multi-scale information. Furthermore, given the significant temporal dependence of vibration signals—that is, the evolution of fault modes within a time series—feature extraction using CNNs alone cannot fully utilize the global temporal characteristics of the signal. Therefore, the Transformer model, due to its powerful sequence modeling capabilities, has been gradually introduced into the field of fault diagnosis to address the limitations of CNNs in temporal feature extraction.

[0004] While MSCNN and Transformer have demonstrated excellent fault feature extraction capabilities when used independently, rolling bearing fault diagnosis is a complex pattern recognition task that requires comprehensive consideration of multiple feature information in vibration signals. Therefore, leveraging their complementary strengths in fault feature extraction while simultaneously reducing computational costs has become a pressing research challenge. Summary of the Invention

[0005] In response to the above technical problems, the present invention proposes a rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder. The multi-scale local features of the rolling bearing signal are extracted through the multi-scale residual attention network, and the feature transfer is enhanced through the residual connection. At the same time, the channel attention mechanism is combined to optimize the feature expression, reduce redundancy and improve computational efficiency; then, an adaptive Transformer encoder is introduced to extract the signal timing features, and through adaptive position encoding, it is enabled to better capture the long-term dependencies and time correlations in the sequence data, thereby improving the fault diagnosis performance.

[0006] The technical solution adopted by the present invention is: The rolling bearing fault diagnosis method based on multi-scale residual attention network and adaptive Transformer encoder includes the following steps: Step 1: Obtaining the original vibration data during the operation of the rolling bearing; Step 2: Split the collected raw vibration data into samples of specified length and divide them into training data set and test data set; Step 3: Input the training dataset into the multi-scale residual attention network for preliminary multi-scale feature extraction; Step 4: Input the feature information extracted by the multi-scale residual attention network into the adaptive Transformer encoder to further obtain temporal features; Step 5: The final feature information is processed by Global Average Pooling (GAP) and then input into the Softmax layer for fault diagnosis; Step 6: Repeat forward and backpropagation calculations to optimize the parameters of the multi-scale residual attention network and the adaptive Transformer encoder model until the diagnostic accuracy and loss of the training dataset reach a stable level; Step 7: Input the test dataset into the trained multi-scale residual attention network and adaptive Transformer encoder model for fault diagnosis to determine the health status of the rolling bearing.

[0007] In step 1, the original vibration data of the rolling bearing during operation is obtained through the data acquisition system. The original vibration data, such as Figure 1 shown.

[0008] In step 2, the collected original vibration data is divided into samples of a specified length. The division of the original vibration data is as follows: Figure 2 shown.

[0009] In step 3, the multi-scale residual attention network is composed of a cascade of multiple multi-scale residual attention modules.X When , the initial feature map is first obtained through a 32×32 wide convolution kernel. x , and then the initial feature map x Divided into n feature map subsets, denoted as x i , i ∈{1, 2,…, n}, n Represents the initial feature map x The number of subsets divided. The spatial dimension of each subset is x Same, but the number of channels is reduced to 1 / n .remove x Except 1, the rest x i All undergo corresponding 5×5 convolution C i ( ) is processed, and its output is expressed as y i ,like Figure 3 As shown in part (a) of .

[0010] In order to fully integrate information at different scales, x i Before convolution, it is first combined with the output of the previous stage y i-1 Perform feature fusion and then input C i ( ).therefore, y i Expressed as: ; in, Indicates the i feature map subset x i The output of the convolution operation performed; Indicates the current input subset x i With the output features of the previous level y i-1 The result of feature fusion and convolution processing.

[0011] Get a new feature map y i After that, they are spliced ​​together to generate a composite feature map y m , and recalibrate the channels through SENet to dynamically adjust the response strength of each channel, and finally obtain the recalibrated feature map y n ,like Figure 3 As shown in part (b) of .

[0012] Then, a 1×1 convolution is used to calibrate the feature map y n Process and generate the initial feature map x Feature maps with the same spatial dimensions and number of channels y ,like Figure 3 As shown in part (c); Finally, x and y Superposition is performed, skip connections are established, and a mixed feature map with multi-scale characteristics is obtained. y s , y s After Batch Normalization (BN) and Rectified Linear Unit (ReLU) activation layer processing, the output is obtained Y ,like Figure 3 As shown in part (d) of .

[0013] Among them, SENet includes three operation stages: compression, excitation and feature reweighting. The overall process is as follows Figure 4 As shown, the specific implementation is as follows: S3.1, Compression stage: Quantize each feature channel into a corresponding real number through global average pooling z c , the calculation formula is: ; in, z c Represents a scalar after compression, u c represents the channel characteristics of the input, W Indicates the length of each channel data, F sp ( ) represents the global average pooling function; Indicates the c In the feature channel i The value of a spatial position.

[0014] S3.2, incentive stage: Generate the weights of each feature channel through two fully connected layers S c , the calculation formula is: ; in, Represents the activation function, which is used to generate channel weights; A scalar representing the output of the compression stage, representing the channel cGlobal information; Represents the joint operation of two fully connected layers; Express Apply the Sigmoid function to the result; Represents the dimensionality reduction result W 1 z c Apply the ReLU activation function; It means that the dimensionality reduction result is first activated by ReLU, then passed through the dimensionality increase fully connected layer, and finally the weight is generated by the Sigmoid function; W 1 represents the weight matrix of the first fully connected layer (dimensionality reduction layer), W 2 represents the weight matrix of the second fully connected layer (dimensionality increase layer).

[0015] S3.3, feature reweighting stage: The generated channel weights S c With the original feature channel input u c Perform channel-by-channel multiplication to obtain the output of SENet: ; in, Represents the features after channel reweighting; Represents the channel scaling operation, which converts the channel weight s c Original feature channel u c Multiply channel by channel.

[0016] In the multi-scale residual attention module, the number of splits n Controlling the initial feature map x The scale dimension, the larger n This may allow learning features with richer receptive field sizes. <j≤i hour, j Indicates that at the current stage i The index of the previous subset involved in the fusion in the convolution operation; i Indicates the currently processed i The feature map subset number. Ci () accepts subsets from all feature maps x j In order to simplify the structure and reduce the calculation, the present invention will n Set to 4.

[0017] In step 4, the input of the adaptive Transformer encoder is composed of the query vector (Query, Q )、Key vector(Key, K) and the value vector (Value, V ), and the calculation formula is: ; in, Represents the output calculated by the attention mechanism; d k represents the dimension of the key vector; Softmax ( ) represents the normalization operation, which is used to generate attention weights; Represents the key vector matrix K The transpose of .

[0018] No. i The attention head is linearly transformed Q 、 K and V Generate, its formula is defined as: ; in, head i Indicates the i The output of an attention head; W i Q 、 W i K and W i V Respectively represent i Attention Head Q、 K and V The linear transformation matrix of Indicates the i The calculation process of an attention head.

[0019] In the multi-head attention mechanism, the outputs of multiple attention heads are concatenated and then linearly transformed. The process is: ; in, Mh ( ) represents the output of the multi-head attention mechanism, h represents the number of attention heads, W O The weight matrix representing the linear transformation; Indicates that the outputs of all attention heads are concatenated in the feature dimension; They represent the output of the attention head respectively.

[0020] The implementation of adaptive position encoding is as follows: The core idea of ​​adaptive position encoding is to introduce a trainable position encoding matrix P, where each row of the matrix corresponds to the encoding vector of a position, which is defined as: ; in, L represents the length of the input sequence, p i Indicates the i Position encoding vectors, d model represents the embedding dimension.

[0021] Embed the original input X With the position encoding matrix P Add element-wise to form the final input representation: ; in, X′ Represents the input of the fused positional encoding.

[0022] During training, the position encoding matrix P As one of the trainable parameters of the Transformer encoder, it is optimized through gradient descent along with other parameters. The loss function of the Transformer encoder model is: ; in, L ( θ ) represents the loss function of the model, θ Represents all trainable parameters of the model, y i Indicates the i The true labels of samples are represents the model prediction value, N Represents the total number of training samples.

[0023] At each gradient update, the optimization rule for position encoding is as follows: ; in, Indicates in t +1 updated position encoding matrix; P (t) Indicates the t The position encoding moment of the iteration; η represents the learning rate; Represents the loss function for the position encoding matrix P gradient.

[0024] The step 5 comprises the following steps: S5.1: Use GAP to perform dimensionality reduction on the fused feature information to reduce the high-dimensional feature vector to the one-dimensional vector required by the classifier; S5.2: Input the final feature vector into the Softmax layer for fault classification; In S5.1, GAP can reduce the spatial dimension of the feature map to a single dimension. The calculation of GAP is determined by the following formula: ; in, Z represents the output of the previous network, N Indicates the number of output channels, L Indicates the length of each channel, V j Indicates the j The dimension reduction output of channels.

[0025] In S5.2, the Softmax function is essentially a normalized exponential function, primarily used to convert a real number vector into a probability distribution vector. In scenarios such as fault diagnosis, Softmax is often used in the last layer of a deep network to convert these outputs into a directly understandable probability form, facilitating fault classification. The calculation formula for the Softmax function is: ; in, P i Indicates the i The probability of the class, n Represents the input vector z Dimensions, z i The output of the model i The scoring value of the category, z j The output of the model j The scoring value of the category, and Represents the output after exponential enhancement.

[0026] In step 6, forward propagation and back propagation calculations are repeatedly performed to optimize the model parameters until the diagnostic accuracy and loss of the training data set reach a stable level; specifically, forward propagation refers to passing the input sample through each network layer in sequence to calculate the predicted output result of each sample; while back propagation is to calculate the gradient of the parameters of each layer according to the loss function, and use the gradient descent algorithm to update and optimize the model parameters, thereby improving the classification performance of the model.

[0027] During the forward propagation process, the input sample passes through the network and the prediction result is output. Finally, the prediction probability of each fault category is obtained through the Softmax function. This process is described in the formula described in step S5.2.

[0028] In order to measure the difference between the model prediction results and the true labels, the cross entropy loss function is introduced as the training target, which can effectively characterize the distance between the predicted distribution and the true distribution. The calculation formula of this loss function is as follows: ; in, L Represents the loss value of the current sample, m Indicates the total number of fault classification categories, y i Indicates the true label in i The value of the class, Indicates that the predicted sample belongs to i The probability of the class.

[0029] In the back propagation process, the model parameters are derived according to the loss function. After obtaining the gradient, the gradient descent algorithm is used to update the model parameters. The update formula is as follows: ; in, θ represents the set of model parameters, represents the learning rate, Represents the gradient of the loss function with respect to the parameters.

[0030] After each round of forward propagation and backpropagation, the classification effect of the model on the training set or validation set needs to be evaluated. The classification accuracy is often used as a performance indicator. The accuracy calculation formula is as follows: ; in, Ac represents the training accuracy, N represents the total number of samples, represents the predicted category, y i represents the true label, I ( ) represents an indicator function.

[0031] The present invention provides a rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder, and the technical effects are as follows: 1) This paper introduces a hierarchical structure and channel-level information fusion to construct a multi-scale residual attention network, which can achieve deep mining of features at different scales while effectively reducing model parameters and improving network operation efficiency.

[0032] 2) This paper uses residual connections to enhance feature transfer capabilities and combines the channel attention mechanism to optimize feature expression, thereby solving the gradient vanishing and gradient exploding problems that are prone to occur during the training process of deep neural networks.

[0033] 3) This paper improves the position encoding method of the Transformer encoder and adopts trainable position encoding parameters to replace the traditional fixed sine-cosine position encoding, thereby constructing an adaptive Transformer encoder. This allows the model to dynamically adjust the position representation according to specific task requirements during training, thereby improving its ability to extract timing features and enhancing its adaptability and diagnostic accuracy in timing-dependent scenarios such as rolling bearing fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The present invention will be further described below with reference to the accompanying drawings and examples: Figure 1 This is the original vibration signal data diagram used in the present invention.

[0035] Figure 2 This is a diagram showing the division of the original vibration signal data used in the present invention.

[0036] Figure 3 This is the structural diagram of the multi-scale residual attention module proposed in this invention.

[0037] Figure 4 This is the SENet operation flow chart applicable to one-dimensional vibration signals proposed in the present invention.

[0038] Figure 5 This is a flow chart of the rolling bearing fault diagnosis method proposed in the present invention.

[0039] Figure 6 This is the structural diagram of the adaptive Transformer encoder proposed in this invention.

[0040] Figure 7 This is a visualization distribution diagram of the original data of the preferred embodiment of the present invention.

[0041] Figure 8 This is a visualization distribution diagram of features extracted according to a preferred embodiment of the present invention.

[0042] Figure 9 This is a confusion matrix diagram of the diagnosis results of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to more clearly and completely illustrate the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0044] The rolling bearing fault diagnosis method based on multi-scale residual attention network and adaptive Transformer encoder is shown in the flowchart. Figure 5 As shown, it mainly includes the following steps: S1: Obtain the original vibration data of the rolling bearing during operation through the data acquisition system; S2: Split the collected data into samples of specified length and divide them into training data set and test data set; S3: Input the training dataset into the multi-scale residual attention network for preliminary multi-scale feature extraction; S4: The feature information extracted by the multi-scale residual attention network is input into the adaptive Transformer encoder to further obtain temporal features; S5: The final feature information is processed by GAP and input into the Softmax layer for fault diagnosis; S6: Repeat forward propagation and backpropagation calculations to optimize model parameters until the diagnostic accuracy and loss of the training set reach a stable level; S7: Input the test data set into the trained model for fault diagnosis to determine the health status of the rolling bearing.

[0045] Furthermore, in S3, the multi-scale residual attention network is composed of a plurality of multi-scale residual attention modules in cascade. The structure diagram of the multi-scale residual attention module is as follows: Figure 3 Specifically, when the input signal X When , the initial feature map is first obtained through a 32×32 wide convolution kernel. x The convolution kernel provides a larger receptive field, which helps the model capture global information and lays the foundation for subsequent multi-scale feature extraction. Then the feature map x Divided into n feature map subsets, denoted as x i ( i ∈{1, 2,…, n}). The spatial dimension of each subset is x Same, but the number of channels is reduced to 1 / n .remove x Except 1, the rest x i All of them go through the corresponding 5×5 convolution (denoted as C i ( )) is processed, and its output is recorded as y i In order to fully integrate information of different scales, x i Before convolution, it is first combined with the output of the previous stage yi-1 Perform feature fusion and then input C i ( ).therefore, y i It can be expressed as: ; Get a new feature map y i After that, they are spliced ​​together to generate a composite feature map y m , and recalibrate the channels through SENet to dynamically adjust the response strength of each channel, and finally obtain the recalibrated feature map y n Then, a 1×1 convolution is used to y n Process and generate the initial feature map x Feature maps with the same spatial dimensions and number of channels y Next, x and y Superposition is performed, skip connections are established, and a mixed feature map with multi-scale characteristics is obtained. y s .at last, y s After Batch Normalization (BN) and Rectified Linear Unit (ReLU) activation layer processing, the output is obtained Y .

[0046] Preferably, in the multi-scale residual attention module, the number of splits n Controlling the initial feature map x The scale dimension, the larger n This may allow learning features with richer receptive field sizes. <j≤i hour, Ci () can accept subsets from all feature maps x j In order to simplify the structure and reduce the calculation, the present invention will n Set to 4.

[0047] Among them, SENet mainly includes three operation stages: compression, excitation and feature reweighting. The specific implementation is as follows: S3.1: Compression stage. Each feature channel is quantized into a corresponding real number through global average pooling z c , the calculation formula is: ; in, zc Represents a scalar after compression, u c represents the channel characteristics of the input, W Indicates the length of each channel data, F sp ( ) represents the global average pooling function.

[0048] S3.2: Excitation stage. Generate the weights of each feature channel through two fully connected layers S c , the calculation formula is: ; in, σ represents the Sigmoid activation function, δ represents the ELU activation function, W 1 represents the weight matrix of the first fully connected layer (dimensionality reduction layer), W 2 represents the weight matrix of the second fully connected layer (dimensionality increase layer).

[0049] S3.3: Feature reweighting stage. The generated channel weights S c With the original feature channel input u c Perform channel-by-channel multiplication to obtain the output of SENet: ; in, x c Represents the features after channel reweighting.

[0050] Furthermore, in S4, the adaptive Transformer encoder structure is as follows: Figure 6 As shown in the figure, by improving the position encoding method of the Transformer encoder and using trainable position encoding parameters instead of the traditional fixed sine-cosine position encoding, the model can dynamically adjust the position representation according to specific task requirements during training, thereby improving its ability to express timing information and enhancing its adaptability and diagnostic accuracy in timing-dependent scenarios such as rolling bearing fault diagnosis.

[0051] The Transformer encoder maps the input sequence into a series of representations, weights the importance of each token in the sequence through the self-attention mechanism, and then performs nonlinear transformations on these representations through a feedforward neural network. The self-attention mechanism is the core component of the Transformer encoder. It distributes weights to each position in the sequence by dynamically adjusting the contribution of each position to the current task. Its input consists of a query vector (Query, Q )、Key vector(Key, K) and the value vector (Value, V ), and the calculation formula is: ; in, d k represents the dimension of the key vector, Softmax ( ) represents the normalization operation, which is used to generate attention weights.

[0052] No. i The attention head is linearly transformed Q 、 K and V Generate, its formula is defined as: ; in, head i Indicates the i The output of an attention head, W i Q 、 W i K and W i V Respectively represent i Attention Head Q、 K and V The linear transformation matrix of .

[0053] In the multi-head attention mechanism, the outputs of multiple attention heads are concatenated and then linearly transformed. The process is: ; in, Mh ( ) represents the output of the multi-head attention mechanism, h represents the number of attention heads, W O A weight matrix representing a linear transformation.

[0054] The implementation process of adaptive position encoding is as follows: 4.1: The core idea of ​​adaptive position encoding is to introduce a trainable position encoding matrix P , where each row of the matrix corresponds to the encoding vector of a position, which is defined as: ; in, L represents the length of the input sequence, d model represents the embedding dimension, p i Indicates the i Position encoding vector.

[0055] 4.2: Embedding the original input X With the position encoding matrix P Add element-wise to form the final input representation: ; in, X′ Represents the input of the fused positional encoding.

[0056] 4.3: During training, the position encoding matrix P As one of the trainable parameters of the Transformer encoder, it is optimized through gradient descent along with other parameters. The loss function of the model is: ; in, L ( θ ) represents the loss function of the model, θ Represents all trainable parameters of the model, y i Indicates the i The true labels of samples are represents the model prediction value, N Represents the total number of training samples.

[0057] 4.4: At each gradient update, the optimization rule for position encoding is as follows: ; in, P (t) Indicates the t The position encoding moment of the iteration, η represents the learning rate, Represents the loss function for the position encoding matrix P gradient.

[0058] Furthermore, the S5 can be divided into the following two steps: S5.1: Use GAP to perform dimensionality reduction on the fused feature information to reduce the high-dimensional feature vector to the one-dimensional vector required by the classifier; S5.2: Input the final feature vector into the Softmax layer for fault classification.

[0059] Preferably, GAP is different from the traditional pooling layer, which calculates the average value of each channel as the channel output. In this way, GAP can reduce the spatial dimension of the feature map to a single dimension. The calculation of GAP can be determined by the following formula: ; in, Z represents the output of the previous network,N Indicates the number of output channels, L Indicates the length of each channel, V j Indicates the j The dimension reduction output of channels.

[0060] The following is a detailed description with reference to specific embodiments: A preferred embodiment of the present invention is an application of a rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder on a bearing fault diagnosis dataset of a mechanical fault simulator (MFS).

[0061] The MFS primarily consists of a drive motor, motor speed controller, digital tachometer, test bearing, and sensors. Experimental data is collected simultaneously in the horizontal and vertical directions of the test bearing using a PCB352C33 uniaxial vibration accelerometer. The MFS bearing dataset covers five operating conditions: normal, inner race fault, outer race fault, ball fault, and combined fault (including the aforementioned three). All faulty bearings are ER12KCL type.

[0062] The speed range of the MFS drive motor is 500rpm to 3600rpm. When the motor speed approaches the maximum speed, it may generate large noise and cause damage to the equipment. Therefore, in this embodiment, the following four speeds are selected for experiments: 1000rpm, 1400rpm, 1800rpm and 2200rpm. Accordingly, the MFS bearing data set is collected at the above four speeds, with a sampling frequency of 16kHz and a sampling time of 10s. The data set corresponding to each speed covers five working states. Each data set is divided into 156 samples, and each sample contains 1024 continuous data points. Among them, 104 samples are randomly selected as training sets, and the remaining samples are used for testing. The specific information of the data set is shown in Table 1.

[0063] Table 1 Description of the MFS bearing dataset

[0064] Furthermore, the training set is input into the deep learning model of the multi-scale residual attention network and adaptive Transformer encoder proposed in the present invention for training. After the diagnostic accuracy rises to a stable level, a trained model is obtained, and then the test set data is input into the model for testing.

[0065] At 1000 rpm, the original data and the features extracted by the proposed model were visualized using the t-Distributed Stochastic Neighbor Embedding (t-SNE) method to verify the feature extraction capability of the proposed model. Figure 7 As shown in , the different fault features of the original data overlap significantly and are difficult to distinguish. Figure 8 As shown in Figure 2, the feature distribution extracted by the model proposed in this invention is clearly separated, and each fault category is effectively distinguished, which shows that the model proposed in this invention has excellent feature extraction capabilities. In addition, the confusion matrix of the test set at different speeds is shown in Figure 2. Figure 9 As shown in the figure, the results show that the proposed model can accurately identify fault samples at various speeds and achieve good diagnostic results.

[0066] Preferably, to further verify the superiority of the proposed model, CNN and ResNet were used as basic comparison models. Furthermore, to demonstrate the superior feature extraction capabilities of the multi-scale residual attention network and the effectiveness of its improvements on the Transformer encoder, ablation experiments were conducted on the multi-scale residual attention network and the Transformer encoder. For ease of documentation, the multi-scale residual attention network is denoted as MSRAN, the Transformer encoder is denoted as TE, and the adaptive Transformer encoder is denoted as ATE. Detailed comparative experimental results are shown in Table 2.

[0067] As can be seen in Table 2, the diagnostic accuracy of the proposed model under all four loads is above 99.8%, which is higher than that of the basic comparison model, demonstrating its superior fault diagnosis performance. Furthermore, although the diagnostic accuracy of the ablation model is lower than that of the proposed model, it still has an advantage over the basic comparison model, demonstrating the excellent feature extraction capabilities of the proposed multi-scale residual attention network and the effectiveness of its improvements to the Transformer encoder.

[0068] Table 2 Experimental results of different models

Claims

1. Rolling bearing fault diagnosis method based on multi-scale residual attention network and adaptive Transformer encoder, characterized by The following steps are involved: Step 1: Obtaining the original vibration data during the operation of the rolling bearing; Step 2: Split the collected raw vibration data into samples of specified length and divide them into training data set and test data set; Step 3: Input the training dataset into the multi-scale residual attention network for preliminary multi-scale feature extraction; Step 4: Input the feature information extracted by the multi-scale residual attention network into the adaptive Transformer encoder to further obtain temporal features; Step 5: The final feature information is processed by global average pooling and then input into the Softmax layer for fault diagnosis; Step 6: Repeat forward and backpropagation calculations to optimize the parameters of the multi-scale residual attention network and the adaptive Transformer encoder model until the diagnostic accuracy and loss of the training dataset reach a stable level; Step 7: Input the test dataset into the trained multi-scale residual attention network and adaptive Transformer encoder model for fault diagnosis to determine the health status of the rolling bearing.

2. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 1 is characterized in that: In step 3, the multi-scale residual attention network is composed of a cascade of multiple multi-scale residual attention modules. X When , the initial feature map is first obtained through a 32×32 wide convolution kernel. x , and then the initial feature map x Divided into n feature map subsets, denoted as x i , i ∈{1, 2,…, n }, n Represents the initial feature map x The number of subsets divided; the spatial dimension of each subset is x Same, but the number of channels is reduced to 1 / n ;remove x Except 1, the rest x i All undergo corresponding 5×5 convolution C i ( ) is processed, and its output is expressed as y i ; In order to fully integrate information at different scales, x i Before convolution, it is first combined with the output of the previous stage y i-1 Perform feature fusion and then input C i ( );therefore, y i Expressed as: ; in, Indicates the i feature map subset x i The output of the convolution operation performed; Indicates the current input subset x i With the output features of the previous level y i-1 The result of feature fusion and convolution processing; Get a new feature map y i After that, they are spliced ​​together to generate a composite feature map y m , and recalibrate the channels through SENet to dynamically adjust the response strength of each channel, and finally obtain the recalibrated feature map y n ; Then, a 1×1 convolution is used to calibrate the feature map y n Process and generate the initial feature map x Feature maps with the same spatial dimensions and number of channels y ; Finally, x and y Superposition is performed, skip connections are established, and a mixed feature map with multi-scale characteristics is obtained. y s , y s After batch normalization and rectified linear unit activation layer processing, the output is Y .

3. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 2 is characterized by: SENet includes three operation stages: compression, excitation, and feature reweighting, as follows: 3.1 Compression stage: Quantize each feature channel into a corresponding real number through global average pooling z c , the calculation formula is: ; in, z c Represents a scalar after compression, u c represents the channel characteristics of the input, W Indicates the length of each channel data, F sp ( ) represents the global average pooling function; Indicates the c In the feature channel i The value of a spatial position; 3.2, Incentive stage: Generate the weights of each feature channel through two fully connected layers S c , the calculation formula is: ; in, Represents the activation function, which is used to generate channel weights; A scalar representing the output of the compression stage, representing the channel c Global information; Represents the joint operation of two fully connected layers; Express Apply the Sigmoid function to the result; Represents the dimensionality reduction result W 1 z c Apply the ReLU activation function; It means that the dimensionality reduction result is first activated by ReLU, then passed through the dimensionality increase fully connected layer, and finally the weight is generated by the Sigmoid function; W 1 represents the weight matrix of the first fully connected layer (dimensionality reduction layer), W 2 represents the weight matrix of the second fully connected layer (dimensionality-increasing layer); 3.

3. Feature reweighting stage: The generated channel weights S c With the original feature channel input u c Perform channel-by-channel multiplication to obtain the output of SENet: ; in, Represents the features after channel reweighting; Represents the channel scaling operation, which converts the channel weight s c Original feature channel u c Multiply channel by channel.

4. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 3 is characterized by: In the multi-scale residual attention module, the number of splits n Controlling the initial feature map x The scale dimension, the larger n May allow learning features with richer receptive field sizes; because when 1 <j≤i hour, j Indicates that at the current stage i The index of the previous subset involved in the fusion in the convolution operation; i Indicates the currently processed i feature map subset number; Ci () accepts subsets from all feature maps x j Characteristic information of n Set to 4.

5. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 4 is characterized in that: In step 4, the input of the adaptive Transformer encoder is composed of the query vector Q , key vector K Sum value vector V Composition, the calculation formula is: ; in, Represents the output calculated by the attention mechanism; d k represents the dimension of the key vector; Softmax ( ) represents the normalization operation, which is used to generate attention weights; Represents the key vector matrix K The transpose of No. i The attention head is linearly transformed Q 、 K and V Generate, its formula is defined as: ; in, head i Indicates the i The output of an attention head; W i Q 、 W i K and W i V Respectively represent i Attention Head Q, K and V The linear transformation matrix of Indicates the i The calculation process of an attention head; In the multi-head attention mechanism, the outputs of multiple attention heads are concatenated and then linearly transformed. The process is: ; in, Mh ( ) represents the output of the multi-head attention mechanism, h represents the number of attention heads, W O The weight matrix representing the linear transformation; Indicates that the outputs of all attention heads are concatenated in the feature dimension; They represent the output of the attention head respectively.

6. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 5, characterized in that: The implementation of adaptive position encoding is as follows: Adaptive Positional Encoding introduces a trainable positional encoding matrix P , where each row of the matrix corresponds to the encoding vector of a position, which is defined as: ; in, L represents the length of the input sequence, p i Indicates the i Position encoding vectors, d model represents the embedding dimension; Embed the original input X With the position encoding matrix P Add element-wise to form the final input representation: ; in, X′ represents the input of the fused positional encoding; During training, the position encoding matrix P As one of the trainable parameters of the Transformer encoder, it is optimized through gradient descent along with other parameters. The loss function of the Transformer encoder model is: ; in, L ( θ ) represents the loss function of the model, θ Represents all trainable parameters of the model, y i Indicates the i The true labels of samples are represents the model prediction value, N represents the total number of training samples; At each gradient update, the optimization rule for position encoding is as follows: ; in, Indicates in t +1 updated position encoding matrix; P (t) Indicates the t The position encoding moment of the iteration; η represents the learning rate; Represents the loss function for the position encoding matrix P gradient.

7. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 6, characterized in that: The step 5 comprises the following steps: S5.1: Use GAP to perform dimensionality reduction on the fused feature information to reduce the high-dimensional feature vector to the one-dimensional vector required by the classifier; S5.2: Input the final feature vector into the Softmax layer for fault classification.

8. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 7, characterized in that: In S5.1, GAP can reduce the spatial dimension of the feature map to a single dimension. The calculation of GAP is determined by the following formula: ; in, Z represents the output of the previous network, N Indicates the number of output channels, L Indicates the length of each channel, V j Indicates the j The dimension reduction output of channels.

9. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 7, characterized in that: In S5.2, the calculation formula of the Softmax function is: ; in, P i Indicates the i The probability of the class, n Represents the input vector z Dimensions, z i The output of the model i The scoring value of the category, z j The output of the model j The scoring value of the category, and Represents the output after exponential enhancement.

10. The rolling bearing fault diagnosis method based on a multi-scale residual attention network and an adaptive Transformer encoder according to claim 9, characterized in that: In step 6, forward propagation and back propagation calculations are repeatedly performed to optimize the model parameters until the diagnostic accuracy and loss of the training data set reach a stable level; specifically: Forward propagation refers to passing the input sample through each network layer in sequence to calculate the predicted output result of each sample; while backpropagation is to calculate the gradient of each layer parameter according to the loss function, and use the gradient descent algorithm to update and optimize the model parameters, thereby improving the classification performance of the model; In the forward propagation process, the input sample passes through the network and the prediction result is output, and finally the prediction probability of each fault category is obtained through the Softmax function; In order to measure the difference between the model prediction results and the true labels, the cross entropy loss function is introduced as the training objective, which can effectively characterize the distance between the predicted distribution and the true distribution; the calculation formula of the loss function is as follows: ; in, L Represents the loss value of the current sample, m Indicates the total number of fault classification categories, y i Indicates the true label in i The value of the class, Indicates that the predicted sample belongs to i class probability; In the back propagation process, the model parameters are derived according to the loss function. After obtaining the gradient, the gradient descent algorithm is used to update the model parameters. The update formula is as follows: ; in, θ represents the set of model parameters, represents the learning rate, Represents the gradient of the loss function with respect to the parameters; After each round of forward propagation and backpropagation, the classification effect of the model on the training set or validation set needs to be evaluated. The classification accuracy is often used as a performance indicator. The accuracy calculation formula is as follows: ; in, Ac represents the training accuracy, N represents the total number of samples, represents the predicted category, y i represents the true label, I ( ) represents an indicator function.

Citation Information

Cited By

  • Equipment fault diagnosis method and device based on AI large model, equipment and medium

    CN121211288A

  • Rolling bearing fault multi-source vibration feature extraction method

    CN121388545A

  • Dynamic leakage fault diagnosis method for flexible hand for deep-sea submersible vehicle

    CN121524814A

  • Dry bulk cargo port rotating machinery fault diagnosis method based on Transform-KAN network

    CN121637165A

  • Few-sample fault diagnosis method and device suitable for core component of rail train

    CN121659022A