A lightweight rotating machinery fault diagnosis method based on SEFormer

By adopting the SEFormer model in rotary mechanical fault diagnosis, combining the separation of multi-scale depth convolution and efficient self-attention module, the problem of difficulty in capturing global dependencies and local features in the existing technology is solved, and efficient and robust fault diagnosis performance is achieved.

CN119293589BActive Publication Date: 2025-05-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411366242.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-05-13
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the global dependence and local characteristics of long-term sequences in rotary mechanical fault diagnosis, especially in the case of noise interference, resulting in poor diagnostic performance.

Method used

A lightweight rotary mechanical fault diagnosis method based on SEFormer is designed to extract and integrate multi-scale features from different channel dimensions of vibration signals and capture key fine-grained features from global scope by separating multi-scale depth convolution SMDC and efficient self-attention ESA modules.

Benefits of technology

It realizes the reduction of computing costs and improves industrial applicability while maintaining the overall performance of the model, and significantly improves the robustness and generalization ability of rotary machinery fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293589B_ABST
    Figure CN119293589B_ABST
Patent Text Reader

Abstract

A lightweight rotating machinery fault diagnosis method based on SEFormer, 1) Data acquisition: collect vibration signals of key transmission components (such as bearings and gears) of rotating machinery in various health states; 2) Data preprocessing: segment the vibration signal into sample sets through sliding window sampling, and divide it into training set, validation set and test set; 3) Model building: design and develop separable multi-scale deep convolution SMDC and efficient self-attention ESA, and build SEFormer model; 4) Model training, verification and evaluation: train the model on the training set and validation set, and evaluate the model performance on the test set; 5) Fault diagnosis: use the trained model to diagnose the fault of rotating machinery. The present invention constructs a lightweight SEFormer model for rotating machinery fault diagnosis, extracts and integrates multi-scale features from different channel dimensions of vibration signals, captures key fine-grained features of vibration signals from a global perspective, and has the advantages of robustness, generalization ability and lightweight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of rotating machinery fault diagnosis, and specifically relates to a lightweight rotating machinery fault diagnosis method based on SEFormer. Background Art

[0002] With the rapid development of modern mechanical systems, rotating machinery has become indispensable in intelligent equipment, and academia and industry have paid great attention to its safety. Under high-intensity working conditions, key transmission components of rotating machinery (such as bearings and gears) will inevitably suffer from wear, corrosion, deformation, fracture and other faults. Faulty transmission components directly affect the operational reliability of rotating machinery, which may lead to serious accidents, economic losses and even casualties. Therefore, it is of great research value to carry out fault diagnosis and predictive maintenance of rotating machinery.

[0003] Traditional data-driven fault diagnosis methods rely on expert experience to manually extract effective fault features of signals, which has certain limitations. Deep learning technology has become a research hotspot in the field of fault diagnosis due to its powerful fault feature extraction capabilities and end-to-end diagnostic characteristics. In particular, methods based on convolutional neural networks (CNNs) have performed well in various fault diagnosis tasks. Although CNNs perform well in extracting local features of short-time series, they lack the ability to establish global dependencies in long-time series. When there is noise interference in the collected signal, it will be difficult to identify effective fault information relying only on local features.

[0004] As a rising star in the fields of natural language processing (NLP) and computer vision (CV), Transformer can capture fine-grained features and establish temporal correlations from long time series by evaluating the similarity between sequences. However, Transformer lacks the local correlation extraction and spatial inductive bias capabilities of CNN. Therefore, a large number of training samples are usually required to ensure the effectiveness of performance, which poses a challenge to fault diagnosis tasks. In addition, vibration signals are periodic and continuous, so local features cannot be ignored.

[0005] Recently, the collaborative method CNN-Transformer, which combines CNN and Transformer, has been explored to jointly capture the local features and global dependencies of the signal, but the cross-channel convolution mechanism of CNN and the self-attention calculation of Transformer make the complexity of the collaborative model too high. This complexity leads to high computational cost and limited industrial applicability. Therefore, it is particularly important to design and develop simple, lightweight and efficient convolution mechanisms and self-attention calculation methods while maintaining the overall performance of the model.

[0006] Technical comparison with patent CN113505830B "rotating machinery fault diagnosis method, system, equipment and storage medium"

[0007] Patent CN113505830B uses a lightweight convolutional neural network to extract rotating machinery fault signals, obtain spatial feature information, and perform feature recognition on the spatial feature information; this patent constructs a lightweight SEFormer model for rotating machinery fault diagnosis, extracts and integrates multi-scale features from different channel dimensions of vibration signals, and captures key fine-grained features of vibration signals from a global perspective. There is an essential difference between the two in terms of technical thinking.

[0008] Patent CN113505830B established a lightweight convolutional neural network fault diagnosis model, which is based on the convolutional neural network CNN architecture and is an improvement in the network structure; this patent designed and developed a separable multi-scale deep convolution SMDC module and an efficient self-attention ESA module, and based on these two modules, built a lightweight SEFormer model for rotating machinery fault diagnosis. The model is based on the CNN-Transformer architecture and is a design and development of the network structure. There is an essential difference between the two at the algorithm level.

[0009] Technical comparison with patent CN117113170A "A lightweight rotating machinery fault diagnosis method based on multi-scale information fusion"

[0010] Patent CN117113170A performs continuous wavelet transform on the collected original vibration data of rotating machinery to obtain a time-frequency diagram data set with enhanced state information, and the sample data is a two-dimensional time-frequency image; this patent directly divides the collected original vibration signal of rotating machinery into a data set through sliding window sampling, and the sample data is a one-dimensional time domain vibration signal. There is an essential difference between the two in the data preprocessing methods of vibration signals.

[0011] Patent CN117113170A builds a dual-stream network structure, including an upper branch structure and a lower branch structure. This model belongs to the dual-stream network architecture; this patent designs and develops a separable multi-scale deep convolution SMDC module and an efficient self-attention ESA module. Based on these two modules, a lightweight SEFormer model is constructed for rotating machinery fault diagnosis. The model is based on the CNN-Transformer architecture. There is an essential difference between the network structures of the two. Summary of the invention

[0012] In order to solve the above technical problems, the present invention proposes a lightweight rotating machinery fault diagnosis method based on SEFormer. The proposed method has two main purposes: (1) to design and develop a simple, lightweight and efficient convolution mechanism and self-attention calculation method; (2) to construct a CNN-Transformer model that has both robustness, generalization ability and lightweight for rotating machinery fault diagnosis.

[0013] To achieve the above object, the technical solution adopted by the present invention is:

[0014] A lightweight rotating machinery fault diagnosis method based on SEFormer comprises the following steps:

[0015] Step 1: Data collection:

[0016] Collect vibration signals of key transmission components of rotating machinery under various health conditions;

[0017] Install at least one sensor at different positions or directions near the rotating machinery transmission parts that need to be diagnosed, and use data acquisition equipment to collect vibration signals under various health conditions;

[0018] Step 2: Data preprocessing:

[0019] The vibration signal is divided into sample sets through sliding window sampling, and then divided into training set, validation set and test set;

[0020] Step 3: Model building:

[0021] Design and develop separable multi-scale deep convolution SMDC and efficient self-attention ESA, and build SEFormer model;

[0022] Step 4: Model training, verification and evaluation:

[0023] Train the model on the training set and validation set, and evaluate the model performance on the test set;

[0024] Step 5: Fault diagnosis:

[0025] Use the trained model to perform fault diagnosis on rotating machinery.

[0026] Furthermore, in the second step of the preprocessing process, in order to avoid test leakage, when the sliding window is sampled, there is no overlap between adjacent windows and a certain interval is maintained; each sample contains 1024 signal points; all samples are normalized by mean-standard deviation.

[0027] Further, in step 3, the SEFormer model is composed of a plurality of feature extraction layers and an output layer connected in sequence, and the number of feature extraction layers is adjusted according to the requirements of different tasks;

[0028] In the feature extraction layer;

[0029] Firstly, separable multi-scale deep convolution (SMDC) is used to extract and integrate multi-scale features from different channel dimensions of vibration signals.

[0030] Secondly, efficient self-attention (ESA) is used to capture the key fine-grained features of vibration signals from a global perspective;

[0031] Then, the residual connection Add is used to reduce the risk of overfitting, and batch normalization BN is used to stabilize the feature distribution;

[0032] Next, the feedforward network FFN is used to perform nonlinear transformation on the features;

[0033] Finally, the extracted features are output through residual connection Add and batch normalization BN;

[0034] The separable multi-scale deep convolution SMDC;

[0035] Firstly, the multi-local receptive field features of the vibration signal are extracted using parallel multi-scale deep convolutions with different kernel sizes;

[0036] Then, the multi-scale features are concatenated along the channel dimension;

[0037] Next, point-wise convolution is used to integrate feature information and capture the correlation between channels;

[0038] Finally, batch normalization BN and Gaussian error linear unit GELU are used to stabilize feature distribution and perform nonlinear mapping, as shown in the following formula:

[0039]

[0040] in, represents the input, C1 and L1 represent the channel dimension and time dimension of the input respectively, k l represents the size of the depth convolution kernel, represents the depth convolution kernel k l The weight of represents the depth convolution kernel k l The output, represents the output of multi-scale depth convolution, represents the weight of the point-by-point convolution kernel, represents the output, n represents the number of different depth convolution kernel sizes, and L2 represents the time dimension of the output;

[0041] The efficient self-attention ESA;

[0042] First, in order to expand the representation capacity of each feature and reduce computation, three parallel depth-wise separable convolutions DSConv are used to generate the input features of self-attention: query matrix Q, key matrix K, and value matrix V;

[0043] Secondly, efficient attention is used to transfer global feature information while reducing the cost of self-attention calculation;

[0044] Then, Softmax is used as a normalization operation suitable for efficient attention. Softmax normalization is performed on the matrix Q in the time dimension and on the matrix K in the channel dimension.

[0045] Finally, the attention weight is calculated through matrix transposition and matrix multiplication operations, as shown in the following formula:

[0046] Efficient attention(Q,K,V)=Softmax L (Q)·[(Softmax C (K) T ·V] (4)

[0047] in, and denote the query matrix, key matrix, and value matrix respectively, L and C denote the time dimension and channel dimension respectively, (·) T and represent matrix transposition and matrix multiplication respectively, Softmax L (·) and Softmax C (·) represents the Softmax normalization in the time dimension and channel dimension respectively;

[0048] The residual connection Add is shown in the following formula:

[0049] x'=f(x)+x (5)

[0050] The batch normalization BN is shown in the following formula:

[0051]

[0052] Among them, h l ={h l(1) ,…,h l(N)} represents the input feature map of the lth layer with a batch size of N, h l(n) ={h1 l (n) ,…,h k l(n)},y j l(n) represents the j-th output feature map of the l-th layer, u j , σ j 2 Respectively represent h j l The mean and variance of , ε represents a small constant used to prevent invalid calculations when the variance is 0, They represent the scale and translation parameters that need to be learned respectively;

[0053] The Gaussian error linear unit GELU is shown in the following formula:

[0054]

[0055] The feedforward network FFN includes two layers of linear transformation and intermediate nonlinear activation, as shown in the following formula:

[0056] FFN(x)=W2σ(W1x+b1)+b2 (11)

[0057] Among them, x represents the input, W1 and W2 represent the weights of the linear transformation of the first and second layers respectively, b1 and b2 represent the bias of the linear transformation of the first and second layers respectively, and σ represents the intermediate nonlinear activation function;

[0058] In the output layer;

[0059] First, global average pooling (GAP) is used to reduce the feature dimension in the time dimension;

[0060] Secondly, point-by-point convolution PConv is used to integrate features in the channel dimension;

[0061] Then, batch normalization (BN) and Gaussian error linear unit (GELU) are used to stabilize feature distribution and perform nonlinear mapping.

[0062] Finally, the high-dimensional features are mapped to the categorical dimension of health status through fully connected FC.

[0063] Furthermore, the training of the SEFormer model adopts a multi-classification cross entropy loss function to calculate the training loss, uses the AdamW optimization algorithm to update the model parameters, and dynamically adjusts the learning rate using an adaptive decay strategy based on the loss of the validation set.

[0064] Compared with the prior art, the advantages of the present invention are:

[0065] (1) A separable multi-scale deep convolution (SMDC) is designed to extract and integrate multi-scale features from different channel dimensions of vibration signals;

[0066] (2) We developed an efficient self-attention (ESA) to capture the key fine-grained features of vibration signals from a global perspective;

[0067] (3) Based on separable multi-scale deep convolution (SMDC) and efficient self-attention (ESA), a SEFormer model with robustness, generalization ability and lightweight is constructed for rotating machinery fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1is a flow chart of the method of the present invention;

[0069] Figure 2 It is a structural schematic diagram of the SEFormer model of the method of the present invention;

[0070] Figure 3 It is a schematic diagram of the structure of the separable multi-scale deep convolution SMDC of the method of the present invention;

[0071] Figure 4 Schematic diagram of the structure of the efficient self-attention ESA of the method of the present invention. DETAILED DESCRIPTION

[0072] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments:

[0073] In this embodiment, if Figure 1 As shown, a lightweight rotating machinery fault diagnosis method based on SEFormer includes the following steps: Step 1, data acquisition: collecting vibration signals of key transmission components of rotating machinery (such as bearings and gears) in various health states; Step 2, data preprocessing: dividing the vibration signal into sample sets through sliding window sampling, and dividing them into training set, verification set and test set; Step 3, model building: designing and developing separable multi-scale deep convolution SMDC and efficient self-attention ESA, and constructing SEFormer model; Step 4, model training, verification and evaluation: training the model on the training set and verification set, and evaluating the model performance on the test set; Step 5, fault diagnosis: using the trained model to diagnose the fault of rotating machinery.

[0074] In step one, one or more sensors are installed at different positions or directions near the rotating mechanical transmission parts (such as bearings and gears) that need to be diagnosed, and vibration signals under various health conditions are collected using data acquisition equipment.

[0075] In step 2, in order to avoid test leakage, when sliding window sampling, there is no overlap between adjacent windows and a certain interval is maintained; each sample contains 1024 signal points; all samples are normalized by mean-standard deviation.

[0076] In this embodiment, acceleration sensors are installed in the X and Y directions of the planetary gearbox housing to collect multi-sensor vibration signals of bearings and gears respectively. During the collection process, the motor speed is 1800r / min and the sampling frequency is 20480Hz. The fault types of the planetary gearbox include four bearing faults and four gear faults. A total of nine vibration signals in healthy states including normal state are collected, as shown in Table 1 below:

[0077]

[0078] In this embodiment, the vibration signal of each health state is divided into 1200 samples through sliding window sampling. Among them, 500 samples are used for training, 300 samples are used for verification, and 400 samples are used for testing. In order to evaluate the anti-noise ability of the model, two types of noise are added to the test samples in turn, as shown in the following formula:

[0079] S i =(S i +α)×β (1)

[0080] Among them, S i represents the i-th signal point of the test sample, α~N(0,σ) and β~N(1,σ) represent additive Gaussian noise and scaled Gaussian noise respectively, σ represents the variance of Gaussian distribution, and the larger the σ value, the more significant the noise. In order to simulate the uncertainty of noise, each test sample has a 50% probability of randomly adding any type of noise.

[0081] In this embodiment, if Figure 2 As shown, in step 3, the SEFormer model is composed of three feature extraction layers and an output layer connected in sequence;

[0082] In the feature extraction layer, first, separable multi-scale deep convolution SMDC is used to extract and integrate multi-scale features from different channel dimensions of the vibration signal. Secondly, efficient self-attention ESA is used to capture the key fine-grained features of the vibration signal from a global perspective. Then, residual connection Add is used to reduce the risk of overfitting, and batch normalization BN is used to stabilize the feature distribution. Next, a feedforward network FFN is used to perform nonlinear transformation on the features. Finally, the extracted features are output through residual connection Add and batch normalization BN;

[0083] like Figure 3 As shown in the figure, the separable multi-scale deep convolution SMDC first uses parallel multi-scale deep convolutions with different kernel sizes to extract multi-local receptive field features of the vibration signal. Then, the multi-scale features are concatenated along the channel dimension. Next, point-by-point convolution is used to integrate feature information and capture the correlation between channels. Finally, batch normalization BN and Gaussian error linear unit GELU are used to stabilize feature distribution and perform nonlinear mapping, as shown in the following formula:

[0084]

[0085] in, represents the input, C1 and L1 represent the channel dimension and time dimension of the input respectively, k l represents the size of the depth convolution kernel, represents the depth convolution kernel k l The weight of represents the depth convolution kernel k l The output, represents the output of multi-scale depth convolution, represents the weight of the point-by-point convolution kernel, represents the output, n represents the number of different depth convolution kernel sizes, and L2 represents the time dimension of the output;

[0086] like Figure 4 As shown, the efficient self-attention ESA, first, in order to expand the representation capacity of each feature and reduce calculations, uses three parallel depth-separable convolutions DSConv to generate the input features of self-attention: query matrix Q, key matrix K and value matrix V. Secondly, efficient attention Efficientattention is used to transfer information of global features while reducing the cost of self-attention calculation. Then, Softmax is used as a normalization operation suitable for the form of efficient attention Efficientattention, and the matrix Q is Softmax normalized in the time dimension, and the matrix K is Softmax normalized in the channel dimension. Finally, the attention weight is calculated by matrix transposition and matrix multiplication operations, as shown in the following formula:

[0087] Efficient attention(Q,K,V)=Softmax L (Q)·[(Softmax C (K) T ·V] (5)

[0088] in, and denote the query matrix, key matrix, and value matrix respectively, L and C denote the time dimension and channel dimension respectively, (·) T and represent matrix transposition and matrix multiplication respectively, Softmax L (·) and Softmax C (·) represents the Softmax normalization in the time dimension and channel dimension respectively;

[0089] The residual connection Add is shown in the following formula:

[0090] x'=f(x)+x (6)

[0091] The batch normalization BN is shown in the following formula:

[0092]

[0093] Among them, h l ={h l(1) ,…,h l(N)} represents the input feature map of the lth layer with a batch size of N, hl(n) ={h1 l (n) ,…,h k l(n)},y j l(n) represents the j-th output feature map of the l-th layer, u j , σ j 2 Respectively represent h j l The mean and variance of , ε represents a small constant used to prevent invalid calculations when the variance is 0, They represent the scale and translation parameters that need to be learned respectively;

[0094] The Gaussian error linear unit GELU is shown in the following formula:

[0095]

[0096] The feedforward network FFN includes two layers of linear transformation and intermediate nonlinear activation, as shown in the following formula:

[0097] FFN(x)=W2σ(W1x+b1)+b2 (12)

[0098] Among them, x represents the input, W1 and W2 represent the weights of the linear transformation of the first and second layers respectively, b1 and b2 represent the bias of the linear transformation of the first and second layers respectively, and σ represents the intermediate nonlinear activation function;

[0099] In the output layer, first, global average pooling (GAP) is used to reduce the feature dimension in the time dimension. Second, point-by-point convolution (PConv) is used to integrate features in the channel dimension. Then, batch normalization (BN) and Gaussian error linear unit (GELU) are used to stabilize feature distribution and perform nonlinear mapping. Finally, the high-dimensional features are mapped to the classification dimension of health status through full connection (FC).

[0100] In this embodiment, the detailed parameter configuration of the SEFormer model is shown in Table 2 below:

[0101]

[0102]

[0103] Note: m represents the number of sensor channels for signal acquisition, d represents the output channel dimension, k l represents the convolution kernel size of SMDC, s represents the convolution stride, p represents the convolution padding, r represents the scaling factor of FFN, and d represents the convolution kernel size of SMDC. g represents the output time dimension of GAP, k represents the convolution kernel size of PConv, and d lrepresents the output time dimension of FC, and C represents the number of health status categories.

[0104] In this embodiment, m=2, C=9.

[0105] In this embodiment, in step 4, the training of the SEFormer model adopts a multi-classification cross entropy loss function to calculate the training loss, uses the AdamW optimization algorithm to update the model parameters, and dynamically adjusts the learning rate using an adaptive decay strategy based on the loss of the validation set.

[0106] In this embodiment, the training hyperparameters of the SEFormer model are configured as follows: the initial learning rate of the AdamW optimization algorithm is 0.001, the batch size is 64, and the number of iterations is 100.

[0107] In this embodiment, seven existing advanced models are selected for comparative analysis with the SEFormer model, including three end-to-end fault diagnosis models based on CNN-Transformer, namely LiConvFormer, Convformer-NSE and CLFormer, and four popular CNN models, namely MobileNetV2, MobileNet, MK-ResCNN and ResNet18. It is worth noting that for fair comparison, all the above models use the same input signal length, training strategy and training hyperparameter configuration. In order to reduce the influence of randomness, each experiment is repeated five times, and the mean value Mean and standard deviation Std of the diagnostic accuracy Accuracy are used as the final experimental results for analysis. In addition, the parameter quantity Params and the number of floating-point operations FLOPs are used as evaluation indicators of the model computational complexity. The diagnostic accuracy and model computational complexity of all models under different noise levels are shown in Table 3 below:

[0108]

[0109]

[0110] As can be seen from the table, when no noise is added, the average accuracy of MK-ResCNN and ResNet18 is slightly higher than that of SEFormer. As the noise level increases, the robustness advantage of SEFormer becomes significant. SEFormer performs well under all three noise levels. Specifically, when σ=0.2 and σ=0.6, SEFormer achieves the highest average accuracy (96.63% and 72.72%) with a small standard deviation (0.89% and 0.59%). When σ=0.4, the average accuracy of SEFormer is slightly lower than that of LiConvFormer, but the standard deviation is smaller than that of LiConvFormer. Considering the overall performance, SEFormer is comparable to LiConvFormer in diagnostic performance, but has a more obvious advantage over LiConvFormer in model computational complexity. Although the model computational complexity of Convformer-NSE and CLFormer is relatively low, their diagnostic performance lags significantly behind other models.

[0111] It can be seen from the above embodiments that the present invention is a lightweight rotating machinery fault diagnosis method based on SEFormer. Compared with the prior art, the advantages of the present invention are: (1) a separable multi-scale deep convolution SMDC is designed to extract and integrate multi-scale features from different channel dimensions of the vibration signal; (2) an efficient self-attention ESA is developed to capture the key fine-grained features of the vibration signal from a global perspective; (3) based on the separable multi-scale deep convolution SMDC and the efficient self-attention ESA, a SEFormer model with robustness, generalization ability and lightweight is constructed for rotating machinery fault diagnosis.

[0112] The above description is only a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent change made based on the technical essence of the present invention still falls within the scope of protection required by the present invention.

Claims

1. A lightweight rotating machinery fault diagnosis method based on SEFormer, characterized by: The steps include: Step 1: Data collection: Collect vibration signals of key transmission components of rotating machinery under various health conditions; Install at least one sensor at different positions or directions near the rotating machinery transmission parts that need to be diagnosed, and use data acquisition equipment to collect vibration signals under various health conditions; Step 2: Data preprocessing: The vibration signal is divided into sample sets through sliding window sampling, and then divided into training set, validation set and test set; Step 3: Model building: Design and develop separable multi-scale deep convolution SMDC and efficient self-attention ESA, and build SEFormer model; In step 3, the SEFormer model is composed of multiple feature extraction layers and an output layer connected in sequence, and the number of feature extraction layers is adjusted according to different task requirements; In the feature extraction layer; Firstly, separable multi-scale deep convolution (SMDC) is used to extract and integrate multi-scale features from different channel dimensions of vibration signals. Secondly, efficient self-attention (ESA) is used to capture the key fine-grained features of vibration signals from a global perspective. Then, residual connection (Add) is used to reduce the risk of overfitting, and batch normalization (BN) is used to stabilize the feature distribution. Next, the feedforward network FFN is used to perform nonlinear transformation on the features; Finally, the extracted features are output through residual connection Add and batch normalization BN; The separable multi-scale deep convolution SMDC; Firstly, parallel multi-scale deep convolution with different kernel sizes is used to extract the multi-local receptive field features of the vibration signal. Then, the multi-scale features are concatenated along the channel dimension. Next, point-wise convolution is used to integrate feature information and capture the correlation between channels; Finally, batch normalization BN and Gaussian error linear unit GELU are used to stabilize feature distribution and perform nonlinear mapping, as shown in the following formula: in, represents the input, C1 and L1 represent the channel dimension and time dimension of the input respectively, k l represents the size of the depth convolution kernel, represents the depth convolution kernel k l The weight of represents the depth convolution kernel k l The output, represents the output of multi-scale depth convolution, represents the weight of the point-by-point convolution kernel, represents the output, n represents the number of different depth convolution kernel sizes, and L2 represents the time dimension of the output; The efficient self-attention ESA; First, in order to expand the representation capacity of each feature and reduce computation, three parallel depth-wise separable convolutions DSConv are used to generate the input features of self-attention: query matrix Q, key matrix K, and value matrix V; Secondly, efficient attention is used to transfer global feature information while reducing the cost of self-attention calculation; Then, Softmax is used as a normalization operation suitable for efficient attention. Softmax normalization is performed on the matrix Q in the time dimension and on the matrix K in the channel dimension. Finally, the attention weight is calculated through matrix transposition and matrix multiplication operations, as shown in the following formula: in, and denote the query matrix, key matrix, and value matrix respectively, L and C denote the time dimension and channel dimension respectively, (·) T and represent matrix transposition and matrix multiplication respectively, Softmax L (·) and Softmax C (·) represents the Softmax normalization in the time dimension and channel dimension respectively; The residual connection Add is shown in the following formula: x'=f(x)+x (5) The batch normalization BN is shown in the following formula: Among them, h l ={h l(1) ,…,h l(N) } represents the input feature map of the lth layer with a batch size of N, h l(n) ={h1 l(n) ,…,h k l (n) },y j l(n) represents the j-th output feature map of the l-th layer, u j , σ j 2 Respectively represent h j l The mean and variance of , ε represents a small constant used to prevent invalid calculations when the variance is 0, They represent the scale and translation parameters that need to be learned respectively; The Gaussian error linear unit GELU is shown in the following formula: The feedforward network FFN includes two layers of linear transformation and intermediate nonlinear activation, as shown in the following formula: FFN(x)=W2σ(W1x+b1)+b2 (11) Among them, x represents the input, W1 and W2 represent the weights of the linear transformation of the first layer and the second layer respectively, b1 and b2 represent the bias of the linear transformation of the first layer and the second layer respectively, and σ represents the intermediate nonlinear activation function; In the output layer; First, global average pooling (GAP) is used to reduce the feature dimension in the time dimension; Secondly, point-by-point convolution PConv is used to integrate features in the channel dimension; Then, batch normalization BN and Gaussian error linear unit GELU are used to stabilize feature distribution and perform nonlinear mapping; Finally, the high-dimensional features are mapped to the classification dimension of health status through fully connected FC; Step 4: Model training, verification and evaluation: Train the model on the training set and validation set, and evaluate the model performance on the test set; Step 5: Fault diagnosis: Use the trained model to perform fault diagnosis on rotating machinery.

2. The SEFormer-based lightweight rotating machinery fault diagnosis method according to claim 1, characterized in that: In the second step, the preprocessing process is to avoid test leakage. When the sliding window is sampled, there is no overlap between adjacent windows and a certain interval is maintained. Each sample contains 1024 signal points. All samples are standardized by mean-standard deviation.

3. The SEFormer-based lightweight rotating machinery fault diagnosis method according to claim 1, characterized in that: In step 4, the training of the SEFormer model uses a multi-classification cross entropy loss function to calculate the training loss, uses the AdamW optimization algorithm to update the model parameters, and uses an adaptive decay strategy to dynamically adjust the learning rate according to the loss of the validation set.

Citation Information

Patent Citations

  • Rotating machinery fault diagnosis method, system, device and storage medium

    CN113505830B

  • Lightweight rotating machine fault diagnosis method based on multi-scale information fusion

    CN117113170A