TCN-SVM rolling bearing fault diagnosis method fusing SE attention mechanism
By embedding the SE attention mechanism in TCN and using SVM classifiers, the problem of insufficient feature redundancy and generalization in bearing fault diagnosis is solved, and higher fault recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510273374.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional timing convolutional networks (TCNs) have problems such as feature redundancy, noise sensitivity and insufficient generalization of Softmax classifiers in bearing fault diagnosis, resulting in reduced feature extraction capabilities and limited classification performance.
The TCN-SVM method that integrates the SE attention mechanism is adopted to optimize the feature channel weight allocation by embedding SE modules in TCN and replacing the Softmax classifier with SVM, improving feature extraction capabilities and classification accuracy.
It significantly improves the accuracy and robustness of rolling bearing fault diagnosis, enhances the dynamic weight allocation capability of key timing characteristics, reduces noise interference, and improves the generalization performance under small samples.
Smart Images

Figure CN120257067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a fault identification method, and more particularly to a rolling bearing fault diagnosis method of TCN-SVM integrating SE attention mechanism. Background Art
[0002] Bearings are one of the most critical components in rotating machinery, and their operating status directly affects the stability and lifespan of equipment. However, during long-term operation under complex and variable conditions, bearings are prone to defects or faults. If not detected in time, it is very likely to cause bearing function failure and trigger equipment faults, and even lead to casualties. Therefore, developing a precise and effective bearing fault diagnosis method is of great significance for improving the safety and reliability of equipment.
[0003] In recent years, fault diagnosis methods based on deep learning have made remarkable progress in the field of mechanical condition monitoring. The Temporal Convolutional Network (TCN), with its dilated convolution structure and residual connections, can effectively capture the long-term temporal dependencies of vibration signals, avoiding the gradient vanishing problem of recurrent neural networks, and is widely used in bearing fault diagnosis. However, traditional TCN models have inherent defects: on the one hand, as the network depth increases, the model generates a large number of redundant feature channels, resulting in a decline in feature expression ability; on the other hand, the TCN has insufficient ability to dynamically allocate weights for key fault features and is vulnerable to interference from irrelevant components under strong noise interference, affecting the pertinence of feature extraction; in addition, traditional TCN mostly uses the Softmax classifier, which is prone to overfitting and limited generalization performance when the training samples are insufficient. To address the above problems, there is an urgent need for a composite fault diagnosis method that can adaptively enhance key temporal features, suppress noise interference, and have efficient classification ability. Therefore, the present invention discloses a rolling bearing fault diagnosis method of TCN-SVM integrating SE attention mechanism. Summary of the Invention
[0004] To solve the classification problem of bearing faults, the present invention proposes a rolling bearing fault diagnosis method of TCN-SVM integrating SE attention mechanism, which optimizes the dynamic weight allocation of feature channels by deeply integrating the SE attention mechanism and TCN; combines SVM to replace the Softmax classifier, and utilizes its structural risk minimization characteristic to improve the generalization performance under small samples, thereby improving the accuracy and robustness of rolling bearing fault diagnosis.
[0005] A rolling bearing fault diagnosis method of TCN-SVM integrating SE attention mechanism provided by the present invention is characterized by including the following steps:
[0006] S1: Collect the vibration period signals x(t) of normal and various fault rolling bearings, where t is the sampling time;
[0007] S2: Construct a sample set containing normal states and multiple fault types, and preprocess the original signal;
[0008] S3: Construct an SE-TCN feature extraction network, including an SE module and a temporal convolutional network (TCN);
[0009] S4: Construct an SVM classifier decision layer to replace the Softmax classifier, and realize fault mode recognition through feature space mapping;
[0010] S5: Evaluate the classification effect of the model using model accuracy, T-SNE model visualization, and confusion matrix.
[0011] In one embodiment of the present invention, in step S2, the construction of the sample set containing normal states and multiple fault types includes the following steps:
[0012] S201: Random sampling. Assume that there are a total of L = 10 normal states and different types of faults. For each fault type, randomly select multiple sample segments of length N from the collected signal x(t), and the length of each sample segment is N. The i-th sample segment can be expressed as:
[0013] x i (t) = [x1, x2,..., x N
[0014] where i = 1, 2,..., N, x1 is the starting point of the signal segment selected by the random number generator, and N is the sample length.
[0015] S202: Generate a sample set. Collect M groups of samples for each fault type. Each group of samples consists of randomly selected signal segments. The total number of samples under each fault type is expressed as:
[0016]
[0017] S203: Label conversion. Ten types of bearing fault types (including normal state, 3 types of inner ring faults, 3 types of outer ring faults, and 3 types of rolling element faults) each correspond to an independent label, and are used together with the input features during model training. Each group of signal samples x i (t) corresponds to a sample fault type label y i , which is used for subsequent supervised learning.
[0018] where the sample fault type label y i ∈(0, 1, 2,...., L).
[0019] S204: Dataset division. To conduct the training, validation, and testing of the model, the constructed sample set is divided into a training set, a validation set, and a testing set according to a certain ratio.
[0020] S205: Fourier transform. To further extract fault features for subsequent model training, the fast Fourier transform (FFT) is used to perform spectral analysis on the signals. The FFT transform is carried out on each sample signal to obtain its frequency-domain sequence X(f):
[0021]
[0022] where j is the imaginary unit (j 2 = -1), N is the signal length, f is the sampling frequency, set to 48000 Hz, t is the time, and n = 0, 1, 2,..., N - 1;
[0023] S206: Calculate the spectrum P(f) of the signal x(t) using the frequency-domain sequence X(f) after Fourier transform:
[0024]
[0025] where N is the signal length;
[0026] In an embodiment of the present invention, in step S3, the construction of the SE-TCN network includes the following steps:
[0027] S301: Design the network structure of SE-TCN, which consists of two temporal convolutional network modules and an SE module. The temporal convolutional network module adopts a double-layer stacked structure with the number of layers D = 2, including a convolutional kernel, layer normalization, and a Dropout layer. Each TCN block contains two dilated causal convolutional layers, and each layer contains 32 convolutional kernels of size 2×2. The dilation factor of the first layer is 1, and the dilation factor of the second layer is 2, forming an exponentially increasing temporal perception ability;
[0028] S302: TCN temporal convolutional layer. A double-layer feature extraction module is constructed using dilated causal convolution. The mathematical expression for the causal convolution is as follows:
[0029]
[0030] where a = 0, 1,..., k - 1, w a is the weight of the convolutional kernel, r is the time step, k is the convolutional kernel size, x(t - d·a) is the value of the input sequence at time step t - d·a, and y(t) is the output at time step t;
[0031] It should be noted that w aIt is randomly assigned an initial value during the initialization of network training. The Adam optimization algorithm is used to update the weights according to the calculated gradients. The formula for weight update is as follows:
[0032]
[0033] where is the updated weight, is the weight before update, η is the learning rate, η ∈ [10 -4 , 10 -2 , and the default value of η is 0.001. E is the cross-entropy loss function;
[0034] The cross-entropy loss function is as follows:
[0035]
[0036] where y i is the type label of the sample, and p(y i ) is the probability of predicting the fault type of the sample;
[0037] By introducing the dilation factor d, the dilated convolution sets the dilation factor in an exponentially increasing manner, expanding the receptive field of the convolutional kernel and enabling the model to capture longer-range dependencies. The value of the dilation factor d is as follows:
[0038] d = 2 D-1
[0039] where D is the number of layers of the network. A double-layer stacked structure is adopted, D = 2. Then the dilation factor of the first-layer network is 1, and the dilation factor of the second-layer network is 2;
[0040] S303: Embed the SE (Squeeze-and-Excitation) attention module between the residual modules of the TCN, which mainly includes the squeeze operation, excitation operation, and feature reweighting. The specific steps for embedding the SE module are as follows:
[0041] Denote the input signal data as a tensor with a shape of H × N × C, where N represents the length of the signal, H represents the width of the signal, and C represents the number of channels of the input data.
[0042] The squeeze operation, that is, perform global pooling on the input features, and synthesize the input features with a shape of H × N × C into a feature description z C :
[0043]
[0044] where x C(m,n) represents the eigenvalue of the C-th channel at the position (m,n), z c is the compressed vector; H is the width of the signal, N is the length of the signal, and C represents the number of channels of the input data;
[0045] The excitation operation, that is, to obtain a more comprehensive channel-level dependence, satisfying the ability of flexibility and being able to learn non-mutually exclusive emphasis, the channel weight calculation s C :
[0046] s C =σ(W2δ(W1z C ))
[0047] where z C is the feature description with the number of channels C, W1 and W2 are the weight matrices of the fully connected layers, r is the compression ratio, used to reduce the computational complexity and the number of parameters of the network, generally taking r = 8, δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function, which normalizes the weights to (0,1);
[0048] Feature reweighting:
[0049]
[0050] where s C is the generated channel weight, x C is the original feature, is the weighted feature;
[0051] In an embodiment of the present invention, in step S4, the construction of the SVM classifier decision layer realizes the fault mode recognition through feature space mapping, including the following steps:
[0052] S401: Extract the feature vector f o , f u ∈R D ;
[0053] S402: Use the RBF kernel function to construct a linear classification hyperplane, map the samples to a higher-dimensional space, and the RBF kernel function K(f o , f u ) is specifically:
[0054] K(f o , f u ) = exp(-γ‖f o -f u ‖ 2 )
[0055] where f o is the feature vector of the training set sample, fu is the feature vector of the test set samples, and γ is the kernel function coefficient that controls the Gaussian kernel width;
[0056] It should be noted that in the formula σ is the standard deviation of the Gaussian distribution of the sample features, and γ = 5000 is taken;
[0057] S403: Find the hyperplane by solving the optimal classification model through structural risk minimization, so as to obtain the classification result. The optimization objective of SVM can be expressed as:
[0058]
[0059] where p is the normal vector of the hyperplane, which is the variable to be solved, C is the penalty coefficient used to control the penalty degree for misclassified samples, and C = 0.01 is taken, ζ i is the slack variable of the i-th sample, representing the distance from the sample to the hyperplane, which is automatically optimized by the algorithm. b is the bias term of the hyperplane, determining the offset of the hyperplane, and is the variable to be solved;
[0060] It should be noted that is the norm of the normal vector of the hyperplane, representing the complexity of the hyperplane. Minimizing this term can maximize the margin of the hyperplane, while is the penalty term for misclassified samples. If the sample is misclassified, ξ i > 0, this term will increase the value of the objective function and make the sample be correctly classified; if the sample is correctly classified, then ξ i = 0;
[0061] Its constraint conditions are:
[0062] y i (p·f o +b)≥1 - ξ i
[0063] where y i is the sample label, p is the normal vector of the hyperplane, which is the variable to be solved, f o is the feature vector of the training set samples, b is the bias term of the hyperplane, and ζ i is the slack variable of the i-th sample;
[0064] Using the Lagrange multiplier method to combine the objective function and the constraint conditions, solve for p and b to obtain the optimal fault classification hyperplane, and use the classification decision function to classify the test samples, so as to obtain the SVM fault type classification result of the samples. The classification decision function is as follows:
[0065]
[0066] where, For the SVM fault classification result of the test sample, sign(·) is the sign function. If p·f u +b > 0, then it is the corresponding bearing fault type. If p·f u +b < 0, then it is not the corresponding bearing fault type, and f u is the feature vector of the test set sample. Description of the Drawings
[0067] Figure 1 is the flowchart of the TCN-SVM rolling bearing fault diagnosis method integrating the SE attention mechanism provided by the embodiment of the present invention;
[0068] Figure 2 is the structural diagram of the TCN-SVM rolling bearing fault diagnosis method integrating the SE attention mechanism provided by the embodiment of the present invention;
[0069] Figure 3 is the sample T-SNE distribution diagram of different models in the embodiment of the present invention;
[0070] Figure 4 is the confusion matrix diagram of different models in the embodiment of the present invention. Detailed Embodiment
[0071] To more clearly illustrate the technical solution of the present invention, the following will clearly and completely describe it in conjunction with the embodiments and their accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. For those of ordinary skill in the art, all other embodiments obtained without creative efforts fall within the scope of protection of the present invention.
[0072] Figure 1 is the structural diagram of a diesel engine multi-sensor multi-scale fusion feature extraction and fault warning method provided by this application, and its detailed process is as follows Figure 2 shown. Refer to Figure 1 , the implementation process and results of the present invention are as follows:
[0073] S1: Collect the vibration period signals x(t) of normal and various fault rolling bearings, where t is the sampling time;
[0074] S2: Construct a sample set including normal states and various fault types, and preprocess the original signal;
[0075] S3: Construct an SE-TCN feature extraction network, including an SE module and a temporal convolutional network (TCN);
[0076] S4: Construct an SVM classifier decision layer to replace the Softmax classifier, and realize fault pattern recognition through feature space mapping;
[0077] S5: Evaluate the classification effect of the model using model accuracy, T-SNE model visualization, and confusion matrix.
[0078] In one embodiment of the present invention, in step S2, the construction of a sample set including normal states and multiple fault types includes the following steps:
[0079] S201: Random sampling. Assume that there are a total of L = 10 types including normal states and different types of faults. For each fault type, randomly select multiple sample segments of length N from the collected signal x(t), and the length of each sample segment is N. The i-th sample segment can be expressed as:
[0080] x i (t) = [x1, x2,..., x N
[0081] where i = 1, 2,..., N, x1 is the starting point of the signal segment selected by the random number generator, and N is the sample length.
[0082] S202: Generate a sample set. Collect M groups of samples for each fault type. Each group of samples consists of randomly selected signal segments, and the total number of samples under each fault type is expressed as:
[0083]
[0084] S203: Label conversion. Ten types of bearing fault types (including normal state, 3 types of inner ring faults, 3 types of outer ring faults, and 3 types of rolling element faults) each correspond to an independent label and are used together with the input features during model training. Each group of signal samples x i (t) corresponds to a sample fault type label y i , for subsequent supervised learning.
[0085] where the sample fault type label y i ∈(0, 1, 2,...., L).
[0086] S204: Dataset division. To perform model training, validation, and testing, divide the constructed sample set into a training set, a validation set, and a test set according to a certain proportion.
[0087] S205: Fourier transform. To further extract fault features and use them for subsequent model training, perform spectral analysis on the signal using the fast Fourier transform (FFT), perform FFT transformation on each sample signal, and obtain its frequency domain sequence X(f):
[0088]
[0089] where j is the imaginary unit (j2 = -1), N is the signal length, f is the sampling frequency, set to 48000Hz, t is the time, n = 0, 1, 2, ..., N - 1;
[0090] S206: Calculate the spectrum P(f) of the signal x(t) using the frequency-domain sequence X(f) after Fourier transform:
[0091]
[0092] where N is the signal length;
[0093] In an embodiment of the present invention, in step S3, the construction of the SE-TCN network includes the following steps:
[0094] S301: Design the network structure of SE-TCN, which consists of two temporal convolutional network modules and an SE module. The temporal convolutional network module adopts a double-layer stacked structure, with the number of layers D = 2, including convolutional kernels, layer normalization, and Dropout layers. Each TCN block contains two dilated causal convolutional layers, and each layer contains 32 convolutional kernels of size 2×2. The dilation factor of the first layer is 1, and the dilation factor of the second layer is 2, forming an exponentially increasing temporal perception ability;
[0095] S302: The TCN temporal convolutional layer uses dilated causal convolution to construct a two-level feature extraction module, and the mathematical expression for the causal convolution is:
[0096]
[0097] where a = 0, 1, ..., k - 1, w a is the weight of the convolutional kernel, r is the time step, k is the size of the convolutional kernel, x(t - d·a) is the value of the input sequence at time step t - d·a, and y(t) is the output at time step t;
[0098] It should be noted that w a is randomly assigned an initial value during the initialization of network training, and the Adam optimization algorithm is used to update the weights according to the calculated gradients. The formula for weight update is:
[0099]
[0100] where, is the updated weight, is the weight before update, η is the learning rate, η ∈ [10 -4 , 10 -2 , and the default value of η is 0.001, and E is the cross-entropy loss function;
[0101] The cross-entropy loss function is as follows:
[0102]
[0103] Among them, y i is the type label of the sample, and p(y i ) is the probability of predicting the failure type of the sample;
[0104] By introducing the dilation factor d and setting the dilation factor in an exponentially increasing manner, dilated convolution expands the receptive field of the convolutional kernel, enabling the model to capture longer-range dependencies. The value of the dilation factor d is as follows:
[0105] d = 2 D-1
[0106] where D is the number of layers of the network, adopting a double-layer stacked structure, D = 2. Then the dilation factor of the first-layer network is 1, and the dilation factor of the second-layer network is 2;
[0107] S303: Embed the SE (Squeeze-and-Excitation) attention module between each residual module of the TCN, which mainly includes the squeeze operation, excitation operation, and feature reweighting. The specific steps for embedding the SE module are as follows:
[0108] Denote the input signal data as a tensor with a shape of H×N×C, where N represents the length of the signal, H represents the width of the signal, and C represents the number of channels of the input data.
[0109] The squeeze operation, that is, perform global pooling on the input features, and synthesize the input features with a shape of H×N×C into a feature description z with a shape of C×1×1 C :
[0110]
[0111] Among them, x C (m,n) represents the feature value at the position (m,n) of the Cth channel, and z c is the compressed vector; H is the width of the signal, N is the length of the signal, and C represents the number of channels of the input data;
[0112] The excitation operation, that is, obtain a more comprehensive channel-level dependency, meet the ability of flexibility and being able to learn non-mutually exclusive emphasis, and calculate the channel weight s C :
[0113] s C = σ(W2δ(W1z C ))
[0114] Among them, z CFor the feature description with the number of channels being C, W1 and W2 are the weight matrices of the fully connected layers. r is the compression ratio, used to reduce the computational complexity and the number of parameters of the network. Generally, r = 8 is taken. δ(·) is the ReLU activation function, and σ(·) is the Sigmoid activation function, which normalizes the weights to (0, 1).
[0115] Feature reweighting:
[0116]
[0117] Among them, s C is the generated channel weight, x C is the original feature, is the weighted feature;
[0118] In an embodiment of the present invention, in step S4, the construction of the SVM classifier decision layer realizes fault mode recognition through feature space mapping, including the following steps:
[0119] S401: Extract the feature vector f o , f u ∈R D ;
[0120] S402: Use the RBF kernel function to construct a linear classification hyperplane, mapping the samples to a higher-dimensional space. The RBF kernel function K(f o , f u ) is specifically:
[0121] K(f o , f u ) = exp(-γ‖f o -f u ‖ 2 )
[0122] Among them, f o is the feature vector of the training set samples, f u is the feature vector of the test set samples, and γ is the kernel function coefficient, controlling the Gaussian kernel width;
[0123] It should be noted that in the formula σ is the standard deviation of the Gaussian distribution of the sample features, and γ = 5000 is taken;
[0124] S403: Solve the optimal classification model through structural risk minimization to find the hyperplane, thereby obtaining the classification result. The optimization objective of SVM can be expressed as:
[0125]
[0126] Among them, p is the normal vector of the hyperplane, is the variable to be solved, C is the penalty coefficient, used to control the penalty degree for misclassified samples, taking C = 0.01, ζ i is the slack variable of the i-th sample, representing the distance from the sample to the hyperplane, which is automatically optimized by the algorithm. b is the bias term of the hyperplane, determining the offset of the hyperplane, and is the variable to be solved;
[0127] It should be noted that is the norm of the hyperplane normal vector, representing the complexity of the hyperplane. Minimizing this term can maximize the margin of the hyperplane, while is the penalty term for misclassified samples. If the sample is misclassified, ξ i > 0, this term will increase the value of the objective function and make the sample be correctly classified; if the sample is correctly classified, then ξ i = 0;
[0128] Its constraint conditions are:
[0129] y i (p·f o +b)≥1 - ξ i
[0130] Among them, y i is the sample label, p is the normal vector of the hyperplane, is the variable to be solved, f o is the feature vector of the training set sample, b is the bias term of the hyperplane, ζ i is the slack variable of the i-th sample;
[0131] Using the Lagrange multiplier method to combine the objective function and the constraint conditions, solve for p and b to obtain the optimal fault classification hyperplane, and use the classification decision function to classify the test samples, so as to obtain the SVM fault type classification result of the samples. The classification decision function is as follows:
[0132]
[0133] Among them, is the SVM fault classification result of the test sample, sign(·) is the sign function. If p·f u +b > 0, then it is the corresponding bearing fault type. If p·f u +b < 0, then it is not the corresponding bearing fault type, f u is the feature vector of the test set sample.
[0134] To verify the effectiveness of the method proposed in the present invention, the acceleration data of the motor drive end of the SKF 6205-2RS deep groove ball bearing from Case Western Reserve University was used. In this paper, data with a rotational speed of 1797 r / min, a load of 0 Nm, and a sampling frequency of 48 kHz was loaded. The single-point diameter damages of the selected bearings were 0.007 mm, 0.014 mm, and 0.021 mm respectively. Each fault diameter contained three types of bearing damages, namely rolling element faults (sample numbers 1, 2, 3), inner race faults (sample numbers 4, 5, 6), and outer race faults (sample numbers 7, 8, 9). In addition, there was also a normal state (sample number 10), which together formed a sample data set of 10 fault types. The data set parameters and the sample numbers of each fault are shown in Table 1.
[0135] Table 1 Experimental parameters of the CWRU bearing data set
[0136]
[0137]
[0138] The fault signals of rolling bearings were trained using the method of the present invention. Since the sample size of the test object of the present invention is 10×200×512, in order to facilitate model input, the samples of each fault were merged. The size of the merged samples was 2000×512, and then they were transformed into single-sided spectrum features of 2000×513 through FFT. Then, the training set, validation set, and test set were divided according to the ratio of 7:2:1.
[0139] The key parameters in the SE-TCN model input during the training phase are shown in Table 2. The TCN module adopts a double-layer stacked structure. Each TCN block contains two 1D convolutional layers with causal convolution characteristics. Each layer contains 32 convolutional kernels of size 2×2. The dilation factor of the first layer is 1, and the dilation factor of the second layer is 2, forming an exponentially growing temporal perception ability. The SE module compresses spatio-temporal features through global average pooling, adopts a channel attention mechanism with a compression ratio of 16, and realizes feature recalibration through two fully connected layers. During model training, 80 rounds of iteration were used, Dropout was set to 0.1, the Adam optimizer with an initial learning rate of 0.001 was used, the performance of the validation set was evaluated every 8 batches, and the generalization ability was enhanced through the data shuffling strategy for each epoch.
[0140] Table 2 Key parameters of the SE-TCN network model
[0141]
[0142] In the present invention, an SVM classifier is adopted to replace the original Softmax classifier, which receives the features extracted by the TCN as input, and realizes end-to-end fault classification through the SVM classifier in the test phase. Its kernel function is a linear kernel function, with the penalty coefficient C = 0.01 and γ = 5000.
[0143] To verify its effectiveness, this paper makes a comparison with TCN and SE-TCN, and the experimental results are shown in Table 3.
[0144] Table 3 Performance comparison of different models
[0145]
[0146] It can be seen that this model has improved by 12.01% and 2.04% respectively compared with the other two unimproved methods, and the training time is reduced by 1.92 s compared with SE-TCN.
[0147] To more intuitively analyze the fault classification effect of the model, we use the T-SNE method to visualize the model results, as Figure 3 shown.
[0148] From Figure 3 it can be seen in (a) that there is a large overlap between data points of different categories, indicating that the distribution of the original data in the feature space is relatively mixed, and it is easy to cause classification confusion when directly used for classification. Figure 3 It can be observed from (b) that the TCN has improved the separation degree of different category samples to a certain extent, but there are still some cross-regions between some categories, especially there are still many overlapping regions between the inner raceway fault and the rolling element fault, indicating that the TCN still has certain limitations in feature extraction. Figure 3 (c) shows the T-SNE distribution after the SE-TCN-SVM model extracts features. Compared with Figure 3 (a) and Figure 3 (b), it can be clearly seen that the sample points of different categories are more concentrated, and the boundaries between them are clearer, indicating that the SE-TCN-SVM effectively enhances the feature extraction ability through the SE attention mechanism, enabling the data of different categories to be better separated in the feature space and improving the classification performance.
[0149] To further analyze the classification performance of the model, this paper draws the confusion matrix of the test set, as Figure 4 shown.
[0150] From Figure 4(a) It can be seen from the confusion matrix of TCN that the performance of this model in the classification task is inferior to that of SE-TCN and SE-TCN-SVM, with more misclassifications. There are significant confusions especially among inner race faults, rolling element faults, and outer race faults. This indicates that there are still certain deficiencies in the feature extraction ability of TCN, and it fails to fully extract the discriminative information between different categories, resulting in lower classification performance. Figure 4 (b) The confusion matrix of SE-TCN shows that the model has a high classification accuracy for most categories, but there are still some misclassification cases, mainly manifested as confusions between some adjacent categories (such as rolling element faults and outer race faults), indicating that it does not combine SVM for classification optimization, resulting in a decrease in the discrimination ability for some categories with similar features. From Figure 4 (c), it can be seen that the SE-TCN-SVM model has the highest classification accuracy, and almost all categories can be correctly identified, with a very low misclassification rate. This indicates that this method has high generalization ability and stability in the classification task of different bearing fault types.
[0151] Generally speaking, SE-TCN-SVM improves the feature extraction ability through the SE module and combines the SVM classifier for optimization, thus significantly improving the classification accuracy and having obvious advantages compared with SE-TCN and TCN.
[0152] The above is only the preferred embodiment of the present application and the description of the applied technical principles. Those skilled in the art should understand that the scope involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept, such as the technical solutions formed by the mutual replacement of the above features and the technical features (but not limited to) with similar functions disclosed in the present application.
[0153] Except for the technical features described in the specification, the remaining technical features are known to those skilled in the art. To highlight the innovative features of the present invention, the remaining technical features are not described herein again.
Claims
1. A TCN-SVM rolling bearing fault diagnosis method integrating SE attention mechanism, characterized in that, It includes the following steps: S1: Collect the vibration period signals x(t) of normal and various faulty rolling bearings, where t is the sampling time; S2: Construct a sample set including normal states and multiple fault types, and preprocess the original signals; S3: Construct a SE-TCN feature extraction network, including a SE module and a temporal convolutional network (TCN); S4: Construct a SVM classifier decision layer to replace the Softmax classifier, and realize fault mode recognition through feature space mapping.
2. The method according to claim 1, wherein in step S2, the construction of the sample set including normal states and multiple fault types includes the following steps: S201: Random sampling. Assume that there are a total of L = 10 normal states and different types of faults. For each fault type, randomly select multiple sample segments of length N from the collected signal x(t), and the length of each sample segment is N; the i-th sample segment is expressed as: x i (t) = [x1, x2,..., x N where i = 1, 2,..., N, x1 is the starting point of the signal segment selected by the random number generator, and N is the sample length; S202: Generate a sample set. Collect M groups of samples for each fault type; each group of samples consists of randomly selected signal segments, and the total number of samples under each fault type is expressed as: S203: Label conversion. The 10 types of bearing fault types include normal state, 3 types of inner ring faults, 3 types of outer ring faults, and 3 types of rolling element faults; each corresponds to an independent label and is used together with the input features during model training; each group of signal samples x i (t) corresponds to a sample fault type label y i ; among them, the sample fault type label y i ∈(0, 1, 2,...., L); S204: Dataset division. In order to train, validate, and test the model, divide the constructed sample set into a training set, a validation set, and a test set according to a certain proportion; S205: Fourier transform. In order to further extract fault features and use them for subsequent model training, use the fast Fourier transform (FFT) to perform spectral analysis on the signal, perform FFT transformation on each sample signal, and obtain its frequency domain sequence X(f): where j is the imaginary unit (j 2 = -1), N is the signal length, f is the sampling frequency, set to 48000 Hz, t is the time, and n = 0, 1, 2, ..., N - 1; S206: Calculate the spectrum P(f) of the signal x(t) using the Fourier-transformed frequency domain sequence X(f): where N is the signal length.
3. The method according to claim 1, wherein in step S3, the construction of the SE-TCN network includes the following steps: S301: Design the network structure of SE-TCN, which consists of two temporal convolutional network modules and a SE module; the temporal convolutional network module adopts a double-layer stacked structure, with the number of layers D = 2, including convolutional kernels, layer normalization, and Dropout layers. Each TCN block contains two dilated causal convolutional layers, and each layer contains 32 convolutional kernels of size 2×2. The dilation factor of the first layer is 1, and the dilation factor of the second layer is 2, forming an exponentially growing temporal perception ability; S302: TCN temporal convolutional layer. Use dilated causal convolution to construct a two-level feature extraction module, where the mathematical expression for causal convolution is: where a = 0, 1, ..., k-1, w a is the weight of the convolution kernel, r is the time step, k is the convolution kernel size, x(t-d·a) is the value of the input sequence at time step t-d·a, and y(t) is the output at time step t; w a It is randomly assigned an initial value during the initialization of network training, and the Adam optimization algorithm is used to update the weights according to the calculated gradients. The formula for weight update is as follows: Among them, is the updated weight, is the weight before update, η is the learning rate, η ∈ [10 -4 , 10 -2 , by default η = 0.001, and E is the cross-entropy loss function; The cross-entropy loss function is as follows: where y i is the type label of the sample, and p(y i ) is the probability of predicting the fault type of the sample; The dilated convolution sets the dilation factor in an exponentially growing manner by introducing the dilation factor d, expanding the receptive field of the convolutional kernel, enabling the model to capture longer-range dependencies. The value of the dilation factor d is as follows: d=2 D-1 where D is the number of layers of the network. Adopting a double-layer stacked structure, D = 2, then the dilation factor of the first-layer network is 1, and the dilation factor of the second-layer network is 2; S303: Embed the SE (Squeeze-and-Excitation) attention module between each residual module of the TCN, including the squeeze operation, excitation operation, and feature reweighting. The specific steps for embedding the SE module are as follows: Denote the input signal data as a tensor with a shape of H×N×C, where N represents the length of the signal, H represents the width of the signal, and C represents the number of channels of the input data; Compression operation, that is, global pooling is performed on the input features, and the input features with the shape of H×N×C are synthesized into a feature description z with the shape of C×1×1 C : where x C (m,n) represents the eigenvalue of the C-th channel at the position (m,n), z c is the compressed vector; H is the width of the signal, N is the length of the signal, and C represents the number of channels of the input data; Activation operation, that is, obtaining a more comprehensive channel-level dependence, satisfying the ability of flexibility and being able to learn non-mutually exclusive emphasis, channel weight calculation s C : s C = σ(W2δ(W1z C )) where z C is the feature description with the number of channels being C, and W1 and W2 are the weight matrices of the fully connected layers. r is the compression ratio, taking r = 8, δ(·) is the ReLU activation function, σ(·) is the Sigmoid activation function, and the weights are normalized to (0, 1). Feature reweighting: Among them, s C is the generated channel weight, x C is the original feature, and is the weighted feature.
4. The method according to claim 1, wherein In step S4, the construction of the decision layer of the SVM classifier realizes fault mode recognition through feature space mapping, including the following steps: S401: Extract the feature vector f output by the SE-TCN network o , f u ∈R D ; S402: Construct a linear classification hyperplane using the RBF kernel function, map the samples to a higher-dimensional space, and the RBF kernel function K(f o , f u ) is specifically as follows: K(f o ,f u ) = exp(-γ‖f o -f u ‖ 2 ) Among them, f o is the feature vector of the training set samples, and f u is the feature vector of the test set samples. γ is the kernel function coefficient that controls the Gaussian kernel width; In the formula σ is the standard deviation of the Gaussian distribution of the sample features, and γ = 5000 is taken; S403: Find the hyperplane by solving the optimal classification model through structural risk minimization to obtain the classification result. The optimization objective of the SVM is expressed as: Among them, p is the normal vector of the hyperplane, is the variable to be solved, C is the penalty coefficient, which is used to control the penalty degree for misclassified samples, and C = 0.01 is taken, ζ i is the slack variable of the i-th sample, representing the distance from the sample to the hyperplane, b is the bias term of the hyperplane, which determines the offset of the hyperplane, and is the variable to be solved; is the norm of the hyperplane normal vector, representing the complexity of the hyperplane; is the penalty term for misclassified samples. If a sample is misclassified, ξ i > 0, and this term will increase the value of the objective function to correctly classify the sample; if the sample is correctly classified, then ξ i = 0; The constraint conditions are: y i (p·f o +b)≥1-ξ i Among them, y i is the sample label, p is the normal vector of the hyperplane, which is the variable to be solved, f o is the feature vector of the training set sample, b is the bias term of the hyperplane, ζ i is the slack variable of the i-th sample; Using the Lagrange multiplier method to combine the objective function and the constraint conditions, solve for p and b to obtain the optimal fault classification hyperplane, and use the classification decision function to classify the test samples, thereby obtaining the SVM fault type classification result of the samples. The classification decision function is as follows: Among them, is the SVM fault classification result of the test sample, sign(·) is the sign function. If p·f u +b > 0, then it is the corresponding bearing fault type. If p·f u +b < 0, then it is not the corresponding bearing fault type. f u is the feature vector of the test set sample.
Citation Information
Cited By
Multi-source information fusion fault diagnosis method for BIT module filter circuit
CN121093161A
State recognition model training method, state recognition method, device and equipment
CN121542725A
Lightweight TCN broadband oscillation suppression method for electric excitation system
CN122246702A