Rotating machinery fault diagnosis method based on residual shrinkage convolution and attention mechanism
By using the multi-channel residual shrinkage convolution unit designed by Meta-ACON and the global second-order pooling attention mechanism in rotary mechanical fault diagnosis, the problem of difficult identification of effective fault characteristics in vibration signals is solved, and more accurate and robust fault diagnosis is achieved.
Patent Information
- Application Number
- CN202510027908.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the complex working environment of strong background noise and variable load, the effective fault characteristics of rotating mechanical vibration signals are difficult to identify. Traditional convolutional neural networks lack mechanisms to ignore invalid features and enhance effective feature similarity, resulting in insufficient feature extraction and insufficient model generalization capabilities.
The multi-channel residual shrinkage convolution unit designed based on Meta-ACON, combined with the global second-order pooling attention mechanism, the noise and irrelevant features are effectively filtered out through adaptive learning of the filter threshold and weighting coefficient, and the feature extraction capability is enhanced through second-order statistical information.
It improves the accuracy of feature extraction and the generalization ability of the model, enhances the fault recognition ability under low signal-to-noise ratio and variable load conditions, and improves the accuracy and robustness of fault diagnosis.
Smart Images

Figure CN119475191B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rotating machinery fault diagnosis, and in particular to a rotating machinery fault diagnosis method based on residual shrinkage convolution and attention mechanism. Background Art
[0002] As key equipment in industrial production, rotating machinery is widely used in various manufacturing and processing industries, and its position in the modern industrial system is indispensable. However, rotating machinery frequently fails during operation, especially the wear, fatigue damage and imbalance of mechanical parts, which has become the focus of continuous attention in the engineering community. These failures not only affect production efficiency, but may also cause serious safety accidents and lead to huge economic losses. Therefore, timely and accurate diagnosis of rotating machinery failures is crucial to ensure safe production and improve production efficiency.
[0003] Rotating machinery fault diagnosis based on vibration signals is to classify fault types by extracting effective feature information from fault signals. Feature extraction is the focus of fault diagnosis. Deep learning has been widely used in the field of fault diagnosis with its end-to-end automatic feature learning capabilities and high-order abstract modeling capabilities for signals, such as deep belief networks, one-dimensional convolutional neural networks (CNN), two-dimensional time-frequency convolutional neural networks, self-attention mechanisms, recurrent neural networks, generative adversarial networks, etc.
[0004] Due to the interference of strong background noise of rotating machinery, the instability of working load, the complexity of vibration transmission path and the intricate and strongly correlated coupling relationship within the components, the effective information of vibration signal is confused, the feature expression is different, which increases the difficulty of effective feature extraction and reduces the generalization ability of the model. Therefore, designing a diagnostic model that can extract more discriminative features under these interference factors has become a key challenge to improve the accuracy and robustness of fault diagnosis.
[0005] In recent years, researchers have actively explored solutions. The wide-core deep convolutional neural network model WDCNN uses wide convolution kernels to suppress high-frequency noise. The MC1-DCNN model uses a multi-scale approach to learn effective features from dynamic process signals in the time and frequency domains. The ECNN model uses dilated convolution to expand the receptive field and comprehensively obtain local and global fault features of vibration signals. The MB-CNN model fuses multiple sensor data into two-dimensional image information and uses bottleneck layer convolutional neural networks to extract richer features. However, in a complex working environment with strong background noise and variable loads, the effective fault features of vibration signals are difficult to identify due to noise interference and the non-stationary characteristics of the signal. Some features in the signal represent fault information, while others are interference information. Traditional convolutional neural networks lack a mechanism to ignore invalid features and enhance the similarity of effective features, which makes the extraction of effective features insufficient, thereby reducing the generalization ability and diagnostic effect of the model.
[0006] As an important way to improve the effectiveness of feature extraction, the attention mechanism focuses on key features and adaptively adjusts the feature response value to improve the model's ability to capture correlation. For example, the MA1DCNN model uses multiple attentions to highlight discriminative features; the DAMN model uses a dual-attention multi-scale module to extract multi-scale and multi-level features; the DCA-BiGRU model uses dual channels with an attention mechanism to capture the spatial and channel relationships of the signal. The attention mechanism used by these methods extracts the first-order statistical features of the feature map, and uses the global mean or global maximum as the basis for adjusting the attention factor, which has the problem of insufficient modeling capabilities.
[0007] As an important way to improve the effectiveness of feature extraction, the attention mechanism focuses on key features and adaptively adjusts the feature response value to improve the model's ability to capture correlations. Researchers have proposed a multi-attention one-dimensional convolutional neural network (MA1DCNN) to highlight discriminative features; a dual-attention multi-scale one-dimensional CNN model network model has been proposed, which uses a dual-attention multi-scale module to extract multi-scale and multi-level features; and a bidirectional gated recurrent network model of the attention mechanism has been proposed to capture the spatial and channel relationships of the signal. The attention mechanism used in these methods extracts the first-order statistical features of the feature map, and uses the global mean or global maximum as the basis for adjusting the attention factor, which has the problem of insufficient modeling ability.
[0008] In addition, as an effective method to suppress noise, the deep residual shrinkage network (DRSN) combines soft threshold filtering technology with deep learning, uses the scaling nonlinear transformation of the channel average to obtain the threshold of each channel, and adaptively adjusts the threshold to set invalid features to zero. However, the single soft threshold filtering lacks a mechanism to highlight key features. Its threshold filter function and derivative relationship are as follows: Figure 1 As shown, Figure 1 (a) is the threshold filter function graph, and (b) is the derivative graph.
[0009] Its threshold filtering function is expressed as:
[0010]
[0011] in, τ is the filtering threshold. τ The input is simply set to zero (i.e. “clipped”), and the derivative of the output with respect to the input is either 1 or 0.
[0012] The above soft threshold filter function is as follows Figure 1 As shown in the figure, it is a hard clipping function that cannot provide a smooth transition when the input value is close to the threshold, and the gradient disappearance phenomenon is prone to occur. In addition, the soft threshold filter function is symmetrical in processing positive and negative inputs. When the input is greater than the threshold, it only "subtracts the threshold" without further "enhancement" and lacks a mechanism to highlight key features.
[0013] Therefore, the above method still has the problems of insufficient feature extraction and insufficient model generalization ability in complex working environments. Summary of the invention
[0014] The purpose of the present invention is to provide a rotating machinery fault diagnosis method combining an improved residual shrinkage convolution and a global second-order pooling attention mechanism, so as to solve the problems that the effective features of the vibration signals of rotating machinery are difficult to extract under strong background noise and variable load working environment, the fault diagnosis accuracy is low, and the generalization ability of the model is poor.
[0015] In order to achieve the above object, the technical solution adopted by the present invention is: a rotating machinery fault diagnosis method based on Meta-ACON residual shrinkage convolution and attention mechanism, using a trained fault diagnosis network model to perform fault detection on the rotating machinery, and the calculation process of the fault diagnosis network model includes:
[0016] S1, inputs the vibration signal of the rotating machinery, performs feature pre-extraction on the vibration signal through a wide convolution preprocessing layer, and outputs a feature map I;
[0017] S2, multi-channel and multi-scale feature extraction is performed on the output feature map I through multiple stacked multi-channel residual shrinkage convolution units. The multi-channel residual shrinkage convolution unit is composed of two channels with residual connections. Each channel includes two one-dimensional convolution layers connected in sequence and an activation layer ACON-FilterNet with a soft threshold filtering function. The output features of the two channels are connected to a 1×1 dimensional convolution after splicing to achieve feature fusion in the channel dimension. Finally, the spliced features are linearly superimposed with the residual connection to obtain the output feature map II;
[0018] S3, sends the output feature map II in S2 to the global second-order pooling module, extracts the second-order statistical information of the channel by calculating the covariance matrix of the feature map, obtains the attention factor of each channel after scaling nonlinear mapping, and finally multiplies the attention factor and the output feature map II channel by channel to complete the attention mechanism operation and output the operation result;
[0019] S4, performs global average pooling on the operation results in S3, and then classifies them through full connection to obtain the category to which the fault signal belongs.
[0020] Preferably, the multi-channel residual shrinkage convolution unit has an activation layer ACON-FilterNet with a soft threshold function, and its activation operation function Using Meta-ACON design, the expression is:
[0021]
[0022] In the formula, τ is the filtering threshold, β is the nonlinear scaling factor, p 1 is the positive dynamic weighting coefficient, p 2 is the negative dynamic weighting coefficient, x is the input feature.
[0023] In the above formula, Calculated by the following formula:
[0024]
[0025]
[0026] In the formula, p , p’ and β represents the learning parameter, σ represents the sigmoid activation function, and t represents the input amount.
[0027] Preferably, the filtering threshold τ and the forward dynamic weighting coefficient p 1 as a channel-level learning parameter, p 2 and β As hyperparameters or neuron parameters.
[0028] Preferably, the activation layer ACON-FilterNet implements the channel-level parameter filtering threshold τ and the forward dynamic weighting coefficient p 1, the specific model structure is designed as follows: first, the input feature map is globally averaged by absolute value pooling to obtain a one-dimensional vector 1× C , Cis the number of channels of the input feature map, and the one-dimensional vector is passed to the compressed fully connected network to obtain a compressed two-dimensional vector 2× C / r , r is the contraction factor, and then the final two-dimensional vector 2× C , where a vector is used as the forward dynamic weighting coefficient p 1, another vector is subjected to the Sigmoid function to obtain a threshold factor less than 1 and greater than zero, which is multiplied by the global absolute value average pooling statistic of each channel to obtain the filtering threshold of each channel τ .
[0029] Preferably, the operation process of the global second-order pooling module in S3 is: first, the output feature map II is reduced in dimension through 1×1 convolution, second-order pooling is performed through covariance operation, the correlation between channels is calculated, and the covariance matrix is obtained, and then row convolution nonlinear operation is performed on the covariance matrix to obtain the attention factor of each channel, and finally the attention factor and the output feature map II are multiplied channel by channel.
[0030] Furthermore, in the shrinkage convolution dual channel, a large receptive field convolution kernel is used to extract low-frequency features on the low-frequency channel, and a small convolution kernel is used to extract high-frequency features on the other high-frequency channel.
[0031] Furthermore, if the feature after dual-channel fusion is different from the input dimension, the residual connection is a direct connection mapping operation; if the feature after dual-channel fusion is the same as the input dimension, the residual connection is taken as an identity operation.
[0032] Furthermore, in S4, the global average pooling operation extracts the global average value of each channel, and after the full connection operation, the Softmax activation function is used to obtain the probability distribution value of the fault category, and the category with the largest probability value is the category to which the fault signal belongs.
[0033] By adopting the above technical solution, the present invention can achieve the following beneficial effects:
[0034] (1) Based on Meta-ACON, the present invention designs an activation function with soft threshold filtering function. The activation function learns key parameters such as filter threshold and weighting coefficient through adaptive methods, and constructs the activation function parameter learning soft threshold filtering module ACON-FilterNet. On this basis, combined with multi-channel and multi-scale convolution technology, a multi-channel residual shrinkage convolution unit is designed. The design can adaptively adjust the filter threshold and forward dynamic weighting coefficient, effectively filter out noise and irrelevant features, and enhance the extraction and expression capabilities of effective features. This effectively reduces the interference of irrelevant features on the diagnosis results and improves the accuracy of feature extraction.
[0035] (2) The present invention introduces the GSoP attention mechanism after the residual shrinkage convolution unit and utilizes the second-order statistical information of the high-level channel feature map to enhance the feature extraction capability of the model, so that the model can still maintain a high fault recognition capability under low signal-to-noise ratio and variable load conditions, thereby improving the accuracy and robustness of the model's fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the threshold filter function and derivative graph in the background technology;
[0037] Figure 2 A schematic diagram of the structure of a fault diagnosis model in an embodiment of the present invention;
[0038] Figure 3 Schematic diagram of the influence of p1, p2 and β on the activation function curve in an embodiment of the present invention;
[0039] Figure 4 A schematic diagram of the activation layer ACON-FilterNet in an embodiment of the present invention;
[0040] Figure 5 In the embodiment of the present invention, in the data set D CWRU Confusion matrix of diagnosis results on;
[0041] Figure 6 The different methods in the embodiments of the present invention are used in the data set D CWRU The noise-resistant diagnostic accuracy;
[0042] Figure 7 The model in the embodiment of the present invention is in the data set D CWRU t-SNE visualization of the features of each layer extracted;
[0043] Figure 8 In the embodiment of the present invention, in the data set D Gear Training and validation acc and loss curves on ;
[0044] Fig. 9 The different methods in the embodiments of the present invention are used in the data set D Gear Anti-noise diagnosis on
[0045] Fig.10 The model in the embodiment of the present invention is in the data set D Gear Grad-CAM++ visualization on . DETAILED DESCRIPTION
[0046] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] Embodiment 1
[0048] A rotating machinery fault diagnosis method based on residual shrinkage convolution and attention mechanism uses a trained fault diagnosis network model to detect faults in rotating machinery. The structural diagram of the fault diagnosis network model is shown in Figure 2 As shown, Figure 2 The signal on the left represents the input rotating machinery vibration signal.
[0049] The fault diagnosis network model designed in this patent consists of five parts: a wide convolution preprocessing layer, multiple stacked multi-channel residual shrinkage convolution units (Multi-Branch Residual Shrinkage Convolution, MBRSC), GSoP (Global Second-OrderPooling, GSoP) attention mechanism, global average pooling GAP layer (Global AveragePooling, GAP) and a fully connected FC layer (Full Connect, FC).
[0050] The calculation process of the fault diagnosis network model is:
[0051] S1, inputs the vibration signal of the rotating machinery, performs feature pre-extraction on the vibration signal through a wide convolution preprocessing layer, and outputs a feature map I;
[0052] S2, multi-channel and multi-scale feature extraction is performed on the output feature map I through multiple stacked multi-channel residual shrinkage convolution units. The multi-channel residual shrinkage convolution unit is composed of two channels with residual connections. Each channel includes two one-dimensional convolution layers connected in sequence and an activation layer ACON-FilterNet with a soft threshold filtering function. The output features of the two channels are connected to a 1×1 dimensional convolution after splicing to achieve feature fusion in the channel dimension. Finally, the spliced features are linearly superimposed with the residual connection to obtain the output feature map II;
[0053] S3, sends the output feature map II in S2 to the global second-order pooling module, extracts the second-order statistical information of the channel by calculating the covariance matrix of the feature map, obtains the attention factor of each channel after scaling nonlinear mapping, and finally multiplies the attention factor and the output feature map II channel by channel to complete the attention mechanism operation and output the operation result;
[0054] S4, performs global average pooling on the operation results in S3, and then classifies them through full connection to obtain the category to which the fault signal belongs.
[0055] In order to effectively filter out irrelevant features and enhance the expressiveness of effective features, the present invention designs an activation operation function with a soft threshold filtering function based on Meta-ACON in step S2. , through the activation layer ACON-FilterNet adaptively learns the activation function parameters and implements the activation function operation, The expression is:
[0056] (1)
[0057] In the formula, τ is the filtering threshold, β is the nonlinear scaling factor, p 1 is the positive dynamic weighting coefficient, p 2 is the negative dynamic weighting coefficient, x is the input feature.
[0058] In formula (1), Calculated by the following formula:
[0059] (2)
[0060] (3)
[0061] In the formula, p , p’ and represents the learning parameter, σ represents the sigmoid activation function, and t represents the input amount.
[0062] Based on the activation function designed by Meta-ACON, the smooth transition between nonlinear and linear activation functions is achieved by dynamically adjusting parameters. This smoothness can effectively improve the training stability of the neural network.
[0063] Figure 3 Shows the threshold τ = 2 o'clock, p 1. p 2 and β The influence of activation function curve, the activation function curve is divided into three parts, in the interval |x| < τ When the output is near zero, noise filtering is achieved; in the interval x>τ When y Depend on p 1 decide; in the interval x <- τ When y Depend on p 2 decisions. β Used to control the degree of nonlinearity, p 1. p 2 respectively determine the input features x Positive and negative enhancement (or inhibition) of activation responses. Figure 3 (a) shows β When the value is fixedp 1. p 2 Impact on the activation function curve, Figure 3 (b) shows different The influence of the value on the activation function curve.
[0064] The activation function uses an adaptive method to learn based on the input features τ , p 1. p 2 and β Parameters, the adaptive function can be designed using layer-level, channel-level and neuron-level learning parameters. τ and the forward dynamic weighting coefficient p 1 as a channel-level learning parameter, negative dynamic weighting coefficient p 2 and nonlinear scaling factor β as hyperparameters or neuron parameters to design the activation layer ACON-FilterNet.
[0065] Specifically, the activation layer ACON-FilterNet implements the channel-level parameter filtering threshold τ and the forward dynamic weighting coefficient p 1 adaptive learning, using a two-level nonlinear network structure to learn channel-level τ , p 1. Its specific model structure is as follows Figure 4 As shown, firstly, the input feature map is globally averaged and pooled to obtain a one-dimensional vector 1× C , C is the number of channels of the input feature map, and then the one-dimensional vector is passed to the compressed fully connected neural network to obtain a compressed two-dimensional vector 2× C / r , r is the contraction factor, and then through the expansion of the fully connected neural network, the two-dimensional vector 2× C , where a vector is used as the forward dynamic weighting coefficient p 1, another vector is subjected to the Sigmoid function to obtain a threshold factor less than 1 and greater than zero, which is multiplied by the global absolute value average pooling statistic of each channel to obtain the filtering threshold of each channel τ .
[0066] The setting of the activation layer realizes adaptive threshold filtering and enhanced (or weakened) activation of input features, effectively suppressing irrelevant features and enhancing effective features. It solves the problem that the traditional soft threshold filter function cannot provide a smooth transition when the input value is close to the threshold, the gradient disappears, and there is a lack of a mechanism to highlight key features.
[0067] Embodiment 2
[0068] Based on the first embodiment, this embodiment further optimizes various parts of the fault diagnosis network model.
[0069] Preferably, in step S1, the implementation process of the wide convolution preprocessing layer is as follows:
[0070] In this embodiment, the wide convolution preprocessing layer includes a large-size 1D convolution layer, a batch normalization layer, a ReLU activation function, and a 1D maximum pooling layer connected in sequence. The wide convolution preprocessing layer uses a large-size 1D convolution kernel and a 1D maximum pooling to preprocess the signal to quickly increase the number of channels, reduce the length of the signal sequence, and improve the representation ability and generalization performance of the model. for:
[0071] (4)
[0072] (5)
[0073] (6)
[0074] In the formula, is a one-dimensional convolution compound operation, represents the ReLU activation function, represents the batch normalization layer, Represents the one-dimensional convolution operation function of the wide convolution layer, C in Represents the input sequence x The number of channels, W is the convolution kernel, * is the cross-correlation operation, k For the input sequence x , j is the channel number of the one-dimensional convolution output feature map.
[0075] Using wide convolution to preprocess the input vibration signal can effectively retain the original feature information and has strong anti-interference ability.
[0076] Preferably, in step S2, the multi-channel residual shrinkage convolution unit is implemented as follows:
[0077] The multi-channel residual shrinkage convolution unit MBRSC uses shrinkage convolution dual channels with residual connections to extract multi-scale features of the signal. On the low-frequency channel, a large receptive field convolution kernel is used to extract low-frequency features, and on the other high-frequency channel, a small convolution kernel is used to extract high-frequency features. The output features of the two channels are concatenated and then connected to a 1×1 convolution to perform feature fusion in the channel dimension. Finally, the residual connection is used. h ( x l ) is linearly superimposed, and its operation expression is shown in formula (7). The model stacks multiple multi-channel residual shrinkage convolution units to extract the deep feature information of the signal.
[0078] (7)
[0079] In the formula, It is a one-dimensional convolution compound operation with a convolution kernel of 1×1. and They are contraction convolution operations on two channels respectively. Concat For the concatenation operation, h ( x l ) is a residual connection.
[0080] If the dual-channel fusion features are consistent with the input x l The dimensions are different. , the residual connection is a direct connection mapping operation; if the dual-channel fused features are consistent with the input x l The dimensions are the same, , the residual connection is taken as the identity operation.
[0081] Preferably, in step S3, the GSoP attention mechanism is implemented as follows:
[0082] The attention mechanism design of this model adopts a channel attention module. The commonly used Squeeze-and-Excitation (SE) channel attention uses global pooling to compress channel statistics, and then obtains the attention factor through scaling nonlinear mapping. Finally, the attention factor is multiplied with the original feature map as a scaling factor to emphasize or weaken the channel characteristics. The global maximum or mean of the channel extracted by global pooling is a simple first-order feature statistics. In order to obtain more discriminative channel features, this embodiment adopts a global second-order pooling module, which extracts the second-order statistical information of the channel by calculating the covariance matrix of the feature map.
[0083] The operation process of the global second-order pooling module is as follows: first, the output feature map of the multi-channel residual shrinkage convolution unit is reduced in dimension through 1×1 convolution, and the second-order pooling is performed through covariance operation to calculate the correlation between channels and obtain the covariance matrix. Then, the row convolution nonlinear operation is performed on the covariance matrix to obtain the attention factor of each channel. Finally, the attention factor and the feature map output by the multi-channel residual shrinkage convolution unit are multiplied channel by channel. The specific operation is shown in formulas (8)-(11):
[0084] (8)
[0085] (9)
[0086] (10)
[0087] (11)
[0088] In the formula, is a 1×1 convolution operation, x represents the output feature map of the multi-channel residual shrinkage convolution unit, is the covariance matrix operation, s is the attention factor, , is the row convolution operation, W is the convolution kernel, y’ It is the result of the global second-order pooling operation.
[0089] Specifically, for a two-dimensional input, the row convolution is a convolution kernel with a size of 1× Col , the step length is Col The two-dimensional convolution of is:
[0090] (12)
[0091] In the formula, is the row convolution operation, , x represents the output of the covariance matrix operation, is the weight matrix of the convolution kernel, m is the index of the output channel, i is the row index, j is the column index, Represents the dot product operation.
[0092] according to Figure 2 As shown, the output characteristics of the multi-channel residual shrinkage convolution unit are , first reduce the dimension to (Formula 8), the second-order pooling is performed through the covariance operation to calculate the correlation between channels, and we get dimensional covariance matrix (Formula 9), and then perform row convolution nonlinear operation on the covariance matrix to obtain the attention factor of each channel (Formula 10, 12). Finally, the attention factor and the input are multiplied channel by channel to complete the attention mechanism operation.
[0093] This embodiment introduces a global second-order pooling attention mechanism after the residual shrinkage convolution, and improves the model's ability to extract discriminative features through the second-order statistical information of the high-level channel feature map.
[0094] Preferably, in step S4, the implementation process of global average pooling and full connection is as follows:
[0095] In this embodiment, the classification layer includes a global average pooling GAP layer and a fully connected FC layer, and its specific operations are as follows:
[0096] (13)
[0097] (14)
[0098] In the formula, x c ’ represents the output result of channel c after the global average pooling layer, which is the average value of all elements on channel c; C is the total number of channels of the feature map, L is the one-dimensional length of the feature map on each channel, l : Index, indicating the first l elements, ranging from 0 to L -1; Indicates the index range of channel c; is the probability value of the fault category, is the Softmax activation function.
[0099] The global average pooling GAP operation extracts the global average pooling value of each channel x c ’ , after fully connected FC operation, use Softmax The activation function finally obtains the probability distribution value of the fault category, and the category with the largest probability value is the category to which the fault signal belongs.
[0100] The specific steps of the rotating machinery fault diagnosis process based on this patent method are:
[0101] (1) Data preprocessing: Collect the original one-dimensional time series signal, segment the signal using a sliding window, and generate a sample set. Perform standardization preprocessing on the sample set to eliminate the impact of numerical differences on model training.
[0102] (2) Dataset division: The preprocessed sample set is randomly divided into training set and test set.
[0103] (3) Model training and parameter adjustment: Build a network model, train the model using the training set, use the cross entropy loss function, optimize the model parameters through the gradient descent optimization algorithm, and save the optimal parameters of the model training.
[0104] (4) Model testing and evaluation: The final performance of the model is evaluated through the test set.
[0105] Experimental verification
[0106] The effectiveness of the present invention is verified using a bearing dataset from Case Western Reserve University and a two-stage gearbox dataset from the University of Connecticut.
[0107] 1. Bearing data set experiment:
[0108] (1) Experimental data description
[0109] The bearing data acquisition test bench of Case Western Reserve University in the United States consists of a motor, a torque sensor and a dynamometer. The SKF6205 motor-driven bearing was used as the experimental object, and the electric spark method was used to machine damaged grooves on the inner ring, rolling element and outer ring surfaces to simulate the wear of the rolling bearing in actual operation. The faulty bearings worked under loads of 0hp, 1hp, 2hp and 3hp, with a speed of 1720RPM-1797RPM. Acceleration data sets were collected at a sampling frequency of 12kHz, marked as data set A, data set B, data set C and data set D respectively. In each data set, the bearing fault is classified according to the damage location: ball, inner ring and outer ring, and the damage diameter: 0.007inch, 0.014inch and 0.021inch. There are 9 fault states and 1 normal state. The correspondence between fault type and label is shown in Table 1. For the original sequence signal , using a sliding window to continuously capture 1024 data points as a sample Under each type of load, 120 samples are taken for each fault type, and each data set contains 1200 samples. At the same time, the design data set D CWRU This is a mixed data set of four loads from 0 to 3 hp, with 4800 samples. Data sets AD and D CWRU The training set and test set are divided into 7:3 ratio.
[0110] Table 1 Fault category label coding
[0111]
[0112] In order to study the noise suppression performance of the model, Gaussian white noise is added to the data set to simulate the operating state of rotating machinery in a heavy noise pollution environment. The signal-to-noise ratio (SNR) is used to characterize the intensity of the added noise. As shown in Formula 15, the smaller the SNR value, the stronger the noise in the characterization signal.
[0113] (15)
[0114] In the formula, is the signal power, is the noise power.
[0115] (2) Model parameter setting
[0116] The low channel convolution kernel size of the multi-channel contraction convolution unit in this embodiment is set to 1*31&1*15, and the multi-channel contraction convolution unit is stacked 4 times. The detailed configuration of the model parameters is shown in Table 2. The model input is a 1*1024 one-dimensional time series vibration signal, and the output is 1*10 classification result data.
[0117] Table 2 Model parameter settings
[0118]
[0119] (3) Model training and visualization analysis
[0120] The model training and verification experiments were conducted under the Pytorch framework using Intel i7-1075 (main frequency 2.6G), memory 32G, and GTX1650 computing conditions. The model training used the cross entropy loss function as the model training loss function, using the Adam optimizer, the initial learning rate was 0.001, the batch data size was 64, and the number of training iterations was 40. Gaussian noise with an SNR of 6 was added to the DCWRU dataset. The fault diagnosis accuracy of the model on the verification set can reach 99.25%. The accuracy and recall of the diagnosis results are shown in Table 3, and the confusion matrix is shown in Figure 5 shown.
[0121] Table 3 Performance indicators of diagnostic results
[0122]
[0123] (4) Ablation experiment
[0124] In order to explore the impact of the shrinkage filter and GSoP attention mechanism on the fault diagnosis performance in the proposed method, this study designed two sets of ablation experiments, respectively targeting the unit module and the filtering method in the model to design a comparative model. CWRU Gaussian noise with an SNR of 4 is added to test the fault diagnosis accuracy of each comparison model to verify the effectiveness of each module and mechanism.
[0125] <1> Analysis of the influence of unit modules on fault diagnosis accuracy
[0126] According to whether the shrink filter module ACON-FilterNet and the GSoP attention mechanism are used, the design comparison model is shown in Table 4. Taking the model of the present invention as the original model, the structure of each model is described as follows:
[0127] Model 1.A: This model removes the ACON-FilterNet module and GSoP attention mechanism from the original model, and only retains the residual connection multi-channel convolution structure.
[0128] Model 1.B: This model removes the GSoP attention module and retains the ACON-FilterNet module.
[0129] Model 1.C: This model removes the ACON-FilterNet module from the original model and uses the ReLU activation function instead of the Thresholder Meta-ACON activation function.
[0130] Model 1.D: This model replaces the GSoP attention mechanism with the Squeeze and Excitation (SE) attention mechanism. This model is used to compare the impact of SE and GSoP attention mechanisms on system performance.
[0131] Table 4 Comparison of model structure and diagnostic accuracy
[0132]
[0133] It can be seen from the experimental data in Table 4 that compared with model 1.A, the diagnostic accuracy of model 1.B is improved by 1.77%. The standard deviation is reduced by 37.30%, indicating that the contraction filter module can improve the diagnostic effect of the model and make the diagnosis more stable. Compared with 1.A, the diagnostic accuracy of model 1.C is improved by 1.17%, and the standard deviation is reduced by 16.67%, indicating that the GSoP attention mechanism can also improve the diagnostic effect to a certain extent. Model 1.D adopts the SE attention module, and the diagnostic accuracy of 96.36 is between model 1.B without the attention mechanism and the method of the present invention using the GSoP attention mechanism. This shows that both the contraction filter module and the GSoP attention mechanism can improve the diagnostic effect of the model. In addition, compared with the SE attention mechanism, the GSoP attention mechanism uses second-order statistical information to enhance the feature expression ability of the model, further improving the accuracy of the diagnosis.
[0134] <2> Analysis of the influence of filtering method on fault diagnosis accuracy
[0135] According to different filtering methods, the design comparison model is shown in Table 5. The specific structure is described as follows:
[0136] Model 2.A: This model does not use filtering operations and uses the ReLU activation function.
[0137] Model 2.B: This model uses the shrinkage filtering operation of the DRSN model, that is, using Figure 1 The filtering operation shown.
[0138] The method of the present invention: The method of the present invention uses the Thresholder Meta-ACON activation function to implement filtering operations.
[0139] Table 5 Comparison of model structure and diagnostic accuracy
[0140]
[0141] It can be seen from the experimental data in Table 5 that compared with model 2.A, the diagnostic accuracy of the method of the present invention is improved by 2.65%. The standard deviation is reduced by 44.76%, indicating that the shrinkage filter operation in the method of the present invention can improve the diagnostic accuracy of the model and make the diagnostic results more stable. Compared with 2.B, the diagnostic accuracy of the method of the present invention is improved by 1.74%, and the standard deviation is reduced by 15.94%, indicating that the soft threshold filter module based on Meta-ACON improvement in the method of the present invention can improve the diagnostic accuracy of the model better than the residual shrinkage module.
[0142] (5) Comparative experiment
[0143] <1> Anti-noise comparison experiment
[0144] In actual work, rotating machinery often works under strong noise interference and variable load conditions. It is particularly important whether the model has strong anti-noise diagnostic performance and generalized diagnostic ability of variable load.
[0145] In order to verify the effectiveness of the fault diagnosis method of the present invention, the experiment selected the first-layer wide convolutional deep neural network WDCNN, the multi-attention one-dimensional convolutional neural network MA1DCNN, the residual network ResNet, and the residual shrinkage network DRSN as the comparison group models, and used the same data set to perform comparative analysis on the fault classification effect evaluation.
[0146] The anti-noise diagnostic performance test experiment is conducted on the dataset D CWRU Gaussian noise with different signal-to-noise ratios was added to simulate the operation of rotating machinery in a heavily polluted environment, and the noise suppression performance of the experimental group model was tested. The experimental results are shown in Figure 6 As shown, the corresponding columnar contrast Figure 6 Each group of bar graphs in the figure represents WDCNN, MA1DCNN, ResNet, DRSN and the patented method from left to right. Figure 6 The error bars represent the standard deviation of the accuracy. The Flops and parameters of each model are shown in Table 6.
[0147] Depend on Figure 6It can be seen that when the SNR is 8dB and above, the diagnostic effects of each model are basically the same. As the background noise increases, the signal-to-noise ratio (SNR) of the data set decreases, the diagnostic accuracy of each model generally decreases, and the standard deviation increases. When the signal-to-noise ratio (SNR) is in the range of [-6~6], the diagnostic accuracy of the MBRS-GSoP-Net method of the present invention is higher than that of the comparison group model, especially when the SNR is in the range of [-6~4], the improvement effect is more significant. It shows that the fault diagnosis method of the present invention has a strong discriminative feature extraction ability and strong anti-interference ability under strong noise interference. This is because the contraction filtering module adopted by the method of the present invention adds a dynamic weighting coefficient on the basis of smooth soft threshold filtering, which can adaptively filter out irrelevant interference features and enhance the discriminative feature learning ability inherent in the signal, thereby showing good anti-noise performance.
[0148] Table 6 Parameters of different methods
[0149]
[0150] <2> Variable load comparison experiment
[0151] In order to verify the diagnostic performance of the fault diagnosis method of the present invention under variable load conditions, data sets A, B, C and D of 0hp, 1hp, 2hp and 3hp loads are selected as experimental data sets. The variable load diagnosis experiment is to train the model parameters on one data set and test on other load data sets, which is expressed as training set → test set. The noise interference of 6db is set to compare the variable load fault diagnosis performance of the five models. The experimental results are shown in Table 5, where the data marked in bold is the maximum value of the task in this row.
[0152] Table 7. Results of different methods on dataset D CWRU The accuracy of variable load diagnosis (%)
[0153]
[0154] From the data in Table 7, it can be seen that under 6dB noise interference, the fault diagnosis method of the present invention can achieve a maximum accuracy of 99.61 in the B→C task, a minimum accuracy of 93.98 in the A→B task, and an average accuracy of 97.36%, which is better than other comparison group methods, indicating that the method of the present invention has good generalization ability of variable load fault diagnosis on the CWRU dataset. The method of the present invention adopts shrinkage filtering operation at the low level and a second-order statistical feature attention mechanism at the high level, which can improve the extraction ability of effective features while suppressing irrelevant features, thereby showing good generalization ability of variable load fault diagnosis.
[0155] (6) t-SNE visualization
[0156] In order to analyze the feature extraction and classification capabilities of the model of the present invention, the t-distributed stochastic neighbor embedding (t-SNE) algorithm is used to reduce the dimension of the feature graph output by each layer of the model to observe the separability of the feature data. Figure 7 Figure (a) is the t-SNE distribution of the original signal, and Figures (b)-(h) are the t-SNE distributions of the feature maps of each layer of the model. The fault types corresponding to labels 0-9 are shown in Table 1. Figure 7 It can be seen that after the layer-by-layer feature extraction of the model, the feature distribution distance of different fault types becomes larger and larger, and the feature distribution of the same fault type becomes more and more convergent.
[0157] 2. Gearbox dataset experiment:
[0158] In order to further verify the advantages of the fault diagnosis method of the present invention, a two-stage gearbox dataset from the University of Connecticut was used for experiments. The dataset was collected by the dSpace system at a sampling frequency of 20kHz. There are nine types of gear faults, namely healthy state, missing teeth, tooth root cracks, tooth surface peeling, and five different levels of tooth tip damage, namely, healthy state, missing teeth, tooth root cracks, tooth surface peeling, and tooth tip cracks L1, ..., tooth tip cracks L5. The corresponding fault types are coded as 0, 1, ..., 8. The original sequence signal contains 936 samples, each with 3600 data points. The same data preprocessing method as CWRU is used, and a sliding window is used to continuously intercept 1024 data points as an experimental sample. 312 samples are collected for each type of gear fault type. The final experimental dataset D Gear There are 2808 sample data in total, which are divided into training set and test set in a ratio of 7:3.
[0159] (1) Model training
[0160] In the dataset D Gear White noise with an SNR of 6 is added, the model training batch data size is 64, and the number of training iterations is 100. The training and verification accuracy acc curves and cross entropy loss loss curves of the model of the present invention are shown in Figure 2. Figure 8 As shown in the figure, the fault diagnosis accuracy of the model of the present invention on the validation set can reach 99.84%.
[0161] (2) Noise resistance test
[0162] In order to further verify the anti-noise diagnosis performance of the proposed model, the gearbox data set D Gear The fault diagnosis performance of the method of the present invention is compared with that of the comparison group methods WDCNN, MA1DCNN, ResNet, and DRSN under different signal-to-noise ratios (SNRs). The experimental results are as follows: Fig. 9 shown.
[0163] Fig. 9 The error bars represent the standard deviation of the accuracy. Fig. 9 The experimental data show that the diagnostic classification accuracy of the method of the present invention is higher than that of the comparison group model when the SNR is in the range of [-6~6], especially in the range of [-6~0], the diagnostic accuracy is significantly improved. This shows that the method of the present invention has strong anti-interference ability on the gearbox data set and has strong discriminative feature extraction ability.
[0164] (3) Grad-CAM++ Visual Analysis
[0165] In order to demonstrate the discriminative features that the method of the present invention focuses on during fault diagnosis, a gradient weighted class activation mapping (ClassActivation Mapping Grad-CAM++) visualization method is used for processing, such as Fig.10 As shown in the figure, Grad-CAM++ shows the contribution distribution of sequence data to the fault prediction type output. The darker the color, the greater the weight, indicating that the corresponding area in the signal has a higher response and contribution to the prediction of the fault. Fig.10 It can be seen that different fault types have different CAM activation features, amplitudes, and receptive fields. For example, tooth tip cracks with different damage degrees have different activation areas, activation impact amplitudes, and receptive fields. The reason is that the present invention uses multi-channel residual shrinkage convolution, which enables the model to analyze global vibration data with a larger receptive field and identify the discriminative features of the impact segment.
[0166] In summary, the present invention proposes a rotating machinery fault diagnosis method that combines improved shrinkage convolution and global second-order pooling attention mechanism.
[0167] (1) This method adopts deep multi-channel residual shrinkage convolution. Meta-ACON designs a soft threshold filter activation function and constructs a soft threshold filter module.
[0168] (2) After deep multi-channel residual shrinkage convolution, this method uses the GSoP attention mechanism to assign different response weights to abstract multi-scale and multi-level features, and uses second-order statistical information to enhance the feature expression ability of the model, further improving the accuracy of diagnosis. Experimental results show that under low signal-to-noise ratio conditions and variable load conditions, this method has good fault identification and generalization capabilities.
[0169] Finally, it should be noted that the parts of the present invention that are not described in detail are all prior art. Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not intended to limit the invention. Although the invention is described in detail with reference to the aforementioned examples, those of ordinary skill in the art can still modify the technical solutions recorded in the aforementioned examples, or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, etc. made within the spirit and principles of the invention should be included in the scope of protection of the invention.
Claims
1. A rotating machinery fault diagnosis method based on residual shrinkage convolution and attention mechanism, characterized in that: The trained fault diagnosis network model is used to detect the fault of the rotating machinery. The calculation process of the fault diagnosis network model includes: S1, inputs the vibration signal of the rotating machinery, performs feature pre-extraction on the vibration signal through a wide convolution preprocessing layer, and outputs a feature map I; S2, multi-channel and multi-scale feature extraction is performed on the output feature map I through multiple stacked multi-channel residual shrinkage convolution units. The multi-channel residual shrinkage convolution unit is composed of two channels with residual connections. Each channel includes two one-dimensional convolution layers connected in sequence and an activation layer ACON-FilterNet with a soft threshold filtering function. The output features of the two channels are connected to a 1×1 dimensional convolution after splicing to achieve feature fusion in the channel dimension. Finally, the spliced features are linearly superimposed with the residual connection to obtain the output feature map II; S3, sends the output feature map II in S2 to the global second-order pooling module, extracts the second-order statistical information of the channel by calculating the covariance matrix of the feature map, obtains the attention factor of each channel after scaling nonlinear mapping, and finally multiplies the attention factor and the output feature map II channel by channel to complete the attention mechanism operation and output the operation result; S4, performs global average pooling on the operation results in S3, and then performs full connection classification to obtain the category to which the fault signal belongs; The activation operation function of the activation layer ACON-FilterNet Using Meta-ACON design, the expression is: In the formula, τ is the filtering threshold, β is the nonlinear scaling factor, p 1 is the positive dynamic weighting coefficient, p 2 is the negative dynamic weighting coefficient, x is the input feature; The activation layer ACON-FilterNet implements channel-level parameter filtering thresholds τ and the forward dynamic weighting coefficient p 1’s adaptive learning, the specific model structure is designed as follows: First, the input feature map is globally averaged by absolute value pooling to obtain a one-dimensional vector 1× C , C is the number of channels of the input feature map, and the one-dimensional vector is passed to the compressed fully connected network to obtain a compressed two-dimensional vector 2× C / r , r is the contraction factor, and then the final two-dimensional vector 2× C , where a vector is used as the forward dynamic weighting coefficient p 1, another vector is subjected to the Sigmoid function to obtain a threshold factor less than 1 and greater than zero, which is multiplied by the global absolute value average pooling statistic of each channel to obtain the filtering threshold of each channel τ .
2. The rotating machinery fault diagnosis method based on residual shrinkage convolution and attention mechanism according to claim 1 is characterized in that: The filtering threshold τ and the forward dynamic weighting coefficient p 1 as a channel-level learning parameter, p 2 and β As hyperparameters or neuron parameters.
3. The rotating machinery fault diagnosis method based on residual shrinkage convolution and attention mechanism according to claim 1 is characterized in that: The operation process of the global second-order pooling module in S3 is as follows: first, the output feature map II is reduced in dimension through 1×1 convolution, second-order pooling is performed through covariance operation, the correlation between channels is calculated, and the covariance matrix is obtained. Then, a row convolution nonlinear operation is performed on the covariance matrix to obtain the attention factor of each channel, and finally, the attention factor and the output feature map II are multiplied channel by channel.
Citation Information
Patent Citations
Fault diagnosis method based on fusion of residual learning and attention mechanism
CN115640531A
Rolling bearing fault diagnosis method and system based on multi-scale residual shrinkage network
CN118603546A