Industrial Data Diagnosis Method Based on SEResNet and Attention Mechanism
By combining SEResNet and attention mechanism in industrial data diagnosis, the problems of insufficient feature extraction and insufficient model stability in the prior art are solved, and efficient fault signal extraction and accurate fault diagnosis are achieved.
Patent Information
- Application Number
- CN202510507221.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the multi-sensor data processing, existing industrial data diagnostic methods have problems such as insufficient feature extraction, serious interference from useless information, and insufficient model stability and generalization capabilities.
Using an industrial data diagnosis method based on SEResNet and attention mechanism, a SEResNet feature extraction network is constructed by combining SE-Net signal reconstruction and ResNet residual feature extraction network, and combining attention mechanism to fuse multiple sensor signals to achieve efficient extraction and diagnosis of fault signals.
It improves the efficiency of extracting useful information of fault signals, reduces useless information interference, enhances the stability and comprehensiveness of fault diagnosis, and significantly improves the diagnostic accuracy.
Smart Images

Figure CN120030334B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an industrial data diagnosis method based on SEResNet and attention mechanism, belonging to the field of mechanical equipment signal processing. Background Art
[0002] With the advent of industrial intelligence, the scale of mechanical equipment has been continuously expanding and the system has been continuously complicated, resulting in an increase in sensor monitoring points. After long-term data collection, a large amount of data has been formed, and these large amounts of data have greatly increased the difficulty of industrial data fault diagnosis.
[0003] At present, common multi-sensor fusion fault diagnosis methods include the CNN model method with multi-channel input, the fusion method based on attention mechanism, and the fusion fault diagnosis method based on image stitching. Among them, the CNN model with multi-channel input collects different types of signals through multi-channels, then uses the CNN model for feature extraction and fusion, and finally realizes fault classification; however, the CNN model has high requirements for data quality and the preprocessing process is complex; and there will also be problems of gradient disappearance or gradient explosion as the number of network layers deepens; the fusion method based on attention mechanism constructs a dynamic graph and uses the attention mechanism to enhance the model's learning ability for key features, but this method is more sensitive to noise and the generalization ability of the model is limited; the fusion method based on image stitching realizes fault diagnosis by stitching and fusing the image data of different sensors, but the feature extraction and fusion of this method are more difficult and it is highly dependent on the environment; in addition, the above methods cannot ensure the stability and generalization of the network model while extracting multi-level and deep-level features.
[0004] In addition, the existing methods only use the residual network as the feature extractor, and cannot effectively eliminate the interference of useless information, and the diagnosis results are not accurate enough. Summary of the Invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides an industrial data diagnosis method based on SEResNet and attention mechanism, and this method includes:
[0006] Step 1: Collect historical data in the industrial process and preprocess it;
[0007] Step 2: Divide the data preprocessed in Step 1 into a training set, a validation set and a test set;
[0008] Step 3: Construct a fault diagnosis model based on the fusion of SEResNet multi-sensor attention mechanism (hereinafter referred to as SRMMF fault diagnosis model) by designing a feature extraction and fusion sub-network and a fault recognition sub-network;
[0009] Step 4: Train the model constructed in Step 3 using the training set obtained in Step 2;
[0010] Step 5: Input the data to be diagnosed into the model trained in Step 4 for fault diagnosis.
[0011] The industrial process data in Step 1 includes: the process data of the waste heat recovery fan in the cold rolling mill, the process data of the descaling pump in the hot rolling mill, the process data of the coal mine air compressor, etc.;
[0012] The preprocessing includes: Fast Fourier Transform (FFT) and Variational Mode Decomposition (VMD);
[0013] In Step 2, the ratio of the training set, validation set, and test set is 7:2:1;
[0014] The SRMMF fault diagnosis model in Step 3 includes a feature extraction and fusion sub-network and a fault identification sub-network; the feature extraction and fusion sub-network can adaptively combine multiple fusion layers through a multi-layer fusion framework and an attention-based fusion strategy, and extract relevant information from multiple sensor features. Then, the fused multi-sensor features are further sent to the fully connected layer in the fault identification sub-network to achieve fault identification.
[0015] The feature extraction and fusion sub-network includes 1 central network and 2 SEResNet branch networks; the central network includes an AMF module, a dimension transformation module, and a global pooling module; the SEResNet branch network includes a convolutional layer, a pooling layer, a SEResNet feature extraction network, and global pooling;
[0016] The feature extraction and fusion sub-network is constructed based on a multi-layer fusion framework. First, the features of different signals are extracted through two SEResNet branch networks. The fusion stage is activated after the convolutional layer, and a fusion point is set after the pooling layer. The pooling layer is mainly used for aggregating effective information. The fusion point is set after the pooling layer to reduce the number of model parameters. The attention fusion algorithm is used to fuse multi-sensor features at each fusion point. Due to the presence of the convolutional layer and the pooling layer, the feature dimensions between different fusion points are not consistent. To fuse features at different levels, the fused features at the current level need to be dimensionally transformed, and the dimension transformation module is used to match the feature dimensions of the next fusion point. All pooling operations are max pooling, and batch normalization is adopted after each convolutional layer.
[0017] The SEResNet branch network automatically extracts the deep features of single-sensor data, fuses the extracted multi-level multi-sensor data features with the central network, significantly enhances the information interaction between multi-sensor data, and realizes the adaptive hierarchical fusion of information.
[0018] Specifically, it includes: The first convolutional layer is used to extract shallow features of the signal. The extracted shallow feature maps are provided to the SEResNet layer, and 16 SEResNet blocks are used to extract deep features to obtain the most important features and suppress redundant features. Then, in order to improve accuracy and reduce computational costs, the shallow features and deep features are fused through global residuals. The features extracted from each SEResNet block are combined and connected through a convolutional layer, and then the effective information is aggregated through a pooling layer for fusion in the central network.
[0019] The fault recognition subnetwork uses a global average pooling layer to receive the high-level fusion features. It should be noted that although recognition results are also required in the SEResNet branch network, only the performance of the branch network needs to be ensured, and the final recognition result is still determined by the central network.
[0020] The SEResNet feature extraction network in the SEResNet branch network combines the signal reconstruction of SE-Net and the ResNet residual feature extraction network. SE-Net can automatically learn a set of weights through a small subnetwork and calculate the weights for each channel of the feature map. In this way, useful feature channels are enhanced, and redundant feature channels are weakened. In addition, the residual network is easy to optimize and alleviates the problem of vanishing gradients due to the increase in depth. Therefore, the residual network is selected as the feature extractor.
[0021] The specific process is as follows:
[0022] SE-Net includes: Squeeze, Excitation, and Scale. Given the input data with the number of feature channels , through a series of convolutional operations, the features with the number of feature channels are obtained . The implementation process is as follows: :
[0023]
[0024] represents the input data, represents the convolutional operation.
[0025] is compressed through global average pooling for , and each feature map is compressed into a real number with a global sensitivity field on the feature map, obtaining a global attention information with a size of . The calculation process is as follows:
[0026]
[0027] Among them, represents the number of feature channels as The feature is the global attention information extracted from the corresponding feature map through an extrusion operation. represents global average pooling.
[0028] Next, excitation is performed on two fully connected layers in the SE-Net to obtain the relationship between feature channels. The training results are used to increase the weight of the more important feature information of the task and reduce the weight of the unimportant feature information. The calculation process is as follows:
[0029]
[0030] where and respectively represent the weight matrices of different fully connected layers in the SE-Net, represents the ReLU activation function, represents the sigmoid activation function, represents the feature weights of each resulting feature map.
[0031] The feature weights of each feature channel are estimated according to the value of the loss function, and its expression is:
[0032]
[0033] where, represents the total number of categories, represents the true value, represents the predicted value.
[0034] The weight parameters are updated through backpropagation of the loss function. According to the error between the predicted value and the true value, the optimal weights are output. Then, based on the feature weights, the useful information is enhanced and the useless information is suppressed, so that the model can obtain better performance.
[0035] Finally, there is the Scale operation, which performs recalibration of the original features in the channel dimension by multiplying the feature weights output by the Excitation operation by the previous feature channel by channel. Therefore, the model can distinguish the characteristics of each channel. The formula is as follows:
[0036]
[0037] where, is the feature map that rescales the original features.
[0038] The SE-Net is used to learn the importance of multi-channel features, enhance the feature learning ability of each channel, and show higher accuracy.
[0039] The residual network contains many residual blocks, and the basic residual learning block is defined as:
[0040]
[0041]
[0042] Among them is the feature map that resizes the original feature, is the residual function, represents the weights of the first layer of convolution in the residual block, represents the weights of the second layer of convolution in the residual block, is the non - linear function ReLU, represents the output after learning the network.
[0043] The central network fuses the features extracted from the two - branch network through the AMF module, and realizes the fusion between the fusion points through dimension conversion, greatly improving the comprehensiveness of feature extraction and fusion;
[0044] The specific process of the AMF module is as follows: First, two single - branch feature extraction networks are used to extract features from a single signal, then the extracted features are input into the AMF module of the central network to fuse the signal features from two different sensors, and finally, the SoftMax classification function is used for fault classification.
[0045] Taking two monitoring signals as an example. Suppose are the signal features extracted from the monitoring signals of two different sensors respectively, represents a three - dimensional real - valued tensor; the specific fusion process is as follows:
[0046] (1) Use the global average pooling operation to compress the signal features from the monitoring signals of different sensors in the spatial dimension:
[0047]
[0048]
[0049]
[0050] Among them and respectively represent the compressed features of the two signal features in the th channel, represents that the spatial dimension of the feature is , represents the number of channels of the feature, i represents the c th feature in the i th channel, represents in the The features of a channel denotes in the features of the -th channel.
[0051] (2) Combine the compressed features of the two sensor signals to generate global representation information ; its expression is:
[0052]
[0053] where respectively denote the signal features after compression;
[0054] In addition, to enable the excitation signal to fully correct the features of each sensor, the feature learning process should be non-linear. Therefore, a fully connected operation is added after the global information features to improve non-linearity; the expression is:
[0055]
[0056] where is the compact feature after dimensionality reduction, represents the dimension within the real number range; and respectively denote the weight and bias; denotes the non-linear function ReLU.
[0057] (3) Based on the compact feature after dimensionality reduction, then generate the excitation signals and with soft attention, where can adaptively select the features of each channel. In addition, a SoftMax function is added to obtain the excitation probability of each channel feature:
[0058]
[0059]
[0060] Here respectively denote the learnable weight matrices corresponding to the two signal features , represents that the dimension of the matrix is rows, columns, and the elements are real numbers. and represent that through two different learnable weight matrices and , the same feature is transformed, and respectively represent and the excitation signals of
[0061] (4) Each feature from the data of different sensors is recalibrated and fused by the excitation signal through a gating mechanism:
[0062]
[0063] wherein represents the fused feature; ⊗ represents the product in the channel direction.
[0064] After the above four steps, the features of the signals of the two sensors can be fused at each fusion point. Due to the existence of the convolutional layer and the pooling layer, the feature dimensions between different fusion points are not consistent.
[0065] To fuse features at different levels, the fused feature of the current level will first pass through a dimension transformation module (DTM). The dimension transformation module includes a 1×1 convolutional layer and a pooling layer. The 1×1 convolutional layer can scale the channel dimension and increase the non-linearity of the network, and the pooling layer can scale the spatial dimension of the feature. When the feature dimensions of two fusion points are consistent, the features of the current layer and the lower layer can be obtained through the following formula:
[0066]
[0067] where represents the output feature of the th fusion point, represents the weight coefficient, represents the multi-sensor feature passing through the AMF module at the th fusion point.
[0068] The dimension transformation module transforms the dimension of the previous fusion point into the same dimension as the next fusion point, and thus the feature fusion of the two fusion points can be realized.
[0069] Step 4: During the training process of the model, by continuously adjusting the iteration times and the learning rate of the model, when the accuracy rate tends to be stable and the loss value curve is stable and no longer decreases, the model training is completed; then the model is verified through the validation set, and when the validation set also reaches the above effects, the trained model can be obtained for fault diagnosis of the data to be diagnosed.
[0070] Step 5: Input the data to be diagnosed into the model trained in Step 4 for fault diagnosis.
[0071] The beneficial effects of the present invention are:
[0072] The present invention proposes an industrial data diagnosis method based on SEResNet and attention mechanism. By combining the signal reconstruction of SE-Net and the ResNet residual feature extraction network, a SEResNet feature extraction network is constructed, which improves the extraction efficiency of useful information of fault signals and solves the problem of interference from useless information in feature extraction. By combining the attention mechanism to fuse signals from two different sensors, the comprehensiveness and stability of fault diagnosis are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0074] Figure 1 It is a schematic structural diagram of SE-Net in the industrial data diagnosis method based on SEResNet and attention mechanism proposed by the present invention;
[0075] Figure 2 It is a structural diagram of the SRMMF fault diagnosis model in the industrial data diagnosis method based on SEResNet and attention mechanism proposed by the present invention;
[0076] Figure 3 It is a flowchart of the industrial data diagnosis method based on SEResNet and attention mechanism proposed by the present invention;
[0077] Figure 4 It is a curve graph of the change of accuracy rate and loss value of the industrial data diagnosis method based on SEResNet and attention mechanism proposed by the present invention;
[0078] Figure 5 It is a curve graph of the change of accuracy rate and loss value of the existing technology MCFCNN method;
[0079] Figure 6 It is a curve graph of the change of accuracy rate and loss value of the existing technology RMF method;
[0080] Figure 7 It is a schematic diagram of the confusion matrix of the industrial data diagnosis method based on SEResNet and attention mechanism proposed by the present invention;
[0081] Figure 8 It is a schematic diagram of the confusion matrix of the existing technology MCFCNN method;
[0082] Figure 9 It is a schematic diagram of the confusion matrix of the existing technology RMF method. Detailed implementation manners
[0083] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0084] Embodiment 1
[0085] The present invention proposes an industrial data diagnosis method based on SEResNet and an attention mechanism. The method specifically includes:
[0086] Step 1: Collect historical data in the industrial process and preprocess it;
[0087] Step 2: Divide the preprocessed data in Step 1 into a training set, a validation set and a test set;
[0088] Step 3: Construct an SRMMF fault diagnosis model by designing a feature extraction and fusion sub-network and a fault identification sub-network;
[0089] Step 4: Use the training set obtained in Step 2 to train the model constructed in Step 3;
[0090] Step 5: Input the data to be diagnosed into the model trained in Step 4 for fault diagnosis.
[0091] The industrial process data in Step 1 includes: the process data of the waste heat recovery fan in the cold rolling mill, the process data of the descaling pump in the hot rolling mill, the process data of the coal mine air compressor, and so on;
[0092] The preprocessing includes: fast Fourier transform (FFT) and variational mode decomposition (VMD);
[0093] In Step 2, the ratio of the training set, the validation set and the test set is 7:2:1;
[0094] The SRMMF fault diagnosis model in Step 3 includes a feature extraction and fusion sub-network and a fault identification sub-network; the feature extraction and fusion sub-network can adaptively combine multiple fusion layers through a multi-layer fusion framework and an attention-based fusion strategy, and extract relevant information from multiple sensor features. Then, the fused multi-sensor features are further sent to the fully connected layer in the fault identification sub-network to achieve fault identification.
[0095] The feature extraction and fusion sub-network includes 1 central network and 2 SEResNet branch networks; the central network includes an AMF module, a dimension conversion module and a global pooling; the SEResNet branch network includes a convolutional layer, a pooling layer, a SEResNet feature extraction network and a global pooling;
[0096] The feature extraction and fusion sub-network is constructed based on a multi-layer fusion framework. First, two SEResNet branch networks are used to extract the features of different signals. The fusion stage is activated after the convolutional layer, and fusion points are set after the pooling layer. The pooling layer is mainly used for aggregating effective information, and the fusion points are set after the pooling layer to reduce the number of model parameters. The attention fusion algorithm is used to fuse multi-sensor features at each fusion point. Due to the presence of the convolutional layer and the pooling layer, the feature dimensions between different fusion points are not consistent. To fuse features at different levels, the fusion features at the current level need to be dimensionally transformed, and the dimensional transformation module is used to match the feature dimensions of the next fusion point. All pooling operations are max pooling, and batch normalization is adopted after each convolutional layer.
[0097] The SEResNet branch network automatically extracts the deep features of single-sensor data, fuses the extracted multi-level multi-sensor data features with the central network, significantly enhances the information interaction between multi-sensor data, and realizes the adaptive hierarchical fusion of information.
[0098] Specifically, the first convolutional layer is used to extract the shallow features of the signal. The extracted shallow feature maps are provided to the SEResNet layer, and 16 SEResNet blocks are used to extract the deep features to obtain the most important features and suppress redundant features. Then, to improve the accuracy and reduce the computational cost, the shallow features and the deep features are fused through global residuals. The features extracted from each SEResNet block are combined and connected through a convolutional layer, and then the effective information is aggregated through a pooling layer for the fusion of the central network.
[0099] The fault identification sub-network uses the global average pooling layer to receive the high-level fusion features. It should be noted that although the recognition results are also required in the SEResNet branch network, it is only necessary to ensure the performance of the branch network, and the final recognition results are still determined by the central network.
[0100] The SEResNet feature extraction network in the SEResNet branch network combines the signal reconstruction of SE-Net and the residual feature extraction network of ResNet. SE-Net can automatically learn a set of weights through a small sub-network and calculate the weights for each channel of the feature map. In this way, the useful feature channels are enhanced, and the redundant feature channels are weakened. In addition, the residual network is easy to optimize and alleviates the problem of gradient disappearance due to the increase in depth. Therefore, the residual network is selected as the feature extractor.
[0101] The specific process is as follows:
[0102] SE-Net includes: squeeze, excitation, and scale. Given the input data with the number of feature channels of , the feature with the number of feature channels obtained through a series of convolution operations is . . The implementation process is as follows:
[0103]
[0104] represents the input data, represents the convolution operation.
[0105] Through global average pooling, is compressed, and each feature map is compressed into a real number with a global sensitivity field on the feature map, obtaining global attention information of size . . The calculation process is as follows:
[0106]
[0107] Among them, represents the feature with the number of feature channels being , is the global attention information extracted from the corresponding feature map through the squeezing operation. represents global average pooling.
[0108] Next, excitation is performed on two fully connected layers in SE-Net to obtain the relationship between feature channels. The training results are used to increase the weights of more important feature information for the task and reduce the weights of unimportant feature information. The calculation process is as follows:
[0109]
[0110] Among them and respectively represent the weight matrices of different fully connected layers in SE-Net, represents the ReLU activation function, represents the sigmoid activation function, represents the feature weights of each resulting feature map.
[0111] Estimate the feature weights of each feature channel according to the value of the loss function, and its expression is:
[0112]
[0113] Among them, represents the total number of categories, represents the true value, represents the predicted value.
[0114] Update the weight parameters by backpropagation through the loss function, and output the optimal weights according to the error between the predicted value and the true value. Then, based on the feature weights, enhance the useful information and suppress the useless information, so that the model can obtain better performance.
[0115] Finally, there is the Scale operation, which re-calibrates the original features in the channel dimension by multiplying the feature weights output by the Excitation operation by the previous feature channel by channel. Therefore, the model can distinguish the characteristics of each channel. The formula is as follows:
[0116]
[0117] Among them, is the feature map that rescales the original features.
[0118] SE-Net is used to learn the importance of multi-channel features, enhance the feature learning ability of each channel, and show higher accuracy.
[0119] The residual network contains many residual blocks, and the basic residual learning block is defined as:
[0120]
[0121]
[0122] Among them is the feature map that rescales the original features, is the residual function, represents the weight of the first layer of convolution in the residual block, represents the weight of the second layer of convolution in the residual block, is the non-linear function ReLU, represents the output after learning the network.
[0123] The central network fuses the features extracted from the two-branch network through the AMF module, and realizes the fusion between the fusion points through dimension conversion, greatly improving the comprehensiveness of feature extraction and fusion;
[0124] The specific process of the AMF module is as follows: First, extract the features of a single signal through two single-branch feature extraction networks, then input the extracted features into the AMF module of the central network to fuse the signal features from two different sensors, and finally perform fault classification through the SoftMax classification function.
[0125] Taking two monitoring signals as an example. Suppose are the signal features extracted from the monitoring signals of two different sensors respectively, represents a three-dimensional real number tensor; the specific fusion process is as follows:
[0126] (1) Use the global average pooling operation to compress the signal features from the monitoring signals of different sensors in the spatial dimension:
[0127]
[0128]
[0129]
[0130] Among them and respectively represent the compressed features of two signal features in the th channel, represents that the spatial dimension of the feature is , represents the number of channels of the feature, i represents the c th feature in the i th channel, represents the feature in the th channel in represents the feature in the th channel in
[0131] (2) Combine the compressed features of the two sensor signals to generate global representation information ; its expression is:
[0132]
[0133] Among them, respectively represent the features after the signal features are compressed;
[0134] In addition, in order to enable the excitation signal to fully correct the features of each sensor, the feature learning process should be non-linear. Therefore, a fully connected operation is added after the global information feature to improve non-linearity; the expression is:
[0135]
[0136] Among them, is the compact feature after dimensionality reduction, represents the dimension within the real number range; and respectively represent the weight and bias; represents the non-linear function ReLU.
[0137] (3) Based on the compact feature after dimensionality reduction, then generate an excitation signal with soft attention and , where can adaptively select the features of each channel. In addition, a SoftMax function is added to obtain the excitation probability of the features of each channel:
[0138]
[0139]
[0140] Here , respectively represent the learnable weight matrices corresponding to the two signals , The dimension of the matrix is rows, columns, and the elements are real numbers. and represent passing through two different learnable weight matrices and , respectively, to transform the same feature . and respectively represent and 's excitation signals.
[0141] (4) Each feature from different sensor data is recalibrated and fused by the excitation signal through a gating mechanism:
[0142]
[0143] Among them, represents the fused feature; ⊗ represents the product in the channel direction.
[0144] After the above four steps, the features of the signals of the two sensors can be fused at each fusion point. Due to the existence of the convolutional layer and the pooling layer, the feature dimensions between different fusion points are not consistent.
[0145] In order to fuse features at different levels, the fused feature of the current level will first pass through a dimension transformation module (DTM). The dimension transformation module includes a 1×1 convolutional layer and a pooling layer. The 1×1 convolutional layer can scale the channel dimension and increase the nonlinearity of the network, and the pooling layer can scale the spatial dimension of the features. When the feature dimensions of two fusion points are consistent, the features of the current layer and the lower layer can be obtained through the following formula:
[0146]
[0147] where represents the output feature of the th fusion point, represents the weight coefficient, Indicates the multi-sensor features passing through the AMF module at the th fusion point.
[0148] The dimension conversion module converts the dimension of the previous fusion point to the same dimension as the next fusion point, so as to realize the feature fusion of the two fusion points.
[0149] Step 4: During the training process of the model, by continuously adjusting the iteration times and learning rate of the model, when the accuracy rate tends to be stable and the loss value curve is stable and no longer decreases, the model training is completed; then the model is verified through the validation set. When the validation set also reaches the above effects, the trained model can be obtained for fault diagnosis of the data to be diagnosed.
[0150] Step 5: Input the data to be diagnosed into the model trained in Step 4 for fault diagnosis.
[0151] Embodiment 2
[0152] This embodiment provides an industrial data diagnosis method based on SEResNet and attention mechanism, which is implemented based on the SRMMF fault diagnosis model described in Embodiment 1.
[0153] Step 1: Collect the vibration signal and current signal during the operation of the bearing and preprocess them;
[0154] Among them, the preprocessing includes fast Fourier transform (FFT) and variational mode decomposition (VMD). The signal is converted from the time domain to the frequency domain through FFT to obtain the spectrum information of the signal, and the signal is decomposed into a series of mode functions through VMD to obtain the time domain feature information of the signal at different scales. Through the above preprocessing methods, the time domain and frequency domain feature information of the signal are obtained simultaneously, thereby improving the globality of feature extraction.
[0155] Step 2: Input the data preprocessed in Step 1 into the SRMMF fault diagnosis model for diagnosis, and output the final diagnosis result.
[0156] To evaluate the performance of the proposed SRMMF fault diagnosis model of the present invention, the model is compared with existing multi-sensor fusion methods, including the CNN model with multi-channel input (MCFCNN) and the multi-sensor fusion method using ResNet alone (RMF). For the CNN model with multi-channel input (MCFCNN), specific reference can be made to the introduction in "Li Hongmei. Research on Intelligent Fault Diagnosis Method Based on Convolutional Neural Network [D]. North University of China, 2021.", and for the multi-sensor fusion method using ResNet alone (RMF), reference can be made to the introduction in "Hong Liang, Yu Qiyuan, Qin Chaoqun, et al. Bearing Fault Diagnosis Based on Information Fusion and Double-Connected Attention Residual Network [J]. Journal of Vibration and Shock, 2023, 42(20): 114-123".
[0157] For the SRMMF fault diagnosis model proposed in the present invention, a network model is built based on the pytorch framework. The network model can output an accuracy graph, a loss value graph, and a confusion matrix graph. Through the above results, the fault diagnosis accuracy and classification effect are analyzed, and then the performance of the proposed network model is evaluated.
[0158] Comparative analysis is carried out from three aspects: the accuracy change curve, the loss value change curve, and the confusion matrix.
[0159] The accuracy and loss value change curves of the SRMMF model in the method proposed in the present invention are as Figure 4 shown. The accuracy and loss value change curves of the MCFCNN model and the RMF model in the prior art are as Figure 5 and Figure 6 shown respectively; in the accuracy curve, the proposed SRMMF model in the present invention tends to be stable after fewer iterations, with reduced fluctuations, and the accuracy reaches more than 99%; the accuracy curve of the MCFCNN model has more fluctuations than that of the SRMMF, but the accuracy can still reach more than 95% after tending to be stable; while the accuracy of the RMF model is similar to that of the SRMMF after stabilization, but the fluctuations are large, and the model does not have better stability. In the loss value curve, the curve of the proposed SRMMF model in the present invention converges faster and is smoother, and more than 99% accurate extraction and recognition of the fault characteristics of different types of bearings are successfully achieved. The curves of the other two models converge slower and have larger fluctuations, and the effect is worse than that of the SRMMF.
[0160] The confusion matrix of the SRMMF model in the method proposed in the present invention is as Figure 7 shown. The confusion matrices of the MCFCNN model and the RMF model in the prior art are as Figure 8 and Figure 9As shown, the confusion matrix is a table used in machine learning to evaluate the performance of classification models. With actual classes and predicted classes as rows and columns, it includes four cases: true positives, true negatives, false positives, and false negatives, and can be used to calculate metrics such as accuracy, precision, recall, etc., to help comprehensively evaluate the model performance. For the confusion matrix diagram of the SRMMF model proposed in this invention, it can be observed that the prediction accuracy of the model for all classes reaches 100%, and there are no misclassifications. While the MCFCNN model has some confusions in label 0 and label 3. In label 0, 1 sample is wrongly predicted as label 2 respectively, and in label 3, 4 samples are wrongly predicted as label 0. The RMF model has confusion in label 0, and 19 samples are wrongly predicted as label 3.
[0161] Since compared with the MCFCNN model and RMF model in the prior art, the SRMMF model in the method proposed in this invention has the advantages of high accuracy and model stability.
[0162] Some steps in the embodiments of this invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0163] The above are only the preferred embodiments of this invention and are not intended to limit this invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this invention shall be included within the protection scope of this invention.
Claims
1. An industrial data diagnosis method based on SEResNet and attention mechanism, characterized in that: The method comprises: Step 1: Collect historical data from industrial processes and preprocess them; Step 2: Divide the preprocessed data in step 1 into training set, validation set and test set; Step 3: By designing the feature extraction and fusion sub-network and the fault identification sub-network, a fusion fault diagnosis model based on the SEResNet multi-sensor attention mechanism is constructed; Step 4: Use the training set obtained in step 2 to train the model constructed in step 3; Step 5: Input the data to be diagnosed into the model trained in step 4 for fault diagnosis; The feature extraction and fusion sub-network in step 3 includes a central network and a SEResNet branch network; The SEResNet branch network includes: a convolutional layer, a pooling layer, a SEResNet feature extraction network and a global pooling; The central network includes: an AMF module, a dimension conversion module and a global pooling; The AMF module extracts signal features from two sensor monitoring signals. Perform global average pooling operation in the spatial dimension to obtain compressed features ; Then compress the features Combined into global representation information F g , and then use the full connection operation and ReLU function to perform nonlinear feature learning to obtain compact features after dimensionality reduction ; Then, using the corresponding two signal features Learnable weight matrix For the compact features after dimensionality reduction Transform and combine with SoftMax function to generate signal features with soft attention and The excitation signal ; Finally, through the gating mechanism, the excitation signal Signal characteristics Recalibrate and fuse to obtain the final fusion feature F ; The fault identification subnetwork includes a fully connected layer and an output, and implements fault identification by receiving the output of the feature extraction and fusion subnetwork; The features of different signals are extracted through the SEResNet feature extraction network in the SEResNet branch network, and the extracted features are fused through the AMF module in the central network. The fusion between different AMF modules is then achieved through the dimension conversion module. The fusion result is output through the fully connected layer to produce the fault identification result.
2. The method according to claim 1, characterized in that The SEResNet feature extraction network is composed of a signal reconstruction of SE-Net and a residual feature extraction network of ResNet; The specific process of SE-Net signal reconstruction is as follows: SE-Net includes: squeezing, excitation and scaling. The number of feature channels is given as Input data , the number of feature channels obtained through convolution operation is Features , the expression is: in, Represents input data, Represents the convolution operation; Through global average pooling, the features Compress each feature map into a real number with a global sensitivity field on the feature map, and obtain a size of Global attention information , the expression is: in, Indicates the number of feature channels Features, is the global attention information extracted from the corresponding feature map through the squeezing operation; represents global average pooling; Excitation is performed on the two fully connected layers of SE-Net to obtain the relationship between feature channels; the expression is: in and Respectively represent the weight matrices of different fully connected layers in SE-Net, represents the ReLU activation function, represents the sigmoid activation function, Represents the feature weight of each resulting feature map; The feature weight of each feature channel is estimated according to the value of the loss function. The expression of the loss function is: in, represents the total number of categories, represents the true value, Represents the predicted value; updates the weight parameters through back propagation of the loss function; Finally, the feature weights output by the Excitation operation are multiplied by the previous feature channel by channel through the Scale operation, and the original features are recalibrated in the channel dimension; its expression is: in, is the feature map that rescales the original features.
3. The method according to claim 2, characterized in that The expression of the ResNet residual feature extraction network is: in, is the feature map that rescales the original features, is the residual function, Represents the weight of the first convolution layer in the ResNet residual feature extraction network, Represents the weight of the second convolution layer in the ResNet residual feature extraction network, is the nonlinear function ReLU, Represents the output of the SEResNet feature extraction network after learning.
4. The method according to claim 3, characterized in that The processing process of the AMF module is: Take two monitoring signals as an example, assuming They are the signal features extracted from two different sensor monitoring signals. Represents a three-dimensional real tensor; the specific fusion process is as follows: Step S1: Use the global average pooling operation to compress the signal features from different sensor monitoring signals in the spatial dimension: in, and Represents two signal characteristics respectively In the The compression characteristics of the channels, The spatial dimension of the feature representation is , The number of channels representing features, i Indicates The first i Features, express Middle The characteristics of the channels, express Middle The characteristics of each channel; Step S2: Combine the compressed features of the two sensor signals to generate global representation information ; Its expression is: in, Respectively represent signal characteristics Features after compression; A full connection operation is added after the global information feature, and its expression is: in, is a compact feature after dimensionality reduction, Represents dimensions in the range of real numbers; and Represent weight and bias respectively; Represents the nonlinear function ReLU; Step S3: Based on the compact features after dimensionality reduction , generating an incentive signal with soft attention and , by adding the SoftMax function, the excitation probability of each channel feature is obtained: in, , which correspond to two features respectively. The learnable weight matrix, The dimensions of the matrix are OK, Column, elements are real numbers; and Represents two different learnable weight matrices and , for the same feature To transform, and Respectively represent the corresponding and The incentive signal; Step S4: Each feature from different sensor data is recalibrated and fused by the excitation signal through a gating mechanism: in, represents fusion features; ⊗ represents channel direction product.
5. The method according to claim 4, characterized in that The dimension conversion module converts the dimensions output by different AMF modules into the same one and fuses the converted results; its expression is: in, Indicates The output features of the fusion points are represents the weight coefficient, Indicated in The dimension conversion module converts the dimension of the previous fusion point into the same dimension as the next fusion point, thereby realizing the feature fusion of the output results of different AMF modules.
6. The method according to claim 5, characterized in that During the model training process of step 4: the model is optimized by adjusting the number of model iterations and the learning rate. When the accuracy and loss value curves do not change with the training process, the model training is completed.
7. The method according to claim 6, characterized in that The historical data of the industrial process in step 1 includes: process data of waste heat recovery fan in cold rolling mill, process data of phosphorus removal pump in hot rolling mill or process data of coal mine air compressor; The preprocessing includes: fast Fourier transform and variational mode decomposition.
8. The method according to claim 7, characterized in that In step 2, the ratio of the training set, the validation set and the test set is 7:2:1.
Citation Information
Patent Citations
Motor vibration data processing and state identification method based on multi-scale SE-Resnet
CN113673346A
Circuit breaker fault assessment method based on multi-domain information fusion and deep learning
CN116403032A