A Chemical Process Fault Diagnosis Method Based on Improved Residual Network
By improving the residual network, deep separation convolution, Inception module, channel attention mechanism and spatial attention mechanism are introduced, which solves the problems of low accuracy and poor generalization ability in chemical process fault diagnosis, and achieves higher fault diagnosis accuracy and better data utilization.
Patent Information
- Application Number
- CN202210954740.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-10
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-08-10
AI Technical Summary
The existing chemical process fault diagnosis methods have problems such as low accuracy, poor generalization ability, frequent missed and false alarms, and have failed to effectively utilize the deep information in the chemical process data.
Using an improved residual network, a chemical process fault diagnosis model that can automatically find important information by introducing deep separation convolution, Inception module, channel attention mechanism and spatial attention mechanism is constructed.
It improves the accuracy of chemical process fault diagnosis, reduces missed and false alarms, enhances the generalization ability of the model, and can more effectively utilize the deep information in chemical process data.
Smart Images

Figure CN115457307B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of supervision algorithms and fault diagnosis, and particularly relates to a chemical process fault diagnosis method based on an improved residual network. Background Art
[0002] In recent years, with the large-scale production and development of modern industrial processes, the operational complexity of chemical processes has been continuously increasing, and the situations of aging and corrosion of operating equipment have become increasingly frequent. As a result, unpredictable safety hazards such as chemical substance leakage, fires, and explosions may occur, and at the same time, serious personal safety problems and property damage accidents such as environmental pollution will also be caused. Therefore, fault diagnosis of chemical processes can prevent the occurrence of catastrophic accidents, reduce casualties, and achieve the purpose of ensuring product quality.
[0003] The residual network is a convolutional neural network proposed by four Chinese scholars from Microsoft Research, and won the championships in image classification and object recognition in the ImageNet Large Scale Visual Recognition Challenge in 2015. The main idea of ResNet is to add a direct connection channel in the network, that is, the idea of the Highway Network. The previous network structure performed a non-linear transformation on the performance input, while the Highway Network allows a certain proportion of the output of the previous network layer to be retained. The idea of ResNet is also very similar to the idea of the Highway Network, allowing the original input information to be directly transmitted to the subsequent layers. In this way, the neural network of this layer does not need to learn the entire output, but learns the residual of the output of the previous network. Therefore, ResNet is also called the residual network.
[0004] The Tennessee-Eastman (TE) chemical process created by Eastman Chemical Company mainly provides an actual industrial process control data set for the research field of fault diagnosis and monitoring of chemical processes. At present, using traditional fault diagnosis methods, including methods based on analytical models and methods based on qualitative empirical knowledge, certain progress and achievements have been made in the research on fault detection and diagnosis of the TE data set.
[0005] Tian Wende et al. proposed an online parameter estimation method for a dynamic model for chemical process fault diagnosis. Tian Wende et al. proposed a model-based chemical pipeline leakage fault diagnosis method. Simani et al. proposed a fault diagnosis method for chemical processes based on a robust model, and the diagnosis system is based on the robust estimation of the process output. Iri et al. first proposed applying SDG (Sign Directed Graph) to chemical process fault diagnosis, and Tsuge proposed a chemical process system fault diagnosis algorithm based on SDG and its improvements, namely "SDG with delay" and "Multi-Level SDG (MSDG)". Shiozaki et al. also proposed an improved SDG chemical process fault diagnosis algorithm based on such research. Li Anfeng et al. from Beijing University of Chemical Technology proposed a method for SDG modeling of a chemical process diagnosis model. Gao Dong et al. proposed a chemical process fault diagnosis method that combines SDG with Qualitative Trend Analysis (QTA), uses a two-way reasoning algorithm based on hypothesis and verification to find all possible fault causes and their corresponding consistent paths in the SDG model, then uses an improved QTA algorithm to extract and analyze the trends of the nodes on the consistent paths found in the previous step, and uses a new consistency rule based on qualitative trends to find the true fault cause from the candidate causes. Qian et al. proposed an expert system fault diagnosis method for real-time fault diagnosis of complex chemical processes. Xu et al. proposed an expert system for online fault diagnosis of industrial lubricating oil refining processes.
[0006] When facing chemical process fault diagnosis, the research of existing papers mainly focuses on how to improve traditional fault diagnosis methods, including methods based on analytical models and methods based on qualitative empirical knowledge, to obtain more comprehensive prior information or expert experience and knowledge, so as to construct a more perfect mathematical model and system. There is a lack of reasonable utilization of a large amount of existing chemical process data. Although there are patents and papers applying deep learning and neural networks to achieve fault diagnosis, the existing models fail to automatically find important features in the data and fully explore different levels of information generated in the chemical process and apply them to fault diagnosis.
[0007] Traditional fault diagnosis methods have defects such as long modeling time, low accuracy of fault diagnosis, frequent occurrence of missed reports and false alarms, and low generalization ability. With the improvement of technical means, a large amount of fault data can be saved. Building a model through a data-driven approach does not require an accurate mechanism model and is very suitable for the current complex process industry. Summary of the Invention
[0008] Objective of the Invention: Aiming at the problems pointed out in the background art, the present invention proposes a chemical process fault diagnosis method based on an improved residual network, which constructs a fault diagnosis model in a data-driven manner. At the same time, depthwise separable convolution is applied to the residual network to reduce the number of parameters and save the time for model construction. Then, the Inception module is used to improve the residual unit of the network, enabling the improved residual network to extract deep information of data from different levels. Moreover, a channel attention mechanism and a spatial attention mechanism are introduced into the residual unit, enabling the model to have the ability to autonomously find the location of important information, thereby solving the problems existing in fault diagnosis in chemical processes and improving the accuracy of fault diagnosis.
[0009] Technical Solution: The present invention discloses a chemical process fault diagnosis method based on an improved residual network, which includes the following steps:
[0010] Step 1: Dataset preparation. Collect data for the dataset generated by the chemical process, and store the data simulated under normal conditions and 21 different faults respectively.
[0011] Step 2: Preprocess the collected data. First, use the polynomial dimension elevation method to elevate the feature vectors of the samples, and process the one-dimensional feature vectors into two-dimensional grayscale images. Then, randomly divide the dataset into a training set and a test set.
[0012] Step 3: Improve the architecture of the residual network. Use the pytorch framework to build a network model, and select normal conditions and several fault conditions from the training set to train the model.
[0013] The architecture of the improved residual network is as follows:
[0014] First, extract features from the grayscale image through a convolutional layer of conv1, perform a round of screening on the features through the max pooling layer pool1, and then connect 4 groups of improved residual units, namely the conv2_x convolutional layer, the conv3_x convolutional layer, the conv4_x convolutional layer, and the conv5_x convolutional layer. The number of bottleneck structures contained in each group of convolutional layers is different. The conv2_x convolutional layer contains 2 residual units, the conv3_x convolutional layer contains 4 residual units, the conv4_x convolutional layer contains 6 residual units, and the conv5_x convolutional layer contains 3 residual units. Then, follow a global average pooling layer to perform another round of screening on the features. Next, connect two fully connected layers FC1 and FC2 to mine the laws hidden in the features. Finally, use the softmax classifier to realize the diagnosis of faults and output the results.
[0015] The improved residual unit includes two branches. One branch sequentially includes a 1×1 convolutional layer processed with batch normalization algorithm and Relu activation function, a depthwise separable convolution processed with batch normalization algorithm and Relu activation function, an Inception module for extracting features of different scales, a channel attention layer, and a spatial attention layer. The other branch is a 1×1 skipconv convolutional layer. Finally, the results of the two branches are summed and then processed through a Relu activation function.
[0016] Step 4: Select the test set data corresponding to several fault states involved in the training to evaluate the model effect.
[0017] Further, the specific method of step 1 is as follows:
[0018] Step 1.1: Collect the original data generated by the chemical process, store the data in the normal state in G00, and store the 21 kinds of fault data in G01, G02, ……, G21 respectively.
[0019] Step 1.2: Extract 960 pieces of data from the dataset G00 storing the normal state and store them in the dataset G0_00.
[0020] Step 1.3: There are 21 kinds of fault states. Introduce a fault signal from the 160th piece of data, and extract the data from the 160th to the 960th piece of data in each fault state and store them in G0_01, ……, G0_21 respectively.
[0021] Further, the specific method of step 2 is as follows:
[0022] Step 2.1: Realize polynomial dimension elevation for each sample data in the datasets G0_00, G0_01, ……, G0_21. The original sample data is a feature vector composed of 52 features, that is where represents the s-th feature of the original data sample, and the variable s ∈ [1, 52]. Use the second-order polynomial dimension elevation method to elevate the feature vector of the original sample to obtain a new feature vector where represents the s-th feature of the data sample after feature dimension elevation, and the variable s ∈ [1, 1430];
[0023] Step 2.2: Process the one-dimensional feature vector data after feature dimension elevation into a two-dimensional grayscale image with a size of 38×38, and fill the positions without numbers with 0.
[0024] Step 2.3: Divide the processed two-dimensional grayscale image data into a training set and a test set. Store the normal state and fault state data used for training the model in the training sets G00_tr, G01_tr, ……, G21_tr respectively, and store the normal state and fault state data for testing the model effect in G00_te, G01_te, ……, G21_te respectively.
[0025] Further, in the said step 3, set the convolution kernel size in the conv1 convolutional layer to 3×3, the number of channels to 64, the stride to 2, and the edge padding to 1; set the size of the pool1 max pooling layer to 3×3, the stride to 2, and the edge padding to 1; set the number of neurons in the FC1 fully connected layer to 1024; set the number of neurons in the FC2 fully connected layer to 5.
[0026] Further, the depthwise separable convolution used in the improved residual unit includes using a DepthWise convolution with a convolution kernel size of 3×3, a stride of 1, and an edge padding of 1 to extract feature map information, then using a PointWise convolution with a convolution kernel size of 1×1 and a stride of 1 to fuse the information, and finally outputting the feature map.
[0027] Further, the Inception module used in the improved residual unit includes four branches:
[0028] The first branch is a 1×1 convolutional layer and is processed using the batch normalization algorithm and the Relu activation function;
[0029] The second branch is successively a 1×1 convolutional layer and is processed using the batch normalization algorithm and the Relu activation function, and a 3×3 convolutional layer and is processed using the batch normalization algorithm and the Relu activation function;
[0030] The third branch is successively a 1×1 convolutional layer and is processed using the batch normalization algorithm and the Relu activation function, and a 5×5 convolutional layer and is processed using the batch normalization algorithm and the Relu activation function;
[0031] The fourth branch successively includes a 3×3 max pooling layer and a 1×1 convolutional layer and is processed using the batch normalization algorithm and the Relu activation function;
[0032] The feature maps generated by the four branches will be uniformly aggregated at the output to form a very deep feature map.
[0033] Further, the stride of the 1×1 convolutional layer used in the Inception module in step 3 is 1; the stride of the 3×3 convolutional layer used is 1 and the edge padding is also 1; the stride of the 5×5 convolutional layer used is 1 and the edge padding is 2; the stride of the 3×3 max pooling layer used is 1 and the edge padding is 1.
[0034] Further, the channel attention layer used in the improved residual unit includes two branches. One branch retains the original feature map F, and the other branch includes a channel attention module to distinguish which channels have more important features, obtaining the channel description F'. c , when outputting, the contents of the two branches are multiplied corresponding to each other to obtain a new feature map F', that is, F' = F × F'. c ; the spatial attention layer includes two branches. One branch retains the received feature map F', and the other branch uses a spatial attention module to distinguish the positions of important information in the feature map, obtaining the access description F'. s , when outputting, the contents of the two branches are multiplied corresponding to each other to obtain a new feature map F", that is, F" = F' × F'. s .
[0035] Further, the specific method of the channel attention module used in the channel attention layer is as follows:
[0036] Step 11.1: Use global average pooling AvgPool to process the feature map F with an input size of H×W×C, AvgPool(F), to obtain a channel description F with a size of 1×1×C a , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map;
[0037] Step 11.2: Use global max pooling MaxPool to process the feature map F with an input size of H×W×C, MaxPool(F), to obtain a channel description F with a size of 1×1×C m , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map;
[0038] Step 11.3: Use a multi-layer perceptron MLP with shared parameters to process F a , that is, MLP(F a ), to obtain a new description F'. a ;
[0039] Step 11.4: Use a multi-layer perceptron MLP with shared parameters to process F m , that is, MLP(F m ), to obtain a new description F'. m ;
[0040] Step 11.5: Add the two new channel descriptions F’ a and F’ m to obtain the new description F add , that is, F add = F a ‘ + F’ m ;
[0041] Step 11.6: Process F add using a Sigmoid activation function to obtain the final channel description F c ’, that is, F c ’ = Sigmoid(F add ).
[0042] Furthermore, the specific method of the spatial attention module used in the spatial attention layer is as follows:
[0043] Step 12.1: Process the feature map F’ with input size H×W×C using average pooling AvgPool along the channel dimension AvgPool(F’), to obtain the channel description F a , that is, F a = AvgPool(F’), where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map;
[0044] Step 12.2: Process the feature map F’ with input size H×W×C using max pooling MaxPool along the channel dimension MaxPool(F’), to obtain the channel description F m , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map;
[0045] Step 12.3: Concatenate F a and F m along the channel dimension to obtain the new channel description F con , where H represents the height of the received feature map, W represents the width of the received feature map, and the depth is 2;
[0046] Step 12.4: Process the description F con using a convolutional layer conv_spa with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 on the edges, to obtain the new channel description F s , that is, F s = conv_spa(F con );
[0047] Step 12.5: Use the Sigmoid function to process the channel description F s to obtain the final weight coefficient F’ s , that is, F’ s = Sigmoid(F s ).
[0048] Beneficial effects:
[0049] Based on the dataset generated by the Tennessee Eastman chemical process, the present invention uses an improved residual network to realize fault diagnosis of the chemical process. First, for the Tennessee Eastman chemical process, the simulation data of the chemical process is collected and sorted, and the data generated under normal conditions and various fault conditions are stored in their respective datasets; then the data is preprocessed, the feature dimension of each sample data is increased, and it is processed into a two-dimensional grayscale image, and then the preprocessed data is randomly divided into a training set and a test set; secondly, the residual network is appropriately improved to make it more suitable for processing the data generated by the chemical process. The depthwise separable convolution is used to replace the original convolution method of the residual network, reducing the number of parameters and saving the time for model construction. The Inception module is used to improve the residual unit of the network, so that the improved residual network can extract the deep information of the data from different levels. The channel attention mechanism and the spatial attention mechanism are introduced into the residual unit, so that the model has the ability to autonomously find the location of important information, thus solving the problems existing in fault diagnosis in the chemical process and improving the accuracy of fault diagnosis. Description of the drawings
[0050] Figure 1 is the overall flow chart of the present invention;
[0051] Figure 2 is the flow chart of dataset preparation;
[0052] Figure 3 is the flow chart of dataset preprocessing;
[0053] Figure 4 is the structural schematic diagram of the residual network model;
[0054] Figure 5 is the structural schematic diagram of the improved residual unit;
[0055] Figure 6 is the schematic diagram of the depthwise separable convolution;
[0056] Figure 7 is the structural schematic diagram of the Inception module;
[0057] Figure 8 is the schematic diagram of the channel attention layer;
[0058] Figure 9 It is a schematic structural diagram of the channel attention module;
[0059] Figure 10 It is a schematic diagram of the spatial attention layer;
[0060] Figure 11 It is a schematic structural diagram of the spatial attention module. Specific implementation manners
[0061] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent forms of modification of the present invention all fall within the scope defined by the appended claims of this application.
[0062] See Appendix Figure 1 to Appendix Figure 11 The present invention discloses a chemical process fault diagnosis method based on an improved residual network, which specifically includes the following steps:
[0063] Step 1: Dataset preparation. Collect data for the Tennessee Eastman chemical process, and store the data simulated under normal conditions and 21 different faults in datasets G0_00, G0_01,..., G0_21 respectively, specifically as Figure 2 shown:
[0064] Step 1.1: Collect the original data of the Tennessee Eastman chemical process, store the data under normal conditions in G00, and store the 21 kinds of fault data in G01, G02,..., G21 respectively;
[0065] Step 1.2: Extract 960 pieces of data from the dataset G00 storing normal conditions and store them in the dataset G0_00;
[0066] Step 1.3: There are 21 kinds of fault states. Introduce a fault signal from the 160th piece of data, and extract the data from the 160th to the 960th pieces of data under each fault state and store them in G0_01,..., G0_21 respectively.
[0067] Step 2: Preprocess the collected data. First, use the polynomial dimension elevation method to elevate the feature vectors of the samples, process the one-dimensional feature vectors into two-dimensional grayscale images, and then randomly divide the dataset into a training set and a test set. The new data is stored in the training sets G00_tr, G01_tr,..., G21_tr according to normal conditions and different fault types respectively, and the test set is stored in the datasets G00_te, G01_te,..., G21_te, specifically as Figure 3 shown:
[0068] Step 2.1: Perform polynomial dimensionality increase for each sample data in the G0_00, G0_01, ..., G0_21 dataset. The original sample data is a feature vector D consisting of 52 features, namely in Represents the sth feature of the original data sample, variable s∈[1,52], and uses the second-order polynomial dimension-raising method to increase the dimension of the feature vector of the original sample to obtain a new feature vector in Represents the mth feature of the data sample after feature dimension upgrading, variable m∈[1,1430];
[0069] Step 2.2: Process the one-dimensional feature vector data after feature dimension upgrading into a two-dimensional grayscale image of size 38×38, and fill the positions without numbers with 0;
[0070] Step 2.3: Divide the processed two-dimensional grayscale image data into a training set and a test set. Store the normal state and fault state data used for training the model in the training set G00_tr, G01_tr, ..., G21_tr, respectively, and store the normal state and fault state data for testing the model effect in G00_te, G01_te, ..., G21_te, respectively.
[0071] Step 3: Improve the architecture of the residual network, use the pytorch framework to build a network model, select the normal state and several fault states from the training set to train the model. The specific method is:
[0072] Improve the structure of the residual network such as Figure 4 As shown in the figure, the specific description is: first, a conv1 convolution layer is used to extract features of the grayscale image, and then a maximum pooling layer pool1 is used to screen the features, followed by 4 groups of improved residual units, namely conv2_x convolution layer, conv3_x convolution layer, conv4_x convolution layer, and conv5_x convolution layer. The number of bottleneck structures contained in each group of convolution layers is different. The conv2_x convolution layer contains 2 residual units, the conv3_x convolution layer contains 4 residual units, the conv4_x convolution layer contains 6 residual units, and the conv5_x convolution layer contains 3 residual units. Then, a global average pooling layer AvgPool is used to screen the features again, followed by two fully connected layers FC1 and FC2 to mine the rules hidden in the features, and finally the soffmax classifier is used to diagnose the fault and output the results.
[0073] In this embodiment, the convolution kernel size in the conv1 convolutional layer is set to 3×3, the number of channels is 64, the stride is 2, and the padding is 1; the size of the pool1 max pooling layer is set to 3×3, the stride is 2, and the padding is 1; the number of neurons in the FC1 fully connected layer is set to 1024; the number of neurons in the FC2 fully connected layer is set to 5.
[0074] In this embodiment, the improved residual unit structure is as Figure 5 shown, including two branches. One of the branches is, in sequence, a 1×1 convolutional layer processed with batch normalization algorithm and Relu activation function, a depthwise separable convolution (DSC) layer processed with batch normalization algorithm and Relu activation function, an Inception module for extracting features of different scales, a channel attention layer, and a spatial attention layer. The other branch is a 1×1 skipconv convolutional layer. Finally, the results of the two branches are summed and then processed through a Relu activation function.
[0075] The depthwise separable convolution, specifically as Figure 6 shown, includes using a DepthWise convolutional layer with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 to extract feature map information, and then using a PointWise convolutional layer with a convolution kernel size of 1×1 and a stride of 1 to fuse the information, and finally outputting the feature map.
[0076] The Inception module, specifically as Figure 7 shown, includes four branches. The first branch is a 1×1 convolutional layer processed with batch normalization algorithm and Relu activation function. The second branch is, in sequence, a 1×1 convolutional layer processed with batch normalization algorithm and Relu activation function, and a 3×3 convolutional layer processed with batch normalization algorithm and Relu activation function. The third branch is, in sequence, a 1×1 convolutional layer processed with batch normalization algorithm and Relu activation function, and a 5×5 convolutional layer processed with batch normalization algorithm and Relu activation function. The fourth branch includes, in sequence, a 3×3 max pooling layer, and a 1×1 convolutional layer processed with batch normalization algorithm and Relu activation function. The feature maps generated by the four branches will be aggregated together at the output to form a very deep feature map.
[0077] Specifically, in this embodiment, the stride of the 1×1 convolutional layer used in the Inception module is 1; the stride of the 3×3 convolutional layer used is 1 and the padding is also 1; the stride of the 5×5 convolutional layer used is 1 and the padding is 2; the stride of the 3×3 max pooling layer used is 1 and the padding is 1.
[0078] The channel attention layer used in the improved residual unit is specifically as follows Figure 8 shown, including two branches. One branch retains the original feature map F, and the other branch includes a channel attention module to distinguish which channels have more important features, obtaining the channel description F' c , and when outputting, the contents of the two branches are multiplied correspondingly to obtain a new feature map F', that is, F' = F × F' c .
[0079] The channel attention module used in the channel attention layer is specifically as follows Figure 9 shown. The method is: First, use global average pooling AvgPool to process the feature map F with an input size of H×W×C, AvgPool(F), to obtain a channel description F with a size of 1×1×C a , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map. At the same time, use global max pooling MaxPool to process the feature map F with an input size of H×W×C, MaxPool(F), to obtain a channel description F with a size of 1×1×C m , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map; Second, use a multi-layer perceptron MLP with shared parameters to process F a , that is, MLP(F a ), to obtain a new description F' a , and at the same time, use a multi-layer perceptron MLP with shared parameters to process F m , that is, MLP(F m ), to obtain a new description F' m ; Second, add the two new channel descriptions F' a and F' m to obtain a new description F add , that is, F add = F‘ a + F’ m ; Finally, use a Sigmoid activation function to process F add to obtain the final channel description F' c , that is, F’ c = Sigmoid(F add ).
[0080] The spatial attention layer used in the improved residual unit is specifically as follows Figure 10 shown, including two branches. One branch retains the received feature map F', and the other branch resolves the position of important information in the feature map through a spatial attention module, obtaining the channel description F' s, when outputting, the contents of the two branches are multiplied correspondingly to obtain a new feature map F", that is, F" = F' × F' s .
[0081] The spatial attention module used in the spatial attention layer, specifically as Figure 11 shown, the method is: First, use an average pooling AvgPool with a channel dimension to process the feature map F' with an input size of H×W×C, AvgPool(F'), to obtain a channel description with a size of H×W×1 That is where H represents the height of the received feature map, W represents the width of the received feature map, C represents the depth of the received feature map. At the same time, use a max pooling MaxPool with a channel dimension to process the feature map F' with an input size of H×W×C, MaxPool(F'), to obtain a channel description with a size of H×W×1 where H represents the height of the received feature map, W represents the width of the received feature map, C represents the depth of the received feature map; then, and are concatenated according to the channel to obtain a new channel description F with a size of H×W×2 con , where H represents the height of the received feature map, W represents the width of the received feature map, and the depth is 2; secondly, use a convolutional layer conv_spa with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 on the edge to process the description F con to obtain a new channel description F s , that is, F s = conv_spa(F con ); finally, use the Sigmoid function to process the channel description F s to obtain the final weight coefficient F' s , that is, F' s = Sigmoid(F s ).
[0082] For step 3, training the improved residual network, the specific method is:
[0083] First, select the chemical process training dataset G00_tr in the normal state, the chemical process training dataset G01_tr in the fault 1 state, the chemical process training dataset G04_tr in the fault 4 state, the training dataset G08_tr in the fault 8 state, and the training dataset G12_tr in the fault 12 state as the training dataset of the model; secondly, use the Adam optimizer and the cross-entropy loss function to optimize the model. Among them, the learning rate of the Adam optimizer is set to 0.0002, and a total of 100 rounds are trained. The formula of the cross-entropy loss function is Among them, m represents the number of samples used in one training; y i represents the true probability distribution; represents the model's predicted probability distribution.
[0084] Step 4: Select the test set data corresponding to several fault states involved in the training to evaluate the model effect. The specific method is as follows:
[0085] Select the chemical process test data set G00_te under normal conditions, the chemical process test data set G01_te under fault 1 state, the chemical process test data set G04_te under fault 4 state, the test data set G08_te under fault 8 state, and the test data set G12_te under fault 12 state to test the model effect. The experimental results show that the upper limit of the accuracy of the improved residual network in fault diagnosis under different states is higher than 94.7%, and there are fewer outliers, proving that the fault diagnosis performance of the model on the chemical process data set is relatively stable and has high accuracy. Faults 8 and 12 are faults generated under different perturbations, but the fault types are both related to random changes. Such faults will show large fluctuations, while in the improved residual network model, the upper limit of the accuracy is higher than 96.3%. The experimental results are good.
[0086]
[0087]
[0088] In summary, the present invention can be combined with the data generated by the chemical process, and utilize the strong information mining ability of the improved residual network for complex data, and can solve the problems of low accuracy, frequent false alarms and missed alarms, and low generalization ability in the fault diagnosis of the process industry.
[0089] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A chemical process fault diagnosis method based on an improved residual network, characterized in that The steps include: Step 1: Dataset preparation, collect data from the dataset generated by the Tennessee Eastman Chemical process, and store the data simulated under normal conditions and 21 different faults separately; Step 2: Preprocess the collected data. First, use the polynomial dimension-raising method to increase the dimension of the sample's feature vector, and process the one-dimensional feature vector into a two-dimensional grayscale image. Then, randomly divide the data set into a training set and a test set. Step 3: Improve the architecture of the residual network, use the pytorch framework to build the network model, and select the normal state and several fault states from the training set to train the model; The architecture of the improved residual network is: First, the grayscale image is feature extracted through a convolution layer of conv1, and the features are screened through a round of the maximum pooling layer pool1, followed by 4 groups of improved residual units, namely conv2_x convolution layer, conv3_x convolution layer, conv4_x convolution layer, and conv5_x convolution layer. Each group of convolution layers contains a different number of bottleneck structures. The conv2_x convolution layer contains 2 residual units, the conv3_x convolution layer contains 4 residual units, the conv4_x convolution layer contains 6 residual units, and the conv5_x convolution layer contains 3 residual units. Then, a global average pooling layer is used to screen the features again, followed by two fully connected layers FC1 and FC2 to mine the rules hidden in the features, and finally the softmax classifier is used to diagnose the fault and output the results. The improved residual unit includes two branches, one branch is a 1×1 convolution layer and is processed by batch normalization algorithm and Relu activation function, a depth-separable convolution and is processed by batch normalization algorithm and Relu activation function, an Inception module to extract features of different scales, a channel attention layer, and a spatial attention layer; the other branch is a 1×1 skipconv convolution layer, and the results of the last two branches are summed and then processed by Relu activation function; Step 4: Select test set data corresponding to several fault conditions involved in training to evaluate the model effect.
2. The chemical process fault diagnosis method based on the improved residual network according to claim 1, wherein, The specific method of step 1 is: Step 1.1: Collect the raw data generated by the Tennessee Eastman Chemical process, store the data in the normal state in G00, and store the 21 types of fault data in G01, G02, ..., G21 respectively; Step 1.2: Extract 960 data from the data set G00 storing the normal state and store them in the data set G0_00; Step 1.3: There are 21 fault states in total. The fault signal is introduced from the 160th data. The 160th to 960th data under each fault state are extracted and stored in G0_01, ..., G0_21 respectively.
3. The chemical process fault diagnosis method based on the improved residual network according to claim 2, wherein, The specific method of step 2 is: Step 2.1: Implement polynomial dimensionality elevation for each sample data in the datasets G0_00, G0_01, ……, G0_21. The original sample data is a feature vector composed of 52 features, that is where represents the s-th feature of the original data sample, and the variable s ∈ [1, 52]. Use the second-order polynomial dimensionality elevation method to perform dimensionality elevation on the feature vector of the original sample, and obtain a new feature vector where represents the s-th feature of the data sample after feature dimensionality elevation, and the variable s ∈ [1, 1430]; Step 2.2: Process the one-dimensional feature vector data after feature dimension upgrading into a two-dimensional grayscale image of size 38×38, and fill the positions without numbers with 0; Step 2.3: Divide the processed two-dimensional grayscale image data into a training set and a test set. Store the normal state and fault state data used for training the model in the training sets G00_tr, G01_tr, ……, G21_tr respectively, and store the normal state and fault state data for testing the model effect in G00_te, G01_te, ……, G21_te respectively.
4. The chemical process fault diagnosis method based on the improved residual network according to claim 1, characterized in that In step 3, set the convolution kernel size in the conv1 convolutional layer to 3×3, the number of channels to 64, the stride to 2, and the padding to 1; set the size of the pool1 max-pooling layer to 3×3, the stride to 2, and the padding to 1; set the number of neurons in the FC1 fully connected layer to 1024; set the number of neurons in the FC2 fully connected layer to 5.
5. The chemical process fault diagnosis method based on the improved residual network according to claim 1, characterized in that The depthwise separable convolution used in the improved residual unit includes using a DepthWise convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 to extract feature map information, then using a PointWise convolution with a convolution kernel size of 1×1 and a stride of 1 to fuse the information, and finally outputting the feature map.
6. The chemical process fault diagnosis method based on the improved residual network according to claim 1, characterized in that, The Inception module used in the improved residual unit includes four branches: The first branch is a 1×1 convolutional layer and is processed using batch normalization algorithm and Relu activation function; The second branch is successively a 1×1 convolutional layer and is processed using batch normalization algorithm and Relu activation function, and a 3×3 convolutional layer and is processed using batch normalization algorithm and Relu activation function; The third branch is successively a 1×1 convolutional layer and is processed using batch normalization algorithm and Relu activation function, and a 5×5 convolutional layer and is processed using batch normalization algorithm and Relu activation function; The fourth branch successively includes a 3×3 max-pooling layer and a 1×1 convolutional layer and is processed using batch normalization algorithm and Relu activation function; The feature maps generated by the four branches will be uniformly aggregated at the output to form a very deep feature map.
7. The chemical process fault diagnosis method based on the improved residual network according to claim 1, characterized in that In step 3, the stride of the 1×1 convolutional layer used in the Inception module is 1; the stride of the 3×3 convolutional layer used is 1 and the padding is also 1; the stride of the 5×5 convolutional layer used is 1 and the padding is 2; The stride of the 3×3 max-pooling layer used is 1 and the padding is 1.
8. The chemical process fault diagnosis method based on the improved residual network according to claim 1, characterized in that The channel attention layer used in the improved residual unit includes two branches. One branch retains the original feature map F, and the other branch includes a channel attention module to distinguish which channels have more important features, obtaining the channel description F c ’. When outputting, the contents of the two branches are multiplied corresponding to each other to obtain a new feature map F', that is, F' = F × F c ’; The spatial attention layer includes two branches. One branch retains the received feature map F', and the other branch uses a spatial attention module to distinguish the positions of important information in the feature map, obtaining the access description F s '. When outputting, the contents of the two branches are multiplied corresponding to each other to obtain a new feature map F”, that is, F” = F' × F s '.
9. The chemical process fault diagnosis method based on the improved residual network according to claim 8, characterized in that, The specific method of the channel attention module used in the channel attention layer is as follows: Step 11.1: Process the feature map F with input size H×W×C using global average pooling AvgPool(F) to obtain the channel description F with size 1×1×C, where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map; a , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map; Step 11.2: Process the feature map F with input size H×W×C using global max pooling MaxPool(F) to obtain a channel description F with size 1×1×C, where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map; m , where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map; Step 11.3: Process F using a multi-layer perceptron MLP with shared parameters a , that is, MLP(F a ), to obtain a new description F a '; Step 11.4: Process F using a multi-layer perceptron MLP with shared parameters m , that is, MLP(F m ), to obtain a new description F m '; Step 11.5: Add the two new channel descriptions F a ' and F m ' to obtain the new description F add , i.e., F add = F a ‘ + F′ m ; Step 11.6: Use a Sigmoid activation function to process F add to obtain the final channel description F c ’, that is, F c ’ = Sigmoid(F add ).
10. The chemical process fault diagnosis method based on the improved residual network according to claim 8, characterized in that The specific method of the spatial attention module used in the spatial attention layer is as follows: Step 12.1: Process the feature map F' with input size H×W×C using average pooling AvgPool in one channel dimension, AvgPool(F'), to obtain the channel description F with size H×W×1, i.e., F a , that is, F a = AvgPool(F'), where H represents the height of the received feature map, W represents the width of the received feature map, and C represents the depth of the received feature map; Step 12.2: Process the feature map F' with input size H×W×C using max pooling MaxPool in one channel dimension MaxPool(F'), to obtain a channel description F with size H×W×1 m , where H represents the height of the receptive feature map, W represents the width of the receptive feature map, and C represents the depth of the receptive feature map; Step 12.3: Concatenate F a and F m along the channels to obtain a new channel description F con with the size of H×W×2, where H represents the height of the received feature map, W represents the width of the received feature map, and the depth is 2; Step 12.4: Process the description Fcon using a convolutional layer conv_spa with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 at the edges to obtain a new channel description F s , that is, F s = conv_spa(F con ); Step 12.5: Process the channel description F s to obtain the final weight coefficient F s ', that is, F s ' = Sigmoid(F s ).