Normalized flow-based zero-fault sample chemical process fault detection method
By adopting a zero-failure sample chemical process fault detection method based on normalized flow in the chemical process, using the TNF network and the improved RealNVP model, the problem of lack of fault data in the traditional method is solved, and efficient fault detection and identification is achieved.
Patent Information
- Application Number
- CN202510212765.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-20
AI Technical Summary
During the chemical process, traditional fully supervised deep learning methods need to have both normal data and fault data during the training stage, making it difficult to collect enough fault data in practical applications to train the model.
Using the zero-failure sample chemical process fault detection method based on normalized flow, the timing characteristics are extracted through the TNF network and the improved RealNVP model using the long and short-term memory network, and the fault detection capability is improved through the improved KAN network model.
It realizes fault detection of chemical process without fault data, reduces data acquisition costs, improves production deployment efficiency, and improves the accuracy of fault detection.
Smart Images

Figure CN120180077A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial process fault detection, and specifically relates to a chemical process fault detection method based on normalizing flow for zero-fault samples. Background Art
[0002] In recent years, the rapid development of science and technology has made major production processes more complex and large-scale. In the context of the big data era, the wide application of automation technology has accumulated a vast amount of data resources at chemical sites, making data-driven method modeling more applicable. This method can directly and effectively perform statistical analysis and information extraction on massive, multi-source, and high-dimensional data. By collecting monitoring data from different sources and of different types, and using various data mining techniques, the implicit useful information is extracted therefrom to characterize the normal mode and fault mode of system operation.
[0003] In the field of fault monitoring and diagnosis in the chemical industry, deep learning technology has shown great potential. However, traditional fully supervised deep learning methods have encountered a challenge that cannot be ignored during the training stage: they require both normal data and fault data. This means that in order to build an effective classification model, researchers not only need to extract the features of normal data, but also cover all possible fault types and extract the features of these fault data. For example, in the "Industrial Process Fault Diagnosis Method Based on Dynamic Time Warping and Graph Convolutional Network" with the publication number CN113110398A, normal data and fault data are required during training. Specifically, the normal data in the simulation experiment dataset is processed into an adjacency matrix through data, and the fault data is processed into a node feature matrix through data. The adjacency matrix and the node feature matrix are used to train and test the DTW-GCN model. In the actual data collection process, normal data, since it represents the normal operation state of the system, is relatively easy to obtain and is in a large quantity. However, faults are often small-probability events that occur during the operation of the system, and once they occur, immediate measures need to be taken for repair to avoid greater damage to the system. Therefore, in practical applications, it is difficult for us to collect a sufficient amount of fault data to train a fully supervised deep learning model.
[0004] Unsupervised anomaly detection mainly relies on a large number of normal samples to build a model. In contrast, supervised learning requires a large amount of labeled data to train a model, including input features and corresponding output labels. It is applicable to situations where abnormal events are relatively few, but once they occur, they may have a major impact on the system or business, while supervised learning is applicable to situations where there is sufficient labeled data and there is a clear relationship between the labels and the input features. For the practical application of chemical processes, unsupervised anomaly detection is more extensive and flexible than supervised learning. Therefore, it is necessary to optimize and improve the fault detection method for chemical processes combined with deep learning. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a zero-fault sample chemical process fault detection method based on normalizing flow for online real-time fault detection and identification of chemical process data.
[0006] To solve the above technical problem, the present invention provides a zero-fault sample chemical process fault detection method based on normalizing flow. The process includes: collecting material parameters in the chemical production process, preprocessing the collected real-time data, and then inputting it into the offline-trained TNF network. The input sample data first obtains temporal features through a long short-term memory network, and then obtains the latent data distribution through an improved normalizing flow model. The latent data distribution is compared with the standard normal distribution. If the difference from the standard normal distribution is less than a given threshold, it is determined as a normal sample; otherwise, it is a fault sample.
[0007] As an improvement of the zero-fault sample chemical process fault detection method based on normalizing flow of the present invention:
[0008] The improved normalizing flow model is based on the RealNVP model. The coupling layer of the RealNVP model is replaced with an improved KAN network to model the scaling function and translation function instead of the original MLP neural network.
[0009] As a further improvement of the zero-fault sample chemical process fault detection method based on normalizing flow of the present invention:
[0010] The improved KAN network includes an input layer, three hidden layers, and an output layer. The number of layers increases step by step and then decreases to the same as the input dimension. The number of input neurons in the input layer is the same as the number of output neurons in the output layer. The number of neurons in the first hidden layer is the same as the number of neurons in the third hidden layer, and the number of neurons in the second hidden layer is the largest.
[0011] As a further improvement of the zero-fault sample chemical process fault detection method based on normalizing flow of the present invention:
[0012] The process of offline training of the TNF network includes:
[0013] Using the Tennessee-Eastman process to generate the TEP dataset, only selecting all normal data from the training set of the TEP dataset as the training set for offline training of the TNF network, and using the test set of the TEP dataset as the test set for offline training of the TNF network; then preprocessing all data in the training set and test set for offline training of the TNF network and manually annotating labels.
[0014] Input the samples in the training set into the TNF network one by one. The optimization algorithm adopts the adaptive momentum estimation algorithm, and an early stopping mechanism is added to prevent overfitting. During the training process, calculate the loss function value and backpropagate to iteratively optimize the model parameters. End the training after reaching the preset number of epochs.
[0015] Input the test set into the trained TNF network, judge the classification result according to the given threshold, compare the difference between the predicted result and the true label, and use the misjudgment rate and the detection rate as evaluation indicators. Select the model with the lowest misjudgment rate and the highest detection rate from the test results for online detection.
[0016] As a further improvement of the zero-fault sample chemical process fault detection method based on normalizing flow of the present invention:
[0017] The loss function is:
[0018] log p(x)=log q(f(x))+log|det▽f(x)| (10)
[0019] Among them, |def▽f(x)| is the Jacobian determinant.
[0020] As a further improvement of the zero-fault sample chemical process fault detection method based on normalizing flow of the present invention:
[0021] The given threshold is: the loss function value of the optimal model obtained during the training process.
[0022] As a further improvement of the zero-fault sample chemical process fault detection method based on normalizing flow of the present invention:
[0023] The preprocessing includes: normalization and moving sliding window interception;
[0024] The calculation of normalization is:
[0025]
[0026] Among them: Data is the fault sample data without preprocessing, μ is the mean of the normal sample data, s is the variance of the normal sample data, Data* is the normalized data, and the value of Data* is in the interval (-1, 1);
[0027] Moving sliding window interception is to slide and intercept along the time series direction with a moving window of fixed width and step size.
[0028] The beneficial effects of the present invention are mainly reflected in:
[0029] 1. The present invention adopts an unsupervised learning method to train the model offline, ensuring that only normal samples are required to train the model, reducing the cost of data collection, and thus improving the production deployment efficiency.
[0030] 2. The TNF model of the present invention uses the KAN network as a non-linear function to improve the fault detection ability, reduce the number of model parameters and computational complexity, and at the same time has good generalization performance.
[0031] 3. Compared with the KPCA, PCA-SVDD, and DSAEN methods, the misjudgment rate of the TNF model of the present invention is comparable, and the detection rates of fault types 3, 9, 10, 11, 13, 15, 18, and 20 have been significantly improved, and the average detection rate has also increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The following further describes the specific embodiments of the present invention in conjunction with the accompanying drawings.
[0033] Figure 1 Schematic diagram of a zero-fault sample chemical process fault detection method based on normalizing flow of the present invention;
[0034] Figure 2 Schematic diagram of the structure of the TNF network of the present invention;
[0035] Figure 3 Structure comparison diagram of MLP network (a) and KAN network (b);
[0036] Figure 4 Network structure diagram of Realnvp;
[0037] Figure 5 Schematic diagram of the improved KAN network structure of the present invention;
[0038] Figure 6 Structure diagram of LSTM;
[0039] Figure 7 Flow chart of offline training and testing of the TNF model of the present invention;
[0040] Figure 8 Visualization result diagram of each fault detected by the TNF model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The following further describes the present invention in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:
[0042] Example 1. A zero-fault sample chemical process fault detection method based on normalizing flow. First, a chemical process fault monitoring network of normalizing flow with time series features of the present invention (hereinafter referred to as TNF network) is constructed. The data for offline modeling is the data in the Tennessee Eastman (TE) process. The data in the TE process is preprocessed and training sets and test sets are constructed. The modeling process is completed through offline training of the model. Then, the trained and tested TNF model is applied to the actual chemical production process for fault detection. The specific process is as Figure 1 shown.
[0043] 1. Construct the TNF network
[0044] The TNF network is constructed for chemical process fault monitoring, including a long short-term memory (LSTM) network and a normalizing flow model. The TNF network structure is as Figure 2 shown.
[0045] 1.1. Normalizing flow model
[0046] Normalizing flow (NF) maps a complex data distribution to a simple distribution (such as a Gaussian distribution), and maintains the reversibility of the transformation during this process. Abnormal samples are identified by comparing the probability densities of latent variables. During the use of an autoencoder (AE), statistics are constructed based on the covariance matrix and the mean of the latent variables, and control limits are set to identify abnormal samples. During the training process, the main focus is on the reconstruction error of the data, rather than learning the probability distribution of the data. As a result, the learned latent variables are random and uncontrollable, which affects the construction of statistics and thus the identification effect. Normalizing flow (NF) can achieve accurate probability density estimation and is differentiable on the probability density function, and can more accurately simulate the data distribution. Therefore, the ability of normalizing flow (NF) to process complex data is higher than that of AE, and the limitation and unique reversibility of normalizing flow (NF) for latent variables enable the model to learn the specific true distribution of the data, with strong interpretability. The present invention uses normalizing flow (NF) technology to further process the time series features obtained by the long short-term memory (LSTM) network.
[0047] RealNVP (Real-valued Non-Volume Preserving Transformations) is a model based on normalizing flow (NF). Its structure mainly includes an input layer, a series of reversible transformation layers (i.e., coupling layers), and an output layer. In RealNVP, the coupling layer plays a core role, mapping the input data to the latent space through reversible transformations while maintaining the probability distribution characteristics of the data. The RealNVP structure is as Figure 4As shown, each coupling layer is an affine transformation. The affine transformation converts the determinant of the derivative of the output corresponding to the input into a diagonal product, achieving the effect of simplifying calculations. Multiple coupling layers are stacked to evenly process each dimension of the input. The input of the coupling layer is the data X(x1:n), and the data X is divided into two parts, X1(x1:d) and X2(xd+1:n). X2 is processed through a coupling layer to obtain s and t, where s represents the scale function and t represents the translate function. To enhance the expressive power, a deep neural network (MLP is used for processing one-dimensional data, and CNN is used for processing image data) is usually used to model the scale function and the translate function. The input of the deep neural network is X2, and the output is the corresponding scale and translation parameters. (In the chemical process fault detection process, RealNVP processes time series data, and an MLP neural network is used in the coupling layer to model the scale function and the translate function). The formula for the affine transformation is:
[0048] Z2 = X2, Z1 = X1 ⊙ exp(s(X2)) + t(X2) (1)
[0049] where ⊙ represents element-wise multiplication.
[0050] The Z1 and Z2 after the affine transformation are recombined to form a new data representation Z. This process can be stacked in multiple layers, and each layer contains one or more coupling layers. In Figure 4 , it is represented by the functions f1, f2, … fn, and each function represents an affine transformation. After the transformation of all coupling layers, the data is mapped to the latent space Z, which is usually assumed to be a Gaussian distribution. This structure allows the model to learn the complex distribution of the data, and due to the reversibility of the transformation, accurate density estimation and sample generation can be performed. According to the probability density transformation, we can obtain:
[0051] p(x) = q(z) × |det▽f(x)| (2)
[0052] where |det▽f(x)| is the Jacobian determinant.
[0053] 1.2. KAN Network
[0054] The multi-layer perceptron (MLP) is a fundamental deep learning model composed of multiple fully connected layers, including at least one input layer, one or more hidden layers, and an output layer. The MLP learns the non-linear features of data by using weights and biases between layers. Different from neural network models with fixed activation functions, inspired by the Kolmogorov-Arnold theorem, KAN replaces the weight parameters on the edges of the MLP with univariate functions and parameterizes them in the form of spline functions, transforming the traditional form of "learnable weights + fixed activation function" into the form of "learnable activation function + function summation". The network structures of the MLP and KAN are as Figure 3 shown.
[0055] The general form of the KAN network can be expressed as
[0056]
[0057] where, represents the univariate function that maps the input variable, represents the trainable activation function.
[0058] In the Kolmogorov-Arnold theorem, the internal function forms a KAN layer with n inputs and 2n + 1 outputs, and the output function forms a KAN layer with 2n + 1 inputs and n outputs. Therefore, KAN(x) essentially consists of two KAN layers. To make KAN easier to optimize, Liu et al. proposed a residual activation strategy [1] , by introducing the basis function b(x), making the trainable activation function the sum of the basis function b(x) and the spline function spline(x):
[0059]
[0060] where, w b and w s are essentially redundant and can be incorporated into the b(x) and spline(x) functions. Liu et al. set b(x) as the SiLU function and parameterize spline(x) as a linear combination of B-spline curves:
[0061]
[0062] spline(x) = ∑ i c i B i (x) (6)
[0063] where, c i represents the trainable parameter, B i(x) represents a certain spline curve. Therefore, the training process of KAN is to learn the weight coefficients c of the spline function i of the process. Since the spline function is more difficult to train than the weights of the MLP, Xu et al. used Fourier coefficients to decompose the spline curve into multiple relatively simple non-linear functions [2] , at this time The expression of is modified to:
[0064]
[0065] In the formula, d represents the feature dimension, a ik and b ik are trainable Fourier coefficients, g represents the grid size, which is used to control the number of terms used in the Fourier series expansion, that is, to determine how many different sine and cosine term functions are included in the Fourier coefficients of each input dimension.
[0066] The improved normalizing flow model of the present invention is based on the RealNVP model, and replaces the traditional MLP neural network in the coupling layer of RealNVP with the improved KAN network, as Figure 5 shown, that is, the coupling layer of RealNVP of the present invention uses the KAN network to model the scale function and the translate function. At the same time, according to the characteristics of the KAN network, a stepped structure in which the number of layers increases step by step and then decreases to the same as the input dimension is designed, which helps the model find a balance between deep feature extraction and final output, and can both capture sufficient rich feature information and avoid the risks of over-parameterization and overfitting. The settings of the input layer, output layer and hidden layer of the MLP network and the KAN network are roughly the same, so the structure can be directly replaced.
[0067] The improved KAN network of the present invention includes an input layer, three hidden layers, and an output layer. The number of input neurons in the input layer is the same as the number of output neurons in the output layer. The number of neurons in the first hidden layer is the same as the number of neurons in the third hidden layer, and the number of neurons in the second hidden layer is the largest. MLP uses fixed activation functions on neurons (nodes), such as ReLU, sigmoid, or tanh, etc. These activation functions are usually non-linear and are used to introduce the non-linear expression ability of the network. KAN, on the other hand, uses learnable activation functions on the edges (weights) of the network, which means that each weight parameter can be replaced by a univariate function, usually parameterized as a spline function. This design makes each weight in KAN a learnable function rather than a fixed value. The complex transformation originally performed in the coupling layer is now achieved by neurons in the KAN network through simple weighted summation and learnable functions. This replacement not only retains the reversibility and probability distribution preservation characteristics of the RealNVP model but also introduces the high interpretability and flexibility of the KAN network. Through this replacement, a new neural network model is constructed that not only retains the advantages of the RealNVP model but also has the high interpretability and accuracy of the KAN network. Most importantly, it can reduce the model parameters.
[0068] 1.3 Long Short-Term Memory (LSTM) Network
[0069] The LSTM network is a special type of Recurrent Neural Network (RNN) specifically designed to address the problem of long sequence dependencies. For a general feedforward neural network, there is no feedback connection between neurons in each layer, that is, the current input does not pay attention to previous information. However, in a time series, the information at different times is very likely to influence each other. Reasonably using the input at time t and previous inputs to process the information at time t can obviously learn features more reasonably and fully. The LSTM network has three inputs: the input value x at the current moment t , the value h of the hidden state at the previous moment t-1 , and the cell state C at the previous moment t-1 . Among them, the cell state is the memory unit in the LSTM network and is used to store previous information and calculate new information. The LSTM network controls the flow of the cell state through three gates, namely the forget gate, the input gate, and the output gate. First, two gates are used to control the content of the cell state. One is the forget gate, which determines how much of the hidden state at the previous moment is retained to the current moment; the other is the input gate, which determines how much of the current input x t is stored in the cell state C t . Next, according to the current input xt and cell state C t , determines what information to output to the next layer or the output layer. The LSTM network can generate a state X at each time step that contains information from all previous time steps. The structure is as Figure 6 shown.
[0070] The calculation process of the LSTM network is shown in Equation (8):
[0071]
[0072] where f t , i t , C t , o t , h t are the values of the forget gate, input gate, candidate memory cell, cell state (long-term memory output), output gate, and hidden state (short-term memory output) at time step t, respectively. W f , W i , W C , W o are the weights of the forget gate, input gate, candidate memory cell, and output gate, respectively. b f , b i , b C , b o are the biases of the forget gate, input gate, candidate memory cell, and output gate, respectively. h t-1 is the value of the hidden state of the LSTM network at time step t - 1, and x t is the input of the LSTM network at time step t.
[0073] Combining historical data with the data to be detected is a common time series analysis method. The core of this method lies in constructing a comprehensive data sample set that can cover the data changes over a period of time, thus providing rich information for the model. The LSTM network can effectively process and analyze time series data and provide temporal information for the model.
[0074] In summary, the input sample data obtains temporal features through a long short-term memory (LSTM) network and then obtains the latent data distribution through a normalizing flow model (improved RealNVP model). After comparison with the standard normal distribution, if the difference from the standard normal distribution is less than a given threshold, it is determined to be a normal sample; otherwise, it is a faulty sample.
[0075] 2. Model Training
[0076] 2.1. Establishment of the Dataset
[0077] The present invention uses the TEP dataset generated by the Tennessee - Eastman process (TEP) provided by Lawrence Ricker et al. of the University of Washington as the dataset for offline training of the TNF network. The TEP dataset contains 52 process variables, including 41 measured variables and 11 manipulated variables. These variables come from multiple simulations of a chemical process consisting of a reactor, a condenser, a compressor, a separator, and a stripper.
[0078] The training set and the test set of the generated TEP dataset each include 1 set of normal data and fault data of 21 different fault types. The training set in the TEP dataset is obtained under a 25 - hour simulation, with data collected once every 3 minutes. Fault samples are introduced 1 hour after the simulation starts. The simulation duration of the test set is 48 hours, and fault samples are introduced 8 hours after the simulation starts. The specific fault types are shown in Table 1:
[0079] Table 1 Fault Descriptions of the Tennessee Eastman Process
[0080]
[0081]
[0082] The present invention only selects all the normal data from the training set of the TEP dataset for training the TNF network, and uses all the test data of the TEP dataset for testing of the present invention. The data for training contains a total of 500 normal - condition data, and the data for testing contains 960 normal - condition data and 21 groups, where the first 160 are normal - condition data and 800 are fault data.
[0083] Before training, all the data in the training set and the test set of the present invention are pre - processed, including normalization and moving - window truncation. The calculation process of normalization is shown in Equation (8):
[0084]
[0085] In the formula: Data is the fault - sample data without pre - processing, μ is the mean of the normal - sample data, s is the variance of the normal - sample data, and Data* is the normalized data. After normalization pre - processing, the values of each variable in the samples of the training set and the test set are all in the interval (-1, 1), reducing the numerical gap between different variables.
[0086] Then, a moving window with a width of L and a step size of S (in the present invention, L = 25, S = 1) is slid along the time series direction for intercepting. In the training set and test set of the present invention, the data corresponding to each fault type for each simulation is processed into (T - L)+1 multivariate time series data with an L×N structure. Wherein, N is the number of process variables (in the present invention, N = 52), T is the number of fault data of each fault type in each simulation. Therefore, there are 476 (500 - 25 + 1) multivariate time series data in the training set of the present invention. The test set has (960 - 25 + 1)*22 = 20592 multivariate time series data.
[0087] 2.2. Offline training and testing process
[0088] The offline training and testing process is as Figure 7 shown.
[0089] Training process: First, the preprocessed training set is used to train the model. The training epoch is set to 500, the batch size is set to 10, the optimization algorithm adopts the adaptive momentum estimation (Adam) algorithm, the preset learning rate is set to 0.001, and an early stopping mechanism is added to prevent overfitting. During the training process, the loss function value is calculated and the model parameters are iteratively optimized by backpropagation. After reaching the preset number of epochs, the training ends.
[0090] Since the training only has normal samples, our model adopts the method of learning the distribution characteristics of the data. A probability density estimation model is used to learn the true distribution of the data of normal samples, and the maximum likelihood estimation loss function is used to approximately learn the true distribution of the data. The output is not one-hot encoding, but probability density estimation. The corresponding loss function is not the cross-entropy function, but the maximum likelihood estimation. The specific loss function is shown in Equation (9):
[0091] log p(x) = log q(f(x)) + log|det▽f(x)| (10)
[0092] Where |det▽f(x)| is the Jacobian determinant, and the Jacobian determinant represents the scaling factor of the probability density transformation.
[0093] Testing process: First, set the threshold, and the size of the threshold is the numerical value of the loss function value obtained from the training. According to the threshold, it is judged whether the input sample is a fault sample. The false positive rate and the detection rate are used as evaluation indicators, and the model with the lowest false positive rate and the highest detection rate is selected from the test results for online detection.
[0094] 3. Online use of the trained TNF model
[0095] To achieve on-line detection of chemical faults, on the chemical production site, material parameters in the chemical production process are regularly collected through technical means such as sensors and computers. The collected real-time data is preprocessed through standardization and moving sliding window, where the width L of the moving window is 25 and the step size S is 1. The preprocessed data is input into the TNF network that can be used on-line obtained in step 2, so as to obtain the fault monitoring results of the real-time data. The collected data is saved for the maintenance and update of the model.
[0096] Experiment:
[0097] 1. Evaluation indicators
[0098] The experimental evaluation indicators include False Alarm Rate and Fault Detection Rate. The formulas are shown in Formulas (10) and (11) as follows:
[0099]
[0100] Among them, n is the number of positive sample data predicted as negative samples, and m is the total number of positive samples.
[0101]
[0102] Among them, p is the number of negative sample data predicted as negative samples, and q is the total number of negative samples.
[0103] 2. Comparative experiment of moving window
[0104] To verify the influence of different width settings of the moving window on the model of the present invention, the widths are respectively set to 10, 20, 25, 30, and 40 for testing. The comparison test results are shown in Table 2.
[0105] Table 2 Performance comparison of moving windows with different settings
[0106]
[0107]
[0108] As can be seen from Table 2, on dataset 1 after sliding and intercepting processing along the time series direction using a moving window with a width of 25, the indicators of the TNF model of the present invention perform the best, and the detection rate reaches 84.53%.
[0109] 3. Comparative experiment
[0110] The TNF model of the present invention was compared with four other mainstream process monitoring models (KPCA, EKPCA, PAC - SVDD, and EKCVA) in terms of performance. The experimental comparison test results are shown in Tables 3 and 4. To reduce the randomness of the experimental results, each experiment was repeated 8 times, and the obtained results were averaged.
[0111] Compared with the classical KPCA, PCA - SVDD combined with machine learning, and the DSAEN method based on AE, the misjudgment rate of the TNF model of the present invention is comparable. The detection rates for more than 20 types of faults such as Fault 3, 9, 10, 11, 13, 15, 18, and 20 have been significantly improved, and the average detection rate has also increased.
[0112] Table 3 Comparison results of the fault detection rates (FDR) of fault samples between the mainstream fault monitoring models
[0113]
[0114]
[0115] Table 4 Comparison results of the false alarm rates (FAR) of normal samples between the mainstream fault monitoring models
[0116]
[0117] 4. Ablation experiments
[0118] The TNF model of the present invention is based on the RealNVP framework. The improved coupling layer uses the KAN network and adds LSTM to extract temporal features, thereby improving the fault detection rate. To verify the effectiveness of these improvements, ablation experiments were conducted to compare the fault detection rates and the number of parameters under different module combinations, so as to verify the improvement effect of these improvements on the model performance, including:
[0119] 1. Baseline model: The original RealNVP model.
[0120] 2. Improved model 1: Replace the coupling layer of RealNVP with the improved KAN network of the present invention (including an input layer, three hidden layers, and an output layer. The number of input neurons in the input layer is the same as the number of output neurons in the output layer. The number of neurons in the first hidden layer is the same as the number of neurons in the third hidden layer, and the number of neurons in the second hidden layer is the largest).
[0121] 3. Improved model 2: Add an LSTM layer in front of the RealNVP model.
[0122] For each model configuration, we conducted tests on the fault detection rate and statistics on the number of parameters. The ablation experiment results are shown in Tables 5 and 6:
[0123] Table 5 Comparison of Fault Detection Rates (FDR (%))
[0124]
[0125]
[0126] Table 5 Comparison of False Alarm Rates (FAR) of Normal Samples and the Number of Model Parameters
[0127]
[0128] As can be seen from Table 5 and Table 6, by replacing the coupling layer of RealNVP with the improved KAN network of the present invention, the model can more effectively capture the key features in the data, thereby improving the fault detection rate. At the same time, the stepped structure in the improved KAN network makes feature extraction more effective and the structure more compact, resulting in a decrease in the number of parameters.
[0129] Adding an LSTM layer enables the TNF model to process time-series data and extract the time-series features therein, which is very helpful for improving the fault detection rate. Although the LSTM layer may increase the complexity of the model, this is acceptable compared to the performance improvement it brings.
[0130] The TNF model of the present invention combining the KAN coupling layer and the LSTM layer reaches the maximum value in terms of the fault detection rate, and at the same time, the number of parameters also reaches the minimum value. This verifies that our improvement strategy is effective.
[0131] The results of this ablation experiment show that by improving the coupling layer of the RealNVP model, using the KAN network, and adding an LSTM layer to extract time-series features, the fault detection rate can be significantly improved and the number of parameters can be reduced. These improvements not only enhance the performance of the model but also make the model more efficient and practical.
[0132] 5. Visualization
[0133] To more intuitively see the effect of each fault, the results are visualized by plotting the fluctuations of the loss function values. The magnitude of this loss function value reflects the difference between the sample data and the standard normal distribution, as Figure 8 shown, where the red line is the set threshold. Each fault consists of 960 pieces of data, with the first 160 pieces of data being normal data and the last 800 pieces of data being fault data.
[0134] As can be seen from the figure, almost all of the first 160 pieces of data do not exceed the threshold red line, and most of the data after 160 pieces exceed the threshold red line.
[0135] Finally, it should also be noted that the above-listed are only several specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments and there can be many variations. All variations that can be directly derived or associated by those of ordinary skill in the art from the disclosed content of the present invention shall be considered within the protection scope of the present invention.
[0136] References:
[0137] [1] Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljacic, M., Hou, T. Y., & Tegmark, M. (2024). KAN: Kolmogorov - Arnold Networks. ArXiv, abs / 2404.19756.
[0138] [2] XU J, CHEN Z, LI J, et al. FourierKAN - GCF: Fourier Kolmogorov - Arnold Network--An Effective and Efficient Feature Transformation for Graph Collaborative Filtering[J]. ArXiv, abs / 2406.01034, 2024.
Claims
1. A zero-fault sample chemical process fault detection method based on normalized flow, characterized by: The process of chemical process detection includes collecting material parameters in the chemical production process, preprocessing the collected real-time data, and then inputting it into the offline trained TNF network. The input sample data first passes through the long short-term memory network to obtain the time series characteristics, and then passes through the improved normalized flow model to obtain the potential data distribution. The potential data distribution is compared with the standard normal distribution. If the difference with the standard normal distribution is less than a given threshold, it is judged as a normal sample, otherwise it is a faulty sample.
2. The method for detecting chemical process faults with zero-fault samples based on normalized flow according to claim 1 is characterized by: The improved normalized flow model is based on the RealNVP model, and the coupling layer of the RealNVP model is replaced by an improved KAN network to replace the original MLP neural network to model the scaling function and the translation function.
3. The method for detecting chemical process faults with zero-fault samples based on normalized flow according to claim 2 is characterized by: The improved KAN network includes an input layer, three hidden layers and an output layer. The number of layers increases in steps and then decreases to the same as the input dimension. The number of input neurons in the input layer is the same as the number of output neurons in the output layer. The number of neurons in the first hidden layer is the same as the number of neurons in the third hidden layer, and the second hidden layer has the largest number of neurons.
4. The method for detecting chemical process faults with zero-fault samples based on normalized flow according to claim 3 is characterized by: The process of offline training of the TNF network includes: The Tennessee-Eastman process is used to generate the TEP dataset. Only all normal data are selected from the training set of the TEP dataset as the training set for the offline training of the TNF network. The test set of the TEP dataset is used as the test set for the offline training of the TNF network. Then, all data of the training set and test set of the offline training of the TNF network are preprocessed and manually labeled. The samples in the training set are input into the TNF network in sequence. The optimization algorithm adopts an adaptive momentum estimation algorithm and adds an early stopping mechanism to prevent overfitting. During the training process, the loss function value is calculated and the model parameters are optimized by back propagation. The training ends after the preset number of epochs is reached. The test set is input into the trained TNF network, and the classification result is obtained according to the given threshold. The difference between the predicted result and the true label is compared, and the false positive rate and detection rate are used as evaluation indicators. The model with the lowest false positive rate and the highest detection rate is selected from the test results for online detection.
5. The method for detecting chemical process faults with zero-fault samples based on normalized flow according to claim 4 is characterized by: The loss function is: in, is the Jacobian determinant.
6. The method for detecting chemical process faults with zero-fault samples based on normalized flow according to claim 5, characterized in that: The given threshold is: the loss function value of the optimal model obtained during the training process.
7. The method for detecting chemical process faults with zero-fault samples based on normalized flow according to claim 6, characterized in that: The preprocessing includes: normalization and moving sliding window interception; The normalized calculation is: Where: Data is the fault sample data without preprocessing, μ is the mean of the normal sample data, s is the variance of the normal sample data, Data* is the normalized data, and the value of Data* is in the range of (-1, 1); Moving sliding window capture is to use a moving window with fixed width and step size to perform sliding capture along the time series direction.
Citation Information
Patent Citations
Industrial process fault diagnosis method based on dynamic time warping and graph convolutional network
CN113110398A