Anomaly Monitoring Method for Sewage Treatment Industrial Process Based on Deep Gaussian Mixture Model

By adopting anomaly monitoring method based on deep Gaussian hybrid model in the sewage treatment industry, using autoencoder and core principal components to analyze and extract features, and optimizing model parameters through deep learning estimation network, the problem of frequent abnormal conditions during sewage treatment is solved, and efficient and accurate abnormal monitoring is achieved.

CN119720055BActive Publication Date: 2025-07-01SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228665.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-01
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

During the sewage treatment, due to factors such as fluctuations in the inlet water quality, equipment failures and operating errors, abnormal working conditions occur frequently, which reduces treatment efficiency, increases operating costs, and may cause equipment damage and environmental pollution.

Method used

Anomaly monitoring method based on deep Gaussian hybrid model is adopted, features are extracted through autoencoder and core principal component analysis, and a deep Gaussian hybrid model is constructed, combined with deep learning to estimate network optimization model parameters, to realize dynamic updates of feature extraction and model determination.

Benefits of technology

It improves the accuracy and robustness of abnormal monitoring of wastewater treatment industry processes, can effectively capture the complex structure and potential relationships of data, improves the sensitivity and accuracy of monitoring, and ensures the safe and stable operation of wastewater treatment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119720055B_ABST
    Figure CN119720055B_ABST
Patent Text Reader

Abstract

The present invention discloses an abnormal monitoring method for sewage treatment industrial processes based on a deep Gaussian mixture model, belonging to the technical field of abnormal monitoring of industrial processes, and comprising the following steps: Step 1, collect industrial process variable data generated during the urban sewage treatment process to construct an original data set; Step 2, design a feature extraction strategy based on an autoencoder and kernel principal component analysis to obtain a low-dimensional feature representation; Step 3, construct a deep Gaussian mixture model for abnormal monitoring of the sewage treatment industrial process; Step 4, design an abnormal monitoring strategy based on a threshold, obtain a threshold evaluation index based on the constructed deep Gaussian mixture model, and further determine whether an abnormal phenomenon occurs at the current moment. The present invention realizes the synchronous optimization and adjustment of both feature extraction and model determination, realizes the effective monitoring of abnormalities occurring in the actual sewage treatment process, and can ensure the stable operation of the sewage treatment industrial process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial process anomaly monitoring, and particularly relates to a sewage treatment industrial process anomaly monitoring method based on a deep Gaussian mixture model. Background Art

[0002] The sewage treatment process has characteristics of high nonlinearity, non-Gaussianity, and high dimensionality. Its operating state is easily affected by various factors such as influent water quality fluctuations, equipment failures, and operation errors, resulting in frequent abnormal conditions. These abnormalities will not only reduce the sewage treatment efficiency and increase the operating costs, but may also cause varying degrees of damage to equipment, and even lead to catastrophic casualties and environmental pollution, causing adverse social impacts. Therefore, developing efficient and accurate sewage treatment process anomaly monitoring technologies has practical application significance for ensuring the stable operation of the sewage treatment system, improving the effluent water quality, and reducing the operating costs.

[0003] With the rapid development of machine learning technologies, data-driven anomaly monitoring methods have been widely applied in the sewage treatment field. These methods can effectively overcome the limitations of traditional methods by learning the normal operation patterns in historical data and constructing anomaly monitoring models. Among them, the Gaussian mixture model (GMM), as a probabilistic generative model, can model data with complex distributions and shows good performance in the field of anomaly monitoring. However, traditional GMM models usually adopt shallow structures and are difficult to capture the deep features and complex relationships in data, which limits their monitoring accuracy. In addition, sewage treatment process data usually has characteristics such as high dimensionality, strong coupling, and nonlinearity. The characteristic information of sewage treatment process data will change with operating conditions, anomalies, etc. Using fixed feature modeling has defects and hinders the improvement of model performance. Therefore, how to effectively extract the key features in data and construct an efficient anomaly monitoring model is an urgent problem to be solved. Therefore, it is necessary to design a precise and effective sewage treatment industrial process anomaly monitoring method to realize the dynamic update of feature extraction and model determination and ensure the effective operation of the actual sewage treatment industrial process. Summary of the Invention

[0004] To solve the above problems, the present invention proposes an abnormal monitoring method for sewage treatment industrial processes based on a deep Gaussian mixture model to achieve abnormal monitoring of sewage treatment industrial processes. The abnormal monitoring method first obtains the reconstruction error features and compression features of sample data through an autoencoder neural network layer, and at the same time performs feature extraction processing through kernel principal component analysis. Then, a deep Gaussian mixture model including a Gaussian mixture model and a deep learning estimation network is constructed. The features extracted by the autoencoder and kernel principal component analysis respectively are used as the feature providers, and then the parameters of the Gaussian mixture model are determined. During the batch update of the parameters of the autoencoder neural network layer and the deep learning estimation network layer, the feature part provided by the kernel principal component analysis will also be provided in batches. Due to the design of the loss function, after obtaining the parameters of the Gaussian mixture model, gradient descent will be used for backpropagation to change the parameters of the network layer in real time, so as to realize the simultaneous optimization and adjustment of feature extraction and model determination, and complete the abnormal monitoring of the sewage treatment industrial process. The present invention finally solves the problem of data distribution change in the actual sewage treatment industrial process and can lay a foundation for the safe and stable operation of the sewage treatment industrial process.

[0005] The technical solution of the present invention is as follows:

[0006] An abnormal monitoring method for sewage treatment industrial processes based on a deep Gaussian mixture model, comprising the following steps:

[0007] Step 1: Collect industrial process variable data generated during the urban sewage treatment process and construct an original data set;

[0008] Step 2: Design a feature extraction strategy based on an autoencoder and kernel principal component analysis to obtain a low-dimensional feature representation;

[0009] Step 3: Construct a deep Gaussian mixture model for abnormal monitoring of sewage treatment industrial processes;

[0010] Step 4: Design a threshold-based abnormal monitoring strategy, obtain a threshold evaluation index based on the constructed deep Gaussian mixture model, and then judge whether an abnormal phenomenon occurs at the current moment.

[0011] Further, in the above Step 1, the industrial process variable data includes the concentration of active heterotrophic bacteria, the concentration of active autotrophic bacteria, the concentration of dissolved nitrogen, the concentration of nitrate nitrogen, the concentration of dissolved oxygen, alkalinity, and the ratio of volatile suspended solids to total suspended solids in some treatment ponds and the effluent pond; the constructed original data set , , where is the th sample in the original data set; is the total number of samples; is the dimension of the process variable; is the transpose symbol.

[0012] Furthermore, the specific process of step 2 is as follows:

[0013] Step 2.1: Construct an autoencoder neural network including an encoder and a decoder, design a feature extraction strategy based on the autoencoder neural network, and obtain compressed features and reconstruction error features;

[0014] Step 2.2: Design a feature extraction strategy based on kernel principal component analysis to obtain kernel principal component analysis features;

[0015] Step 2.3: Combine the compressed features, reconstruction error features, and kernel principal component analysis features to form a low-dimensional feature representation of the sample data.

[0016] Furthermore, the specific process of step 2.1 is as follows:

[0017] Step 2.1.1: The encoder part compresses the data layer by layer to obtain compressed features:

[0018] (1);

[0019] (2);

[0020] where is the compressed feature; is the encoding function; are the neural network layer parameters of the encoder part; is the activation function; and are the weight and bias of the encoding function respectively;

[0021] Step 2.1.2: The compressed features enter the decoder part for data expansion to obtain a reconstruction data set with the same shape as the original data set , , is the th sample in the reconstruction data set; the whole process is the reconstruction process of the sample data:

[0022] (3);

[0023] (4);

[0024] where is the decoding function; are the neural network layer parameters of the decoder part; and are the weight and bias of the decoding function respectively;

[0025] Step 2.1.3. Calculate the reconstruction error between the original data and the reconstructed data. The formula is as follows:

[0026] (5);

[0027] where, is and the reconstruction error between them; is the th sample in the original dataset; is the th sample in the reconstructed dataset;

[0028] Step 2.1.4. Take the reconstruction error as the reconstruction error feature. The reconstruction error feature consists of two parts, namely the relative Euclidean distance and the cosine similarity, which are specifically as follows:

[0029] (6);

[0030] (7);

[0031] (8);

[0032] where, is the reconstruction error feature; is the relative Euclidean distance; is the cosine similarity; is the th sample's reconstruction error feature; is the relative Euclidean distance calculation function; is the cosine similarity calculation function; is the two-norm calculation symbol.

[0033] Furthermore, the specific process of Step 2.2 is as follows:

[0034] Step 2.2.1. Calculate the kernel matrix between pairwise input sample data through the Gaussian kernel function;

[0035] (9);

[0036] where, is the kernel matrix in the th row and th column's element value, ; is the Gaussian kernel function; and are respectively the th sample, the th sample in the original dataset, the One sample corresponds to the rd row and the th sample corresponds to the th column; is the exponential function with base e; is the parameter of the Gaussian kernel function;

[0037] Step 2.2.2: Centralize the kernel matrix to obtain the centralized matrix:

[0038] (10);

[0039] where is the centralized matrix; is the all-ones vector, and its calculation formula is:

[0040] (11);

[0041] where is 's all-ones column vector;

[0042] Step 2.2.3: Solve the eigenvalues of the centralized matrix to obtain the eigenvalues and the corresponding eigenvectors, and perform regularization on the eigenvectors:

[0043] (12);

[0044] (13);

[0045] where is the th eigenvalue; is the eigenvector corresponding to the th eigenvalue; is the eigenvector corresponding to the th eigenvalue after regularization;

[0046] Step 2.2.4: Select the eigenvectors corresponding to the top largest eigenvalues to form the dimensionality reduction matrix, and obtain the dimensionality-reduced features according to the dimensionality reduction matrix:

[0047] (14);

[0048] where are the dimensionality-reduced features; is the dimensionality reduction matrix;

[0049] Step 2.2.5: Use the dimensionality-reduced features as the kernel principal component analysis features:

[0050] (15);

[0051] Among them, is the kernel principal component analysis feature; is the th feature after dimensionality reduction of the sample; is the kernel principal component analysis feature of the th sample.

[0052] Furthermore, the formula in step 2.3 is:

[0053] (16);

[0054] Among them, is the low-dimensional feature representation; is the th compressed feature of the sample; is the low-dimensional feature representation of the th sample.

[0055] Furthermore, in step 3, the deep Gaussian mixture model includes a Gaussian mixture model and a deep learning estimation network. The Gaussian mixture model is used to implement anomaly monitoring, and the deep learning estimation network is used to solve the parameters in the Gaussian mixture model. The specific process is as follows:

[0056] Step 3.1, the Gaussian mixture model is:

[0057] (17);

[0058] Among them, is the probability density function of the input data under the given model parameters , , is the mixing probability of the th sub-Gaussian distribution, is the total number of sub-Gaussian distributions; is the probability density function of the th sub-Gaussian distribution. The model parameters of the th sub-Gaussian distribution , , are the mean and covariance of the th sub-Gaussian distribution respectively, and are specifically expressed as:

[0059] (18);

[0060] Step 3.2, the parameters in the Gaussian mixture model include the mixing probability, mean, and covariance matrix of each sub-Gaussian distribution. A parameter estimation strategy based on a neural network is designed to solve the parameters, which is specifically expressed as:

[0061] (19);

[0062] (20);

[0063] Among them, is the low-dimensional feature representation which is the output of a multi-layer neural network parameterized by ; are the parameters of the deep learning estimation network; is the multi-layer neural network; is the membership degree prediction; is the softmax layer;

[0064] Step 3.3, the calculation formulas for the mixing probability, mean, and covariance are:

[0065] (21);

[0066] (22);

[0067] (23);

[0068] Among them, is the membership degree prediction of the -th sample belonging to the -th sub-Gaussian distribution; is the low-dimensional feature representation of the -th sample;

[0069] Step 3.4, calculate the likelihood function value:

[0070] (24);

[0071] Among them, is the likelihood function value; is the inverse matrix of the covariance;

[0072] Step 3.5, combine the reconstruction error and the likelihood function of the Gaussian mixture model to design a joint loss function to achieve the optimization and adjustment of the feature extraction and anomaly monitoring model, and construct it as follows:

[0073] (25);

[0074] Among them, is the joint loss function; , are respectively two hyperparameters that control the , weights during the training process; is a penalty term to prevent the covariance matrix from being irreversible, and the formula is:

[0075] (26);

[0076] where is the dimension of the low-dimensional feature representation; is the -th diagonal element in the covariance matrix of the -th sub-Gaussian distribution;

[0077] Step 3.6: Take the joint loss function as the objective function, and backpropagate through the gradient descent algorithm while minimizing the reconstruction error and the likelihood function value in the Gaussian mixture model. The formula is:

[0078] (27);

[0079] (28);

[0080] where is the gradient of the joint loss function ; are the parameters used in the whole model, including , , ; is the learning rate; are the parameters updated by the gradient descent algorithm; after the parameters are updated, the autoencoder neural network layer and the deep learning estimation network will extract features again, and at the same time, the parameters of the Gaussian mixture model will also change, realizing the dynamic optimization adjustment of feature extraction and model determination.

[0081] Furthermore, the specific process of Step 4 is as follows:

[0082] Step 4.1: Divide the original dataset into a training set, a test set, and a validation set;

[0083] Step 4.2: Use the training set to train the deep Gaussian mixture model, and after training, obtain the probability density function of the Gaussian mixture model to which the training data belongs;

[0084] Step 4.3: Calculate the likelihood function values of all validation samples on the trained deep Gaussian mixture model using the validation set, sort the likelihood function values of all validation samples, and since the proportion of abnormal samples in the validation set is known in advance, set the likelihood function value at the same position as the abnormal sample proportion value in the sorting as the threshold;

[0085] Step 4.4: Calculate the likelihood function value using the test set on the trained deep Gaussian mixture model, and then compare it with the threshold to determine whether there is an abnormal phenomenon;

[0086] Step 4.5: Calculate the likelihood function value of the sample data in the current sewage treatment process. If the likelihood function value is greater than the set threshold, it is determined that an abnormality has occurred; otherwise, there is no abnormal phenomenon.

[0087] The beneficial technical effects brought by the present invention are as follows.

[0088] 1. According to the problems in the actual sewage treatment industrial process, the present invention proposes an abnormal monitoring method for the sewage treatment industrial process based on a deep Gaussian mixture model, realizing accurate monitoring of whether an abnormality occurs in the sewage treatment industrial process: due to the change in the distribution of sewage data, the accuracy of the model generated by selecting fixed features is limited. Therefore, an end-to-end model structure is adopted to solve the drawback that the feature extraction and model determination in the existing abnormal monitoring methods are limited in their connection with each other. The gradient descent backpropagation algorithm is used to realize the synchronous update of features and the model, improving the efficiency of the deep Gaussian mixture model in abnormal monitoring, and the feasibility of the proposed method is verified through tests on multiple actual sewage treatment process data sets.

[0089] 2. The present invention introduces an additional kernel principal component analysis feature extraction part, enhancing the feature extraction ability of the model, facilitating the capture of the complex structure and potential relationship of the data, and improving the sensitivity of features to abnormal monitoring; at the same time, in the estimation strategy of the neural network, the principal components extracted by kernel principal component analysis are combined with the reconstruction error of the autoencoder and the compressed low-dimensional features, which is beneficial to improving the accuracy and robustness of abnormal monitoring. The generated Gaussian mixture model can also more accurately predict the membership degree and optimize parameters, thereby improving the overall performance of the abnormal monitoring system, realizing effective monitoring of abnormalities occurring in the actual sewage treatment process, and ensuring the stable operation of the sewage treatment industrial process. Description of the Drawings

[0090] Figure 1 is the schematic diagram of the abnormal monitoring method for the sewage treatment industrial process based on the deep Gaussian mixture model of the present invention.

[0091] Figure 2 is the effect diagram of the deep Gaussian mixture model for abnormal monitoring under fault condition 4 in the experiment of the present invention.

[0092] Figure 3 is the confusion matrix diagram of the effect of the deep Gaussian mixture model for abnormal monitoring under fault condition 4 in the experiment of the present invention. Detailed Embodiment

[0093] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0094] The abnormal monitoring method of the present invention first obtains the reconstructed error sample features and compressed features of industrial process data through the autoencoder neural network layer, and at the same time extracts features through kernel principal component analysis. Then, the features provided by both are jointly used as the input of the deep learning estimation network under the deep Gaussian mixture model framework to obtain the parameters of the Gaussian mixture model, and further calculate the likelihood function value of the sample data in the model. By comparing with the threshold, the abnormal monitoring of the actual industrial process is realized. The key point of the present invention lies in the design of the joint loss function. By updating the parameters of the network layer through gradient descent backpropagation, and then affecting the determination of the model again, the synchronous optimization adjustment of both feature extraction and model determination is realized. Finally, the problem that it is difficult to accurately monitor abnormalities due to the changing data distribution caused by non-stationarity in the actual sewage treatment industrial process is solved, which can lay a foundation for the safe and stable operation of the actual sewage treatment industrial process.

[0095] As Figure 1 shown, an abnormal monitoring method for the sewage treatment industrial process based on a deep Gaussian mixture model realizes the optimization adjustment of both feature extraction and the abnormal monitoring model through the design of the joint loss function, and specifically includes the following steps:

[0096] Step 1: Collect the industrial process variable data generated during the urban sewage treatment process to construct the original data set , , where is the th sample in the original data set; is the total number of samples; is the dimension of the process variable; is the transpose symbol;

[0097] The industrial process variable data includes the concentration of active heterotrophic bacteria, the concentration of active autotrophic bacteria, the concentration of dissolved nitrogen, the concentration of nitrate nitrogen, the concentration of dissolved oxygen, alkalinity, and the proportion of volatile suspended solids in total suspended solids in some treatment ponds and the effluent pond.

[0098] Step 2: Design a feature extraction strategy based on the autoencoder and kernel principal component analysis to obtain a low-dimensional feature representation; the specific process is as follows:

[0099] Step 2.1: Construct an autoencoder neural network including an encoder and a decoder, design a feature extraction strategy based on the autoencoder neural network to obtain compressed features and reconstructed error features; the specific process is as follows:

[0100] Step 2.1.1: First, the encoder part compresses the data layer by layer to obtain compressed features:

[0101] (1);

[0102] (2);

[0103] Among them, is the compression feature; is the encoding function; are the parameters of the neural network layer in the encoder part; is the activation function; and are the weight and bias of the encoding function respectively;

[0104] Step 2.1.2. Then, the compression feature enters the decoder part for data expansion to obtain a reconstructed dataset with the same shape as the original dataset , , is the th sample in the reconstructed dataset; The whole process can be regarded as the reconstruction process of sample data:

[0105] (3);

[0106] (4);

[0107] Among them, is the decoding function; are the parameters of the neural network layer in the decoder part; and are the weight and bias of the decoding function respectively;

[0108] Step 2.1.3. Calculate the reconstruction error between the original data and the reconstructed data. The formula is:

[0109] (5);

[0110] Among them, is the and reconstruction error between; is the th sample in the original dataset; is the th sample in the reconstructed dataset;

[0111] Step 2.1.4. Take the reconstruction error as the reconstruction error feature; The reconstruction error feature consists of two parts, namely the relative Euclidean distance and the cosine similarity, as follows:

[0112] (6);

[0113] (7);

[0114] (8);

[0115] Among them, is the reconstruction error feature; is the relative Euclidean distance; is the cosine similarity; is the reconstruction error feature of the th sample; is the relative Euclidean distance calculation function; is the two-norm calculation symbol;

[0116] Step 2.2: Design a feature extraction strategy based on kernel principal component analysis to obtain kernel principal component analysis features; the specific process is as follows:

[0117] Step 2.2.1: Calculate the kernel matrix between pairwise input sample data through the Gaussian kernel function;

[0118] (9);

[0119] Among them, is the kernel matrix in the th row and th column of the element value, ; is the Gaussian kernel function; , are the th sample, th sample, th sample in the original dataset corresponding to the th row, and the th sample corresponding to the th column; is the exponential function with base e; is the parameter of the Gaussian kernel function, set to 1.0;

[0120] Step 2.2.2: Then centralize the kernel matrix to obtain the centralized matrix:

[0121] (10);

[0122] Among them, is the centralized matrix; is the all-ones vector, and the calculation formula is:

[0123] (11);

[0124] Among them, is the all-ones column vector;

[0125] Step 2.2.3. Subsequently, solve the eigenvalues of the centralized matrix to obtain the eigenvalues and the corresponding eigenvectors, and perform regularization processing on the eigenvectors:

[0126] (12);

[0127] (13);

[0128] where, is the th eigenvalue; is the eigenvector corresponding to the th eigenvalue; is the eigenvector corresponding to the th eigenvalue after regularization;

[0129] Step 2.2.4. Select the eigenvectors corresponding to the top eigenvalues to form a dimensionality reduction matrix, and obtain the dimensionality-reduced features according to the dimensionality reduction matrix:

[0130] (14);

[0131] where, is the dimensionality-reduced feature; is the dimensionality reduction matrix;

[0132] Step 2.2.5. Use the dimensionality-reduced features as the kernel principal component analysis features:

[0133] (15);

[0134] where, is the kernel principal component analysis feature; is the dimensionality-reduced feature of the th sample; is the kernel principal component analysis feature of the th sample;

[0135] After calculating the kernel matrix of the input data, before the eigenvalue solving step, relevant sample data needs to be indexed according to the batch size of the autoencoder neural network processing to ensure data unity, and the batch size is set to 32;

[0136] Step 2.3. Combine the compressed features, reconstruction error features, and kernel principal component analysis features to form the low-dimensional feature representation of sample data:

[0137] (16);

[0138] where, is the low-dimensional feature representation; is the compressed feature of the th sample; is the low-dimensional feature representation of the th sample;

[0139] Step 3: Construct a deep Gaussian mixture model for anomaly monitoring of the sewage treatment industrial process. The deep Gaussian mixture model includes a Gaussian mixture model and a deep learning estimation network. The Gaussian mixture model is used to achieve anomaly monitoring, and the deep learning estimation network is used to solve the parameters in the Gaussian mixture model. The specific process is as follows:

[0140] Step 3.1: After obtaining the low-dimensional feature representation of the sample data set, the Gaussian mixture model is used to achieve the purpose of anomaly monitoring. The Gaussian mixture model is a mixed probability distribution model, which can be specifically expressed as:

[0141] (17);

[0142] where is the probability density function of the input data under the given model parameters , , is the mixing probability of the th sub-Gaussian distribution, is the total number of sub-Gaussian distributions; , ; is the probability density function of the th sub-Gaussian distribution. The model parameters of the th sub-Gaussian distribution, , are the mean and covariance of the th sub-Gaussian distribution respectively, and can be specifically expressed as:

[0143] (18);

[0144] Step 3.2: The parameters in the Gaussian mixture model include the mixing probability, mean, and covariance matrix of each sub-Gaussian distribution. The maximum expectation algorithm is usually used to solve the parameters. In the deep Gaussian mixture model, in order to obtain the above parameters, a parameter estimation strategy based on neural network is designed, and a deep learning estimation network is constructed. Its main role is to replace the E-step in the maximum expectation algorithm and can combine the loss function with feature extraction. After obtaining the low-dimensional feature representation , the corresponding mixing membership degree can be obtained, which can be used to directly estimate the parameters of the Gaussian mixture model, and can be specifically expressed as:

[0145] (19);

[0146] (20);

[0147] Among them, is a low-dimensional feature representation which is the output of a multi-layer neural network parameterized by with the dimension set to 4; is the parameter of the deep learning estimation network; is the multi-layer neural network; is the membership prediction, which is a dimensional probability value, and the output of the network is converted into a probability distribution via the softmax layer, which is used to simulate the probability that the input sample data belongs to sub-Gaussian distributions and the parameters of each sub-Gaussian distribution; is the softmax layer;

[0148] Step 3.3. Given the data sample set and knowing their low-dimensional feature representations and membership predictions, the mixing probability, mean, and covariance in the Gaussian mixture model can be further obtained by the following formulas:

[0149] (21);

[0150] (22);

[0151] (23);

[0152] Among them, is the membership prediction that the th sample belongs to the th sub-Gaussian distribution; is the low-dimensional feature representation of the th sample; the number of sub-Gaussian distributions in the set Gaussian mixture model is 4;

[0153] Step 3.4. When all these parameters are known, the likelihood function value can be known:

[0154] (24);

[0155] Among them, is the likelihood function value; is the inverse matrix of the covariance;

[0156] Step 3.5: Combine the reconstruction error and the likelihood function of the Gaussian mixture model to design a joint loss function for optimizing and adjusting the feature extraction and anomaly monitoring model, which is constructed as follows:

[0157] (25);

[0158] where, is the joint loss function; is the low-dimensional feature representation of the -th sample; and are two hyperparameters that control the and weights during the training process, and are set to 0.1 and 0.005 respectively; is a penalty term to prevent the covariance matrix from being irreversible, and the formula is:

[0159] (26);

[0160] where, is the dimension of the low-dimensional feature representation; is the -th diagonal element in the covariance matrix of the -th sub-Gaussian distribution;

[0161] Step 3.6: Use the joint loss function as the objective function, and backpropagate through the gradient descent algorithm to simultaneously minimize the reconstruction error and the likelihood function value in the Gaussian mixture model, so that the feature extraction task and the model determination task can be carried out collaboratively in the same framework to achieve the purpose of synchronous optimization and adjustment;

[0162] (27);

[0163] (28);

[0164] where, is the gradient of the joint loss function , and the gradient is a vector composed of the partial derivatives of the joint loss function with respect to each parameter; are the parameters used in the entire model, including , , ; is the learning rate, set to 0.01; are the parameters updated by the gradient descent algorithm; after the parameters are updated, the autoencoder neural network layer and the deep learning estimation network will extract features again, and at the same time, the parameters of the Gaussian mixture model will also change, which can achieve the purpose of dynamic optimization and adjustment of feature extraction and model determination;

[0165] In addition, the kernel principal component analysis part takes into account the variance distribution information of the original data, which promotes the optimization of the overall model. It should be noted that the features obtained by extracting the features of the training data through kernel principal component analysis During subsequent training, the same index window size is set according to the size of the batch of the autoencoder neural network layer. At the same time, the step size is also set to the index window size, and the batch size is set to 32;

[0166] Specifically, the neural network layer parameters are set as follows: The autoencoder neural network layer runs as FC(30,15,tanh)-FC(15,10,tanh)-FC(10,5,tanh)-FC(5,1,none)-FC(1,5,tanh)-FC(5,10,tanh)-FC(10,15,tanh)-FC(15,30,tanh), and the deep learning estimation network runs as FC(5,10,tanh)-Drop(0.5)-FC(10,4,softmax); Specifically, FC(a,b,f) represents a fully connected layer, parameterized as the number of input neurons a, the number of output neurons b, and the activation function f; Drop(p) represents a dropout layer, and p represents the retention probability during training. Among them, none represents no activation function.

[0167] Step 4: Design a threshold-based anomaly monitoring strategy. Based on the constructed deep Gaussian mixture model, obtain the threshold evaluation index, and then judge whether an abnormal phenomenon occurs at the current moment. The specific process is as follows:

[0168] Step 4.1: Divide the original data set into a training set, a test set, and a validation set.

[0169] Step 4.2: Use the training set to train the deep Gaussian mixture model. After training, the probability density function of the Gaussian mixture model to which the training data belongs can be obtained;

[0170] Step 4.3: Use the validation set to calculate the likelihood function values of all validation samples on the trained deep Gaussian mixture model, sort the likelihood function values of all validation samples, and the proportion of abnormal samples in the validation set is known in advance. Set the likelihood function value at the same position as the abnormal sample proportion value in the sorting as the threshold; for example, if the proportion of abnormal samples is 70%, then select the likelihood function value at 70% in the sorting as the threshold;

[0171] Step 4.4: Use the test set to calculate the likelihood function value on the trained deep Gaussian mixture model, and then compare it with the threshold to judge whether an abnormal phenomenon occurs;

[0172] The reason for choosing the normal likelihood function value as the threshold is that during the monitoring process, the distance between the low-dimensional representation of the sample and all sub-Gaussian distributions in the Gaussian mixture model can be used to measure the probability of the sample being abnormal or normal. The smaller the distance, the higher the probability that the sample belongs to the normal category, and at this time, the likelihood function value of the sample is smaller; the larger the distance, the higher the probability that the sample belongs to the abnormal category, and the larger the likelihood function value of the sample. Therefore, the normal likelihood function value is used as an indicator to judge whether an abnormality has occurred;

[0173] Step 4.5: Calculate the likelihood function value of the sample data in the current sewage treatment process. If the likelihood function value is greater than the set threshold, it can be determined that an abnormality has occurred; otherwise, no abnormal phenomenon occurs.

[0174] To prove the feasibility and superiority of the present invention, the process variable data generated during the urban sewage treatment process in the benchmark simulation model No. 1 (BSM1) was selected to analyze the operating characteristics of the urban sewage treatment process. Thirty specific relevant process variables were selected: the concentration of active heterotrophic bacteria in the 1st / 2nd / 3rd / 4th / 5th / effluent tank, the concentration of active autotrophic bacteria in the 1st / 2nd / 3rd / 4th / 5th / effluent tank, the dissolved nitrogen concentration in the 1st / 2nd / 3rd / 4th / effluent tank, the nitrate nitrogen concentration in the 1st / 2nd / 3rd / 4th effluent tank, the dissolved oxygen concentration in the 1st / 2nd / 3rd / 4th / 5th / effluent tank, the alkalinity in the 5th effluent tank, and the ratio of volatile suspended solids to total suspended solids in the 3rd / 5th effluent tank. At the same time, a fault condition based on BSM1 was designed to simulate some common faults, and 4035 sample data with a period of 14 days were recorded under each fault condition.

[0175] Figure 2 Fig. shows the effect of the deep Gaussian mixture model on anomaly monitoring under fault condition 4. Fault condition 4 represents a fault situation where the decay rate of heterotrophic bacteria in Unit2 changes from 0.62 to 0.2, and the set fault occurrence time is the 7th day. From Figure 2 it can be seen that the set fault occurrence time is the 7th day, that is, the sample point around the 2017th one, Figure 2 the overall trend of the sample likelihood function values monitored for the input sample data in Fig. conforms to the set fault, that is, it can be obtained that under this fault condition, the model can monitor the occurrence of anomalies.

[0176] Figure 3 Fig. shows the confusion matrix diagram of the effect of the deep Gaussian mixture model on anomaly monitoring under fault condition 4. From Figure 3 it can be seen that under this fault condition, the false discovery rate (FDR) of the model's anomaly monitoring is 91.68%, and the false alarm rate (FAR) is 8.24%. From the results, it can be seen that the deep Gaussian mixture model of the present invention can monitor the occurrence of abnormal situations and has a low false alarm rate.

[0177] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions or substitutions made by those skilled in the art to the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for monitoring abnormalities in industrial processes of sewage treatment based on a deep Gaussian mixture model, characterized in that: The steps include: Step 1: Collect industrial process variable data generated in the process of urban sewage treatment and construct the original data set; Step 2: Design a feature extraction strategy based on autoencoder and kernel principal component analysis to obtain low-dimensional feature representation; The specific process is: Step 2.1, construct an autoencoder neural network including an encoder and a decoder, design a feature extraction strategy based on the autoencoder neural network, and obtain compression features and reconstruction error features; The specific process is: Step 2.1.1: The encoder compresses the data layer by layer to obtain the compression features: z c =h(X;θ e ) (1); h=α(wX+b) (2); Among them, z c is the compression feature; h is the encoding function; θ e are the parameters of the neural network layer of the encoder; α is the activation function; w and b are the weight and bias of the encoding function respectively; Step 2.1.2: The compressed features enter the decoder part for data expansion to obtain a reconstructed data set X′=[x1′,x2′,…,x′] with the same shape as the original data set. N ] T , x′ N To reconstruct the Nth sample in the data set; the whole process is the reconstruction process of the sample data: X′=g(z c ;θ d ) (3); g=α(w′z c +b′) (4); Among them, g is the decoding function; θ d are the parameters of the neural network layer of the decoder; w′ and b′ are the weight and bias of the decoding function respectively; Step 2.1.3: Calculate the reconstruction error between the original data and the reconstructed data. The formula is: F(x i ,x i ′)=‖x i -x i ′‖ 2 (5); Among them, F(x i ,x i ′) is x i With x i ′; x i is the i-th sample in the original data set; x i ′ is the i-th sample in the reconstructed data set; Step 2.1.4: Use the reconstruction error as the reconstruction error feature; the reconstruction error feature consists of two parts, namely the relative Euclidean distance and the cosine similarity, as follows: With r =[z1,z2] T =[of r1 ,With r2 ,…,With rN ] T (6); Among them, z r is the reconstruction error feature; z1 is the relative Euclidean distance; z2 is the cosine similarity; z rN is the reconstruction error feature of the Nth sample; f1(·) is the relative Euclidean distance calculation function; f2(·) is the cosine similarity calculation function; ‖·‖2 is the symbol for the two-norm calculation; Step 2.2: Design a feature extraction strategy based on kernel principal component analysis to obtain kernel principal component analysis features. The specific process is as follows: Step 2.2.1, calculate the kernel matrix between the input sample data pairs by using the Gaussian kernel function; Among them, K ij is the element value of the i-th row and j-th column in the kernel matrix K, is the Gaussian kernel function; x i 、x j are the i-th sample and j-th sample in the original data set, respectively. The i-th sample corresponds to the i-th row, and the j-th sample corresponds to the j-th column. exp(·) is an exponential function with e as the base. γ is the parameter of the Gaussian kernel function. Step 2.2.2: Centralize the kernel matrix to obtain the centralized matrix: K c =K-1 N K-K1 N +1 N K1 N (10); in, is the centralized matrix; is a vector of all 1s, and the calculation formula is: Among them, e N is an N×1 column vector of all ones; Step 2.2.3, solve the eigenvalue of the centered matrix, obtain the eigenvalue and the corresponding eigenvector, and regularize the eigenvector: K c m l =λ l m l ,l=1,2,…,N (12); Among them, λ l is the lth eigenvalue; μ l is the eigenvector corresponding to the lth eigenvalue; μ′ l is the eigenvector corresponding to the lth eigenvalue after regularization; Step 2.2.4, select the eigenvectors corresponding to the largest first d eigenvalues ​​to form a dimensionality reduction matrix, and obtain the reduced dimensionality features according to the dimensionality reduction matrix: X k =K c μ′ (14); in, is the feature after dimensionality reduction; is the dimension reduction matrix; Step 2.2.5: Use the reduced dimension features as kernel principal component analysis features: z k =X k =[X k1 ,X k2 ,…,X kN ] T =[z k1 ,z k2 ,…,z kN ] T (15); Among them, z k is the kernel principal component analysis feature; X kN is the feature of the Nth sample after dimensionality reduction; z kN is the kernel principal component analysis feature of the Nth sample; Step 2.3: Combine the compression features, reconstruction error features, and kernel principal component analysis features to form a low-dimensional feature representation of N sample data; the formula is: Among them, Z is the low-dimensional feature representation; z cN is the compression feature of the Nth sample; Z N is the low-dimensional feature representation of the Nth sample; Step 3: Construct a deep Gaussian mixture model to monitor abnormalities in the sewage treatment industrial process; The deep Gaussian mixture model includes a Gaussian mixture model and a deep learning estimation network. The Gaussian mixture model is used to realize abnormality monitoring, and the deep learning estimation network is used to solve the parameters in the Gaussian mixture model. The specific process is as follows: Step 3.1, Gaussian mixture model is: Where P(y|ψ) is the probability density function of the input data y under the given model parameter ψ, For the The probability of a mixture of sub-Gaussian distributions, is the total number of sub-Gaussian distributions; For the The probability density function of the Gaussian distribution is Model parameters of the sub-Gaussian distribution Respectively The mean and covariance of the sub-Gaussian distribution are specifically expressed as: Step 3.2: The parameters in the Gaussian mixture model include the mixing probability, mean, and covariance matrix of each sub-Gaussian distribution. A parameter estimation strategy based on a neural network is designed to solve the parameters, which is specifically expressed as: L=MLN(z;θ m ) (19); Among them, L is the low-dimensional feature representation z through θ m The output of the parameterized multi-layer neural network; θ m It is the parameter estimation of the deep learning network; MLN(·) is a multi-layer neural network; is the membership prediction; softmax(·) is the softmax layer; Step 3.3, the calculation formulas for mixing probability, mean, and covariance are: in, The i-th sample belongs to Membership prediction of sub-Gaussian distribution; i is the low-dimensional feature representation of the i-th sample; Step 3.4, calculate the likelihood function value: Where, E(·) is the likelihood function value; is the inverse matrix of the covariance; Step 3.5: Combine the reconstruction error and the likelihood function of the Gaussian mixture model to design a joint loss function to achieve optimal adjustment of the feature extraction and anomaly detection model. The structure is as follows: Among them, J is the joint loss function; λ1 and λ2 are respectively used to control E(z i ), P(∑) weights; P(∑) is a penalty term to prevent the covariance matrix ∑ from being irreversible, and the formula is: Among them, H is the dimension of the low-dimensional feature representation; For the The covariance matrix of the sub-Gaussian distribution is diagonal elements; Step 3.6: Use the joint loss function as the objective function and back-propagate through the gradient descent algorithm to minimize the reconstruction error and the likelihood function value in the Gaussian mixture model. The formula is: in, is the gradient of the joint loss function J; θ is the parameter used in the entire model, including θ e ,θ d ,θ m ; η is the learning rate; θ′ is the updated parameter of the gradient descent algorithm; after the parameters are updated, the autoencoder neural network layer and the deep learning estimation network will extract features again, and the parameters of the Gaussian mixture model will also change, realizing dynamic optimization and adjustment of feature extraction and model determination; Step 4: Design a threshold-based anomaly monitoring strategy, obtain the threshold evaluation index based on the constructed deep Gaussian mixture model, and then determine whether an abnormal phenomenon occurs at the current moment.

2. The method for monitoring abnormalities in industrial processes of sewage treatment based on a deep Gaussian mixture model according to claim 1, characterized in that: In step 1, the industrial process variable data include the concentration of active heterotrophic bacteria, active autotrophic bacteria, dissolved nitrogen, nitrate nitrogen, dissolved oxygen, alkalinity, and the proportion of volatile suspended solids to total suspended solids in some treatment tanks and effluent tanks; the constructed original data set X = [x1, x2, ..., x N ] T , Among them, x N is the Nth sample in the original data set; N is the total number of samples; D is the dimension of the process variable; T is the transposition symbol.

3. The method for monitoring abnormalities in industrial processes of sewage treatment based on a deep Gaussian mixture model according to claim 2, characterized in that: The specific process of step 4 is as follows: Step 4.1, divide the original data set into training set, test set and validation set; Step 4.2: Use the training set to train the deep Gaussian mixture model, and after training, obtain the probability density function of the Gaussian mixture model to which the training data belongs; Step 4.3: Use the validation set to calculate the likelihood function values ​​of all validation samples on the trained deep Gaussian mixture model, sort the likelihood function values ​​of all validation samples, and set the likelihood function value at the same position as the abnormal sample ratio in the sorting as the threshold if the abnormal sample ratio in the validation set is known in advance; Step 4.4: Use the test set to calculate the likelihood function value on the trained deep Gaussian mixture model, and then compare it with the threshold to determine whether any abnormal phenomenon occurs; Step 4.5: Calculate the likelihood function value of the sample data in the current sewage treatment process. If the likelihood function value is greater than the set threshold, it is determined that an abnormality has occurred; otherwise, no abnormality has occurred.

Citation Information

Patent Citations

  • Sewage treatment process fault monitoring method based on variational auto-encoder model

    CN112631255A

  • Power consumption anomaly detection method based on depth auto-encoder Gaussian mixture model

    CN113902581A