Ammonia nitrogen concentration prediction method for multi-sampling-rate sewage treatment process

Through the HMWS model, the MCLC strategy and GWSNET network were used to construct supervised and unsupervised sub-networks, which solved the problem of insufficient prediction accuracy of ammonia nitrogen concentration of multi-sampling rate data in the sewage treatment process and achieved higher prediction accuracy and data utilization efficiency.

CN120706270APending Publication Date: 2025-09-26WUXI TSINGDA BIOTECH EPE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510876660.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing soft sensor models cannot effectively deal with the modeling problem of multi-sampling rate data of process variables in the sewage treatment process, resulting in insufficient prediction accuracy of ammonia nitrogen concentration.

Method used

A hierarchical masked weak supervision (HMWS) soft sensing model is adopted, combined with the masked coarse-grained linear completion (MCLC) strategy and the generative weak supervision network (GWSNET). Through the construction of supervised and unsupervised sub-networks, the missing values ​​of auxiliary variables are filled and the data perception ability of the model is improved.

Benefits of technology

The effective use of unlabeled data in the multi-sampling rate process improves the prediction accuracy of ammonia nitrogen concentration and the applicability of the model, alleviates the problem of missing data values, and broadens the field of data perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706270A_ABST
    Figure CN120706270A_ABST
Patent Text Reader

Abstract

The invention discloses an ammonia nitrogen concentration prediction method for a multi-sampling-rate sewage treatment process, and belongs to the technical field of sewage treatment. According to the method, a mask coarseness linear complementation strategy (MCLC) is provided to input missing values of auxiliary variables, a generative weakly supervised network (GWSNET) is designed to relieve the shortage problem of marked data, and hierarchical learning of multi-sampling-rate data is achieved. And the MCLC strategy performs interpolation on the auxiliary variables based on the mask matrix and convolution coarse granularity complementation to generate weak supervision data so as to expand the field of data perception and realize subsequent weak supervision modeling. In addition, the GWSNET establishes supervised and unsupervised generative branch sub-networks with similar latent variable distribution. According to the design, the GWSNET is allowed to utilize limited labeled samples to guide unlabeled samples, latent variable distribution capable of accurately representing process characteristics is obtained, and online prediction of the ammonia nitrogen concentration of the sewage treatment process is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an ammonia nitrogen concentration prediction method for a multi-sampling rate sewage treatment process, and belongs to the technical field of sewage treatment. Background Art

[0002] Wastewater treatment is a crucial branch of environmental protection, with effluent quality often used to characterize treatment effectiveness. A key parameter among effluent quality indicators is ammonia nitrogen concentration, which refers to the total amount of nitrogen in the water in the form of ammonia (NH3) and ammonium ions (NH4+). When ammonia nitrogen levels in water exceed a certain range, it can lead to eutrophication. Furthermore, due to its high oxygen demand, aquatic life can be affected. Furthermore, under certain conditions, ammonia nitrogen can be converted into nitrite, which reacts with proteins to produce nitrosamines, a potential carcinogen and a serious threat to human health. Therefore, real-time monitoring of ammonia nitrogen concentration is necessary. However, due to harsh environments, technical limitations, chemical properties, and other factors, ammonia nitrogen concentration is difficult to measure directly using hardware sensors. Soft sensing technology addresses this problem by estimating difficult-to-measure key variables (quality variables) using easily measurable process variables (auxiliary variables).

[0003] Existing soft measurement models mainly include white-box and black-box models. Among them, black-box-based soft measurement models have the advantage of low dependence on process mechanism knowledge and have developed rapidly in recent years. Black-box models have a powerful ability to analyze data structures. Therefore, black-box-based models are called data-driven models. However, under data-driven models, each process variable needs to maintain a consistent sampling rate. However, due to factors such as sensor limitations, chemical properties, and distributed control systems, it is difficult for each process variable in the sewage treatment process to maintain a consistent sampling rate. This heterogeneity seriously undermines the integrity of the data, thereby increasing the complexity and error of data-driven soft measurement modeling.

[0004] In the field of soft sensing, chemical industrial processes with three or more sampling rates are usually called multi-sampling rate processes. In addition to the existence of multiple sampling rates between process variables in such processes, there is another problem. Due to the difficulty of measuring quality variables, they are usually sampled at a lower sampling rate than auxiliary variables, resulting in a shortage of labeled samples. This also poses a challenge to traditional data-driven soft sensing models. To address the challenge of shortage of labeled samples, some scholars have adopted semi-supervised or weakly supervised methods to improve the learning ability of soft sensing models. Shi et al. (Shi X, Kang Q, Zhou M, et al. Soft sensing of nonlinear and multimode processes based on semi-supervised weighted gaussian regression [J]. IEEE Sensors Journal, 2020, 20 (21): 12950-12960.) introduced a weighted Gaussian model in the semi-supervised model, which can approximate the nonlinearity of input and output variables and compensate for the scarcity of labeled samples. Sui et al. (Sui L, Zhang CL and Wu J. Salvage of supervision in weakly supervised object detection and segmentation [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 (8): 10394-10408.) utilized every potential useful information in the weakly supervised task and significantly improved the performance and applicability of the model. The above-mentioned semi-supervised or weakly supervised methods can effectively use existing samples to predict quality variables, but they cannot be directly applied to unlabeled samples in multi-sampling rate data of process variables. This is because traditional weakly supervised methods assume that auxiliary variables are fully sampled, but in actual multi-sampling rate processes, auxiliary variables are inevitably missing.

[0005] To sum up, the existing soft sensing methods are unable to cope with the modeling problem of multi-sampling rate data of process variables in industrial processes. Therefore, there is still room for improvement in the prediction accuracy of ammonia nitrogen concentration in sewage treatment processes. Summary of the Invention

[0006] In order to solve the problem of multiple sampling rates of process variables faced by current soft sensor models in sewage treatment and further improve the prediction accuracy of key variables, this paper proposes a hierarchical mask weekly-supervised (HMWS) soft sensor model based on the mask coarse-grained linear complementation (MCLC) strategy and the generative weekly-supervised network (GNN). In order to predict the ammonia nitrogen concentration of sewage treatment data with multiple sampling rates, the MCLC strategy is first used to fill in the auxiliary variables to facilitate the subsequent weak supervision method modeling. Because the weak supervision method mainly focuses on the problem of label scarcity, the impact of the vacancies of auxiliary variables on modeling is mainly that it limits the data perception field of the model, makes the model overly dependent on the only data, and reduces the modeling accuracy of the model. After completing the auxiliary variable filling work, the data with only the vacancies of quality variables show the characteristics of weak supervision learning. In order to achieve accurate modeling, a generative weak supervision model with supervised subnetwork and unsupervised subnetwork is constructed to learn supervised data and unsupervised data respectively, and the two are combined to form a generative weak supervision model.

[0007] The first object of the present invention is to provide a method for predicting ammonia nitrogen concentration in a multi-sampling rate sewage treatment process, the method comprising:

[0008] Step 1: Obtain the input variables X and quality variables Y of the sewage treatment process, and divide the input variables into labeled samples and unlabeled samples based on whether the quality variables are included;

[0009] Step 2: construct a hierarchical masked weakly supervised HMWS soft sensing model based on the masked coarse-grained linear filling MCLC strategy and the generative weakly supervised network GWSNET; wherein the generative weakly supervised network consists of two sub-networks, a supervised sub-network and an unsupervised sub-network;

[0010] Step 3: Calculate the unsampled points of each auxiliary variable obtained in step 1 by using the MCLC strategy and the mask matrix, and use convolution coarse-grained linear calculation to complete the missing value filling;

[0011] Step 4, constructing a hierarchical masked weakly supervised HMWS soft sensing model based on the data training after filling the missing values ​​in step 3;

[0012] Step 5: In real time, the input variable X is collected and input into the trained hierarchical mask weak supervision HMWS soft measurement model to obtain the corresponding quality variable prediction value.

[0013] Optionally, the supervised subnetwork includes a pre-neural network and a generator network, wherein the pre-neural network is used to generate pseudo labels for unlabeled samples, so that the pseudo labels and labeled samples and their corresponding labels are input as training data into the generator network for training, thereby obtaining latent variable distribution information and determining encoding and decoding parameters of the generator network;

[0014] The unsupervised sub-network constructs its own encoder and decoder according to the latent variable distribution information obtained by the supervised sub-network, and uses unlabeled samples and their corresponding pseudo labels for training to determine its own encoder and decoder parameters.

[0015] Optionally, step 3 includes:

[0016] The initial multi-sampling rate data auxiliary variable set is denoted as X o , then the MCLC strategy first detects whether the data is missing and marks it, marks the data that exists in the data set, and records it as a null value if it does not exist, and generates the corresponding mask matrix X mask , which ensures that only the data at the missing value location is obtained, while the data points with actual values ​​remain unchanged. The specific calculation process of MCLC is:

[0017] X s =L(X mask ,X o )+X o (1)

[0018] L(x1,x2)=WC(x2)*(x1*x2) (1)

[0019] Among them, X s is the final calculation result after filling, L(x1, x2) is the linear operation function of the convolution kernel element, WC(x) means that after the convolution operation on the input data at the coarse-grained level, linear weighting is completed according to the distance between the elements in the convolution kernel and the point to be calculated to fill the data at the missing value position, and * is the corresponding multiplication of matrix elements of the same dimension.

[0020] Optionally, step 4 includes:

[0021] Step 4.1: Use the pre-processing neural network in the supervised sub-network to generate pseudo labels for the unlabeled samples. Use the generated pseudo labels and the labeled samples and their labels to train the supervised sub-network. Use the encoder of the generative network to encode the input samples and calculate the latent variable distribution. Then use the corresponding decoder to obtain the quality variable prediction value of some samples. Calculate the loss and repeat the cycle multiple times until the latent variable distribution close to the distribution characteristics of the real industrial process data is obtained.

[0022] Step 4.2: The pseudo-labels generated by the pre-neural network of the supervised sub-network are passed to the unsupervised sub-network as supplementary information to guide the construction of its encoder and decoder, giving the unsupervised sub-network a latent variable distribution close to the true distribution; and the loss function is obtained through variational inference to improve the accuracy of the predicted value, ultimately obtaining the predicted value of the quality variable that can accurately characterize the process characteristics, and realizing the online prediction of ammonia nitrogen concentration in the wastewater treatment process with multiple sampling rates.

[0023] Optionally, the input variables X of the sewage treatment process include the concentration of active heterotrophic bacteria, the concentration of insoluble particulate non-biodegradable organic matter, the concentration of inert substances in biological solids attenuation, the concentration of nitrate nitrogen, the concentration of active autotrophic bacteria, the concentration of soluble biodegradable organic nitrogen, the concentration of dissolved oxygen, the concentration of soluble rapidly biodegradable organic matter, and the concentration of insoluble particulate biodegradable organic nitrogen; the quality variable Y of the sewage treatment process is the effluent ammonia nitrogen concentration.

[0024] Optionally, the sampling frequency of active autotrophic bacteria in the input variable X of the sewage treatment process is different from that of other parameters, and each input variable is obtained through hardware sensors set at various parts of the sewage treatment process; the quality variable Y is obtained through offline laboratory analysis, and its sampling frequency is different from that of all input variables X.

[0025] Optionally, the loss function of the supervised sub-network is:

[0026]

[0027] Among them, y represents the label of the labeled data, y f represents the pseudo labels of unlabeled data, and θ are Gaussian distribution parameters, p θ represents the posterior Gaussian distribution, and represents the mean vector and the dth element of the covariance matrix of the latent variable distribution, x l represents the input variable vector of the standardized multivariate Gaussian distribution, D is the dimension of the latent variable and I represents the identity matrix.

[0028] Optionally, the loss function of the unsupervised sub-network is:

[0029]

[0030] in, and are the mean vector and covariance matrix obtained in the unsupervised sub-network.

[0031] The second object of the present invention is to provide an ammonia nitrogen concentration prediction system for a multi-sampling rate sewage treatment process, the system comprising a hardware sensor and a processor, wherein the hardware sensor is used to collect the concentration of active heterotrophic bacteria, the concentration of insoluble particulate non-biodegradable organic matter, the concentration of inert substances in biological solids decay, the concentration of nitrate nitrogen, the concentration of active autotrophic bacteria, the concentration of soluble biodegradable organic nitrogen, the concentration of dissolved oxygen, the concentration of soluble rapidly biodegradable organic matter, and the concentration of insoluble particulate biodegradable organic nitrogen; the processor is used to use the above method to realize the prediction of the effluent ammonia nitrogen concentration.

[0032] The present invention also provides an application of the above method in the field of sewage treatment, for realizing the prediction of water quality parameters in the field of sewage treatment, wherein the water quality parameters include dissolved oxygen concentration, nitrate nitrogen concentration, total nitrogen and total phosphorus.

[0033] The beneficial effects of the present invention are:

[0034] 1. A HMWS soft sensing model is proposed. Through coarse-grained complementarity and weakly supervised learning, HMWS can learn explicit synthetic representations of multi-sampling rate process dynamics, thereby alleviating the problem of missing values ​​in process data.

[0035] 2. Through mask matrices and convolutional coarse-grained computation, the proposed MCLC strategy can perceive missing data and fill in missing values, assisting the model in extracting complementary information between auxiliary variables. This MCLC strategy can broaden the scope of data perception and extend its application to various weakly supervised learning models.

[0036] 3. Design a GWSNET network, constructing an unsupervised branch subnetwork to refine the overall feature representation and a supervised subnetwork to strengthen the input-output relationship. GWSNET constructs two subnetworks and makes their latent variable distributions as close as possible, effectively utilizing unlabeled data and maximizing the prediction accuracy of effluent ammonia nitrogen concentration. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0038] Figure 1 This is a flow chart of the multi-sampling rate ammonia nitrogen concentration prediction method based on hierarchical mask weak supervision provided by this application;

[0039] Figure 2 It is a schematic diagram of the linear complement of auxiliary variable mask convolution.

[0040] Figure 3 It is the structure diagram of the generative weakly supervised network GWSNET.

[0041] Figure 4 It is a process flow chart of sewage treatment process.

[0042] Figure 5 It is the RMSE heat map of the ammonia nitrogen concentration prediction based on the number of iterations of each sub-network.

[0043] Figure 6 This is a simulation diagram of the ammonia nitrogen concentration prediction results of the HMWS model under the optimal parameters.

[0044] Figure 7 These are error line graphs for predicting ammonia nitrogen concentration using the method of the present application and five existing methods. DETAILED DESCRIPTION

[0045] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0046] First, the relevant basic knowledge involved in this application is introduced as follows:

[0047] 1. The concept of multiple sampling rates;

[0048] Variables in chemical industrial processes are often sampled at different rates. For example, process variables measured by hardware sensors are often recorded multiple times per second, while quality variables are typically obtained through offline laboratory analysis, with each measurement taking a long time. Process data, typically with three or more sampling rates, is considered multi-rate data. In this context, auxiliary variables suffer from information loss, leading to an over-reliance on existing samples for modeling. Consequently, data availability is severely limited during modeling, hindering data-driven soft sensor models from accurately capturing the underlying process data characteristics.

[0049] The important tasks of soft sensing models for multi-sampling rate data are to alleviate the impact of data imbalance caused by different sampling rates, unify multiple variables into a common space, and capture quality-related information from low-sampling rate data to characterize system dynamic characteristics and learn multi-sampling rate process characteristics.

[0050] 2. Sewage treatment process;

[0051] Wastewater treatment generally includes physical treatment methods, chemical treatment methods, and biological treatment methods. Physical treatment methods mainly remove suspended solids, grease, and larger particles in wastewater through physical action. Chemical treatment methods remove dissolved pollutants, heavy metals, and nutrients in wastewater through chemical reactions. Biological treatment methods generally use the metabolic action of microorganisms to decompose organic matter in wastewater. In actual treatment plans, a combination of multiple methods may be involved, such as the activated sludge process:

[0052] The activated sludge process is a biological sewage treatment method that relies on a variety of microbial communities to decompose soluble and colloidal pollutants in sewage into harmless substances, thereby achieving sewage purification. Figure 4 As shown, the sewage to be treated passes through two anaerobic tanks (tanks 1 and 2), three aerobic tanks (tanks 3, 4, and 5), and a secondary sedimentation tank for treatment in sequence. The process can be divided into four stages: primary treatment, secondary treatment, tertiary treatment, and sludge treatment. Primary treatment mainly involves preliminary filtration and purification of sewage by physical means. Secondary treatment is the main biochemical reaction step in the sewage treatment process, relying on microbial activity to eliminate degradable substances in sewage and reduce ammonia nitrogen concentrations to purify sewage. Its main processes include a variety of technologies, including anoxic-aerobic, anaerobic-anoxic-aerobic, etc. The core principle is to use microorganisms in activated sludge to biosorb and decompose organic pollutants, removing pollutants such as nitrogen and phosphorus. Tertiary treatment mainly involves reprocessing the sewage after treatment in the secondary sedimentation tank to remove residual pollutants. The final sludge treatment is to fully utilize sludge resources and reduce environmental impact.

[0053] The following embodiments of the present invention are introduced using the activated sludge method for treating sewage as an example.

[0054] Example 1

[0055] This embodiment provides a method for predicting ammonia nitrogen concentration in a sewage treatment process with multiple sampling rates, which is used to predict ammonia nitrogen concentration in a sewage treatment process. Figure 1 As shown, the method includes:

[0056] Step 1: Obtain the input variables X and quality variables Y of the sewage treatment process, and divide the input variables into labeled samples and unlabeled samples based on whether the quality variables are included;

[0057] The input variables X involved in the wastewater treatment process include the concentration of active heterotrophic bacteria, the concentration of insoluble particulate non-biodegradable organic matter, the concentration of inert matter in biosolids decay, the concentration of nitrate nitrogen, the concentration of active autotrophic bacteria, the concentration of soluble biodegradable organic nitrogen, the concentration of dissolved oxygen, the concentration of soluble rapidly biodegradable organic matter, and the concentration of insoluble particulate biodegradable organic nitrogen. The quality variable Y is the ammonia nitrogen concentration in the supernatant after treatment in the secondary sedimentation tank. The parameters in the input variable X are obtained through hardware sensors installed in various parts of the wastewater treatment process, while the quality variable Y is obtained through offline laboratory analysis.

[0058] In this embodiment, the sampling frequency of active autotrophic bacteria in the input variable X is different from that of other parameters. Since the quality variable Y is obtained through offline laboratory analysis, its sampling frequency is different from that of all input variables X. That is, the present invention is targeted at chemical industrial processes with multiple sampling rates.

[0059] The data containing both the input variable X and the corresponding quality variable Y is regarded as a labeled sample, and the input variable X data without the corresponding quality variable Y is regarded as an unlabeled sample.

[0060] Step 2: Based on the Mask Coarse-Grained Linear Complementation (MCLC) strategy and the Generative Weekly-Supervised Network (GWSNet), a hierarchical Mask Weekly-Supervised (HMWS) soft-sensing model is constructed.

[0061] The hierarchical masked weakly supervised HMWS soft sensor model constructed in this step includes a convolution coarse-grained linear padding part and a generative weakly supervised network part. The convolution coarse-grained linear padding part adopts convolution coarseness calculation and uses a mask matrix to confirm the sampling of auxiliary variables. The features of the convolution kernel are linearly operated to obtain the patched weakly supervised samples. The generative weakly supervised network GWSNET part contains supervised and unsupervised sub-networks. A supervised sub-network is constructed by taking labeled samples as input to improve the encoder's regression representation ability. The pseudo-labels of the unlabeled samples estimated by the pre-neural network regressor are injected into the supervised module as a supplement, so as to fully utilize the unlabeled data while constraining the generated values ​​to be close to the actual labels. The distribution of the supervised sub-network is then extracted to construct the unsupervised sub-network. This similar distribution gives the unsupervised sub-network the ability to model a large amount of unlabeled data. Finally, the two are combined to construct a generative weakly supervised network GWSNET to achieve prediction of low-quality variable sampling rate data. The basic model of the modeling of the two sub-networks is the conditional variational autoencoder.

[0062] Step 3: Calculate the missing data of each auxiliary variable obtained in step 1 by using the MCLC strategy and the mask matrix, and use MCLC calculation to complete the missing value filling;

[0063] like Figure 2 As shown in Figure 1, through mask matrices and coarse-grained convolution computation, the MCLC strategy can capture missing data and assist the hierarchical masked weakly supervised HMWS soft sensing model in extracting complementary information between auxiliary variables. The MCLC strategy has two major advantages: first, it can expand the scope of data perception; second, it has good model universality and is adaptable to various weakly supervised learning frameworks.

[0064] The initial multi-sampling rate data auxiliary variable set is denoted as X o , then the MCLC strategy first detects whether the data is missing and marks it, marks the data that exists in the data set, and records it as a null value if it does not exist, and generates the corresponding mask matrix X mask , which ensures that only the data at the missing value location is obtained, while the data points with actual values ​​remain unchanged. The specific calculation process of MCLC is:

[0065] X s =L(X mask ,X o )+X o (1)

[0066] L(x1,x2)=WC(x2)*(x1*x2) (2)

[0067] Among them, X sThe final calculation result after filling is shown in Figure 1. L(x1, x2) is the linear operation function of the convolution kernel elements. WC(x) represents the linear weighting of the input data based on the distance between the elements within the convolution kernel and the point to be calculated after performing a coarse-grained convolution operation on the input data to fill the missing value locations. * represents the corresponding multiplication of matrix elements of the same dimension. After inputting auxiliary variables in this way, the resulting data includes quality variables sampled at a lower rate and auxiliary variables sampled at a uniform rate. To distinguish it from semi-supervised data, data processed by the MCLC strategy is referred to as weakly supervised data. This strategy expands the scope of data perception and facilitates soft sensing models to learn data at multiple sampling rates in a hierarchical manner.

[0068] Step 4, constructing a hierarchical masked weakly supervised HMWS soft sensing model based on the data training after filling the missing values ​​in step 3;

[0069] The two branch subnetworks are the supervised and unsupervised subnetworks, respectively. The supervised subnetwork consists of a pre-neural network and a generator network. First, the pre-neural network is used to generate pseudo-labels for unlabeled samples. These pseudo-labels and labeled samples are then used to train the supervised subnetwork to enhance the encoder's feature learning capabilities. The unsupervised subnetwork is trained using unlabeled data, and the pseudo-labels generated by the pre-neural network in the supervised subnetwork are transferred to the unsupervised subnetwork as supplementary information to ensure full utilization of the unlabeled data while maintaining consistency between the predicted output and the actual label. The unsupervised subnetwork is constructed using the latent variable distribution information obtained by the supervised subnetwork. This distribution approximation enables GWSNET to effectively model unlabeled data. Finally, the two subnetworks are combined to construct GWSNET, which improves the prediction accuracy of low-sampled quality variable data.

[0070] Specifically, this step includes:

[0071] Step 4.1: Use the pre-processing neural network in the supervised sub-network to generate pseudo labels for the unlabeled samples. Use the generated pseudo labels and the labeled samples and their labels to train the supervised sub-network. Use the encoder of the generative network to encode the input samples and calculate the latent variable distribution. Then use the corresponding decoder to obtain the quality variable prediction value of some samples. Calculate the loss and repeat the cycle multiple times until the latent variable distribution close to the distribution characteristics of the real industrial process data is obtained.

[0072] The supervised sub-network includes a pre-neural network and a generator network (e.g., a conditional variational autoencoder). The pre-neural network estimates the pseudo-label y of the unlabeled data. f , helps capture the nonlinear relationship between input x and output y, and acts as a constraint on the latent variable z, ensuring that data sampled from a Gaussian distribution are closer to the original data. fAs an additional condition, the log-likelihood of the supervised subnetwork can be expressed as:

[0073]

[0074] Among them, x r represents the input data with actual labels, Obey the parameters Gaussian distribution, used to estimate the true posterior Gaussian distribution p θ The log-likelihood function can be expressed as the ELBO term and the KL divergence. The ELBO term can be further written as:

[0075]

[0076] Among them, p θ (z|y f ) is the prior distribution, p θ (x r ,y|y f ,z) is the distribution obtained from the decoder. Therefore, the learning objective of the supervised sub-network can be expressed as the following maximization problem:

[0077]

[0078] In KL divergence In the item, p θ (z|y f ) is a standardized multivariate Gaussian distribution, is an approximate posterior that follows a multivariate Gaussian distribution, so the KL divergence term can be further written as:

[0079]

[0080] in, and Represents the mean vector of the latent variable distribution and the dth element of the covariance matrix. To distinguish it from the previous calculation steps, x l A vector of auxiliary variables representing the normalized Gaussian distribution.

[0081] By expanding the KL divergence term, the loss function of the supervised subnetwork can be expressed as:

[0082]

[0083] in, D is the dimension of the latent variable.

[0084] Step 4.2: The pseudo labels generated by the front neural network in the supervised sub-network are transmitted to the unsupervised sub-network as supplementary information, and the latent variable distribution information obtained by the supervised sub-network is used to construct the unsupervised sub-network;

[0085] The architecture of the unsupervised sub-network is similar to that of the supervised sub-network, including an encoder and a decoder. The main difference between the two is the input and output. Specifically, the unsupervised sub-network uses data y f , does not involve y. The learning objective of the unsupervised sub-network is given by:

[0086]

[0087] Among them, the ELBO term can be expressed as:

[0088]

[0089] To ensure that the supervised and unsupervised sub-networks have as similar latent variable distributions as possible, the unsupervised sub-network is designed to solve the following maximization problem:

[0090]

[0091] The KL divergence term in the above formula can be written as:

[0092]

[0093] in, and is the mean vector and covariance matrix obtained from the unsupervised sub-network. To ensure that the latent variable distributions of the supervised and unsupervised sub-networks are as similar as possible, and is passed from the supervised sub-network to the unsupervised sub-network. Therefore, the loss function of the unsupervised sub-network can be expressed as:

[0094]

[0095] Among them, the latent variable It is sampled by the reparameterized Monte Carlo method.

[0096] Step 4.3: Combine the encoder of the unsupervised sub-network and the decoder of the supervised sub-network to construct the complete generative weakly supervised network GWSNET;

[0097] like Figure 3 As shown in Figure 2, due to the high latent variable similarity between the supervised and unsupervised subnetworks, they can be integrated to construct the GWSNET. In the GWSNET, pseudo labels for unlabeled samples are generated by a pre-neural network regressor, allowing the model to fully utilize unlabeled data and address the label shortage problem. Simultaneously, these pseudo labels are used as constraints to effectively guide the generation process in the GWSNET, thereby improving the accuracy of soft sensor modeling.

[0098] GWSNET is constructed by combining the encoder of the unsupervised sub-network with the decoder of the supervised sub-network. This architecture is beneficial for predicting data with low sampling rate quality variables.

[0099] Step 5: Real-time collection of input variables and input into GWSNET to obtain the predicted value of ammonia nitrogen concentration Y ^ ;

[0100] Example 2

[0101] This embodiment provides a method for predicting ammonia nitrogen concentration in a multi-sampling rate sewage treatment process, the method comprising:

[0102] Step 1, obtain input variable X and quality variable y;

[0103] This embodiment takes the activated sludge process as an example. The collected input variables X and auxiliary variables include the concentration of active heterotrophic bacteria, the concentration of insoluble particulate non-biodegradable organic matter, the concentration of inert substances in biosolids decay, the concentration of nitrate nitrogen, the concentration of active autotrophic bacteria, the concentration of soluble biodegradable organic nitrogen, the concentration of dissolved oxygen, the concentration of soluble rapidly biodegradable organic matter, and the concentration of insoluble particulate biodegradable organic nitrogen. The sampling interval for the active autotrophic bacteria data is 30 minutes, the sampling interval for the other auxiliary variables is 15 minutes, and the sampling period for the quality variable Y, ammonia nitrogen concentration, is 75 minutes. The input variables X and the quality variable Y are shown in Table 1 below:

[0104] Table 1 Description of variables in the biochemical reaction pool sampling process

[0105]

[0106]

[0107] As can be seen from Table 1 above, in the multi-sampling rate sewage treatment process targeted by the present invention, among the easily collected auxiliary variables X, the sampling rate of the active autotrophic bacteria data is different from that of other auxiliary variables, and the sampling rate of the ammonia nitrogen concentration as the quality variable Y is different from that of all the auxiliary variables X. That is, this embodiment targets a chemical industrial process with three sampling rates.

[0108] Step 2: Construct a hierarchical masked weakly supervised HMWS model based on the MCLC strategy and GWSNET network;

[0109] The construction process of the hierarchical masked weakly supervised HMWS model can be referred to the introduction in the above embodiment 1, and will not be repeated in this embodiment. This embodiment uses two real chemical industrial processes to verify the ability of the constructed hierarchical masked weakly supervised HMWS model to predict the effluent ammonia nitrogen concentration in the sewage treatment process with multiple sampling rates:

[0110] Specifically, in this embodiment, the root mean square error (RMSE), mean absolute error (MAE) and correlation coefficient (R 2 ) is used as a performance evaluation indicator to evaluate the prediction results.

[0111] Step 3: Calculate the unsampled points of each auxiliary variable through the MCLC strategy combined with the mask matrix, and use convolution coarse-grained linear calculation to complete the missing value filling;

[0112] Step 4: Build the initial offline model of the two branch sub-networks, initialize the parameters, and train the front neural network;

[0113] Model parameters have a significant impact on modeling capabilities and predictive performance. In order to explore the impact of key parameters on the performance of the model of this application and obtain their appropriate setting values, parameter optimization experiments were conducted on the latent variable dimension, supervised sub-network and unsupervised sub-network iteration times by combining empirical methods with grid search methods.

[0114] Table 2 Ammonia nitrogen concentration prediction evaluation index of HMWS model under different latent variable dimensions

[0115]

[0116] Table 2 lists the prediction performance indicators of the HMWS model for ammonia nitrogen concentration under different latent variable dimensions; Figure 5 The following is a heatmap of the RMSE of the prediction results for different iteration numbers of the two branch subnetworks. Through multiple experiments combined with performance metric analysis, the optimal latent variable dimension was determined to be 16, the number of iterations for the supervised subnetwork to be 250, and the number of iterations for the unsupervised subnetwork to be 1500. The latent variable dimension determines the model's ability to learn to represent the latent variable subspace, and the number of iterations for each of the two subnetworks is strongly correlated with the performance of the final weakly supervised network.

[0117] like Figure 3 As shown in the figure, the encoder of the unsupervised sub-network and the decoder of the supervised sub-network are combined to construct a complete generative weakly supervised network GWSNET;

[0118] Step 5: Real-time collection of input variables and input into GWSNET to obtain the predicted value of ammonia nitrogen concentration Y ^ ;

[0119] In order to fully verify the effectiveness of the proposed HMWS, five representative multi-sampling rate methods were selected for comparison: Method 1, a soft sensing model based on k-nearest neighbor (KNN) filling, can be found in X. Zhang, Y. Nojima, H. Ishibuchi, W. Hu and S. Wang, "Prediction by Fuzzy Clustering and KNN on Validation Data With Parallel Ensemble of Interpretable TSK Fuzzy Classifiers," IEEE Transactions on Systems, Man, and Cybernetics: Systems., vol. 52, no. 1, pp. 400-414, Jan. 2022.;

[0120] Method 2: Interpolation soft sensing model based on random forest (MF). For reference, see F. Aracri, M. Giovanna Bianco, A. Quattrone and A. Sarica, “Imputation of missing clinical, cognitive and neuroimaging data of Dementia using missForest, a Random Forest based algorithm,” 2023 IEEE 36th International Symposium on Computer-Based Medical Systems (CBMS), L'Aquila., pp. 684-688, 2023.

[0121] Method 3, Variational Progressive-Transfer Network (VPTN), can be found in Z. Chai, C. Zhao and B. Huang, “Variational Progressive-Transfer Network for Soft Sensing of Multirate Industrial Processes,” IEEE Transactions on Cybernetics, vol. 52, no. 12, pp. 12882-12892, Dec. 2022.

[0122] Method 4, flexible clock recurrent neural network (FCW-RNN), can be found in S. Chang, X. Chen, and C. Zhao, “Flexible Clockwork Recurrent Neural Network for multirate industrial softsensor,” Journal of Process Control, vol. 119, pp. 86-100, Nov. 2022.

[0123] Method 5: Time-aware long short-term memory network (TLSTM), see IMBaytas, et al. "Patient subtyping via time-aware LSTM networks," Proceedings of the Proceedings of the 23rd ACM SIGKDDInternational Conference on Knowledge Discovery and DataMining, pp. 65-74, 2017.

[0124] Among them, Method 1 KNN and Method 2 MF use data interpolation models based on machine learning technology. Method 4 FCW-RNN, Method 3 VPTN, and Method 5 TLSTM use models that are specially constructed by modifying the network structure for multi-sampling rate data.

[0125] Table 3 Ammonia nitrogen concentration prediction performance evaluation indicators of the present method and five existing methods

[0126]

[0127]

[0128] The comparison of the ammonia nitrogen concentration prediction results of each multi-sampling rate method is shown in Table 3. It is observed that in the multi-sampling rate sewage data modeling task, the correlation coefficient of the prediction method using the HMWS model proposed in this application is 0.9986, and the RMSE is the smallest. First, a comprehensive comparison shows that KNN, as a machine learning filling method, is affected by the inherent precision limit of its results; and although the MF method is also limited by its own precision, it has the ability to update parameters, and the results are improved. Secondly, the FCW-RNN model maps different sampling rates to a unified latent space; the VPTN method divides the multi-sampling rate data into blocks and then unifies the modeling through a deep variational structure; the TLSTM model combines the sampling interval to reduce the impact of the low proportion of quality-related information. The above three models can overcome the impact of multiple sampling rates on the soft measurement model to a certain extent, but the prediction results are not optimal. After filling in the missing auxiliary variables, the HMWS model proposed in this application realizes online modeling of multi-sampling rate data through weak supervision. Compared with the representative multi-sampling rate methods, the prediction performance is improved.

[0129] In order to further evaluate the effectiveness of the HMWS model proposed in this invention in predicting different sample points, Figure 6 The prediction error line graphs for each method are shown in Figure 2. The closer the colored curves are to the zero-point grid line, the better the fit. The results show that the proposed method achieves a smaller deviation between predicted and true values ​​than other representative multi-sampling rate methods, enabling better online monitoring of multi-sampling rate industrial processes.

[0130] The HMWS model consists of two key modules: MCLC and GWSNET. Ablation experiments were conducted to verify the effectiveness of each module. Specifically, because MCLC is an auxiliary variable imputation strategy, it was utilized in conjunction with the TLSTM model for validation. In contrast, the GWSNET module, with its own online modeling capabilities, served as an independent experimental unit. These ablation experiments provide a reference for verifying the individual contribution of each module to the overall performance of the HMWS model.

[0131] Table 4 Ammonia nitrogen concentration prediction and evaluation indicators of ablation test

[0132]

[0133] The results are shown in Table 4, and it can be seen that when either module is eliminated, the prediction performance degrades. The MCLC module expands the existing modeling space, while the GWSNET module provides HMWS with weakly supervised online modeling capabilities. MCLC and GWSNET play a key role in improving the modeling accuracy of HMWS, highlighting their soft sensor modeling capabilities for multi-rate data.

[0134] This paper proposes a novel soft-sensor modeling framework, HMWS, for predicting ammonia nitrogen concentration in wastewater treatment processes under multi-sampling rate conditions. First, in the proposed MCLC strategy, multi-sampling rate data are hierarchically coarse-grained and estimated using a mask matrix and convolutional computations. In the designed GWSNET, supervised and unsupervised branching subnetworks are constructed to simultaneously capture the overall feature representation and input-output relationships. Experiments conducted in a wastewater treatment process demonstrate the precise performance and advantages of HMWS in soft-sensor modeling of multi-sampling rate processes.

[0135] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.

[0136] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for predicting ammonia nitrogen concentration in a multi-sampling rate sewage treatment process, characterized in that: The method comprises: Step 1: Obtain the input variables X and quality variables Y of the sewage treatment process, and divide the input variables into labeled samples and unlabeled samples based on whether the quality variables are included; Step 2: construct a hierarchical masked weakly supervised HMWS soft sensing model based on the masked coarse-grained linear filling MCLC strategy and the generative weakly supervised network GWSNET; wherein the generative weakly supervised network consists of two sub-networks, a supervised sub-network and an unsupervised sub-network; Step 3: Calculate the missing data of each auxiliary variable obtained in step 1 by using the MCLC strategy and the mask matrix, and use the convolution kernel elements to linearly fill in the missing values; Step 4, constructing a hierarchical masked weakly supervised HMWS soft sensing model based on the data training after filling the missing values ​​in step 3; Step 5: In real time, the input variable X is collected and input into the trained hierarchical mask weak supervision HMWS soft measurement model to obtain the corresponding quality variable prediction value.

2. The method according to claim 1, characterized in that The supervised subnetwork includes a pre-neural network and a generator network, wherein the pre-neural network is used to generate pseudo labels for unlabeled samples, so that the pseudo labels and labeled samples and their corresponding labels are input as training data into the generator network for training, thereby obtaining latent variable distribution information and determining encoding and decoding parameters of the generator network; The unsupervised sub-network constructs its own encoder and decoder based on the latent variable distribution information obtained by the supervised sub-network, and uses unlabeled samples and their corresponding pseudo labels for training to determine its own encoder and decoder parameters.

3. The method according to claim 2, characterized in that The step 3 comprises: Through the MCLC strategy, combined with the mask matrix, the data missing state in step 1 is obtained and the missing values ​​are filled. Through the mask matrix combined with the linear calculation of the convolution kernel elements, the auxiliary hierarchical mask weakly supervised HMWS soft measurement model is used to extract the complementary information between the auxiliary variables and supplement the missing values ​​of the samples to increase the available data.

4. The method according to claim 3, characterized in that The step 4 comprises: Step 4.1: Use the pre-processing neural network in the supervised sub-network to generate pseudo labels for the unlabeled samples. Use the generated pseudo labels and the labeled samples and their labels to train the supervised sub-network. Use the encoder of the generative network to encode the input samples and calculate the latent variable distribution. Then use the corresponding decoder to obtain the quality variable prediction value of some samples. Calculate the loss and repeat the cycle multiple times until the latent variable distribution close to the distribution characteristics of the real industrial process data is obtained. Step 4.2: The pseudo-labels generated by the supervised subnetwork's pre-processing neural network are passed to the unsupervised subnetwork as supplementary information to guide the construction of its encoder and decoder, giving the unsupervised subnetwork a latent variable distribution close to the true distribution. A loss function is then calculated using variational inference to improve the accuracy of the predicted values. Ultimately, predicted values ​​of the quality variables are obtained that accurately characterize the process characteristics, enabling online prediction of ammonia nitrogen concentration in wastewater treatment processes with multiple sampling rates. Step 4.3: Combine the encoder of the unsupervised sub-network and the decoder of the supervised sub-network to construct the complete generative weakly supervised network GWSNET.

5. The method according to claim 4, characterized in that The input variables X of the sewage treatment process include the concentration of active heterotrophic bacteria, the concentration of insoluble particulate non-biodegradable organic matter, the concentration of inert substances in biosolids attenuation, the concentration of nitrate nitrogen, the concentration of active autotrophic bacteria, the concentration of soluble biodegradable organic nitrogen, the concentration of dissolved oxygen, the concentration of soluble rapidly biodegradable organic matter, and the concentration of insoluble particulate biodegradable organic nitrogen; the quality variable Y of the sewage treatment process is the effluent ammonia nitrogen concentration.

6. The method according to claim 5, characterized in that The sampling frequency of active autotrophic bacteria in the input variable X of the sewage treatment process is different from that of other parameters. Each input variable is obtained through hardware sensors installed in various parts of the sewage treatment process; the quality variable Y is obtained through offline laboratory analysis, and its sampling frequency is different from that of all input variables X.

7. The method according to claim 6, characterized in that The loss function of the supervised sub-network is: Among them, y represents the label of the labeled data, y f represents the pseudo labels of unlabeled data, and θ are Gaussian distribution parameters, p θ represents the posterior Gaussian distribution, and represents the mean vector and the dth element of the covariance matrix of the latent variable distribution, x l represents the input variable vector of the standardized multivariate Gaussian distribution, D is the dimension of the latent variable and I represents the identity matrix.

8. The method according to claim 7, characterized in that The loss function of the unsupervised sub-network is: in, and are the mean vector and covariance matrix obtained in the unsupervised sub-network.

9. An ammonia nitrogen concentration prediction system for multi-sampling rate sewage treatment process, characterized by: The system includes a hardware sensor and a processor, wherein the hardware sensor is used to collect the concentration of active heterotrophic bacteria, the concentration of insoluble particulate non-biodegradable organic matter, the concentration of inert substances in biological solids attenuation, the concentration of nitrate nitrogen, the concentration of active autotrophic bacteria, the concentration of soluble biodegradable organic nitrogen, the concentration of dissolved oxygen, the concentration of soluble rapidly biodegradable organic matter, and the concentration of insoluble particulate biodegradable organic nitrogen; and the processor is used to use the method described in any one of claims 1 to 8 to predict the ammonia nitrogen concentration in the effluent.

10. Application of the method according to any one of claims 1 to 8 in the field of sewage treatment.

Citation Information

Cited By

  • Multi-time-scale dynamic soft measurement modeling method for multi-sampling-rate data

    CN121211990A