Sewage treatment plant effluent prediction method based on multi-task learning
By using a multi-task learning-based effluent prediction model that comprehensively considers multi-dimensional data from wastewater treatment plants, the model solves the problems of complex calibration of mechanistic models and high computational costs of machine learning models. It achieves efficient and accurate multi-indicator prediction, thereby improving the intelligent management level of wastewater treatment plants.
Patent Information
- Application Number
- CN202511029835.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-28
AI Technical Summary
In existing wastewater treatment plant effluent prediction technologies, the calibration of mechanistic model parameters is complex, time-consuming, and prone to distortion, while machine learning models have high computational costs. The modeling of multiple indicators ignores the mutual influence of processes, resulting in insufficient prediction accuracy and generalization ability.
A multi-task learning-based method is adopted to establish an effluent prediction model. Through a data sharing layer, multiple attention network layers, and an output layer, the model comprehensively considers the temporal dependence and mutual influence of influent water quality, process, environment, and unit data, and uses long short-term memory networks and multi-head attention mechanisms for prediction.
It significantly improves the accuracy of effluent water quality prediction and process control efficiency, reduces computational redundancy, adapts to the multi-indicator prediction needs in complex industrial scenarios, and enhances the robustness and interpretability of the model.
Smart Images

Figure CN121034470A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of technology, specifically to a method for predicting effluent from wastewater treatment plants based on multi-task learning. Background Technology
[0002] Wastewater treatment plants, as a crucial component of urban infrastructure, bear the key responsibility of purifying municipal sewage and industrial wastewater. Their operational efficiency directly impacts water quality and ecological security. While current wastewater treatment infrastructure is nearly complete, achieving almost 100% treatment rates, most plants suffer from low levels of automation and relatively inefficient operation and management, resulting in higher energy and chemical consumption. Therefore, it is necessary to understand the wastewater treatment process and construct effluent prediction models to achieve intelligent prediction of the wastewater treatment process, thereby laying a solid foundation for promoting intelligent operation and management of wastewater treatment plants.
[0003] Existing wastewater treatment plant effluent prediction technologies include mechanistic models and machine learning models. Mechanistic models are based on prior knowledge of the wastewater treatment system and describe its behavior and dynamic characteristics through a set of mathematical equations. Taking the activated sludge series models proposed by the International Water Association as an example, the activated sludge model breaks down the entire system into sub-processes such as microbial growth and nitrification, and constructs a set of differential equations. By solving these equations, the concentration of pollutants in the effluent is analyzed and predicted. Machine learning models include traditional machine learning models such as decision trees and support vector machines, as well as deep learning models such as convolutional neural networks and recurrent neural networks. Deep learning models have strong nonlinear fitting capabilities and higher prediction accuracy, and are increasingly used for effluent prediction. Machine learning models can learn the mapping relationship between input and output indicators based on historical data, and can better cope with the nonlinearity, multiple couplings, and strong interference characteristics of the wastewater treatment process, achieving accurate prediction of effluent quality.
[0004] However, the performance of mechanistic models is affected by the accuracy of their parameters. If the parameters are inaccurate, the accuracy of the mechanistic model will be significantly negatively impacted. Therefore, accurate calibration of mechanistic model parameters is crucial for application. However, mechanistic models contain a large number of parameters, many of which require complex and time-consuming experimental calibration processes. The overall calibration process is time-consuming and labor-intensive. Furthermore, when the influent characteristics and biochemical reaction characteristics of wastewater treatment plants change, the mechanistic model parameters calibrated based on historical conditions will gradually become distorted, leading to a decrease in model prediction accuracy. This necessitates complex and time-consuming recalibration to restore the model's accuracy, making it unsuitable for the operation and control of wastewater treatment plants.
[0005] Wastewater treatment is a complex system involving the simultaneous removal of organic matter and nitrogen- and phosphorus-containing nutrients. These sub-processes interact and compete with each other. For example, influent COD is an important carbon source for heterotrophic microorganisms, and denitrifying bacteria compete for it with other heterotrophic bacteria. This leads to the interaction and coupling of carbon removal and denitrification processes, and excessive carbon removal may result in insufficient carbon source for denitrification. However, current machine learning methods for modeling multiple effluent indicators from biological treatment units often involve building a separate model for each indicator. This firstly ignores the interactions between processes, negatively impacting the model's accuracy and generalization ability; secondly, using multiple models to predict multiple indicators, especially when using deep learning models with a large number of parameters, incurs high computational and storage costs. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a method for predicting wastewater effluent from wastewater treatment plants based on multi-task learning.
[0007] A wastewater treatment plant effluent prediction method based on multi-task learning includes the following steps:
[0008] Acquire wastewater treatment data; wastewater treatment data includes influent water quality data, process data, environmental data, wastewater treatment unit data, and effluent index data;
[0009] An effluent prediction model is established based on wastewater treatment data. The inputs of the effluent prediction model are influent water quality data, process data, environmental data, and wastewater treatment unit data, and the output is the predicted effluent index data. The effluent prediction model includes a data sharing layer, multiple attention network layers, and multiple output layers connected in sequence. The data sharing layer is used to describe the temporal dependencies between the various inputs in the effluent prediction model. The attention network layers are used to extract the inputs related to each effluent index from the data sharing layer or the data sharing layer and other attention network layers. The output layers are used to fit the effluent index data to the extracted inputs based on the effluent indexes in the multi-task attention layer.
[0010] Based on the inputs of the water discharge prediction model and the actual water discharge prediction model, the water discharge index data for future times are predicted.
[0011] Note: The prediction model constructed by the above method can cover multi-dimensional data such as influent water quality, process, environment, and unit, ensuring that the model can fully capture various factors affecting the effluent. At the same time, it can comprehensively consider the temporal dependence and mutual influence between various input data, and extract relevant data that are more important to each effluent indicator for prediction. Through multi-task learning, multiple indicators are predicted simultaneously, which reduces computational redundancy and meets the actual process requirements for collaborative optimization of multiple objectives, significantly improving the accuracy of effluent water quality prediction and the efficiency of process control.
[0012] Furthermore, the influent water quality data includes influent ammonia nitrogen, total phosphorus, pH, COD, and SS; process data includes aeration volume, sludge concentration, return ratio, and dissolved oxygen concentration; environmental data includes plant ambient temperature and plant ambient humidity; wastewater treatment unit data includes monitoring data from sensors in each wastewater treatment unit; and effluent index data includes COD, ammonia nitrogen, and total phosphorus.
[0013] Note: The above content lists wastewater operation parameters, providing highly targeted and practical input features for subsequent model training. This enables the model to accurately capture the intrinsic correlation between various factors and effluent indicators. The monitoring data from the sensors of each wastewater treatment unit include: surface loading rate and sludge interface height in the primary sedimentation tank; ORP (oxidation-reduction potential) and VFA (volatile fatty acid online chromatography) in the anaerobic tank; DO concentration, sludge concentration (ultrasonic probe), and aeration pipe pressure in the aerobic tank; and transmembrane pressure and membrane flux in the MBR membrane tank.
[0014] Furthermore, the attention network layer includes a COD attention network layer, an ammonia nitrogen attention network layer, and a total phosphorus attention network layer, with each attention network layer separated from the others.
[0015] Note: The above-mentioned network layer corresponds to each output indicator, which can improve the model's ability to process complex water quality data and its prediction accuracy. It enables the model to learn and distinguish the correlation and differences between different water quality indicators more effectively, providing more accurate support for water quality prediction and management.
[0016] Furthermore, methods for establishing effluent prediction models based on wastewater treatment data include:
[0017] S1. Building the model structure:
[0018] A data sharing layer is established based on a long short-term memory network or a gated recurrent unit; all inputs in the data sharing layer are arranged in temporal order; an attention network layer is established using a multi-head attention mechanism; and the output layer is a fully connected layer.
[0019] S2, Training the Model:
[0020] The wastewater treatment data was divided into a training set and a test set in an 8:2 ratio.
[0021] The training set data is input into the water discharge prediction model, passing through the data sharing layer, attention network layer, and output layer in sequence to obtain the predicted water discharge index data. The error between the predicted value and the true value is calculated based on the loss function. The error is propagated from the output layer to the front layer by layer through the backpropagation algorithm, and the gradient of the parameters of each layer is calculated. The loss function is the mean squared error and the mean absolute error.
[0022] The Adam optimizer is used to update the model parameters based on the calculated gradients. The process of forward propagation, back propagation and parameter update is repeated until the preset maximum number of iterations is reached.
[0023] Then, the water discharge prediction model is evaluated using test set data. Commonly used evaluation metrics include root mean square error and coefficient of determination. The smaller the root mean square error, the smaller the prediction error of the model; the closer the coefficient of determination is to 1, the better the model fits the data.
[0024] The hyperparameters of the water discharge prediction model were tuned using Bayesian optimization methods. Once the optimal combination of hyperparameters was found, the model training was completed.
[0025] Explanation: The above method combines a data-sharing layer of a Long Short-Term Memory network or a gated recurrent unit with an attention network layer of a multi-head attention mechanism. This allows the model to effectively process time-series data and focus on key features, thus achieving a hierarchical fusion of time-series data and multi-head attention. Combined with a parameter-sharing mechanism of multi-task learning, it achieves a balance between efficiency, accuracy, and interpretability in multi-indicator prediction in complex industrial scenarios, thereby improving prediction accuracy. By using the Adam optimizer and Bayesian optimization methods to fine-tune the model, its reliability in practical applications is ensured.
[0026] Furthermore, the attention network layer decomposes the data of the data sharing layer into multiple subspaces, determines the attention weight of each subspace, and captures the inputs related to each water output indicator based on the attention weight of each subspace; wherein, the subspace refers to the feature association between at least two input data.
[0027] Explanation: By capturing multiple sets of feature association patterns in parallel, different subspaces can be modeled with diverse feature interactions such as linear / nonlinear and local / global features (e.g., one subspace focuses on the product effect of influent flow rate and temperature, while another subspace analyzes the threshold correlation between pH and dissolved oxygen). This allows for more accurate adaptation to the complex and heterogeneous driving relationships in multi-index prediction tasks. At the same time, weight visualization provides decision-making basis for process control (e.g., identifying the strong dependence of COD prediction on the dissolved oxygen-flow rate combination), significantly improving the robustness and interpretability of the model.
[0028] Furthermore, in the fully connected layer, the output nodes of the fully connected layer are determined according to the number of water output indicators.
[0029] Note: Independent output nodes ensure that the model can learn the mapping relationship of each indicator in a differentiated manner, improve prediction accuracy and generalization ability, and facilitate the addition of new indicators to adapt to dynamic business needs.
[0030] Furthermore, a temporal attention layer is provided between the attention network layer and the output layer to enhance the influence of key time points. This temporal attention layer is constructed using a deep learning framework.
[0031] Explanation: By dynamically allocating time step weights, the model's adaptability to long-range dependencies and sudden events is enhanced, thereby improving prediction robustness and interpretability.
[0032] Furthermore, the deep learning framework is either PyTorch or TensorFlow.
[0033] Note: The above functionality can be achieved using PyTorch or TensorFlow.
[0034] Furthermore, the processing steps of the temporal attention layer include:
[0035] Obtain the hidden state of each time step in the data sharing layer and calculate the attention weight of each time step. The hidden state of the time step refers to the change of the state of the data in the data sharing layer as arranged over time.
[0036] The hidden state is weighted according to the attention weight at each time step to generate a context vector that can focus on key time points;
[0037] The context vector is connected to the output layer. The output layer fits the corresponding water discharge index data based on the water discharge index in the multi-task attention layer, the extracted input, and the context vector.
[0038] Note: The above method, through spatiotemporal feature decoupling and dynamic weight allocation, can improve the model's ability to model long-term time-series dependencies and sudden events, thus benefiting the accuracy and practicality of handling complex industrial time-series forecasting tasks.
[0039] Furthermore, the effluent prediction model is used to predict the effluent indicators of the biochemical treatment unit in the wastewater treatment plant. The influent water quality data, process data, environmental data, and wastewater treatment unit data in the input of the effluent prediction model are all upstream data of the biochemical treatment unit, and the output is the predicted effluent indicator data of the biochemical treatment unit.
[0040] The beneficial effects of this invention are:
[0041] This invention combines a prediction model constructed using multi-task learning. This prediction model can cover multi-dimensional data such as influent water quality, process, environment, and unit, ensuring that the model can fully capture various factors affecting the effluent. At the same time, it can comprehensively consider the temporal dependencies and mutual influences between various input data based on the model, and extract relevant data that are more important to each effluent indicator for prediction. By predicting multiple indicators simultaneously through multi-task learning, it not only reduces computational redundancy, but also meets the actual process requirements for collaborative optimization of multiple objectives, significantly improving the accuracy of effluent water quality prediction and the efficiency of process control. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the water discharge prediction model structure according to an embodiment of the present invention;
[0043] Figure 2 This is a schematic diagram illustrating the prediction effect of the effluent prediction model in an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram illustrating the performance improvement of the water effluent prediction model based on multi-task learning in this embodiment of the invention compared to the model based on a single task. Detailed Implementation
[0045] To further illustrate the methods and effects of this invention, the technical solution of this invention will be clearly and completely described below in conjunction with experiments.
[0046] Example 1: A wastewater treatment plant effluent prediction method based on multi-task learning, comprising the following steps:
[0047] S1. Obtain wastewater treatment data; wastewater treatment data includes influent water quality data, process data, environmental data, wastewater treatment unit data, and effluent indicator data;
[0048] The aforementioned influent water quality data includes influent ammonia nitrogen, total phosphorus, pH, COD, and SS; process data includes aeration volume, sludge concentration, return ratio, and dissolved oxygen concentration; environmental data includes plant ambient temperature and humidity; wastewater treatment unit data includes monitoring data from sensors in each wastewater treatment unit; the wastewater treatment units include physical treatment units, biological treatment units, chemical treatment units, and sludge-water separation units. Sensor monitoring data in the physical treatment units includes surface loading rate of the primary sedimentation tank and sludge interface height; sensor monitoring data in the biological treatment units includes ORP (oxidation-reduction potential) and VFA (volatile fatty acids) in the anaerobic tank. Linear chromatograph; DO concentration, sludge concentration (ultrasonic probe), aeration pipe pressure, transmembrane pressure, and membrane flux in the aerobic tank; sensor monitoring data for the chemical treatment unit includes coagulant dosage, floc particle size, UV transmittance, residual chlorine concentration, and contact time in the coagulation / flocculation tank; sensor monitoring data for the sludge-water separation unit includes sludge settling ratio and sludge level in the secondary sedimentation tank; the above sensor monitoring data are all conventional data in wastewater treatment plants. It should be understood that sensor monitoring data from treatment units in the technical field can also be used to establish the prediction model in this embodiment of the invention, which will not be listed here; effluent index data includes COD, ammonia nitrogen, and total phosphorus;
[0049] The aforementioned attention network layers include a COD attention network layer, an ammonia nitrogen attention network layer, and a total phosphorus attention network layer, with each attention network layer separated (e.g., Figure 1 The example only shows two attention network layers (task A and task B);
[0050] Preferably, the data obtained above is preprocessed, including ① missing value processing: using sliding window interpolation to fill in missing data; ② normalization processing: Z-score standardization to eliminate dimensional differences and enhance model convergence; ③ noise filtering: removing abnormal fluctuations based on wavelet transform; the preprocessed data is then used for subsequent model training.
[0051] S2. Based on wastewater treatment data, establish an effluent prediction model; the inputs of the effluent prediction model are influent water quality data, process data, environmental data, and wastewater treatment unit data, and the output is the predicted effluent index data;
[0052] like Figure 1 As shown, the water discharge prediction model consists of a data sharing layer, an attention network layer, and an output layer connected in sequence. Figure 1The TCN network is the shared network layer. TCN Layer 1-4 refers to multiple data layers in the shared network layer. Task A and B refer to the attention network layers of each task (i.e., the metric). The attention network layer structure and connection with the shared network of other tasks are the same as those of Task A and B in the figure. The prediction layer in the figure refers to the output layer.
[0053] The aforementioned data sharing layer is used to describe the temporal dependencies between various inputs in the water discharge prediction model. The attention network layer is used to extract the inputs related to each water discharge index from the data sharing layer. The output layer is used to fit the corresponding water discharge index data based on the water discharge index in the multi-task attention layer and the extracted inputs.
[0054] Methods for establishing effluent prediction models based on wastewater treatment data include:
[0055] 1) Building the model structure:
[0056] A data sharing layer is established based on a long short-term memory network or a gated recurrent unit; all inputs in the data sharing layer are arranged in chronological order.
[0057] In this embodiment of the invention, taking a gated loop unit as an example, the structure of each level is briefly described as follows:
[0058] Input to the data sharing layer: time step × feature dimension, for example, 24 hours × 50 dimensions (number of input data); Output of the data sharing layer: global feature vector; The global feature vector is specifically a vector that captures the common patterns of process parameters and water quality changes (1 × hidden layer dimension);
[0059] An attention network layer is established using a multi-head attention mechanism. The input of the attention network layer is a global feature vector, and the output is a feature vector related to the indicator. The data of the data sharing layer is decomposed into multiple subspaces, and the attention weight of each subspace is determined. The input related to each water output indicator is captured according to the attention weight of each subspace. The subspace refers to the feature association between at least two input data.
[0060] For example, the COD attention network layer focuses on features related to organic matter degradation, such as aeration volume and sludge age, through a soft attention mechanism; the ammonia nitrogen attention network layer focuses on parameters related to nitrification, such as dissolved oxygen, pH, and temperature; and the total phosphorus attention network layer strengthens the weight of features such as chemical phosphorus removal agent dosage and mixing intensity.
[0061] The output layer adopts a fully connected layer; in the fully connected layer, the output node of the fully connected layer is determined according to the number of effluent indicators; and the input of the output layer is the dedicated feature vector of each indicator, and the output is the predicted value of the effluent indicator (such as COD, ammonia nitrogen, and total phosphorus concentration).
[0062] Based on the above, the model algorithm specifically includes the following steps 1-1 to 1-6:
[0063] 1-1. Initialization GRU settings for the data sharing layer: The calculation of the hidden layer dimension H (e.g., 128) satisfies the formula:
[0064] h t =GRU(x t ,h t-1 )
[0065] In the formula, x t For the input at time t (influent water quality data, process data, environmental data, wastewater treatment unit data, and effluent index data at time t), h t Let h be the hidden state input at time t. t-1 The hidden state is the input at time t-1;
[0066] 1-2. Timing processing of the data sharing layer:
[0067] The input sequence is processed sequentially according to time steps, resulting in the final hidden state h. T Includes global timing information; uses a bidirectional GRU to merge the forward and reverse final states to satisfy the formula:
[0068] h g =concat(h f T h b T )
[0069] In the formula, h g h is the global feature vector. f T For the positive final state, h b T This is the reverse final state;
[0070] Understandably, the above content performed state verification on the time series data of the input (influent water quality data, process data, environmental data, wastewater treatment unit data, and effluent index data) to obtain the global feature vector of the input data. Based on the global feature vector, a multi-head attention mechanism can be calculated to obtain the input feature vector related to the model output index.
[0071] 1-3. Multi-head attention decomposition in attention network layers:
[0072] Let the number of attention heads num_heads = 3 (corresponding to COD, ammonia nitrogen, and total phosphorus), and the dimension d of each head. k =H / num_heads;
[0073] h g Projecting onto the query (Q), key (K), and value (V) space: as shown in the following equation,
[0074] Qi, Ki, Vi = Linear(h g ), i∈[1,num_heads]
[0075] 1-4. Subspace Attention Calculation:
[0076] First, compute the scaled dot product attention for the head of each attention:
[0077] Then merge the multihead outputs (Multihead(h)) g = concat(head1,...,head) numheads W o W o To output the projection matrix;
[0078] Then, specific attention enhancements were applied to each num_heads indicator: for COD, high weights were assigned to features such as aeration volume and sludge age (through learnable attention masks); for ammonia nitrogen, the correlation between dissolved oxygen, pH, and temperature was emphasized; for total phosphorus, the weights of chemical phosphorus removal agent dosage and mixing intensity were strengthened.
[0079] 1-5. Soft attention processing, the input of which is the multi-head attention output (1×(num_heads×d)). k The output is a feature vector specific to the indicator (such as COD features); taking COD as an example, a learnable weight matrix W is used. COD Calculate attention score: α COD =softmax(W COD ·h multihead +b COD Features with higher scores (such as aeration volume) are given greater weight; then, a weighted aggregation is performed to generate a COD-related feature vector: z COD =∑α COD,i ·h multihead,i ;
[0080] The specific method of the soft attention mechanism mentioned above is as follows: the input of a single module is composed of shared features u. (j) Task-specific features of the previous module They are cascaded together; the modular structure consists of three batch-normalized convolutional layers, denoted as f.(j) , Where f (j) It is responsible for matching the form between the input and output of the module. Responsible for learning the soft attention mask a i (j) The process is shown in Equation 3-1, and then combined with the shared feature p (j) Extracting the output of this module by element-wise multiplication The process is shown in Equation 3-2, where ⊙ represents element-wise multiplication:
[0081]
[0082] The structure of this model allows it to learn universal features in a shared network and automatically extract corresponding features in an attention network;
[0083] 1-6. The output layer processing steps include first inputting the specific feature vectors for each index (such as z). COD (Specific vector features for the three indicators), using an independent fully connected layer for each indicator, i.e. Then use linear activation for the output;
[0084] In this embodiment of the invention, a multi-task attention network and a gated recurrent unit network are used to predict the effluent water quality indicators of multiple aeration tanks. The basic structure of the model is as follows: Figure 1 As shown, it is divided into a shared network part (data sharing layer) shared between prediction tasks and an attention network part (attention network layer) specific to each task. The attention network part is connected to the final output layer. The number of attention networks depends on the number of metrics that the model needs to predict. For example, in the embodiment of this invention, there are three.
[0085] For example, the above model can be constructed using PyTorch and / or TensorFlow. For instance, when using PyTorch, the model implementation includes: a data sharing layer: PyTorch's GRU module, specifically nn.GRU, which can directly process time-series inputs (such as 24×50-dimensional data) and output global feature vectors; a multi-head attention network layer: using the nn.MultiheadAttention module to implement subspace decomposition, supporting custom attention head count and feature dimensions, adapting to feature association modeling of indicators such as COD, ammonia nitrogen, and total phosphorus; and a fully connected layer: using the nn.Linear module to dynamically configure the output layer according to the number of effluent indicators (such as COD, ammonia nitrogen, and total phosphorus, a total of 3 nodes).
[0086] 2) Training the model:
[0087] S2-1. Divide the wastewater treatment data into a training set and a test set according to an 8:2 ratio;
[0088] S2-2. Input the training set data into the water discharge prediction model, which passes through the data sharing layer, attention network layer and output layer in sequence to obtain the predicted water discharge index data. Calculate the error between the predicted value and the true value according to the loss function. Propagate the error from the output layer forward layer by layer through the backpropagation algorithm and calculate the gradient of the parameters of each layer. The loss function is the mean squared error and the mean absolute error.
[0089] S2-3. Use the Adam optimizer to update the model parameters based on the calculated gradients, and repeat the above forward propagation, back propagation and parameter update process until the preset maximum number of iterations is reached.
[0090] S2-4. Then, the water discharge prediction model is evaluated using test set data. Commonly used evaluation indicators include root mean square error and coefficient of determination. The smaller the root mean square error, the smaller the prediction error of the model; the closer the coefficient of determination is to 1, the better the model fits the data.
[0091] S2-5. The hyperparameters of the water discharge prediction model are tuned using the Bayesian optimization method. After finding the hyperparameter combination that optimizes the model performance, the model training is complete.
[0092] S3. Based on the inputs of the water discharge prediction model and the actual water discharge prediction model, predict the water discharge index data for future times.
[0093] In this embodiment of the invention, taking the aeration tank of a wastewater treatment plant in southern China as an example, a multi-task learning effluent prediction model established based on the plant's historical data can simultaneously and accurately predict the effluent TN, effluent COD, effluent TP, and effluent NH4 by inputting influent parameters and aeration tank process parameters. + -N. In the experiment, the model achieved a linearity factor above 0.90 for each effluent indicator, demonstrating good predictive performance. The model's predictive performance is as follows: Figure 2 As shown;
[0094] Compared to single-task models trained separately for each task, this multi-task learning-based water emission prediction model achieves a certain degree of improvement in each evaluation metric for each task, with the improvement being as follows: Figure 3 As shown, the performance has been significantly improved compared to traditional methods;
[0095] In summary, the embodiments of the present invention establish a multi-task learning effluent prediction model based on a multi-task attention network to simultaneously predict multiple effluent water quality indicators. This model can utilize the correlation between pollutant removal sub-processes to improve the prediction accuracy of each effluent water quality indicator.
[0096] Compared to traditional methods, this model not only improves predictive capabilities but also saves on storage and computational costs, laying a solid foundation for intelligent management and control of wastewater treatment. The shared network can learn shared features among various tasks, placing the changes in each indicator within a unified domain. This allows the model to learn the correlations between indicators, i.e., the correlations between pollutant removal sub-processes, thereby more comprehensively assessing the influencing factors of each indicator task and improving the model's prediction accuracy for each effluent water quality indicator. It solves the problem of traditional mechanistic models requiring complex calibration to maintain accuracy; and avoids the traditional approach of building a separate model for each indicator, instead establishing a single multi-task model to predict multiple output indicators simultaneously, reducing computational and storage costs.
[0097] Example 2: This example differs from Example 1 in that a temporal attention layer is set between the attention network layer and the output layer to enhance the influence of key time points. The temporal attention layer is constructed using a deep learning framework.
[0098] The deep learning framework is either PyTorch or TensorFlow.
[0099] The processing steps of the time attention layer include:
[0100] Obtain the hidden state of each time step in the data sharing layer and calculate the attention weight of each time step. The hidden state of the time step refers to the change of the state of the data in the data sharing layer as arranged over time.
[0101] The hidden state is weighted according to the attention weight at each time step to generate a context vector that can focus on key time points;
[0102] The context vector is connected to the output layer. The output layer fits the corresponding water discharge index data based on the water discharge index in the multi-task attention layer, the extracted input, and the context vector.
[0103] Example 3: The effluent prediction model is used to predict the effluent indicators of the biochemical treatment unit in the wastewater treatment plant. The input of the effluent prediction model includes influent water quality data, process data, environmental data, and wastewater treatment unit data, all of which are data from the biochemical treatment unit / upstream data. The output is the effluent indicator data of the biochemical treatment unit.
Claims
1. A wastewater treatment plant effluent prediction method based on multi-task learning, characterized in that, Includes the following steps: Acquire wastewater treatment data; wastewater treatment data includes influent water quality data, process data, environmental data, wastewater treatment unit data, and effluent index data; Based on wastewater treatment data, an effluent prediction model is established; the inputs of the effluent prediction model are influent water quality data, process data, environmental data, and wastewater treatment unit data, and the output is the predicted effluent index data. The water discharge prediction model includes a data sharing layer, multiple attention network layers, and multiple output layers connected in sequence. The data sharing layer is used to describe the temporal dependencies between various inputs in the water discharge prediction model. The attention network layers are used to extract inputs related to each water discharge index from the data sharing layer or the data sharing layer and other attention network layers. The output layers are used to fit the corresponding water discharge index data based on the water discharge index in the multi-task attention layer and the extracted inputs. Based on the inputs of the water discharge prediction model and the actual water discharge prediction model, the water discharge index data for future time are predicted.
2. The method as described in claim 1, characterized in that, The influent water quality data includes influent ammonia nitrogen, total phosphorus, pH, COD, and SS; process data includes aeration volume, sludge concentration, return ratio, and dissolved oxygen concentration; environmental data includes plant ambient temperature and plant ambient humidity; wastewater treatment unit data includes monitoring data from sensors in each wastewater treatment unit; and effluent index data includes COD, ammonia nitrogen, and total phosphorus.
3. The method as described in claim 1, characterized in that, The attention network layer includes three types: COD attention network layer, ammonia nitrogen attention network layer, and total phosphorus attention network layer.
4. The method as described in claim 1, characterized in that, Methods for establishing effluent prediction models based on wastewater treatment data include: S1. Building the model structure: A data sharing layer is established based on a long short-term memory network or a gated recurrent unit; all inputs in the data sharing layer are arranged in temporal order; an attention network layer is established using a multi-head attention mechanism; and the output layer is a fully connected layer. S2, Training the Model: The wastewater treatment data was divided into a training set and a test set in an 8:2 ratio. The training data is input into the water discharge prediction model, passing through the data sharing layer, attention network layer, and output layer in sequence to obtain the predicted water discharge index data. The error between the predicted value and the true value is calculated according to the loss function. The error is propagated from the output layer to the front layer by layer through the backpropagation algorithm, and the gradient of the parameters of each layer is calculated. The loss function includes mean squared error and mean absolute error. The Adam optimizer is used to update the model parameters based on the calculated gradients. The process of forward propagation, back propagation and parameter update is repeated until the preset maximum number of iterations is reached. Then, the water discharge prediction model is evaluated using test set data. Commonly used evaluation metrics include root mean square error and coefficient of determination. The smaller the root mean square error, the smaller the prediction error of the model; the closer the coefficient of determination is to 1, the better the model fits the data. The hyperparameters of the water discharge prediction model were tuned using Bayesian optimization methods. Once the optimal combination of hyperparameters was found, the model training was completed.
5. The method as described in claim 1, characterized in that, The attention network layer decomposes the data from the data sharing layer into multiple subspaces, determines the attention weight of each subspace, and captures the inputs related to each water output indicator based on the attention weight of each subspace; wherein, the subspace refers to the feature association between at least two input data.
6. The method as described in claim 1, characterized in that, A temporal attention layer is set between the attention network layer and the output layer to enhance the influence of key time points. The temporal attention layer is constructed using a deep learning framework.
7. The method as described in claim 6, characterized in that, The deep learning framework is either PyTorch or TensorFlow.
8. The method as described in claim 7, characterized in that, The processing steps of the time attention layer include: Obtain the hidden state of each time step in the data sharing layer and calculate the attention weight of each time step. The hidden state of the time step refers to the change of the state of the data in the data sharing layer as arranged over time. The hidden state is weighted according to the attention weight at each time step to generate a context vector that can focus on key time points; The context vector is connected to the output layer. The output layer fits the corresponding water discharge index data based on the water discharge index in the multi-task attention layer, the extracted input, and the context vector.
9. The application of the method as described in any one of claims 1 to 7, characterized in that, The effluent prediction model is used to predict the effluent indicators of the biochemical treatment unit in the wastewater treatment plant. The inputs of the effluent prediction model are the influent water quality data, process data, environmental data, and wastewater treatment unit data, which are all data of the biochemical treatment unit / upstream data. The output is the effluent indicator data of the biochemical treatment unit.
Citation Information
Cited By
Real-time monitoring and closed-loop regulation and control method and device for construction quality of asphalt pavement
CN121351037A
Asphalt pavement construction quality real-time monitoring and closed-loop control method and device
CN121351037B
Sewage purification process intelligent optimization method and system using deep learning
CN121615887A