Sewage plant carbon source accurate adding method based on online water quality monitoring and automatic machine learning
By establishing a carbon source addition prediction model through online water quality monitoring and automatic machine learning, the problem of insufficient carbon source addition accuracy in wastewater treatment plants has been solved, achieving precise carbon source addition, improving denitrification efficiency and system stability, and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN CHENGJIAN UNIV
- Filing Date
- 2025-07-31
- Publication Date
- 2026-05-26
AI Technical Summary
When wastewater treatment plants improve nitrogen removal efficiency, insufficient precision in carbon source dosing control leads to limitations in the denitrification process and a decrease in nitrogen removal efficiency. Excessive dosing increases costs and affects system stability. Existing intelligent dosing control systems lack adaptive capabilities and their models lack universality.
A regression model was established using online water quality monitoring and automated machine learning. Influent and effluent water quality indicators and pollutant removal rate indicators were used as explanatory variables. Automated machine learning algorithms (such as gradient boosters) were used to predict carbon source dosage, and the prediction model for carbon source dosage was optimized.
It enables accurate prediction of carbon source dosage, reduces carbon source waste, lowers operating costs, improves denitrification efficiency and system stability, and adapts to water quality changes in different wastewater treatment plants.
Smart Images

Figure CN122091017A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment, specifically to a method for precise carbon source dosing in wastewater treatment plants based on online water quality monitoring and automated machine learning. Background Technology
[0002] Currently, wastewater treatment plants generally face the problem of insufficient precision in controlling carbon source addition during the process of improving nitrogen removal efficiency and meeting stricter emission standards. Insufficient carbon source addition will limit the denitrification process and reduce nitrogen removal efficiency; conversely, excessive addition will not only increase operating costs but may also lead to a series of operational risks such as carbon source waste and decreased system stability. For example, excessive carbon source (such as sodium acetate) may inhibit the nitrification process, cause nitrite accumulation, and even cause sludge bulking and loose floc structure, thereby affecting effluent stability.
[0003] Currently, some wastewater treatment plants use automatic dosing control systems based on programmable logic controllers (PLCs) to add carbon sources according to preset rules or empirical setpoints. While these systems achieve mechanization and continuous dosing, they are essentially still based on static rules or fixed thresholds, lacking the ability to adapt to fluctuations in influent water quality and the dynamic characteristics of the biochemical process, easily leading to insufficient or excessive carbon source dosing. Therefore, intelligent dosing control technology has emerged. For example, some studies have used XGBoost to predict total nitrogen removal rate and microbial growth, thereby optimizing carbon source dosing and predicting excess sludge volume, and have built integrated models to optimize carbon source dosing and predict excess sludge volume, demonstrating the significant advantages of machine learning in carbon source dosing optimization. However, limited computing power restricts data processing speed, resulting in expensive model training time costs. Furthermore, existing models focus on building intelligent carbon source dosing models for single wastewater treatment plants and their specific processes. These models are limited by the specificity of wastewater treatment plants and lack generalizability and universality.
[0004] Automated Machine Learning (AutoML) completes feature selection, model construction, and hyperparameter optimization without human intervention. It supports the automatic integration of multiple modeling algorithms, effectively improving the model's adaptability and generalization ability to different wastewater treatment plant data structures, and solving the problems of poor transferability and weak generalization of traditional modeling methods. On the other hand, compared with existing carbon source addition models that only use multiple online monitored influent and effluent water quality indicators (such as COD, NH3-N, ORP, etc.) or real-time monitoring and feedback control for prediction, this invention comprehensively considers the pollutant removal effect, more closely reflects the changing patterns of carbon source demand during actual operation, and thus significantly improves the accuracy of prediction and the scientific nature of decision-making. Currently, there is no model method that uses automated machine learning to predict the optimal carbon source addition based on online water quality monitoring indicators of wastewater treatment plants (including influent and effluent water quality indicators and pollutant removal rate indicators). Summary of the Invention
[0005] Therefore, the present invention aims to provide a method for precise carbon source addition based on online water quality monitoring and artificial intelligence algorithms.
[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms, comprising the following steps: (1) The influent and effluent water quality indicators and pollutant removal rate indicators are used as explanatory variables, including influent and effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, FLOW; 5 removal rate indicators: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, TP removal rate; and carbon source dosage is used as the target variable. An automatic machine learning method is used to establish a regression model as a prediction model for carbon source dosage. (2) Set the effluent quality to the local effluent quality standard limit and use the carbon source addition prediction model to predict the optimal carbon source addition for the wastewater treatment plant.
[0007] Furthermore, in step (1), the regression model is established using the h2o package in R software.
[0008] Further, in step (2), the method for predicting the carbon source dosage includes: using influent and effluent water quality indicators and pollutant removal rate indicators as explanatory variables, including influent and effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, FLOW; 5 removal rate indicators: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, TP removal rate; using the carbon source dosage as the target variable, and using an automatic machine learning method to establish a regression model as the prediction model for the carbon source dosage, the obtained target variable value is the prediction result of the carbon source dosage to be added.
[0009] Further, in step (2), the method for predicting the optimal carbon source dosage includes: replacing the effluent quality with the local effluent quality standard limit, and using the carbon source dosage prediction model to predict the optimal carbon source dosage for the wastewater treatment plant.
[0010] Furthermore, in step (1), the carbon source precision addition method based on online water quality monitoring and artificial intelligence algorithm further includes: dividing the influent and effluent water quality index datasets into training set and test set, using the training set to establish the prediction model, and using the test set to verify the prediction ability of the prediction model. Preferably, 80% of the dataset is used as the training set and 20% of the dataset is used as the test set.
[0011] Furthermore, using the fitting coefficient R...2 Measuring the predictive power of the prediction model: When R 2 When the value is ≤0.3, the predicted value fits the observed value poorly, and the prediction model has poor predictive ability. When 0.3 < R 2 When the value is ≤0.4, the predicted value fits the observed value poorly, and the prediction ability of the prediction model is weak. When 0.4 < R 2 When the value is ≤0.6, the predicted value fits the observed value moderately, and the prediction ability of the prediction model is moderate. When 0.6 < R 2 When the value is ≤1.0, the predicted value fits the observed value well, indicating strong predictive ability of the prediction model. The technical solution of this invention has the following advantages: This invention provides a method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms. It establishes a predictive model through automatic machine learning using water quality monitoring indicators, thereby predicting the amount of carbon source to be added to wastewater treatment plants. Specifically, it uses 21 online water quality monitoring indicators as explanatory variables, including 8 influent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW; and 8 effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW. Five removal rate indicators are calculated based on the monitoring indicators: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, and TP removal rate. The carbon source dosing amount is used as the target variable. Verification shows that the predictive model established using the method provided by this invention has a strong fitting effect and can accurately predict the carbon source dosing amount using only water quality monitoring results. On the one hand, it can solve the problem that insufficient or excessive carbon source addition will lead to poor treatment effect, and excessive addition will cause carbon source waste, increase the cost of sewage treatment plants and carbon emissions; on the other hand, it can analyze the response relationship between water quality monitoring indicators and carbon source addition based on interpretable machine learning methods. Attached Figure Description
[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 This describes the model training and testing results of using water quality monitoring indicators to measure carbon source dosage in this embodiment of the invention. Figure 2This is a partial dependency diagram between the target variable and each explanatory variable in the prediction model for carbon source addition. Detailed Implementation
[0014] The following embodiments are provided to better understand the present invention and are not limited to the preferred embodiments described. They do not constitute a limitation on the content and scope of protection of the present invention. Any product that is the same as or similar to the present invention, derived by any person under the guidance of the present invention or by combining the features of the present invention with other prior art, falls within the protection scope of the present invention.
[0015] The carbon source addition control method involved in the embodiments of this application is mainly based on water quality data input and artificial intelligence algorithm calculation. For those embodiments that do not explicitly list specific software training parameters, model structures or algorithm hyperparameter configurations, conventional settings can be adopted based on publicly available literature or engineering experience in this field.
[0016] This embodiment provides a method for establishing a predictive model for carbon source addition, the specific steps of which are as follows: (1) Water quality indicators monitored at wastewater treatment plants A wastewater treatment plant in Tianjin, using sodium acetate as the carbon source, was selected as the research subject. Twenty-one water quality monitoring indicators were used as explanatory variables, including eight influent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW; and eight effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW. Five removal rate indicators were calculated based on the monitoring indicators: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, and TP removal rate. A total of 366 samples were collected. The dataset was complete, without outliers, and the data were accurate and reasonable. The full names and abbreviations of each monitoring indicator are shown in Table 1.
[0017] Table 1. Full names and abbreviations of monitoring indicators
[0018] (2) Obtaining the best machine learning algorithm This study utilizes the AutoML algorithm from the h2o package in R to conduct research and analysis, specifically employing the h2o.automl function to automatically select the best machine learning algorithm. h2o AutoML can automatically choose the optimal algorithm, including deep learning, random forest (RF), generalized linear model (GLM), gradient boosting machine (GBM), and XGBoost. The best algorithm selected by h2o AutoML is Gradient Boosting Machine (GBM).
[0019] (3) Use the gradient boosting machine algorithm to establish a prediction model The above monitoring data comprised a total of 366 samples. The aforementioned 21 water quality monitoring indicators were used as explanatory variables, including 8 influent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW; and 8 effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW. Five removal rate indicators were calculated based on the monitoring indicators: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, and TP removal rate. The carbon source dosage was used as the target variable, and a prediction model was established using the gradient booster algorithm.
[0020] After comparing dataset split ratios (the ratio of training set to test set samples) of 0.7, 0.8, 0.85, and 0.9, a ratio of 0.8 was determined to provide the best prediction performance. A loop was used with 30 random seeds, running 30 times. Each learned model developed based on the training set was then used on the test set to verify the model's generalization ability.
[0021] The mean squared error of prediction (MSE) and the fitting coefficient (R²) based on the training set 2 The performance of the model is evaluated using a test set. MSE is generally used to compare prediction results for the same metric. For predictions of the same metric, generally, the smaller the MSE, the better the R². 2 The larger the value, the more appropriate the R value is. 2 R is used as the primary evaluation parameter. 2 <0.3, 0.3≤R 2 <0.4, 0.4≤R 2 <0.6, 0.6≤R 2 ≤1.0 represents poor, weak, moderate, and strong fit between the predicted and measured values, respectively.
[0022] (5) Model training performance based on training and test sets The fitting results of the predicted and observed values of carbon source injection based on the training and test sets using 24 monitoring indicators (Table 1) are shown below. Figure 1 R using the training and test sets 2 To evaluate the model's performance (Table 2), the training set R of the optimal model is used. 2 ≥ 0.9, test set R 2 A value ≥ 0.8 indicates a strong fit.
[0023] Table 2 Model performance evaluation results
[0024] The prediction model established in Example 1 is used to predict the amount of carbon source added, and the method is as follows: The effluent quality was set to the Tianjin Municipal Standard as specified in the Integrated Wastewater Discharge Standard (DB12 / 356-2018), as shown in Table 3. The optimal carbon source prediction model established in Example 1 was then used for prediction. The carbon source dosage was ultimately calculated from the predicted carbon source dosage, resulting in the optimized carbon source dosage. Monthly data are shown in Table 4. It can be seen that the model exhibits significant energy-saving and emission-reduction effects, saving 13.5%-26.9% of carbon source per month, and a comprehensive annual carbon source saving of 18.5%. The optimization of carbon source dosage is more pronounced during the low-temperature season, indicating that the wastewater treatment plant is overdosing during this period.
[0025] Table 3 Target Effluent Water Quality
[0026] Table 4 Monthly Carbon Source Savings at Wastewater Treatment Plants
[0027] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms, characterized in that, Includes the following steps: Step 1: Monitor the water quality indicators of the wastewater treatment plant, including monitoring 8 influent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW; And 8 effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW; Five removal rate indicators were calculated based on monitoring metrics: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, and TP removal rate. Step 2: Use influent and effluent water quality indicators and pollutant removal rate indicators as explanatory variables, including influent and effluent indicators: CODcr, NH3-N, SS, TN, TP, C / N, pH, and FLOW; And five removal rate indicators: CODcr removal rate, NH3-N removal rate, SS removal rate, TN removal rate, and TP removal rate; Using carbon source input as the target variable, an automatic machine learning method is used to establish a regression model as a prediction model for carbon source input. Step 3: Set the effluent quality to the local effluent standard limit, and use the carbon source addition prediction model to predict the optimal carbon source addition for the wastewater treatment plant.
2. The method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms according to claim 1, characterized in that, In step 2, a regression model is established using the h2o package in R software.
3. The method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms according to claim 2, characterized in that, In step 3, the method for predicting the carbon source dosage includes: substituting the explanatory variables from step 2 into the precision dosing model to obtain the predicted carbon source dosage.
4. The method for predicting carbon source dosage based on wastewater quality according to claim 3, characterized in that, It also includes dividing the influent and effluent water quality index dataset into a training set and a test set, using the training set to build the prediction model, and using the test set to verify the prediction ability of the prediction model.
5. The method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms according to claim 4, characterized in that, 80% of the dataset is used as the training set, and 20% of the dataset is used as the test set.
6. The method for precise carbon source dosing based on online water quality monitoring and artificial intelligence algorithms according to claim 4, characterized in that, The predictive power of the prediction model is measured by the fitting coefficient R². When R² ≤ 0.3, the predicted values fit the observed values poorly, indicating poor predictive ability of the prediction model. When 0.3 < R² ≤ 0.4, the predicted value fits the observed value poorly, and the prediction ability of the prediction model is weak. When 0.4 < R² ≤ 0.6, the predicted value fits the observed value moderately, and the prediction ability of the prediction model is moderate. When 0.6 < R² ≤ 1.0, the predicted value fits the observed value well, and the prediction model has strong predictive ability.