A method and system for guiding carbon replacement in biological activated carbon tanks of water supply plants based on machine learning.
Patent Information
- Application Number
- CN202510515760.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-04-23
AI Technical Summary
[0015]根据本发明所涉及的一种基于机器学习指导给水厂生物活性炭池换炭的方法及系统,能够科学、准确地预测生物活性炭的性能,确定换炭时间,避免了传统方法的滞后性和不确定性。通过机器学习算法,可以充分利用大量的历史数据,提高预测的准确性和可靠性。有助于给水厂合理安排换炭计划,降低运行成本,精准调控生物活性炭池。
Smart Images

Figure CN120622591B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water treatment technology, specifically to a method and system for guiding carbon replacement in a biological activated carbon tank of a water treatment plant based on machine learning. Background Technology
[0002] In the field of water treatment, biological activated carbon is widely used to remove organic matter, ammonia nitrogen, and trace pollutants from water. Through multiple mechanisms such as physical adsorption, chemical adsorption, and biodegradation, biological activated carbon can effectively improve water quality and ensure the safety of drinking water for residents. However, with continuous use, the adsorption performance of biological activated carbon gradually declines, requiring the replacement of the activated carbon particles in the carbon bath. Currently, water treatment plants determine the replacement time of biological activated carbon mainly based on experience, regular testing of water quality indicators, or the failure of a single performance indicator of the activated carbon to meet standards. Experience-based judgment is often highly subjective and uncertain, easily leading to untimely or premature carbon replacement. While regular testing of water quality indicators can reflect changes in the performance of biological activated carbon to some extent, this method fluctuates with changes in the raw water. Relying on replacing activated carbon based on a single substandard performance indicator ignores the contribution ratio and synergistic effect among the performance components of activated carbon, making it rather one-sided.
[0003] In summary, existing methods for replacing activated carbon in biological systems have many shortcomings, and there is an urgent need for a more scientific and accurate method to guide the replacement of activated carbon in water treatment plants in order to achieve optimized operation and scientific management of water treatment plants. Summary of the Invention
[0004] This invention is made to solve the above-mentioned problems. Its purpose is to provide a method and system for carbon replacement in biological activated carbon tanks of water supply plants based on machine learning. This method can scientifically and accurately determine the carbon replacement time of biological activated carbon, achieve refined management of biological activated carbon tanks in water supply plants, and realize energy saving, consumption reduction and high-quality water supply.
[0005] This invention provides a machine learning-guided method for replacing activated carbon in a biological activated carbon tank at a water treatment plant. The method comprises the following steps: S1, collecting historical testing data from the biological activated carbon tank at the water treatment plant, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, and effective particle size, and recording the usage time of the activated carbon under the corresponding data to obtain initial performance data for the activated carbon; S2, cleaning and normalizing the initial performance data to obtain pre-processed performance data for the activated carbon; S3, performing correlation analysis on the pre-processed performance data to select relatively independent characteristic performance features. S4. Obtain relatively independent performance data of bio-activated carbon; divide the relatively independent performance data of bio-activated carbon into training set and test set, select multiple machine learning algorithms, and use the data of the training set to train the machine learning model to obtain training set model; S5. Use the training set model to predict the data of the training set, and calculate the evaluation index of the training set model on the test set. Adjust the parameters of the training set model according to the evaluation results to optimize the performance of the training set model and obtain the performance index of the optimized training set model; S6. Compare the performance indexes of different optimized training set models and select the optimal carbon replacement model.
[0006] The method for replacing carbon in a biological activated carbon tank of a water plant based on machine learning provided by this invention may also have the following feature: the cleaning process includes interpolating missing data and deleting erroneous data.
[0007] The method for replacing activated carbon in a biological activated carbon tank of a water plant based on machine learning provided by the present invention may also have the following feature: wherein step S2 includes the following sub-step: checking whether there are missing values in the dataset; if there are missing values, performing multivariate imputation by chained equations on the dataset.
[0008] The method for carbon replacement in a biological activated carbon tank of a water plant based on machine learning provided in this invention may also have the following feature: the correlation analysis adopts the Spearman rank correlation coefficient analysis method.
[0009] The method for replacing activated carbon in a biological activated carbon tank of a water plant based on machine learning provided by this invention may also have the following feature: wherein the pre-data of biological activated carbon performance is divided into a training set and a test set, and the ratio of the amount of data in the training set to the amount of data in the test set is 80:(10-30).
[0010] The method for carbon replacement in the biological activated carbon tank of a water plant based on machine learning provided in this invention may also have the following features: the machine learning algorithm includes four machine learning models: CatBoost, LightGBM, XGBoost, and Random Forest.
[0011] The method for carbon replacement in a biological activated carbon tank of a water treatment plant based on machine learning, provided by this invention, may also have the following feature: the evaluation indicators include root mean square error (RMSE) and coefficient of determination (R²). 2 The formula for calculating the root mean square error of the training set model is as follows: The unit used to measure the difference between the predicted values and the true values of the training set model is the same as the original unit of the data, where n is the number of samples and yi is the actual value. These are predicted values. The formula for calculating the coefficient of determination of the training set model is as follows: The coefficient of determination is used to measure how well the training set model fits the data. The coefficient ranges from 0 to 1, where n is the number of samples and yi is the actual value. This is a predicted value.
[0012] The method for carbon replacement in a biological activated carbon tank of a water plant based on machine learning provided by this invention may also have the following feature: the optimal carbon replacement model is the model with the smallest root mean square error and the largest coefficient of determination among the four machine learning models.
[0013] This invention also provides a system for guiding carbon replacement in a biological activated carbon tank of a water treatment plant based on machine learning. The system comprises: a data collection module that collects historical test data from the biological activated carbon tank of the water treatment plant, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, and effective particle size, and records the usage time of the biological activated carbon under the corresponding data to obtain initial performance data of the biological activated carbon; a data preprocessing module that cleans and normalizes the initial performance data of the biological activated carbon to obtain preprocessed performance data of the biological activated carbon; and a feature selection module that performs correlation analysis on the preprocessed performance data of the biological activated carbon, selects relatively independent feature performances, and obtains the biological activated carbon performance data. The system comprises the following modules: a bio-activated carbon performance pre-independent data set; a machine learning model construction module, which divides the pre-independent data into a training set and a test set, selects multiple machine learning algorithms, and trains the machine learning model using the training set data to obtain a training set model; a model evaluation and optimization module, which uses the training set model to predict the data in the training set, calculates the evaluation index of the training set model on the test set, adjusts the parameters of the training set model based on the evaluation results, optimizes the performance of the training set model, and obtains the performance index of the optimized training set model; and a carbon replacement model determination module, which compares the performance indices of different optimized training set models and selects the optimal carbon replacement model.
[0014] The role and effect of invention
[0015] The present invention relates to a method and system for guiding carbon replacement in a biological activated carbon tank of a water treatment plant based on machine learning. This method can scientifically and accurately predict the performance of biological activated carbon and determine the carbon replacement time, avoiding the lag and uncertainty of traditional methods. Through machine learning algorithms, a large amount of historical data can be fully utilized to improve the accuracy and reliability of predictions. This helps water treatment plants to rationally plan carbon replacement schedules, reduce operating costs, and precisely control the biological activated carbon tank. Attached Figure Description
[0016] Figure 1 This is an architectural diagram of the carbon exchange system for the biological activated carbon tank in a water supply plant, as described in an embodiment of the present invention.
[0017] Figure 2 The Spearman rank correlation coefficients of the physicochemical properties of bio-activated carbon over the years in the embodiments of the present invention; and
[0018] Figure 3 This is a comparison chart of actual and predicted values of different machine learning models in embodiments of the present invention, along with a comparison of evaluation metrics. Detailed Implementation
[0019] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0020] To make the technical means, creative features, objectives and effects of this invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the method and system of carbon replacement in a biological activated carbon tank of a water supply plant guided by machine learning.
[0021] Figure 1 This is an architectural diagram of the carbon exchange system for the biological activated carbon tank in a water supply plant, as described in an embodiment of the present invention.
[0022] like Figure 1 As shown in the figure, the system for guiding carbon replacement in the biological activated carbon tank of the water supply plant based on machine learning in this embodiment specifically includes: a data collection module, a data preprocessing module, a feature selection module, a machine learning model construction module, a model evaluation and optimization module, and a carbon replacement model determination module.
[0023] The data collection module collects historical test data of the biological activated carbon tank in the water supply plant, including iodine adsorption value, methylene blue adsorption value, strength, uniformity coefficient, and effective particle size, and records the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon.
[0024] This embodiment collected historical physicochemical performance data of activated carbon pools from three advanced water treatment plants with different water sources in Shanghai, including performance characteristics such as iodine adsorption value, methylene blue adsorption value, strength, uniformity coefficient, and effective particle size. The test data were all obtained from carbon particles captured on-site from the biological activated carbon pools, based on the "Test Methods for Coal-based Granular Activated Carbon" (GB / T 7702-1997, GB / T7702-2008), and the usage time of the biological activated carbon under the corresponding data was recorded.
[0025] The data preprocessing module cleans and normalizes the initial performance data of the bio-activated carbon to obtain preprocessed performance data of the bio-activated carbon.
[0026] Data cleaning includes imputing missing data and removing erroneous data. Specifically, it involves checking the dataset for missing values. If missing values are found, multivariate imputation by chained equations is performed. This can be implemented using the following code in R:
[0027]
[0028] The `is.na(data)` function checks each element in the `data` object to determine if it contains a missing value. The `mice` function is an important function provided by the `mice` package for handling missing data; `data` is the original data object to be imputed; `m=5` indicates 5 imputations will be performed; `maxit=50` limits the maximum number of iterations to 50; `method="pmm"` indicates that the "Predictive Mean Matching" (PMM) method is selected. The basic principle is to predict missing values based on the relationship between other variables, and then select appropriate values from observations that are close to the predicted values to fill the missing values.
[0029] The feature selection module performs correlation analysis on the pre-processed data of the bio-activated carbon performance, selects relatively independent feature performances, and obtains relatively independent pre-processed data of bio-activated carbon performance.
[0030] Figure 2 , is the Spearman rank correlation coefficient of the physicochemical properties of bio-activated carbon over the years in the embodiments of the present invention.
[0031] Correlation analysis was performed using Spearman's rank correlation coefficient method, such as... Figure 2 As shown, in the Spearman rank correlation coefficients of the pre-processed data of bio-activated carbon performance, all features except iodine adsorption values and methylene blue adsorption values showed weak correlations (Spearman rank correlation coefficients < 0.7). Furthermore, the Spearman rank correlation coefficient between iodine adsorption values and methylene blue adsorption values was 0.85, less than 0.9, indicating a weak correlation. This suggests that the feature dimensions of the two datasets are relatively independent, without excessive linear or monotonic relationships. The model trained in this way can make better predictions when faced with new data.
[0032] The machine learning model building module divides the relatively independent pre-data of the bio-activated carbon performance into a training set and a test set, selects multiple machine learning algorithms, and uses the data in the training set to train the machine learning model to obtain the training set model.
[0033] Specifically, the pre-independent performance data of bio-activated carbon was divided into training and test sets at an 80:20 split ratio, with random seeds set to ensure reproducibility of results. Four machine learning models—CatBoost, LightGBM, XGBoost, and RandomForest—were selected and trained using data from the training set of the pre-independent performance data of bio-activated carbon.
[0034] The model evaluation and optimization module uses the training set model to predict the data in the training set and calculates the evaluation metrics of the training set model on the test set. The evaluation metrics include the root mean square error (RMSE) and the coefficient of determination (R²). 2 Based on the evaluation results, adjust the parameters of the training set model to minimize the RMSE and R0. 2 To optimize the performance of the training set model, the performance index of the optimized training set model is obtained by making the index as close to 1 as possible.
[0035] The root mean square error (RMSE) of the training set model is calculated using the following formula: RMSE = The unit used to measure the difference between the predicted values and the true values of the training set model is the same as the original unit of the data, where n is the number of samples and yi is the actual value. This is a predicted value.
[0036] The formula for calculating the determination coefficient of the training set model is as follows: The coefficient of determination is used to measure how well the training set model fits the data. The coefficient ranges from 0 to 1, where n is the number of samples and yi is the actual value. This is a predicted value.
[0037] Figure 3 This is a comparison chart of actual and predicted values of different machine learning models in embodiments of the present invention, along with a comparison of evaluation metrics.
[0038] The carbon replacement model determination module compares the performance metrics of different training set optimization models and selects the optimal carbon replacement model.
[0039] The optimal carbon replacement model is the one with the smallest root mean square error and the largest coefficient of determination among the four machine learning models, such as... Figure 3 As shown, the RMSE and R-values of different machine learning models after optimization are compared. 2 For the test data, the R-values of the four models are... 2 Test The values were 0.90 (CatBoost), 0.99 (LightGBM), 0.88 (XGBoost), and 0.93 (Random Forest), respectively; the RMSE values for the four models were... Test The values were 0.75 (CatBoost), 0.97 (LightGBM), 0.83 (XGBoost), and 0.71 (Random Forest), respectively. Comparing the coefficients of determination, the LightGBM model achieved a higher R-value on the test set. 2The test result closest to 1 indicates that this model best fits the relationship between the physicochemical properties of activated carbon and the service life of carbon particles. Comparing the root mean square error (RMSE), the RandomForest model achieves the best RMSE on the test set. Test The smallest value indicates that the average deviation between the predicted and actual values is the smallest, while the LightGBM model performs poorly.
[0040] RMSE and R 2 Different conclusions were reached when evaluating the model's fit. In this invention, the focus is on the model's ability to explain the overall data, i.e., R0. 2 The final metric for the model is selected, that is, the LightGBM model trained is selected as the final carbon replacement model.
[0041] This invention also discloses a method for guiding carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, which specifically includes the following steps:
[0042] S1. Collect historical test data of the biological activated carbon tank in the water supply plant, including iodine adsorption value, methylene blue adsorption value, strength, uniformity coefficient, and effective particle size, and record the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon.
[0043] S2, the initial performance data of the bio-activated carbon is cleaned and normalized to obtain pre-processed performance data of the bio-activated carbon.
[0044] S3, perform correlation analysis on the pre-processed performance data of the bio-activated carbon, select relatively independent characteristic performances, and obtain relatively independent pre-processed performance data of the bio-activated carbon.
[0045] S4, the relatively independent data of the bio-activated carbon performance are divided into a training set and a test set. Multiple machine learning algorithms are selected, and the machine learning model is trained using the data in the training set to obtain the training set model.
[0046] S5, use the training set model to predict the data in the training set, calculate the evaluation index of the training set model on the test set, adjust the parameters of the training set model according to the evaluation results, optimize the performance of the training set model, and obtain the performance index of the optimized training set model.
[0047] S6. Compare the performance metrics of different training set optimization models and select the optimal carbon replacement model.
[0048] The role and effect of the embodiments
[0049] The present invention relates to a method and system for guiding carbon replacement in a biological activated carbon tank of a water treatment plant based on machine learning. This method can scientifically and accurately predict the performance of biological activated carbon and determine the carbon replacement time, avoiding the lag and uncertainty of traditional methods. Through machine learning algorithms, a large amount of historical data can be fully utilized to improve the accuracy and reliability of predictions. This helps water treatment plants to rationally plan carbon replacement schedules, reduce operating costs, and precisely control the biological activated carbon tank.
[0050] This invention uses a data preprocessing module to perform preprocessing operations such as cleaning and normalization on the collected data, and to impute missing data and delete erroneous data, which can effectively improve data quality.
[0051] This invention calculates the root mean square error and coefficient of determination of the test set, adjusts the model parameters, and selects the trained LightGBM model as the final carbon-changing model by comparing the conclusions obtained from evaluating the model fitting effect using the coefficient of determination and root mean square error, which shows the best performance.
[0052] Those skilled in the art should understand that this invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to this invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for carbon replacement in a biological activated carbon tank of a water treatment plant based on machine learning, characterized in that, Specifically, the steps include the following: S1. Collect the test data of the biological activated carbon tank of the water supply plant over the years, including iodine adsorption value, methylene blue adsorption value, strength, uniformity coefficient, and effective particle size, and record the usage time of the biological activated carbon under the corresponding data to obtain the initial data of the performance of the biological activated carbon. S2, the initial performance data of the bio-activated carbon is cleaned and normalized to obtain pre-processed performance data of the bio-activated carbon. S3, perform correlation analysis on the pre-processed performance data of the bio-activated carbon, select relatively independent characteristic performance, and obtain relatively independent pre-processed performance data of bio-activated carbon. S4, the pre-relatively independent data of the performance of the bio-activated carbon are divided into a training set and a test set. Multiple machine learning algorithms are selected, and the machine learning model is trained using the data in the training set to obtain the training set model. S5, use the training set model to predict the data in the training set, calculate the evaluation index of the training set model on the test set, adjust the parameters of the training set model according to the evaluation results, optimize the performance of the training set model, and obtain the performance index of the optimized training set model. S6. Compare the performance metrics of different training set optimization models and select the optimal carbon replacement model.
2. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, as described in claim 1, is characterized in that: in, The cleaning process includes imputing missing data and deleting erroneous data.
3. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, as described in claim 2, is characterized in that: in, Step S2 includes the following sub-steps: checking whether there are missing values in the dataset; if there are missing values, performing multivariate imputation on the dataset using the Multivariate Imputation by Chained Equations method.
4. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, as described in claim 1, is characterized in that: in, The correlation analysis employed the Spearman rank correlation coefficient analysis method.
5. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, as described in claim 1, is characterized in that: in, The pre-relatively independent data of the bio-activated carbon performance were divided into a training set and a test set, with the ratio of the amount of data in the training set to the amount of data in the test set being 80:(10-30).
6. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning as described in claim 1, characterized in that: in, The machine learning algorithms include four machine learning models: CatBoost, LightGBM, XGBoost, and Random Forest.
7. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, as described in claim 6, is characterized in that: in, The evaluation metrics include root mean square error (RMSE) and coefficient of determination (R²). 2 ), The formula for calculating the root mean square error of the training set model is as follows: The unit used to measure the difference between the predicted values and the true values of the training set model is the same as the original unit of the data, where n is the number of samples and yi is the actual value. It is a predicted value. The formula for calculating the determination coefficient of the training set model is as follows: The coefficient of determination is used to measure how well the training set model fits the data. The coefficient ranges from 0 to 1, where n is the number of samples and yi is the actual value. This is a predicted value.
8. The method for carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, as described in claim 7, is characterized in that: in, The optimal carbon replacement model is the one with the smallest root mean square error and the largest coefficient of determination among the four machine learning models.
9. A system for guiding carbon replacement in a biological activated carbon tank of a water supply plant based on machine learning, characterized in that, include: The data collection module collects historical test data of the biological activated carbon tank in the water supply plant, including iodine adsorption value, methylene blue adsorption value, strength, uniformity coefficient, and effective particle size, and records the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon. The data preprocessing module cleans and normalizes the initial performance data of the bio-activated carbon to obtain preprocessed performance data of the bio-activated carbon. The feature selection module performs correlation analysis on the pre-processed performance data of the bio-activated carbon, selects relatively independent feature performances, and obtains relatively independent pre-processed performance data of the bio-activated carbon. The machine learning model building module divides the relatively independent pre-data of the bio-activated carbon performance into a training set and a test set, selects multiple machine learning algorithms, and uses the data in the training set to train the machine learning model to obtain the training set model. The model evaluation and optimization module uses the training set model to predict the data in the training set, calculates the evaluation index of the training set model on the test set, adjusts the parameters of the training set model according to the evaluation results, optimizes the performance of the training set model, and obtains the performance index of the optimized training set model. The carbon replacement model determination module compares the performance metrics of different optimized models on the training set and selects the optimal carbon replacement model.
Citation Information
Patent Citations
Operation evaluation method of biological activated carbon filter tank
CN108226400A
Chemical material adsorption performance prediction method and device based on automatic machine learning
CN112966447A