Method and system for guiding charcoal change of biological activated carbon pool of water supply plant based on machine learning
By using machine learning methods to process the detection data of the biological activated carbon pool of the water supply plant and select the optimal carbon replacement model, the problem of uncertainty in carbon replacement time in the existing method is solved, and scientific and accurate carbon replacement time prediction and cost optimization are achieved.
Patent Information
- Application Number
- CN202510515760.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing biological activated carbon replacement methods rely on empirical judgment or regular testing, which are subjective and uncertain, resulting in untimely or premature carbon replacement, and the inability to scientifically and accurately determine the carbon replacement time.
Using machine learning methods, by collecting and processing the detection data of the biological activated carbon pool of the water plant, correlation analysis and feature selection are carried out, multiple machine learning models are trained, model parameters are optimized, and the optimal carbon replacement model is selected to achieve scientific and accurate carbon replacement time prediction.
It improves the accuracy and reliability of carbon replacement prediction, helps water plants to rationally arrange carbon replacement plans, reduce operating costs, and achieve precise control of biological activated carbon pools.
Smart Images

Figure CN120622591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water treatment, and in particular to a method and system for guiding carbon replacement in a biological activated carbon pool of a water plant based on machine learning. Background Art
[0002] In the water treatment sector, biological activated carbon is widely used to remove organic matter, ammonia nitrogen, and trace pollutants from water. Through various mechanisms, including physical adsorption, chemical adsorption, and biodegradation, biological activated carbon can effectively improve water quality and ensure drinking water safety for residents. However, with continued use, its adsorption performance gradually declines, necessitating replacement of the carbon granules in the tank. Currently, water plants primarily rely on empirical judgment, regular water quality testing, or failure to meet individual activated carbon performance indicators to determine when to replace biological activated carbon. Empirical judgments are often highly subjective and uncertain, leading to untimely or premature carbon replacement. While regular water quality testing can reflect changes in biological activated carbon performance to a certain extent, this method fluctuates with changes in the raw water. Relying on a single substandard activated carbon performance indicator to determine carbon replacement ignores the contribution ratio and synergistic effects of various activated carbon properties and is a one-sided approach.
[0003] In summary, the existing biological activated carbon replacement method has many shortcomings, and there is an urgent need for a more scientific and accurate method to guide the replacement of biological activated carbon in water plants to achieve optimized operation and scientific management of water plants. Summary of the Invention
[0004] The present invention is carried out to solve the above-mentioned problems. Its purpose is to provide a method and system for guiding the carbon replacement of the biological activated carbon pool of a water supply plant based on machine learning, which can scientifically and accurately determine the carbon replacement time of the biological activated carbon, achieve refined management of the biological activated carbon pool of the water supply plant, and realize energy saving and consumption reduction and high-quality water supply.
[0005] The present invention provides a method for guiding the replacement of biological activated carbon pools in water supply plants based on machine learning, which has the following characteristics and specifically includes the following steps: S1, collecting the detection data of the biological activated carbon pools in the water supply plants over the years, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, and effective particle size, and recording the usage time of the biological activated carbon under the corresponding data to obtain the initial data of the biological activated carbon performance; S2, cleaning and normalizing the initial data of the biological activated carbon performance to obtain the pre-processed data of the biological activated carbon performance; S3, performing correlation analysis on the pre-processed data of the biological activated carbon performance, and selecting relatively independent characteristic performance , obtain relatively independent data on the performance of biological activated carbon; S4, divide the relatively independent data on the performance of biological activated carbon into a training set and a test set, select multiple machine learning algorithms, use the data of the training set to train the machine learning model, and obtain a training set model; S5, use the training set model to predict the data of the training set, and calculate the evaluation index of the training set model in the test set, adjust the parameters of the training set model according to the evaluation results, optimize the performance of the training set model, and obtain the performance index of the training set optimization model; S6, compare the performance indexes of different training set optimization models, and select the optimal carbon replacement model.
[0006] The method for replacing carbon in a biological activated carbon pool of a water plant based on machine learning guidance provided by the present invention may also have the following characteristics: wherein, cleaning includes interpolating missing data and deleting erroneous data.
[0007] The method for replacing carbon in a biological activated carbon pool of a water plant based on machine learning provided by the present invention may also have the following features: wherein, step S2 includes the following sub-steps: checking whether there are missing values in the data set; if there are missing values, performing multivariate interpolation on the data set using the Multivariate Imputation by Chained Equations method.
[0008] The method for replacing carbon in a biological activated carbon pool of a water plant based on machine learning guidance provided by the present invention may also have the following characteristics: wherein, the correlation analysis adopts the Spearman rank correlation coefficient analysis method.
[0009] In the method provided by the present invention for carbon replacement in the biological activated carbon pool of a water plant based on machine learning guidance, it can also have the following characteristics: wherein, relatively independent data on the performance of the biological activated carbon are divided into a training set and a test set, and the ratio of the data volume of the training set to the data volume of the test set is 80:(10~30).
[0010] The method for guiding carbon replacement in a biological activated carbon pool of a water plant based on machine learning provided by the present invention may also have the following characteristics: wherein the machine learning algorithm includes four machine learning models: CatBoost, LightGBM, XGBoost, and Random Forest.
[0011] In the method for replacing carbon in the biological activated carbon pool of a water plant based on machine learning provided by the present invention, the method may also have the following characteristics: wherein the evaluation indicators include the root mean square error (RMSE) and the coefficient of determination (R 2 ), the calculation formula of the root mean square error of the training set model is as follows: It is used to measure the difference between the predicted value of the training set model and the true value. Its unit is the same as the original unit of the data, where n is the number of samples, yi is the actual value, is the predicted value, and the calculation formula of the determination coefficient of the training set model is as follows: It is used to measure the degree of fit of the training set model to the data. The value range of the determination coefficient is between 0 and 1, where n is the number of samples and yi is the actual value. is the predicted value.
[0012] The method for carbon replacement in a biological activated carbon pool of a water plant based on machine learning guidance provided by the present invention may also have the following characteristics: wherein the optimal carbon replacement model is the model with the smallest root mean square error and the largest determination coefficient among the four machine learning models.
[0013] The present invention also provides a system for guiding the replacement of biological activated carbon pools in water supply plants based on machine learning, which has the following characteristics: a data collection module, which collects the detection data of the biological activated carbon pools in water supply plants over the years, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, effective particle size, and records the usage time of the biological activated carbon under the corresponding data to obtain the initial data of the biological activated carbon performance; a data preprocessing module, which cleans and normalizes the initial data of the biological activated carbon performance to obtain the preprocessed data of the biological activated carbon performance; a feature selection module, which performs correlation analysis on the preprocessed data of the biological activated carbon performance, selects relatively independent feature performance, and obtains the biological activated carbon performance data. Relatively independent data on the performance of biological activated carbon; a machine learning model construction module, which divides the relatively independent data on the performance of biological activated carbon into a training set and a test set, selects a variety of machine learning algorithms, and uses the data of the training set to train the machine learning model to obtain a training set model; a model evaluation and optimization module, which uses the training set model to predict the data of the training set, and calculates the evaluation index of the training set model in the test set, adjusts the parameters of the training set model according to the evaluation results, optimizes the performance of the training set model, and obtains the performance index of the training set optimization model; a carbon replacement model determination module, which compares the performance indexes of different training set optimization models and selects the optimal carbon replacement model.
[0014] Functions and effects of the invention
[0015] The present invention provides a method and system for guiding carbon replacement in water plant biological activated carbon tanks based on machine learning. This method can scientifically and accurately predict the performance of biological activated carbon and determine the replacement time, avoiding the lag and uncertainty of traditional methods. The machine learning algorithm leverages extensive historical data to improve the accuracy and reliability of predictions. This helps water plants rationally plan carbon replacements, reduce operating costs, and precisely control biological activated carbon tanks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is an architectural diagram of a carbon replacement system for a biological activated carbon pool in a water plant according to an embodiment of the present invention;
[0017] Figure 2 is the Spearman rank correlation coefficient of the physical and chemical properties of the biological activated carbon in the embodiments of the present invention over the years; and
[0018] Figure 3 It is a comparison chart of actual values and predicted values of different machine learning models and evaluation index comparison in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections, electrical connections, or mutual communication; they can refer to direct connections or indirect connections through an intermediate medium; they can refer to internal communication between two components or the interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on specific circumstances.
[0020] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments, combined with the accompanying drawings, specifically illustrate the method and system of the present invention for machine learning to guide carbon replacement in the biological activated carbon pool of a water plant.
[0021] Figure 1 It is an architectural diagram of a carbon replacement system for a biological activated carbon pool in a water plant according to an embodiment of the present invention.
[0022] like Figure 1 As shown, the system for guiding carbon replacement in the biological activated carbon pool of a water plant based on machine learning in this embodiment specifically includes: a data collection module, a data preprocessing module, a feature selection module, a machine learning model construction module, a model evaluation and optimization module, and a carbon replacement model determination module.
[0023] The data collection module collects the test data of the biological activated carbon pool in the water supply plant over the years, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, effective particle size, and records the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon.
[0024] In this example, historical physical and chemical performance data for activated carbon pools at three deep-water treatment plants in Shanghai, each with different water sources, was collected. These performance characteristics included iodine adsorption, methylene blue adsorption, strength, uniformity coefficient, and effective particle size. The test data were obtained from on-site carbon granules collected from biological activated carbon pools, according to the "Test Method for Coal-Based Granular Activated Carbon" (GB / T 7702-1997, GB / T 7702-2008). The age of the activated carbon was also recorded for the corresponding data.
[0025] The data preprocessing module cleans and normalizes the initial data on the performance of the biological activated carbon to obtain preprocessed data on the performance of the biological activated carbon.
[0026] Cleaning includes interpolating missing data and deleting erroneous data. Specifically, check whether the dataset has missing values. If so, use the Multivariate Imputation by Chained Equations method to perform multivariate imputation on the dataset. For example, this can be achieved using the following code based on R programming:
[0027]
[0028] Here, is.na(data) checks each element in the data object to determine whether it is a missing value. The mice function, provided by the mice package, is an important function for handling missing data. data is the original data object for missing value interpolation. m = 5 indicates that five interpolations will be performed. maxit = 50 limits the number of iterations to 50. method = "pmm" indicates that the Predictive Mean Matching (pmm) method is selected. The basic principle is to predict missing values based on the relationship between other variables and then select appropriate values from observations close to the predicted values to match and fill the missing values.
[0029] The feature selection module performs correlation analysis on the pre-processed data of the biological activated carbon performance, selects relatively independent feature performances, and obtains relatively independent pre-processed data of the biological activated carbon performance.
[0030] Figure 2 is the Spearman rank correlation coefficient of the physical and chemical properties of biological activated carbon in the examples of the present invention over the years.
[0031] The correlation analysis was performed using the Spearman rank correlation coefficient analysis method. Figure 2 As shown in the Spearman rank correlation coefficients for the pretreated biological activated carbon performance data, all features, except for iodine adsorption and methylene blue adsorption, were weakly correlated (Spearman rank correlation coefficient < 0.7). Furthermore, the Spearman rank correlation coefficient between iodine adsorption and methylene blue adsorption was 0.85, less than 0.9, and not in the strong correlation range. This indicates that the feature dimensions of the two data sets are relatively independent, lacking excessive linear or monotonic relationships. This allows the trained model to better predict new data.
[0032] The machine learning model construction module divides the relatively independent data on the performance of the biological activated carbon into a training set and a test set, selects multiple machine learning algorithms, and uses the data of the training set to train the machine learning model to obtain a training set model.
[0033] Specifically, the relatively independent data on activated carbon performance was split into training and test sets at an 80:20 split ratio, with a random seed set to ensure reproducible results. Four machine learning models, CatBoost, LightGBM, XGBoost, and RandomForest, were selected and trained using the training set of the relatively independent data.
[0034] The model evaluation and optimization module uses the training set model to predict the data of the training set and calculates the evaluation indicators of the training set model in the test set, including the root mean square error (RMSE) and the coefficient of determination (R 2 ), adjust the parameters of the training set model according to the evaluation results to make the RMSE as small as possible, R 2 As close to 1 as possible, the performance of the training set model is optimized, and the performance index of the training set optimization model is obtained.
[0035] The calculation formula of the root mean square error of the training set model is as follows: RMSE = It is used to measure the difference between the predicted value of the training set model and the true value. Its unit is the same as the original unit of the data, where n is the number of samples, yi is the actual value, is the predicted value.
[0036] The calculation formula of the determination coefficient of the training set model is as follows: It is used to measure the degree of fit of the training set model to the data. The value range of the determination coefficient is between 0 and 1, where n is the number of samples and yi is the actual value. is the predicted value.
[0037] Figure 3 It is a comparison chart of actual values and predicted values of different machine learning models and evaluation index comparison in an embodiment of the present invention.
[0038] The carbon replacement model determination module compares the performance indicators of different training set optimization models and selects the optimal carbon replacement model.
[0039] The optimal carbon replacement model is the model with the smallest root mean square error and the largest coefficient of determination among the four machine learning models, such as Figure 3 As shown, the RMSE and R of different machine learning models after optimization are compared. 2 For the test data, the R 2 Test The values are 0.90 (CatBoost), 0.99 (LightGBM), 0.88 (XGBoost), and 0.93 (Random Forest); the RMSE of the four models Test The values are 0.75 (CatBoost), 0.97 (LightGBM), 0.83 (XGBoost), and 0.71 (Random Forest). 2The test is closest to 1, indicating that this model has the best fitting effect in fitting the physical and chemical properties of activated carbon to the service life of carbon particles. Test The smallest, indicating that the average deviation between the predicted value and the true value is the smallest, while the LightGBM model performs poorly.
[0040] RMSE and R 2 Different conclusions were drawn when evaluating the model fitting effect. In this paper, we focus on the model's ability to explain the data as a whole, i.e., R 2 The final indicator for the model is selected, that is, the trained LightGBM model is selected as the final carbon conversion model.
[0041] The present invention also discloses a method for guiding carbon replacement in a biological activated carbon pool of a water plant based on machine learning, which specifically includes the following steps:
[0042] S1. Collect the test data of the biological activated carbon pool in the water supply plant over the years, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, and effective particle size, and record the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon.
[0043] S2, cleaning and normalizing the initial data of the biological activated carbon performance to obtain preprocessed data of the biological activated carbon performance.
[0044] S3, performing correlation analysis on the pre-processed data of the biological activated carbon performance, selecting relatively independent characteristic performances, and obtaining relatively independent pre-processed data of the biological activated carbon performance.
[0045] S4, dividing the relatively independent data on the performance of the biological activated carbon into a training set and a test set, selecting a plurality of machine learning algorithms, and using the data of the training set to train the machine learning model to obtain a training set model.
[0046] S5, using the training set model to predict the data of the training set, and calculating the evaluation index of the training set model in the test set, adjusting the parameters of the training set model according to the evaluation results, optimizing the performance of the training set model, and obtaining the performance index of the training set optimization model.
[0047] S6, comparing the performance indicators of different training set optimization models and selecting the optimal carbon replacement model.
[0048] Functions and Effects of the Embodiments
[0049] The present invention provides a method and system for guiding carbon replacement in water plant biological activated carbon tanks based on machine learning. This method can scientifically and accurately predict the performance of biological activated carbon and determine the replacement time, avoiding the lag and uncertainty of traditional methods. The machine learning algorithm leverages extensive historical data to improve the accuracy and reliability of predictions. This helps water plants rationally plan carbon replacements, reduce operating costs, and precisely control biological activated carbon tanks.
[0050] The present invention uses a data preprocessing module to perform preprocessing operations such as cleaning and normalization on the collected data, interpolate missing data, and delete erroneous data, which can effectively improve data quality.
[0051] The present invention calculates the root mean square error and determination coefficient of the test set, adjusts the model parameters, and obtains conclusions by comparing the determination coefficient and root mean square error when evaluating the model fitting effect. The trained LightGBM model is selected as the final carbon replacement model, which has the best effect.
[0052] Those skilled in the art will appreciate that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for guiding carbon replacement in a biological activated carbon pool of a water plant based on machine learning, characterized in that: The specific steps include: S1, collect the test data of the biological activated carbon pool in the water supply plant over the years, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, effective particle size, and record the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon; S2, cleaning and normalizing the initial data of the biological activated carbon performance to obtain pre-processed data of the biological activated carbon performance; S3, performing correlation analysis on the pre-processed data of the biological activated carbon performance, selecting relatively independent characteristic performances, and obtaining relatively independent pre-processed data of the biological activated carbon performance; S4, dividing the relatively independent data on the performance of the biological activated carbon into a training set and a test set, selecting a plurality of machine learning algorithms, and using the data of the training set to train the machine learning model to obtain a training set model; S5, using the training set model to predict the data of the training set, and calculating the evaluation index of the training set model in the test set, adjusting the parameters of the training set model according to the evaluation results, optimizing the performance of the training set model, and obtaining the performance index of the training set optimization model; S6, comparing the performance indicators of different training set optimization models and selecting the optimal carbon replacement model.
2. The method for replacing carbon in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 1, characterized in that: in, The cleaning includes interpolating missing data and deleting erroneous data.
3. The method for carbon replacement in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 2, characterized in that: in, The step S2 includes the following sub-steps: checking whether there are missing values in the data set; if there are missing values, performing multivariate interpolation on the data set using the Multivariate Imputation by Chained Equations method.
4. The method for replacing carbon in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 1, characterized in that: in, The correlation analysis was performed using the Spearman rank correlation coefficient analysis method.
5. The method for carbon replacement in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 1, characterized in that: in, The relatively independent data on the performance of the biological activated carbon are divided into a training set and a test set, and the ratio of the data volume of the training set to the data volume of the test set is 80:(10-30).
6. The method for carbon replacement in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 1, characterized in that: in, The machine learning algorithms include four machine learning models: CatBoost, LightGBM, XGBoost, and Random Forest.
7. The method for replacing carbon in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 6, characterized in that: in, The evaluation indicators include root mean square error (RMSE) and coefficient of determination (R 2 ), The calculation formula of the root mean square error of the training set model is as follows: It is used to measure the difference between the predicted value of the training set model and the true value. Its unit is the same as the original unit of the data, where n is the number of samples, yi is the actual value, is the predicted value, The calculation formula of the determination coefficient of the training set model is as follows: It is used to measure the degree of fit of the training set model to the data. The value range of the determination coefficient is between 0 and 1, where n is the number of samples and yi is the actual value. is the predicted value.
8. The method for replacing carbon in a biological activated carbon pool of a water supply plant based on machine learning guidance according to claim 7, characterized in that: in, The optimal carbon replacement model is the model with the smallest root mean square error and the largest determination coefficient among the four machine learning models.
9. A system based on machine learning to guide carbon replacement in biological activated carbon pools of water plants, characterized in that: include: The data collection module collects the test data of the biological activated carbon pool in the water supply plant over the years, including iodine adsorption value, methylene blue adsorption value, intensity, uniformity coefficient, effective particle size, and records the usage time of the biological activated carbon under the corresponding data to obtain the initial performance data of the biological activated carbon; A data preprocessing module cleans and normalizes the initial data of the biological activated carbon performance to obtain preprocessed data of the biological activated carbon performance; A feature selection module performs correlation analysis on the pre-processed data of the biological activated carbon performance, selects relatively independent feature performances, and obtains relatively independent pre-processed data of the biological activated carbon performance; A machine learning model construction module divides the relatively independent data on the performance of the biological activated carbon into a training set and a test set, selects multiple machine learning algorithms, and uses the data of the training set to train the machine learning model to obtain a training set model; A model evaluation and optimization module uses the training set model to predict the data of the training set, calculates the evaluation index of the training set model in the test set, adjusts the parameters of the training set model according to the evaluation results, optimizes the performance of the training set model, and obtains the performance index of the training set optimization model; The carbon replacement model determination module compares the performance indicators of different training set optimization models and selects the optimal carbon replacement model.
Citation Information
Patent Citations
Operation evaluation method of biological activated carbon filter tank
CN108226400A
Chemical material adsorption performance prediction method and device based on automatic machine learning
CN112966447A
Biomass activated carbon methylene blue adsorption performance prediction method based on deep neural network
CN114202060A
Granular active carbon use state detection and replacement method
CN117491238A
Activated carbon adsorption device real-time detection method and system based on deep learning
CN119064052A