A Machine Learning-Driven Method for Predicting the Phosphorus Removal Performance of Metal-Modified Biochar

Through machine learning models, the preparation process parameters of metal-modified biochar are optimized, and the preparation problem of biochar phosphorus removal materials lacking in the existing technology for different water quality treatment targets is solved, and efficient and accurate phosphorus removal performance prediction and material optimization are achieved.

CN118230850BActive Publication Date: 2025-07-22AGRO ENVIRONMENTAL PROTECTION INST OF MIN OF AGRI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410586065.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-07-22
Estimated Expiration
2044-05-13

AI Technical Summary

Technical Problem

The prior art lacks a biochar phosphorus removal material preparation scheme for different water quality treatment targets, and machine learning is insufficient in the preparation of phosphate adsorption materials, making it difficult to effectively guide the recommendation and preparation process of targeted biochar types under different water quality management.

Method used

Using a machine learning-driven method, the preparation process parameters of metal-modified biochar are optimized through random forests, gradient enhancement, extreme gradient enhancement, support vector machine, ridge regression and artificial neural network models, combined with Bayesian optimization and 50% cross-validation, to predict its phosphorus removal performance.

Benefits of technology

It provides a targeted application solution for the biochar preparation process under different water quality management goals, improves the phosphorus removal efficiency and accuracy of the material, and reduces the experimental cost and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118230850B_ABST
    Figure CN118230850B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting the phosphorus removal performance of machine learning-driven metal-modified biochar, comprising the following steps: obtaining the original data under the loading of Mg or La on the biochar to form a data set, the data set including input features and output variables; performing data preprocessing and input feature analysis on the data set; dividing the data set into a training set and a test set, training six machine learning models including random forest (RF), gradient boosting (GBR), extreme gradient boosting (XGB), support vector machine (SVM), ridge regression, and artificial neural network (ANN), and determining the training model after evaluation; performing hyperparameter optimization on the training model through Bayesian optimization combined with five-fold cross-validation to obtain the optimized training model; using the optimized training model to predict the phosphorus removal performance of the metal-modified biochar, and predicting the corresponding preparation process parameters with the adsorption capacity and the remaining phosphate concentration as the targets, providing a new idea for the directional application of the biochar preparation process under different water quality management objectives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a method for predicting the phosphorus removal performance of metal-modified biochar driven by machine learning. Background Art

[0002] Phosphorus (P) plays an important role in the biogeochemical cycle and human production and life. However, a large amount of phosphorus enters natural water bodies, ultimately leading to water environment problems such as water eutrophication. To protect the health of the water environment, the World Health Organization stipulates that the maximum phosphorus concentration in wastewater discharge is limited to 0.5 - 1.0 mg / L, and the total phosphorus concentration in the effluent of urban sewage treatment plants in China is lower than 0.5 mg / L. The critical phosphorus concentration causing water eutrophication is 0.02 mg / L. Therefore, removing or recovering excessive phosphate in water bodies is crucial for controlling pollution. The adsorption method has the advantages of simple operation, environmental friendliness, and low cost, and is widely used in removing phosphate in water bodies. However, due to the huge differences in the phosphate concentration composition of water bodies from different sources, for example, the phosphate concentration in aquaculture wastewater can be as high as several hundred mg / L, while the phosphate in agricultural runoff tailwater is usually only a few to more than a dozen mg / L. The treatment objectives and ideas for water bodies with different phosphate concentrations are different. The treatment of high-concentration phosphate water bodies aims at efficient removal, which is reflected in the construction of materials with high adsorption capacity; the treatment of low-concentration phosphate water bodies aims at meeting the discharge standards, which is reflected in the construction of materials with low residual phosphate concentration. However, there is currently a lack of a construction method for adsorption materials targeted at water quality treatment objectives.

[0003] Biochar has the advantages of a porous structure, rich surface groups, and easy metal loading on the surface. It is a potential phosphorus removal carrier, and the biomass type and pyrolysis temperature are important factors affecting its performance. A large number of studies have shown that Mg and La are metals with a strong affinity for phosphate. The phosphorus removal ability of biochar after loading is much higher than that of the original biochar, which is a research hotspot in current metal-modified biochars. The metal type, loading method, loading amount, etc. directly affect the phosphorus removal performance of metal-modified biochar. In addition, adsorption reaction conditions such as the initial phosphate concentration, solution pH, and adsorbent dosage are also important factors affecting the adsorption performance.

[0004] Current research mainly starts from the material itself to explore the improvement of the phosphate adsorption performance of biochar before and after modification. There has not been any research work on specifically proposing an adsorbent preparation scheme based on water quality treatment objectives. Whether the phosphate concentration meets the relevant standards is also an important indicator for evaluating and quantifying the phosphorus removal performance of adsorbents.

[0005] Machine learning, as a data-driven mathematical method, provides a framework that can map the complex relationships between various input parameters and output variables. It has advantages such as a wide variety of algorithm choices, efficient processing of big data, saving experimental operations and costs, and has been applied to adsorption prediction and material optimization of heavy metals, organic pollutants, emerging pollutants, etc. Machine learning models efficiently analyze the non-linear relationships among numerous characteristic variables affecting adsorption performance during training, understand the material preparation process, predict the adsorption performance of materials, and optimize synthesis process parameters, thereby significantly reducing experimental time and costs, improving efficiency, and making up for the deficiencies of traditional material preparation experiments. However, in the field of phosphate adsorption material preparation, previous studies have one-sidedly pursued high adsorption capacity and lacked consideration of targeted preparation processes based on water quality characteristics. Applying machine learning to guide the recommendation of targeted biochar types and the application of preparation processes under different water quality management objectives is an effective way to break through this bottleneck. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a method for predicting the phosphorus removal performance of metal-modified biochar driven by machine learning.

[0007] The present invention is realized through the following technical solutions:

[0008] A method for predicting the phosphorus removal performance of metal-modified biochar driven by machine learning includes the following steps:

[0009] S1. Obtain the original data under the condition of biochar loaded with Mg or La to form a data set, wherein the data set includes input features and output variables;

[0010] S2. Perform data preprocessing and input feature analysis on the data set;

[0011] S3. Divide the data set into a training set and a test set, train six machine learning models including random forest (RF), gradient boosting (GBR), extreme gradient boosting (XGB), support vector machine (SVM), ridge regression (Ridge-Regression), and artificial neural network (ANN), and determine the training model after evaluation;

[0012] S4. Perform hyperparameter optimization on the training model through Bayesian optimization combined with five-fold cross-validation to obtain an optimized training model;

[0013] S5. Use the optimized training model to predict the phosphorus removal performance of metal-modified biochar.

[0014] According to the above technical solution, preferably, in step S1, the original data includes biomass type, pyrolysis temperature, metal type, metal loading, metal loading method, pH, solid-liquid ratio, initial phosphate concentration of the solution, phosphate adsorption capacity, and residual phosphate concentration.

[0015] According to the above technical solution, preferably, in step S1, the input features are biomass type, pyrolysis temperature, metal type, metal loading, metal loading method, pH, solid-liquid ratio, and initial phosphate concentration of the solution, and the output variables are phosphate adsorption capacity or residual phosphate concentration.

[0016] According to the above technical solution, preferably, in step S2, the data preprocessing includes: performing a normal distribution test on the original data set using the Kolmogorov-Smirnov test, performing data transformation on the features that do not conform to the normal distribution using the Yeo-Johnson transformation, and simultaneously performing data anomaly detection and deleting outliers according to plus or minus three standard deviations.

[0017] According to the above technical solution, preferably, in step S2, Pearson correlation test is used to determine whether there are redundant features among the input features.

[0018] According to the above technical solution, preferably, in step S3, the data set is divided into 50 cycles during the training process, and the coefficient of determination (R 2 )), root mean square error (RMSE), and mean absolute error (MAE) are used to evaluate the model.

[0019] According to the above technical solution, preferably, in step S3, the determined training models after evaluation are random forest (RF), gradient boosting (GBR), and extreme gradient boosting (XGB).

[0020] According to the above technical solution, preferably, in step S4, the optimized training model determines that gradient boosting (GBR) is the optimal model for predicting the phosphorus removal performance of metal-modified biochar.

[0021] According to the above technical solution, preferably, in step S5, feature importance analysis is performed through the EFI, PFI methods and Shap model. It is determined that the most important input features affecting the output variable of phosphate adsorption capacity are metal loading, pH, and initial phosphate concentration of the solution, and the most important input features affecting the output variable of residual phosphate concentration are initial phosphate concentration of the solution, metal loading, and solid-liquid ratio.

[0022] The beneficial effects of the present invention are:

[0023] The present invention is based on six machine learning models, namely Random Forest (RF), Gradient Boosting Regression (GBR), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Ridge Regression, and Artificial Neural Network (ANN). Combining Bayesian optimization and five-fold cross-validation, it optimizes and screens the model with the best performance in predicting phosphate adsorption capacity and residual concentration, and uses the adsorption capacity and residual phosphate concentration as targets to predict the corresponding preparation process parameters, providing new ideas for the directional application of biochar preparation processes under different water quality management objectives. Description of the Drawings

[0024] Figure 1 are the p-values, skewness, and kurtosis of Dataset 1 and Dataset 2 before and after Yeo-Johnson transformation under orthogonal tests.

[0025] Figure 2 Among them, (a)-(e) are the histograms of the data distributions of LoadC, WV, pH, C0, and T in Dataset 1; (f)-(j) are the histograms of the data of LoadC, WV, pH, C0, and T in Dataset 1 after Yeo-Johnson transformation; (k)-(o) are the histograms of the data distributions of LoadC, WV, pH, C0, and T in Dataset 2; (p)-(t) are the histograms of the data of LoadC, WV, pH, C0, and T in Dataset 2 after Yeo-Johnson transformation.

[0026] Figure 3 Among them, (a)-(j) are the Q-Q plots comparing the data cleaning of the features LoadC, WV, pH, C0, and T in Dataset 1 before and after; (k)-(t) are the Q-Q plots comparing the data cleaning of the features LoadC, WV, pH, C0, and T in Dataset 2 before and after.

[0027] Figure 4 are the means and standard deviations of the numerical features in Dataset 1' and Dataset 2'.

[0028] Figure 5 are the results of the feature data distribution and statistical analysis in the present invention.

[0029] Figure 6 Among them, (a) is the heat map of Pearson correlation analysis between the input features of Dataset 1'; (b) is the heat map of Pearson correlation analysis between the input features of Dataset 2'.

[0030] Figure 7 Among them, (a) is the R of the training set obtained by splitting Dataset 1' into fifty cycles during the training of six machine learning models 2Box plot of; (b) Box plot of RMSE of the training sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 1'; (c) Box plot of MAE of the training sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 1';

[0031] (d) R of the test sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 1' 2 Box plot of; (e) Box plot of RMSE of the test sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 1'; (f) Box plot of MAE of the test sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 1';

[0032] (g) R of the training sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 2' 2 Box plot of; (h) Box plot of RMSE of the training sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 2'; (i) Box plot of MAE of the training sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 2';

[0033] (j) R of the test sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 2' 2 Box plot of; (k) Box plot of RMSE of the test sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 2'; (l) Box plot of MAE of the test sets obtained from fifty cyclic partitions during the training of six machine learning models on dataset 2'.

[0034] Figure 8 are the average R of the training sets and test sets, RMSE, and MAE of dataset 1' and dataset 2' under the training of six models 2 、RMSE and MAE.

[0035] Figure 9 are the R 2 、RMSE and MAE before and after Bayesian optimization and the corresponding optimal hyperparameters.

[0036] Figure 10 In, (a) is the scatter plot of the predicted values and true values of the RF model based on dataset 1' after hyperparameter optimization; (b) is the scatter plot of the predicted values and true values of the GBR model based on dataset 1' after hyperparameter optimization; (c) is the scatter plot of the predicted values and true values of the XGB model based on dataset 1' after hyperparameter optimization;

[0037] (d) Scatter plot of the predicted values and true values of the RF model based on dataset 2' after hyperparameter optimization; (e) Scatter plot of the predicted values and true values of the GBR model based on dataset 2' after hyperparameter optimization; (f) Scatter plot of the predicted values and true values of the XGB model based on dataset 2' after hyperparameter optimization.

[0038] Figure 11 Among them, (a) shows the top five feature importances of EFI in dataset 1'; (b) shows the top five feature importances of PFI in dataset 1'; (c) shows the top five feature importances of the Shap model in dataset 1'.

[0039] (d) shows the top five feature importances of EFI in dataset 2'; (e) shows the top five feature importances of PFI in dataset 2'; (f) shows the top five feature importances of the Shap model in dataset 2'.

[0040] Figure 12 Among them, (a)-(e) are the PDP and ICE calculated based on GBR for the input features T, LoadC, WV, pH, and C0 in dataset 1; (f)-(j) are the PDP and ICE calculated based on GBR for the input features T, LoadC, WV, pH, and C0 in dataset 2.

[0041] Figure 13 is the optimal Q based on GBR e and the lowest C e Preparation process parameter optimization of and Qe and C e predicted values and experimental values.

[0042] Figure 14 are the Langmuir and Freundlich isotherms for MgBC to adsorb phosphate (conditions: C0 = 10–500 mg / L, tcontact = 24 h, m / V = 1 g / L, pH = 4, and T = 25 °C). Detailed implementation manner

[0043] To enable those skilled in the art of the present technology to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and the best embodiments. Based on the embodiments of the invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the invention.

[0044] As shown in the figure, the present invention includes the following steps:

[0045] S1. Obtain the original data under biochar loaded with Mg or La to form a dataset.

[0046] Among them, the original data includes biomass type (Biochar), pyrolysis temperature (T), metal type (Metal), metal loading (LoadC, the molar amount of metal loaded per 1 g of biomass), metal loading method (LoadM), pH, solid-liquid ratio (WV), initial phosphate concentration in the solution (C0), phosphate adsorption capacity (Q e ), and residual phosphate concentration (C e ). Based on the above data, two datasets are established: Dataset 1 and Dataset 2. The input features are all set as Biochar, T, Metal, LoadC, LoadM, pH, WV, and C0. The output variables of Dataset 1 and Dataset 2 are Q e and C e respectively. The data is sourced from relevant literature on Mg or La-loaded biochar for phosphorus removal in the Web of Science database. All relevant data is directly extracted from the original tables and charts of the papers using WebPlot Digitizer to ensure data accuracy, without any modification or deviation. 427 valid data groups are obtained for Dataset 1, and 441 valid data groups are obtained for Dataset 2.

[0047] S2. Perform data preprocessing and input feature analysis on the dataset.

[0048] Among them, the data preprocessing includes: using the Kolmogorov-Smirnov test to perform a normal distribution test on the original dataset. For features that do not conform to the normal distribution, data transformation is performed using the Yeo-Johnson transformation. At the same time, data anomaly detection is performed based on plus or minus three standard deviations, and outliers are deleted.

[0049] The KS test (Kolmogorov-Smirnov test) is a non-parametric test method used to test whether two probability distributions are the same. It judges whether there are significant differences between two distributions based on the maximum difference between the cumulative distribution function (CDF) of the sample and the theoretical distribution. Hypothesis testing problem:

[0050] H0: The overall distribution of the sample follows a normal distribution.

[0051] H1: The overall distribution of the sample does not follow a normal distribution.

[0052] Let the probability distribution function to be tested be F(x), and the probability distribution function of the theoretical distribution be G(x). Arrange the sample data in ascending order, and calculate the empirical distribution function value Fn(x) corresponding to each observed value. Calculate the KS statistic Dn = max|Fn(x) - G(x)|, calculate the p-value in combination with the KS statistic and the sample size, and compare it with the significance level (0.05) to judge: if p > 0.05, then accept H0; otherwise, reject H0 and accept H1.

[0053] Yeo-Johnson Transformation:

[0054] For a random variable on the given real number field , the Yeo-Johnson transformation is a piecewise function composed of power functions, and its expression is:

[0055]

[0056] where is the data before transformation, is the data after transformation, is the transformation coefficient.

[0057] Figure 1 , 2 shows the p-values, skewness, and kurtosis of five numerical features (T, LoadC, pH, WV, and C0) of the dataset before and after the Yeo-Johnson transformation under orthogonal tests. Since the dataset used in the present invention comes from various biochar materials, the five numerical features before transformation do not conform to the normal distribution. However, the Yeo-Johnson transformation significantly improves the data distribution, the p-values increase significantly, and the data is more approaching the normal distribution structure. After the Yeo-Johnson transformation, the skewness and kurtosis of dataset 1 and dataset 2 are significantly reduced, the overall skewness is closer to 0, and the kurtosis is less than 3, indicating that the data symmetry is significantly improved and shows a flat distribution feature with a lower kurtosis.

[0058] Datasets 1' and 2' are obtained by deleting data outliers from datasets 1 and 2 through the Yeo-Johnson transformation and the method of plus or minus three times the standard deviation. Figure 3 The Q-Q scatter plots of dataset 1 and 1' and dataset 2 and 2' are compared. The Q-Q plot drawn for data conforming to the normal distribution will approach a straight line , and the slope is the standard deviation, and the intercept is the mean. Figure 3 It shows that after data cleaning, the data of the two datasets changes from an irregular distribution to approaching a straight line, and the slope and intercept of the straight line equation are closer to the standard deviation and the mean respectively ( Figure 4 ), verifying the improvement of data normality again. The subsequent model training uses datasets 1' and 2' as input data.

[0059] In addition, the Pearson correlation coefficient (PCC) is the linear correlation between any two continuous variables, and the Pearson correlation test is used to determine whether there are redundant features among the input features.

[0060] The Pearson correlation coefficient is calculated as follows:

[0061]

[0062] wherein, is the PCC value of two input features; and are the input variables and the output variable respectively. takes values between -1 and 1, where positive and negative values indicate positive or negative correlation, and 0 indicates no linear correlation.

[0063] In this example, the input features include biomass type (Biochar), pyrolysis temperature (T), metal type (Metal), metal loading (LoadC, the molar amount of metal loaded per 1 g of biomass), metal loading method (LoadM), pH, solid-liquid ratio (WV), and initial phosphate concentration in the solution (C0). The data distribution and statistical analysis results of all features are as Figure 5 shown. The commonly used biomass types are divided into 5 types: fruit shell (GK), fallen leaves (LY), straw (JG), wood (MZ), and sludge (WN). The main metal types are Mg and La, and the metal loading methods mainly include dipping (Dipping) and coprecipitation (Precip). The range of biochar firing temperature in the dataset is 300 - 800 °C, the metal loading is between 0.05 - 20.00 mmol / g, the solid-liquid ratio is between 0.10 - 4.00 mg / L, the solution pH is mainly between 2 - 12, and the initial phosphate concentration is between 2 - 400 mg / L.

[0064] The independence of input features is crucial for the prediction efficiency of the model. PCC quantifies the linear correlation between each numerical feature and the output variable ( Figure 6 ), and the results show that the PCC values are all at a low level, indicating that there is a weak linear correlation between any two features and they have good independence. Therefore, these input features can well avoid redundant information and overfitting problems.

[0065] S3. Divide the dataset into a training set and a test set, train six machine learning models: Random Forest (RF), Gradient Boosting (GBR), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Ridge Regression, and Artificial Neural Network (ANN), and determine the training model after evaluation.

[0066] In this example, 80% of the data was randomly selected as the training set, and the remaining 20% of the data in the dataset was used as the test set. Six machine learning models, namely Random Forest (RF), Gradient Boosting (GBR), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Ridge Regression, and Artificial Neural Network (ANN), were trained based on Scikit-Learn (version 1.3.2) in Python 3.9. During the training process, the dataset was divided into 50 loops to examine the stability of the models. Box plots were drawn to statistically analyze the evaluation metrics of each model for each split. The coefficient of determination (R 2 ), Root Mean Square Error (RMSE), and Mean Absolute Error (MAE) were used to evaluate the models, and the models with better performance were further optimized using hyperparameter tuning.

[0067] The coefficient of determination (R 2 ), Root Mean Square Error (RMSE), and Mean Absolute Error (MAE) are indicators for evaluating the accuracy of each model. The higher the R 2 and the lower the RMSE and MAE values, the better the simulation effect of the model. The calculation formulas for R 2 , RMSE, and MAE are as follows:

[0068]

[0069]

[0070]

[0071] Among them, represents the number of test sets, is the predicted value, is the true value, is the average value of the true values.

[0072] The R 2 , RMSE, and MAE of the training set and test set of these six models were evaluated in 50 loop splits as shown in Figure 7 , 8 . Figure 7 a-c show the characteristics of the six models on the training set of dataset 1': The R 2 of RF, GBR, and XGB are all higher than 0.9, and the RMSE and MAE are much lower than those of SVM, Ridge, and ANN. Figure 7 d-f show the performance of the six models on the test set of dataset 1'. RF, GBR, and XGB also have relatively high R 2 and relatively low RMSE and MAE. In the later stage, the performance difference between the training set and the test set can be reduced through hyperparameter optimization. Figure 7g-l shows the R of 6 models on the training set and test set of dataset 2'. 2 , RMSE and MAE. RF, GBR and XGB are still the top 3 models with better performance. In the training set, GBR shows absolute prediction advantages, followed by RF and XGB. However, in the test set, the R 2 is generally low, and cross-validation and hyperparameter optimization need to be used to improve the generalization ability. In summary, in order to build a model with better performance and stronger generalization ability, after evaluation, random forest (RF), gradient boosting (GBR), and extreme gradient boosting (XGB) are determined as the subsequent training models.

[0073] S4. Hyperparameter optimization of the training model is carried out by combining Bayesian optimization with five-fold cross-validation to obtain an optimized training model.

[0074] Hyperparameter tuning aims to automatically optimize the hyperparameter configuration of the model through a designed algorithm to obtain the best performance under a given dataset. Its core goal is to pursue higher model prediction accuracy. The Bayesian optimization algorithm can continuously update the prior according to the parameter information, with fewer iteration times and faster calculation speed. It is still robust even for non-convex problems and has the characteristics of high efficiency, automation, and intelligence, and has more advantages for complex hyperparameter spaces and noisy data. Combining Bayesian optimization with five-fold cross-validation avoids the blindness of random sampling and direct evaluation, reduces the risk of overfitting, improves the efficiency and accuracy of hyperparameter optimization, and enhances the model generalization ability.

[0075] In the mathematical process of Bayesian optimization, the following steps are mainly executed:

[0076] (1) Define f(x) to be estimated and the domain of x;

[0077] (2) Take out the values of f(x) corresponding to these x by taking the values of n finite x (solving the observed values);

[0078] (3) Estimate the function based on the finite observed values (this hypothesis is called the prior knowledge in Bayesian optimization), and obtain the target value (maximum or minimum value) on this estimated f*;

[0079] (4) Define a certain rule to determine the next observation point to be calculated.

[0080] Continuously loop in steps 2-4 until the target value on the assumed distribution reaches our standard, or all computing resources are used up (for example, at most m observations, or at most allowed to run for t minutes).

[0081] In addition, K-fold cross-validation divides the original data into K groups. Each subset of data is used as a test set once, and the remaining K - 1 subsets of data are used as the training set, resulting in K models. These K models are evaluated in the validation set respectively, and the final errors are summed and averaged to obtain the cross-validation error, thus eliminating the adverse effects caused by unbalanced data division in a single division.

[0082] Specifically, the specific operation steps of Bayesian optimization combined with K-fold (five-fold) cross-validation in this example are as follows:

[0083] (1) Introduce the BayesianOptimization function in the bayes_opt library;

[0084] (2) Define the optimization target models (RF, GBR, and XGB respectively);

[0085] (3) Define the scoring of the model by five-fold cross-validation (using R 2 as the evaluation index);

[0086] (4) Define the parameter space of the Bayesian optimization object;

[0087] RF: n_estimators, max_depth, min_samples_split, min_samples_leaf

[0088] GBR: n_estimators, max_depth, min_samples_split, min_samples_leaf, learning_rate

[0089] XGB: n_estimators, max_depth, learning_rate, subsample

[0090] (5) Call the BayesianOptimization function for Bayesian optimization;

[0091] (6) Retrain the target models (RF, GBR, and XGB respectively) with the optimized parameters;

[0092] (7) Output the optimized hyperparameters and the evaluation indexes of the optimized models on the training set and the test set.

[0093] RF#1’: 'max_depth' = 13,'min_samples_leaf' = 4,'min_samples_split' = 9, 'n_estimators' = 151

[0094] GBR#1’: 'learning_rate' = 0.437147064543167,'max_depth' = 6,'min_samples_leaf' = 2,'min_samples_split' = 3, 'n_estimators' = 155

[0095] XGB#1’: 'learning_rate' = 0.07212155094905877,'max_depth' = 5, 'n_estimators' = 162,'subsample' = 0.507900831830210

[0096] RF#2’: 'max_depth' = 15,'min_samples_leaf' = 1,'min_samples_split' = 2, 'n_estimators' = 95

[0097] GBR#2’: 'learning_rate' = 0.7605969085323319,'max_depth' = 6,'min_samples_leaf' = 1,'min_samples_split' = 4, 'n_estimators' = 146

[0098] XGB#2’: 'learning_rate' = 0.06021504687955905,'max_depth' = 15, 'n_estimators' = 104,'subsample' = 0.5327101881334091

[0099] Figure 9 Shows the best combinations of R 2 , RMSE and MAE of RF, GBR and XGB before and after Bayesian optimization, as well as the hyperparameters n_estimators and max_depth. The R 2 of RF, GBR and XGB imported with dataset 1' after Bayesian optimization all reached above 0.95, generally increasing by 0.1, while RMSE and MAE were also reduced to varying degrees after optimization. Generally speaking, the prediction performance ranking of the three models is GBR > XGB > RF. The one with the most excellent prediction performance is GBR, and the corresponding n_estimators and max_depth are 155 and 6 respectively. Combining cross-validation and Bayesian optimization to optimize the RF, GBR and XGB models for dataset 2', the R 2 of RF, GBR and XGB after optimization all increased to above 0.94, and the R 2The values are close, solving the overfitting problem and showing an ideal fitting effect. At the same time, RMSE and MAE decrease within a small range. According to Figure 5 it can be seen that it is mainly caused by the large numerical differences in the output variables of different data sets. The output variable Q of data set 1’ e has a significantly larger mean than the output variable C of data set 2’. e The prediction performance of the three models for data set 2’ is ranked as GBR>XGB>RF. The best prediction effect is still GBR, and the corresponding best hyperparameter combination is n_estimators = 146, max_depth = 6.

[0100] The distributions of the predicted values and the true values of RF, GBR, and XGB after training with the best hyperparameters on data sets 1’ and 2’ are shown in Figure 10 . The distribution of the data scatter points of the GBR model is closer to the dashed line y = x, indicating that the predicted values are closer to the true values ( Figure 10 b), while the RF and XGB models are relatively inferior ( Figure 10 a and Figure 10 c). Similarly, on data set 2’, the trends of the predicted values and the true values of the three models are similar to those of data set 1’ ( Figure 10 d-f), which also verifies the conclusion of the prediction performance ranking. In summary, among the six models, gradient boosting (GBR) is the most suitable model for predicting the phosphate adsorption performance of modified biochar.

[0101] S5. Use the optimized training model to predict the phosphorus removal performance of metal-modified biochar.

[0102] While using the optimized training model to predict the phosphorus removal performance of metal-modified biochar, feature importance analysis can also be carried out through the EFI, PFI methods and the Shap model, and partial dependence analysis can be carried out using the PDP and ICE methods at the same time.

[0103] Feature importance is a metric used to evaluate the contribution of each feature to the prediction results of a machine learning model, which helps to further optimize the model, screen features, and optimize the preparation process parameters. In this paper, three methods, Embedded Feature Importance (EFI), Permutation Feature Importance (PFI), and Shapley Additive explanations (Shap), are used to comprehensively analyze feature importance. EFI evaluates and directly obtains the importance results of each feature to the model's prediction performance by calling the algorithm built into the model to calculate feature importance. PFI evaluates the importance of features by randomly permuting the feature values and then recalculating the change in the model's prediction performance. In the Shap model, the Shap value of each feature is calculated, and a higher Shap value indicates that the feature contributes more to the model prediction.

[0104] The visualization analysis results of the feature importance of the GBR model by the EFI, PFI methods, and the Shap model are shown in Figure 11 . The calculation results of EFI and PFI show that the most important features affecting the modified biochar Q e are the metal loading LoadC, pH, and the initial phosphate concentration C0 of the solution ( Figure 11 a and 11b). The sum of the contributions of the three features is 0.833 and 0.856 respectively. The prediction results of the Shap model also consider that the metal loading LoadC and pH are the most significant features affecting the metal-modified biochar Q e ( Figure 11 c). Thus, it can be seen that the difference in metal loading during the preparation of the modified biochar significantly affects the phosphate adsorption capacity of the biochar. The most important features affecting the remaining concentration C of phosphate adsorbed by the metal-modified biochar e are the initial phosphate concentration C0 of the solution, the metal loading LoadC, and the solid-liquid ratio WV ( Figure 11 d and 11e). The sum of the feature contributions calculated by EFI and PFI is 0.741 and 0.796 respectively, which is exactly the same as the results of the Shap model predicting feature importance ( Figure 11 f). However, for the actual water body, C0 is an uncontrollable factor, and WV belongs to the adsorption process conditions. Therefore, the optimization focus of the preparation process should be placed more on LoadC.

[0105] Partial dependence analysis usually uses two visualization tools, Partial Dependence Plot (PDP) and Individual Conditional Expectation (ICE). PDP reveals the marginal impact of the average value of all samples of a single feature on the output variable. ICE reveals the marginal impact of all sample values of a single feature on the predicted output variable.

[0106] PDP and ICE were plotted based on the best trained GBR model, and the partial dependence of numerical features on the average sample and individual sample levels was analyzed. Metal loading LoadC is considered to affect the Q of metal-modified biochar. e and C e One of the most important features ( Figure 12 b and Figure 12 g), LoadC to Q e The overall effect showed a trend of first increasing and then stabilizing. When LoadC gradually increased to 5mmol / g, Q e Gradually increase, LoadC exceeds 10 mmol / g e The reason why it is basically stable is that when the metal loading is small, the biochar surface is more conducive to metal binding sites, and the adsorption capacity continues to increase with the increase of the loading. When the metal loading increases to a certain extent, the active sites on the biochar surface are covered, and even metal stacking occurs, which makes the adsorption capacity stable or slightly decreased. e The effect of Q e The changes are affected by the amount of active sites. The pH range is 5-7 with dense Q e sample( Figure 12 d), indicating that the developed biochar is more suitable for the pH range of 5-7, and the solution pH between 3-8 is more conducive to adsorption by the adsorbent, thereby reducing C e ( Figure 12 i). The overall range of C0 changes relatively smoothly ( Figure 12 e), indicating the effect of initial concentration on Q e The effect of C0 on the adsorption capacity is not significant, which again shows that the adsorption capacity is a property of the material itself. e The increase ( Figure 12 j), and the influence trend is significant, which is consistent with the objective experimental phenomenon. From the PDP average line, we can see that T and WV have a significant influence on Q e or C e The effect was not significant ( Figure 12 a, f, c, h). PDP and ICE help analyze the trend of the independent effects of input features and the performance of a single feature in individual samples.

[0107] In addition, based on the machine learning model with better performance after hyperparameter optimization, the highest adsorption capacity (Q e ) and the minimum residual concentration (C e ) is used as the output target to predict the preparation process parameters of modified biochar. Modified biochars were prepared according to the parameter combinations optimized by the model, and batch adsorption experiments were carried out to test the adsorption capacity and residual concentration to verify the reliability of the model optimization parameters.

[0108] The preparation process of metal-modified biochar with the highest phosphate adsorption capacity (Q e ) and the lowest residual concentration (C e ) is output by the trained RF, GBR, and XGB models ( Figure 13 ). The results of parameter optimization of the three models with Q e as the output variable are highly consistent. The biomass type is fallen leaves (LY), the pyrolysis temperature is 550 °C, the metal type is Mg, the loading amount is 8.33 mmol / g, the loading method is impregnation, the solid-liquid ratio is 1 g / L, and the predicted Q e is 387.24 - 396.61 mg / g at pH 4. The preparation process parameters of metal-modified biochar with the lowest residual concentration (C e ) predicted by the GBR and XGB models are: the biomass type is wood, the pyrolysis temperature is 600 °C, the metal type is La, the loading amount is 3.56 mmol / g, the loading method is coprecipitation, the solid-liquid ratio is 0.4 g / L, and the predicted C e is 0 mg / L.

[0109] According to the optimized process, materials with the highest phosphate adsorption capacity, MgBC (Mg-loaded modified biochar material), and materials with the lowest residual concentration, LaBC (La-loaded modified biochar material), are prepared respectively, and batch adsorption tests are carried out. The Langmuir model can better describe the behavior of MgBC adsorbing phosphate ( Figure 14 ), and the R 2 values are 0.96 and 0.88 respectively, and the maximum adsorption amount of phosphate is 372.01 ± 17.40 mg / g. The residual concentration of LaBC after adsorption equilibrium in 10 mg / L initial phosphate solution at pH 3 and 5 is 0 mg / L, which is completely consistent with the prediction results of the machine learning model. Therefore, the material recommended for high-concentration phosphate wastewater phosphorus recovery is Mg-based biochar (Mg-loaded modified biochar material), and the material recommended for controlling low-concentration phosphate wastewater or preventing water eutrophication is La-based biochar (La-loaded modified biochar material).

[0110] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A machine learning-driven prediction method for the phosphorus removal performance of metal-modified biochar, characterized in that, It includes the following steps: S1. Obtain the original data on the adsorption of phosphate in water by Mg- or La-loaded biochar to form a data set. The original data includes biomass type, pyrolysis temperature, metal type, metal loading, metal loading method, pH, solid-liquid ratio, initial phosphate concentration in the solution, phosphate adsorption capacity, and residual phosphate concentration. The data set includes input features and output variables. The input features are biomass type, pyrolysis temperature, metal type, metal loading, metal loading method, pH, solid-liquid ratio, and initial phosphate concentration in the solution. The output variables are phosphate adsorption capacity or residual phosphate concentration. S2. Perform data preprocessing and input feature analysis on the data set. The data preprocessing includes: using the Kolmogorov-Smirnov test to perform a normal distribution test on the original data set, performing data transformation on the features that do not conform to the normal distribution using the Yeo-Johnson transformation, and simultaneously performing data anomaly detection and deleting outliers according to plus or minus three standard deviations, so that the data of the data set changes from an irregular distribution to approaching a straight line, and the slope and intercept of the straight line equation are closer to the standard deviation and the mean respectively. Judge whether there are redundant features among the input features through Pearson correlation test to make any two input features independent of each other. S3. Divide the data set into a training set and a test set, train six machine learning models including random forest, gradient boosting, extreme gradient boosting, support vector machine, ridge regression, and artificial neural network, and determine the training model after evaluation. S4. Perform hyperparameter optimization on the training model through Bayesian optimization combined with five-fold cross-validation to obtain the optimized training model. S5. Use the optimized training model to perform feature importance analysis through the EFI, PFI methods and the Shap model. It is determined that the most important input features affecting the output variable of phosphate adsorption capacity are metal loading, pH, and initial phosphate concentration in the solution. It is determined that the most important input features affecting the output variable of residual phosphate concentration are initial phosphate concentration in the solution, metal loading, and solid-liquid ratio. Among them, the metal loading is defined as one of the most significant input features affecting the output variable of phosphate adsorption capacity and the output variable of residual phosphate concentration, rather than the initial phosphate concentration in the solution. Predict the phosphorus removal performance of metal-modified biochar. Specifically, predict the preparation process parameters of modified biochar with the highest phosphate adsorption capacity and the lowest residual phosphate concentration as the output targets respectively. Among them, based on the optimized training model, it is predicted that Mg-loaded modified biochar material is recommended when the highest phosphate adsorption capacity is used as the output target, and La-loaded modified biochar material is recommended when the lowest residual phosphate concentration is used as the output target.

2. The method for predicting the phosphorus removal performance of machine learning-driven metal-modified biochar according to claim 1, wherein, In step S3, the data set is divided into 50 cycles during the training process, and the model is evaluated using the coefficient of determination, root mean square error, and mean absolute error.

3. The method for predicting the phosphorus removal performance of machine learning-driven metal-modified biochar according to claim 2, wherein In step S3, the determined training models after evaluation are random forest, gradient boosting, and extreme gradient boosting.

4. The method for predicting the phosphorus removal performance of machine learning-driven metal-modified biochar according to claim 3, wherein, In step S4, the optimized training model determines that gradient boosting is the optimal model for predicting the phosphorus removal performance of metal-modified biochar.

Citation Information

Patent Citations

  • Gas high-pressure isothermal adsorption curve prediction method and system, storage medium and terminal

    CN111625953A

  • Machine learning prediction method for As (III) and As (V) adsorption of biochar

    CN116564440A