Prediction method and system for lead adsorption of biochar based on machine learning model

Optimizing the prediction method of biochar lead adsorption through machine learning models, the time-consuming and cost-effective problems in the existing technology are solved, and more efficient adsorption performance optimization and wastewater treatment quality improvement are achieved.

CN120299557APending Publication Date: 2025-07-11SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376325.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the experimental method of biochar adsorbing lead is time-consuming, labor-intensive and costly, making it difficult to accurately predict its adsorption capacity through empirical models, and traditional modeling methods cannot consider the interaction between features, resulting in low efficiency in optimizing adsorption performance.

Method used

Using machine learning model, by obtaining the data set of biochar adsorbed lead, data visualization and preprocessing are carried out, grid search technology and early stop mechanism are introduced, machine learning model is optimized hyperparameters, screening the optimal model, and SHAP feature importance analysis and graphical user interface application to establish a prediction system.

Benefits of technology

The prediction research efficiency of biochar lead adsorption is improved, the adsorption performance is optimized, the wastewater treatment quality is improved, and more accurate prediction and optimization results are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299557A_ABST
    Figure CN120299557A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for predicting lead adsorption of biochar based on a machine learning model. The method comprises the following steps: acquiring a data set of lead adsorption of biochar and determining an input variable of the data set and an output variable of the data set; performing data visualization and preprocessing to obtain a preprocessed data set; introducing a grid search technology and an early stop mechanism, performing hyper-parameter optimization and training on the machine learning model through the preprocessed data set, and screening to obtain an optimal machine learning model; sHAP feature importance analysis and graphical user interface application program establishment are carried out, and an importance sequence of adsorption capacity prediction influenced by input features and a graphical user interface are obtained. The research efficiency of the prediction model can be improved through a machine learning method, so that the adsorption efficiency of the biochar is improved, and the wastewater treatment quality is improved. The method and the system for predicting the lead adsorbed by the biochar based on the machine learning model can be widely applied to the technical field of prediction of the lead adsorbed by the biochar.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of predicting lead adsorption by biochar, and in particular to a method and system for predicting lead adsorption by biochar based on a machine learning model. Background Art

[0002] Due to the diversity of biochar types and the variation of adsorption test conditions, there are significant differences in the adsorption capacities of different biochars. Therefore, it is often necessary to optimize the adsorption process of biochar to achieve better adsorption performance. In recent years, optimization studies have been widely carried out through batch experiments to improve the adsorption performance of Pb(II) on biochar adsorbents. However, conducting scientific experiments to explore the best characteristics and conditions for improving Pb adsorption efficiency is time-consuming, labor-intensive, and expensive. Traditional labor-intensive laboratory-scale experiments have limited effectiveness in enhancing the design of biochar production. Most biochar adsorption studies follow similar experimental procedures, involving quantitative comparisons of adsorption capacities, biochar characteristics, and environmental impacts under various conditions. Then, kinetic and thermodynamic models are constructed to analyze the adsorption results and explain the mechanisms based on biochar characteristics and physical and chemical principles. However, the contributions and impacts of these characteristics on the biochar adsorption capacity cannot be clearly elucidated, which hinders accurate prediction through empirical models. In addition, the one-factor-at-a-time method in experiments can only briefly examine the influence of one characteristic and cannot consider the interactions between characteristics. Moreover, the modeling methods in related technologies for in-depth study of adsorbent surface characteristics and their interactions with pollutants, such as linear correlation, multiple linear regression, and response surface optimization, etc., but their applicability and accuracy in practical applications are often limited. In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention

[0003] In order to solve the above technical problems, the object of the present invention is to provide a method and system for predicting lead adsorption by biochar based on a machine learning model, which can improve the research efficiency of the prediction model through machine learning methods, thereby improving the adsorption efficiency of biochar and the quality of wastewater treatment.

[0004] The first technical solution adopted by the present invention is: A method for predicting lead adsorption by biochar based on a machine learning model, comprising the following steps:

[0005] Obtain a dataset of lead adsorption by biochar and determine the input variables and output variables of the dataset;

[0006] Perform data visualization and preprocessing on the dataset of lead adsorption by biochar to obtain a preprocessed dataset;

[0007] The grid search technique and early stopping mechanism are introduced to optimize the hyperparameters and train the machine learning model through the preprocessed dataset, and the optimal machine learning model is selected through screening.

[0008] Based on the optimal machine learning model, SHAP feature importance analysis and graphical user interface application program are established to obtain the importance ranking of input features affecting the adsorption capacity prediction and the graphical user interface.

[0009] Furthermore, the input variables of the dataset include preparation conditions, biochar properties, and adsorption conditions. The preparation conditions include pyrolysis temperature and pyrolysis time. The biochar properties include pH value, surface area, pore volume, average pore diameter, C(%), H(%), O(%), N(%), H / C, O / C, (O + N) / C, aromaticity index, and double bond equivalent. The adsorption conditions include initial Pb(II) concentration, biochar dosage, solution pH value, adsorption time, stirring speed, and adsorption temperature. The output variable of the dataset is the adsorption capacity of biochar for lead.

[0010] Furthermore, the step of performing data visualization and preprocessing on the dataset of biochar adsorbing lead to obtain the preprocessed dataset specifically includes:

[0011] A box plot is drawn for the visualization of the dataset, and Pearson correlation test is used to determine whether there are redundant features among the input features and the redundant features are deleted to obtain the dataset after deleting redundant features;

[0012] The average value of the dataset after deleting redundant features is obtained to replace the missing values in the dataset after deleting redundant features to obtain the replaced dataset;

[0013] The replaced dataset is successively subjected to standardization and partitioning processing to obtain the preprocessed dataset.

[0014] Furthermore, the step of introducing the grid search technique and early stopping mechanism to optimize the hyperparameters and train the machine learning model through the preprocessed dataset, and screening to obtain the optimal machine learning model specifically includes:

[0015] A machine learning model is obtained, and the machine learning model includes CatBoost model, ANN model, and BPNN model;

[0016] The grid search technique and early stopping mechanism are introduced. First, the best combination of hyperparameters is found, and then the CatBoost model is iteratively trained through the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and the trained CatBoost model is output;

[0017] The grid search technique and early stopping mechanism are introduced. First, the best combination of hyperparameters is found, and then the ANN model is iteratively trained with the preprocessed dataset until the early stopping condition is met or the preset number of iterations is reached, and the trained ANN model is output;

[0018] The grid search technique and early stopping mechanism are introduced. First, the best combination of hyperparameters is found, and then the BPNN model is iteratively trained with the preprocessed dataset until the early stopping condition is met or the preset number of iterations is reached, and the trained BPNN model is output;

[0019] The coefficient of determination and root mean square error values of the test sets of the trained CatBoost model, the trained ANN model, and the trained BPNN model are obtained and compared respectively, and the trained CatBoost model is selected as the optimal machine learning model.

[0020] Furthermore, the step of performing SHAP feature importance analysis and establishing a graphical user interface application based on the optimal machine learning model to obtain the importance ranking of input features affecting the adsorption capacity prediction and the graphical user interface specifically includes:

[0021] Performing SHAP feature importance analysis based on the optimal machine learning model to obtain the average absolute SHAP value of each input feature;

[0022] Calculating the importance value of each input feature according to the average absolute SHAP value and performing sorting processing to obtain the feature importance index of the dataset;

[0023] Based on the optimal machine learning model, use Tkinter to create a graphical user interface that allows users to input feature values for prediction.

[0024] The second technical solution adopted by the present invention is: A prediction system for biochar adsorption of lead based on a machine learning model, including:

[0025] The first module is used to obtain the dataset of biochar adsorption of lead and determine the input variables and output variables of the dataset;

[0026] The second module is used to perform data visualization and preprocessing on the dataset of biochar adsorption of lead to obtain the preprocessed dataset;

[0027] The third module is used to introduce the grid search technique and early stopping mechanism, optimize the hyperparameters and train the machine learning model with the preprocessed dataset, and select the optimal machine learning model;

[0028] The fourth module is used to perform SHAP feature importance analysis and establish a graphical user interface application based on the optimal machine learning model, obtaining the importance ranking of input features affecting the adsorption capacity prediction and the graphical user interface.

[0029] The beneficial effects of the method and system of the present invention are as follows: By obtaining the dataset of biochar adsorbing lead and determining the input variables and output variables of the dataset, and then performing data preprocessing, and introducing an early stopping mechanism, training the machine learning model with the preprocessed dataset, screening to obtain the optimal machine learning model. The prediction model established based on the machine learning method can improve the research efficiency. Finally, performing SHAP feature importance analysis and display processing on the dataset of biochar adsorbing lead with the optimal machine learning model. By learning and analyzing various parameters in the wastewater adsorption treatment process, the algorithm can accurately predict and optimize the treatment results, ultimately improving the adsorption efficiency of biochar and the quality of wastewater treatment. Description of the Drawings

[0030] Figure 1 is the step flow chart of a prediction method for biochar adsorbing lead based on a machine learning model of the present invention;

[0031] Figure 2 is the structural block diagram of a prediction system for biochar adsorbing lead based on a machine learning model of the present invention;

[0032] Figure 3 is the schematic diagram of the process framework for the prediction of biochar adsorbing lead provided by a specific embodiment of the present invention;

[0033] Figure 4 is the box plot schematic diagram of the data distribution of the data input and output variables provided by a specific embodiment of the present invention;

[0034] Figure 5 is the heat map schematic diagram of the Pearson correlation matrix between any two variables provided by a specific embodiment of the present invention;

[0035] Figure 6 is the schematic diagram of the CatBoost model structure provided by a specific embodiment of the present invention;

[0036] Figure 7 is the schematic diagram of the ANN model structure provided by a specific embodiment of the present invention;

[0037] Figure 8 is the schematic diagram of the BPNN model structure provided by a specific embodiment of the present invention;

[0038] Figure 9 is the scatter plot schematic diagram of the predicted value and actual value of the CatBoost model provided by a specific embodiment of the present invention;

[0039] Figure 10 It is a scatter diagram of the predicted values and actual values of the ANN model provided by a specific embodiment of the present invention;

[0040] Figure 11 It is a scatter diagram of the predicted values and actual values of the BPNN model provided by a specific embodiment of the present invention;

[0041] Figure 12 It is a comparison schematic diagram of the predicted sample results of the CatBoost model provided by a specific embodiment of the present invention;

[0042] Figure 13 It is a comparison schematic diagram of the predicted sample results of the ANN model provided by a specific embodiment of the present invention;

[0043] Figure 14 It is a comparison schematic diagram of the predicted sample results of the BPNN model provided by a specific embodiment of the present invention;

[0044] Figure 15 It is a schematic diagram of the data feature importance index provided by a specific embodiment of the present invention;

[0045] Figure 16 It is a distribution schematic diagram of the feature SHAP values provided by a specific embodiment of the present invention;

[0046] Figure 17 It is a display schematic diagram of the graphical user interface provided by a specific embodiment of the present invention;

[0047] Figure 18 It is a display schematic diagram of the prediction result pop-up window of the graphical user interface provided by a specific embodiment of the present invention. Detailed implementation manners

[0048] The following further elaborates on the present invention in detail with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0049] First of all, it should be noted that rapid urbanization and industrialization have disrupted the entire ecosystem and significantly affected the water environment, leading to the emergence of harmful pollutants, including heavy metal (HM) ions. Lead is a major source of heavy metal pollution, and a large amount of high-concentration lead-containing wastewater is generated from various industrial processes. Polluted wastewater can easily accumulate in the environment and be absorbed by organisms, resulting in harmful effects. Among various treatment technologies, adsorption is considered a robust, economical, non-toxic and efficient method that can rapidly remove various heavy metal pollutants in water. Biochar has abundant surface functional groups, a high surface area, a developed porous structure and excellent tunability, and has been proven effective in removing various organic and inorganic pollutants in water bodies.

[0050] The rapid development of information technology has given rise to breakthroughs in academic and technological fields. In the field of environmental research, machine learning algorithms have shown significant potential in multiple applications, including environmental quality modeling and prediction, catalyst and adsorbent design, and process optimization, to improve research efficiency and reveal hidden patterns between variables, thereby facilitating decision-making, resource allocation, and environmental management. In recent years, machine learning has become increasingly popular in big data mining and analysis for predicting unknown data and revealing hidden information, such as feature correlations and contributions in experimental data.

[0051] Based on this, with reference to Figure 1 and Figure 3 , the present invention provides a prediction method for lead adsorption by biochar based on a machine learning model, and the method comprises the following steps:

[0052] S100. Obtain a dataset of biochar lead adsorption and determine the input variables and output variables of the dataset;

[0053] Specifically, the input variables of the dataset include preparation conditions, biochar properties, and adsorption conditions. The preparation conditions include pyrolysis temperature and pyrolysis time. The biochar properties include pH value, surface area, pore volume, average pore diameter, C(%), H(%), O(%), N(%), H / C, O / C, (O + N) / C, aromaticity index, and double bond equivalent. The adsorption conditions include initial Pb(II) concentration, biochar dosage, solution pH value, adsorption time, stirring speed, and adsorption temperature. The output variable of the dataset includes the adsorption capacity of biochar for lead.

[0054] In this embodiment, the original dataset (1120 data points) for predicting the adsorption capacity is obtained from the experimental results of published literature and contains the adsorption data of 83 biochars produced by pyrolysis equipment such as tube furnaces, muffle furnaces, and microwave ovens. The input data is directly extracted from the tables in the published literature, while the output data (adsorption capacity) is mainly obtained by extracting graphic information through WebPlot Digitizer software. The input variables include 21 parameters, which can be divided into preparation conditions, biochar properties, and adsorption conditions. The preparation conditions of biochar include pyrolysis temperature (°C) and pyrolysis time (min); the properties of biochar include pH value, surface area (m 2 / g), pore volume (cm 3 / g), average pore diameter (nm), C (%), H (%), O (%), N (%), as well as H / C, O / C, (O+N) / C, aromaticity index (AI), and double bond equivalent (DBE); adsorption conditions include initial Cd(II) concentration (mg / L), biochar dosage (g / L), solution pH value, adsorption time (min), stirring speed (rpm), and adsorption temperature (°C); the key output variable is the adsorption capacity of biochar for lead (mg / g).

[0055] S200. Visualize and preprocess the dataset of biochar adsorption of lead to obtain the preprocessed dataset;

[0056] Specifically, visualize the linear dependence between any two variables in the dataset of biochar adsorption of lead to obtain the visualized dataset; obtain the average value of the visualized dataset and replace the missing values in the visualized dataset to obtain the replaced dataset; perform standardization and partitioning on the replaced dataset in sequence to obtain the preprocessed dataset.

[0057] In this embodiment, first, Python is used as a programming tool, and third-party libraries provided by the community such as Numpy, Pandas, seaborn, SHAP, Matplotlib, etc. are utilized to perform visualization, preprocessing, and analysis according to the characteristics of the original dataset. Draw box plots as shown in Figure 4 and heat maps as shown in Figure 5 for data visualization of variables and identification of relevant input variables. Among them, as shown in Figure 4 , it can be seen that the correlation coefficients between O / C and (O+N) / C, and H and DBE are 1 and -0.96 respectively. To avoid collinearity, and the absolute value of the correlation coefficient between (O+N) / C and the adsorption capacity is greater than that of O / C, so O / C is deleted. Similarly, the feature DBE is deleted. Therefore, the remaining 19 features are used as input variables.

[0058] Further perform data preprocessing. There are a small number of missing values in pore volume, average pore diameter, and stirring speed in the dataset. To enhance the stability of the model, these missing values are replaced with the average values of the corresponding parameters in the entire dataset. Use data standardization to eliminate the differences in dimensions and magnitudes of various feature variables, so that they enter a unified range and are on the same scale. Finally, divide the preprocessed input data into a training set and a test set in a ratio of 8:2.

[0059] S300. Introduce grid search technology and early stopping mechanism, and perform hyperparameter optimization and training on the machine learning model through the preprocessed dataset to screen and obtain the optimal machine learning model;

[0060] Specifically, obtain machine learning models, including CatBoost model, ANN model, and BPNN model; introduce grid search technology and early stopping mechanism, first find the best combination of hyperparameters, and then perform iterative training on the CatBoost model with the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and output the trained CatBoost model; introduce grid search technology and early stopping mechanism, first find the best combination of hyperparameters, and then perform iterative training on the ANN model with the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and output the trained ANN model; introduce grid search technology and early stopping mechanism, first find the best combination of hyperparameters, and then perform iterative training on the BPNN model with the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and output the trained BPNN model; respectively obtain the coefficient of determination and root mean square error values of the test sets of the trained CatBoost model, the trained ANN model, and the trained BPNN model and perform comparison processing, and screen to obtain the trained CatBoost model as the optimal machine learning model.

[0061] In this embodiment, for the development and establishment of the machine learning model, CatBoost, ANN, and BPNN in the ML (machine learning) model are selected to predict the adsorption capacity of biochar for Pb(II) in water, and the parameters related to the ML model algorithm are optimized to find the best set of hyperparameters. Specifically, the optimal hyperparameters are obtained by using the five-fold cross-validation method on the training set through grid search technology, and the evaluation index is specified as the negative mean square error. The structure of the CatBoost neural network model is as Figure 6 shown, the structure of the ANN neural network model is as Figure 7 shown, and the structure of the BPNN neural network model is as Figure 8As shown below. The fine-tuned hyperparameters include depth (3, 5, 6, 7, 9), learning_rate (0.01, 0.05, 0.1, 0.2, 0.3), and l2_leaf_reg (1, 3, 4, 5, 7) of the CatBoost model; hidden_layer_sizes [(50,), (100,), (200,), (50, 50), (100, 50)], activation ('relu', 'tanh', 'logistic', 'identity'), alpha (0.0001, 0.001, 0.01, 0.1, 1.0), and solver ('adam', 'lbfgs','sgd') of the BPNN model; activation ('relu', 'tanh', 'logistic', 'identity'), hidden_layer_sizes [(128, 32), (64, 32), (32, 16), (64, 64, 32)], solver ('adam', 'lbfgs','sgd'), and learning_rate ('constant', 'adaptive', 'invscaling') of the ANN model. The best hyperparameters of the ML models are summarized in Table 1.

[0062] Table 1 Data table of the best hyperparameters of three ML models

[0063]

[0064]

[0065] In addition, in the CatBoost model, the maximum number of iterations for model training is set to 1000, and the early stopping mechanism is enabled. When the performance of the validation set does not improve in 50 consecutive iterations, the training will stop; in the ANN model, the maximum number of iterations for training is set to 1000, the early stopping mechanism is enabled, and 10% of the training data is randomly selected during the training process to verify the model performance to evaluate whether early stopping is needed; in the BPNN model, the maximum number of iterations is set to 1000, the early stopping mechanism is enabled, and 10% of the training data is randomly selected during the training process to verify the model performance to evaluate whether early stopping is needed. When the performance of the validation set does not improve in 10 consecutive iterations, early stopping will be triggered.

[0066] Furthermore, model evaluation is carried out to evaluate the R 2 and RMSE of the training set and test set of each model, where the training set R 2 of the CatBoost model is 0.997 and the RMSE is 5.893, and the test set R 2= 0.979, RMSE = 17.310; R of the BPNN model training set 2 = 0.993, RMSE = 9.153, R of the test set 2 = 0.946, RMSE = 27.839; R of the ANN model training set 2 = 0.994, RMSE = 8.351, R of the test set 2 = 0.954, RMSE = 25.750. It can be concluded from the performance on the test set that the CatBoost model has the best performance. The scatter plot of the predicted values and the actual values of the model is as shown in Figure 9 、 Figure 10 and Figure 11 shown, and the comparison chart of the predicted sample results is as shown in Figure 12 、 Figure 13 and Figure 14 shown.

[0067] S400. Based on the optimal machine learning model, perform SHAP feature importance analysis and establish a graphical user interface application to obtain the importance ranking of input features affecting the adsorption capacity prediction and the graphical user interface.

[0068] Specifically, based on the optimal machine learning model, perform SHAP feature importance analysis to obtain the average absolute SHAP value of each input feature; calculate the importance value of each input feature according to the average absolute SHAP value and perform sorting to obtain the feature importance index of the data set; based on the optimal machine learning model, use Tkinter to create a graphical user interface that allows users to input feature values for prediction.

[0069] In the embodiment of the present invention, first perform SHAP feature importance analysis. Based on the constructed CatBoost model, evaluate the contribution of each feature to the model prediction result, and calculate the importance value of each feature according to the average absolute SHAP value. The results are as shown in Figure 15 and Figure 16 , and the importance ranking is as follows: initial lead content > adsorption time > nitrogen content > average pore diameter > biochar pH value > solution pH > biochar dosage > surface area > (O + N) / C > carbon content > AI > hydrogen content > H / C > O > pore volume > pyrolysis time > stirring speed > pyrolysis temperature > adsorption temperature. In addition, the contributions of adsorption conditions, biochar properties, and pyrolysis conditions account for 54.9%, 43.1%, and 2.0% of the overall impact, respectively. Therefore, when developing efficient adsorbents and optimizing the Pb removal adsorption process, attention should be focused on adsorption conditions and biochar properties.

[0070] Further, a graphical user interface (GUI) application is created. Based on the constructed CatBoost model, a graphical user interface (GUI) is created using Tkinter, allowing users to input feature values for prediction. In the graphical user interface, as Figure 17 shown, multiple input boxes are provided for users to input different features. After the users input the features, they click the "Predict Adsorption Capacity" button. The program will standardize the input features and use the trained model for prediction, and display the prediction results in a pop-up window, as Figure 18 shown.

[0071] In summary, the CatBoost prediction model established in the embodiments of the present invention has excellent performance (the R 2 and RMSE of the test set are 0.979 and 17.310 respectively). In addition, the SHAP feature importance analysis reveals that the priority order of the input feature types is: adsorption conditions (54.9%) > biochar properties (43.1%) > preparation conditions (2.0%). The adsorption design and experimental conditions are further optimized through the SHAP feature importance analysis. A graphical user interface (GUI) application is created, which is convenient for users to use. The present invention provides a beneficial modeling and optimization strategy for the adsorption capacity of Pb(II) on biochar from the perspective of machine learning, and provides valuable help for the prediction modeling research of pollutant adsorption in the water environment.

[0072] Referring to Figure 2 , a prediction system for biochar to adsorb lead based on a machine learning model includes:

[0073] The first module 201 is used to obtain the dataset of biochar adsorbing lead and determine the input variables and output variables of the dataset;

[0074] The second module 202 is used to perform data visualization and preprocessing on the dataset of biochar adsorbing lead to obtain the preprocessed dataset;

[0075] The third module 203 is used to introduce the grid search technique and the early stopping mechanism, and optimize and train the machine learning model through the preprocessed dataset to screen out the optimal machine learning model;

[0076] The fourth module 204 is used to perform SHAP feature importance analysis and establish a graphical user interface application based on the optimal machine learning model to obtain the importance ranking of the input features affecting the adsorption capacity prediction and the graphical user interface.

[0077] The content in the above method embodiments is applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0078] The above is a specific description of the preferred embodiments of the present invention. However, the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A prediction method for biochar adsorption of lead based on a machine learning model, characterized in that, It includes the following steps: Obtain the dataset of biochar adsorbing lead and determine the input variables and output variables of the dataset; Perform data visualization and preprocessing on the dataset of biochar adsorbing lead to obtain the preprocessed dataset; Introduce the grid search technique and early stopping mechanism, and optimize and train the hyperparameters of the machine learning model through the preprocessed dataset to screen out the optimal machine learning model; Based on the optimal machine learning model, conduct SHAP feature importance analysis and establish a graphical user interface application program to obtain the importance ranking of input features affecting adsorption capacity prediction and the graphical user interface.

2. The prediction method for biochar to adsorb lead based on a machine learning model according to claim 1, wherein The input variables of the dataset include preparation conditions, biochar properties, and adsorption conditions. The preparation conditions include pyrolysis temperature and pyrolysis time. The biochar properties include pH value, surface area, pore volume, average pore diameter, C(%), H(%), O(%), N(%), H / C, O / C, (O+N) / C, aromaticity index, and double bond equivalent. The adsorption conditions include initial Pb(II) concentration, biochar dosage, solution pH value, adsorption time, stirring speed, and adsorption temperature. The output variable of the dataset is the adsorption capacity of biochar for lead.

3. The prediction method for biochar to adsorb lead based on a machine learning model according to claim 2, wherein The step of performing data visualization and preprocessing on the dataset of biochar adsorbing lead to obtain the preprocessed dataset specifically includes: Draw box plots for dataset visualization, judge whether there are redundant features among input features through Pearson correlation test and delete the redundant features to obtain the dataset after deleting redundant features; Obtain the average value of the dataset after deleting redundant features and replace the missing values in the dataset after deleting redundant features to obtain the replaced dataset; Perform standardization and partitioning processing on the replaced dataset in sequence to obtain the preprocessed dataset.

4. The prediction method for biochar to adsorb lead based on a machine learning model according to claim 3, characterized in that, The step of introducing the grid search technique and early stopping mechanism, optimizing and training the hyperparameters of the machine learning model through the preprocessed dataset to screen out the optimal machine learning model specifically includes: Obtain machine learning models, which include CatBoost model, ANN model, and BPNN model; Introduce the grid search technique and early stopping mechanism, first find the best combination of hyperparameters, and then perform iterative training on the CatBoost model through the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and output the trained CatBoost model; Introduce the grid search technique and early stopping mechanism, first find the best combination of hyperparameters, and then perform iterative training on the ANN model through the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and output the trained ANN model; Introduce the grid search technique and early stopping mechanism, first find the best combination of hyperparameters, and then perform iterative training on the BPNN model through the preprocessed dataset until the early stopping condition or the preset number of iterations is met, and output the trained BPNN model; Respectively obtain the coefficient of determination and root mean square error values of the test sets of the trained CatBoost model, the trained ANN model, and the trained BPNN model, and perform comparison processing to select the trained CatBoost model as the optimal machine learning model.

5. The prediction method for biochar to adsorb lead based on a machine learning model according to claim 4, characterized in that, The step of performing SHAP feature importance analysis and establishing a graphical user interface application based on the optimal machine learning model to obtain the importance ranking of input features affecting adsorption capacity prediction and the graphical user interface specifically includes: Perform SHAP feature importance analysis based on the optimal machine learning model to obtain the average absolute SHAP value of each input feature; Calculate the importance value of each input feature according to the average absolute SHAP value and perform sorting processing to obtain the feature importance index of the data set; Based on the optimal machine learning model, use Tkinter to create a graphical user interface that allows users to input feature values for prediction.

6. A prediction system for biochar adsorption of lead based on a machine learning model, characterized in that, It includes the following modules: The first module is used to obtain the data set of biochar adsorbing lead and determine the input variables and output variables of the data set; The second module is used to perform data visualization and preprocessing on the data set of biochar adsorbing lead to obtain the preprocessed data set; The third module is used to introduce grid search technology and early stopping mechanism, optimize and train the machine learning model through the preprocessed data set, and select the optimal machine learning model; The fourth module is used to perform SHAP feature importance analysis and establish a graphical user interface application based on the optimal machine learning model to obtain the importance ranking of input features affecting adsorption capacity prediction and the graphical user interface.