Method for predicting extraction efficiency of soil heavy metal eluting agent based on machine learning
Through the machine learning-based eluent extraction efficiency prediction model, the resource waste problem in eluent selection and condition optimization in soil remediation is solved, and efficient and accurate soil heavy metal extraction rate prediction is achieved, reducing soil treatment costs.
Patent Information
- Application Number
- CN202510656903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-19
AI Technical Summary
Existing soil remediation technologies consume a lot of time and resources in terms of eluent selection and condition optimization, and lack systematicity and predictability, resulting in limited heavy metal leaching efficiency.
Based on the machine learning algorithm, soil heavy metal extraction rate data was used to optimize the machine learning algorithm through grid search, and a prediction model for eluent extraction efficiency was established, including data collection, preprocessing and model building, and the extreme gradient boosting decision tree algorithm was used for prediction.
It improves soil remediation efficiency, reduces resource consumption, has high prediction accuracy, is easy to operate, is suitable for different regions and types of eluents, and reduces soil remediation costs.
Smart Images

Figure CN120671865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the extraction efficiency of a soil heavy metal leaching agent based on machine learning, and specifically belongs to the technical field of soil remediation. Background Art
[0002] With the acceleration of industrialization and urbanization, the problem of heavy metal pollution in soil is becoming increasingly serious, posing a major threat to the ecological environment, agricultural production and human health. Heavy metals such as lead (Pb), cadmium (Cd), zinc (Zn), etc., due to their non-degradability and bioaccumulation, will continue to accumulate in the soil, destroying soil structure, reducing soil fertility, and being transmitted step by step through the food chain, ultimately endangering human health.
[0003] A variety of soil remediation technologies exist, including solidification / stabilization (S / S), chemical leaching, thermal desorption, electrokinetic remediation, and bioremediation. Chemical leaching is considered one of the most promising remediation methods due to its ability to effectively remove the total amount of heavy metals. Chemical leaching enhances the solubility and mobility of heavy metals by adding inorganic / organic acids, chelating agents, or surfactants, thereby achieving heavy metal removal. For example, Xu et al. tested nearly 30 leaching agents and achieved extraction rates of 85.06% for Cd and 56.04% for Cu, respectively. Furthermore, Liang et al., studying paddy soil near a lead-zinc smelter, determined, through extensive experiments, that optimal removal rates for Cd and Pb were 60.7% and 61.8%, respectively. Traditional leaching remediation techniques often require significant time and resources for leaching agent selection and condition optimization, and their efficiency is limited. Furthermore, existing technology development research often relies on empirical evidence or single-factor analysis, lacking systematicity and predictability, making the results difficult to generalize and optimize. Therefore, an efficient method is urgently needed to predict the leaching efficiency of soil heavy metals by eluents in order to improve remediation efficiency and reduce resource consumption and experimental costs. Summary of the Invention
[0004] In view of the above situation, in order to overcome the limitations of the existing technology in the extraction efficiency of heavy metals in soil using eluents and further improve the efficiency of soil remediation, the purpose of the present invention is to provide a method based on a machine learning algorithm to predict the extraction rate of heavy metals in soil using soil heavy metal extraction rate data.
[0005] The present invention provides a method for predicting the extraction efficiency of soil heavy metal eluents based on machine learning. The method collects soil heavy metal extraction rate data in a database as a data set for a machine learning model. After preprocessing, the method optimizes the machine learning algorithm using grid search to establish an eluent extraction efficiency prediction model to predict the extraction rate of the heavy metal to be measured. The method specifically comprises the following steps: Step 1: Data Collection Using soil, heavy metals, EDTA, chelating agents, or elution agents as keywords, a literature search was conducted in the Web of Science, CNKI, VIP, and Wanfang databases. Software was used to collect soil heavy metal extraction rate data on soil physical and chemical properties, elution agent types, experimental conditions, and heavy metal types from the retrieved literature icons. A sample database was constructed, and a dataset of soil heavy metal leaching rate data was compiled. Step 2: Dataset preprocessing The units in the dataset samples were converted to be consistent, the extraction pH was calculated, missing data were deleted, and Z-score normalization was performed. The dataset samples with missing features were deleted, and the Z-score normalization method was applied to convert each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1. Its linear scale was normalized to the range of (0, 1). The dataset was randomly divided into two parts, a training set and a test set, with a ratio of 80% and 20% respectively. Step 3: Establishment of eluent extraction efficiency prediction model Utilize grid search to optimize machine learning algorithms and establish a model to predict eluent extraction efficiency. The specific machine learning algorithms used include k-nearest neighbor, random forest, extreme gradient boosting, and artificial neural network. Determine the hyperparameters of the machine learning algorithm. Hyperparameters are determined on the training set and the model is optimized using grid search and ten-fold cross validation. An optimized machine learning algorithm was used to establish a soil heavy metal extraction efficiency prediction model on the training set, and the test set was used to determine the reliability of the soil heavy metal extraction efficiency model. Evaluation indicators included R², MAE, MSE, and RMSE. Step 4: Predict the extraction rate of heavy metals to be tested Determine the data to be predicted for the soil heavy metal leaching agent extraction rate, and use the trained soil heavy metal leaching efficiency model to make predictions.
[0006] The soil physical and chemical properties include pH, cation exchange content, organic matter content, clay percentage, silt percentage and sand percentage.
[0007] The experimental conditions include eluent concentration, liquid-to-solid ratio, extraction pH and reaction time.
[0008] The software is GetData Graph Digitizer 2.25.
[0009] The R², MAE, MSE, and RMSE are calculated using the sklearn.metrics.r2_score, sklearn.metrics.mean_squared_error, and sklearn.metrics.mean_absolute_error functions, respectively.
[0010] The present invention has the following benefits: 1. The method, based on soil extraction efficiency data, uses machine learning to predict extraction efficiency and is applicable to regions with varying heavy metal concentrations and different eluent types. Compared to traditional field sampling followed by laboratory leaching, the present method significantly saves manpower and material resources, resulting in less workload, lower costs, and higher efficiency.
[0011] 2. The model established by the method of the present invention is easy to operate, can optimize experimental parameters, improve soil remediation efficiency, and predict the extraction rate of heavy metals in soil with high accuracy and good reliability. It is superior to traditional methods in terms of computational efficiency and analytical capabilities, and can provide a basis for reducing soil remediation costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 : A schematic diagram of a process for predicting soil heavy metal extraction rates using machine learning provided by the present invention; Figure 2 : Distribution diagram of soil physical and chemical property loss ratios modeled by the present invention; Figure 3 : Visualization of the observed and predicted values of the present invention on the training set and test set; Figure 4 : Comparison chart of observed values and predicted values on the test set of the present invention. DETAILED DESCRIPTION
[0013] Example 1 like Figure 1 As shown, based on the heavy metal extraction rate data, machine learning is used to predict the heavy metal extraction efficiency, including the following steps: Step 1 (S1): Data collection A literature search was conducted in Web of Science, CNKI, VIP, and Wanfang databases using the keyword [TS=“soil” OR “heavy metal” OR “EDTA” OR “chelating agent” OR “elution agent”], and 114 articles published from 1998 to 2024 were screened.
[0014] Known soil physical and chemical properties, elution agent types, experimental conditions, and heavy metal species were collected from a database as a dataset for the machine learning model. Specifically, GetData Graph Digitizer 2.25 software was used to extract data from retrieved literature graphs and tables. A dataset containing the leaching rates of four heavy metals—lead (Pb), cadmium (Cd), copper (Cu), and zinc (Zn)—was constructed. This dataset included 6,464 sample points, 158 soil sampling sites, and 87 elution agent types.
[0015] Figure 2 The distribution ratio of missing data is shown, with missing data for the CEC, OM, Slit, Clay, and Sand features marked in black. Numerical data are normalized to the (0, 1) interval, and the magnitude of each column is represented using a grayscale image: light colors represent smaller values, and dark colors represent larger values.
[0016] Step 2 (S2): Dataset preprocessing The units in the dataset samples were converted to consistent and extracted pH calculations, and the dataset samples containing missing features were deleted; the Z-score normalization method was applied to the dataset, and each feature was converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, and its linear scale was normalized to the range of (0, 1); the dataset was randomly divided into a training set and a test set with a ratio of 80% and 20% respectively.
[0017] The Z-score standardization calculation formula is as follows: ,in: is the original eigenvalue, is the characteristic mean, is the standard deviation of the feature.
[0018] Step 3 (S3): Establishment of eluent extraction efficiency prediction model The Extreme Gradient Boosting (XGBoost) model was selected as the prediction model for soil heavy metal extraction efficiency. XGBoost integrates multiple weak learners to gradually optimize the loss function, which is characterized by high efficiency and high accuracy.
[0019] To determine the optimal hyperparameters, we used a grid search method combined with ten-fold cross-validation to systematically explore multiple hyperparameter combinations and evaluate the performance of each. This method resulted in the following optimal hyperparameter configuration: max_depth = 20, n_estimators = 900, and min_child_weight = 10.
[0020] To evaluate the model performance, four evaluation indicators, R², RMSE, MSE, and MAE, are used; These metrics are calculated by calling the sklearn.metrics.r2_score, sklearn.metrics.mean_squared_error, and sklearn.metrics.mean_absolute_error functions respectively.
[0021] The calculation formulas for the four evaluation indicators R², RMSE, MSE and MAE are as follows: and ; in: represents the observed value of heavy metal leaching rate, represents the model prediction value, represents the average of the actual observations, and n represents the number of samples.
[0022] After optimization in this embodiment, the maximum depth of the extreme gradient boosting decision tree model is 20, the number of weak learners is 900, and the minimum sum of subnode weights is 10. The optimal extreme gradient boosting decision tree model is trained using the entire training set, and the performance of the gradient boosting decision tree model on the training set and the test set is tested. Figure 3 As shown, the extreme gradient boosting decision tree model predicts heavy metal extraction rates with R² of 1 and 0.9 for the training and test sets, respectively, demonstrating excellent accuracy for both. The extreme gradient boosting decision tree model that achieves the best prediction across the entire training set is used as the heavy metal extraction rate prediction model of the present invention.
[0023] Step 4 (S4), predicting the extraction rate of the heavy metal to be tested The optimal hyperparameter configuration determined by the grid search method combined with ten-fold cross validation was used to build the best XGBoost model. Subsequently, the standardized test set was input into the trained model to predict the heavy metal extraction rate of each test sample and record the prediction results. Figure 4 As shown in the figure, the R² between the model's predicted soil heavy metal extraction rates and the experimental values on the test set reached 0.9. The close proximity between the predicted and true values demonstrates the feasibility of using machine learning to predict heavy metal extraction efficiency using soil physical and chemical properties and experimental conditions as factors influencing heavy metal extraction rates.
Claims
1. A method for predicting the extraction efficiency of soil heavy metal eluents based on machine learning, characterized by: The method collects soil heavy metal extraction rate data from a database as a data set for a machine learning model. After preprocessing, the method uses grid search to optimize the machine learning algorithm, establishes a leaching agent extraction efficiency prediction model, and predicts the extraction rate of the heavy metal to be tested. The method specifically includes the following steps: Step 1: Data Collection Using soil, heavy metals, EDTA, chelating agents, or elution agents as keywords, a literature search was conducted in the Web of Science, CNKI, VIP, and Wanfang databases. Software was used to collect soil heavy metal extraction rate data on soil physical and chemical properties, elution agent types, experimental conditions, and heavy metal types from the retrieved literature icons. A sample database was constructed, and a dataset of soil heavy metal leaching rate data was compiled. Step 2: Dataset preprocessing The units in the dataset samples were converted to be consistent, the extraction pH was calculated, missing data were deleted, and Z-score normalization was performed. The dataset samples with missing features were deleted, and the Z-score normalization method was applied to convert each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1. Its linear scale was normalized to the range of (0, 1). The dataset was randomly divided into two parts, a training set and a test set, with a ratio of 80% and 20% respectively. Step 3: Establishment of eluent extraction efficiency prediction model Utilize grid search to optimize machine learning algorithms and establish a model to predict eluent extraction efficiency. The specific machine learning algorithms used include k-nearest neighbor, random forest, extreme gradient boosting, and artificial neural network. Determine the hyperparameters of the machine learning algorithm. Hyperparameters are determined on the training set and the model is optimized using grid search and ten-fold cross validation. An optimized machine learning algorithm was used to establish a soil heavy metal extraction efficiency prediction model on the training set, and the test set was used to determine the reliability of the soil heavy metal extraction efficiency model. Evaluation indicators included R², MAE, MSE, and RMSE. Step 4: Predict the extraction rate of heavy metals to be tested Determine the data to be predicted for the soil heavy metal leaching agent extraction rate, and use the trained soil heavy metal leaching efficiency model to make predictions.
2. The method for predicting the extraction efficiency of heavy metals from soil based on machine learning according to claim 1, wherein: The soil physical and chemical properties include pH, cation exchange content, organic matter content, clay percentage, silt percentage and sand percentage.
3. The method for predicting soil heavy metal extraction efficiency based on machine learning according to claim 1, characterized in that: The experimental conditions include eluent concentration, liquid-to-solid ratio, extraction pH and reaction time.
4. The method for predicting soil heavy metal extraction efficiency based on machine learning according to claim 1, wherein: The software is GetData Graph Digitizer 2.
25.
5. The method for predicting soil heavy metal extraction efficiency based on machine learning according to claim 1, characterized in that: The R², MAE, MSE, and RMSE are calculated using the sklearn.metrics.r2_score, sklearn.metrics.mean_squared_error, and sklearn.metrics.mean_absolute_error functions, respectively.