Method for identifying LDHs passivated soil heavy metal main control factors and boundary conditions based on integrated machine learning
By integrating machine learning methods to identify the main controlling factors and boundary conditions for LDHs passivation of heavy metals in soil, this approach solves the problems of heterogeneity in passivation effects and poor lateral comparability in existing technologies. It enables robust parameter screening and engineering design in highly heterogeneous data, supporting the analysis of multi-metal and complex contaminated soils.
Patent Information
- Application Number
- CN202511748466.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-03-03
AI Technical Summary
Existing research on LDHs passivation of heavy metals in soil suffers from problems such as highly heterogeneous passivation effects, poor cross-sectional comparisons between studies, over-reliance on repeated experiments, and a lack of systematic quantification of the main controlling factors and parameter boundaries, leading to difficulties in engineering promotion.
An integrated machine learning approach was adopted, which involved multi-source data collection and processing, combined with feature subset determination, machine learning model selection, and hyperparameter optimization, to identify the main controlling factors and boundary conditions of heavy metals in LDHs passivated soil. This included the application of models such as kernel extreme learning machine and extreme gradient boosting algorithm.
It enables robust determination of the optimal controlling factors and boundary conditions for LDHs passivation of heavy metals in soils with highly heterogeneous data, significantly shortening the screening cycle for materials and engineering design, and supporting the analysis of multimetallic and complex contaminated soils.
Smart Images

Figure CN121598008A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of soil pollution remediation technology, specifically relating to a method for identifying the main controlling factors and boundary conditions of heavy metals in layered bimetallic hydroxide passivated soil. Background Technology
[0002] Layered bimetallic hydroxides (LDHs) are a class of metal hydroxide minerals with high specific surface area and excellent ion exchange and adsorption properties. Because their layer composition can be artificially controlled to regulate their metal passivation properties, they are currently widely used in soil heavy metal pollution remediation projects. Existing research mainly focuses on the kinetics, thermodynamics and cultivation of LDHs under laboratory conditions. However, the heavy metal passivation effect of LDHs in field trials often differs significantly from that in laboratory settings. The main reason is that real field environmental conditions are complex and variable, and the factors that dominate the passivation performance in different soil environments are not clear. To maximize the heavy metal passivation effect of LDHs in the field, the main challenges are: (1) the high heterogeneity of the soil environment itself, but the preparation of LDHs does not take into account the variability of its environment, resulting in significant differences in passivation performance across studies, media and conditions; (2) there are inconsistencies in the variable scope and indicators among existing cases, making it difficult to conduct horizontal comparisons of passivation performance; (3) there is a lack of systematic quantification of the main control factors and parameter boundaries, and engineering promotion relies on repeated experiments. Therefore, establishing a method and system that can determine the optimal controlling factors and boundary conditions for LDH passivation performance based on current highly heterogeneous multi-source data through integrated machine learning is an important way to improve the passivation efficiency of heavy metals under different soil environmental conditions. Summary of the Invention
[0003] The purpose of this invention is to provide a method for identifying the main controlling factors and boundary conditions of heavy metals in LDHs passivated soil based on integrated machine learning, so as to solve the problems of high heterogeneity of passivation effect, poor cross-comparison between studies, and excessive reliance on repeated experiments in existing research and practice on LDHs passivated soil heavy metals.
[0004] The method for identifying the main controlling factors and boundary conditions of heavy metals in LDH-passivated soil based on ensemble machine learning provided by this invention mainly includes multi-source data collection and processing, and the construction of ensemble machine learning methods; the process is as follows: Figure 1 As shown, the specific steps are as follows:
[0005] The first step is to construct a multi-source dataset: systematically search and screen publicly available literature globally to construct a multi-source dataset that includes variables such as material synthesis methods, material structural composition information, environmental factors, and passivation indices.
[0006] The second step is to standardize the multi-source dataset; the passivation index in the multi-source dataset obtained in the first step is standardized, the effect size RR, variance and 95% confidence interval are calculated, and heterogeneity test and publication bias assessment are performed.
[0007] The third step is to complete the standardized dataset. Specifically, a strategy combining group imputation and multiple chain equation imputation is used to complete the missing variables in the dataset.
[0008] The fourth step is to determine the feature subset for the completed dataset. Specifically, the feature subset is determined by comprehensively screening the variables based on the correlation between variables and algorithms such as the "fully correlated" feature selection algorithm (Boruta) based on the importance test of random forest, taking into account both the importance of variables and redundancy control.
[0009] The fifth step is to select machine learning models based on feature subsets. Specifically, at least one model from other machine learning models, including Kernel Extreme Learning Machine (KELM), Extreme Gradient Boosting Algorithm (XGBoost), Random Forest, and Lightweight Gradient Boosting Machine (LightGBM), is used for modeling. The optimal hyperparameters and model are selected through hyperparameter optimization and k-fold cross-validation.
[0010] The sixth step, based on the obtained optimal model and dataset, combines methods such as the importance of permutation variables and multiple test correction or general importance analysis corresponding to other models, quantile regression, local weighted regression smoothing (LOESS), weighted kernel density estimation, etc., to output the main control factor and boundary conditions of the heavy metal passivation effect, and gives the effective conditions;
[0011] The seventh step involves identifying and weighing the common and differential controlling factors and boundary conditions obtained in the sixth step for complex pollution of two or more heavy metals using multi-objective analysis methods, including but not limited to principal component analysis and cluster analysis, which require specific analysis in conjunction with the data.
[0012] The soil heavy metals involved in this invention include As, Cd, Cr, Cu, Hg, Ni, Pb, and Zn.
[0013] Furthermore:
[0014] The first step is to construct a multi-source dataset. The systematic global retrieval and screening of publicly available literature involves selecting more relevant and compliant literature based on inclusion and exclusion criteria after the initial retrieval. Specifically:
[0015] Inclusion criteria include: (1) the literature must include a control treatment: experiments with added LDHs and blank control experiments without added LDHs; (2) the literature must include a quantitative value of the passivation effect: passivation rate, passivation amount, or residual amount after passivation; (3) the mean, standard deviation (SD) or error (SE), and the number of repetitions for the control and experimental groups are obtained from text, tables, or digital charts; (4) the full text of the literature is available.
[0016] Exclusion criteria include: (1) literature where experimental data is unavailable, or data is incomplete or cannot be converted; (2) literature where the research subject is only aqueous solution or mechanism study; (3) research type is non-randomized controlled trial, including review literature, meta-analysis, conference; (4) duplicate publication or poor quality literature.
[0017] The second step, standardization of multi-source datasets, includes unifying indicators and selecting factors. Specifically:
[0018] Considering that the passivation effect of materials is affected by multiple factors, during the data collection and processing, the data volume and the correlation between variables were comprehensively considered. The factors affecting the passivation effect were summarized into three categories: synthetic factors, material properties, and environmental factors. In order to facilitate data processing and analysis in a programming language environment, the variables were simultaneously coded in English and the units of continuous variables were unified. In addition, the passivation amount, passivation residue, and passivation rate were uniformly converted into the passivation rate, i.e., the effect size RR.
[0019] The removal rate, residual amount, and material dosage were standardized, and the relative response rate (RR) was used as the effect size. Variance and 95% confidence intervals were derived based on the statistics of the treatment and control groups. Specific variables included: synthesis method, intercalated ions, loading material, divalent cations, trivalent cations, cation molar ratio, interlayer anions, and specific surface area (m²). 2 / g), pore volume (cm³) 3 / g), average pore size (nm), interlayer spacing (nm), experimental subject, material concentration (%), initial pollution concentration, initial pH, experimental temperature, competing ions, reaction time (min), soil moisture (%).
[0020] The third step involves completing the standardized dataset, including dataset evaluation and imputation. Specifically:
[0021] (1) Heterogeneity and publication bias: The heterogeneity test statistic Q, the heterogeneity ratio index I², and the heterogeneity variance Tau² indicate the heterogeneity of the data; publication bias is evaluated to ensure the robustness of the conclusions.
[0022] (2) Missing values: First, statistical imputation is performed by grouping by metal type and medium type. Then, the continuous and classification features are completed by using the multiple chain equation imputation method. The results of multiple imputation are summarized.
[0023] The fourth step is to determine the feature subset for the completed dataset. Specifically, the Boruta algorithm is used to screen variables that have a significant effect on the response value. Then, based on the results of Spearman correlation analysis between numerical variables, one-way ANOVA between numerical and categorical variables, and chi-square test based on contingency tables between categorical variables, and combined with the actual meaning of the variables, highly correlated or highly collinear variables are removed if they are meaningful or have related properties.
[0024] The fifth step is to select machine learning models based on feature subsets, using the model R from the test set (i.e., 20% of the dataset). 2 To determine the best model, 10-fold cross-validation was used during parameter tuning and training.
[0025] In step six, based on the obtained optimal model, the impact of each variable on model performance is evaluated to obtain comparable "group permutation importance," and the p-value of the one-sided permutation test and the q-value after multiple test correction are given; important and significant variables are selected as the main controlling factors for the passivation of heavy metals in soil by LDHs materials; among them,
[0026] (1) Grouping strategy based on importance of grouping by permutation:
[0027] a. Grouping principle: Numerical features are grouped independently by "column"; Categorical features are aggregated by "original factor" and all its unique columns are treated as a whole group for intervention, thereby evaluating the overall contribution of the original factor at the group level and ensuring the comparability of different types of variables.
[0028] b. Evaluation dataset: Baseline performance is calculated using independent validation / test sets to avoid optimistic bias on parameter tuning data;
[0029] c. Importance metric: based on "loss increment" "Characteristic variables (groups) of importance; definition" Loss_base is the baseline loss, and Loss_perm is the loss after permutation;
[0030] d. Permutation method: Randomly rearrange all columns of a given variable group simultaneously, repeating nsim times (defined as the number of repetitions as nsim), calculating each time. Group importance takes all The mean, and stability is characterized by standard deviation;
[0031] (2) Significance determination of one-sided permutation test:
[0032] a. Null hypothesis H0: This variable (group) makes no positive contribution to the model performance, i.e. The expected value is ≤ 0;
[0033] b. Calculation of one-sided p-value: If the p-value is smaller, the evidence to reject H0 is stronger, indicating that the loss increases significantly after the replacement and the variable does indeed contribute.
[0034] (3) Multiple test correction (BH-FDR):
[0035] a. Apply the Benjamini-Hochberg method to the one-sided permutation test p-values of all variables for FDR control to obtain the corrected q-values;
[0036] b. Significance threshold and labeling: Statistical significance is determined by q < 0.05;
[0037] Step 6: Determining the boundary conditions for heavy metals in the soil, specifically:
[0038] (1) For numerical variables, the candidate interval generation method is as follows:
[0039] Threshold distribution method: In samples where RR≥t, take the quartile interval [Q25,Q75] of the independent variable (representing the 25th and 75th percentiles of the value), and generate the interval using the highest threshold t;
[0040] LOESS peak preservation method: fitting Find the peak effect y* of the LOESS fit, and take the value that satisfies the condition. The interval of independent variables, where α is the peak preservation coefficient;
[0041] Weighted kernel density method: The kernel density of the estimated variable is estimated by using RR as the weight, and the range is taken as the proportion of the central interval of the cumulative probability of the kernel density;
[0042] Then, interval fusion is performed: the final interval is determined in the order of "all intersections → widest pairwise intersections → median aggregation"; finally, the range of the interval and the RR statistic within the interval are output to characterize the optimal range of values for the independent variable and the expected effect.
[0043] (2) The optimal method for determining the values of categorical variables is to calculate the effect and frequency of each category, perform normalized weighting to obtain the comprehensive score, one-way ANOVA, and Tukey test; First, categorical statistics and scoring: for each category c of each categorical variable, calculate: mean_Effect(c), quartile interval of RR RR_Q25 / Q75, sample size n, and the proportion of that category in all samples Frequency(c); after normalization, we get:
[0044] Overall rating: , used for sorting.
[0045] Then, significance and grouping are performed: one-way ANOVA is used to test the overall difference (ANOVA statistic F, significance probability p), and Tukey's significance test results are used for pairwise comparisons.
[0046] Finally, the optimal set is selected: find the letter g* of the category c* with the highest mean, and the candidate set C0 = a group with the same letter as c*; apply a robust threshold to C0: n≥5 and Mean_Effect≥0.8 to obtain C1; if C1 is empty, take the one with the highest comprehensive score in the whole set as a fallback; output C1 in descending order of comprehensive score, which is the best category.
[0047] The present invention achieves the following beneficial effects:
[0048] 1. This invention provides a machine learning method based on highly heterogeneous data integration, which can determine the optimal controlling factors and boundary conditions for the passivation performance of heavy metals in soil by LDHs, thus achieving the goal of directly guiding practice from data collection.
[0049] 2. Compared with traditional methods that rely too much on experiments and experience, this invention can still obtain robust conclusions from multi-source, highly heterogeneous data, and can significantly shorten the parameter screening and control cycle for materials and other experiments, engineering designs, and practical applications.
[0050] 3. This invention supports multi-metal, cross-media scenarios involving As, Cd, Pb, etc., and can be extended to the synergistic analysis of soils contaminated with multiple heavy metals. Attached Figure Description
[0051] Figure 1 This is a flowchart of the method for identifying the main controlling factors and boundary conditions of heavy metals in LDHs passivated soil according to the present invention.
[0052] Figure 2 This is a schematic diagram of the PRISMA data collection process.
[0053] Figure 3 A funnel plot for evaluating bias in data quality testing.
[0054] Figure 4 This is a schematic diagram of the Boruta analysis results for heavy metal As in LDH-passivated soil during the variable screening stage.
[0055] Figure 5 This is a schematic diagram of the Boruta analysis results for LDHs passivated soil heavy metal Cd during the variable screening stage.
[0056] Figure 6Figure showing the comparison of the performance of LDHs passivating heavy metal As in soil on the test set during model training.
[0057] Figure 7 Figure 1 shows the performance comparison of LDHs passivating soil heavy metal Cd on the test set during model training.
[0058] Figure 8 Ranking the importance of relevant variables of heavy metal As in LDHs passivated soil under the SVM_RBF model.
[0059] Figure 9 Ranking the importance of relevant variables related to LDHs passivating soil heavy metal Cd under the Random_Forest model.
[0060] Figure 10 A comprehensive score for the optimal classification variable values for LDHs passivated soil As.
[0061] Figure 11 Multi-method analysis of the optimal range of values for As in LDHs passivated soil.
[0062] Figure 12 A comprehensive score for the optimal categorical variable value of Cd in LDHs-passivated soil.
[0063] Figure 13 Multi-method analysis of the optimal range of values for Cd in LDH-passivated soil (Part 1).
[0064] Figure 14 Multi-method analysis of the optimal range of values for Cd in LDH-passivated soil (Part 2). Detailed Implementation
[0065] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. These embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0066] Example 1
[0067] To collect data more broadly, the data collection in this invention includes, but is not limited to, searches in databases such as Web of Science, Scopus, and CNKI, and must be screened according to the PRISMA process. Variables selected for inclusion in the dataset may include synthetic factors, material properties, environmental factors, pollutant removal indicators, basic literature information, and the latitude and longitude of sampling points.
[0068] To obtain a valid and usable dataset, the different response quantities among studies were uniformly converted from passivation quantity, passivation efficiency, and passivation residual quantity into passivation efficiency. Furthermore, the initial screening of variables mainly relied on the fully correlated feature screening algorithm and correlation analysis. Correlation analysis methods such as Spearman correlation analysis, one-way ANOVA, and Monte Carlo tests were used to determine the correlations.
[0069] In the process of ensemble machine learning, it is necessary to integrate linear models and tree models, including but not limited to: LightGBM, CatBoost, ElasticNet, MARS, RandomForest, XGBoost, GBDT, AdaBoost, SVM_RBF, KNN, GaussianProcess, LinearRegression, and Kriging, and adjust the parameters according to the model performance.
[0070] (a) Data acquisition and variable selection: Effective and standardized datasets are obtained through the following methods:
[0071] 1. Data retrieval, filtering, and dataset construction.
[0072] Web of Science, Scopus, CNKI, see attached Figure 2 The PRISMA process was used for screening, and the following are the search terms for each database: Web of Science search term: (TS=(soil) AND TS=(Layered Double Hydroxides) AND TS=(cadmium or heavy metal)); Scopus search term: (TITLE-ABS-KEY(soil) ANDTITLE-ABS-KEY(ldhs) OR TITLE-ABS-KEY (layered AND double AND hydroxides) ANDTITLE-ABS-KEY (material) AND ALL (cadmium) OR ALL (heavy AND metal) OR ALL (heavy AND ions)); CNKI search term: (Subject=soil) AND (Subject=heavy metal) AND (Subject=layered double hydroxides).
[0073] After obtaining the literature through the initial search, further screening of more relevant and compliant literature was conducted based on the inclusion and exclusion criteria:
[0074] Inclusion criteria include, but are not limited to: (1) The literature must include a control treatment: experiments with added LDHs and blank control experiments without added LDHs; (2) The literature must include a quantitative value of the passivation effect: passivation rate, passivation amount, or residual amount after passivation; (3) The mean, standard deviation (SD) or error (SE) and the number of repetitions of the control group and the experimental group can be obtained from text, tables or digital charts; (4) The full text of the literature can be obtained.
[0075] Exclusion criteria include, but are not limited to: (1) literature where experimental data is unavailable, or data is incomplete or cannot be converted; (2) literature where the research object is only aqueous solution or mechanism research; (3) research type is non-randomized controlled trial, such as review literature, meta-analysis, conference, etc.; (4) literature that is repeatedly published or of poor quality.
[0076] A total of 79 articles were selected, yielding 1,825 relevant data points, including 30 variables used for subsequent modeling (covering basic information, synthesis and modification methods, material properties, environmental factors, and various passivation indices).
[0077] 2. Data set standardization: unification of indicators and selection of factors.
[0078] Considering that the passivation effect of materials is influenced by multiple factors, during data collection and processing, the factors affecting the passivation effect were summarized into three categories, taking into account the amount of data and the correlation between variables: synthetic factors, material properties, and environmental factors. Furthermore, to facilitate data processing and analysis in a programming language environment, the variables were simultaneously coded in English, and the units for continuous variables were standardized. In addition, passivation amount, passivation residue, and passivation rate were uniformly converted into passivation rate, i.e., the effect size RR.
[0079] The removal rate, residual amount, and material dosage were standardized, and the relative response (RR) was used as the effect size. The variance and 95% CI were derived based on the statistics of the treatment and control groups. Specific variables included: synthesis method, intercalated ions, loading material, divalent cations, trivalent cations, cation molar ratio, interlayer anions, and specific surface area (m²). 2 / g), pore volume (cm³) 3 / g), average pore size (nm), interlayer spacing (nm), experimental subject, material concentration (%), initial pollution concentration (soil mg / kg, solution mg / L), initial pH, experimental temperature, competing ions, reaction time (min), soil moisture (%).
[0080] 3. Dataset Evaluation and Imputation
[0081] (1) Heterogeneity and publication bias: The heterogeneity test statistic Q, the heterogeneity ratio index I², and the heterogeneity variance Tau² indicate extremely high true heterogeneity (I²≈99.99%); publication bias assessment is conducted to ensure the robustness of the conclusions. (See attached...) Figure 3 As shown in the funnel plot, the collected data exhibits a good symmetrical distribution.
[0082] (2) Missing values: First, statistical imputation is performed by grouping by metal type and medium type, and then MICE multiple imputation joint modeling is used to complete continuous and classification features; the results of multiple imputation are summarized.
[0083] 4. Determine the feature subset
[0084] The Boruta algorithm, based on the importance test of random forest, screens variables with significant effects on response values. Then, based on the results of Spearman correlation analysis between numerical variables, one-way ANOVA between numerical and categorical variables, and chi-square test based on contingency tables between categorical variables, and considering the actual meaning of the variables, highly correlated or highly collinear variables are removed if they are meaningful or have a strong correlation. (See attached...) Figure 4 As shown, the left side presents Boruta's analysis results. Green indicates important features, statistically significantly stronger than the maximum importance of "shadow features"; yellow indicates undetermined features, where evidence is insufficient to determine whether they are stronger or weaker than the maximum shadow value, and their retention will be decided based on correlation results; red indicates unimportant features, weaker than the maximum shadow value; blue represents shadow features (randomly shuffled independent variables), whose maximum value (shadowMax) is used as the baseline / baseline distribution. Combining the importance and correlation results of the above data, the As dataset will retain the 14 variables on the right side of the figure for subsequent analysis. (See attached image.) Figure 5 Similarly, based on the Boruta and correlation analysis results, the 17 variables on the right side of the figure will be retained in the final Cd dataset for subsequent analysis.
[0085] (II) Identification of controlling factors and boundary conditions in integrated machine learning models:
[0086] 1. Model Optimization and Selection
[0087] Due to the characteristics of the dataset, a very small number of models may fail to train during the training process, but this will not affect the subsequent analysis of the overall data. Finally, based on the test set (20% of the dataset), the model R... 2 To determine the optimal model, 10-fold cross-validation was used during parameter tuning and training. The process variables As and Cd in LDHs-passivated soil were modeled separately, and the results are as follows:
[0088] Successfully constructed models for passivated soil As include: SVM_RBF(R 2=0.487), Elastic_Net (R 2 =0.460), GPR_Radial (R 2 =0.445), Linear_Regression (R 2 =0.436), Random_Forest(R 2 =0.421), GBDT (R 2 =0.419), MARS (R 2 =0.334), XGBoost (R 2 =0.273), KNN (R 2 =0.247), see Figure 6 Among them, the best model is the Support Vector Machine with Radial Basis Function Kernel (SVM_RBF), which has a higher performance than the previous model.
[0089] Successful models for passivating Cd in soil include: Random_Forest (R 2 =0.716), GBDT (R 2 =0.700), XGBoost (R 2 =0.677), SVM_RBF (R 2 =0.670), GPR_Radial (R 2 =0.612), MARS (R 2 =0.570), Elastic_Net (R 2 =0.529), Linear_Regression (R 2 =0.522), KNN(R) 2 =0.476), see Figure 7 Overall, the fit was good, with the best model being Random Forest, which performed better than the previous model.
[0090] 2. Identification of the main controlling factors of LDHs passivation of heavy metals in soil
[0091] Based on the optimal model obtained in step 1 of this example, the impact of each variable on model performance is evaluated to obtain comparable "group permutation importance," and the p-value of the one-sided permutation test and the q-value after multiple test correction are given. Important and significant variables are selected as the main controlling factors for the passivation of heavy metals in soil by materials such as LDHs.
[0092] (2) Grouping strategy based on importance of grouping
[0093] a. Grouping principle: Numerical features are grouped independently by "column"; categorical features are aggregated by "original factor" and all its unique columns are treated as a whole group for intervention, thereby evaluating the overall contribution of the original factor at the group level and ensuring the comparability of different types of variables.
[0094] b. Evaluation dataset: Baseline performance is calculated using independent validation / test sets to avoid optimistic bias on hyperparameter tuning data.
[0095] c. Importance metric: based on "loss increment" "Characteristic variables (groups) are defined as..." Loss_base is the baseline loss, and Loss_perm is the loss after permutation.
[0096] d. Replacement method: Randomly rearrange all columns of a variable group simultaneously (preserving the correspondence between columns within the group), repeating nsim times (default 300), calculating each time. Group importance takes all The mean and stability are characterized by standard deviation.
[0097] (2) Significance determination of one-sided permutation test:
[0098] a. Null hypothesis H0: This variable (group) makes no positive contribution to the model performance (i.e., The expected value is ≤ 0.
[0099] b. Calculation of one-sided p-value: The smaller the p-value, the stronger the evidence to reject H0, indicating that the loss increased significantly after the replacement and that the variable did indeed contribute.
[0100] (3) Multiple test correction:
[0101] a. Apply the Benjamini-Hochberg method to the one-sided permutation test p-values of all variables for FDR control to obtain the corrected q-values.
[0102] b. Significance threshold and labeling: Statistical significance is determined by q < 0.05.
[0103] As attached Figure 8 As shown, the importance of the relevant variables for LDHs passivation of heavy metal As in soil is ranked as follows: average pore size > interlayer anions > synthesis method > passivation time > composite materials > competing ions > pore volume > divalent cations > divalent / trivalent cation ratio > initial pH > material concentration > specific surface area > passivation temperature > initial pollution concentration. Furthermore, considering the significant indicators * in the figure, the main controlling factors for significant and important variables are: average pore size, interlayer anions, synthesis method, composite materials, and pore volume.
[0104] As attached Figure 9 As shown, the importance of the relevant variables for LDHs passivation of heavy metal Cd in soil is ranked as follows: average pore size > initial soil pH > passivation time > synthesis method > material concentration > initial contamination concentration > pore volume > trivalent cations > competing ions > soil moisture > specific surface area > ratio of divalent to trivalent cations > composite material > intercalated ions > divalent cations > passivation temperature > interlayer anions. Combined with the significant markers * in the figure, the main controlling factors are average pore size, initial pH, passivation time, synthesis method, material concentration, contamination concentration, pore volume, and competing ions.
[0105] 3. Determination of boundary conditions for LDHs passivation of heavy metals in soil
[0106] (1) For numerical variables, the candidate interval generation method is as follows:
[0107] Threshold distribution method: In samples where RR ≥ t, the quartile interval [Q25, Q75] of the independent variable is selected, prioritizing the interval generated by the highest threshold t; LOESS peak preservation method: Fitting... Find the peak value y*, and take the value that satisfies the condition. The independent variable interval is determined using the weighted kernel density method: the kernel density of the variable is estimated using RR as the weight, and the central p% interval of the cumulative density is taken (default 80%). Then, interval fusion is performed: the final interval [Lower, Upper] is determined in the order of "all intersections → widest pairwise intersections → median aggregation". The final output is [Lower, Upper] and the RR statistic within the interval, which is used to characterize the optimal range of values for the independent variable and the expected effect.
[0108] (2) The best method for determining the class variables is to calculate the effect and frequency of each class, perform normalization and weighting to obtain the comprehensive score, one-way ANOVA, and Tukey test; First, class statistics and scoring: for each class c of each class variable, calculate: mean_Effect(c), quartile interval of RR RR_Q25 / Q75, sample size n, and the proportion of this class in all samples Frequency(c); after normalization, we get:
[0109] Overall rating: , used for sorting.
[0110] Then, significance and grouping were performed: a univariate ANOVA test was used to assess the overall difference, and Tukey HSD was used for pairwise comparisons.
[0111] Finally, the optimal set (discrete boundary conditions) is selected: find the letter g* of the class c* with the highest mean, and the candidate set C0 = a group with the same letter as c*; apply a robust threshold within C0: n≥5 and Mean_Effect≥0.8 to obtain C1; if C1 is empty, take the one with the highest comprehensive score among all the classes; output C1 in descending order of comprehensive score, which is the best class.
[0112] Combine steps 2 and 3:
[0113] The results are attached. Figure 10 , 11 Therefore, the optimal range of values for the main controlling factor of LDHs passivation of heavy metal As in soil is: synthesized by hydrothermal coprecipitation (HyCop), with interlayer anions. - Furthermore, the composite material loaded with magnetic biochar has an average pore size of approximately 3.8 nm and a pore volume of 0.133 cm³. 3 / g of LDHs has a good passivation effect on heavy metal As in soil.
[0114] The results are attached. Figure 12 , 13 From 14, we can obtain the optimal range of values for the main controlling factor of LDHs passivating soil heavy metal Cd: synthesized using the formamide intercalation exfoliation (FIE) method, with an average pore size of approximately 17.37 nm and a pore volume of 0.16-0.21 cm³. 3 / g of LDHs material was added at a concentration of 0.3-2% to soil with Cd contamination concentration of 1.1-50 mg / kg and initial pH of 6.1-6.6, with additional competitive ions added. When the passivation time is 0.5-22 days, the passivation effect on soil heavy metal Cd is good.
[0115] 4. Identification and trade-off of complex heavy metal pollution in LDH-passivated soils
[0116] Extended analysis of As and Cd combined soil pollution, combining the main controlling factors and boundary conditions under the individual presence of As and Cd, and considering the overall trend of the original data, reveals that to achieve a satisfactory passivation effect of LDHs in soils contaminated with both, the following measures are necessary: synthesis using a hydrothermal co-precipitation method, and the use of Ca as the divalent cation. 2+ Trivalent cations are made using Al 3+ The ratio of divalent to trivalent cations is approximately 2.8-3.0. The composite material carrier, serving as the interlayer anion, is magnetic biochar, and the material has a specific surface area of 37.2-63.7 m². 2 / g; passivation is performed in soil with neutral pH.
[0117] The method of this invention maintains robust output in highly heterogeneous data; provides directly callable parameter ranges and boundary conditions for engineering design; can be extended to other heavy metal and compound pollution scenarios, and has versatility and transferability; compared with traditional experiments to screen for optimal conditions, this invention significantly reduces experimental costs and time.
Claims
1. A method for identifying the main controlling factors and boundary conditions of heavy metals in LDH-passivated soils based on ensemble machine learning, characterized in that, This includes multi-source data collection and processing, as well as the construction of integrated machine learning methods; the specific steps are as follows: The first step is to construct a multi-source dataset: systematically search and screen publicly available literature globally to construct a multi-source dataset that includes material synthesis methods, material structural composition information, environmental factors, and passivation indices; The second step is to standardize the multi-source datasets; The passivation index in the multi-source dataset obtained in the first step was standardized, the effect size RR, variance and 95% confidence interval were calculated, and heterogeneity test and publication bias assessment were performed. The third step is to complete the standardized dataset. Specifically, a strategy combining group imputation and multiple chain equation imputation is used to complete the missing variables in the dataset. The fourth step is to determine the feature subset for the completed dataset. Specifically, based on the correlation between variables and the full correlation feature selection algorithm, a comprehensive selection of variable importance and redundancy control is performed to determine the feature subset used for modeling. The fifth step is to select machine learning models based on feature subsets. Specifically, at least one model from the machine learning models will be used for modeling, and the optimal hyperparameters and model will be selected through hyperparameter optimization and k-fold cross-validation. Step 6: Based on the obtained optimal model and dataset, combine the importance of the permutation variables with multiple test correction or other general importance analysis, quantile regression, LOESS smoothing, and weighted kernel density estimation methods to output the main control factor and boundary conditions of the heavy metal passivation effect, and give the effective conditions. The seventh step involves identifying and weighing the common and differential controlling factors and boundary conditions obtained in the sixth step for complex pollution of two or more heavy metals using a multi-objective analysis method.
2. The method according to claim 1, characterized in that, The heavy metals listed in the soil include As, Cd, Cr, Cu, Hg, Ni, Pb, and Zn.
3. The method according to claim 1, characterized in that, The first step in constructing the multi-source dataset, namely the systematic retrieval and screening of publicly available literature globally, involves selecting more relevant and compliant literature based on inclusion and exclusion criteria after the initial retrieval. Specifically: Inclusion criteria: (1) The literature must include a control treatment: experiments with added LDHs and blank control experiments without added LDHs; (2) The literature must include a quantitative value of the passivation effect: passivation rate, passivation amount, or residual amount after passivation; (3) The mean, standard deviation (SD) or error (SE), and the number of repetitions for the control and experimental groups must be obtained from text, tables, or digital charts; (4) The full text of the literature must be obtained. Exclusion criteria include: (1) literature where experimental data is unavailable, or data is incomplete or cannot be converted; (2) literature where the research subject is only aqueous solution or mechanism study; (3) research type is non-randomized controlled trial, including review literature, meta-analysis, conference; (4) duplicate publication or poor quality literature.
4. The method according to claim 1, characterized in that, The second step, standardization of multi-source datasets, includes unifying indicators and selecting factors. Specifically: Considering that the passivation effect of materials is affected by multiple factors, during the data collection and processing, the data volume and the correlation between variables were comprehensively considered. The factors affecting the passivation effect were summarized into three categories: synthetic factors, material properties, and environmental factors. In order to facilitate data processing and analysis in a programming language environment, the variables were simultaneously coded in English and the units of continuous variables were unified. In addition, the passivation amount, passivation residue, and passivation rate were uniformly converted into the passivation rate, i.e., the effect size RR. The removal rate, residual amount, and material dosage were standardized, and the relative response (RR) was used as the effect size. The variance and 95% confidence interval were derived based on the statistics of the treatment group and the control group. Specific variable parameters used include: synthesis method, intercalated ions, supporting materials, divalent cations, trivalent cations, cation molar ratio, interlayer anions, and specific surface area (m²). 2 / g), pore volume (cm³) 3 / g), average pore size (nm), interlayer spacing (nm), experimental subject, material concentration (%), initial pollution concentration, initial pH, experimental temperature, competing ions, reaction time (min), soil moisture (%).
5. The method according to claim 1, characterized in that, The third step, dataset completion after standardization, includes dataset evaluation and imputation, specifically: (1) Heterogeneity and publication bias: The heterogeneity test statistic Q, the heterogeneity ratio index I², and the heterogeneity variance Tau² indicate the heterogeneity of the data; publication bias evaluation is conducted to ensure the robustness of the conclusions. (2) Missing values: First, statistical imputation is performed by grouping by metal type and medium type, and then continuous and classification features are completed by using multiple chain equation imputation method in combination modeling. Summarize the results of multiple interpolation.
6. The method according to claim 1, characterized in that, The fourth step is to determine the feature subset of the completed dataset. Specifically, the "fully correlated" feature screening algorithm based on the importance test of random forest is used to screen variables that have a significant effect on the response value. Then, based on the results of Spearman correlation analysis between numerical variables, one-way ANOVA between numerical and categorical variables, and chi-square test based on contingency tables between categorical variables, combined with the actual meaning of the variables, highly correlated or highly collinear variables are removed if they are meaningful or have related properties.
7. The method according to claim 1, characterized in that, In the fifth step, which involves selecting machine learning models based on feature subsets, the model R is selected based on the test set, which represents 20% of the dataset. 2 To determine the best model, 10-fold cross-validation was used during parameter tuning and training.
8. The method according to claim 1, characterized in that, In step six, based on the obtained optimal model, the impact of each variable on model performance is evaluated to obtain comparable "group permutation importance," and the p-value of the one-sided permutation test and the q-value after multiple test correction are given; important and significant variables are selected as the main controlling factors for the passivation of heavy metals in soil by LDHs materials; among them, (1) Grouping strategy based on importance of grouping by permutation: a. Grouping principle: Numerical features are grouped independently by "column"; Categorical features are aggregated by "original factor" and all its unique columns are treated as a whole group for intervention, thereby evaluating the overall contribution of the original factor at the group level and ensuring the comparability of different types of variables. b. Evaluation dataset: Baseline performance is calculated using independent validation / test sets to avoid optimistic bias on parameter tuning data; c. Importance metric: based on "loss increment" "The importance of characterizing variables (groups);" d. Permutation method: Randomly rearrange all columns of a given variable group simultaneously, repeating NIMS times, calculating each time. Group importance takes all The mean, and stability is characterized by standard deviation; (2) Significance determination of one-sided permutation test: a. Null hypothesis H0: This variable (group) makes no positive contribution to the model performance, i.e. The expected value is ≤ 0; b. Calculation of one-sided p-value: If the p-value is smaller, the evidence to reject H0 is stronger, indicating that the loss increases significantly after the replacement and the variable does indeed contribute. (3) Multiple test correction (BH-FDR): a. Apply the Benjamini-Hochberg method to control for false discovery rate (FDR) in the one-sided permutation test p-values of all variables to obtain the corrected q-values; b. Significance threshold and labeling: Statistical significance is determined by q < 0.
05.
9. The method according to claim 8, characterized in that, In step six, the boundary conditions for heavy metals in the soil are determined: (1) For numerical variables, the candidate interval generation method is as follows: Threshold distribution method: In samples where RR≥t, the interquartile range [Q25,Q75] of the independent variable is taken, and the interval generated by the highest threshold t is used; LOESS peak preservation method: fitting Find the peak value y*, and take the value that satisfies the condition. The interval of the independent variable; Weighted kernel density method: The kernel density of the estimated variable is estimated by using RR as the weight, and the center of the cumulative density is taken as a certain percentage interval; Then, interval fusion is performed: the final interval is determined in the order of "all intersections → widest pairwise intersections → median aggregation"; the final output is the interval range and the RR statistic within the interval, which is used to characterize the optimal range of values for the independent variable and the expected effect. (2) The optimal method for determining the values of categorical variables is to calculate the effect and frequency of each category, perform normalized weighting to obtain the comprehensive score, one-way ANOVA, and Tukey test; First, categorical statistics and scoring: for each category c of each categorical variable, calculate: mean_Effect(c), quartile interval of RR RR_Q25 / Q75, sample size n, and the proportion of this category in all samples Frequency(c); after normalization, we get: Overall rating: Used for sorting; Next, significance and grouping were performed: the results of the univariate ANOVA test for overall differences and the Tukey significance test were compared pairwise. Finally, the optimal set is selected: find the letter g* of the category c* with the highest mean, and the candidate set C0 = a group with the same letter as c*; apply a robust threshold to C0: n≥5 and Mean_Effect≥0.8 to obtain C1; if C1 is empty, take the one with the highest comprehensive score in the whole set as a fallback; output C1 in descending order of comprehensive score, which is the best category.