An ecological risk prediction method for plastic additive-like organic matters in farmland soil
By constructing a random forest model and Bayesian optimization, combined with SHAP value analysis, the problem of accuracy in predicting the ecological risks of plastic additives in farmland soil was solved, achieving efficient identification of high-risk areas and improving prediction accuracy and model stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING NORMAL UNIVERSITY
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot accurately predict the ecological risks of plastic additives in farmland soil. Traditional regression models are unable to capture the complex nonlinear relationships between pollutant characteristics and multiple factors, and are not sensitive enough to outliers.
A machine learning-based random forest model was constructed, and hyperparameters were tuned using Bayesian optimization. By screening key variables and visualizing SHAP values, the relationship between the concentration of organic matter such as plastic additives in farmland soil and potential variables was established. Ecological risk assessment was conducted using risk entropy.
It enables accurate ecological risk prediction of organic compounds such as plastic additives in farmland soil, identifies high-risk areas, improves prediction accuracy and model stability, and reduces costs and complexity.
Smart Images

Figure CN122114615A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for predicting the ecological risks of organic matter, and more particularly to a method for predicting the ecological risks of plastic additives in farmland soil. Background Technology
[0002] Phthalic acid esters (PAEs) and organophosphate esters (OPEs), widely used as plastic additives in industry, agriculture, medical, and construction, have become a significant emerging pollutant in farmland soil. The accumulation of these organic compounds in farmland soil is absorbed by crops and subsequently amplifies through the food chain, posing a threat to food security and endangering human health. Therefore, identifying and scientifically delineating potential pollution risk areas for PAEs and OPEs in farmland soil is of great importance.
[0003] Current technologies cannot accurately predict the ecological risks of plastic additives in farmland soil. Traditional regression models (such as multiple linear regression) struggle to capture the complex nonlinear relationships between pollutant characteristics and multiple factors, and are not sensitive enough to outliers. Random Forest (RF) models, on the other hand, are a supervised machine learning algorithm. Their core concept is to reduce the inherent overfitting risk of individual decision trees by constructing an ensemble of multiple decision trees, while significantly improving prediction accuracy and effectively mitigating problems such as overfitting, missing values, and multicollinearity. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a method for predicting the ecological risks of plastic additive organic matter in farmland soil that is highly accurate, interpretable, has good model stability, and high generalization ability.
[0005] Technical solution: The method for predicting the ecological risk of plastic additives in farmland soil includes the following steps:
[0006] Step 1: Screening potential factors influencing pollutant occurrence from five categories of factors: pollutant properties (n=11), soil physicochemical properties (n=7), agricultural activities (n=3), climatic conditions (n=2), and geographical coordinates (n=2); selecting sampling points, and obtaining sample data of organic plastic additives in farmland soil from the sampling points, along with multi-source data of the five categories of factors, as preliminary sample data; the sample data of organic plastic additives in farmland soil includes the types and concentrations of target pollutants, including phthalates and organophosphates;
[0007] Step 2: Integrate the data from the multiple source datasets, retaining the valid variables as the input dataset;
[0008] Step 3: Establish a random forest model based on machine learning to establish the relationship between the concentration of plastic additive organic matter in farmland soil and latent variables; perform hyperparameter tuning through Bayesian optimization; evaluate the model; screen key variables affecting the accumulation of plastic additive organic matter in farmland soil through random forest importance scores; and visualize and analyze key variables using SHAP values.
[0009] Step four: Use the model constructed in step three to predict the concentration and spatial pollution characteristics of the target pollutants in the target farmland soil;
[0010] Step 5: Convert the concentration into the ecological risk level of the target pollutant using risk entropy, and conduct an ecological risk assessment of the target farmland soil.
[0011] In step one, the pollutant properties preferably include molecular weight, octanol-water partition coefficient, logarithm of vapor pressure, highest occupied molecular orbital, lowest unoccupied molecular orbital, charge, maximum and minimum electrical topological state indices, number of rotatable bonds, Kappa3 index, and FractionCSP3; the soil physicochemical properties preferably include soil organic carbon, pH, cation exchange capacity, clay, sand / silt ratio, total nitrogen, and total phosphorus; the climatic conditions preferably include the logarithm of 1 / T and precipitation; the agricultural activities include average annual application of agricultural fertilizers, average annual pesticide use, and average annual use of agricultural plastic film; and the geographical coordinates preferably include longitude and latitude.
[0012] In step one, the sampling points are preferably obtained from literature on the detection of PAEs and OPEs pollution in farmland soil, from which the coordinates of soil sampling points and pollutant concentration data are obtained; or the sampling point data are obtained from Bigemap Pro (version 4.5.2) software.
[0013] In step one, the preferred sources of multi-source data for the five types of factors are databases such as SoilGrids 2.0, the China High-Resolution National Soil Information Grid Basic Attribute Data Set, the National Glacier, Permafrost and Desert Scientific Data Center, and the China Statistical Yearbook; some data are calculated using software such as the U.S. Environmental Protection Agency EPI Suite™ (version 4.11), Gaussian software (version 16), and Python's RDKit library.
[0014] In step two, the data integration includes methods such as correlation analysis, weighted average, arithmetic average, taking the logarithm, and taking the reciprocal.
[0015] In step three, hyperparameter tuning of the model is performed based on Bayesian optimization. The hyperparameters include the number of trees, the random number seed, the maximum depth, and the maximum eigenvalue. According to R... 2Model evaluation was performed using RMSE and MAE; Sankey diagrams of feature importance were plotted based on the selected key variables; and SHAP values were calculated using the shap library in Python to visualize and analyze the key variables.
[0016] In step four, it is preferable to form a spatial pollution characteristic distribution map of pollutants in farmland soil.
[0017] In step five, the preferred method for calculating the risk entropy is as follows:
[0018] RQ = MEC / PNECsoil Formula 1
[0019] In Formula 1, RQ is the risk entropy, MEC is the concentration of pollutants in the soil, and PNECsoil is the concentration of pollutants in the soil that has no effect.
[0020] PNECsoil=Foc×Koc×PNECaqua Formula 2
[0021] In Formula 2, Foc is the weight fraction of organic carbon in the soil, Koc is the partition coefficient between organic carbon and water, and PNECaqua is the concentration of pollutants in the aquatic environment that has no effect.
[0022] In step five, an ecological risk level distribution map of plastic additive organic matter in farmland soil is preferably generated based on risk entropy.
[0023] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: By constructing a random forest model of the relationship between the concentration of plastic additive organic matter (PAEs) in farmland soil and latent variables, this invention can accurately identify the spatial pollution characteristics and key influencing factors of PAEs and OPEs in farmland soil, and rapidly screen potential high-ecological-risk points of PAEs and OPEs in farmland soil. This method provides a new approach for identifying the characteristics of organic pollution in farmland soil and screening potential high-ecological-risk points in my country. Attached Figure Description
[0024] Figure 1 A flowchart for predicting the ecological risk of organic matter such as plastic additives in farmland soil;
[0025] Figure 2 This is a schematic diagram showing the regression scattering points between the predicted and actual values of the data (a represents PAEs, b represents OPEs).
[0026] Figure 3 Sankey diagram for the importance of key variables;
[0027] Figure 4 Heatmaps showing the contribution of SHAP features to key variables (a represents PAEs, b represents OPEs).
[0028] Figure 5 Map showing the coordinate distribution of farmland soil sampling points in five cities in southern Jiangsu;
[0029] Figure 6 This is a spatial distribution map of pollutant characteristics based on predicted concentrations (a represents PAEs, b represents OPEs).
[0030] Figure 7 This is a distribution map of soil ecological risk classification based on risk entropy (a represents PAEs, b represents OPEs). Detailed Implementation
[0031] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0032] Example 1
[0033] This embodiment provides a method for predicting the ecological risk of plastic additives in farmland soil, the process of which is as follows: Figure 1 As shown, the steps are as follows.
[0034] Step 1: Screen potential factors affecting pollutant occurrence from five categories of factors: pollutant properties, soil physicochemical properties, agricultural activities, climate conditions, and geographical coordinates; find farmland soil sampling point information from the selected literature, covering all provinces in China; and obtain sample data of organic matter such as plastic additives in farmland soil in the study area and multi-source data of the five categories of factors from relevant literature and authoritative databases based on the sampling point information.
[0035] The data on organic matter samples of plastic additives in farmland soil includes pollutant types and concentrations; the multi-source dataset of variables includes data on pollutant properties (n=11), soil physicochemical properties (n=7), agricultural activities (n=3), meteorological conditions (n=2), and geographic coordinates (n=2); the target pollutants include 6 phthalic acid esters (PAEs) and 11 organophosphate esters (OPEs).
[0036] The six PAEs mentioned above are: di(2-ethylhexyl) phthalate (DEHP), dimethyl phthalate (DMP), diethyl phthalate (DEP), di-n-butyl phthalate (DBP), benzyl butyl phthalate (BBP), and di-n-octyl phthalate (DNOP).
[0037] The 11 OPEs mentioned above are: tris(2-butoxyethyl) phosphate (TBOEP), tris(1-chloro-2-propyl) phosphate (TCIPP), tributyl phosphate (TnBP), tris(2-ethylhexyl) phosphate (TEHP), diphenyl (2-ethylhexyl) phosphate (EHDPP), tris(2-chloroethyl) phosphate (TCEP), triethyl phosphate (TEP), tris(2-chloropropyl) phosphate (TCPP), triphenylphosphine oxide (TPPO), triphenyl phosphate (TPhP), and triphenylphosphine (TPP).
[0038] The study area covers all provinces of China, with the prediction area being the southern Jiangsu region (see [link]). Figure 5 The soil sampling point coordinates are obtained from the relevant papers or Bigemap Pro (version 4.5.2) based on the locations described in the selected articles.
[0039] The data sources for the 25 variables across the 5 major categories of factors are detailed in Table 1.
[0040] Table 1
[0041] .
[0042] Step 2: Integrate the initial sample dataset to obtain the research sample dataset.
[0043] The initial sample data is integrated to obtain the research sample dataset: The corresponding input dataset is obtained by integrating the data, including methods such as correlation analysis, weighted average, arithmetic average, logarithm, and reciprocal.
[0044] After the above steps of screening and integration, the finally determined independent variables are the effective variables. Ineffective variables and redundant variables with strong collinearity are deleted from the data or are no longer included in the subsequent random forest model analysis, and are not used as research sample data.
[0045] Step 3: Based on the research sample data, establish a random forest model of the relationship between the concentration of organic matter such as plastic additives in farmland soil and latent variables using machine learning.
[0046] In some instances, the specific steps include: building the model in Python 3.11.9 (Visual Studio Code); after preparing the input dataset, splitting it randomly into training and validation sets in an 8:2 ratio; and then using the coefficient of determination (R²) to... 2The robustness and generalization ability of the RF model are evaluated using the root mean square error (RMSE) and mean absolute error (MAE), where R0 is the most significant factor. 2 In the interval between 0 and 1, R is generally... 2 The closer the value is to 1, the better the fit. The smaller the RMSE and MAE, the smaller the prediction bias, indicating better model prediction performance.
[0047] The hyperparameters of the RF model were tuned using Bayesian optimization, and the model was repeated 400 times. The hyperparameters included the number of trees (n_estimators), the random number seed (random_state), the maximum depth (max_depth), and the maximum number of features (max_features). Taking the coefficient of determination as an example, the goal was to maximize R0. 2 For this purpose, the number of trees in the RF model is 350 / 243; the random number seed of the RF model is 135 / 135; the maximum depth of the RF model is 12 / 14; and the maximum feature value of the RF model is 0.5 / 0.2653. Then, the objective function value R is calculated on the validation set. 2 In this study, the R of the PAEs dataset... 2 =0.70, RMSE=0.47, MAE=0.34, R value of the OPEs dataset 2 =0.69, RMSE=0.41, MAE=0.31, indicating that this RF model has high robustness and generalization ability (see Figure 2 ).
[0048] Random Forest Importance Scores (RFISs) were used to comprehensively assess the importance of input variables to the target vector, identifying key variables influencing the accumulation of plastic additive-like organic matter in farmland soil. A Sankey diagram of feature importance was then constructed based on the described eigenvalues (see [link to diagram]). Figure 3 This demonstrates the degree of contribution of different key variables to the target variable.
[0049] Based on the aforementioned random forest model, key variables are selected. The contribution rate of these key variables to the concentration of plastic additive organic matter in farmland soil is visualized and analyzed using a SHAP cellar diagram. The influence and direction of each feature on the model's prediction results are analyzed. Specific steps include: using the trained random forest model and corresponding training and test sets, and calculating SHAP values using the shap library in Python. This method uses a SHAP cellar diagram, such as... Figure 4As shown, by observing the arrangement order, color distribution, curve trend, and arrow direction and length of features in different visualizations, key variables and their contribution to the concentration of plastic additive organic matter in farmland soil can be determined. At the same time, by comparing the magnitude and sign of the SHAP values of different variables, the relative importance of each key variable and their promoting or inhibiting effects on organic matter concentration can be analyzed.
[0050] Step 4: Predict the concentration of organic matter from plastic additives in farmland soil and spatial pollution characteristics based on RF models and algorithms.
[0051] Based on the aforementioned feature importance ranking and SHAP value, the key variables selected are used to predict the concentration of PAEs and OPEs in farmland soil using the RF model, forming a spatial pollution characteristic map of plastic additive organic matter in farmland soil in southern Jiangsu.
[0052] In some instances, the specific steps included: collecting the coordinates of 216 sampling points from farmland soil in southern Jiangsu using the geographic information software Bigemap Pro (see...). Figure 5 The model uses data on five major categories of influencing factors—pollutant properties, soil physicochemical properties, climate conditions, agricultural activities, and geographic coordinates—from relevant databases as input data. Based on the selected key variables and the model's optimal hyperparameters, and considering the characteristics of the original data, a machine learning model is used to predict the concentrations of PAEs and OPEs in farmland soils in southern Jiangsu. Based on the prediction results, Kriging interpolation in ArcGIS software is used to generate a spatial pollution characteristic distribution map of plastic additive organic matter in farmland soils in southern Jiangsu (see...). Figure 6 ).
[0053] Step 5: Calculate the risk entropy of plastic additive organic matter in farmland soil in southern Jiangsu, and use it as an ecological risk assessment standard for plastic additive organic matter in farmland soil in southern Jiangsu.
[0054] Based on the concentration of organic matter from plastic additives in farmland soil in southern Jiangsu Province predicted by the RF model, the predicted concentration is converted into the ecological risk level of the pollutants in farmland soil in southern Jiangsu Province through the risk entropy (RQ), thereby conducting an ecological risk assessment of farmland soil in southern Jiangsu Province.
[0055] The ecological risks of PAEs and OPEs in the farmland soil were calculated based on RQ:
[0056] RQ = MEC / PNECsoil (1)
[0057] Wherein, MEC is the soil pollutant concentration (ng / g dw), and PNECsoil is the pollutant no-effect concentration in the soil (ng / g dw). According to commonly used recommended standards, the risk level can be divided into three levels: RQ < 0.1 is low risk; 0.1≤RQ< 1 is medium risk or adverse reaction; and RQ≥1 is high risk.
[0058] PNECsoil=Foc×Koc×PNECaqua (2)
[0059] Wherein, Foc is the weight fraction of organic carbon in the soil (using 10%), Koc is the partition coefficient between organic carbon and water (L / kg), and PNECaqua is the concentration of pollutants in the aquatic environment with no effect.
[0060] Based on the calculated RQ, an ecological risk level distribution map of plastic additive organic matter in farmland soil in southern Jiangsu was generated using the Kriging interpolation method in ArcGIS software. Figure 7 As shown.
[0061] This embodiment uses RF simulation to measure the concentration levels of organic matter (PAEs) in farmland soil, identifies key variables affecting the accumulation of PAEs and OPEs, predicts the concentrations of PAEs and OPEs, and uses Kriging interpolation in ArcGIS software to generate spatial pollution characteristics based on the prediction results, thereby conducting an ecological risk assessment of the soil. This method overcomes the drawbacks of traditional soil sampling qualitative and quantitative analysis, such as high cost, increased complexity, and time-consuming processes. It also overcomes the shortcomings of traditional linear algorithms, such as low accuracy and weak generalization ability in concentration prediction. This method offers advantages such as high prediction accuracy, good interpretability, good model stability, and excellent generalization ability, providing a new approach for identifying organic pollution characteristics and screening potential high-ecological-risk points in farmland soil in my country.
Claims
1. A method for predicting the ecological risk of plastic additives in farmland soil, characterized in that, Includes the following steps: Step 1: Screening potential factors affecting pollutant occurrence from five categories of factors: pollutant properties, soil physicochemical properties, agricultural activities, climate conditions, and geographical coordinates; selecting sampling points and obtaining sample data of organic plastic additives in farmland soil from the sampling points, along with multi-source data of the five categories of factors, as preliminary sample data; the sample data of organic plastic additives in farmland soil includes the types and concentrations of target pollutants, including phthalates and organophosphates; Step 2: Integrate the data from the multiple source datasets, retaining the valid variables as the input dataset; Step 3: Establish a random forest model based on machine learning to establish the relationship between the concentration of plastic additive organic matter in farmland soil and latent variables; perform hyperparameter tuning through Bayesian optimization; evaluate the model; screen key variables affecting the accumulation of plastic additive organic matter in farmland soil through random forest importance scores; and visualize and analyze key variables using SHAP values. Step four: Use the model constructed in step three to predict the concentration and spatial pollution characteristics of the target pollutants in the target farmland soil; Step 5: Convert the predicted concentration into the ecological risk level of the target pollutant using risk entropy, and conduct an ecological risk assessment of the target farmland soil.
2. The method according to claim 1, characterized in that, In step one, the properties of the pollutant include molecular weight, octanol-water partition coefficient, vapor pressure logarithm, highest occupied molecular orbital, lowest unoccupied molecular orbital, charge, maximum electrotopic state index, minimum electrotopic state index, number of rotatable bonds, Kappa3 index, and FractionCSP3. The soil physicochemical properties include soil organic carbon, pH, cation exchange capacity, clay, sand / silt, total nitrogen, and total phosphorus. The climatic conditions include the logarithm of 1 / T and precipitation; The agricultural activities include the average annual application of agricultural fertilizers, the average annual use of pesticides, and the average annual use of agricultural plastic film; The geographical coordinates mentioned include longitude and latitude.
3. The method according to claim 1, characterized in that, In step one, the sampling points are selected from literature that has detected PAEs and OPEs pollution in farmland soil, and the coordinates of the soil sampling points and pollutant concentration data are obtained from them; or the sampling point data are obtained from Bigemap Pro (version 4.5.2) software.
4. The method according to claim 1, characterized in that, In step one, the multi-source data for the five types of factors are obtained from SoilGrids 2.0, the basic attribute dataset of the China High-Resolution National Soil Information Grid, the National Data Center for Glacier, Permafrost and Desert Science, or the China Statistical Yearbook; or calculated using the U.S. Environmental Protection Agency's EPI Suite™ (version 4.11), Gaussian software (version 16), or Python's RDKit library.
5. The method according to claim 1, characterized in that, In step two, the data integration methods include correlation analysis, weighted average, arithmetic average, logarithm, and reciprocal.
6. The method according to claim 1, characterized in that, In step three, the hyperparameters include the number of trees, the random number seed, the maximum depth, and the maximum eigenvalue; according to R... 2 RMSE and MAE are used to evaluate the model.
7. The method according to claim 1, characterized in that, In step three, a Sankey diagram of feature importance is drawn based on the key variables; the SHAP value is calculated using the shap library in Python, and the key variables are visualized and analyzed.
8. The method according to claim 1, characterized in that, In step four, a spatial pollution characteristic distribution map of pollutants in farmland soil is generated.
9. The method according to claim 1, characterized in that, In step five, the risk entropy is calculated as follows: RQ = MEC / PNECsoil Formula 1 In Formula 1, RQ is the risk entropy, MEC is the concentration of pollutants in the soil, and PNECsoil is the concentration of pollutants in the soil that has no effect. PNECsoil=Foc×Koc×PNECaqua Formula 2 In Formula 2, Foc is the weight fraction of organic carbon in the soil, Koc is the partition coefficient between organic carbon and water, and PNECaqua is the concentration of pollutants in the aquatic environment that has no effect.
10. The method according to claim 1, characterized in that, In step five, an ecological risk level distribution map of plastic additive organic matter in farmland soil is generated based on risk entropy.