Soil heavy metal area risk probability assessment method, system and device and storage medium
By employing technical means to assess regional risks of heavy metals in soil, the lack of differentiated risk standards and population assessments in existing technologies has been addressed, enabling precise risk and population assessments and providing a scientific basis for environmental management.
Patent Information
- Application Number
- CN202510995556.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Existing regional risk prediction models for heavy metals in soil lack differentiated risk standards in large-scale multi-receptor regions, making it difficult to integrate multidimensional environmental variables and quantitatively assess at-risk populations. This results in inaccurate risk assessment results and a lack of scientific decision support.
The collected data is cleaned and standardized by collecting land use types. Risk standards are set based on land use types. Multidimensional environmental variables and random forest classifiers are introduced to construct a risk assessment model. Combined with population spatial distribution data, the risk cutoff probability is determined and the number of at-risk people is estimated.
It enables precise risk assessment for different land use types, improves the adaptability and prediction accuracy of the model, provides quantitative risk population assessment, and provides scientific decision support for environmental management.
Smart Images

Figure CN121034469A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental risk assessment technology, specifically to a method, system, device, and storage medium for assessing the probability of heavy metal risk in soil regions. Background Technology
[0002] Assessing the regional environmental risk of heavy metals in soil is fundamental to developing regional risk control strategies. Regional environmental risk assessment is a method for analyzing and comparing the environmental impacts of different pollutants at a macro-scale geographical area, including the risk of sudden environmental pollution incidents and the probability of cumulative environmental pollution hazards. Among these, regional environmental risk assessment of heavy metals evolved from the health risk assessment of heavy metals in contaminated soil at contaminated sites, primarily examining the static risk of a specific pollutant at a specific point in time.
[0003] However, regional environmental risks are influenced by a combination of environmental risk subsystems, including risk source release, translocation, and risk receptor exposure, and these subsystems change significantly with social development. Most existing spatiotemporal risk prediction models require extensive regional measured environmental pollution data and historical data, focusing on the regional heavy metal pollution risk status of a single site or a site-specific event. This scale difference means that existing models and research findings are not suitable for large-scale, multi-receptor regional environmental risk prediction and simulation.
[0004] Furthermore, most existing studies on soil heavy metal risk prediction do not consider the impact of land use patterns on exposure models, relying solely on a single risk exposure model for evaluation. Ignoring the impact of land use change on the spatiotemporal prediction of soil heavy metal risk may lead policymakers to overestimate or underestimate the environmental risks in certain regions. Currently, there is a lack of accurate and effective analytical methods to address the problems of poor predictive performance, inflexibility, and adaptability of existing regional soil heavy metal risk prediction models for areas with high spatial variability. New theoretical foundations and data models are needed to resolve this current predicament.
[0005] Chinese patent document CN114511239B discloses a method, device, electronic device and medium for delineating soil heavy metal pollution risk zones. It discloses a technical solution for predicting the risk level of soil heavy metal pollution based on the random forest algorithm. It has the technical effect of improving the accuracy of risk zone delineation boundaries and solving the problem of strong dependence on the representativeness and quantity of sampling points in geostatistical spatial interpolation methods. However, it still has the problems of not considering the differences in risk receptors and exposure modes under different land use types, lacking the function of quantitative assessment of the number of at-risk populations, and failing to provide direct quantitative support for environmental management decisions.
[0006] Traditional methods for assessing heavy metal risk in soil face numerous challenges in dealing with complex spatial variability and multi-factor influences. On the one hand, existing methods struggle to effectively integrate multi-dimensional environmental variables such as climate factors, anthropogenic activity factors, and soil geological conditions, making it impossible to establish comprehensive risk prediction models. On the other hand, the lack of differentiated risk assessment standards for different land use types limits the accuracy and applicability of risk assessment results.
[0007] Meanwhile, existing technologies have significant shortcomings in assessing the size of at-risk populations. While high-risk areas can be identified, there is a lack of effective methods to combine risk areas with population distribution, failing to provide quantitative decision support for environmental health risk management. This leaves government departments without a scientific basis when formulating environmental management policies, making it difficult to achieve precise risk control. Summary of the Invention
[0008] The purpose of this invention is to provide a method, system, device, and storage medium for assessing the regional risk probability of heavy metals in soil, which solves the problems of lacking differentiated risk standards, difficulty in integrating multidimensional environmental variables, and inability to quantitatively assess at-risk populations in the prior art, and achieves accurate risk assessment of heavy metals in soil and evaluation of population exposure.
[0009] To achieve the above objectives, the present invention provides the following technical solution: A method for assessing the probability of heavy metal risk in soil areas includes the following steps: S1: Collect basic information on soil heavy metal content, environmental variables, and land use types; clean, filter, format, and standardize the collected data. S2: Soil samples collected are classified based on land use type, and corresponding standard risk values for heavy metals in soil are set for different risk receptor groups and their exposure pathways under different land use types. S3: Based on the relationship between the release and accumulation of heavy metals in soil, collect spatially continuous environmental variable data as predictive factors. The environmental variable data includes the dimensions of climate factors, human activity factors, soil geological conditions, and weathering factors. S4: A pollution risk assessment model is constructed using a random forest classifier, with the soil heavy metal risk value set as the response variable, and the original dataset is randomly divided into a training set and a test set. S5: The model is trained and tested based on the training subset and the test subset to obtain the regional risk hazard probability distribution of heavy metals in the regional soil. S6: Introduce the receiver operating characteristic curve to determine the risk cutoff probability. The region where the probability of harm is greater than the risk cutoff probability is a high-risk region, and the region where the probability of harm is less than the risk cutoff probability is a low-risk region. S7: Combine the regional soil heavy metal risk probability distribution map with population spatial distribution data to estimate the number of at-risk people living in high-risk areas.
[0010] Furthermore, in S2, the collected soil samples are divided into three categories based on land use type: agricultural land, construction land, and forest and grassland. For agricultural land, different soil risk screening values are selected as standard risk values according to different pH ranges. For construction land, the construction land screening value is selected as the standard risk value. For forest and grassland, the geochemical baseline of heavy metals in the soil is calculated using the relative cumulative frequency distribution method, and the lower limit of the abnormal concentration of heavy metals is used as the standard risk value.
[0011] Furthermore, the climate factor dimension in S3 includes annual average rainfall and annual average temperature, while the anthropogenic activity factor dimension includes land use type, population density, GDP per square kilometer, distance between sampling point and polluting enterprise, distance between sampling point and road, distance between sampling point and railway, and distance between sampling point and tailings dam.
[0012] Furthermore, the soil geological conditions dimension in S3 includes parent material type, soil type, altitude, slope, and landform type, while the weathering factor dimension includes chemical alteration index, silicon-aluminum ratio, and silicon-aluminum-iron index.
[0013] Furthermore, in S4, the environmental factor weights are obtained using the average decrease in precision.
[0014] Furthermore: the formula for calculating the average decrease in accuracy is as follows:
[0015] In the formula, OOB refers to out-of-bag data; MSE(OOB) ' iv The expression () represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree; while MSE(OOB) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree. iv ) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable of the i-th tree after shuffling and reordering into the i-th tree.
[0016] Furthermore, in S6, the prediction results of the prediction model are evaluated using the confusion matrix statistical method and ROC curve. High-risk areas and low-risk areas are distinguished by setting the cutoff probability, which is determined by the Youden index.
[0017] Furthermore: In step S7, by determining high-risk areas and population spatial distribution maps, the number of at-risk individuals in high-risk areas is estimated, wherein the formula for calculating the at-risk population is:
[0018] POPaffected It refers to the number of at-risk individuals in high-risk areas. >risk value It is the probability that the concentration of heavy metals in the soil exceeds the risk value, and Cut is the cutoff probability that distinguishes between high risk and low risk.
[0019] A machine learning-based system for assessing the regional risk probability of heavy metals in soil, comprising: The data collection and preprocessing module is used to collect basic information on soil heavy metal content, environmental variables, and land use status, and to clean, filter, format, and standardize the collected data. The risk standard setting module is used to classify the collected soil samples based on land use type and set corresponding soil heavy metal standard risk values for different land use types. The environmental variable collection module is used to collect spatially continuous environmental variable data as predictors; The risk assessment module is used to construct a pollution risk assessment model using a random forest classifier to obtain the regional risk hazard probability distribution of heavy metals in regional soil. The risk zone division module is used to determine the risk cutoff probability and divide high-risk and low-risk zones. The risk population assessment module is used to estimate the number of people at risk by combining risk probability distribution maps with population spatial distribution data.
[0020] Furthermore, the risk standard setting module divides soil samples into three categories: agricultural land, construction land, and forest and grassland, and adopts corresponding standard risk value setting methods for different types.
[0021] A soil heavy metal regional risk probability assessment device includes a processor, which is capable of executing a computer program that can implement the above-mentioned soil heavy metal regional risk probability assessment method.
[0022] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the program is executed by a processor, it implements the above-described method for assessing the regional risk probability of heavy metals in soil.
[0023] Compared with the prior art, the present invention has the following advantages: I. This invention constructs a targeted regional environmental risk prediction model for heavy metals in soil by comprehensively considering the spatial variability of heavy metals in regional soils and the differences in different land use types. Corresponding risk assessment standards are set for three different land use types: agricultural land, construction land, and forest and grassland. This fully considers the differences in risk receptors and exposure modes on different land types, making the risk assessment results more scientific and accurate, and providing reliable technical support for the formulation of precise soil pollution control and environmental management policies.
[0024] Second, this invention effectively solves the technical problems of existing soil heavy metal regional risk prediction models, such as poor prediction performance, lack of flexibility and adaptability in areas with high spatial variability. By introducing multi-dimensional environmental variable data and random forest machine learning algorithms, the model can clearly distinguish the transition characteristics from high-risk to low-risk areas, accurately capture the spatial volatility of data, and significantly improve the prediction accuracy and model adaptability for areas with strong spatial variability.
[0025] Third, this invention realizes a complete technical chain from risk identification to population impact assessment. By combining the probability distribution of heavy metal risks in soil with spatial distribution data of the population, it quantitatively estimates the number of people exposed to high-risk areas, providing a direct basis for decision-making in environmental health risk management. Compared with existing technologies that can only identify risk areas, this invention further quantifies the potential impact of risks on the population, achieving an effective transformation from technical assessment to management application.
[0026] Fourth, this invention establishes a hierarchical risk management zone division system and a risk change trend prediction mechanism. By using ROC curve analysis and Youden's index optimization to determine the optimal risk cutoff probability, it achieves a scientific division between high-risk and low-risk areas. This enables effective early warning of the environmental risk status of heavy metals in the soil of the study area, providing important theoretical basis and technical means for dynamic monitoring and proactive prevention and control of regional environmental risks. Attached Figure Description
[0027] Figure 1 A flowchart of a soil heavy metal regional risk probability assessment method based on machine learning provided by the present invention; Figure 2 The double logarithmic distribution of relative cumulative density and heavy metal elements in a machine learning-based soil heavy metal regional risk probability assessment method provided by this invention; Figure 3 ROC curve in a machine learning-based method for assessing the risk probability of heavy metals in soil regions provided by this invention; Figure 4 The Youden index curve is shown in the soil heavy metal regional risk probability assessment method provided by the present invention based on machine learning. Detailed Implementation
[0028] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Example 1 like Figure 1As shown: This invention provides a method for assessing the probability of heavy metal risk in soil areas based on machine learning, comprising the following steps: S1: Collect basic information on soil heavy metal content, environmental variables, and land use types; clean, screen, format, and standardize the collected data. S2: Classify the collected soil samples based on land use types; set corresponding standard risk values for soil heavy metals for different risk receptor groups and their exposure pathways under different land use types. S3: Based on the relationship between the release and accumulation of heavy metals in soil, spatially continuous environmental variable data are collected as predictors. The environmental variable data include the dimensions of climate factors, human activity factors, soil geological conditions, and weathering factors. S4: A pollution risk assessment model is constructed using a random forest classifier. The soil heavy metal risk value is set as the response variable, and the original dataset is randomly divided into a training set and a test set. S5: The model is trained and tested based on the training subset and the test subset to obtain the regional risk hazard probability distribution of heavy metals in the regional soil. S6: Introduce the receiver operating characteristic curve to determine the risk cutoff probability. The region where the probability of harm is greater than the risk cutoff probability is a high-risk region, and the region where the probability of harm is less than the risk cutoff probability is a low-risk region. S7: Combine the regional soil heavy metal risk probability distribution map with population spatial distribution data to estimate the number of at-risk people living in high-risk areas.
[0030] Specifically, in step S1, the basic data is regional soil heavy metals, which can be obtained using a grid-based sampling method. The indicators include the total concentration of heavy metals in the soil and soil pH. The data collection process includes obtaining soil heavy metal content data through sampling point deployment, while simultaneously collecting environmental variable information and land use type information for the corresponding locations. The data preprocessing stage involves quality checks on the collected raw data, removing outliers and missing values, and standardizing data of different dimensions to ensure data format consistency and usability. In other words, through systematic data collection and preprocessing, high-quality basic data support is provided for the subsequent construction of the risk assessment model.
[0031] In step S2, soil samples are classified according to the "Land Use Status Classification" standard. A differentiated standard risk value setting system is established for different risk receptor groups and exposure pathways.
[0032] Specifically, based on the "Classification of Current Land Use Status" (GB / T 21010-2017), the collected soil samples are divided into three categories, namely agricultural land, construction land, and forest and grassland. When setting different standard risk values according to different risk receptor groups and exposure pathways. For agricultural land, different soil risk screening values are selected as its standard risk value (C0) according to different pH ranges of the soil; when pH ≤ 5.5, the standard risk value is C0 = A; when 5.5 < pH ≤ 6.5, the standard risk value is C0 = B; when 6.5 < pH ≤ 7.5, the standard risk value is C0 = C; when pH > 7.5, the standard risk value is C0 = D; for construction land, the construction land screening value is selected as its standard risk value. For the first type of land, the standard risk value is C0 = E, and for the second type of land, the standard risk value is C0 = F; for forest and grassland, the relative cumulative frequency distribution method (CFD) is used to calculate the geochemical baseline (baseline, BL) of each heavy metal in the soil. The heavy metal content in the soil in the area not strongly disturbed by human activities generally shows a log-normal distribution. In the double logarithmic distribution diagram of the relative cumulative density and heavy metal elements, the lower inflection point of the distribution curve is the upper limit of the natural background concentration, and the higher inflection point is the lower limit of the abnormal concentration. The lower limit of the heavy metal abnormal concentration is used as the standard risk value of the forest and grassland soil C0 = G There are significant differences in the exposure methods of the population under different land use types. Agricultural land is mainly exposed through the food chain, construction land is mainly exposed through direct contact and inhalation, and forest and grassland mainly consider the natural background level of the ecosystem. By establishing targeted risk standards, the actual risk status of different regions can be more accurately reflected, providing a scientific basis for judgment for subsequent risk assessment.
[0033] In step S3, the collection of environmental variable data is based on the migration and transformation laws of soil heavy metals in the environment. Climate factors affect the weathering and release process of heavy metals, human activity factors reflect the contribution of human activities to heavy metal pollution, soil geological conditions determine the occurrence form and migration ability of heavy metals, and weathering factors characterize the influence of geological background on heavy metal content. By constructing a multi-dimensional environmental variable system, various factors affecting the distribution of soil heavy metals can be comprehensively reflected, providing rich feature information for the machine learning model. That is, the introduction of multi-dimensional environmental variables significantly improves the prediction accuracy and applicability of the risk assessment model.
[0034] In step S4, the construction process of the random forest classifier includes feature selection, model parameter setting, and dataset partitioning. Soil heavy metal risk values are used as the target variable, and environmental variables are used as input features. The dataset is randomly partitioned into a training set (90%) and a test set (10%) to ensure the model's generalization ability. The random forest algorithm, by constructing multiple decision trees and combining their prediction results, can effectively handle high-dimensional data and nonlinear relationships, improving the stability and accuracy of prediction results. Through reasonable model construction, a technical foundation is laid for accurately predicting soil heavy metal risk.
[0035] In step S5, the model training process uses 500 repeated training and test subsets for cross-validation. During training, key parameters of the random forest, including the number of decision trees, maximum number of features, and maximum depth, are adjusted to optimize model performance. In the testing phase, the model's predictive ability is evaluated using an independent test set, and the risk probability value for each sampling point is calculated. By averaging the results of multiple training iterations, a stable and reliable regional risk hazard probability distribution map is obtained. In other words, through a systematic model training and validation process, the scientific validity and reliability of the risk assessment results are ensured.
[0036] In step S6, receiver operating characteristic (ROC) curve analysis is used to determine the optimal risk cutoff probability. The ROC curve reflects the model's sensitivity and specificity across all possible cutoff probabilities. The overall model performance is evaluated by calculating the area under the curve (AUC), with an AUC value closer to 1 indicating stronger discriminative ability. The risk cutoff probability is determined using the Youden index method, selecting the probability value that maximizes the sum of sensitivity and specificity as the optimal cutoff point. Based on the determined cutoff probability, the study area is divided into high-risk and low-risk levels, providing clear spatial boundaries for risk management.
[0037] In step S7, the estimation of the at-risk population is achieved through overlay analysis. The high-risk area distribution map obtained in step S6 is spatially overlaid with the population spatial distribution data, and the population located within the high-risk areas is counted. The calculation process considers the spatial distribution characteristics of risk probability and the spatial differences in population density, and obtains an accurate estimate of the at-risk population through weighted calculation. The risk population assessment results provide a quantitative basis for environmental health risk management and support the formulation of precise risk prevention and control measures. By combining the risk assessment results with population distribution, an effective transformation from technical assessment to management application is achieved.
[0038] In a specific implementation of this embodiment, the collected soil samples are classified into three categories: agricultural land, construction land, and forest and grassland, based on the land use type. For agricultural land, different soil risk screening values are selected as the standard risk values according to different pH ranges of the soil. For construction land, the construction land screening value is selected as the standard risk value. For forest and grassland, the geochemical baseline of heavy metals in the soil is calculated using the relative cumulative frequency distribution method, and the lower limit of the abnormal concentration of heavy metals is used as the standard risk value. Specifically, for agricultural land, when the soil pH ≤ 5.5, the standard risk value is A; when 5.5 < pH ≤ 6.5, the standard risk value is B; when 6.5 < pH ≤ 7.5, the standard risk value is C; when pH > 7.5, the standard risk value is D. For construction land, the standard risk value for the first type of land is E, and the standard risk value for the second type of land is F. For forest and grassland, the soil heavy metal content in areas not strongly disturbed by human activities generally follows a lognormal distribution. As Figure 2 shown, in the double logarithmic distribution graph of the relative cumulative density and heavy metal elements, the upper limit of the natural background concentration is at the low inflection point of the distribution curve, and the lower limit of the abnormal concentration is at the high inflection point. The lower limit of the abnormal concentration of heavy metals is used as the standard risk value G for forest and grassland soils. By setting differentiated risk standards, it is possible to fully consider the exposure characteristics and risk levels under different land use types, and improve the pertinence and accuracy of risk assessment.
[0039] In a specific implementation of this embodiment, the climate factor dimension includes the annual average rainfall and the annual average temperature, and the human activity factor dimension includes the land use type, population density, gross domestic product per square kilometer, distance between the sampling point and the polluting enterprise, distance between the sampling point and the road, distance between the sampling point and the railway, and distance between the sampling point and the tailings pond. Specifically, the annual average rainfall and the annual average temperature are used to extract the climate data of the location of the sampling point through the Arcgis pro software. The land use type is extracted using Arcgis pro for the location of the sampling point, and the population density and gross domestic product per square kilometer are obtained by extracting the statistical data within the buffer range around the sampling point. The distances between the sampling point and the polluting enterprise, road, railway, and tailings pond are obtained by calculating the distance between the sampling point and the nearest facility through the Arcgis pro software. That is, through the integration of multi-source data, a comprehensive index system reflecting the intensity of human activity influence is constructed, providing rich characteristic information for the model.
[0040] In one specific embodiment of this example, the soil geological conditions dimension includes parent material type, soil type, altitude, slope, and landform type; the weathering factor dimension includes chemical alteration index, silica-alumina ratio, and silica-alumina-iron index. Specifically, the parent material type includes intermediate-acidic igneous rocks, Quaternary sandy clay, marine sedimentary clastic rocks, carbonate rocks, and terrestrial sedimentary clastic rocks; the soil type includes lateritic red soil, red soil, yellow soil, limestone soil, paddy soil, alluvial soil, and red soil; altitude and slope are extracted from DEM data using ArcGIS Pro software. The chemical alteration index (CIA) is calculated using the formula CIA = Al2O3 / (Al2O3 + CaO + Na2O + K2O) × 100%, the silica-alumina ratio (SA) is calculated using the formula SA = SiO2 / Al2O3, and the silica-alumina-iron index (SAF) is calculated using the formula SAF = SiO2 / (Al2O3 + Fe2O3), where each term represents the mole fraction of the corresponding oxide. By introducing geological background factors, the natural background level and geological control factors of heavy metals in the soil can be accurately reflected, thereby improving the scientific nature of risk assessment.
[0041] In one specific implementation of this embodiment, the average precision decline value is used to obtain the environmental factor weights. Specifically, the average precision decline value in random forests measures the importance of variables by calculating the difference in mean squared error. For each environmental factor, an importance score is calculated by comparing the change in model prediction accuracy when each factor is included and excluded. A higher importance score indicates a greater contribution of the factor to the model's prediction results. By calculating the environmental factor weights, the key factors that have the most significant impact on soil heavy metal risk can be identified, providing a scientific basis for determining the focus of risk prevention and control.
[0042] In one specific embodiment of this example, the formula for calculating the average decrease in accuracy is:
[0043] In the formula, OOB represents out-of-bag data, used to replace the test set error estimation method; MSE(OOB) ' iv The expression () represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree; while MSE(OOB) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree. iv ) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree after shuffling and reordering into the i-th tree; therefore, the higher the importance of the variable, the larger the MSE after rearranging the out-of-bag data.
[0044] Specifically, out-of-bag data refers to the sample data that was not selected when constructing each decision tree, and can be used to evaluate model performance instead of an independent test set. The contribution of each variable to the model's predictive accuracy is quantified by calculating the difference in out-of-bag error before and after randomizing the variables. During the calculation, the results of each tree are averaged to obtain a comprehensive importance score for each environmental factor. In other words, rigorous mathematical calculations ensure the objectivity and accuracy of the environmental factor weight assessment.
[0045] In one specific implementation of this embodiment, the prediction results of the prediction model are evaluated using a confusion matrix statistical method and ROC curves. High-risk and low-risk regions are distinguished by setting cutoff probabilities, which are determined by the Youden index. Specifically, receiver operating characteristics (ROC) can reflect the sensitivity and specificity of all possible cutoff probabilities. Sensitivity reflects the prediction model's ability to identify high-risk cases, i.e., the percentage of high-risk cases correctly identified in the prediction results; specificity refers to the prediction model's ability to identify low-risk cases, i.e., the percentage of low-risk cases correctly identified in the prediction results.
[0046] These indicators are calculated using a confusion matrix. The regional risk probability prediction model results can be categorized into three types: True Positive (TP), meaning a sample that is actually high-risk is also predicted to be high-risk; False Negative (FN), meaning a sample that is actually high-risk is predicted to be low-risk; False Positive (FP), meaning a sample that is actually low-risk is predicted to be high-risk; and True Negative (TN), meaning a sample that is actually low-risk is also predicted to be low-risk. The formulas for Sensitivity and Specificity are as follows: ; ; Table 1. Confusion Matrix and Related Classification Indicators
[0047] like Figure 3As shown, the ROC curve comprehensively evaluates the classification performance of a model by plotting the relationship between sensitivity and specificity under different cutoff probabilities. The closer the ROC curve is to the top left corner, the higher the model's accuracy. The point on the ROC curve closest to the top left corner has the fewest false positive (FP) and false negative (FN) samples, meaning the model is less likely to make an error. Calculating the area under the ROC curve (AUC) can intuitively represent the model's performance in the risk prediction region. AUC is generally between 0.5 (no predictive ability) and 1 (full predictive ability). An AUC between 0.5 and 0.7 indicates poor model discrimination, 0.7-0.85 indicates moderate model discrimination, and greater than 0.85 indicates very good discrimination.
[0048] like Figure 4 As shown, the cutoff probability of the predictive model at the point closest to the top left corner of the ROC curve is determined by the Youden index, a commonly used statistical indicator for evaluating diagnostic efficacy. The maximum value of the Youden index is defined as the optimal cutoff probability, balancing sensitivity and specificity to obtain the optimal threshold. Its calculation formula is:
[0049] The probability value that maximizes the Youden index is chosen as the optimal cutoff probability. A scientific model evaluation method ensures the accuracy of risk zone delineation.
[0050] In one specific embodiment of this example, by determining high-risk areas and a population spatial distribution map, the number of at-risk individuals in the high-risk areas is estimated, wherein the formula for calculating the at-risk population is:
[0051] POP affected It refers to the number of at-risk individuals in high-risk areas. >risk value It is the probability that the concentration of heavy metals in the soil exceeds the risk value, and Cut is the cutoff probability that distinguishes between high risk and low risk.
[0052] Specifically, the risk population calculation first determines the spatial scope of high-risk areas, and then extracts population distribution data within the corresponding areas. The calculation process considers the spatial heterogeneity of risk probabilities, calculating the total risk population through probability weighting. The calculation results provide a quantitative assessment of population exposure for environmental health risk assessment, supporting the formulation and optimization of risk management decisions. By transforming technical assessment results into management indicators, it achieves an effective transformation between scientific research and practical application.
[0053] Example 2 This invention also provides a machine learning-based system for assessing the regional risk probability of heavy metals in soil, comprising: The data collection and preprocessing module is used to collect basic information on soil heavy metal content, environmental variables, and land use status, and to clean, filter, format, and standardize the collected data. The risk standard setting module is used to classify the collected soil samples based on land use type and set corresponding soil heavy metal standard risk values for different land use types. The environmental variable collection module is used to collect spatially continuous environmental variable data as predictors; The risk assessment module is used to construct a pollution risk assessment model using a random forest classifier to obtain the regional risk hazard probability distribution of heavy metals in regional soil. The risk zone division module is used to determine the risk cutoff probability and divide high-risk and low-risk zones. The risk population assessment module is used to estimate the number of people at risk by combining risk probability distribution maps with population spatial distribution data.
[0054] Specifically, the data collection and preprocessing module receives and processes data through standardized data interfaces and processing procedures.
[0055] The risk standard setting module automatically selects the corresponding risk assessment standard based on the land use type to ensure the scientific nature of the assessment results.
[0056] The environmental variable collection module extracts environmental characteristic data from each sampling point through spatial analysis.
[0057] The risk assessment module, in conjunction with the data collection and preprocessing module, the risk standard setting module, and the environmental variable collection module, organizes and analyzes the data to train and optimize the model.
[0058] The risk zone delineation module automatically determines the optimal risk cutoff probability through statistical analysis.
[0059] The risk population assessment module uses spatial overlay analysis to automatically match and calculate risk areas with population distribution. Through modular system design, it achieves automated and standardized processing of soil heavy metal risk assessment.
[0060] In one specific implementation of this embodiment, the risk standard setting module categorizes soil samples into three types: agricultural land, construction land, and forest / grassland, and employs corresponding standard risk value setting methods for each type. Specifically, the module has a built-in risk assessment standard library for different land use types, including pH grading standards for agricultural land, land use type standards for construction land, and geochemical baseline calculation methods for forest / grassland. Based on the input land use type data, the system automatically calls the corresponding standard setting algorithm to calculate the standard risk value for each sampling point.
[0061] Example 3 In a specific application embodiment, the Guangxi region was selected as the research area, and the method of the present invention was used to conduct a soil heavy metal risk assessment.
[0062] After 500 model calculations, the final AUC scores for Cr and Ni in Guangxi soil were 0.85 and 0.86, respectively, demonstrating excellent model discrimination. This indicates that the regional risk hazard probability model we constructed can effectively distinguish between high-risk and low-risk areas for Cr and Ni in Guangxi soil. The cut-off probabilities for Cr and Ni in the prediction model were 0.23 and 0.24, respectively, confirming that areas in Guangxi soil with hazard probabilities greater than 0.23 and 0.24 can be considered high-risk areas.
[0063] Table 2. Population statistics of different prefecture-level cities exposed to high-risk areas of heavy metals (unit: 10,000 people)
[0064] Based on the regional risk probability of heavy metals and the population density map of Guangxi in 2020, the population density of different high-risk areas for heavy metals in 2020 was obtained. The number of people in different prefecture-level cities exposed to high-risk areas for heavy metals was also estimated. The statistical results are shown in Table 2. Overall, the number of people in the province at high risk of heavy metal exposure is ranked as follows: Cd > As > Ni > Cr > Zn > Pb > Cu > Hg. Approximately 3.3714 million people may be threatened by Cd, requiring strict control measures. Overall, Chongzuo, Nanning, Hechi, and Baise cities have a relatively large population affected by various heavy metals. Chongzuo, in particular, has a higher number of people at high risk from Ni, Cr, Zn, Cu, Cd, and As compared to other cities. In the province, approximately 1.4751 million people are exposed to high-risk areas for Ni, with about 405,500 in Nanning and 489,900 in Chongzuo. The remaining cities have fewer people at risk. Approximately 1.4436 million people in the province are threatened by Cr, with about 273,600 in Laibin, 289,900 in Hechi, 242,600 in Nanning, and 316,000 in Chongzuo being significantly affected. Approximately 1,358,100 people across the province are exposed to high-risk areas for the heavy metal Zn, with about 265,900 in Hechi City, 271,200 in Baise City, and 334,800 in Chongzuo City. The remaining cities have a smaller population at risk. The population affected by Pb is relatively small across the province, approximately 586,600. Baise City (112,400) and Chongzuo City (100,600) have relatively larger affected populations. Similarly, the impact of Hg is relatively small across the province, mainly concentrated in Hechi City (254,400). The same applies to Cu, mainly concentrated in Chongzuo City (313,600). Cd has little impact on Wuzhou City, Yulin City, Fangchenggang City, Beihai City, Qinzhou City, and Hezhou City, but a significant population in the remaining cities is threatened by Cd. Liuzhou City, in particular, has a population threatened by Cd of 702,300. As is mainly affected in Nanning City, with 1,073,300 people affected.
[0065] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent transformations or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for assessing the regional risk probability of heavy metals in soil, characterized in that, Includes the following steps: S1: Collect basic information on soil heavy metal content, environmental variables, and land use types; clean, filter, format, and standardize the collected data. S2: Soil samples collected are classified based on land use type, and corresponding standard risk values for heavy metals in soil are set for different risk receptor groups and their exposure pathways under different land use types. S3: Based on the relationship between the release and accumulation of heavy metals in soil, collect spatially continuous environmental variable data as predictive factors. The environmental variable data includes the dimensions of climate factors, human activity factors, soil geological conditions, and weathering factors. S4: A pollution risk assessment model is constructed using a random forest classifier, with the soil heavy metal risk value set as the response variable, and the original dataset is randomly divided into a training set and a test set. S5: The model is trained and tested based on the training subset and the test subset to obtain the regional risk hazard probability distribution of heavy metals in the regional soil. S6: Introduce the receiver operating characteristic curve to determine the risk cutoff probability. The region where the probability of harm is greater than the risk cutoff probability is a high-risk region, and the region where the probability of harm is less than the risk cutoff probability is a low-risk region. S7: Combine the regional soil heavy metal risk probability distribution map with population spatial distribution data to estimate the number of at-risk people living in high-risk areas.
2. The method according to claim 1, characterized in that, In S2, the collected soil samples are divided into three categories based on land use type: agricultural land, construction land, and forest and grassland. For agricultural land, different soil risk screening values are selected as standard risk values according to different pH ranges. For construction land, the construction land screening value is selected as the standard risk value. For forest and grassland, the geochemical baseline of heavy metals in the soil is calculated using the relative cumulative frequency distribution method, and the lower limit of the abnormal concentration of heavy metals is used as the standard risk value.
3. The method according to claim 1, characterized in that, The S3 climatic factor dimension includes annual average rainfall and annual average temperature, while the anthropogenic factor dimension includes land use type, population density, GDP per square kilometer, distance between sampling point and polluting enterprise, distance between sampling point and road, distance between sampling point and railway, and distance between sampling point and tailings dam.
4. The method according to claim 1, characterized in that, The soil geological conditions dimension in S3 includes parent material type, soil type, altitude, slope, and landform type, while the weathering factor dimension includes chemical alteration index, silicon-aluminum ratio, and silicon-aluminum-iron index.
5. The method according to claim 1, characterized in that, The environmental factor weights are obtained using the average decrease in precision in S4.
6. The method according to claim 5, characterized in that, The formula for calculating the average decrease in accuracy is as follows: In the formula, OOB refers to out-of-bag data; MSE(OOB) ' iv The expression () represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree; while MSE(OOB) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree. iv ) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable of the i-th tree after shuffling and reordering into the i-th tree.
7. The method according to claim 1, characterized in that, In S6, the confusion matrix statistical method and ROC curve are used to evaluate the prediction results of the prediction model. High-risk areas and low-risk areas are distinguished by setting the cutoff probability, which is determined by the Youden index.
8. The method according to claim 7, characterized in that, In step S7, by identifying high-risk areas and their spatial distribution, the number of at-risk individuals in high-risk areas is estimated. The formula for calculating the at-risk population is as follows: POP affected It refers to the number of at-risk individuals in high-risk areas. >risk value It is the probability that the concentration of heavy metals in the soil exceeds the risk value, and Cut is the cutoff probability that distinguishes between high risk and low risk.
9. A soil heavy metal regional risk probability assessment system based on machine learning, characterized in that, include: The data collection and preprocessing module is used to collect basic information on soil heavy metal content, environmental variables, and land use status, and to clean, filter, format, and standardize the collected data. The risk standard setting module is used to classify the collected soil samples based on land use type and set corresponding soil heavy metal standard risk values for different land use types. The environmental variable collection module is used to collect spatially continuous environmental variable data as predictors; The risk assessment module is used to construct a pollution risk assessment model using a random forest classifier to obtain the regional risk hazard probability distribution of heavy metals in regional soil. The risk zone division module is used to determine the risk cutoff probability and divide high-risk and low-risk zones. The risk population assessment module is used to estimate the number of people at risk by combining risk probability distribution maps with population spatial distribution data.
10. The system according to claim 9, characterized in that, The risk standard setting module divides soil samples into three categories: agricultural land, construction land, and forest and grassland, and adopts corresponding standard risk value setting methods for different types.
11. A soil heavy metal regional risk probability assessment device, characterized in that, The device includes a processor capable of executing a computer program that implements the soil heavy metal regional risk probability assessment method according to any one of claims 1-8.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Methods, devices, electronic equipment and media for delineating soil heavy metal pollution risk zones
CN114511239B
Soil heavy metal pollution risk zoning and grading method based on machine learning
CN119313154A
Method for sorting environment risk of abandoned plants
US20160041141A1