A soil heavy metal regional risk probability evaluation method, system, device and storage medium

A regional risk assessment method for heavy metals in soil, constructed using machine learning, combines multi-dimensional environmental variables and land use types. This method addresses the problem of inaccurate assessments in existing models at large scales, achieving high-precision risk area delineation and population assessment, and supporting the formulation of environmental management policies.

CN121034469BActive Publication Date: 2026-03-03CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing regional risk prediction models for heavy metals in soil lack flexibility and adaptability in large-scale, multi-receptor regions. They fail to effectively integrate multi-dimensional environmental variables and land use patterns, resulting in inaccurate risk assessment results and difficulty in providing scientific decision support.

Method used

A machine learning-based approach was adopted to construct a pollution risk assessment model using a random forest classifier. By combining multi-dimensional environmental variables and land use types, differentiated risk standards were set, and risk cutoff probabilities were determined by combining receiver operating characteristic curves and Youden index to estimate the population size in high-risk areas.

Benefits of technology

It has achieved accurate soil heavy metal risk assessment, improved the prediction accuracy and model adaptability in spatially variable areas, provided quantitative risk population assessment, and provided direct decision-making basis for environmental health risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034469B_ABST
    Figure CN121034469B_ABST
Patent Text Reader

Abstract

The application provides a soil heavy metal regional risk probability evaluation method, system and device and a storage medium, relates to the technical field of environmental risk assessment, and comprises the following steps: collecting information about the content of soil heavy metals; classifying the collected soil samples based on land use types; collecting spatially continuous environmental variable data as prediction factors; constructing a pollution risk evaluation model by using a random forest classifier; training and testing the model to obtain a risk hazard probability distribution; and combining the risk probability distribution graph with population spatial distribution data to estimate the number of risk populations living in high-risk areas. The application solves the problems of lacking differentiated risk standards, being difficult to integrate multidimensional environmental variables and being unable to quantitatively evaluate risk populations in the prior art, and realizes precise soil heavy metal risk evaluation and population exposure evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental risk assessment technology, specifically to a method, system, device, and storage medium for assessing the probability of heavy metal risk in soil regions. Background Technology

[0002] Assessing the regional environmental risk of heavy metals in soil is fundamental to developing regional risk control strategies. Regional environmental risk assessment is a method for analyzing and comparing the environmental impacts of different pollutants at a macro-scale geographical area, including the risk of sudden environmental pollution incidents and the probability of cumulative environmental pollution hazards. Among these, regional environmental risk assessment of heavy metals evolved from the health risk assessment of heavy metals in contaminated soil at contaminated sites, primarily examining the static risk of a specific pollutant at a specific point in time.

[0003] However, regional environmental risks are influenced by a combination of environmental risk subsystems, including risk source release, translocation, and risk receptor exposure, and these subsystems change significantly with social development. Most existing spatiotemporal risk prediction models require extensive regional measured environmental pollution data and historical data, focusing on the regional heavy metal pollution risk status of a single site or a site-specific event. This scale difference means that existing models and research findings are not suitable for large-scale, multi-receptor regional environmental risk prediction and simulation.

[0004] Furthermore, most existing studies on soil heavy metal risk prediction do not consider the impact of land use patterns on exposure models, relying solely on a single risk exposure model for evaluation. Ignoring the impact of land use change on the spatiotemporal prediction of soil heavy metal risk may lead policymakers to overestimate or underestimate the environmental risks in certain regions. Currently, there is a lack of accurate and effective analytical methods to address the problems of poor predictive performance, inflexibility, and adaptability of existing regional soil heavy metal risk prediction models for areas with high spatial variability. New theoretical foundations and data models are needed to resolve this current predicament.

[0005] Chinese patent document CN114511239B discloses a method, device, electronic device and medium for delineating soil heavy metal pollution risk zones. It discloses a technical solution for predicting the risk level of soil heavy metal pollution based on the random forest algorithm. It has the technical effect of improving the accuracy of risk zone delineation boundaries and solving the problem of strong dependence on the representativeness and quantity of sampling points in geostatistical spatial interpolation methods. However, it still has the problems of not considering the differences in risk receptors and exposure modes under different land use types, lacking the function of quantitative assessment of the number of at-risk populations, and failing to provide direct quantitative support for environmental management decisions.

[0006] Traditional methods for assessing heavy metal risk in soil face numerous challenges in dealing with complex spatial variability and multi-factor influences. On the one hand, existing methods struggle to effectively integrate multi-dimensional environmental variables such as climate factors, anthropogenic activity factors, and soil geological conditions, making it impossible to establish comprehensive risk prediction models. On the other hand, the lack of differentiated risk assessment standards for different land use types limits the accuracy and applicability of risk assessment results.

[0007] Meanwhile, existing technologies have significant shortcomings in assessing the size of at-risk populations. While high-risk areas can be identified, there is a lack of effective methods to combine risk areas with population distribution, failing to provide quantitative decision support for environmental health risk management. This leaves government departments without a scientific basis when formulating environmental management policies, making it difficult to achieve precise risk control. Summary of the Invention

[0008] The purpose of this invention is to provide a method, system, device, and storage medium for assessing the regional risk probability of heavy metals in soil, which solves the problems of lacking differentiated risk standards, difficulty in integrating multidimensional environmental variables, and inability to quantitatively assess at-risk populations in the prior art, and achieves accurate risk assessment of heavy metals in soil and evaluation of population exposure.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A method for assessing the probability of heavy metal risk in soil areas includes the following steps:

[0011] S1: Collect basic information on soil heavy metal content, environmental variables, and land use types; clean, filter, format, and standardize the collected data.

[0012] S2: Soil samples collected are classified based on land use type, and corresponding standard risk values ​​for heavy metals in soil are set for different risk receptor groups and their exposure pathways under different land use types.

[0013] S3: Based on the relationship between the release and accumulation of heavy metals in soil, collect spatially continuous environmental variable data as predictive factors. The environmental variable data includes the dimensions of climate factors, human activity factors, soil geological conditions, and weathering factors.

[0014] S4: A pollution risk assessment model is constructed using a random forest classifier, with the soil heavy metal risk value set as the response variable, and the original dataset is randomly divided into a training set and a test set.

[0015] S5: The model is trained and tested based on the training subset and the test subset to obtain the regional risk hazard probability distribution of heavy metals in the regional soil.

[0016] S6: Introduce the receiver operating characteristic curve to determine the risk cutoff probability. The region where the probability of harm is greater than the risk cutoff probability is a high-risk region, and the region where the probability of harm is less than the risk cutoff probability is a low-risk region.

[0017] S7: Combine the regional soil heavy metal risk probability distribution map with population spatial distribution data to estimate the number of at-risk people living in high-risk areas.

[0018] Furthermore, in S2, the collected soil samples are divided into three categories based on land use type: agricultural land, construction land, and forest and grassland. For agricultural land, different soil risk screening values ​​are selected as standard risk values ​​according to different pH ranges. For construction land, the construction land screening value is selected as the standard risk value. For forest and grassland, the geochemical baseline of heavy metals in the soil is calculated using the relative cumulative frequency distribution method, and the lower limit of the abnormal concentration of heavy metals is used as the standard risk value.

[0019] Furthermore, the climate factor dimension in S3 includes annual average rainfall and annual average temperature, while the anthropogenic activity factor dimension includes land use type, population density, GDP per square kilometer, distance between sampling point and polluting enterprise, distance between sampling point and road, distance between sampling point and railway, and distance between sampling point and tailings dam.

[0020] Furthermore, the soil geological conditions dimension in S3 includes parent material type, soil type, altitude, slope, and landform type, while the weathering factor dimension includes chemical alteration index, silicon-aluminum ratio, and silicon-aluminum-iron index.

[0021] Furthermore, in S4, the environmental factor weights are obtained using the average decrease in precision.

[0022] Furthermore: the formula for calculating the average decrease in accuracy is as follows:

[0023]

[0024] In the formula, OOB refers to out-of-bag data; MSE(OOB) ' iv The expression () represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree; while MSE(OOB) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree. iv ) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable of the i-th tree after shuffling and reordering into the i-th tree.

[0025] Furthermore, in S6, the prediction results of the prediction model are evaluated using the confusion matrix statistical method and ROC curve. High-risk areas and low-risk areas are distinguished by setting the cutoff probability, which is determined by the Youden index.

[0026] Furthermore: In step S7, by determining high-risk areas and population spatial distribution maps, the number of at-risk individuals in high-risk areas is estimated, wherein the formula for calculating the at-risk population is:

[0027]

[0028] POP affected It refers to the number of at-risk individuals in high-risk areas. >risk value It is the probability that the concentration of heavy metals in the soil exceeds the risk value, and Cut is the cutoff probability that distinguishes between high risk and low risk.

[0029] A machine learning-based system for assessing the regional risk probability of heavy metals in soil, comprising:

[0030] The data collection and preprocessing module is used to collect basic information on soil heavy metal content, environmental variables, and land use status, and to clean, filter, format, and standardize the collected data.

[0031] The risk standard setting module is used to classify the collected soil samples based on land use type and set corresponding soil heavy metal standard risk values ​​for different land use types.

[0032] The environmental variable collection module is used to collect spatially continuous environmental variable data as predictors;

[0033] The risk assessment module is used to construct a pollution risk assessment model using a random forest classifier to obtain the regional risk hazard probability distribution of heavy metals in regional soil.

[0034] The risk zone division module is used to determine the risk cutoff probability and divide high-risk and low-risk zones.

[0035] The risk population assessment module is used to estimate the number of people at risk by combining risk probability distribution maps with population spatial distribution data.

[0036] Furthermore, the risk standard setting module divides soil samples into three categories: agricultural land, construction land, and forest and grassland, and adopts corresponding standard risk value setting methods for different types.

[0037] A soil heavy metal regional risk probability assessment device includes a processor, which is capable of executing a computer program that can implement the above-mentioned soil heavy metal regional risk probability assessment method.

[0038] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the program is executed by a processor, it implements the above-described method for assessing the regional risk probability of heavy metals in soil.

[0039] Compared with the prior art, the present invention has the following advantages:

[0040] I. This invention constructs a targeted regional environmental risk prediction model for heavy metals in soil by comprehensively considering the spatial variability of heavy metals in regional soils and the differences in different land use types. Corresponding risk assessment standards are set for three different land use types: agricultural land, construction land, and forest and grassland. This fully considers the differences in risk receptors and exposure modes on different land types, making the risk assessment results more scientific and accurate, and providing reliable technical support for the formulation of precise soil pollution control and environmental management policies.

[0041] Second, this invention effectively solves the technical problems of existing soil heavy metal regional risk prediction models, such as poor prediction performance, lack of flexibility and adaptability in areas with high spatial variability. By introducing multi-dimensional environmental variable data and random forest machine learning algorithms, the model can clearly distinguish the transition characteristics from high-risk to low-risk areas, accurately capture the spatial volatility of data, and significantly improve the prediction accuracy and model adaptability for areas with strong spatial variability.

[0042] Third, this invention realizes a complete technical chain from risk identification to population impact assessment. By combining the probability distribution of heavy metal risks in soil with spatial distribution data of the population, it quantitatively estimates the number of people exposed to high-risk areas, providing a direct basis for decision-making in environmental health risk management. Compared with existing technologies that can only identify risk areas, this invention further quantifies the potential impact of risks on the population, achieving an effective transformation from technical assessment to management application.

[0043] Fourth, this invention establishes a hierarchical risk management zone division system and a risk change trend prediction mechanism. By using ROC curve analysis and Youden's index optimization to determine the optimal risk cutoff probability, it achieves a scientific division between high-risk and low-risk areas. This enables effective early warning of the environmental risk status of heavy metals in the soil of the study area, providing important theoretical basis and technical means for dynamic monitoring and proactive prevention and control of regional environmental risks. Attached Figure Description

[0044] Figure 1 A flowchart of a soil heavy metal regional risk probability assessment method based on machine learning provided by the present invention;

[0045] Figure 2 The double logarithmic distribution of relative cumulative density and heavy metal elements in a machine learning-based soil heavy metal regional risk probability assessment method provided by this invention;

[0046] Figure 3 ROC curve in a machine learning-based method for assessing the risk probability of heavy metals in soil regions provided by this invention;

[0047] Figure 4 The Youden index curve is shown in the soil heavy metal regional risk probability assessment method provided by the present invention based on machine learning. Detailed Implementation

[0048] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Example 1

[0050] like Figure 1 As shown: This invention provides a method for assessing the probability of heavy metal risk in soil areas based on machine learning, comprising the following steps:

[0051] S1: Collect basic information on soil heavy metal content, environmental variables, and land use types; clean, screen, format, and standardize the collected data. S2: Classify the collected soil samples based on land use types; set corresponding standard risk values ​​for soil heavy metals for different risk receptor groups and their exposure pathways under different land use types.

[0052] S3: Based on the relationship between the release and accumulation of heavy metals in soil, spatially continuous environmental variable data are collected as predictors. The environmental variable data include the dimensions of climate factors, human activity factors, soil geological conditions, and weathering factors. S4: A pollution risk assessment model is constructed using a random forest classifier. The soil heavy metal risk value is set as the response variable, and the original dataset is randomly divided into a training set and a test set.

[0053] S5: The model is trained and tested based on the training subset and the test subset to obtain the regional risk hazard probability distribution of heavy metals in the regional soil.

[0054] S6: Introduce the receiver operating characteristic curve to determine the risk cutoff probability. The region where the probability of harm is greater than the risk cutoff probability is a high-risk region, and the region where the probability of harm is less than the risk cutoff probability is a low-risk region.

[0055] S7: Combine the regional soil heavy metal risk probability distribution map with population spatial distribution data to estimate the number of at-risk people living in high-risk areas.

[0056] Specifically, in step S1, the basic data is the heavy metals in regional soil. The data can be obtained by grid sampling. The indicators include the total concentration of heavy metals in the soil and the soil pH value. The data collection process includes obtaining the heavy metal content data through the layout of sampling points, and at the same time collecting the environmental variable information and land use type information at the corresponding locations. In the data preprocessing stage, quality inspection is carried out on the collected original data, outliers and missing values are removed, and data with different dimensions are standardized to ensure the consistency and availability of the data format. That is, through systematic data collection and preprocessing, high-quality basic data support is provided for the subsequent construction of the risk assessment model.

[0057] In step S2, the soil samples are classified according to the "Classification of Current Land Use" standard. A differentiated standard risk value setting system is established for different risk receptor groups and exposure pathways.

[0058] Specifically, based on the "Classification of Current Land Use" (GB / T 21010-2017), the collected soil samples are divided into three categories, namely agricultural land, construction land, and forest and grassland. When setting different standard risk values according to different risk receptor groups and exposure pathways. For agricultural land, different soil risk screening values are selected as its standard risk value (C0) according to different pH ranges of the soil; when pH≤5.5, the standard risk value is C0 = A; when 5.5 < pH≤6.5, the standard risk value is C0 = B; when 6.5 < pH≤7.5, the standard risk value is C0 = C; when pH>7.5, the standard risk value is C0 = D; for construction land, the construction land screening value is selected as its standard risk value. For the first type of land, the standard risk value is C0 = E, and for the second type of land, the standard risk value is C0 = F; for forest and grassland, the relative cumulative frequency distribution method (CFD) is used to calculate the geochemical baseline (baseline, BL) of each heavy metal in the soil. In the areas where the soil heavy metal content is not strongly disturbed by human activities, the overall distribution is log-normal. In the double logarithmic distribution diagram of the relative cumulative density and heavy metal elements, the lower inflection point of the distribution curve is the upper limit of the natural background concentration, and the higher inflection point is the lower limit of the abnormal concentration. The lower limit of the heavy metal abnormal concentration is used as the standard risk value of forest and grassland soil C0 = G

[0059] There are significant differences in the exposure methods of the population under different land use types. Agricultural land is mainly exposed through the food chain, construction land is mainly exposed through direct contact and inhalation, and forest and grassland mainly consider the natural background level of the ecosystem. By establishing targeted risk standards, the actual risk status of different regions can be more accurately reflected, providing a scientific basis for judgment for subsequent risk assessment.

[0060] In step S3, the collection of environmental variable data is based on the migration and transformation patterns of heavy metals in the soil environment. Climatic factors influence the weathering and release process of heavy metals, anthropogenic factors reflect the contribution of human activities to heavy metal pollution, soil geological conditions determine the occurrence forms and migration capabilities of heavy metals, and weathering factors characterize the influence of geological background on heavy metal content. By constructing a multi-dimensional environmental variable system, various factors affecting the distribution of heavy metals in soil can be comprehensively reflected, providing rich feature information for machine learning models. In other words, the introduction of multi-dimensional environmental variables significantly improves the predictive accuracy and applicability of the risk assessment model.

[0061] In step S4, the construction process of the random forest classifier includes feature selection, model parameter setting, and dataset partitioning. Soil heavy metal risk values ​​are used as the target variable, and environmental variables are used as input features. The dataset is randomly partitioned into a training set (90%) and a test set (10%) to ensure the model's generalization ability. The random forest algorithm, by constructing multiple decision trees and combining their prediction results, can effectively handle high-dimensional data and nonlinear relationships, improving the stability and accuracy of prediction results. Through reasonable model construction, a technical foundation is laid for accurately predicting soil heavy metal risk.

[0062] In step S5, the model training process uses 500 repeated training and test subsets for cross-validation. During training, key parameters of the random forest, including the number of decision trees, maximum number of features, and maximum depth, are adjusted to optimize model performance. In the testing phase, the model's predictive ability is evaluated using an independent test set, and the risk probability value for each sampling point is calculated. By averaging the results of multiple training iterations, a stable and reliable regional risk hazard probability distribution map is obtained. In other words, through a systematic model training and validation process, the scientific validity and reliability of the risk assessment results are ensured.

[0063] In step S6, receiver operating characteristic (ROC) curve analysis is used to determine the optimal risk cutoff probability. The ROC curve reflects the model's sensitivity and specificity across all possible cutoff probabilities. The overall model performance is evaluated by calculating the area under the curve (AUC), with an AUC value closer to 1 indicating stronger discriminative ability. The risk cutoff probability is determined using the Youden index method, selecting the probability value that maximizes the sum of sensitivity and specificity as the optimal cutoff point. Based on the determined cutoff probability, the study area is divided into high-risk and low-risk levels, providing clear spatial boundaries for risk management.

[0064] In step S7, the estimation of the number of at-risk people is achieved through overlay analysis. The high-risk area distribution map obtained in step S6 is spatially overlaid with the population spatial distribution data, and the number of people located within the high-risk areas is counted. The calculation process takes into account the spatial distribution characteristics of the risk probability and the spatial differences in population density, and an accurate estimated value of the at-risk population is obtained through weighted calculation. The at-risk population assessment results provide a quantitative decision-making basis for environmental health risk management and support the formulation of precise risk prevention and control measures. By combining the risk assessment results with the population distribution, an effective transformation from technical assessment to management application is achieved.

[0065] In a specific implementation manner of this embodiment, the collected soil samples are divided into three categories: agricultural land, construction land, and forest and grassland, based on the land use type. For agricultural land, different soil risk screening values are selected as the standard risk values according to different pH ranges of the soil. For construction land, the construction land screening value is selected as the standard risk value. For forest and grassland, the geochemical baseline of heavy metals in the soil is calculated using the relative cumulative frequency distribution method, and the lower limit of the abnormal concentration of heavy metals is used as the standard risk value. Specifically, for agricultural land, when the soil pH ≤ 5.5, the standard risk value is A; when 5.5 < pH ≤ 6.5, the standard risk value is B; when 6.5 < pH ≤ 7.5, the standard risk value is C; when pH > 7.5, the standard risk value is D. For construction land, the standard risk value for the first type of land is E, and the standard risk value for the second type of land is F. For forest and grassland, the heavy metal content in the soil in areas not strongly disturbed by human activities generally follows a lognormal distribution. As Figure 2 shown, in the double logarithmic distribution graph of the relative cumulative density and heavy metal elements, the upper limit of the natural background concentration is at the low inflection point of the distribution curve, and the lower limit of the abnormal concentration is at the high inflection point. The lower limit of the abnormal concentration of heavy metals is used as the standard risk value G for forest and grassland soils. By setting differentiated risk standards, it is possible to fully consider the exposure characteristics and risk levels under different land use types and improve the pertinence and accuracy of risk assessment.

[0066] In one specific implementation of this embodiment, the climate factor dimension includes annual average rainfall and annual average temperature, while the anthropogenic activity factor dimension includes land use type, population density, GDP per square kilometer, distance between the sampling point and polluting enterprises, distance between the sampling point and roads, distance between the sampling point and railways, and distance between the sampling point and tailings ponds. Specifically, annual average rainfall and annual average temperature are extracted from climate data at the sampling point location using ArcGIS Pro software. Land use type is extracted from the land use type at the sampling point location using ArcGIS Pro, and population density and GDP per square kilometer are obtained by extracting statistical data within the buffer zone surrounding the sampling point. The distances between the sampling point and polluting enterprises, roads, railways, and tailings ponds are calculated using ArcGIS Pro software to determine the distance between the sampling point and the nearest facility. In other words, by integrating multi-source data, a comprehensive indicator system reflecting the intensity of anthropogenic activity impacts is constructed, providing rich feature information for the model.

[0067] In one specific embodiment of this example, the soil geological conditions dimension includes parent material type, soil type, altitude, slope, and landform type; the weathering factor dimension includes chemical alteration index, silica-alumina ratio, and silica-alumina-iron index. Specifically, the parent material type includes intermediate-acidic igneous rocks, Quaternary sandy clay, marine sedimentary clastic rocks, carbonate rocks, and terrestrial sedimentary clastic rocks; the soil type includes lateritic red soil, red soil, yellow soil, limestone soil, paddy soil, alluvial soil, and red soil; altitude and slope are extracted from DEM data using ArcGIS Pro software. The chemical alteration index (CIA) is calculated using the formula CIA = Al2O3 / (Al2O3 + CaO + Na2O + K2O) × 100%, the silica-alumina ratio (SA) is calculated using the formula SA = SiO2 / Al2O3, and the silica-alumina-iron index (SAF) is calculated using the formula SAF = SiO2 / (Al2O3 + Fe2O3), where each term represents the mole fraction of the corresponding oxide. By introducing geological background factors, the natural background level and geological control factors of heavy metals in the soil can be accurately reflected, thereby improving the scientific nature of risk assessment.

[0068] In one specific implementation of this embodiment, the average precision decline value is used to obtain the environmental factor weights. Specifically, the average precision decline value in random forests measures the importance of variables by calculating the difference in mean squared error. For each environmental factor, an importance score is calculated by comparing the change in model prediction accuracy when each factor is included and excluded. A higher importance score indicates a greater contribution of the factor to the model's prediction results. By calculating the environmental factor weights, the key factors that have the most significant impact on soil heavy metal risk can be identified, providing a scientific basis for determining the focus of risk prevention and control.

[0069] In one specific embodiment of this example, the formula for calculating the average decrease in accuracy is:

[0070]

[0071] In the formula, OOB represents out-of-bag data, used to replace the test set error estimation method; MSE(OOB) ' iv The expression () represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree; while MSE(OOB) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree into the i-th tree. iv ) represents the mean squared error calculated by substituting the out-of-bag data of the j-th variable in the i-th tree after shuffling and reordering into the i-th tree; therefore, the higher the importance of the variable, the larger the MSE after rearranging the out-of-bag data.

[0072] Specifically, out-of-bag data refers to the sample data that was not selected when constructing each decision tree, and can be used to evaluate model performance instead of an independent test set. The contribution of each variable to the model's predictive accuracy is quantified by calculating the difference in out-of-bag error before and after randomizing the variables. During the calculation, the results of each tree are averaged to obtain a comprehensive importance score for each environmental factor. In other words, rigorous mathematical calculations ensure the objectivity and accuracy of the environmental factor weight assessment.

[0073] In one specific implementation of this embodiment, the prediction results of the prediction model are evaluated using a confusion matrix statistical method and ROC curves. High-risk and low-risk regions are distinguished by setting cutoff probabilities, which are determined by the Youden index. Specifically, receiver operating characteristics (ROC) can reflect the sensitivity and specificity of all possible cutoff probabilities. Sensitivity reflects the prediction model's ability to identify high-risk cases, i.e., the percentage of high-risk cases correctly identified in the prediction results; specificity refers to the prediction model's ability to identify low-risk cases, i.e., the percentage of low-risk cases correctly identified in the prediction results.

[0074] These indicators are calculated using a confusion matrix. The regional risk probability prediction model results can be categorized into three types: True Positive (TP), meaning a sample that is actually high-risk is also predicted to be high-risk; False Negative (FN), meaning a sample that is actually high-risk is predicted to be low-risk; False Positive (FP), meaning a sample that is actually low-risk is predicted to be high-risk; and True Negative (TN), meaning a sample that is actually low-risk is also predicted to be low-risk. The formulas for Sensitivity and Specificity are as follows:

[0075] ;

[0076] ;

[0077] Table 1. Confusion Matrix and Related Classification Indicators

[0078]

[0079] like Figure 3 As shown, the ROC curve comprehensively evaluates the classification performance of a model by plotting the relationship between sensitivity and specificity under different cutoff probabilities. The closer the ROC curve is to the top left corner, the higher the model's accuracy. The point on the ROC curve closest to the top left corner has the fewest false positive (FP) and false negative (FN) samples, meaning the model is less likely to make an error. Calculating the area under the ROC curve (AUC) can intuitively represent the model's performance in the risk prediction region. AUC is generally between 0.5 (no predictive ability) and 1 (full predictive ability). An AUC between 0.5 and 0.7 indicates poor model discrimination, 0.7-0.85 indicates moderate model discrimination, and greater than 0.85 indicates very good discrimination.

[0080] like Figure 4 As shown, the cutoff probability of the predictive model at the point closest to the top left corner of the ROC curve is determined by the Youden index, a commonly used statistical indicator for evaluating diagnostic efficacy. The maximum value of the Youden index is defined as the optimal cutoff probability, balancing sensitivity and specificity to obtain the optimal threshold. Its calculation formula is:

[0081]

[0082] The probability value that maximizes the Youden index is chosen as the optimal cutoff probability. A scientific model evaluation method ensures the accuracy of risk zone delineation.

[0083] In one specific embodiment of this example, by determining high-risk areas and a population spatial distribution map, the number of at-risk individuals in the high-risk areas is estimated, wherein the formula for calculating the at-risk population is:

[0084]

[0085] POP affected It refers to the number of at-risk individuals in high-risk areas. >risk value It is the probability that the concentration of heavy metals in the soil exceeds the risk value, and Cut is the cutoff probability that distinguishes between high risk and low risk.

[0086] Specifically, the risk population calculation first determines the spatial scope of high-risk areas, and then extracts population distribution data within the corresponding areas. The calculation process considers the spatial heterogeneity of risk probabilities, calculating the total risk population through probability weighting. The calculation results provide a quantitative assessment of population exposure for environmental health risk assessment, supporting the formulation and optimization of risk management decisions. By transforming technical assessment results into management indicators, it achieves an effective transformation between scientific research and practical application.

[0087] Example 2

[0088] This invention also provides a machine learning-based system for assessing the regional risk probability of heavy metals in soil, comprising:

[0089] The data collection and preprocessing module is used to collect basic information on soil heavy metal content, environmental variables, and land use status, and to clean, filter, format, and standardize the collected data.

[0090] The risk standard setting module is used to classify the collected soil samples based on land use type and set corresponding soil heavy metal standard risk values ​​for different land use types.

[0091] The environmental variable collection module is used to collect spatially continuous environmental variable data as predictors;

[0092] The risk assessment module is used to construct a pollution risk assessment model using a random forest classifier to obtain the regional risk hazard probability distribution of heavy metals in regional soil.

[0093] The risk zone division module is used to determine the risk cutoff probability and divide high-risk and low-risk zones.

[0094] The risk population assessment module is used to estimate the number of people at risk by combining risk probability distribution maps with population spatial distribution data.

[0095] Specifically, the data collection and preprocessing module receives and processes data through standardized data interfaces and processing procedures.

[0096] The risk standard setting module automatically selects the corresponding risk assessment standard based on the land use type to ensure the scientific nature of the assessment results.

[0097] The environmental variable collection module extracts environmental characteristic data from each sampling point through spatial analysis.

[0098] The risk assessment module, in conjunction with the data collection and preprocessing module, the risk standard setting module, and the environmental variable collection module, organizes and analyzes the data to train and optimize the model.

[0099] The risk zone delineation module automatically determines the optimal risk cutoff probability through statistical analysis.

[0100] The risk population assessment module uses spatial overlay analysis to automatically match and calculate risk areas with population distribution. Through modular system design, it achieves automated and standardized processing of soil heavy metal risk assessment.

[0101] In one specific implementation of this embodiment, the risk standard setting module categorizes soil samples into three types: agricultural land, construction land, and forest / grassland, and employs corresponding standard risk value setting methods for each type. Specifically, the module has a built-in risk assessment standard library for different land use types, including pH grading standards for agricultural land, land use type standards for construction land, and geochemical baseline calculation methods for forest / grassland. Based on the input land use type data, the system automatically calls the corresponding standard setting algorithm to calculate the standard risk value for each sampling point.

[0102] Example 3

[0103] In a specific application embodiment, the Guangxi region was selected as the research area, and the method of the present invention was used to conduct a soil heavy metal risk assessment.

[0104] After 500 model calculations, the final AUC scores for Cr and Ni in Guangxi soil were 0.85 and 0.86, respectively, demonstrating excellent model discrimination. This indicates that the regional risk hazard probability model we constructed can effectively distinguish between high-risk and low-risk areas for Cr and Ni in Guangxi soil. The cut-off probabilities for Cr and Ni in the prediction model were 0.23 and 0.24, respectively, confirming that areas in Guangxi soil with hazard probabilities greater than 0.23 and 0.24 can be considered high-risk areas.

[0105] Table 2. Population statistics of different prefecture-level cities exposed to high-risk areas of heavy metals (unit: 10,000 people)

[0106]

[0107] Based on the regional risk probability of heavy metals and the population density map of Guangxi in 2020, the population density of different high-risk areas for heavy metals in 2020 was obtained. The number of people in different prefecture-level cities exposed to high-risk areas for heavy metals was also estimated. The statistical results are shown in Table 2.

[0108] Overall, the number of people in the province at high risk of heavy metal exposure is ranked as follows: Cd > As > Ni > Cr > Zn > Pb > Cu > Hg. Approximately 3.3714 million people may be threatened by Cd, requiring strict control measures. Overall, Chongzuo, Nanning, Hechi, and Baise cities have a relatively large population affected by various heavy metals. Chongzuo, in particular, has a higher number of people at high risk from Ni, Cr, Zn, Cu, Cd, and As compared to other cities. In the province, approximately 1.4751 million people are exposed to high-risk areas for Ni, with about 405,500 in Nanning and 489,900 in Chongzuo. The remaining cities have fewer people at risk. Approximately 1.4436 million people in the province are threatened by Cr, with about 273,600 in Laibin, 289,900 in Hechi, 242,600 in Nanning, and 316,000 in Chongzuo being significantly affected. Approximately 1,358,100 people across the province are exposed to high-risk areas for the heavy metal Zn, with about 265,900 in Hechi City, 271,200 in Baise City, and 334,800 in Chongzuo City. The remaining cities have a smaller population at risk. The population affected by Pb is relatively small across the province, approximately 586,600. Baise City (112,400) and Chongzuo City (100,600) have relatively larger affected populations. Similarly, the impact of Hg is relatively small across the province, mainly concentrated in Hechi City (254,400). The same applies to Cu, mainly concentrated in Chongzuo City (313,600). Cd has little impact on Wuzhou City, Yulin City, Fangchenggang City, Beihai City, Qinzhou City, and Hezhou City, but a significant population in the remaining cities is threatened by Cd. Liuzhou City, in particular, has a population threatened by Cd of 702,300. As is mainly affected in Nanning City, with 1,073,300 people affected.

[0109] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent transformations or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for assessing regional risk probability of heavy metals in soil, characterized in that, The method comprises the following steps: S1: collecting basic information of soil heavy metal content, environmental variables and land use types, and performing cleaning, screening, formatting and standardization processing on the collected data; S2: classifying the collected soil samples based on land use types, and setting corresponding soil heavy metal standard risk values for different risk receptor groups and their exposure pathways under different land use types; S3: collecting spatially continuous environmental variable data as prediction factors based on the release and accumulation relationship of heavy metals in soil, wherein the environmental variable data includes climate factor dimension, human activity factor dimension, soil geological condition dimension and weathering factor dimension; S4: constructing a pollution risk assessment model by using a random forest classifier, setting soil heavy metal standard risk values as response variables, and randomly dividing an original data set into a training set and a test set; S5: training and testing the model based on the training subset and the test subset to obtain a regional risk hazard probability distribution of regional soil heavy metals; S6: introducing a receiver operating characteristic curve to determine a risk cutoff probability, wherein a region with a hazard probability greater than the risk cutoff probability is a high-risk region, and a region with a hazard probability less than the risk cutoff probability is a low-risk region; S7: combining the regional soil heavy metal risk probability distribution map with population spatial distribution data to estimate the number of risk populations living in the high-risk region; In S6, the prediction results of the prediction model are evaluated by using a confusion matrix statistical method and a receiver operating characteristic curve, a cutoff probability is set to distinguish the high-risk region and the low-risk region, and the cutoff probability is determined by a Youden index; In S7, the number of risk populations in the high-risk region is estimated by determining the high-risk region and the population spatial distribution map, wherein the formula for calculating the risk population is: where POP affected is the number of risk population in the high-risk area, Prob >risk value is the probability of the heavy metal concentration in the soil exceeding the risk value, and Cut is the cut-off probability to distinguish the high-risk from the low-risk.

2. The method of claim 1, wherein, In S2, the collected soil samples are divided into three categories of agricultural land, construction land and forest and grassland based on land use types, different soil risk screening values are selected as standard risk values for agricultural land according to different pH ranges of soil, a construction land screening value is selected as a standard risk value for construction land, and a relative cumulative frequency distribution method is used to calculate the geochemical baseline of heavy metals in soil, and the lower limit of the abnormal concentration of heavy metals is taken as the standard risk value.

3. The method of claim 1, wherein, In S3, the climate factor dimension includes annual average rainfall and annual average temperature, the human activity factor dimension includes land use type, population density, total output value per square kilometer, distance between the sampling point and the pollution enterprise, distance between the sampling point and the road, distance between the sampling point and the railway, and distance between the sampling point and the tailings.

4. The method of claim 1, wherein, In S3, the soil geological condition dimension includes parent material type, soil type, altitude, slope and landform type, and the weathering factor dimension includes chemical alteration index, silicon-aluminum ratio and silicon-aluminum-iron index.

5. The method of claim 1, wherein, In S4, the precision average drop value is used to obtain the weight of the environmental factor.

6. The method of claim 5, wherein, The calculation formula of the precision average drop value is: wherein, OOB in the formula is out-of-bag data; MSE(OOB ' iv ) represents the average square error calculated after the out-of-bag data of the jth variable of the ith tree is substituted into the ith tree; and MSE(OOB iv ) represents the average square error calculated after the out-of-bag data of the jth variable of the ith tree after being shuffled and reordered is substituted into the ith tree.

7. A machine learning based regional risk probability assessment system for heavy metals in soil, applying the method of any one of claims 1 to 6, characterized in that, It comprises: a data collection and preprocessing module, which is used for collecting basic information of soil heavy metal content, environmental variables and land use conditions, and performing cleaning, screening, formatting and standardization processing on the collected data; The risk standard setting module is configured to classify the collected soil samples based on land use types and set corresponding soil heavy metal standard risk values for different land use types; The environmental variable collection module is configured to collect spatially continuous environmental variable data as a predictor; The risk assessment module is configured to construct a pollution risk assessment model by using a random forest classifier to obtain a regional risk hazard probability distribution of the regional soil heavy metals; The risk region division module is configured to determine a risk cutoff probability and divide the region into a high-risk region and a low-risk region; The risk population assessment module is configured to combine the risk probability distribution map with population spatial distribution data to estimate the number of risk populations.

8. The system of claim 7, wherein, The risk standard setting module divides the soil samples into three categories of agricultural land, construction land and forest and grassland, and sets the standard risk values for different types by using corresponding setting methods.

9. A device for assessing regional risk probability of heavy metals in soil, characterized by, The computer program can implement the soil heavy metal regional risk probability assessment method of any one of claims 1-6.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Methods, devices, electronic equipment and media for delineating soil heavy metal pollution risk zones

    CN114511239B

  • Soil heavy metal pollution risk zoning and grading method based on machine learning

    CN119313154A

  • Method for sorting environment risk of abandoned plants

    US20160041141A1