A system and method for constructing a driving model for soil heavy metal accumulation

By constructing a system and method for driving model for soil heavy metal accumulation, the problem of difficult to identify and quantify soil heavy metal driving factors in the prior art is solved, and more accurate analysis and control of the soil heavy metal accumulation formation process is achieved.

CN116187075BActive Publication Date: 2025-05-16KUNMING UNIV OF SCI & TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310204766.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-06
Publication Date
2025-05-16
Estimated Expiration
2043-03-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and quantify drivers in soil heavy metal accumulation formation, especially when dealing with data on categorical variables and nonlinear relationships.

Method used

A construction system and method for forming a driving model for soil heavy metal accumulation is adopted, including information acquisition unit, matrix determination unit, potential driving factor acquisition unit, main control driving factor determination unit and driving model establishment unit. The soil heavy metal data was analyzed through PMF software, and the potential driver factors were obtained using GIS core density analysis and spatial extraction technology, and the main driver factors were determined using geodetectors, and a driving model was established through the gradient enhancement decision tree method.

Benefits of technology

This method can effectively identify and quantify the main control drivers formed by soil heavy metal accumulation, make up for the shortcomings of traditional methods when dealing with categorical variables and nonlinear relationship data, and provides more accurate source control of soil heavy metal pollution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116187075B_ABST
    Figure CN116187075B_ABST
Patent Text Reader

Abstract

The invention relates to a system and method for constructing a soil heavy metal accumulation driving model. The system comprises an information acquisition unit, a matrix determination unit, a potential driving factor acquisition unit, a main driving factor determination unit and a driving model establishment unit. The information acquisition unit is used to acquire soil heavy metal data, soil physical and chemical property data and point information of multiple sampling points in a region. The potential driving factor acquisition unit is used to acquire the name and spatial position of the potential driving factor in the region by using GIS kernel density analysis and spatial extraction technology. The main driving factor determination unit is used to analyze the independent effect and interaction of regional potential driving factors on the spatial distribution of hot areas by using geographic detectors, and determine the main driving factor. The driving model establishment unit is used to establish a soil heavy metal accumulation driving model based on the main driving factor by using a gradient boosting decision tree method, and acquire the driving force contribution of the main driving factor to the soil heavy metal accumulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil pollution prevention and control in the ecological environment and environmental protection industry, and in particular to a system and method for constructing a driving model for the accumulation of heavy metals in soil. Background Art

[0002] The following content only represents the background technology known to the applicant and does not constitute an admission that it is prior art known in the art and readily available to the public.

[0003] The accumulation and formation process of heavy metals in soil is complex, mainly including human activities and natural processes. Human activities such as industrial production, agricultural activities and transportation have exceeded natural processes such as weathering of soil parent materials to some extent. After heavy metals enter the soil, they exist in the soil in different forms through processes such as exchange, adsorption and precipitation. In the process of heavy metal accumulation in soil, it is particularly important to correctly identify and quantify the driving force and driving contribution rate of external driving factors for the control of the source of heavy metal pollution in soil. At present, the analysis methods for the driving mechanism of heavy metal accumulation in soil mainly include multivariate statistical analysis, graphical methods, geochemical simulation methods and other methods. These methods mainly use linear mathematical statistics to distinguish the type of driving factor, but the data statistical method has high requirements on the type and amount of data, requiring the independent variable and the dependent variable to be linearly related, and cannot process driving factor data such as categorical variables. Therefore, the present invention provides a system and method for constructing a driving model for the accumulation and formation of heavy metals in soil, which can analyze and process the driving factors of categorical variables, and can also process the driving factors of nonlinear relationships, making up for the shortcomings of traditional methods with high data requirements. Summary of the invention

[0004] The present invention aims to provide a system and method for constructing a driving model for the accumulation of heavy metals in soil, so as to solve the problem that the driving factors in the process of accumulation of heavy metals in soil are difficult to identify and quantify.

[0005] The purpose of the present invention is to solve the deficiencies of the prior art and provide a system for constructing a driving model for the accumulation of heavy metals in soil, including an information acquisition unit, a matrix determination unit, a potential driving factor acquisition unit, a main control driving factor determination unit and a driving model establishment unit; the information acquisition unit is used to obtain soil heavy metal data, soil physical and chemical property data and point information of multiple sampling points in a region; the matrix determination unit is used to analyze the soil heavy metal content of multiple sampling points in a region according to a first formula using PMF software, and determine the potential driving factor category matrix and potential driving factor score matrix for the accumulation of heavy metals in the region;

[0006] The first formula is:

[0007]

[0008] Where i represents multiple sampling points in the region, with a total of n sampling points, j represents soil heavy metals in the region, with a total of m soil heavy metals, k represents the category of potential driving factors, with a total of p categories, and x ij Indicates the content of heavy metals in the soil of sampling point j in the region, in mg / kg, g ik represents the contribution of the k-type driving factor in the i-th sampling point in the region, that is, the score matrix of the potential driving factor category, f kj represents the content of heavy metals in soil j in category k of potential driving factors, e ij is the residual; u ij represents the uncertainty of soil heavy metals at sampling point j in the region; Q represents the minimum value of the model; ik ≥0, f kj Under the constraint of ≥0, Q is solved by iterative minimization algorithm to determine the driving factor score matrix g ik and the driving factor category matrix f kj ;

[0009] The potential driving factor acquisition unit is used to use GIS kernel density analysis and spatial extraction technology to obtain the name and spatial location of the potential driving factor in the region according to the second formula;

[0010] The second formula is:

[0011]

[0012] Where f(g) is the kernel density estimate of the potential driving factor score; h is the search radius of the calculation, also known as the bandwidth; n is the number of feature points within the region whose distance is less than h around the potential driving factor score g at a certain spatial location; g i is the mean of the potential driving factor scores, g is the observed value of the potential driving factor scores, and the K function represents the spatial calculation weight function; the larger the search radius or bandwidth, the smoother the kernel density function curve and the lower the estimation accuracy; conversely, the smaller the bandwidth, the less smooth the kernel density function curve and the higher the estimation accuracy;

[0013] The main driving factor determination unit is used to use the geographic detector to analyze the independent effects and interactive effects of the regional potential driving factors on the spatial distribution of the factor score hotspot according to the third formula to determine the main driving factor;

[0014] The third formula is:

[0015]

[0016] Where h = 1, 2, ..., L is the stratification of potential driving factor scores, N his the number of units in layer h, N is the number of units in the region; σ h 2 is the variance of the driver factor score of layer h; σ 2 is the variance of the driver factor score for the entire region;

[0017] The driving model establishment unit is used to establish a driving model for soil heavy metal accumulation based on the main driving factor using a gradient boosting decision tree method, and obtain the driving force contribution of the main driving factor to soil heavy metal accumulation;

[0018] The driving model for soil heavy metal accumulation is as follows:

[0019] Y i =a1f1(x0,x i ,...x m )+a2f2(x0,x i ,...x m )+…+a n f n (x0,x i ,...x m )+a0;

[0020] where a1, a2, ...a n Represents the weight coefficient of n decision trees, a0 represents the initial value, x0, x1, x2…x m represents m master driving factors, fn(x0,x1,x2,…x m ) represents the function of m main driving factors, that is, the predicted value of n trees;

[0021] The driving force contribution of the main driving factor to the accumulation of heavy metals in soil is calculated according to the fifth formula, which is:

[0022]

[0023] Where i represents the decision tree, the number is n, j represents the non-leaf node in the decision tree, the number is l, s ij Represents the reduction in squared loss of j non-leaf nodes in i trees.

[0024] The present invention also provides a method for constructing a soil heavy metal accumulation driving model, which mainly includes the following steps:

[0025] (1) Obtain soil heavy metal content, soil physical and chemical properties content and location information at multiple sampling points in the region;

[0026] (2) Obtain the category matrix and score matrix of potential driving factors for regional soil heavy metal accumulation through PMF software;

[0027] (3) Based on the ArcGIS platform, the hot zone analysis of the driving factor categories was carried out through kernel density analysis and spatial analysis technology, and the potential driving factor data of the region was spatialized to obtain the spatial data of the potential driving factors, and the names and spatial locations of the driving factors for soil heavy metal accumulation were obtained;

[0028] (4) Use geographic detectors to analyze the independent and interactive effects of spatial data of potential driving factors and the physical and chemical properties of sampling points on the spatial distribution of potential driving factor scores, and determine the main controlling driving factors for soil heavy metal accumulation;

[0029] (5) Based on the main driving factors and the soil heavy metal accumulation data at multiple sampling points in the region, a soil heavy metal accumulation driving model was established using the gradient boosting decision tree method to obtain the driving force contribution rate of the main driving factors.

[0030] Optionally, the basic data acquisition steps are as follows:

[0031] Field sampling, obtaining soil samples from multiple sampling points in the area and accurately locating them using GPS;

[0032] Determine the heavy metal contents of soil samples at multiple sampling points, including As, Ba, Cu, Cd, Cr, Cu, Cs, Mn, Mo, Ni, Ni, Sr, Sn, Sb, Pb and Zn, soil pH, soil CEC, soil Fe2O3, TN, TC and TP and other physical and chemical properties data, and record the location information of multiple sampling points.

[0033] Optionally, the steps to obtain the potential driver category matrix and score matrix are as follows:

[0034] The PMF model was used to process the soil heavy metal content data of multiple sampling points to obtain the potential driving factor category matrix and potential driving factor score matrix of soil heavy metal accumulation.

[0035] Optionally, the steps for obtaining the specific names and spatial locations of multiple potential driving factors in the region are as follows:

[0036] The spatial distribution of hotspots of potential driving factor scores was obtained using kernel density analysis in ArcGIS platform;

[0037] The potential driving factor category raster data was converted into spatial vectors through ArcGIS spatial analysis;

[0038] Spatial data were extracted using ArcGIS technology. If the potential driving factor was a continuous variable, the nearest neighbor analysis technique was used to extract the distance from the potential factor to the sampling point. If the potential driving factor was a categorical variable, spatial analysis was used to extract the attribute data of the sampling point.

[0039] According to the hot zone distribution of potential driving factor scores and the spatialization data of potential driving factors, the specific name and spatial location of the potential driving factors in the region are determined.

[0040] Optionally, the specific steps for obtaining the regional master driving factor are as follows:

[0041] Taking the specific spatial data of potential driving factors in the region as the independent variable and the potential driving factor score as the dependent variable, the geographic detector factor detection module is used to analyze the influence of potential driving factors on factor scores, and obtain the main controlling driving factors of soil heavy metal accumulation in the region.

[0042] Optionally, the steps for establishing a driving model for soil heavy metal accumulation are as follows:

[0043] Using the spatial data of the main driving factors and the physical and chemical properties of the soil as independent variables and the soil heavy metal content data of multiple sampling points as the dependent variable, the gradient boosting decision model was used to establish a driving model for the accumulation of heavy metals in soil, and the contribution of the main driving factors to the driving force of the accumulation of heavy metals in soil was obtained.

[0044] The coefficient of determination and the average prediction root mean square error were calculated to verify the accuracy of the model.

[0045] The present invention also provides a method for constructing a soil heavy metal accumulation driving model, comprising the following steps:

[0046] The first step is to obtain soil heavy metal data, soil physical and chemical properties data and point information at multiple sampling points in the area;

[0047] The second step is to use PMF software to analyze the soil heavy metal content of multiple sampling points in the region according to the first formula, and determine the potential driving factor category matrix and potential driving factor score matrix for heavy metal accumulation in the region;

[0048] The first formula is:

[0049]

[0050] Among them, i represents multiple sampling points in the region, with a total of n sampling points, j represents soil heavy metals in the region, with a total of m soil heavy metals, k represents the category of potential driving factors, with a total of p categories, and x ij Indicates the content of heavy metals in the soil of sampling point j in the region, in mg / kg, g ik represents the contribution of the k-type driving factor in the i-th sampling point in the region, that is, the score matrix of the potential driving factor category, f kj represents the content of heavy metal j in category k of potential driving factors, e ij is the residual; u ijrepresents the uncertainty of heavy metals at sampling point j in the region i; Q represents the minimum value of the model; ik ≥0, f kj Under the constraint of ≥0, Q is solved by iterative minimization algorithm to determine the driving factor score matrix g ik and the driving factor category matrix f kj ;

[0051] Step 3: Using GIS kernel density analysis and spatial extraction technology, the names and spatial locations of potential driving factors in the region are obtained according to the second formula;

[0052] The second formula is:

[0053]

[0054] Where f(g) is the kernel density estimate of the potential driving factor score; h is the search radius of the calculation, also known as the bandwidth; n is the number of feature points within the region whose distance is less than h around the potential driving factor score g at a certain spatial location; g i is the mean of the potential driving factor scores, g is the observed value of the potential driving factor scores, and the K function represents the spatial calculation weight function; the larger the search radius or bandwidth, the smoother the kernel density function curve and the lower the estimation accuracy; conversely, the smaller the bandwidth, the less smooth the kernel density function curve and the higher the estimation accuracy;

[0055] Step 4: Using the geographic detector, analyze the independent effects and interactive effects of the potential driving factors of the region on the spatial distribution of the hot zone according to the third formula to determine the main driving factors;

[0056] The third formula is:

[0057]

[0058] Where h = 1, 2, ..., L is the stratification of potential driving factor scores, N h is the number of units in layer h, N is the number of units in the region; σ h 2 is the variance of the driver factor score of layer h; σ 2 is the variance of the driver factor score for the entire region;

[0059] Step 5: Use the gradient boosting decision tree method to establish a driving model for soil heavy metal accumulation based on the main driving factors, and obtain the driving force contribution of the main driving factors to soil heavy metal accumulation;

[0060] The driving model for soil heavy metal accumulation is as follows:

[0061] Y i =a1f1(x0,xi ,...x m )+a2f2(x0,x i ,...x m )+…+a n f n (x0,x i ,...x m )+a0;

[0062] Where a1, a2, ...an represent the weight coefficients of n decision trees, a0 represents the initial value, x0, x1, x2 ... xm represent m main driving factors, and fn (x0, x1, x2, ... xm) represents the function of m main driving factors, that is, the predicted value of n trees;

[0063] The driving force contribution of the main driving factor to the accumulation of heavy metals in soil is calculated according to the fifth formula, which is:

[0064]

[0065] Where i represents the decision tree, the number is n, j represents the non-leaf node in the decision tree, the number is l, s ij Represents the reduction in squared loss of j non-leaf nodes in i trees.

[0066] Preferably, the method further comprises: using the coefficient of determination and the average predicted root mean square error to verify the model accuracy of the soil heavy metal accumulation driven model.

[0067] Preferably, the determination coefficient R is calculated according to the sixth formula, which is:

[0068]

[0069] Among them, y p is the predicted value of multiple sampling points, y m is the measured value of multiple sampling points, y a It is the average value of the measured values ​​at multiple sampling points.

[0070] Preferably, the average predicted root mean square error is calculated according to the seventh formula, which is:

[0071]

[0072] Among them, n is the validation data sample set, y p is the predicted value, y m is the measured value.

[0073] Beneficial Effects

[0074] Compared with the prior art, the system and method for constructing a soil heavy metal accumulation driving model according to the present invention have the following beneficial effects:

[0075] (1) The present invention preliminarily determines the categories and scores of potential driving factors for soil heavy metal accumulation through the PMF model; obtains the hot zone distribution of potential driving factor scores and the spatial data of potential driving factors through ArcGIS kernel density analysis and spatial extraction technology, and obtains the names and spatial locations of potential driving factors, thereby reducing the interference of human factors in the selection of potential driving factors in the traditional analysis process and making the identification results more accurate.

[0076] (2) The present invention analyzes the influence of potential driving factors on the spatial distribution of thermal zones by utilizing the geographic detector factor detector method, identifies the main driving factors, and makes up for the shortcoming of traditional analysis methods that cannot handle categorical variables.

[0077] (3) The present invention constructs a soil heavy metal accumulation driving model based on the main driving factors, providing a scientific basis for the prevention, control and management of soil heavy metal pollution. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the specific implementation methods of the present application, they are used to explain the technical solution of the present invention and do not constitute a limitation on the technical solution of the present invention.

[0079] Figure 1 It is a flow chart of the method for constructing the soil heavy metal accumulation driving model described in the present invention.

[0080] Figure 2 It is a schematic diagram of the spatial analysis of latent factor scores.

[0081] Figure 3 This is a schematic diagram for identifying the main driving factors for the accumulation of heavy metals in soil. DETAILED DESCRIPTION

[0082] The present invention is described in more detail hereinafter to facilitate understanding of the present invention.

[0083] The present invention provides a system for constructing a driving model for the accumulation of heavy metals in soil, comprising an information acquisition unit, a matrix determination unit, a potential driving factor acquisition unit, a main control driving factor determination unit and a driving model establishment unit; the information acquisition unit is used to acquire soil heavy metal data, soil physical and chemical property data and point information of multiple sampling points in a region; the matrix determination unit is used to analyze the soil heavy metal content of multiple sampling points in a region according to a first formula using PMF software, and determine a potential driving factor category matrix and a potential driving factor score matrix for the accumulation of heavy metals in the region;

[0084] The first formula is:

[0085]

[0086] Where i represents multiple sampling points in the region (a total of n sampling points), j represents soil heavy metals in the region (a total of m heavy metals), k represents the category of potential driving factors (a total of p categories), x ij represents the content of heavy metal j in sampling point i within the region (mg / kg), g ik represents the contribution of the k-type driving factor in the i-th sampling point in the region, that is, the score matrix of the potential driving factor category, f kj represents the content of heavy metal j in category k of potential driving factors, e ij is the residual; u ij represents the uncertainty of heavy metals at sampling point j in the region i; Q represents the minimum value of the model. ik ≥0, f kj Under the constraint of ≥0, Q is solved by iterative minimization algorithm to determine the driving factor score matrix g ik and the driving factor category matrix f kj .

[0087] The potential driving factor acquisition unit is used to use GIS kernel density analysis and spatial extraction technology to obtain the name and spatial location of the potential driving factor in the region according to the second formula;

[0088] The second formula is:

[0089]

[0090] Where f(g) is the kernel density estimate of the potential driving factor score; h is the search radius of the calculation, also known as the bandwidth; n is the number of feature points within the region whose distance is less than h around the potential driving factor score g at a certain spatial location; g i is the mean of the potential driving factor scores, g is the observed value of the potential driving factor scores, and the K function represents the spatial calculation weight function. The larger the search radius or bandwidth, the smoother the kernel density function curve and the lower the estimation accuracy; conversely, the smaller the bandwidth, the less smooth the kernel density function curve and the higher the estimation accuracy.

[0091] The main driving factor determination unit is used to use the geographic detector to analyze the independent effects and interactive effects of the regional potential driving factors on the spatial distribution of the hot zone according to the third formula to determine the main driving factor;

[0092] The third formula is:

[0093]

[0094] Where h = 1, 2, ..., L is the stratification of potential driving factor scores, N h is the number of units in layer h, N is the number of units in the region; σ h 2 is the variance of the driver factor score of layer h; σ 2 is the variance of the driving factor score for the entire region.

[0095] The driving model establishment unit is used to establish a driving model for soil heavy metal accumulation based on the main driving factor using a gradient boosting decision tree method, and obtain the driving force contribution of the main driving factor to soil heavy metal accumulation;

[0096] The driving model for soil heavy metal accumulation is as follows:

[0097] Y i =a1f1(x0,x i ,...x m )+a2f2(x0,x i ,...x m )+…+a n f n (x0,x i ,...x m )+a0;

[0098] where a1, a2, ...a n Represents the weight coefficient of n decision trees, a0 represents the initial value, x0, x1, x2…x m represents m master driving factors, fn(x0,x1,x2,…x m ) represents the function of m main driving factors, that is, the predicted value of n trees.

[0099] The driving force contribution of the main driving factor to the accumulation of heavy metals in soil is calculated according to the fifth formula, which is:

[0100]

[0101] Where i represents the decision tree (the number is n), j represents the non-leaf node in the decision tree (the number is l), and s ij Represents the reduction in squared loss of j non-leaf nodes in i trees.

[0102] Preferably, the system for constructing the soil heavy metal accumulation driven model also includes a model accuracy verification unit, which is used to verify the model accuracy of the soil heavy metal accumulation driven model using the determination coefficient and the average predicted root mean square error.

[0103] Preferably, the determination coefficient R is calculated according to the sixth formula, which is:

[0104]

[0105] Among them, y p is the predicted value of multiple sampling points, y m is the measured value of multiple sampling points, y a It is the average value of the measured values ​​at multiple sampling points.

[0106] Preferably, the average predicted root mean square error is calculated according to the seventh formula, which is:

[0107]

[0108] Among them, n is the validation data sample set, y p is the predicted value, y m is the measured value.

[0109] Preferably, the acquisition of soil heavy metal data, soil physicochemical property data and point information of multiple sampling points in the area includes: field sampling, obtaining soil samples in the study area, and accurately locating them with GPS; laboratory measurement of heavy metal content and soil physicochemical properties of soil samples at multiple points to obtain soil heavy metal content, soil physicochemical property content and point information data.

[0110] One of the objects of the present invention is to provide a simple, accurate and feasible method, which is a simulation method for identifying and quantifying driving factors of soil heavy metal accumulation.

[0111] The basic flow chart of the method of the present invention is shown in Figure 1 The following is a detailed description of each step of the method of the present invention in conjunction with specific examples.

[0112] Step 1: Collect soil environmental samples from multiple sampling points (140) in the study area, determine the content of soil heavy metals and soil physical and chemical properties, and record the point information.

[0113] Step 1-1: Field sampling: collect soil samples in the study area and accurately locate them using GPS.

[0114] Step 1-2: Determine the heavy metal content and soil physical and chemical properties in soil samples from multiple sampling points in the laboratory to obtain heavy metal content, soil physical and chemical properties and point information data. The specific content data will not be repeated here. Table 1 is an overview of the statistical data of soil heavy metal content, and Table 2 is an overview of the statistical data of soil physical and chemical properties.

[0115] Table 1 Statistical values ​​of heavy metal content in soil (mg / kg)

[0116] Statistics As Ba Cu Cr Cd Cs Mn Mo Ni Sr Sn Sb Pb Zn Minimum 2.11 100.00 16.82 26.30 0.07 285.3 353.0 0.28 20.21 191.4 1.60 0.30 17.78 48.57 Maximum 19.62 1350 567.8 379.0 66.47 961.8 975.8 1.57 68.34 616.1 23.90 4.29 203.4 390.7 Median 10.99 755.5 56.22 71.84 0.21 750.1 720.5 0.88 45.15 322.2 3.86 1.13 36.41 135.8 average value 10.91 746.6 61.17 80.24 0.69 738.6 715.6 0.86 44.75 328.2 5.05 1.14 42.70 144.8 Standard Deviation 3.10 181.9 49.04 36.40 5.60 116.1 104.1 0.26 9.11 56.74 3.73 0.45 23.59 51.63

[0117] Table 2 Statistical data of soil physical and chemical properties

[0118] Statistics pH TC(%) TN(%) AP(mg / kg) <![CDATA[Fe2O3(%)]]> CaO(%) <![CDATA[K2O(%)]]> Minimum 7.69 0.91 0.03 11.34 2.56 1.65 0.78 Maximum 8.98 5.70 0.65 26.50 9.30 18.89 4.37 Median 7.99 2.34 0.17 21.50 6.29 7.84 3.58 average value 8.02 2.43 0.17 21.25 6.19 7.92 3.43 Standard Deviation 0.13 0.79 0.09 3.090 1.02 2.66 0.52

[0119] Step 2: Conduct PMF analysis on soil heavy metal content at multiple sampling points in the region.

[0120] After obtaining the soil heavy metal contents at multiple sampling points in the region, PMF analysis was performed on the soil heavy metal contents at multiple sampling points in the region to determine the category matrix and score matrix of potential driving factors affecting the accumulation of soil heavy metals.

[0121] The potential driving factor category matrix is ​​the coefficient of the fitting equation with each driving factor as the independent variable and the heavy metal content of multiple sampling points as the dependent variable, which reflects the contribution of the potential driving factor category to the heavy metal content of multiple sampling points, and represents the potential driving factor category. The factor category matrix is ​​shown in Table 3. As can be seen from Table 3, PMF identified 4 potential driving factors, among which the first potential driving factor contributed 74.96, 40.10, 40.50, 32.20, 43.91, 37.44, and 54.49 to Ba, Cr, Cs, Mn, Mo, Ni, and Sr; the second potential driving factor contributed 53.28, 64.09, 53.72, 47.08, and 56.76 to Cu, Cd, Sb, Pb, and Zn; the third potential driving factor contributed 52.67 and 39.53 to Sn and Pb; and the fourth potential driving factor contributed 48.0 to As.

[0122] The potential driving factor score data is mainly the coefficient of the fitting equation with the heavy metal content of each sampling point as the independent variable and the potential driving factor as the dependent variable, which reflects the contribution of the driving factor category to the heavy metal content of multiple sampling points and represents the contribution rate of the potential driving factor. The factor score data of multiple sampling points will not be repeated here, and their statistical values ​​are listed (Table 4).

[0123] Table 3 Factor category matrix

[0124]

[0125]

[0126] Table 4 Factor score matrix

[0127] Factor 1 Factor 2 Factor 3 Factor 4 Minimum -0.20 -0.20 -0.20 -0.20 Maximum 2.03 2.48 7.17 2.65 average value 1.00 1.00 0.99 1.00 Median 1.04 1.01 0.77 1.04 Standard Deviation 0.39 0.49 0.98 0.50

[0128] Step 3: Based on the ArcGIS platform, the factor scores are analyzed through kernel density analysis, and the potential driving factor categories of the region are spatialized using the proximity extraction technology to obtain spatial data;

[0129] Step 3-1: According to step 2, the potential driving factor scores corresponding to multiple sampling points in the region are obtained, and the factor scores are analyzed using the ArcGIS platform kernel density analysis to obtain the hot zone distribution of the potential driving factor scores ( Figure 2 ).

[0130] Step 3-2: According to the potential driving factor category, the driving factor category is retrieved in the hot zone distribution area of ​​the potential driving factor score, and ArcGIS is further used to spatially digitize the potential driving factors (Table 5). For example, the hot zone corresponding to potential driving factor category 1 is relatively scattered, mainly considering the soil background factors and the spatial data of the soil parent material where the sample points are located; the hot zone corresponding to potential driving factor category 2 mainly has industrial activities, and the ArcGIS platform is further used to extract the distances from multiple sampling points to the enterprises; the hot zone corresponding to potential driving factor category 3 is mainly in the main traffic arteries, and the ArcGIS platform is further used to extract the distances from multiple sampling points to the roads; the hot zone corresponding to potential driving factor category 4 mainly appears in the area close to the residential area, and the distances from multiple sampling points to the residential area are further extracted through ArcGIS.

[0131] Table 5 Spatial data of potential driving factors

[0132]

[0133]

[0134] Step 4: Use geographic detectors to determine the main driving factors for the accumulation of heavy metals in regional soils.

[0135] In order to obtain the main driving factors of regional soil heavy metal accumulation, the potential driving factor score was used as the dependent variable, and the spatial data of the potential driving factor was used as the independent variable. The geographic detector factor detection template was used to analyze the influence of the potential driving factor on the accumulation of heavy metals in soil.

[0136] Step 4-1: Input the factor scores of potential driving factors into the dependent variable (Y) in the Geographic Detector;

[0137] Step 4-2: Enter the independent variables (X), distance from residential areas (X1), distance from roads (X2), distance from enterprises (X3), soil parent material data (X4), soil pH value (X5), soil Fe2O3 (X6), soil CEC (X7) in the geographic detector;

[0138] Step 4-3: Obtain the impact of potential driving factors and soil physical and chemical properties on the driving factor scores by running the geographic detector ( Figure 3 ).from Figure 3It can be seen that the distance from residential areas, the distance from roads, the distance from enterprises, soil parent material and soil Fe2O3 are the main driving factors for the accumulation of soil heavy metals.

[0139] Step 5: Establish a soil heavy metal accumulation driving model based on the main driving factors.

[0140] Based on the data of the main driving factors, 70% of the data were used as the training set and 30% of the data were used as the validation set to establish a gradient boosting decision tree model. The main driving factor data were used as independent variables, and the soil heavy metal content at multiple sampling points in the region was used as the dependent variable. The gradient decision tree software package of python software was used to construct a driving model for soil heavy metal accumulation, and the driving force contribution of the main driving factors to the accumulation of soil heavy metals was quantified (Table 6). The average prediction root mean square error and determination coefficient were used to characterize the model accuracy (Table 6).

[0141] Table 6 Accuracy of soil heavy metal accumulation model

[0142]

[0143]

[0144] The above describes the preferred embodiments of the present invention, but it is not intended to limit the present invention. Those skilled in the art may make improvements and changes to the embodiments disclosed herein without departing from the scope and spirit of the present invention.

Claims

1. A system for constructing a driving model for soil heavy metal accumulation, characterized in that: The construction system of the driving model for the accumulation of heavy metals in soil comprises an information acquisition unit, a matrix determination unit, a potential driving factor acquisition unit, a main control driving factor determination unit and a driving model establishment unit; the information acquisition unit is used to obtain soil heavy metal data, soil physical and chemical property data and point information of multiple sampling points in a region; the matrix determination unit is used to analyze the soil heavy metal content of multiple sampling points in a region according to a first formula using PMF software, and determine the potential driving factor category matrix and potential driving factor score matrix for the accumulation of heavy metals in the region; The first formula is: Where i represents multiple sampling points in the region, with a total of n sampling points, j represents soil heavy metals in the region, with a total of m soil heavy metals, k represents the category of potential driving factors, with a total of p categories, and x ij Indicates the content of heavy metals in the soil of sampling point j in the region, in mg / kg, g ik represents the contribution of the k-type driving factor in the i-th sampling point in the region, that is, the score matrix of the potential driving factor category, f kj represents the content of heavy metals in soil j in category k of potential driving factors, e ij is the residual; u ij represents the uncertainty of soil heavy metals at sampling point j in the region; Q represents the minimum value of the model; ik ≥0, f kj Under the constraint of ≥0, Q is solved by iterative minimization algorithm to determine the driving factor score matrix g ik and the driving factor category matrix f kj ; The potential driving factor acquisition unit is used to use GIS kernel density analysis and spatial extraction technology to obtain the name and spatial location of the potential driving factor in the region according to the second formula; The second formula is: Where f(g) is the kernel density estimate of the potential driving factor score; h is the search radius of the calculation, also known as the bandwidth; n is the number of feature points within the region whose distance is less than h around the potential driving factor score g at a certain spatial location; g i is the mean of the potential driving factor scores, g is the observed value of the potential driving factor scores, and the K function represents the spatial calculation weight function; the larger the search radius or bandwidth, the smoother the kernel density function curve and the lower the estimation accuracy; conversely, the smaller the bandwidth, the less smooth the kernel density function curve and the higher the estimation accuracy; The main driving factor determination unit is used to use the geographic detector to analyze the independent effects and interactive effects of the regional potential driving factors on the spatial distribution of the factor score hotspot according to the third formula to determine the main driving factor; The third formula is: Where h = 1, 2, ..., L is the stratification of potential driving factor scores, N h is the number of units in layer h, N is the number of units in the region; σ h 2 is the variance of the driver factor score of layer h; σ 2 is the variance of the driver factor score for the entire region; The driving model building unit is used to establish a driving model for soil heavy metal accumulation based on the main driving factor using a gradient boosting decision tree method, and obtain the driving force contribution of the main driving factor to soil heavy metal accumulation; The driving model for soil heavy metal accumulation is as follows: Y i =a1f1(x0,x i ,...x m )+a2f2(x0,x i ,...x m )+…+a n f n (x0,x i ,...x m )+a0; where a1, a2, ...a n Represents the weight coefficient of n decision trees, a0 represents the initial value, x0, x1, x2…x m represents m master driving factors, fn(x0,x1,x2,…x m ) represents the function of m main driving factors, that is, the predicted value of n trees; The driving force contribution of the main driving factor to the accumulation of heavy metals in soil is calculated according to the fifth formula, which is: Where i represents the decision tree, the number is n, j represents the non-leaf node in the decision tree, the number is l, s ij Represents the reduction in squared loss of j non-leaf nodes in i trees.

2. The system for constructing a soil heavy metal accumulation driving model according to claim 1, characterized in that: The system for constructing the soil heavy metal accumulation driven model also includes a model accuracy verification unit, which is used to verify the model accuracy of the soil heavy metal accumulation driven model using the determination coefficient and the average prediction root mean square error.

3. The system for constructing a soil heavy metal accumulation driving model according to claim 2, characterized in that: The determination coefficient R is calculated according to the sixth formula, which is: Among them, y p is the predicted value of multiple sampling points, y m is the measured value of multiple sampling points, y a It is the average value of the measured values ​​at multiple sampling points.

4. The system for constructing a soil heavy metal accumulation driving model according to claim 2, characterized in that: The average prediction root mean square error is calculated according to the seventh formula, which is: Among them, n is the validation data sample set, y p is the predicted value, y m is the measured value.

5. The system for constructing a soil heavy metal accumulation driving model according to claim 1, characterized in that: The acquisition of soil heavy metal data, soil physical and chemical properties data and point information of multiple sampling points in the area includes: field sampling, obtaining soil samples in the study area, and accurately locating them with GPS; laboratory measurement of heavy metal content and soil physical and chemical properties of soil samples at multiple points to obtain soil heavy metal content, soil physical and chemical properties content and point information data.

6. The system for constructing a soil heavy metal accumulation driving model according to claim 1, characterized in that: The steps of obtaining the name and spatial location of the potential driving factor in the region include: (1) obtaining the hot zone spatial distribution of the potential driving factor score through ArcGIS kernel density analysis; (2) spatially vectorizing the potential driving factor category spatial data through ArcGIS spatial analysis; (3) extracting spatial data through ArcGIS, if the potential driving factor is a continuous variable type, using GIS nearest neighbor analysis technology to extract the distance from the potential factor to the sampling point, if the potential driving factor is a categorical variable, using GIS spatial analysis to extract the attribute data of the sampling point location; (4) determining the name and spatial location of the potential driving factor in the region based on the hot zone spatial distribution of the potential driving factor score, the potential driving factor category spatial data and the soil physical and chemical property data.

7. The system for constructing a soil heavy metal accumulation driving model according to claim 1, characterized in that: The steps to determine the main driving factors are as follows: using the spatial data of potential driving factors extracted by GIS as the independent variable, the potential driving factor scores obtained by PMF analysis as the dependent variable, and using the geographic detector factor detection module to analyze the independent and interactive effects of potential driving factors on the potential driving factor scores, to determine the main driving factors for the accumulation of soil heavy metals in the region.

8. The system for constructing a soil heavy metal accumulation driving model according to claim 1, characterized in that: The steps to establish a soil heavy metal accumulation driving model based on the main driving factors are as follows: using the spatial data of the main driving factors as the independent variable and the accumulated content of soil heavy metals at multiple sampling points as the dependent variable, a gradient boosting decision tree is used to establish a soil heavy metal accumulation driving model based on the main driving factors, and the contribution of the main driving factors to the driving force of soil heavy metal accumulation is calculated, and the determination coefficient and the average prediction root mean square error are used to verify the model accuracy of the soil heavy metal accumulation driving model.

9. The system for constructing a soil heavy metal accumulation driving model according to claim 1, characterized in that: The steps to obtain soil heavy metal data, soil physical and chemical properties data and point information of multiple sampling points in the area are as follows: Field sampling, obtaining soil samples from multiple sampling points in the area and accurately locating them using GPS; Determine the heavy metal content of soil samples at multiple sampling points, including the content of As, Ba, Cu, Cd, Cr, Cu, Cs, Mn, Mo, Ni, Ni, Sr, Sn, Sb, Pb and Zn, soil pH, soil CEC, soil Fe2O3, TN, TC and TP, and record the point information of multiple sampling points.

10. The method for constructing a soil heavy metal accumulation driving model according to any one of claims 1 to 9, characterized in that: The method comprises the following steps: The first step is to obtain soil heavy metal data, soil physical and chemical properties data and point information at multiple sampling points in the area; The second step is to use PMF software to analyze the soil heavy metal content of multiple sampling points in the region according to the first formula, and determine the potential driving factor category matrix and potential driving factor score matrix for heavy metal accumulation in the region; The first formula is: Where i represents multiple sampling points in the region, with a total of n sampling points, j represents soil heavy metals in the region, with a total of m soil heavy metals, k represents the category of potential driving factors, with a total of p categories, and x ij Indicates the content of heavy metals in the soil of sampling point j in the region, in mg / kg, g ik represents the contribution of the k-type driving factor in the i-th sampling point in the region, that is, the score matrix of the potential driving factor category, f kj represents the content of heavy metals in soil j in category k of potential driving factors, e ij is the residual; u ij represents the uncertainty of soil heavy metals at sampling point j in the region; Q represents the minimum value of the model; ik ≥0, f kj Under the constraint of ≥0, Q is solved by iterative minimization algorithm to determine the driving factor score matrix g ik and the driving factor category matrix f kj ; Step 3: Using GIS kernel density analysis and spatial extraction technology, the names and spatial locations of potential driving factors in the region are obtained according to the second formula; The second formula is: Where f(g) is the kernel density estimate of the potential driving factor score; h is the search radius of the calculation, also known as the bandwidth; n is the number of feature points within the region whose distance is less than h around the potential driving factor score g at a certain spatial location; g i is the mean of the potential driving factor scores, g is the observed value of the potential driving factor scores, and the K function represents the spatial calculation weight function; the larger the search radius or bandwidth, the smoother the kernel density function curve and the lower the estimation accuracy; conversely, the smaller the bandwidth, the less smooth the kernel density function curve and the higher the estimation accuracy; Step 4: Using the geographic detector, analyze the independent effects and interactive effects of the potential driving factors of the region on the spatial distribution of the hot zone according to the third formula to determine the main driving factors; The third formula is: Where h = 1, 2, ..., L is the stratification of potential driving factor scores, N h is the number of units in layer h, N is the number of units in the region; σ h 2 is the variance of the driver factor score of layer h; σ 2 is the variance of the driver factor score for the entire region; Step 5: Use the gradient boosting decision tree method to establish a driving model for soil heavy metal accumulation based on the main driving factors, and obtain the driving force contribution of the main driving factors to soil heavy metal accumulation; The driving model for soil heavy metal accumulation is as follows: Y i =a1f1(x0,x i ,...x m )+a2f2(x0,x i ,...x m )+…+a n f n (x0,x i ,...x m )+a0; where a1, a2, ...a n Represents the weight coefficient of n decision trees, a0 represents the initial value, x0, x1, x2…x m represents m master driving factors, fn(x0,x1,x2,…x m ) represents the function of m main driving factors, that is, the predicted value of n trees; The driving force contribution of the main driving factor to the accumulation of heavy metals in soil is calculated according to the fifth formula, which is: Where i represents the decision tree, the number is n, j represents the non-leaf node in the decision tree, the number is l, s ij Represents the reduction in squared loss of j non-leaf nodes in i trees.

Citation Information

Patent Citations

  • Method for identifying key factors influencing accumulation and migration of heavy metals in soil of heavily-polluted site

    CN115359185A