Method for predicting spatial distribution of heavy metal arsenic in soil based on Bayesian maximum entropy
By combining Bayesian maximum entropy, geographic weighted regression and ordinary Kriging model, the problems of nonlinear relationship and multi-source data uncertainty in the spatial distribution prediction of soil heavy metal arsenic are solved, and high-precision spatial distribution prediction of soil heavy metal arsenic is achieved.
Patent Information
- Application Number
- CN202510212565.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has failed to effectively solve the complex nonlinear relationship between independent variables and auxiliary variables and the uncertainty of multi-source data in the prediction of spatial distribution of heavy metals in soil, making the prediction accuracy difficult to meet actual needs.
A mixed method combining Bayesian maximum entropy (BME), geographic weighted regression (GWR) and ordinary kriging (OK) models is adopted to integrate multi-source environmental variables through multiple statistical models and BME methods, and complex nonlinear relationships between soil arsenic and environmental factors are processed, and prediction accuracy is improved through spatial interpolation and data integration.
Effectively integrating multi-source environment variables improves the ability to process uncertainty of multi-source data, enhances the model's adaptability in complex environments, and significantly improves the accuracy of predicting spatial distribution of soil heavy metal arsenic.
Smart Images

Figure CN120145328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soil environmental science, and particularly relates to a method for predicting the spatial distribution of heavy metal arsenic in soil based on multi-source environmental variables and Bayesian maximum entropy (BME), geographically weighted regression (GWR), and ordinary kriging (OK) models. Background Art
[0002] Soil heavy metal pollution has become a global environmental problem, posing a serious threat to the ecosystem and human health. Due to their high toxicity, non-degradability, and bioaccumulation, especially in agricultural paddy soil, the main pollution sources of heavy metals include industrial emissions, agricultural inputs (such as chemical fertilizers and pesticides), and atmospheric deposition, etc. Among them, arsenic is a toxic and carcinogenic element among heavy metals. Excessive arsenic will enter the body along with agricultural products, leading to a series of health problems. Moreover, excessive arsenic will also cause abnormal growth of plants and damage the ecological balance. Therefore, accurately predicting its spatial distribution is crucial.
[0003] Currently, the prediction of the spatial distribution of heavy metal arsenic in soil is mainly based on establishing prediction models with multi-source auxiliary variables, aiming to deeply explore the non-linear relationship between auxiliary variables and soil heavy metal pollutants, and then improve the prediction accuracy and accurately draw the spatial distribution map of soil heavy metal pollution. In the prior art, Xie Xuefeng et al. used geostatistical methods, and Zhang C T et al. adopted the geographically weighted regression (GWR) model, combined with multi-source environmental auxiliary variables such as topographic factors, soil properties (natural factors), and human activities. By analyzing the spatial characteristic factors closely related to elements, the prediction accuracy of the spatial distribution of soil heavy metal pollution was improved.
[0004] Although the above technologies have made certain progress in improving the prediction accuracy of spatial distribution, however, these methods have not effectively solved the complex non-linear relationship between independent variables and auxiliary variables and the uncertainty problem of multi-source data, and there is still room for improvement in the prediction accuracy.
[0005] Therefore, the present invention proposes a hybrid method combining BME, GWR, and OK models to achieve high-precision prediction of the spatial distribution of heavy metal arsenic in soil and accurately draw its spatial distribution map. Summary of the Invention
[0006] Aiming at the non-linear relationship between heavy metal arsenic and auxiliary variables and the uncertainty of multi-source data during the prediction of the spatial distribution of heavy metal arsenic in soil, which makes it difficult for the accuracy of traditional prediction methods to meet the actual needs, the present invention proposes a method for predicting the spatial distribution of heavy metal arsenic in soil based on Bayesian maximum entropy, which improves the prediction accuracy of the content of heavy metal arsenic in soil.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] A method for predicting the spatial distribution of heavy metal arsenic in soil based on Bayesian maximum entropy, comprising the following steps:
[0009] Step S1: Collect the data of heavy metal arsenic content in soil and relevant environmental auxiliary variable data in the research area;
[0010] Step S2: Based on the arsenic content data and relevant auxiliary variable data obtained in Step S1, through Pearson correlation coefficient and variance inflation factor analysis, screen out the environmental auxiliary variables that are significantly correlated with the heavy metal arsenic content in soil, so as to obtain the dataset for modeling;
[0011] Step S3: On the basis of the data screened in Step S2, use the geographically weighted regression model to construct a local regression model of the heavy metal content in soil, obtain the regression residuals through this model, and quantitatively analyze the spatial heterogeneity in combination with environmental auxiliary variables, and calculate the local regression coefficients and predicted values of each sample point;
[0012] Step S4: Combine the ordinary Kriging model to perform spatial interpolation on the GWR regression residuals obtained in Step S3, and then add the obtained results to the predicted values in Step S3 to obtain the preliminary spatial prediction result of the heavy metal arsenic content in soil;
[0013] Step S5: Convert the preliminary prediction result obtained in Step S4 into an interval distribution form;
[0014] Step S6: Take the arsenic content data in Step S1 as hard data, and the prediction result in the interval distribution form in Step S5 as soft data, and use the Bayesian maximum entropy model to integrate the hard and soft data to obtain the final prediction result of the heavy metal arsenic content in soil;
[0015] Step S7: Make a spatial distribution map of heavy metal arsenic in soil according to the prediction data of the heavy metal arsenic content obtained in Step S6.
[0016] In Step S1, the environmental auxiliary variables include terrain, natural conditions, remote sensing information, vegetation index and human activities.
[0017] The terrain data includes: elevation, terrain undulation, slope, slope change rate, curvature, aspect, aspect change rate data corresponding to each sampling point;
[0018] The natural conditions include: atmospheric dustfall, ditch bottom mud, ditch bottom mud pH, annual average wind speed, annual relative humidity, annual average temperature, annual precipitation, soil pH data corresponding to each sampling point;
[0019] The remote sensing information includes: Band1, Band2, Band3, Band4, Band5, Band6, Band7 data corresponding to each sampling point;
[0020] The vegetation indices include: the difference vegetation index, ratio vegetation index, red-green ratio index, normalized difference vegetation index, greenness vegetation index, atmospheric impedance vegetation index, clay mineral ratio, and water stress index corresponding to each sampling point;
[0021] The human activity data include: the GDP, the heavy metal contents in commercial organic fertilizers, the heavy metal contents in biological organic fertilizers, the heavy metal contents in organic-inorganic compound fertilizers, the heavy metal contents in superphosphate, and the heavy metal content data in compound fertilizers corresponding to each sampling point.
[0022] In step S1, the soil heavy metal data are obtained by collecting surface soil samples of 0-20 cm and measuring the heavy metal contents using acid digestion method, and the environmental auxiliary variable data are obtained through remote sensing images, geospatial data clouds, and on-site investigations.
[0023] In step S2, it specifically includes the following steps:
[0024] Step S2-1: Import the heavy metal content data and the auxiliary variable data into two arrays respectively;
[0025] Step S2-2: Select the heavy metal content data and the environmental auxiliary variable data and substitute them into the Pearson correlation coefficient calculation formula to obtain the correlation analysis results;
[0026] Step S2-3: Delete the environmental auxiliary variable data with low correlation through the correlation analysis results in step S2-2;
[0027] Step S2-4: Further screen the variables using the variance inflation factor (VIF) method, eliminate the variables with high VIF values (VIF>10), obtain the auxiliary variables significantly correlated with the soil heavy metal arsenic content, and construct a dataset for modeling.
[0028] In step S3, the GWR model uses an adaptive bandwidth kernel function, takes the corrected Akaike information criterion (AICc) as the evaluation index for selecting the optimal bandwidth, generates the local regression coefficients of each sample point, and the GWR formula is as follows:
[0029]
[0030] In the formula: y GWR(i) is the predicted value of the model at position i; β 0(i) is the intercept; β j(i) is the regression coefficient of the j covariate at position i; X j(i) is the value of the j covariate at position i; n is the number of covariates; ε (i) is the regression residual.
[0031] Step S4 specifically includes the following steps:
[0032] Step S4-1: Import the regression residuals obtained in Step S3 into a file;
[0033] Step S4-2: Interpolate the regression residuals using the ordinary Kriging (OK) method;
[0034] Step S4-3: Superimpose the interpolation result on the GWR model result in S3. The formula is as follows:
[0035] y GWR-OK = y gwr + r OK #(2)
[0036] In the formula: y GWR-OK is the predicted value of the heavy metal element by the model; y gwr is the predicted value of the GWR model; r OK is the value after interpolating the GWR model regression residuals using OK;
[0037] Step S4-4: Obtain the output result y GWR-OK which is the preliminary predicted result of the arsenic content in soil heavy metals.
[0038] In Step S6, the soft data integrated by the BME model is the predicted value calculated by the GWR-OK model, and the hard data is the measured arsenic content data of soil heavy metals. The model optimizes the prediction accuracy through covariance fitting.
[0039] When using the BME model, the following steps are adopted:
[0040] S6-1: Based on the principle of maximum entropy, using the general knowledge G and specific knowledge S, with statistical moments (such as mean, variance, and covariance) as constraint conditions, obtain the prior probability density function (prior pdf) f G , and the formula is as follows:
[0041] H(f G (x map )) = -∫f G (x map )logf G (x map )dx map #(3)
[0042] In the formula: H(f G (x map )) represents the entropy expression about f G (x map ); x map = [x data , x k , where xdata = [x hard , x soft represents the set consisting of the hard data value x hard and the soft data value x soft ; x k is the predicted value at a certain point p k to be predicted;
[0043]
[0044] In the formula: represents the expected value of the constraint condition corresponding to the point p map ; g α (x map ) is a function based on the prior knowledge G; g = {g α , α = 1,..., N c} is the vector composed of each constraint condition; N c is the total number of constraint conditions, related to the total number of x map ;
[0045] S6-2: Under certain constraint conditions, make the formula (4) obtain the maximum value. At this time, the prior probability density function is:
[0046]
[0047] In the formula: A is the partition function, and its role is to normalize f G (x map ); μ = {μ α , α = 1,..., N c} is the coefficient related to g, that is, the Lagrange multiplier.
[0048] S6-3: Organize the hard data and soft data (i.e., specific knowledge S) into a suitable form;
[0049] S6-4: Combine S and f G , and under the generalized Bayesian condition, obtain the posterior probability density function (posterior pdf) f k .
[0050]
[0051] S6-5: Obtain the interval-type soft data by calculating the prediction result of the GWR-OK model, and substitute it into formula (7) to obtain the posterior pdf of the variable distribution of the prediction points in the study area:
[0052]
[0053] In the formula: l and u are the lower bound and upper bound of the interval respectively.
[0054] S6-6: According to f k the predicted value of the prediction point can be calculated The formula is as follows:
[0055]
[0056] In the formula: is the predicted value of the BMEGK model.
[0057] S6-7: Obtain the output result is the final predicted result.
[0058] The formula adopted in step S2-2 is as follows:
[0059]
[0060] In the formula, N is the number of samples; X represents the auxiliary variable data; Y represents the soil heavy metal content data; Z x is the covariance of the auxiliary variable data; Z y is the covariance of the soil heavy metal content data; S x is the standard deviation of the auxiliary variable data; S y is the standard deviation of the soil heavy metal content data; is the average value of the auxiliary variable data; is the average value of the soil heavy metal content data; i is the i-th sample point; X i is the auxiliary variable data at the i-th sample point; Y i is the soil heavy metal content data at the i-th sample point.
[0061] Compared with the prior art, the present invention has the following technical effects:
[0062] 1) The present invention can effectively integrate environmental auxiliary variables from different sources (such as terrain, climate, remote sensing information, etc.), express the uncertainty of data in spatial prediction through the Bayesian maximum entropy model, thereby improving the processing ability of the uncertainty of multi-source data. The fusion of such multi-source data not only optimizes the prediction results but also enhances the adaptability of the model in complex environments, making the spatial distribution prediction of soil heavy metal arsenic more accurate under different geographical environments and pollution backgrounds.
[0063] 2) By adopting a variety of statistical models and BME methods, the present invention effectively processes the complex non-linear relationship between soil arsenic and environmental factors and can more accurately predict the spatial distribution of soil arsenic. This method avoids the problems of over-smoothing and local underestimation in traditional methods by refining the predictions in different regions.
[0064] 3) The present invention fully exploits the potential of multi-source data by combining the Bayesian Maximum Entropy (BME) model, the Geographically Weighted Regression (GWR) model, and the Ordinary Kriging (OK) model. The Bayesian Maximum Entropy model can accurately characterize the uncertainty in the distribution of soil heavy metal arsenic by integrating hard data and soft data, making up for the deficiencies of traditional methods. The Geographically Weighted Regression model can better capture the characteristics of soil arsenic pollution at different geographical locations by dynamically considering spatial heterogeneity. The Ordinary Kriging model can perform smooth interpolation on limited sampling data, improving the prediction accuracy of local areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The present invention will be further described below in conjunction with the drawings and embodiments:
[0066] Figure 1 is a flowchart of the method of the present invention;
[0067] Figure 2 is a spatial distribution map of soil heavy metal arsenic drawn based on the method of the present invention;
[0068] Figure 3 is a distribution map of sampling points in the study area. DETAILED DESCRIPTION OF THE INVENTION
[0069] As Figure 1 shown, a method for predicting the spatial distribution of soil heavy metal arsenic based on Bayesian maximum entropy includes the following steps:
[0070] Step S1: Collect data on the content of soil heavy metal arsenic and related environmental auxiliary variable data in the study area;
[0071] Step S2: Based on the arsenic content data and related auxiliary variable data obtained in Step S1, through Pearson correlation coefficient and variance inflation factor analysis, screen out the environmental auxiliary variables that are significantly correlated with the content of soil heavy metal arsenic, so as to obtain a data set for modeling;
[0072] Step S3: On the basis of the data screened in Step S2, use the Geographically Weighted Regression model to construct a local regression model of soil heavy metal content, obtain regression residuals through this model, and quantitatively analyze the spatial heterogeneity in combination with environmental auxiliary variables, and calculate the local regression coefficients and predicted values of each sample point;
[0073] Step S4: Combine the Ordinary Kriging model to perform spatial interpolation on the GWR regression residuals obtained in Step S3, and then add the obtained results to the predicted values in Step S3 to obtain a preliminary spatial prediction result of the soil heavy metal arsenic content;
[0074] Step S5: Convert the preliminary prediction result obtained in Step S4 into an interval distribution form;
[0075] Step S6: Use the arsenic content data in Step S1 as hard data and the prediction results in the form of interval distribution in Step S5 as soft data, and integrate the hard and soft data using the Bayesian maximum entropy model to obtain the final prediction result of the arsenic content in soil heavy metals;
[0076] Step S7: Make a spatial distribution map of soil heavy metal arsenic based on the predicted data of soil heavy metal arsenic content obtained in Step S6.
[0077] The environmental auxiliary variables in Step S1 include terrain, natural conditions, remote sensing information, vegetation index, and human activities;
[0078] Among them, terrain: elevation, terrain undulation, slope, slope variability, curvature, aspect, and aspect variability data corresponding to each sampling point;
[0079] Natural conditions: atmospheric dustfall, ditch bottom sediment, ditch bottom sediment pH, annual average wind speed, annual relative humidity, annual average temperature, annual precipitation, and soil pH data corresponding to each sampling point;
[0080] Remote sensing information: Band1, Band2, Band3, Band4, Band5, Band6, and Band7 data corresponding to each sampling point;
[0081] Vegetation index: difference vegetation index, ratio vegetation index, red-green ratio index, normalized difference vegetation index, greenness vegetation index, atmospheric impedance vegetation index, clay mineral ratio, and water stress index corresponding to each sampling point;
[0082] Human activities: GDP, heavy metal contents in commercial organic fertilizers, heavy metal contents in biological organic fertilizers, heavy metal contents in organic-inorganic compound fertilizers, heavy metal contents in superphosphate, and heavy metal contents in compound fertilizers corresponding to each sampling point;
[0083] In Step S1, the soil heavy metal data is obtained by collecting 0-20 cm surface soil samples and measuring the heavy metal content using acid digestion method, and the environmental auxiliary variable data is obtained through remote sensing images, geospatial data clouds, and on-site investigations.
[0084] The detailed steps in Step S2 include:
[0085] Step S2-1: Import the heavy metal content data and auxiliary variable data into two arrays respectively;
[0086] Step S2-2: Select the heavy metal content data and environmental auxiliary variable data and substitute them into the Pearson correlation coefficient calculation formula to obtain the correlation analysis result. The formula is as follows:
[0087]
[0088] In the formula, N is the number of samples; X represents the auxiliary variable data; Y represents the soil heavy metal content data; Z x is the covariance of the auxiliary variable data; Z y is the covariance of the soil heavy metal content data; S x is the standard deviation of the auxiliary variable data; S y is the standard deviation of the soil heavy metal content data; is the average value of the auxiliary variable data; is the average value of the soil heavy metal content data; i is the i-th sample point; X i is the auxiliary variable data at the i-th sample point; Y i is the soil heavy metal content data at the i-th sample point;
[0089] Step S2-3: Delete the environmental auxiliary variable data with low correlation according to the correlation analysis results in Step S2-2;
[0090] Step S2-4: Further screen variables using the variance inflation factor (VIF) method, eliminate variables with high VIF values (VIF>10), obtain the auxiliary variables significantly correlated with the arsenic content of soil heavy metals, and construct a dataset for modeling.
[0091] In Step S3, the GWR model uses an adaptive bandwidth kernel function, and the corrected Akaike information criterion (AICc) is used as the evaluation index for selecting the optimal bandwidth to generate the local regression coefficient of each sample point. The GWR formula is as follows:
[0092]
[0093] In the formula: y GWR(i) is the predicted value of the model at position i; β 0(i) is the intercept; β j(i) is the regression coefficient of the j covariate at position i; X j(i) is the value of the j covariate at position i; n is the number of covariates; ε (i) is the regression residual.
[0094] The detailed steps in Step S4 include:
[0095] Step S4-1: Import the regression residuals obtained in Step S3 into a file;
[0096] Step S4-2: Interpolate the regression residuals using the ordinary Kriging (OK) method;
[0097] Step S4-3: Superimpose the interpolation result with the GWR model result in S3. The formula is as follows:
[0098] y GWR-OK= y gwr + r OK #(3)
[0099] Where: y GWR-OK is the predicted value of the heavy metal element by the model; y gwr is the predicted value of the GWR model; r OK is the value obtained by interpolating the regression residuals of the GWR model using OK;
[0100] Step S4-4: Obtain the output result y GWR-OK which is the preliminary predicted result of the arsenic content in soil heavy metals;
[0101] In step S6, the soft data integrated by the BME model is the predicted value calculated by the GWR-OK model, and the hard data is the measured data of the arsenic content in soil heavy metals. The model optimizes the prediction accuracy through covariance fitting.
[0102] The detailed steps of the BME model method include:
[0103] S5-1: Based on the principle of maximum entropy, using the general knowledge G and specific knowledge S, with statistical moments (such as mean, variance, and covariance) as constraints, obtain the prior probability density function (prior pdf) f G , and the formula is as follows:
[0104] H(f G (x map )) = -∫f G (x map ) log f G (x map ) dx map #(4)
[0105] Where: H(f G (x map )) represents the entropy expression about f G (x map ); x map = [x data , x k , where x data = [x hard , x soft represents the set composed of the hard data value x hard and the soft data value x soft ; x k is the predicted value at a certain point p k ;
[0106]
[0107] Where: represents the point p mapThe expected value of the corresponding constraint; g α (x map ) is a function based on the prior knowledge G; g = {g α , α = 1, …, N c} is the vector composed of each constraint; N c is the total number of constraints, related to the total number of x map .
[0108] S5-2: Under certain constraints, make Equation (4) achieve the maximum value. At this time, the prior probability density function is:
[0109]
[0110] In the formula: A is the partition function, and its role is to normalize f G (x map ); μ = {μ α , α = 1, …, N c} is the coefficient related to g, that is, the Lagrange multiplier.
[0111] S5-3: Organize the hard data and soft data (i.e., specific knowledge S) into a suitable form;
[0112] S5-4: Combine S and f G , and under the generalized Bayesian condition, obtain the posterior probability density function (posterior pdf) f k .
[0113]
[0114] S5-5: Obtain the interval-type soft data by calculating the prediction result of the GWR-OK model, and substitute it into Equation (8) to obtain the posterior pdf of the variable distribution of the prediction points in the study area:
[0115]
[0116] In the formula: l and u are the lower and upper bounds of the interval respectively.
[0117] S5-6: According to f k , the predicted value of the prediction point can be calculated The formula is as follows:
[0118]
[0119] In the formula: is the predicted value of the BMEGK model.
[0120] S5-7: Obtain the output result is the final predicted result.
[0121] In step S7, a spatial distribution map of the arsenic content in soil within the space is drawn according to the output result as Figure 2 shown.
[0122] The method proposed by the present invention not only improves the accuracy of heavy metal spatial prediction, but also provides scientific support for regional soil pollution monitoring, risk assessment and precise treatment, and can provide an important reference for optimizing land use planning and formulating soil pollution prevention and control policies.
[0123] Data example:
[0124] The present invention was tested in the paddy fields in the key concerned areas of Yichang City as the research area. Among them, the research area was divided into ordinary areas, suspected areas, concerned areas and key areas. Scientific sampling points were arranged mainly by the grid method, and the plough layer soil samples (0 - 20 cm) were collected by methods such as plum blossom points, checkerboards, and serpents according to the specific conditions of the sampling area, and intensive sampling was carried out in the concerned areas and key areas. Finally, 660 samples were obtained after processing, and the distribution of sampling points in the research area is as Figure 3 .
[0125] The 660 sample points of the present invention are randomly allocated according to 7:3. Among them, 70% are used as the model training set, and the remaining 30% are used as the test set. During spatial prediction, first, the training set and environmental covariates are used for modeling, and then the predicted values of the test set sample points are compared with the measured values, and the mean absolute error (MAE) and root mean square error (RMSE) are calculated to compare the model effects. The lower the MAE and RMSE, the better the model prediction effect, and vice versa. The calculation formulas are as follows:
[0126]
[0127] In the formula: n is the number of samples; y i is the i-th measured value; is the i-th predicted value.
[0128] Table 1 shows the auxiliary variables that are significantly correlated with soil heavy metal arsenic after Pearson correlation analysis and variance inflation factor (VIF) screening. Through the result analysis, it can be seen that the content of soil heavy metal arsenic has a significant correlation with multiple environmental factors. As is mainly negatively correlated with the difference vegetation index, annual average wind speed and water stress index, and is positively correlated with annual relative humidity and the As content in the ditch bottom mud.
[0129] Table 1 Correlation between soil heavy metal arsenic and the screened auxiliary variables
[0130]
[0131] Table 2 shows the MAE and RMSE of the prediction of soil heavy metal arsenic by different models (OK, GWR, and BMEGK). Through the result analysis, it can be obtained that compared with the traditional prediction method OK, the MAE of BMEGK in the prediction of As decreased by 8.98%, and the RMSE decreased by 11.39%. This result indicates that the method (BMEGK) proposed by the present invention can not only effectively process multi-source environmental variables under a large dataset and complex environmental conditions, but also more accurately predict the spatial distribution of soil heavy metal arsenic.
[0132] Table 2 Experimental results of different models
[0133]
[0134] By comparing the prediction results of different models under multi-source datasets, the method proposed by the present invention can more precisely depict the spatial heterogeneity and complex non-linear relationship of soil heavy metal arsenic by comprehensively considering multi-source environmental variables such as terrain, natural factors, and remote sensing data. Compared with the traditional OK model and other geostatistical methods, the BMEGK model shows stronger robustness and accuracy in dealing with this spatial heterogeneity and non-linear relationship. This advantage makes the BMEGK method have higher applicability in the prediction of the spatial distribution of soil heavy metal arsenic pollution, especially in areas with complex and diverse environmental variables.
[0135] Therefore, the experimental results not only verify the effectiveness of the method of the present invention in improving the prediction accuracy, but also prove the wide application potential of this method in the fields of soil heavy metal pollution monitoring, precise treatment, environmental risk assessment, and land use planning. This achievement provides a feasible and efficient technical solution for future soil pollution spatial prediction based on multi-source data.
Claims
1. A method for predicting the spatial distribution of heavy metal arsenic in soil based on Bayesian maximum entropy, characterized in that: The following steps are involved: Step S1: Collect soil heavy metal arsenic content data and related environmental auxiliary variable data in the study area; Step S2: Based on the arsenic content data and related auxiliary variable data obtained in step S1, the environmental auxiliary variables significantly correlated with the soil heavy metal arsenic content are screened out through Pearson correlation coefficient and variance inflation factor analysis, thereby obtaining a data set for modeling; Step S3: Based on the data filtered in step S2, a local regression model of soil heavy metal content is constructed using a geographically weighted regression model. The regression residuals are obtained through the model, and the spatial heterogeneity is quantitatively analyzed in combination with environmental auxiliary variables to calculate the local regression coefficient and predicted value of each sample point. Step S4: spatially interpolate the GWR regression residuals obtained in step S3 in combination with the ordinary Kriging model, and then add the obtained result to the predicted value in step S3 to obtain a preliminary spatial prediction result of soil heavy metal arsenic content; Step S5: converting the preliminary prediction result obtained in step S4 into an interval distribution form; Step S6: The arsenic content data in step S1 is used as hard data, and the prediction result in the interval distribution form in step S5 is used as soft data. The Bayesian maximum entropy model is used to integrate the hard and soft data to obtain the final prediction result of the soil heavy metal arsenic content; Step S7: Produce a spatial distribution map of soil heavy metal arsenic based on the predicted data of soil heavy metal arsenic content obtained in step S6.
2. The method according to claim 1, characterized in that In step S1, environmental auxiliary variables include topography, natural conditions, remote sensing information, vegetation index and human activities.
3. The method according to claim 2, characterized in that The terrain data include: the elevation, terrain relief, slope, slope variation rate, curvature, slope aspect, and slope aspect variation rate data corresponding to each sampling point; Natural conditions include: atmospheric dust, ditch sediment, ditch sediment pH, annual average wind speed, annual relative humidity, annual average temperature, annual precipitation, and soil pH data corresponding to each sampling point; Remote sensing information includes: Band 1, Band 2, Band 3, Band 4, Band 5, Band 6, Band 7 data corresponding to each sampling point; Vegetation indexes include: difference vegetation index, ratio vegetation index, red-green ratio index, normalized difference vegetation index, greenness vegetation index, atmospheric resistance vegetation index, clay mineral ratio, and water stress index corresponding to each sampling point; Human activity data include: GDP corresponding to each sampling point, the content of heavy metals in commercial organic fertilizers, the content of heavy metals in biological organic fertilizers, the content of heavy metals in organic-inorganic compound fertilizers, the content of heavy metals in superphosphate, and the content of heavy metals in compound fertilizers.
4. The method according to any one of claims 1 to 3, characterized in that: In step S1, soil heavy metal data are obtained by collecting 0-20 cm surface soil samples and determining the heavy metal content using the acid digestion method, and environmental auxiliary variable data are obtained through remote sensing images, geospatial data cloud and field surveys.
5. The method according to claim 1, characterized in that In step S2, the following steps are specifically included: Step S2-1: Import the heavy metal content data and auxiliary variable data into two arrays respectively; Step S2-2: Select heavy metal content data and environmental auxiliary variable data and substitute them into the Pearson correlation coefficient calculation formula to obtain correlation analysis results; Step S2-3: Deleting environmental auxiliary variable data with low correlation according to the correlation analysis results in step S2-2; Step S2-4: Use the variance inflation factor (VIF) method to further screen variables and eliminate variables with higher VIF values. Here, variables with VIF greater than 10 are eliminated to obtain auxiliary variables that are significantly correlated with the heavy metal arsenic content in soil and construct a data set for modeling.
6. The method according to claim 1, characterized in that In step S3, the GWR model uses an adaptive bandwidth kernel function and the corrected Akaike information criterion AICc as the evaluation index for selecting the optimal bandwidth to generate the local regression coefficient of each sample point. The GWR formula is as follows: Where: y GWR(i) is the predicted value of the model at position i; β 0(i) is the intercept; β j(i) is the regression coefficient of the j covariate at position i; X j(i) is the value of the j covariate at position i; n is the number of covariates; ε (i) is the regression residual.
7. The method according to claim 1, characterized in that Step S4 specifically includes the following steps: Step S4-1: Import the regression residuals obtained in step S3 into a file; Step S4-2: interpolating the regression residuals using the ordinary kriging method; Step S4-3: Overlay the interpolation result with the GWR model result in S3. The formula is as follows: y GWR-OK =y gwr +r OK #(2) Where: y GWR-OK is the model's predicted value for heavy metal elements; gwr is the predicted value of GWR model; r OK is the value after interpolation of the GWR model regression residual using OK; Step S4-4: Get output result y GWR-OK This is the preliminary prediction result of the heavy metal arsenic content in the soil.
8. The method according to claim 1, characterized in that In step S6, the soft data integrated by the BME model are the predicted values calculated by the GWR-OK model, and the hard data are the measured soil heavy metal arsenic content data. The model optimizes the prediction accuracy through covariance fitting.
9. The method according to claim 8, characterized in that When using the BME model, the following steps are taken: S6-1: Based on the maximum entropy principle, using generalized knowledge G and specific knowledge S, and taking statistical moments such as mean, variance and covariance as constraints, we can obtain the prior probability density function f G , the formula is as follows: [x hard , x soft ] represents the hard data value x hard and the soft data value x soft The set composed of; x k For a point p to be predicted k The predicted value on ; Where: Indicates point p map The corresponding constraint condition expected value; g α (x map ) is a function based on prior knowledge G; g = {g α , α=1,…,N c } is the vector composed of various constraints; N c is the total number of constraints, and x map The total number is related; S6-2: Under certain constraints, formula (4) is maximized. The prior probability density function is: Where: A is the partition function, which is used to convert f G (x map ) normalized; μ = {μ α , α=1,…,N c } is the coefficient related to g, i.e., the Lagrange multiplier; S6-3: Arrange hard and soft data into appropriate forms; S6-4: Combining S and f G , under the generalized Bayesian condition, we get the posterior probability density function f k ; S6-5: By calculating the prediction results of the GWR-OK model, we obtain interval soft data, which are then put into formula (7) to obtain the posterior pdf of the variable distribution of the prediction point in the study area: Where: l and u are the lower and upper bounds of the interval respectively; S6-6: According to f k The predicted value of the predicted point can be calculated The formula is as follows: Where: This is the predicted value of the BMEGK model; S6-7: Get the output result This is the final prediction result.
10. The method according to claim 5, characterized in that The formula used in step S2-2 is as follows: Where N is the number of samples; X represents the auxiliary variable data; Y represents the soil heavy metal content data; Z x is the covariance of the auxiliary variable data; Z y is the covariance of soil heavy metal content data; S x is the standard deviation of the auxiliary variable data; S y is the standard deviation of soil heavy metal content data; is the mean value of auxiliary variable data; is the average value of soil heavy metal content data; i is the i-th sample point; X i is the auxiliary variable data at the i-th sample point; Y i is the soil heavy metal content data at the i-th sample point.
Citation Information
Cited By
Basalt area soil heavy metal background value fitting method based on main element geochemical constraint
CN122337376A