A method for soil pollution monitoring site selection based on geographic information

By using a soil pollution monitoring method based on geographic information systems, combining the spatial distribution of pollutants and influencing factors, sampling points can be accurately deployed, solving the problems of waste of sampling resources and inaccurate results in existing methods, and achieving efficient soil pollution monitoring.

CN121805556BActive Publication Date: 2026-05-26INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610267561.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-06
Publication Date
2026-05-26
Estimated Expiration
2046-03-06

AI Technical Summary

Technical Problem

Existing soil pollution monitoring methods ignore the spatial distribution characteristics of pollutants and their influencing factors, resulting in sampling results that cannot fully and objectively reflect the spatial distribution of pollutants, and also lead to a waste of sampling resources.

Method used

Based on geographic information systems, and combined with the spatial distribution and influencing factors of pollutants, an initial monitoring network is constructed. The grid is divided by spatial variation thresholds and variation indicators, and sampling points are precisely deployed to match the distribution characteristics of pollutants.

Benefits of technology

This approach enables the increase of sampling points in high-pollution-risk areas and the reduction of sampling points in low-risk areas, thereby reducing resource waste and providing scientific sampling results that serve as a basis for soil pollution remediation and safe agricultural production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121805556B_ABST
    Figure CN121805556B_ABST
Patent Text Reader

Abstract

This invention discloses a method for soil pollution monitoring site selection based on geographic information, belonging to the field of soil pollution monitoring technology. It addresses the problem that existing site selection methods neglect the distribution characteristics of pollutants and their influencing factors. The method includes: S1, acquiring the spatial distribution of pollutants in the soil within the monitoring area, and determining the influencing factors of spatial distribution based on the geographic information of the monitoring area; S2, determining the initial monitoring network and the spatial variation threshold of pollutants in the monitoring area based on the spatial distribution and influencing factors; S3, determining the spatial variation index of pollutants in each grid of the initial monitoring network based on the influencing factors, and dividing grids with spatial variation indices greater than the spatial variation threshold and grid sizes greater than a preset size into multiple next-level grids of the same size; S4, when all grids meet the preset conditions, deploying sampling points for soil pollution monitoring based on all the divided grids. This invention is used for soil pollution monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for determining soil pollution monitoring sites based on geographic information, belonging to the field of soil pollution monitoring technology. Background Technology

[0002] Currently, soil pollution monitoring sampling point placement generally relies on existing methods such as uniform grid placement or random placement. However, due to the significant spatial autocorrelation and heterogeneity of pollutants such as heavy metals in soil, existing placement methods ignore the non-uniform spatial distribution characteristics of pollutants. This leads to insufficient sampling in high-pollution-risk areas, failing to accurately characterize pollution distribution, while over-sampling in low-pollution-risk areas results in a waste of sampling resources. Furthermore, factors such as arable land distribution, river distribution, mineral resource distribution, topography, and wind fields are closely related to the migration and accumulation of pollutants, but existing placement methods do not consider the impact of these factors on the spatial distribution of pollutants, resulting in sampling results that cannot comprehensively and objectively reflect the spatial distribution of pollutants. Summary of the Invention

[0003] This invention provides a method for soil pollution monitoring sampling based on geographic information, which can solve the problem that existing sampling methods ignore the distribution characteristics of pollutants and their influencing factors, resulting in sampling results that cannot comprehensively and objectively reflect the spatial distribution of pollutants.

[0004] This invention provides a method for soil pollution monitoring site selection based on geographic information, the method comprising:

[0005] S1. Obtain the spatial distribution of pollutants in the soil within the monitoring area, and determine the influencing factors of the spatial distribution based on the geographical information of the monitoring area;

[0006] S2. Determine the initial monitoring network for the monitoring area and the spatial variation threshold of the pollutant in the monitoring area based on the spatial distribution and the influencing factors; the initial monitoring network consists of multiple grids of the same size.

[0007] S3. Determine the spatial variation index of the pollutant in each grid according to the influencing factor, and divide the grids with spatial variation index greater than the spatial variation threshold and grid size greater than the preset size into multiple next-level grids of the same size.

[0008] S4. When all grids meet the preset conditions, sampling points for soil pollution monitoring are deployed based on all the divided grids; the preset conditions are that the spatial variation index is less than or equal to the spatial variation threshold, or the grid size is less than or equal to the preset size.

[0009] Optionally, in S2, the initial monitoring network for the monitoring area is determined based on the spatial distribution and the influencing factors, specifically including:

[0010] The spatial autocorrelation distance of the pollutants is determined based on the spatial distribution, and the influence effect of the influence factors on the spatial distribution is determined based on the spatial distribution and the influence factors.

[0011] The initial monitoring network for the monitoring area is determined based on the spatial autocorrelation distance and the influence effect.

[0012] Optionally, determining the spatial autocorrelation distance of the pollutants based on the spatial distribution specifically includes:

[0013] Based on the spatial distribution, the spatial autocorrelation distance of the pollutants is determined using the Moran index and the semivariogram.

[0014] Optionally, determining the impact of the influencing factor on the spatial distribution specifically includes:

[0015] The correlation between the spatial distribution and the influencing factors was determined using regression analysis.

[0016] Based on the aforementioned correlation, a geographic detector is used to determine the impact of the influencing factors on the spatial distribution.

[0017] Optionally, there are multiple influencing factors; in S2, determining the spatial variation threshold of the pollutant in the monitoring area based on the spatial distribution and the influencing factors specifically includes:

[0018] Based on the spatial distribution and multiple influencing factors, a prediction model is constructed using a machine learning model; the prediction model is used to predict the spatial distribution based on all influencing factors.

[0019] The predictive model is interpreted using an interpretability model, and the spatial variation threshold of the pollutant in the monitoring area under the combined effect of all influencing factors is determined based on the interpretation results.

[0020] Optionally, each grid includes multiple geographic units; in S3, the spatial variation index of the pollutant in each grid is determined based on the influencing factors, specifically including:

[0021] The pollution risk index for each geographic unit is determined based on all influencing factors;

[0022] The spatial variation index of the pollutant in the corresponding grid is determined based on the pollution risk index of all geographical units in each grid.

[0023] Optionally, a pollution risk index for each geographic unit is determined based on all influencing factors, specifically including:

[0024] Determine the quantitative indicators for each geographic unit corresponding to each influencing factor;

[0025] The pollution risk index for each geographic unit is determined based on the quantitative indicators of all influencing factors corresponding to each geographic unit.

[0026] Optionally, sampling points for soil pollution monitoring are deployed in S4 based on all the divided grids, specifically including:

[0027] The minimum allowable spacing between adjacent sampling points is determined based on the spatial autocorrelation distance.

[0028] The sampling points are arranged according to the minimum allowable spacing and the spatial variation index.

[0029] Optionally, the sampling points are arranged according to the minimum allowable spacing and the spatial variation index, specifically including:

[0030] The sampling points are set within the grid according to the spatial variation index of the grid;

[0031] When the distance between adjacent sampling points is less than the minimum allowable spacing, the placement of the sampling points is adjusted within a preset radius range until the distance between adjacent sampling points is greater than or equal to the minimum allowable spacing.

[0032] Optionally, the sampling points are set within the grid according to the spatial variation index of the grid, specifically including:

[0033] When the spatial variation index of the grid is greater than the spatial variation threshold, the sampling point is set at the geographical unit with the highest pollution risk index in the grid.

[0034] Alternatively, when the spatial variation index of the grid is less than or equal to the spatial variation threshold, the sampling point is set at the geometric center of the grid.

[0035] The beneficial effects that this invention can produce include:

[0036] This invention determines an initial monitoring network for a monitoring area based on the spatial distribution of pollutants and their influencing factors. It also establishes a spatial variability threshold to measure the spatial variability of pollutants within the monitoring area. The initial monitoring network is then divided into several grids based on this threshold, and spatial variability indices for pollutants in each grid are determined according to the influencing factors. The grids are further divided based on these indices, and finally, sampling points are deployed based on the divided network. This approach effectively combines the influence of influencing factors on the spatial distribution of pollutants, ensuring that the placement of sampling points matches the distribution characteristics of the pollutants. Specifically, sampling points are increased in areas with high spatial variability and pollution risk to accurately characterize the spatial distribution of pollutants, while sampling points are reduced in areas with low spatial variability and pollution risk to minimize redundant costs caused by oversampling. This allows the sampling results to comprehensively and objectively reflect the spatial distribution of pollutants, providing a scientific basis for soil pollution remediation and agricultural safety production. Attached Figure Description

[0037] Figure 1 A flowchart illustrating the method for soil pollution monitoring site selection based on geographic information provided in this embodiment of the invention. Detailed Implementation

[0038] The present invention will now be described in detail with reference to the embodiments, but the present invention is not limited to these embodiments.

[0039] This invention provides a method for soil pollution monitoring site selection based on geographic information, such as... Figure 1 As shown, the method includes:

[0040] S1. Obtain the spatial distribution of pollutants in the soil within the monitoring area, and determine the influencing factors of spatial distribution based on the geographical information of the monitoring area.

[0041] Among them, pollutants include heavy metals, pesticides and radioactive substances; the spatial distribution of pollutants includes the concentration of pollutants in the soil at different locations within the monitoring area; geographical information includes the geological conditions, climate conditions and topographic conditions of the monitoring area; influencing factors include at least one of the following: distribution of cultivated land, distribution of rivers, distribution of mineral resources, topography and wind field. This embodiment uses multiple influencing factors as an example for illustration.

[0042] S2. Determine the initial monitoring network and spatial variability thresholds of pollutants in the monitoring area based on spatial distribution and influencing factors, specifically including:

[0043] 1. Determine the initial monitoring network for the monitoring area.

[0044] 1) Collect data on influencing factors such as the distribution of cultivated land, rivers, mineral resources, topography, and wind fields within the monitoring area, and then perform preprocessing on all data, including coordinate system 1, resolution matching, and noise removal.

[0045] 2) Determine the spatial autocorrelation distance of pollutants based on the spatial distribution data of pollutants (i.e., the concentration of pollutants in soil at different locations), specifically including:

[0046] First, based on the pollutant concentrations in the soil at different locations within the monitoring area, the spatial aggregation and spatial heterogeneity of pollutants within the monitoring area were quantitatively analyzed using the Global Moran's I and Local Moran's I indices.

[0047] Then, using the detrended pollutant concentration sequences in soil at different locations as a sample set, based on the aforementioned spatial clustering or spatial heterogeneity, and according to the step size... Calculate the semivariogram and obtain its value.

[0048] The semi-variogram can be expressed as:

[0049] (1)

[0050] In formula (1), This represents the value of the semi-mutation function; and They represent the distances between each other in the sample sets. The The spatial point and the first A spatial point; and They represent the first and second samples in the sample set, respectively. spatial points and the spatial points The corresponding sample data values; and Form a set of sample pairs; This indicates that the mutual distances satisfy the distance interval. The number of sample pairs; The step size is determined based on the spatial clustering or spatial heterogeneity obtained above.

[0051] The determination process is an existing technology, which specifically involves: determining the peak value of spatial clustering based on spatial clustering; determining the spatial autocorrelation scale (range) based on the distance corresponding to the peak value; and then dividing the range into several equal intervals (usually 10 to 15), with the distance interval between each interval being the step size. Alternatively, determine the scale of hotspot regions based on spatial heterogeneity, then calculate the average nearest neighbor distance between spatial points based on the scale of the hotspot regions, and use the average nearest neighbor distance as the step size. .

[0052] Subsequently, spherical, exponential, and Gaussian models were used respectively, and the semivariogram values ​​were fitted based on the least squares method. Then, the root mean square error (RMSE) and the Akaike information criterion (AIC) are used to cross-validate the fitting effect of each model, and the best fitting result is selected.

[0053] Finally, the spatial autocorrelation distance is determined based on the optimal fitting result. Specifically, if the semivariogram in the optimal fitting result shows anisotropy, the range in the main range direction is determined as the spatial autocorrelation distance of the pollutants; if the semivariogram in the optimal fitting result shows isotropy, the corresponding effective range of the model corresponding to the optimal fitting result is determined as the spatial autocorrelation distance of the pollutants.

[0054] 3) Determine the impact of influencing factors on spatial distribution based on spatial distribution data and influencing factor data, specifically including:

[0055] Based on pollutant concentration and influencing factor data, regression analysis was used to determine the correlation between pollutant concentration and each influencing factor, including linear and nonlinear relationships. Then, based on these correlations, a geodetector spatial probing method was used to analyze the interaction between pollutant concentration and each influencing factor, determining the impact of each influencing factor on pollutant concentration, including linear and nonlinear enhancement effects.

[0056] 4) Determine the initial monitoring network for the monitoring area based on spatial autocorrelation distance and impact effects, specifically including:

[0057] First, based on the spatial autocorrelation distance and the influence of various factors on pollutant concentration, multiple suitable candidate monitoring networks are constructed in a Geographic Information System (GIS) or Python analysis environment (such as PyKrige or geodetector). Each candidate monitoring network consists of multiple square grids of the same size, but the grid size of each candidate monitoring network is different. For example, the grid size can be 2km×2km, 4km×4km, 8km×8km, 16km×16km, etc. These grid sizes can initially meet the monitoring requirements.

[0058] Then, a comprehensive evaluation is conducted on the spatial variance explanation rate, pollution prediction error, and sampling cost of each candidate monitoring network to determine the monitoring accuracy and economy of each candidate monitoring network.

[0059] Finally, based on the actual situation, a candidate monitoring network that can achieve the optimal balance between monitoring accuracy and economy is selected as the initial monitoring network. For example, the results of the verification in the monitoring area of ​​the Xiangjiang River Basin show that a monitoring network with a grid size of 8km×8km can effectively capture the spatial distribution characteristics of the heavy metal cadmium (Cd) in the soil.

[0060] 2. Determine the spatial variability threshold of pollutants in the monitoring area.

[0061] 1) Construct a dataset based on the spatial distribution data of pollutants (i.e., the concentration of pollutants in soil at different locations) and the data of each influencing factor, and divide the dataset into a training set and a test set.

[0062] 2) Construct a prediction model based on the training set and prediction set, specifically including:

[0063] First, a machine learning model (such as a random forest model or a support vector machine) is trained using a training set to enable the model to establish a mapping relationship between the spatial distribution of pollutants and all influencing factors. Simultaneously, the hyperparameters of the machine learning model are optimized during training using a random grid search method.

[0064] After training, the trained machine learning model is tested using a test set. The coefficient of determination (R-squared, or R for short) is used to determine the model's performance. 2 When evaluation metrics such as root mean square error (RMSE) and mean absolute error (MAE) all achieve optimal values, it indicates that the machine learning model has the best predictive performance. This machine learning model is used as a predictive model for the spatial distribution of pollutants, and it is used to predict pollutant concentrations at different locations based on data from all influencing factors.

[0065] 3) Interpretive models (such as SHAP and LIME models) are used to interpret the prediction model. Interpretive models can decompose the prediction results (i.e., spatial distribution of pollutants) into the contribution of each feature (i.e., each influencing factor) and obtain interpretive results. These interpretive results include the positive, negative, and average effects of different influencing factors on pollutant concentrations at different locations. Based on these interpretive results, the spatial variability threshold of pollutants in the monitoring area under the combined effect of all influencing factors can be determined. This spatial variability threshold is used in subsequent steps to measure the spatial heterogeneity and pollution risk of pollutants at each grid in the initial monitoring network.

[0066] S3. Determine the spatial variability index of pollutants in each grid based on the influencing factors. Divide the grids with spatial variability indices greater than the spatial variability threshold and grid sizes larger than the preset size into multiple next-level grids of the same size. Specifically, this includes:

[0067] 1. Determine the quantitative indicators for each geographic unit corresponding to the influencing factors.

[0068] Each grid of the initial monitoring network consists of multiple geographic units (such as pixels). Based on the data of each influencing factor, the quantitative indicators corresponding to each geographic unit for each influencing factor can be determined.

[0069] For example, the arable land density and type of each geographic unit can be determined based on arable land distribution data; the distance between each geographic unit and the river, as well as the river's direction and velocity, can be determined based on river distribution data; the distance between each geographic unit and the mineral resources can be determined based on mineral resource distribution data; the topographic relief, slope, and runoff accumulation of each geographic unit can be determined based on topographic data; and the wind direction and speed of each geographic unit can be determined based on wind field data.

[0070] 2. Determine the pollution risk index of the corresponding geographical unit based on the quantitative indicators of all influencing factors corresponding to each geographical unit.

[0071] The pollution risk index of a geographical unit can be expressed by a standardized function as follows:

[0072] (2)

[0073] In formula (2), A pollution risk index representing a specific geographical unit; They represent the 1st to the 1st. Each influencing factor corresponds to a quantitative indicator for that geographic unit. They represent the 1st to the 1st. The weight coefficients of each influencing factor can be determined by methods such as expert experience, analytic hierarchy process, or machine learning algorithms. The standardization function can map the calculated pollution risk index to a numerical range of [0,1] or [0,100].

[0074] Since the quantitative indicators of each influencing factor in a specific geographic unit can reflect the impact of that factor on the pollutant concentration in that unit—for example, higher arable land density, proximity to rivers, proximity to mineral resources, higher wind speeds, or more dramatic topographic relief in a geographic unit all have a significant positive impact on pollutant deposition in that unit, leading to increased pollutant concentrations—the pollution risk indicators calculated based on all influencing factors can reflect the potential pollution risk of a geographic unit under the combined effect of all influencing factors.

[0075] 3. Determine the spatial variation index of pollutants in the corresponding grid based on the pollution risk index of all geographic units in each grid.

[0076] The spatial variability index of a grid is used to describe the dispersion of the pollution risk index of all geographic units in the grid. It can be the standard deviation, variance or coefficient of variation of the pollution risk index of all geographic units in the grid, and can reflect the spatial variability of pollutants in the grid.

[0077] 4. Subdivide the initial monitoring network based on all grid spatial variation indicators.

[0078] In the initial monitoring network, grids with spatial variability indices greater than the spatial variability threshold and grid sizes larger than the preset size are divided into multiple next-level grids of the same size. Then, the spatial variability index of each next-level grid is determined based on the quantitative index of all geographic units in each next-level grid, and the next-level grid is further divided based on the spatial variability index of the next-level grid. This process of dividing the grids level by level continues until all grids meet the preset conditions.

[0079] The preset conditions are that the spatial variability index is less than or equal to the spatial variability threshold, or the grid size is less than or equal to the preset size. The preset size can be flexibly set according to monitoring needs; in this embodiment, the preset size is the minimum area determined based on the feasibility of actual ground sampling.

[0080] This embodiment can use a quadtree coding algorithm to divide the grid level by level. The pseudocode example of the algorithm is as follows:

[0081] Step 1: Input the initial monitoring network And set a spatial variation threshold. ;

[0082] Step 2: Initial monitoring network As the root node, compute the values ​​for each of its grid cells. Pollution risk index of all pixels The variance is used to obtain the variance of each grid cell. Spatial variability index ;

[0083] Step 3: For the current grid ,like > If the grid size is larger than the preset size, then the current grid will be... Divided into 4 sub-grids;

[0084] Step 4: Repeat Step 2 and Step 3 for each newly generated next-level mesh until every mesh satisfies the requirements. ≤ Or the grid size is less than or equal to the preset size;

[0085] Step 5: Output the final layered quadtree node set to obtain the completed grid.

[0086] S4. When all grids meet the preset conditions, sampling points for soil pollution monitoring are deployed based on all the divided grids, specifically including:

[0087] 1. Determine the minimum allowable distance between adjacent sampling points based on the spatial autocorrelation distance of pollutants.

[0088] For example, the spatial autocorrelation distance of pollutants is denoted as Let the side length of a certain grid be denoted as The minimum allowable spacing between sampling points within the grid can be expressed as: ,

[0089] 2. Set sampling points within the grid based on the spatial variation index of the grid, specifically including:

[0090] 1) When the spatial variation index of the grid is greater than the spatial variation threshold (i.e.) > (Time), indicating that the spatial variability of pollutants within this grid is significant, and a pollution risk index can be established within this grid. Sampling points are set at the largest geographic unit.

[0091] 2) When the spatial variation index of the grid is less than or equal to the spatial variation threshold (i.e.) ≤ (When), it indicates that the spatial variability of pollutants within the grid is not significant, and sampling points can be set at the geometric center of the grid.

[0092] 3. After setting up the sampling points according to the above steps, if the distance between adjacent sampling points is greater than or equal to the minimum allowable spacing, it indicates that the sampling point layout is reasonable; conversely, if the distance between adjacent sampling points is less than the minimum allowable spacing, it indicates that the sampling point layout is unreasonable. The layout of the sampling points can be adjusted within the preset radius range until the distance between adjacent sampling points is greater than or equal to the minimum allowable spacing. The preset radius range can be flexibly set according to the actual situation.

[0093] For example, when adjusting the location of sampling points, this embodiment can use a circle centered on the sampling point and... Find the nearest location within the radius's neighborhood that satisfies the minimum allowable spacing. If this still cannot be satisfied, the radius can be increased to... Then, the location of the sampling points is shifted to the nearest available location.

[0094] In addition, in order to ensure that the layout of sampling points within the same grid can balance randomness and uniformity, this embodiment can also use the method of "regular grid center + slight disturbance" to lay out the sampling points.

[0095] For example, equidistant grid centers within the grid boundary can be used as candidate locations for sampling points to ensure the overall uniformity of the grid; then... direction and Apply no more than 100% of the force to the candidate deployment positions in each direction. The sampling point is randomly perturbed, and then the distance between the perturbed sampling point and the nearest sampling point (nearest neighbor distance) is checked. If the nearest neighbor distance is greater than or equal to the minimum allowable distance, the perturbed position is taken as the final sampling point location; if the nearest neighbor distance is less than the minimum allowable distance, the candidate location before the perturbed position is taken as the final sampling point location. Through the above steps, the sampling point layout can achieve a repeatable balance of "primarily uniform, supplemented by random"

[0096] After the above adjustment process is completed, this embodiment can also perform interpolation verification and necessary fine-tuning of the sampling point layout. For example, the sampling point layout scheme can be interpolated and verified once using the Kriging interpolation method or the inverse distance weighted interpolation method, and the RMSE can be calculated; if the RMSE exceeds the preset threshold, one sampling point is added in the grid with the larger error and the adjustment process of S4 is repeated; if the RMSE meets the preset threshold, the sampling point layout scheme passes the verification.

[0097] In practical applications, this embodiment will output a vector map of sampling points after the sampling point layout scheme has been verified, including the coordinates of each sampling point, its corresponding quadtree level, and the selection criteria (“ Maximum / geometric center / offset markers) and nearest neighbor distance, etc., are used for field implementation and quality traceability.

[0098] Furthermore, since soil pollution exhibits spatiotemporal evolution characteristics, this embodiment can also dynamically adjust the monitoring site selection scheme based on the latest monitoring data, enabling the scheme to have dynamic adaptive capabilities. The specific process is as follows:

[0099] 1. Regularly collect the latest monitoring data on the spatial distribution of pollutants and compare it with existing data to identify areas where pollutant concentrations or spatial variability have changed significantly;

[0100] 2. Recalculate the spatial variation index of each grid based on the latest monitoring data, and adjust the quadtree structure of the monitoring network according to the spatial variation index. For example, further subdivide the grid in areas where the spatial variation index increases, or merge some subdivided grids in areas where the spatial variation index decreases.

[0101] 3. The updated site selection plan is automatically fed back to the sampling deployment system to achieve real-time optimization.

[0102] This embodiment determines the initial monitoring network of the monitoring area based on the spatial distribution of pollutants and their influencing factors, and determines the spatial variability threshold for measuring the spatial variability of pollutants in the monitoring area. Then, based on the spatial variability threshold, the initial monitoring network is divided into several grids, and the spatial variability index of pollutants in each grid is determined according to the influencing factors. The grids are then divided according to the spatial variability index, and finally, sampling points are deployed based on the divided network. This fully combines the influence of influencing factors on the spatial distribution of pollutants, matching the deployment of sampling points with the distribution characteristics of pollutants. That is, more sampling points are added in areas with high spatial variability and pollution risk to accurately characterize the spatial distribution of pollutants, while fewer sampling points are added in areas with low spatial variability and pollution risk to reduce the redundant costs caused by oversampling. Thus, the sampling results can comprehensively and objectively reflect the spatial distribution of pollutants, providing a scientific basis for soil pollution remediation, agricultural safety production, and other fields.

[0103] The above description is merely a few embodiments of this application and is not intended to limit this application in any way. Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any changes or modifications made by those skilled in the art without departing from the scope of the technical solution of this application using the disclosed technical content are equivalent to equivalent implementation cases and fall within the scope of the technical solution.

Claims

1. A method for soil pollution monitoring site selection based on geographic information, characterized in that, The method includes: S1. Obtain the spatial distribution of pollutants in the soil within the monitoring area, and determine the influencing factors of the spatial distribution based on the geographical information of the monitoring area; S2. Determine the spatial autocorrelation distance of the pollutants based on the spatial distribution; determine the influence effect of the influence factors on the spatial distribution based on the spatial distribution and the influence factors; construct multiple candidate monitoring networks based on the spatial autocorrelation distance and the influence effect; select one of the multiple candidate monitoring networks as the initial monitoring network based on monitoring accuracy and economy; the initial monitoring network consists of multiple grids of the same size; construct a prediction model based on all influence factors to predict the spatial distribution using a machine learning model; interpret the prediction model using an interpretability model; and determine the spatial variation threshold of the monitoring area based on the interpretation results. S3. Determine the quantitative index of each geographical unit in the initial monitoring network corresponding to each influencing factor, determine the pollution risk index of each geographical unit according to the quantitative index, calculate the standard deviation, variance or coefficient of variation according to the pollution risk index of all geographical units in each grid, use the calculation result as the spatial variation index of the corresponding grid, and divide the grid with the spatial variation index greater than the spatial variation threshold and the grid size greater than the preset size into multiple next-level grids with the same size. S4. When all grids meet the preset conditions, sampling points for soil pollution monitoring are deployed based on all the divided grids; the preset conditions are that the spatial variation index is less than or equal to the spatial variation threshold, or the grid size is less than or equal to the preset size; the preset size is the minimum area determined based on the actual sampling feasibility on the ground.

2. The method according to claim 1, characterized in that, Determining the spatial autocorrelation distance of the pollutants based on their spatial distribution specifically includes: Based on the spatial distribution, the spatial autocorrelation distance of the pollutants is determined using the Moran index and the semivariogram.

3. The method according to claim 1, characterized in that, Determining the impact of the influencing factors on the spatial distribution specifically includes: The correlation between the spatial distribution and the influencing factors was determined using regression analysis. Based on the aforementioned correlation, a geographic detector is used to determine the impact of the influencing factors on the spatial distribution.

4. The method according to claim 1, characterized in that, S4 uses sampling points for soil pollution monitoring based on all the grids after division, specifically including: The minimum allowable spacing between adjacent sampling points is determined based on the spatial autocorrelation distance. The sampling points are arranged according to the minimum allowable spacing and the spatial variation index.

5. The method according to claim 4, characterized in that, The sampling points are arranged according to the minimum allowable spacing and the spatial variation index, specifically including: The sampling points are set within the grid according to the spatial variation index of the grid; When the distance between adjacent sampling points is less than the minimum allowable spacing, the placement of the sampling points is adjusted within a preset radius range until the distance between adjacent sampling points is greater than or equal to the minimum allowable spacing.

6. The method according to claim 5, characterized in that, The sampling points are set within the grid according to the spatial variation index of the grid, specifically including: When the spatial variation index of the grid is greater than the spatial variation threshold, the sampling point is set at the geographical unit with the highest pollution risk index in the grid. Alternatively, when the spatial variation index of the grid is less than or equal to the spatial variation threshold, the sampling point is set at the geometric center of the grid.

Citation Information

Patent Citations

  • Staged and partitioned sampling method for agricultural land soil pollution investigation

    CN111707490A

  • Large-area soil heavy metal detection and spatial and temporal distribution characteristic analysis method and system

    CN112986538A