A method, system and storage medium for soil remediation

By constructing a prior probability distribution model and using a Bayesian fusion method, the three-dimensional location of soil sampling points is optimized, solving the problems of inaccurate soil sampling and high cost in existing technologies, and realizing efficient and economical sampling of complex contaminated sites.

CN122287360APending Publication Date: 2026-06-26ZHENGZHOU UNIVERSITY OF AERONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610424047.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-01
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing soil sampling methods cannot effectively integrate prior knowledge, are difficult to represent uncertainty, and cannot collaboratively optimize the three-dimensional location of sampling points, resulting in inaccurate representation of complex contaminated sites and high costs.

Method used

By constructing a prior probability distribution model, updating the posterior probability distribution field using a Bayesian fusion method, calculating the three-dimensional uncertainty field, generating non-uniform Voronoi diagram sampling cells, identifying candidate key sampling depths, calculating sampling priority scores, and optimizing sampling point configuration.

Benefits of technology

It improves the accuracy and reliability of predicting the spatial distribution of pollutants, focuses on key areas, avoids redundant sampling, optimizes the allocation of sampling resources, and reduces exploration costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287360A_ABST
    Figure CN122287360A_ABST
Patent Text Reader

Abstract

This invention provides a stratified sampling method, system, and storage medium for soil remediation. The method includes: integrating historical knowledge base and field survey data; generating a three-dimensional posterior probability distribution and uncertainty field of pollutants using a Bayesian model; generating sampling cells using non-uniform Voronoi partitioning based on the uncertainty intensity plane map; automatically identifying key candidate sampling depths on the vertical probability profile of each cell by analyzing local extrema and curvature changes; calculating priority scores by comprehensively considering the uncertainty, probability value, and derivative amplitude of each depth; and selecting the optimal sampling points after global sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of stratified sampling, and in particular relates to a stratified sampling method, system and storage medium for soil remediation. Background Technology

[0002] Soil pollution remediation is a crucial component of environmental protection. The primary step is conducting on-site investigations and sample collection and analysis of contaminated sites to determine the types, concentrations, and three-dimensional spatial distribution characteristics of pollutants. Traditional soil sampling point placement methods include judgmental placement, systematic placement, and random placement. Judgmental placement relies on the experience of professionals, is highly subjective, and has poor repeatability. While systematic placement provides uniform spatial coverage, its rigid approach can lead to the omission of key data points or the placement of too many invalid points in clean areas, resulting in high sampling costs and low efficiency. Random placement may result in uneven distribution of sampling points, failing to comprehensively represent the overall pollution status. Traditional methods often determine vertical depth by using equidistant stratification or setting several fixed depths based on experience, failing to consider the detailed migration patterns of pollutants along depth. This makes it difficult to detect pollution enrichment layers or boundary layers formed due to changes in soil texture and hydrogeological conditions.

[0003] Most of these methods focus on the optimal estimation of pollutant concentrations, while paying insufficient attention to the uncertainty of the prediction results. They fail to prioritize the allocation of sampling resources to the areas with the greatest model uncertainty to obtain new information. Furthermore, a systematic and quantitative optimization mechanism has not yet been established for deciding on vertical sampling depth, making it impossible to automatically identify and locate critical depth points that are crucial for understanding the vertical distribution patterns of pollutants. Therefore, existing technologies urgently need a novel stratified sampling method that can integrate prior knowledge, represent uncertainty, and collaboratively optimize the three-dimensional spatial location of sampling points to achieve a more accurate and economical representation of complex contaminated sites. Summary of the Invention

[0004] To address the problem that existing technologies fail to integrate prior knowledge, struggle to represent uncertainty, and fail to collaboratively optimize the three-dimensional location of sampling points, thus hindering the accurate and economical representation of complex contaminated sites.

[0005] In the first aspect, the present invention proposes a stratified sampling method for soil remediation, comprising the following steps:

[0006] Using historical data on pollutant types, concentrations, spatial distributions, and geological and hydrological conditions stored in the sampling knowledge base, a prior probability distribution model representing the vertical migration patterns of pollutants in different soil media is constructed. Combined with preliminary survey data of the site to be remediated, the prior probability distribution model is updated using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants for the site to be remediated.

[0007] Based on the posterior probability distribution field, the three-dimensional pollution distribution uncertainty field of the entire site is calculated; by integrating or averaging the three-dimensional uncertainty field along the depth direction, a two-dimensional uncertainty intensity plane map is generated; based on the two-dimensional uncertainty intensity plane map, a set of non-uniformly distributed Voronoi diagram generation points are generated, and the site to be repaired is spatially partitioned based on the generation points to generate a set of non-uniform sampling cells.

[0008] For each sampling cell, extract the probability profile of the posterior probability distribution field on the vertical axis passing through the Voronoi generation point; calculate the local extrema and the zero of the second derivative of the profile curve, and define the depths corresponding to the local extrema and the zero of the second derivative as a set of candidate key sampling depths within the cell.

[0009] Calculate a sampling priority score for each candidate key sampling depth; globally sort all candidate key sampling depths of all cells according to their priority scores; select several candidate key sampling depths with the highest priority scores from the sorted list according to the preset total number of sampling points, and determine the position of the sampling depth as the sampling point to obtain a hierarchical sampling scheme.

[0010] Optionally, updating the prior probability distribution model using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated includes:

[0011] The prior probability distribution model is used as the prior distribution P(θ) for Bayesian inference, and the pollutant concentration values ​​obtained from the preliminary survey data of the land to be restored are used as the observation data D.

[0012] Construct a Gaussian likelihood function P(D|θ) based on the observed data D;

[0013] Based on Bayes' theorem P(θ|D)∝P(D|θ)P(θ), a numerical solution is performed to calculate the posterior probability distribution field P(θ|D) representing the probability of pollutant concentration distribution in three-dimensional space.

[0014] Optionally, calculating the three-dimensional pollution distribution uncertainty field of the entire land parcel based on the posterior probability distribution field includes:

[0015] For each three-dimensional spatial grid node in the posterior probability distribution field, extract the posterior probability distribution of pollutant concentration at the node;

[0016] Calculate the standard deviation of the posterior probability distribution and use the standard deviation value as the pollution distribution uncertainty value at this node;

[0017] By traversing all grid nodes, the three-dimensional uncertainty field of the entire land parcel, composed of the uncertainty values ​​of each node, is obtained.

[0018] Optionally, generating a set of non-uniformly distributed Voronoi diagram generation points based on the two-dimensional uncertainty intensity plane diagram includes:

[0019] The spatial density of the generated points is set to be proportional to the two-dimensional uncertainty intensity value of the location of the generated points;

[0020] Based on the density relationship, a set of non-uniformly distributed Voronoi diagram generation points are generated on the two-dimensional uncertainty intensity plane using the Poisson disk sampling algorithm.

[0021] Optionally, the step of calculating the local extrema and the zeros of the second derivative of the profile curve, and defining the depths corresponding to the local extrema and the zeros of the second derivative as a set of candidate key sampling depths within the cell, includes:

[0022] The vertical probability profile curve is discretized into a series of depth-probability data point pairs. ;

[0023] Identify and satisfy and depth As a local maximum point, identify the condition that satisfies and depth As a local minimum point;

[0024] Calculating the second derivative using the central difference method and identify those that meet the requirements. Depth < 0 As the zero of the second derivative;

[0025] By summing up all identified depth points, a set of candidate key sampling depths for each cell is obtained.

[0026] Optionally, calculating a sampling priority score for each candidate key sampling depth includes:

[0027] For the j-th candidate key sampling depth point in the i-th cell Priority score Calculated using the weighted formula: ;

[0028] in, This is the normalized value of the uncertainty intensity of cell i containing the candidate key sampling depth point. This is the normalized value of the mean posterior probability of the depth point. This is the normalized value of the absolute value of the first or second derivative of the depth point. , , These are the preset non-negative weighting coefficients.

[0029] Optionally, the step of constructing a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media using historical landfill pollutant types, concentrations, spatial distributions, and geological and hydrological data stored in the sampling knowledge base includes:

[0030] The sampling knowledge base includes, but is not limited to: geological borehole data from similar historical sites, soil physicochemical properties, pollutant concentration monitoring data, hydrogeological parameters, and parameters of validated pollutant migration models.

[0031] In another aspect, the present invention also proposes a stratified sampling system for soil remediation, comprising the following modules:

[0032] The update module is used to construct a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media by utilizing historical pollutant types, concentrations, spatial distributions, and geological and hydrological data stored in the sampling knowledge base; combined with the preliminary survey data of the site to be remediated, the prior probability distribution model is updated by a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated.

[0033] The generation module is used to calculate the three-dimensional pollution distribution uncertainty field of the entire plot based on the posterior probability distribution field; generate a two-dimensional uncertainty intensity plane map by integrating or averaging the three-dimensional uncertainty field along the depth direction; generate a set of non-uniformly distributed Voronoi diagram generation points based on the two-dimensional uncertainty intensity plane map, and spatially partition the plot to be repaired based on the generation points to generate a set of non-uniform sampling cells.

[0034] The calculation module is used to extract the probability profile of the posterior probability distribution field on the vertical axis passing through the Voronoi generation point for each sampling cell; calculate the local extrema and the zero of the second derivative of the profile curve, and define the depths corresponding to the local extrema and the zero of the second derivative as a set of candidate key sampling depths within the cell.

[0035] The determination module is used to calculate the sampling priority score for each candidate key sampling depth; to globally sort all candidate key sampling depths of all cells according to the priority score; and to select several candidate key sampling depths with the highest priority scores from the sorted list according to the preset total number of sampling points, and to determine the position of the sampling depth as the sampling point to obtain the hierarchical sampling scheme.

[0036] Preferably, the step of updating the prior probability distribution model using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated includes:

[0037] The prior probability distribution model is used as the prior distribution P(θ) for Bayesian inference, and the pollutant concentration values ​​obtained from the preliminary survey data of the land to be restored are used as the observation data D.

[0038] Construct a Gaussian likelihood function P(D|θ) based on the observed data D;

[0039] Based on Bayes' theorem P(θ|D)∝P(D|θ)P(θ), a numerical solution is performed to calculate the posterior probability distribution field P(θ|D) representing the probability of pollutant concentration distribution in three-dimensional space.

[0040] Preferably, calculating the three-dimensional pollution distribution uncertainty field of the entire site based on the posterior probability distribution field includes:

[0041] For each three-dimensional spatial grid node in the posterior probability distribution field, extract the posterior probability distribution of pollutant concentration at the node;

[0042] Calculate the standard deviation of the posterior probability distribution and use the standard deviation value as the pollution distribution uncertainty value at this node;

[0043] By traversing all grid nodes, the three-dimensional uncertainty field of the entire land parcel, composed of the uncertainty values ​​of each node, is obtained.

[0044] Preferably, the step of generating a set of non-uniformly distributed Voronoi diagram generation points based on the two-dimensional uncertainty intensity plane diagram includes:

[0045] The spatial density of the generated points is set to be proportional to the two-dimensional uncertainty intensity value of the location of the generated points;

[0046] Based on the density relationship, a set of non-uniformly distributed Voronoi diagram generation points are generated on the two-dimensional uncertainty intensity plane using the Poisson disk sampling algorithm.

[0047] Preferably, the step of calculating the local extrema and the zeros of the second derivative of the profile curve, and defining the depths corresponding to the local extrema and the zeros of the second derivative as a set of candidate key sampling depths within the cell, includes:

[0048] The vertical probability profile curve is discretized into a series of depth-probability data point pairs. ;

[0049] Identify and satisfy and depth As a local maximum point, identify the condition that satisfies and depth As a local minimum point;

[0050] Calculating the second derivative using the central difference method and identify those that meet the requirements. Depth < 0 As the zero of the second derivative;

[0051] By summing up all identified depth points, a set of candidate key sampling depths for each cell is obtained.

[0052] Preferably, calculating the sampling priority score for each candidate key sampling depth includes:

[0053] For the j-th candidate key sampling depth point in the i-th cell Priority score Calculated using the weighted formula: ;

[0054] in, This is the normalized value of the uncertainty intensity of cell i containing the candidate key sampling depth point. This is the normalized value of the mean posterior probability of the depth point. This is the normalized value of the absolute value of the first or second derivative of the depth point. , , These are the preset non-negative weighting coefficients.

[0055] Preferably, the step of constructing a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media using historical site pollutant types, concentrations, spatial distributions, and geological and hydrological data stored in the sampling knowledge base includes:

[0056] The sampling knowledge base includes, but is not limited to: geological borehole data from similar historical sites, soil physicochemical properties, pollutant concentration monitoring data, hydrogeological parameters, and parameters of validated pollutant migration models.

[0057] This invention improves the accuracy and reliability of predicting the spatial distribution of pollution by integrating historical sampling knowledge bases with preliminary survey data of sites to be remediated, constructing and updating a three-dimensional posterior probability distribution field of pollutants. By representing the overall uncertainty and generating non-uniform sampling cells, it achieves focus on key areas, avoiding redundant sampling in low-risk areas, thereby optimizing the allocation of sampling resources and reducing survey costs. By identifying local extrema and zeros of the second derivative on the vertical probability profile as candidate critical depths, it can detect the vertical boundary and internal structural features of the pollution plume. By establishing a comprehensive sampling priority evaluation system and performing global ranking optimization, it ensures that the maximum amount of site information is obtained with a limited number of sampling points. Attached Figure Description

[0058] Figure 1 This is a flowchart of Example 1;

[0059] Figure 2 This is a schematic diagram of the Bayesian fusion process;

[0060] Figure 3 This is a schematic diagram illustrating the calculation of sampling priority score. Detailed Implementation

[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0062] Example 1 provides a stratified sampling method for soil remediation, such as... Figure 1 As shown, it includes the following steps:

[0063] S1. Using historical data on pollutant types, concentrations, spatial distributions, and geological and hydrological conditions stored in the sampling knowledge base, a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media is constructed. Combined with preliminary survey data of the site to be restored, the prior probability distribution model is updated using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be restored.

[0064] Historical data was cleaned and integrated using the Pandas library in Python. Pollutant concentration was used as the target variable, and three-dimensional spatial coordinates (X, Y, Z) and geological and hydrological parameters such as soil type and permeability coefficient were used as feature variables. The Gaussian Process Regressor algorithm from the scikit-learn library was used as the training set to construct a prior probability distribution model. This model can predict the mean and variance of pollutant concentration at any spatial point. Preliminary survey data of the site to be remediated was used as new observation data and merged with historical data to form a new training set. The `fit` method of the Gaussian Process Regression model was called again to update the model. Then, the `predict` method was called on the three-dimensional grid points to obtain the posterior mean and posterior variance. Alternatively, a Bayesian recursive update formula was used, with the preliminary survey data as conditional data to directly update the predicted distribution of the Gaussian process, thereby calculating the posterior probability distribution, i.e., the posterior mean and posterior variance, for each grid point on the three-dimensional grid of the site. Figure 2 As shown, a three-dimensional posterior probability distribution field of pollutants is formed.

[0065] In an optional embodiment, updating the prior probability distribution model using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants for the site to be remediated includes:

[0066] The prior probability distribution model is used as the prior distribution P(θ) for Bayesian inference, and the pollutant concentration values ​​obtained from the preliminary survey data of the land to be restored are used as the observation data D.

[0067] Construct a Gaussian likelihood function P(D|θ) based on the observed data D;

[0068] Based on Bayes' theorem P(θ|D)∝P(D|θ)P(θ), a numerical solution is performed to calculate the posterior probability distribution field P(θ|D) representing the probability of pollutant concentration distribution in three-dimensional space.

[0069] The site to be remediated is discretized into a three-dimensional grid, and the model state vector θ is the set of pollutant concentration values ​​at all grid nodes. The prior distribution P(θ) is given from the sampling knowledge base, for example, a knowledge base with a specified mean vector. Covariance Matrix The multidimensional Gaussian distribution is used, where the covariance matrix represents the spatial correlation of pollutants in a specific geological medium. The observational data D obtained from the preliminary exploration will be used to construct the likelihood function P(D|θ), which typically assumes that the measurement error follows a normal distribution.

[0070] One implementation method is to use MCMC algorithms, such as Metropolis-Hastings, which generate a large set of θ samples conforming to the posterior distribution through tens of thousands of iterations. The posterior probability distribution field is then obtained by statistically analyzing the sample set. Another computationally more efficient implementation method is to use ensemble Kalman filtering, which maintains a set of N model state vectors. , , The set is composed to approximate the prior distribution. When new observation data D becomes available, each set member is calculated using Kalman gain. Updated to Where K is the Kalman gain, Let H be the observations with perturbations, and H be the observation operator. The updated set. , , This represents the posterior probability distribution field, the mean and variance of which can be obtained from set statistics, avoiding the lengthy convergence process required by the MCMC method, and is especially suitable for high-dimensional state spaces.

[0071] In an optional embodiment, the step of constructing a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media using historical site pollutant types, concentrations, spatial distributions, and geological and hydrological data stored in the sampling knowledge base includes:

[0072] The sampling knowledge base includes, but is not limited to: geological borehole data from similar historical sites, soil physicochemical properties, pollutant concentration monitoring data, hydrogeological parameters, and parameters of validated pollutant migration models.

[0073] The knowledge base is a structured geospatial database containing: geological borehole data (recording borehole coordinates and stratification information, such as "0-2.5m is miscellaneous fill, 2.5-6.0m is silty clay"); soil physicochemical properties (linked to specific soil layers, storing parameters such as permeability coefficient and porosity); pollutant concentration monitoring data (three-dimensional coordinates of historical sampling points, sampling time, pollutant name, and concentration value); hydrogeological parameters (regional groundwater flow field information); migration model parameters (diffusion rate, adsorption coefficient, etc., validated by historical projects); and historical land use and pollution source data (e.g., historical aerial photographs showing (250.5, 680.1)). "There used to be an underground oil storage tank at this location" is used to guide the setting of high-risk areas in the prior model; Earth exploration data, such as "resistivity tomography profiles show a low-resistivity anomaly that is distinctly different from the background at a depth of 5-8m", is used to construct the initial morphology of the prior geological structure and spatial distribution of pollutants. Among them, geological stratification and soil physicochemical property data can be used to construct the spatial covariance structure; Earth exploration data can be transformed into prior constraints on pollution distribution through geostatistical inversion, for example, as soft data for indicator kriging; Historical pollution source location information can be directly used as a covariate of the prior mean function, or used to set initial high-probability areas.

[0074] S2, Based on the posterior probability distribution field, calculate the three-dimensional pollution distribution uncertainty field of the entire plot; by integrating or averaging the three-dimensional uncertainty field along the depth direction, generate a two-dimensional uncertainty intensity plane map; based on the two-dimensional uncertainty intensity plane map, generate a set of non-uniformly distributed Voronoi diagram generation points, and spatially partition the plot to be repaired based on the generation points to generate a set of non-uniform sampling cells;

[0075] Extract the three-dimensional posterior variance field output from process S1, and use this posterior variance field as the three-dimensional pollution distribution uncertainty field. Use the mean or sum function from the NumPy library, setting the axis parameter to the depth dimension, to perform dimensionality reduction on the three-dimensional uncertainty field matrix, obtaining a two-dimensional uncertainty intensity plane map. After normalizing the two-dimensional intensity plane map, use it as the probability density function, and use the rejection sampling algorithm or the Metropolis-Hastings algorithm to generate a set of non-uniformly distributed generation points on the plane. Use the generation points as input, call the spatial.Voronoi function from the SciPy library to generate a Voronoi map of the land parcel plane, where each polygonal region of the map is a sampling cell.

[0076] In an optional embodiment, calculating the three-dimensional pollution distribution uncertainty field of the entire site based on the posterior probability distribution field includes:

[0077] For each three-dimensional spatial grid node in the posterior probability distribution field, extract the posterior probability distribution of pollutant concentration at the node;

[0078] Calculate the standard deviation of the posterior probability distribution and use the standard deviation value as the pollution distribution uncertainty value at this node;

[0079] By traversing all grid nodes, the three-dimensional uncertainty field of the entire land parcel, composed of the uncertainty values ​​of each node, is obtained.

[0080] For any node in the 3D mesh Having a set of samples representing the posterior probability distribution, such as the concentration values ​​of N set members. , , Calculate the standard deviation of the N sample values. This is a common method for representing uncertainty, indicating the absolute range of fluctuation in the predicted concentration. For example, if the standard deviation of one node is 25.8 mg / kg, much higher than that of another node (1.7 mg / kg), it indicates that the former has a higher prediction uncertainty. The coefficient of variation can also be calculated. ,in The index, representing the posterior mean, indicates relative uncertainty and facilitates fair comparisons across regions with different concentration levels. The standard deviation is chosen as the uncertainty value. By performing this calculation on all grid nodes, a three-dimensional uncertainty matrix U is obtained, where the high-value regions represent the spatial locations where the model's prediction of pollutant concentrations is most uncertain.

[0081] In an optional embodiment, generating a two-dimensional uncertainty intensity plane map by integrating or averaging the three-dimensional uncertainty field along the depth direction includes:

[0082] For each (x,y) position on the horizontal plane, extract the uncertainty profile U(x,y,z) in the vertical direction;

[0083] Calculate the sum, mean, or maximum value of the uncertainty values ​​at all depth points in the uncertainty profile, and use it as the two-dimensional uncertainty intensity value I(x,y) at the (x,y) location;

[0084] By traversing all horizontal positions, the two-dimensional uncertainty intensity plane diagram is generated.

[0085] The three-dimensional uncertainty field U is reduced in dimension along the depth z direction to generate a two-dimensional uncertainty intensity plane diagram I(x,y). One method is the integration method. This represents the total uncertainty on the vertical cross-section of the stated position. Another method is the maximum value method. This method is particularly sensitive to identifying locations containing thin, high-uncertainty regions, ensuring that even if the prediction uncertainty of a key contamination layer is high, that horizontal location will still be given priority.

[0086] In an optional embodiment, generating a set of non-uniformly distributed Voronoi diagram generation points based on the two-dimensional uncertainty intensity plane diagram includes:

[0087] The spatial density of the generated points is set to be proportional to the two-dimensional uncertainty intensity value of the location of the generated points;

[0088] Based on the density relationship, a set of non-uniformly distributed Voronoi diagram generation points are generated on the two-dimensional uncertainty intensity plane using the Poisson disk sampling algorithm.

[0089] Define a sampling radius function r(x,y) that is related to the uncertainty intensity I(x,y), for example ,in It is the intensity value normalized to the [0,1] interval. and This refers to the set minimum and maximum sampling point spacing. The Poisson disk sampling algorithm uses this variable radius to ensure accuracy in regions of high uncertainty. Approaching 1, r approaches The generation points are densely distributed; while in regions of low uncertainty, Approaching 0, r approaches The generated points are sparsely distributed.

[0090] S3, For each sampling cell, extract the probability profile of the posterior probability distribution field on the vertical axis passing through the Voronoi generation point; calculate the local extrema and the zero of the second derivative of the profile curve, and define the depths corresponding to the local extrema and the zero of the second derivative as a set of candidate key sampling depths within the cell.

[0091] For each sampled cell, based on the XY coordinates of the Voronoi generated point, a one-dimensional array along the Z-axis depth direction is indexed and extracted from the mean field of the three-dimensional posterior probability distribution field; this array is the probability profile. The gradient function of the NumPy library is used to numerically differentiate the one-dimensional array to calculate the first and second derivatives. The signal.find_peaks function of the SciPy library is called to find extreme points at the positive and negative changes of the first derivative, and all zeros, i.e., inflection points, of the second derivative are found by analyzing the sign changes of the second derivative. The depth Z values ​​corresponding to all the found extreme points and inflection points are summarized into a candidate key sampling depth set for that cell.

[0092] In an optional embodiment, the calculation of local extrema and zeros of the second derivative of the profile curve, and the definition of the depths corresponding to the local extrema and zeros of the second derivative as a set of candidate key sampling depths within the cell, includes:

[0093] The vertical probability profile curve is discretized into a series of depth-probability data point pairs. ;

[0094] Identify and satisfy and depth As a local maximum point, identify the condition that satisfies and depth As a local minimum point;

[0095] Calculating the second derivative using the central difference method and identify those that meet the requirements. Depth < 0 As the zero of the second derivative;

[0096] By summing up all identified depth points, a set of candidate key sampling depths for each cell is obtained.

[0097] For a given sampling cell, extract the posterior average concentration values ​​of each depth grid node on the vertical axis of the Voronoi generator point, with depth intervals. =0.5 meters, forming discrete profile data. To reduce the impact of noise, it is advisable to first... The sequence is smoothed using methods such as Savitzky-Golay. The smoothed profile is then traversed to identify local maxima and minima. The formula is then used... Calculate the second derivative. If it is found... and Different signs, then in and There exists a zero point of the second derivative, which marks the region where the pollutant concentration gradient changes most rapidly, typically corresponding to the vertical boundary of the pollutant plume. The depths corresponding to all identified maxima, minima, and zero points of the second derivative are aggregated to form the candidate key sampling depths for that cell.

[0098] In another embodiment, local minimum points are not used as depth collection points; that is, the soil layers corresponding to local minimum points are not collected. However, those skilled in the art should know that this can be adjusted according to actual needs.

[0099] S4, calculate the sampling priority score for each candidate key sampling depth; sort all candidate key sampling depths of all cells globally according to their priority scores; select several candidate key sampling depths with the highest priority scores from the sorted list according to the preset total number of sampling points, and determine the position of the sampling depth as the sampling point to obtain the hierarchical sampling scheme.

[0100] Define a sampling priority score for the j-th candidate key sampling depth point in the i-th cell. Through weighted sum formula Perform calculations, such as Figure 3 As shown, where This represents the normalized uncertainty intensity value of the cell containing that point. This represents the normalized mean posterior probability of a point at that depth. This is the normalized value of the absolute value of the first derivative of the posterior probability profile at that depth point. , , Set the preset non-negative weight coefficients; create a list containing all candidate key sampling depth points, their corresponding location coordinates, and priority scores; use Python's built-in `sorted` function or NumPy's `argsort` function to sort the samples according to their priority scores. The list is sorted in descending order globally; N points with a preset total number of sampling points are selected from the top of the sorted list, and the three-dimensional coordinates (XYZ) of these points constitute the sampling point set.

[0101] In an optional embodiment, calculating a sampling priority score for each candidate key sampling depth includes:

[0102] For the j-th candidate key sampling depth point in the i-th cell Priority score Calculated using the weighted formula: ;

[0103] in, This is the normalized value of the uncertainty intensity of cell i containing the candidate key sampling depth point. This is the normalized value of the mean posterior probability of the depth point. This is the normalized value of the absolute value of the first or second derivative of the depth point. , , These are the preset non-negative weighting coefficients.

[0104] When calculating the score, the overall uncertainty value is calculated for all cells i. For example, the two-dimensional uncertainty intensity value, obtained by performing minimum-maximum normalization. Collect the posterior mean concentration of all candidate depth points. and gradient information representing the cross-sectional shape And perform global normalization respectively to obtain and . It can be the absolute value of the first derivative. The normalized value, with high values ​​indicating the location at the pollution boundary; it can also be the absolute value of the second derivative. The normalized value is used, with high values ​​indicating peaks or valleys at the core of pollution. In an alternative embodiment, if it is an extreme point, it is scored based on the absolute magnitude of its mean concentration; if it is an inflection point, it is scored based on the magnitude of its first derivative. Alternatively, the score can be uniformly calculated using the absolute value of the first derivative.

[0105] The weighting coefficients are set according to the remediation target: if the target is to depict high-concentration areas, then set... Higher, such as =0.2, =0.6, =0.2; If the goal is to delineate the pollution area, then the focus should be on boundary points with large gradients, and the setting should be adjusted accordingly. Higher, such as =0.3, =0.2, =0.5. After calculating the scores of all candidate points, perform a global sort. If the budget allows for the collection of 100 samples, select the 100 points with the highest scores as sampling locations.

[0106] Example 2 provides a stratified sampling system for soil remediation, comprising the following modules:

[0107] The update module is used to construct a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media by utilizing historical pollutant types, concentrations, spatial distributions, and geological and hydrological data stored in the sampling knowledge base; combined with the preliminary survey data of the site to be remediated, the prior probability distribution model is updated by a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated.

[0108] The generation module is used to calculate the three-dimensional pollution distribution uncertainty field of the entire plot based on the posterior probability distribution field; generate a two-dimensional uncertainty intensity plane map by integrating or averaging the three-dimensional uncertainty field along the depth direction; generate a set of non-uniformly distributed Voronoi diagram generation points based on the two-dimensional uncertainty intensity plane map, and spatially partition the plot to be repaired based on the generation points to generate a set of non-uniform sampling cells.

[0109] The calculation module is used to extract the probability profile of the posterior probability distribution field on the vertical axis passing through the Voronoi generation point for each sampling cell; calculate the local extrema and the zero of the second derivative of the profile curve, and define the depths corresponding to the local extrema and the zero of the second derivative as a set of candidate key sampling depths within the cell.

[0110] The determination module is used to calculate the sampling priority score for each candidate key sampling depth; to globally sort all candidate key sampling depths of all cells according to the priority score; and to select several candidate key sampling depths with the highest priority scores from the sorted list according to the preset total number of sampling points, and to determine the position of the sampling depth as the sampling point to obtain the hierarchical sampling scheme.

[0111] In this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise limited, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this document, "a," "an," "the," "the," and "its" may also include plural forms unless the context clearly indicates otherwise. "Multiple" refers to at least two, such as 2, 3, 5, or 8, etc. "And / or" includes any and all combinations of the associated listed items.

[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0113] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A stratified sampling method for soil remediation, characterized in that, Includes the following steps: Using historical data on pollutant types, concentrations, spatial distributions, and geological and hydrological conditions stored in the sampling knowledge base, a prior probability distribution model representing the vertical migration patterns of pollutants in different soil media is constructed. Combined with preliminary survey data of the site to be remediated, the prior probability distribution model is updated using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants for the site to be remediated. Based on the posterior probability distribution field, the three-dimensional pollution distribution uncertainty field of the entire site is calculated; by integrating or averaging the three-dimensional uncertainty field along the depth direction, a two-dimensional uncertainty intensity plane map is generated; based on the two-dimensional uncertainty intensity plane map, a set of non-uniformly distributed Voronoi diagram generation points are generated, and the site to be repaired is spatially partitioned based on the generation points to generate a set of non-uniform sampling cells. For each sampling cell, extract the probability profile of the posterior probability distribution field on the vertical axis passing through the Voronoi generation point; calculate the local extrema and the zero of the second derivative of the profile curve, and define the depths corresponding to the local extrema and the zero of the second derivative as a set of candidate key sampling depths within the cell. Calculate a sampling priority score for each candidate key sampling depth; globally sort all candidate key sampling depths of all cells according to their priority scores; select several candidate key sampling depths with the highest priority scores from the sorted list according to the preset total number of sampling points, and determine the position of the sampling depth as the sampling point to obtain a hierarchical sampling scheme.

2. The method according to claim 1, characterized in that, The step of updating the prior probability distribution model using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated includes: The prior probability distribution model is used as the prior distribution P(θ) for Bayesian inference, and the pollutant concentration values ​​obtained from the preliminary survey data of the land to be restored are used as the observation data D. Construct a Gaussian likelihood function P(D|θ) based on the observed data D; Based on Bayes' theorem P(θ|D)∝P(D|θ)P(θ), a numerical solution is performed to calculate the posterior probability distribution field P(θ|D) representing the probability of pollutant concentration distribution in three-dimensional space.

3. The method according to claim 2, characterized in that, The calculation of the three-dimensional pollution distribution uncertainty field for the entire site based on the posterior probability distribution field includes: For each three-dimensional spatial grid node in the posterior probability distribution field, extract the posterior probability distribution of pollutant concentration at the node; Calculate the standard deviation of the posterior probability distribution and use the standard deviation value as the pollution distribution uncertainty value at this node; By traversing all grid nodes, the three-dimensional uncertainty field of the entire land parcel, composed of the uncertainty values ​​of each node, is obtained.

4. The method according to claim 1, characterized in that, The step of generating a set of non-uniformly distributed Voronoi diagrams based on the two-dimensional uncertainty intensity plane diagram includes: The spatial density of the generated points is set to be proportional to the two-dimensional uncertainty intensity value of the location of the generated points; Based on the density relationship, a set of non-uniformly distributed Voronoi diagram generation points are generated on the two-dimensional uncertainty intensity plane using the Poisson disk sampling algorithm.

5. The method according to claim 1, characterized in that, The calculation of the local extrema and the zeros of the second derivative of the profile curve, and the definition of the depths corresponding to the local extrema and the zeros of the second derivative as a set of candidate key sampling depths within the cell, includes: The vertical probability profile curve is discretized into a series of depth-probability data point pairs. ; Identify and satisfy and depth As a local maximum point, identify the condition that satisfies and depth As a local minimum point; Calculating the second derivative using the central difference method and identify those that meet the requirements. Depth < 0 As the zero of the second derivative; By summing up all identified depth points, a set of candidate key sampling depths for each cell is obtained.

6. The method according to claim 1, characterized in that, The calculation of sampling priority score for each candidate key sampling depth includes: For the j-th candidate key sampling depth point in the i-th cell Priority score Calculated using the weighted formula: ; in, This is the normalized value of the uncertainty intensity of cell i containing the candidate key sampling depth point. This is the normalized value of the mean posterior probability of the depth point. This is the normalized value of the absolute value of the first or second derivative of the depth point. , , These are the preset non-negative weighting coefficients.

7. The method according to claim 5, characterized in that, The method utilizes historical data on pollutant types, concentrations, spatial distributions, and geological and hydrological conditions stored in a sampling knowledge base to construct a prior probability distribution model representing the vertical migration patterns of pollutants in different soil media, including: The sampling knowledge base includes, but is not limited to: geological borehole data from similar historical sites, soil physicochemical properties, pollutant concentration monitoring data, hydrogeological parameters, and parameters of validated pollutant migration models.

8. A stratified sampling system for soil remediation, characterized in that, Includes the following modules: The update module is used to construct a prior probability distribution model representing the vertical migration pattern of pollutants in different soil media by utilizing historical pollutant types, concentrations, spatial distributions, and geological and hydrological data stored in the sampling knowledge base; combined with the preliminary survey data of the site to be remediated, the prior probability distribution model is updated by a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated. The generation module is used to calculate the three-dimensional pollution distribution uncertainty field of the entire plot based on the posterior probability distribution field; generate a two-dimensional uncertainty intensity plane map by integrating or averaging the three-dimensional uncertainty field along the depth direction; generate a set of non-uniformly distributed Voronoi diagram generation points based on the two-dimensional uncertainty intensity plane map, and spatially partition the plot to be repaired based on the generation points to generate a set of non-uniform sampling cells. The calculation module is used to extract the probability profile of the posterior probability distribution field on the vertical axis passing through the Voronoi generation point for each sampling cell; calculate the local extrema and the zero of the second derivative of the profile curve, and define the depths corresponding to the local extrema and the zero of the second derivative as a set of candidate key sampling depths within the cell. The determination module is used to calculate the sampling priority score for each candidate key sampling depth; All candidate key sampling depths of all cells are globally sorted according to priority scores; based on the preset total number of sampling points, several candidate key sampling depths with the highest priority scores are selected from the sorted list, and the positions of the sampling depths are determined as sampling points to obtain a hierarchical sampling scheme.

9. The system according to claim 8, characterized in that, The step of updating the prior probability distribution model using a Bayesian fusion method to obtain the three-dimensional posterior probability distribution field of pollutants in the site to be remediated includes: The prior probability distribution model is used as the prior distribution P(θ) for Bayesian inference, and the pollutant concentration values ​​obtained from the preliminary survey data of the land to be restored are used as the observation data D. Construct a Gaussian likelihood function P(D|θ) based on the observed data D; Based on Bayes' theorem P(θ|D)∝P(D|θ)P(θ), a numerical solution is performed to calculate the posterior probability distribution field P(θ|D) representing the probability of pollutant concentration distribution in three-dimensional space.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-7.