A method and system for dynamic risk zoning of soil heavy metal pollution
Patent Information
- Application Number
- CN202610682231.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]而目前传统的区划方法依赖静态插值与固定阈值分级,所获区划精细度低,难以精准表示重金属污染的区域异质性,忽视了电磁感应与重金属污染之间的非线性耦合关系及反演过程的动态更新能力
[0045] This scheme generates Halton low-difference sequences and combines them with batch dynamic measurements on a mobile platform to obtain spatially uniformly distributed total soil measurement points. By embedding a forward mapping model with residual constraints of Maxwell's equations and introducing a Bayesian compressed sensing model, it achieves high-precision nonlinear mapping from electromagnetic induction data to heavy metal concentration, dynamic iterative updates of the inversion process, and effective modeling of the nonlinear coupling relationship between electromagnetic induction and heavy metal pollution. Through the inner and outer loops of lower-level and upper-level agents, it achieves coordinated optimization of local adjustment of superpixel risk levels and global measurement decisions, balancing measurement batches and zoning accuracy as needed, thereby improving the accuracy and adaptability of soil heavy metal pollution risk zoning.
Smart Images

Figure CN122595019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of risk zoning technology, and more specifically, to a method and system for dynamic risk zoning of heavy metal pollution in soil. Background Technology
[0002] Soil heavy metal pollution risk zoning relies on spatial modeling and hierarchical zoning of the soil heavy metal concentration measurement area. It is widely used in scenarios such as farmland, industrial sites, and mining areas, aiming to achieve accurate division and dynamic adjustment of the pollution range. Taking the heavy metal pollution risk zoning of industrial sites as an example, the distribution of pollution sources within the site often has characteristics such as locality and concealment. Heavy metal pollution may not reach the detection threshold, but it can still cause abnormal changes in electromagnetic induction multi-frequency response data.
[0003] Current traditional zoning methods rely on static interpolation and fixed threshold grading, resulting in low zoning precision. They are unable to accurately represent the regional heterogeneity of heavy metal pollution and neglect the nonlinear coupling relationship between electromagnetic induction and heavy metal pollution, as well as the dynamic updating capability of the inversion process. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for dynamic risk zoning of soil heavy metal pollution, the method comprising:
[0005] A physical information neural network is constructed based on the original dataset, and residual constraint terms of Maxwell's equations are embedded to obtain an electromagnetic forward modeling model.
[0006] After obtaining the total number of soil measurement points, the total number of soil measurement points are measured in real time in batches based on the mobile platform to obtain the electromagnetic raw dataset.
[0007] The original electromagnetic dataset is imported into the electromagnetic forward model to obtain the heavy metal concentration estimation set. At the same time, the concentration posterior distribution is generated after inversion by combining the Bayesian compressed sensing model. The risk level probability field is obtained based on the concentration posterior distribution, and a graph structure is constructed.
[0008] The graph structure is input into the lower-level agent to obtain a provisional risk zoning scheme and global state features. The upper-level agent performs binary decision-making based on the global state features and outputs the risk zoning scheme and global probability field. The lower-level agent executes based on the inner loop, and the upper-level agent executes based on the outer loop.
[0009] Pollution risk zoning maps are obtained based on risk zoning schemes and global probability fields.
[0010] Specifically, a physical information neural network is constructed based on the original dataset, and residual constraint terms of Maxwell's equations are embedded to obtain an electromagnetic forward modeling model, including:
[0011] Ground calibration points were set up to obtain a raw dataset that includes at least heavy metal concentration data, electromagnetic induction data, soil moisture data, and temperature data. The raw dataset was then preprocessed and normalized.
[0012] A physical information neural network consisting of an input layer, a hidden layer, and an output layer is constructed. The residual constraint term of Maxwell's equations is embedded in the loss function as a physical constraint. The Adam optimizer is used to minimize the loss function. The physical information neural network is trained with the original dataset to obtain an electromagnetic forward mapping model with the heavy metal concentration estimation set as the output target.
[0013] Specifically, after obtaining the total number of soil measurement points, real-time measurements are performed on the total number of soil measurement points in batches using a mobile platform to obtain the raw electromagnetic dataset, including:
[0014] The total number of measurement points is determined based on the area of the measurement area and the preset measurement density. Two-dimensional coordinate points covering the measurement area are generated based on the Halton low difference sequence as the total soil measurement points. The total soil measurement points are divided into multiple batches in the order of generation and the measurement path is determined.
[0015] Based on the measurement path, a mobile platform equipped with an electromagnetic signal measuring instrument, a soil moisture sensor, and an infrared temperature sensor is controlled in batches to measure and acquire electromagnetic raw datasets, which include electromagnetic induction data, soil moisture data, temperature data, and GPS coordinates.
[0016] Once the first batch of raw electromagnetic datasets is acquired, the measurement is paused and the system waits for a decision from the upper-level agent. If the upper-level agent decides to continue the measurement, the system continues to acquire raw electromagnetic datasets until the upper-level agent decides to end the zoning process.
[0017] Specifically, the raw electromagnetic dataset is imported into an electromagnetic forward mapping model to obtain an estimated set of heavy metal concentrations. Simultaneously, a Bayesian compressed sensing model is used for inversion to generate a posterior concentration distribution. Based on this posterior distribution, a risk level probability field is obtained, and a graph structure is constructed, including:
[0018] Electromagnetic induction data, soil moisture data, and temperature data from the original electromagnetic dataset are input into the electromagnetic forward model to obtain a set of heavy metal concentration estimates.
[0019] Based on the GPS coordinates and heavy metal concentration estimation set in the original electromagnetic dataset, the concentration posterior distribution is obtained by inversion through a Bayesian compressed sensing model. The concentration posterior distribution includes the posterior mean, which characterizes the best estimate of heavy metal concentration, and the posterior variance, which characterizes the uncertainty.
[0020] The measurement area is discretized into original grids. Based on the set risk screening value and risk control value, and the probability of each original grid belonging to each risk level is calculated based on the concentration posterior distribution to obtain the risk level probability field.
[0021] The original grid in the risk level probability field is aggregated into multiple superpixels according to spatial adjacency. A graph structure is constructed with each superpixel as a graph node and spatial adjacency as an edge.
[0022] Specifically, for Bayesian compressed sensing models, including:
[0023] The GPS coordinates in the original electromagnetic dataset were used as the observation locations, and the corresponding heavy metal concentration estimates were used as the observation data.
[0024] The observation location and observation data are input into the Bayesian compressed sensing model, which uses a sparse variational Gaussian process as the inference framework and is configured with an anisotropic squared exponential kernel function with initial hyperparameters.
[0025] After receiving the input, the Bayesian compressed sensing model uses the anisotropic squared exponential kernel function to represent the spatial correlation prior and obtains the posterior probability distribution through the sparse variational inference algorithm.
[0026] In this process, after each inversion of the Bayesian compressed sensing model, the hyperparameters of the anisotropic squared exponential kernel function are updated using the expectation-maximization algorithm. The hyperparameters include spatial correlation length and signal variance.
[0027] Specifically, a graph structure is input into the lower-level agent to obtain a provisional risk zoning scheme and global state features. The upper-level agent performs a binary decision based on the global state features and outputs the risk zoning scheme and global probability field. The lower-level agent executes based on an inner loop, and the upper-level agent executes based on an outer loop, including:
[0028] The risk level of each superpixel in the graph structure is obtained based on the risk level probability field.
[0029] The average value of the risk level probability vector is calculated as the risk composition feature, the variance of the risk level probability is calculated as the uncertainty feature, and the density of all measured original grids is statistically analyzed as the measurement density feature to obtain the risk feature set.
[0030] The lower-level intelligent agent selects and executes actions through the graph structure until the inner loop ends, in order to obtain a provisional risk zoning plan and lower-level confidence indicators.
[0031] By combining the risk feature set with the lower-level confidence index, the global state features are obtained. The upper-level intelligent agent executes binary decision-making based on the global state features until the outer loop ends, in order to output the risk zoning scheme.
[0032] Specifically, for lower-level intelligent agents and upper-level intelligent agents, including:
[0033] The lower-level intelligent agent consists of a graph attention encoder, a temporal convolutional encoder, a policy network based on a flexible actor-critic algorithm, and a value network based on a linear weighted multi-objective reward function. The spatial feature vector and temporal context vector output by the graph attention encoder and the temporal convolutional encoder are used as the state space, and the operations of adjusting, maintaining or lowering the risk level and ending the inner loop are used as the action space.
[0034] The upper-layer intelligent agent consists of a policy network based on a flexible actor-critic algorithm and a value network based on a linearly weighted multi-objective reward function, with global state features as the state space and continued measurement and termination of the division as the action space.
[0035] Specifically, the lower-level agents select and execute actions through a graph structure until the inner loop ends, in order to obtain a provisional risk zoning plan and lower-level confidence indicators, including:
[0036] The graph structure and historical action sequence are used as inputs to the graph attention encoder and the temporal convolutional encoder, respectively, to obtain spatial feature vectors and temporal context vectors, which are then concatenated to form the state space. The state space is used as input to the policy network to obtain the probability of up-adjusting, maintaining, or down-adjusting the risk level of each superpixel in the graph structure and to execute the action corresponding to the highest probability. At the same time, the immediate reward is calculated and the parameters of the lower network are optimized.
[0037] The lower-level agent outputs the end of the inner loop when the risk level of all superpixels no longer changes or the risk level of all superpixels reaches the threshold. It outputs the current zoning scheme as a provisional risk zoning scheme and the average action value of the last inner loop as the lower-level confidence index.
[0038] Specifically, by combining the risk feature set with lower-level confidence indicators to obtain global state features, the upper-level agent executes binary decisions based on these global state features until the outer loop ends, outputting a risk zoning scheme, including:
[0039] Based on the risk feature set, the average value of the uncertainty features is taken and concatenated with the risk composition features, measurement density features and lower-level confidence indicators to obtain the global state features.
[0040] The upper-level agent takes global state features as input, makes binary decisions based on the policy network, including continuing measurement and ending the division, and updates the upper-level network parameters based on the immediate reward after each decision.
[0041] When the upper-level agent outputs a decision to continue measurement, it returns to start the next batch of measurements and performs subsequent processing;
[0042] When the upper-level agent outputs the decision to end the zoning, it outputs the current provisional risk zoning scheme as the risk zoning scheme.
[0043] Secondly, embodiments of the present invention also provide a dynamic risk zoning system for soil heavy metal pollution, comprising: the dynamic risk zoning system for soil heavy metal pollution includes a processor and a memory, the memory and the processor are connected, the memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the above-mentioned dynamic risk zoning method for soil heavy metal pollution.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] This scheme generates Halton low-difference sequences and combines them with batch dynamic measurements on a mobile platform to obtain spatially uniformly distributed total soil measurement points. By embedding a forward mapping model with residual constraints of Maxwell's equations and introducing a Bayesian compressed sensing model, it achieves high-precision nonlinear mapping from electromagnetic induction data to heavy metal concentration, dynamic iterative updates of the inversion process, and effective modeling of the nonlinear coupling relationship between electromagnetic induction and heavy metal pollution. Through the inner and outer loops of lower-level and upper-level agents, it achieves coordinated optimization of local adjustment of superpixel risk levels and global measurement decisions, balancing measurement batches and zoning accuracy as needed, thereby improving the accuracy and adaptability of soil heavy metal pollution risk zoning. Attached Figure Description
[0046] Figure 1 This is a flowchart illustrating the steps of a method for dynamic risk zoning of heavy metal pollution in soil according to the present invention.
[0047] Figure 2 This is a flowchart of a method for dynamic risk zoning of heavy metal pollution in soil according to the present invention.
[0048] Figure 3 This is a schematic diagram of a pollution zoning map in the dynamic risk zoning method for heavy metal pollution in soil according to the present invention. Detailed Implementation
[0049] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating the steps of a dynamic risk zoning method for heavy metal pollution in soil according to the present invention. Figure 2 This is a flowchart of a method for dynamic risk zoning of heavy metal pollution in soil according to the present invention. Figure 3 This is a schematic diagram of a pollution zoning map in a dynamic risk zoning method for soil heavy metal pollution according to the present invention. The following is a detailed introduction to this dynamic risk zoning method for soil heavy metal pollution.
[0050] Specifically, a method for dynamic risk zoning of soil heavy metal pollution includes the following steps:
[0051] Step S1: Construct a physical information neural network based on the original dataset and embed residual constraint terms of Maxwell's equations to obtain an electromagnetic forward model.
[0052] In this embodiment, step S1 specifically includes steps S1-1 to S1-2:
[0053] Step S1-1: Deploy ground calibration points to obtain a raw dataset that includes at least heavy metal concentration data, electromagnetic induction data, soil moisture data, and temperature data. Preprocess and normalize the raw dataset.
[0054] Specifically, based on the soil type distribution map, historical data, and suspected pollution range of the measurement area, a small number of ground calibration points, such as 5-10, are set up in the measurement area. The locations of the ground calibration points should cover different soil types and land use types.
[0055] In practical implementation, for example, within a composite measurement area of farmland and forestland, five ground calibration points need to be set up inside the farmland, inside the forest, at the boundary between farmland and forest, at the boundary between farmland and pollution sources such as roads, and at the boundary between forest and pollution sources such as roads.
[0056] Heavy metal concentration detectors, electromagnetic signal measuring instruments, soil moisture sensors, and temperature sensors were deployed at the five ground calibration points mentioned above to acquire heavy metal concentration data, electromagnetic induction data, soil moisture data, and temperature data, respectively. The data were then integrated to generate a raw dataset, which was then cleaned, time-aligned, and normalized.
[0057] The heavy metal concentration data refers to the concentration values of heavy metals, such as lead concentration and other heavy metal elements. The electromagnetic induction data refers to the in-phase component reflecting soil conductivity and the quadrature component reflecting soil magnetic susceptibility measured under electromagnetic fields of different frequencies. For example, the in-phase component measured at a frequency of 10 kHz is 0.523 mV and the quadrature component is 0.187 mV. The soil moisture data and temperature data refer to soil volumetric water content and temperature.
[0058] Step S1-2: Construct a physical information neural network consisting of an input layer, a hidden layer, and an output layer. Embed the residual constraint term of Maxwell's equations as a physical constraint in the loss function. Simultaneously, use the Adam optimizer to minimize the loss function. Train the physical information neural network with the original dataset to obtain an electromagnetic forward mapping model with the heavy metal concentration estimation set as the output target.
[0059] The input layer takes electromagnetic induction data, soil moisture data, and temperature data from the original electromagnetic dataset as input.
[0060] The hidden layer is used to learn complex nonlinear mappings to extract the deep relationship between the input and the heavy metal concentration.
[0061] The output of the output layer is the estimated heavy metal concentration at the corresponding soil measurement point.
[0062] The loss function consists of a data fitting term and a physical loss, and is minimized by the Adam optimizer. The physical loss is expressed as a physical constraint using the residual constraint term of Maxwell's equations. The residual constraint term of Maxwell's equations enables the output of the physical information neural network to satisfy the propagation law of electromagnetic field in underground medium, and at the same time enables the electromagnetic forward modeling model to make reasonable inferences based on physical consistency in areas lacking training data, thereby improving the extrapolation ability, generalization and interpretability of the model in sparse data areas. The data fitting term is expressed as the mean square error between the heavy metal concentration estimate output by the electromagnetic forward modeling model and the measured concentration value of the corresponding ground calibration point in the original dataset.
[0063] For training the physical information neural network, the original dataset is randomly divided into a training set for driving parameter updates and a validation set for evaluating extrapolation and generalization capabilities according to a preset ratio. The training set is then divided into different batches. The first batch of training sets is then input into the physical information neural network. After calculating the total loss, the Adam optimizer is used to optimize the network along the gradient descent direction. Subsequently, the entire validation set is input into the physical information neural network, and the total loss on the validation set is calculated as the validation loss.
[0064] If the validation loss continues to decrease, it indicates that the generalization ability of the physical information neural network is improving. Continue to input the next batch of training set and repeat the above steps until the validation loss no longer decreases in multiple consecutive batches. At this point, the physical information neural network is used as the electromagnetic forward modeling model.
[0065] Step S2: After obtaining the total soil measurement points, perform real-time measurements on the total soil measurement points in batches using a mobile platform to obtain the electromagnetic raw dataset.
[0066] In this embodiment, step S2 specifically includes steps S2-1 to S2-3:
[0067] Step S2-1: Determine the total number of measurement points based on the area of the measurement area and the preset measurement density. Generate two-dimensional coordinate points covering the measurement area based on the Halton low difference sequence as the total soil measurement points. Divide the total soil measurement points into multiple batches in the order of generation and determine the measurement path.
[0068] The measurement area is determined and its boundary coordinates and the minimum bounding rectangle boundary coordinates are obtained. The minimum bounding rectangle of the measurement area can be obtained through the boundary coordinates and the minimum bounding rectangle boundary coordinates. The total number of measurement points is determined by the area of the measurement area and the preset measurement density. The measurement density is set based on the complexity of the pollution in the area. If the pollution distribution is highly heterogeneous (such as point source pollution or strip distribution), a higher density (80-100 points / hectare) is required to capture details. If the pollution is relatively uniform (such as non-point source pollution), a lower density (20-30 points / hectare) can be used. The total number of measurement points = measurement area × measurement density.
[0069] Two-dimensional coordinate points with spatial uniformity and minimum bounding rectangle covering the measurement area are generated based on the Halton low-difference sequence. Two-dimensional coordinate points outside the measurement area are removed by geographic information algorithm, and two-dimensional coordinate points covering the measurement area are retained. Two-dimensional coordinate points are generated again based on the Halton low-difference sequence until the number of two-dimensional coordinate points in the measurement area equals the total number of measurement points. At this point, linear scaling and translation are performed to obtain a total number of soil measurement points with spatial uniformity. The total number of soil measurement points is divided into multiple batches according to the generation order and preset batches, and the measurement path is obtained according to the generation order of the two-dimensional coordinate points. The preset batches are determined based on the total number of soil measurement points.
[0070] For example, for a polygonal suburban plot with an area of 1 hectare, the minimum bounding rectangle boundary is first determined. Based on the preset measurement density of 50 points / hectare, the total number of measurement points is calculated to be 50, and they are divided into 5 batches for measurement. Then, two-dimensional coordinate points are generated using prime numbers 2 and 3 as bases and based on the Halton low-dispersion sequence. Simultaneously, they are linearly mapped to the range of the minimum bounding rectangle. Two-dimensional coordinate points falling within the polygonal suburban plot are selected through geographic information algorithms. After supplementation, 50 two-dimensional coordinate points are obtained as the total soil measurement points. Finally, the soil measurement points are divided into 5 batches of 10 soil measurement points each according to the generation order of the soil measurement points. The measurement order of the 10 soil measurement points in each batch is determined according to the generation order of the two-dimensional coordinate points. The measurement path of the mobile platform is planned based on the measurement order.
[0071] The proposed measurement method uses a mobile platform to perform batch measurements and leverages the low variance and determinism of Halton low-difference sequences. Compared to traditional measurement methods, it can automatically pause after each batch of measurements is completed and dynamically decide whether to start the next batch of measurements based on the decision instructions of the upper-level intelligent agent. This allows for the acquisition of a more comprehensive raw electromagnetic dataset, enabling adaptive stopping of the measurement process and on-demand control of measurement costs. Simultaneously, it achieves more uniform spatial coverage with fewer sampling points, thereby solving the problem of uneven sampling and improving the mapping accuracy of the electromagnetic forward model.
[0072] Step S2-2: Based on the measurement path, control the mobile platform equipped with an electromagnetic signal measuring instrument, soil moisture sensor and infrared temperature sensor in batches to measure and obtain the electromagnetic raw dataset, which includes electromagnetic induction data, soil moisture data, temperature data and GPS coordinates.
[0073] Electromagnetic signal measuring instruments, soil moisture sensors, and infrared temperature sensors are deployed on mobile platforms such as all-terrain vehicles. The system moves sequentially to the soil measurement point according to the measurement path and, after stabilization, transmits all preset frequencies in sequence, such as 0.5kHz to 40kHz, including at least 8 different frequencies. The in-phase and quadrature components at each frequency are recorded to obtain electromagnetic induction data. Simultaneously, soil moisture data is obtained through the soil moisture sensor, and surface temperature data is obtained through the infrared temperature sensor. The data is then correlated with the GPS coordinates of the current soil measurement point. The electromagnetic induction data, soil moisture data, temperature data, and GPS coordinates are integrated into the first batch of raw electromagnetic datasets, and the measurement is paused. After transmitting the raw electromagnetic datasets back, the system awaits decision-making from the upper-level intelligent agent.
[0074] Step S2-3: After the first batch of electromagnetic raw datasets is acquired, the measurement is paused and the upper-level agent is waited for a decision. If the upper-level agent outputs a decision to continue the measurement, the electromagnetic raw dataset is acquired again until the upper-level agent outputs a decision to end the zoning.
[0075] After acquiring the first batch of raw electromagnetic datasets, the measurement is paused and the system waits for a decision from the upper-level agent. When the upper-level agent outputs a decision to continue the measurement, the mobile platform moves to the next batch of soil measurement points according to the measurement path and performs the measurement. After acquiring the next batch of raw electromagnetic datasets, the measurement is paused again. After transmitting the raw electromagnetic datasets back, the system waits for a decision from the upper-level agent.
[0076] The measurement ends when the mobile platform completes all preset batch measurements according to the measurement path or when the upper-level agent outputs a decision to end the zoning.
[0077] Step S3: Import the original electromagnetic dataset into the electromagnetic forward mapping model to obtain the heavy metal concentration estimation set. At the same time, combine the Bayesian compressed sensing model to generate the concentration posterior distribution after inversion. Based on the concentration posterior distribution, obtain the risk level probability field and construct the graph structure.
[0078] In this embodiment, step S3 specifically includes steps S3-1 to S3-4:
[0079] Step S3-1: Input the electromagnetic induction data, soil moisture data and temperature data from the original electromagnetic dataset into the electromagnetic forward modeling model to obtain the heavy metal concentration estimation set.
[0080] Based on the original electromagnetic dataset, electromagnetic induction data, soil moisture data, and temperature data of each soil measurement point are extracted, sequentially concatenated, and input into the electromagnetic forward model to obtain the heavy metal concentration estimates of all soil measurement points. The heavy metal concentration estimates are composed of all the heavy metal concentration estimates, which are approximate estimates of the heavy metal element concentration at the corresponding soil measurement point.
[0081] Step S3-2: Based on the GPS coordinates and heavy metal concentration estimation set in the original electromagnetic dataset, inversion is performed using a Bayesian compressed sensing model to obtain the posterior distribution of concentration, where the posterior distribution of concentration includes the posterior mean used to characterize the best estimate of heavy metal concentration and the posterior variance used to characterize uncertainty.
[0082] Specifically, the Bayesian compressed sensing model includes:
[0083] A Bayesian compressed sensing model is constructed, which uses a sparse variational Gaussian process as the inference framework and is configured with an anisotropic squared exponential kernel function with initial hyperparameters. The sparse variational Gaussian process can reduce the order of the complete Gaussian process, represent the entire region with a small number of induced points, and approximate the real Gaussian process through mathematical optimization.
[0084] Among them, hyperparameters include spatial correlation length. and and signal variance The initial hyperparameters were set based on the soil type distribution map, historical data, and suspected pollution range of the measurement area.
[0085] Specifically, Typically, the length is taken as 1 / 3 to 1 / 2 of the spindle length. Typically, half of the soil type variation zone is taken, and less than... , Typically, the square of half the concentration variation range is taken.
[0086] For example, in a 10-hectare farmland area, there is a seasonal river in the northwest-southeast direction. Historical monitoring shows that the heavy metal Cd pollution mainly comes from upstream mine drainage. The pollution band along the river is about 200m long and decays rapidly perpendicular to the river. The soil type change zone is about 50m long. The historical lowest Cd concentration was 0.1mg / kg and the historical highest Cd concentration was 1.5mg / kg.
[0087] From this we can obtain , , .
[0088] Specifically, we first define the Gaussian process prior, and then obtain the covariance from the anisotropic squared exponential kernel function:
[0089] ;
[0090] in and These are represented as the coordinates of two points in space. Represented as signal variance, it is used to control the overall variation of the function output. Let them be the length scales in the x and y directions, respectively, and let them be... , As a hyperparameter.
[0091] Then, M induction points are introduced, and the function values at the induction points are denoted as vectors. A Gaussian variational distribution is introduced for the function values at the induced points:
[0092] ;
[0093] Where m∈R M Let S ∈ R be the variational mean. M×M Let be the variational covariance matrix, and the Gaussian variational distribution be used. Approximate True Posterior .
[0094] Next, we construct a lower bound on the evidence as the optimization objective. The lower bound on the evidence is expressed as:
[0095] ;
[0096] in Through The predicted distribution is calculated using the conditional Gaussian distribution. For divergence, This indicates that, given a latent function value At that time, the observed value The conditional probability density function is obtained, and the lower bound of evidence is maximized by a maximization algorithm to simultaneously optimize the variational parameters (m, S) and hyperparameters. .
[0097] Finally, the posterior mean and posterior variance are obtained, and expressed as follows:
[0098] ;
[0099] ;
[0100] in It is the covariance vector between the predicted point and the induced point. It is the covariance matrix between induced points. For signal variance, The posterior mean is... This represents the posterior variance.
[0101] Using GPS coordinates from the original electromagnetic dataset as observation locations and corresponding heavy metal concentration estimates as observation data, the observation locations and observation data are input into a Bayesian compressed sensing model. After receiving the input, the Bayesian compressed sensing model uses the anisotropic squared exponential kernel function to characterize the spatial correlation prior, and obtains the posterior mean to characterize the best estimate of heavy metal concentration and the posterior variance to characterize the uncertainty through a sparse variational inference algorithm, thereby obtaining the posterior concentration distribution.
[0102] In this process, after each inversion of the Bayesian compressed sensing model, the hyperparameters of the anisotropic squared exponential kernel function are updated using the expectation-maximization algorithm to automatically adjust the spatial correlation, enabling the Bayesian compressed sensing model to better fit the anisotropy and spatial scale of the actual pollution distribution, thereby improving the inversion accuracy and zoning reliability.
[0103] In the traditional method of inverting electromagnetic induction data into heavy metal concentration values, there are problems such as the measurement data dimension being much smaller than the number of grid points to be determined and the lack of spatial prior constraints. However, the Bayesian compressed sensing model introduces a sparse variational Gaussian process and transforms the inversion problem into Bayesian inference. At the same time, it uses the anisotropic squared exponential kernel function to characterize the spatial correlation prior, which can accurately quantify the posterior distribution of concentration through the heavy metal concentration estimate.
[0104] Step S3-3: Discretize the measurement area into original grids, and calculate the probability of each original grid belonging to each risk level based on the set risk screening value and risk control value and the concentration posterior distribution to obtain the risk level probability field.
[0105] The grid side length is set according to the measurement accuracy requirements to determine the resolution (e.g., 10m×10m). The side length of the smallest bounding rectangle of the measurement area is divided by the grid side length, and the result is rounded up to obtain the number of columns and rows. The total number of grids = number of rows × number of columns. The lower left corner of the smallest bounding rectangle is used as the origin. The coordinates of the lower left corner and the center point of each grid are generated at a fixed step size. A unique index is assigned to each grid based on the GPS coordinates in the original electromagnetic dataset to obtain the original grid.
[0106] Risk screening and control values are set based on GB15618-2018 and GB36600-2018. Multiple risk levels are obtained according to the classification requirements. The probability of each original grid cell belonging to each risk level is calculated using the posterior distribution of concentration. The calculation formula is as follows:
[0107] ;
[0108] in The cumulative distribution function is expressed as a standard normal distribution. This represents the heavy metal concentration threshold corresponding to level k. and These are represented as the posterior mean and posterior standard deviation, respectively. After obtaining the risk level probability vector for each original grid, the risk level probability field is obtained.
[0109] For example, in a specific embodiment, based on the risk screening value T1 and the risk control value T2, it is necessary to obtain the probability of the first, second, and third risk levels of each grid in the target area to obtain the risk level probability field.
[0110] First, for each grid j, based on the posterior mean μ... j and posterior standard deviation σ j ;
[0111] When the concentration of heavy metal pollution When the risk level is low, it indicates that protection should be prioritized, and this is classified as Level 1 risk (Q1). ;
[0112] when Heavy metal pollution concentration It should be used safely at this time, and is classified as a level 2 risk (Q2). ;
[0113] A high concentration of heavy metal pollution indicates a high health risk and requires strict control; in this case, it is classified as risk level Q3 (Level 3). ;
[0114] This yields the risk level probability vector for each grid cell. The risk level probability vector of each grid is stored according to the grid index to obtain the risk level probability field covering the entire target area.
[0115] Step S3-4: Aggregate the original grid in the risk level probability field into multiple superpixels according to spatial adjacency, and construct a graph structure with each superpixel as a graph node and spatial adjacency as an edge.
[0116] The number of original grids contained in each superpixel is set according to the measurement area and the number of original grids. The number of superpixels is usually controlled between 100 and 500 to balance information preservation and computational efficiency. For example, if the number of original grids is 4000 and the number of superpixels is preset to 200 according to the measurement area, then each superpixel contains 20 original grids. Based on spatial adjacency and using simple linear iterative clustering as the superpixel segmentation algorithm, the original grids in the risk level probability field are adaptively aggregated into multiple superpixels. Among them, grids that are less than a complete block at the boundary of the measurement area are included in the adjacent superpixels.
[0117] Specifically, initial centers are uniformly sampled on the original grid plane according to a preset superpixel spacing C. Each initial center contains the GPS coordinates and probability vector of the original grid. Within a 2C×2C area around each initial center, the total distance D from each original grid to the initial center is calculated. The original grid is assigned to the initial center with the closest total distance, and the GPS coordinates and probability vector of each initial center are updated. The above steps are repeated until only isolated small regions exist. The isolated small regions are then merged into the adjacent largest superpixel.
[0118] Each superpixel is treated as a graph node in the graph structure. The Chebyshev distance between each superpixel is calculated. If there is a grid pair with a Chebyshev distance of 1, it indicates that there is a common edge or common point in the spatial adjacency relationship between the two superpixels. An undirected edge is established in this way. The above steps are repeated until a complete graph structure is obtained. The graph structure discretizes the geographical region into a spatial adjacency network, which can represent the diffusion path and dependence of heavy metal pollution.
[0119] Step S4: Input the graph structure into the lower-level agent to obtain a provisional risk zoning scheme and global state features. The upper-level agent performs binary decision-making based on the global state features and outputs the risk zoning scheme and global probability field. The lower-level agent executes based on the inner loop, and the upper-level agent executes based on the outer loop.
[0120] In this embodiment, step S4 specifically includes steps S4-1 to S4-4:
[0121] Step S4-1: Obtain the risk level of each superpixel in the graph structure based on the risk level probability field.
[0122] Step S4-2: Calculate the average value of the risk level probability vector as the risk composition feature, calculate the variance of the risk level probability as the uncertainty feature, and simultaneously count the density of all measured original grids as the measurement density feature to obtain the risk feature set.
[0123] Based on the risk level probability field, the risk level probability vector corresponding to all original grids within each superpixel is statistically analyzed, the average probability of each risk level is calculated, and the risk level with the highest average probability is taken as the risk level of the corresponding superpixel.
[0124] The average value of the risk level probability vector is calculated as the risk composition feature to represent the average degree to which the superpixel as a whole belongs to each risk level. The variance of the risk level probability is calculated as the uncertainty feature to represent the reliability of the risk level determination within the superpixel. The density of all measured original grids is statistically analyzed as the measurement density feature to represent the completeness of the electromagnetic original dataset. The risk feature set is composed of the risk composition feature, the uncertainty feature, and the measurement density feature.
[0125] Step S4-3: The lower-level agent selects and executes actions through the graph structure until the inner loop ends, in order to obtain the pause risk zone delineation plan and the lower-level confidence index.
[0126] For lower-level intelligent agents, specifically including:
[0127] First, a graph attention encoder consisting of two graph attention layers and one global pooling layer is constructed. Each node in each layer is determined based on the spatial adjacency relationship between superpixels. The input of the first graph attention layer is a graph structure, and the input of each subsequent layer is the node feature matrix output by the previous layer. The graph attention layer is used to update the features of each node and encode the spatial dependencies into the new feature representation.
[0128] Understandably, in each graph attention layer, for each undirected edge, based on the risk composition features of the node and the features of the neighboring nodes, the attention coefficient of the source node to the target node is calculated through a single-layer feedforward neural network. The attention coefficients of all incoming edges of each target node are normalized using the Softmax function so that the sum of the coefficients of all incoming edges is 1. Then, the old features of the neighboring nodes are weighted and summed according to the attention coefficients, and then passed through the ReLU function to obtain the updated node features. The average of the updated node features is taken as the output node feature matrix.
[0129] A global pooling layer composed of pooling functions is constructed after the graph attention layer to aggregate the features of all nodes. The input of the global pooling layer is the node feature matrix output by the second-layer graph attention layer, and the output is a spatial feature vector. The pooling function adopts average pooling, which can retain global statistical information. The arithmetic mean of all nodes in each feature dimension of the node feature matrix is calculated to obtain multiple averages. These averages are arranged in order into a multi-dimensional vector, which is the spatial feature vector.
[0130] Then, a temporal convolutional encoder consisting of three dilated convolutional layers and one global pooling layer is constructed. The input of the first dilated convolutional layer is the historical action sequence, and the input of each subsequent layer is the temporal feature sequence output by the previous layer. The dilated convolutional layer is used to update the features at each time step, encoding the temporal dependencies into the new feature representation.
[0131] Understandably, in each dilated convolutional layer, the input is first padded causally to ensure that the output at the current time step depends only on the current and previous inputs. The padding length is (f-1)×d, where f represents the convolution kernel size and d represents the dilation rate of the current layer. Then, convolution operations are performed through the convolution kernel to output the original features of each time step. Next, batch normalization is performed on the original features of each time step to accelerate training stabilization. Then, nonlinearity is introduced through the ReLU activation function to obtain the updated time step features. All the updated time step features are arranged in chronological order to form a temporal feature sequence with the length remaining unchanged.
[0132] Following the dilated convolutional layer is a global average pooling layer, which aggregates features from all time steps. The input to the global average pooling layer is the temporal feature sequence output from the third dilated convolutional layer, and the output is a temporal context vector. The pooling function uses average pooling to preserve global temporal statistics. For each feature dimension of the temporal feature sequence, the arithmetic mean of all time steps in the corresponding dimension is calculated and arranged in order into a multi-dimensional vector, which is the temporal context vector.
[0133] After constructing the graph attention encoder and the temporal convolutional encoder, the spatial feature vector and the temporal context vector are obtained respectively and concatenated to form the state space. At the same time, the operation of adjusting, maintaining or lowering the risk level and ending the inner loop are used as the action space. Thus, a policy network based on the flexible actor-critic algorithm and a value network based on the linear weighted multi-objective reward function can be set.
[0134] In the action space, raising the risk level indicates that the lower-level agent judges that the actual risk of the current superpixel is higher than the risk level; maintaining the risk level indicates that the lower-level agent considers the current risk level of the superpixel to be reasonable; lowering the risk level indicates that the lower-level agent judges that the actual risk of the current superpixel is lower than the risk level; and ending the inner loop indicates that the lower-level agent believes that the risk levels of all superpixels have been adjusted to a relatively stable state or that the risk levels of all superpixels have reached the threshold.
[0135] Furthermore, for the policy network, the policy network is responsible for outputting the action probability distribution based on the current state space, while the value network is responsible for evaluating the state-action value to guide policy updates.
[0136] The policy network consists of two fully connected hidden layers and an output layer, using ReLU as the activation function. It progressively maps state features to a high-dimensional space, then passes the output layer and uses the Softmax function to obtain the probability value of each action in the action space. The action with the highest probability is then selected for execution. The policy network aims to maximize the weighted sum of the expected cumulative reward and the entropy regularization term. The loss function is:
[0137] ;
[0138] in, Let be the mathematical expectation of a randomly sampled state s in the experience replay buffer D. Represented as the current state Next, based on the current policy network Output action probability distribution Action Take the expected value. Represented as the entropy regularization coefficient, The larger the value, the more the strategy tends to favor random actions. The smaller the value, the more certain the strategy. This represents the current policy network in state. Select action The logarithmic probability, Represented as in state Execute action Then, the expected cumulative reward that the agent can obtain.
[0139] Furthermore, for the value network, it consists of two fully connected hidden layers and an output layer, with ReLU as the activation function. The output layer is a single node with no activation function, directly outputting the average action value of the current inner loop. The value network aims to accurately estimate the expected cumulative reward, and the loss function is:
[0140]
[0141] in, This means randomly sampling an empirical tuple from the empirical playback buffer D. And take the mathematical expectation, This is indicated as an instant reward. Represented as a discount factor, Represented as the next state Based on the current policy network Output action probability distribution for actions Take the expected value. Represented as the state-action value output by the value network. Represented as in state Select action The logarithmic probability.
[0142] The immediate reward is obtained by a linearly weighted multi-objective reward function, which is:
[0143] ;
[0144] in This is expressed as a weighting coefficient for risk loss. This is expressed as the change in expected risk loss before and after the action is performed. This is represented as the weighting coefficient for adjusting costs. This is represented as an indicator function, which is 0 when the inner loop finishes execution, and 1 otherwise. This is expressed as the weight coefficient of the boundary smoothing penalty. Represented as boundary changes, all the above weight coefficients are set based on the degree of importance attached to the corresponding target.
[0145] The lower-level agent is constructed by sequentially connecting the graph attention encoder, the temporal convolutional encoder, the policy network, and the value network. The graph structure and the historical actions of the lower-level agent are used as inputs to the graph attention encoder and the temporal convolutional encoder, respectively, to obtain spatial feature vectors and temporal context vectors, which are then concatenated to form the state space. The state space is used as input to the policy network to obtain the probability of upsetting, maintaining, or downsetting the risk level of each superpixel in the graph structure and to execute the action corresponding to the highest probability. At the same time, the immediate reward is calculated and the parameters of the lower-level network are optimized.
[0146] The lower-level agent ends its inner loop when the risk level of all superpixels no longer changes or the risk level of all superpixels reaches the threshold. It outputs the current zoning scheme as a provisional risk zoning scheme and the average action value of the last inner loop as a lower-level confidence index. The lower-level agent can capture the spatiotemporal dependency of risk levels between superpixels, iteratively optimize the risk level of all superpixels through reinforcement learning, and finally output an accurate and reliable provisional risk zoning scheme.
[0147] Step S4-4: Combine the risk feature set with the lower-level confidence index to obtain global state features. The upper-level agent executes binary decision-making based on the global state features until the outer loop ends, in order to output the risk zoning scheme.
[0148] Based on the risk feature set, the uncertainty features are averaged and concatenated with the risk composition features, measurement density features and lower-level confidence indicators to obtain the global state features.
[0149] For upper-level intelligent agents, specifically including:
[0150] We construct a policy network based on a flexible actor-critic algorithm and a value network based on a linearly weighted multi-objective reward function.
[0151] The construction methods for the policy network and value network are the same as those for the lower-level agent. The state space of the policy network is a global state feature, and the action space is a binary decision including continuing measurement and ending the zoning. Continuing measurement indicates that the upper-level agent believes the current provisional risk zoning scheme is inaccurate and needs to obtain more raw electromagnetic data to reduce uncertainty and improve the reliability of the risk zoning scheme. Ending the zoning indicates that the upper-level agent believes the current provisional risk zoning is accurate enough and needs to obtain more raw electromagnetic data. The linearly weighted multi-objective reward function is expressed as follows:
[0152] ;
[0153] in, This is expressed as a weighting coefficient for risk loss. This is represented as the final risk loss. This is expressed as a weighting coefficient for measurement costs. This is expressed as cumulative measurement cost.
[0154] The upper-layer agent takes global state features as input, makes binary decisions based on the policy network, including continuing measurement and ending the division, and updates the upper-layer network parameters based on the immediate reward after each decision.
[0155] When the upper-level agent outputs a decision to continue measurement, it returns to start the next batch of measurements and loops through steps S2-2 to S4-4.
[0156] When the upper-layer agent outputs the decision to end the zoning, it outputs the current provisional risk zoning scheme as the risk zoning scheme, and arranges the risk level probability vectors of all original grids according to their spatial positions to obtain the global probability field. The upper-layer agent is used to evaluate the accuracy and credibility of the provisional risk zoning scheme, and continuously iterates and optimizes the provisional risk zoning scheme through binary decision-making to obtain the final risk zoning scheme.
[0157] Step S5: Obtain a pollution risk zoning map based on the risk zoning scheme and the global probability field.
[0158] The risk level of each superpixel in the risk zoning scheme is mapped back to the original grid, and a pollution risk zoning map is obtained based on the global probability field. Standard color codes are used to color and mark areas with different risk levels.
[0159] In one possible embodiment, such as Figure 3 As shown, after processing through steps S1 to S5, the final pollution risk zoning map is generated. In the map, area A represents a low-risk area, area B represents a medium-risk area, area C represents a high-risk area, and area D represents an extremely high-risk area.
[0160] Furthermore, this invention also provides a dynamic risk zoning system for soil heavy metal pollution, comprising: a processor and a memory, the memory and the processor being connected, the memory being used to store programs, instructions or code, and the processor being used to execute the programs, instructions or code in the memory to implement the aforementioned dynamic risk zoning method for soil heavy metal pollution.
[0161] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0162] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for dynamic risk zoning of soil heavy metal pollution, characterized in that, include: A physical information neural network is constructed based on the original dataset, and residual constraint terms of Maxwell's equations are embedded to obtain an electromagnetic forward modeling model. After obtaining the total number of soil measurement points, the total number of soil measurement points are measured in real time in batches based on the mobile platform to obtain the electromagnetic raw dataset. The original electromagnetic dataset is imported into the electromagnetic forward model to obtain the heavy metal concentration estimation set. At the same time, the concentration posterior distribution is generated after inversion by combining the Bayesian compressed sensing model. The risk level probability field is obtained based on the concentration posterior distribution, and a graph structure is constructed. The graph structure is input into the lower-level agent to obtain a provisional risk zoning scheme and global state features. The upper-level agent performs binary decision-making based on the global state features and outputs the risk zoning scheme and global probability field. The lower-level agent executes based on the inner loop, and the upper-level agent executes based on the outer loop. Pollution risk zoning maps are obtained based on risk zoning schemes and global probability fields.
2. The method for dynamic risk zoning of soil heavy metal pollution according to claim 1, characterized in that, A physical information neural network is constructed based on the original dataset, and residual constraint terms of Maxwell's equations are embedded to obtain an electromagnetic forward model, including: Ground calibration points were set up to obtain a raw dataset that includes at least heavy metal concentration data, electromagnetic induction data, soil moisture data, and temperature data. The raw dataset was then preprocessed and normalized. A physical information neural network consisting of an input layer, a hidden layer, and an output layer is constructed. The residual constraint term of Maxwell's equations is embedded in the loss function as a physical constraint. The Adam optimizer is used to minimize the loss function. The physical information neural network is trained with the original dataset to obtain an electromagnetic forward mapping model with the heavy metal concentration estimation set as the output target.
3. The method for dynamic risk zoning of soil heavy metal pollution according to claim 1, characterized in that, After obtaining the total number of soil measurement points, real-time measurements of the total soil measurement points are performed in batches using a mobile platform to obtain the raw electromagnetic dataset, including: The total number of measurement points is determined based on the area of the measurement area and the preset measurement density. Two-dimensional coordinate points covering the measurement area are generated based on the Halton low difference sequence as the total soil measurement points. The total soil measurement points are divided into multiple batches in the order of generation and the measurement path is determined. Based on the measurement path, a mobile platform equipped with an electromagnetic signal measuring instrument, a soil moisture sensor, and an infrared temperature sensor is controlled in batches to measure and acquire electromagnetic raw datasets, which include electromagnetic induction data, soil moisture data, temperature data, and GPS coordinates. Once the first batch of raw electromagnetic datasets is acquired, the measurement is paused and the system waits for a decision from the upper-level agent. If the upper-level agent decides to continue the measurement, the system continues to acquire raw electromagnetic datasets until the upper-level agent decides to end the zoning process.
4. The method for dynamic risk zoning of soil heavy metal pollution according to claim 1, characterized in that, The raw electromagnetic dataset is imported into an electromagnetic forward model to obtain an estimated set of heavy metal concentrations. Simultaneously, a Bayesian compressed sensing model is used for inversion to generate a posterior concentration distribution. Based on this posterior distribution, a risk level probability field is obtained, and a graph structure is constructed, including: Electromagnetic induction data, soil moisture data, and temperature data from the original electromagnetic dataset are input into the electromagnetic forward model to obtain a set of heavy metal concentration estimates. Based on the GPS coordinates and heavy metal concentration estimation set in the original electromagnetic dataset, the concentration posterior distribution is obtained by inversion through a Bayesian compressed sensing model. The concentration posterior distribution includes the posterior mean, which characterizes the best estimate of heavy metal concentration, and the posterior variance, which characterizes the uncertainty. The measurement area is discretized into original grids. Based on the set risk screening value and risk control value, and the probability of each original grid belonging to each risk level is calculated based on the concentration posterior distribution to obtain the risk level probability field. The original grid in the risk level probability field is aggregated into multiple superpixels according to spatial adjacency. A graph structure is constructed with each superpixel as a graph node and spatial adjacency as an edge.
5. The method for dynamic risk zoning of soil heavy metal pollution according to claim 4, characterized in that, For Bayesian compressed sensing models, including: The GPS coordinates in the original electromagnetic dataset were used as the observation locations, and the corresponding heavy metal concentration estimates were used as the observation data. The observation location and observation data are input into the Bayesian compressed sensing model, which uses a sparse variational Gaussian process as the inference framework and is configured with an anisotropic squared exponential kernel function with initial hyperparameters. After receiving the input, the Bayesian compressed sensing model uses the anisotropic squared exponential kernel function to represent the spatial correlation prior and obtains the posterior probability distribution through the sparse variational inference algorithm. In this process, after each inversion of the Bayesian compressed sensing model, the hyperparameters of the anisotropic squared exponential kernel function are updated using the expectation-maximization algorithm. The hyperparameters include spatial correlation length and signal variance.
6. The method for dynamic risk zoning of soil heavy metal pollution according to claim 1, characterized in that, A graph structure is input into the lower-level agent to obtain a provisional risk zoning scheme and global state features. The upper-level agent performs a binary decision based on the global state features and outputs a risk zoning scheme and a global probability field. The lower-level agent executes based on an inner loop, and the upper-level agent executes based on an outer loop, including: The risk level of each superpixel in the graph structure is obtained based on the risk level probability field. The average value of the risk level probability vector is calculated as the risk composition feature, the variance of the risk level probability is calculated as the uncertainty feature, and the density of all measured original grids is statistically analyzed as the measurement density feature to obtain the risk feature set. The lower-level intelligent agent selects and executes actions through the graph structure until the inner loop ends, in order to obtain a provisional risk zoning plan and lower-level confidence indicators. By combining the risk feature set with the lower-level confidence index, the global state features are obtained. The upper-level intelligent agent executes binary decision-making based on the global state features until the outer loop ends, in order to output the risk zoning scheme.
7. The method for dynamic risk zoning of soil heavy metal pollution according to claim 6, characterized in that, For lower-level and upper-level intelligent agents, including: The lower-level intelligent agent consists of a graph attention encoder, a temporal convolutional encoder, a policy network based on a flexible actor-critic algorithm, and a value network based on a linear weighted multi-objective reward function. The spatial feature vector and temporal context vector output by the graph attention encoder and the temporal convolutional encoder are used as the state space, and the operations of adjusting, maintaining or lowering the risk level and ending the inner loop are used as the action space. The upper-layer intelligent agent consists of a policy network based on a flexible actor-critic algorithm and a value network based on a linearly weighted multi-objective reward function, with global state features as the state space and continued measurement and termination of the division as the action space.
8. The method for dynamic risk zoning of soil heavy metal pollution according to claim 6, characterized in that, The lower-level agent selects and executes actions through a graph structure until the inner loop ends, in order to obtain a provisional risk zoning plan and lower-level confidence indicators, including: The graph structure and historical action sequence are used as inputs to the graph attention encoder and the temporal convolutional encoder, respectively, to obtain spatial feature vectors and temporal context vectors, which are then concatenated to form the state space. The state space is used as input to the policy network to obtain the probability of up-adjusting, maintaining, or down-adjusting the risk level of each superpixel in the graph structure and to execute the action corresponding to the highest probability. At the same time, the immediate reward is calculated and the parameters of the lower network are optimized. The lower-level agent outputs the end of the inner loop when the risk level of all superpixels no longer changes or the risk level of all superpixels reaches the threshold. It outputs the current zoning scheme as a provisional risk zoning scheme and the average action value of the last inner loop as the lower-level confidence index.
9. The method for dynamic risk zoning of soil heavy metal pollution according to claim 6, characterized in that, By combining the risk feature set with the lower-level confidence index to obtain global state features, the upper-level agent executes binary decisions based on these global state features until the outer loop ends, outputting a risk zoning scheme, including: Based on the risk feature set, the average value of the uncertainty features is taken and concatenated with the risk composition features, measurement density features and lower-level confidence indicators to obtain the global state features. The upper-level agent takes global state features as input, makes binary decisions based on the policy network, including continuing measurement and ending the division, and updates the upper-level network parameters based on the immediate reward after each decision. When the upper-level agent outputs a decision to continue measurement, it returns to start the next batch of measurements and performs subsequent processing; When the upper-level agent outputs the decision to end the zoning, it outputs the current provisional risk zoning scheme as the risk zoning scheme.
10. A dynamic risk zoning system for soil heavy metal pollution, characterized in that, include: The dynamic risk zoning system for soil heavy metal pollution includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the dynamic risk zoning method for soil heavy metal pollution as described in any one of claims 1-9.