Soil heavy metal pollution interpolation prediction method of variation coupling graph attention mechanism

Through the variational coupled graph attention mechanism, the problem of insufficient utilization of spatial correlation in traditional interpolation methods is solved, and more accurate soil heavy metal pollution prediction is achieved. It adapts to different geographical conditions and pollution scenarios, and improves the sensitivity of polluted areas and the prediction accuracy of unsampled points.

CN120596871APending Publication Date: 2025-09-05CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510731178.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional interpolation methods are unable to fully utilize the spatial correlation and environmental characteristics between sampling points, resulting in inaccurate interpolation results of soil heavy metal pollution.

Method used

The variational coupled graph attention mechanism is adopted to calculate the Euclidean distance, dynamically divide the distance interval, perform empirical mode decomposition, construct the graph structure and integrate the attention mechanism. The multi-head attention mechanism and multi-layer GAT structure are used to iteratively process the graph structure to generate the interpolation prediction results of soil heavy metal pollution.

Benefits of technology

It improves the understanding of soil heavy metal pollution patterns, adapts to different geographical conditions and pollution scenarios, provides more accurate and reliable prediction results, and improves the sensitivity of severely polluted areas and the prediction accuracy of heavy metal concentrations in unsampled points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596871A_ABST
    Figure CN120596871A_ABST
Patent Text Reader

Abstract

The invention discloses a soil heavy metal pollution interpolation prediction method of a variation coupling graph attention mechanism, and relates to the field of soil heavy metal detection, and the method comprises the following steps: collecting a target area soil sample, and calculating the Euclidean distance between every two sampling points; dividing a plurality of distance intervals; carrying out empirical mode decomposition on the heavy metal concentration and verifying an intrinsic part; calculating semi-variation values of the heavy metal concentration intrinsic part under different distances; fitting a theoretical variation function model; constructing a graph structure; calculating an attention weight coefficient between the nodes; obtaining node-level mapping features; and taking the node-level mapping characteristics output by the last layer of GAT structure as the input of a pre-trained multi-layer perceptron, and outputting a soil heavy metal pollution interpolation prediction result. According to the method, the variation function and the distance are integrated into the GAT model, the understanding of the soil heavy metal pollution mode is enhanced, the method can adapt to different geographical conditions and pollution scenes, and a more accurate and reliable prediction result is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of soil heavy metal detection, and in particular to a soil heavy metal pollution interpolation prediction method based on a variational coupled graph attention mechanism. Background Art

[0002] Heavy metals are the primary inorganic pollutants in contaminated sites and have attracted considerable attention due to their high toxicity and bioaccumulation. Heavy metal pollution in contaminated sites primarily originates from mining, industrial emissions, solid waste, and sewage irrigation. These heavy metals enter the soil through air deposition and runoff, even contaminating local groundwater. Therefore, it is imperative to characterize soil contamination at contaminated sites and implement remedial measures for heavy metal-contaminated soils. Accurately predicting the distribution of heavy metals in soil is crucial. However, traditional interpolation methods (such as kriging) struggle to fully exploit the spatial correlation and environmental characteristics between sampling points, resulting in inaccurate interpolation results for heavy metal contamination in soils. This makes it difficult to provide more accurate data support for fields such as earth science, environmental science, and resource management. Summary of the Invention

[0003] In response to the above-mentioned deficiencies in the prior art, the soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism provided by the present invention solves the problem that traditional interpolation methods are unable to fully utilize the spatial correlation and environmental characteristics between sampling points, resulting in inaccurate soil heavy metal pollution interpolation results.

[0004] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: A soil heavy metal pollution interpolation prediction method based on a variational coupled graph attention mechanism is provided, which includes the following steps: Collect soil samples from the target area, record the geographic coordinates and heavy metal concentrations of each sampling point, and calculate the Euclidean distance between any two sampling points. Dynamically divide the distance intervals based on the distribution characteristics of the Euclidean distance between any two sampling points. Perform empirical mode decomposition on heavy metal concentrations to obtain intrinsic mode functions and residual terms; Verify the intrinsic properties of each intrinsic mode function and residual term, and select the intrinsic mode function and residual term that meet the intrinsic conditions as the intrinsic part of heavy metal concentration; According to the experimental semivariogram, the semivariogram of the intrinsic part of heavy metal concentration at different distances is calculated to generate data pairs with the structure of (distance, semivariogram), which are recorded as experimental semivariogram data pairs; Fit the experimental semivariogram data pairs to the theoretical variogram model, and select the theoretical variogram with the highest fitting degree and the smallest error as the variogram of the intrinsic part of heavy metal concentration; Constructing a graph structure: taking sampling points as nodes of the graph attention network, the heavy metal concentrations of the sampling points as node attributes, and the variogram and distance of the intrinsic part of the heavy metal concentration between two nodes as edge attributes; The graph structure is input into the graph attention network, the heavy metal concentration, the variogram of the intrinsic part of the metal concentration, and the distance are integrated into the attention mechanism to calculate the attention weight coefficient between nodes; The graph structure is iteratively processed through a multi-head attention mechanism and a multi-layer GAT structure to obtain node-level mapping features; The node-level mapping features output by the last layer of the GAT structure are used as the input of the pre-trained multi-layer perceptron to output the interpolation prediction results of soil heavy metal pollution.

[0005] Further, heavy metals include cadmium, arsenic, chromium, copper, mercury, manganese, lead, zinc and nickel.

[0006] Furthermore, the Euclidean distance between two sampling points is calculated as follows:

[0007] in and For sampling points With sampling point geographical coordinates; For sampling points With sampling point The Euclidean distance of .

[0008] Furthermore, according to the distribution characteristics of the Euclidean distance between two sampling points, the specific method of dynamically dividing the distance interval is as follows: According to the distribution characteristics of the Euclidean distance between two sampling points, several equal distance intervals are divided to ensure that each distance interval contains enough sampling points for stable calculation of the semi-variance value.

[0009] Furthermore, the expression for the empirical mode decomposition of heavy metal concentration is:

[0010] in Indicates sampling point Heavy metal concentrations; is the eigenmode term, and all eigenmode terms constitute the eigenmode function; For sampling points All residual values ​​constitute the residual term.

[0011] Furthermore, the intrinsic properties of each intrinsic mode function and residual term are verified, and the method of selecting the intrinsic mode function and residual term that meet the intrinsic properties condition as the intrinsic part of heavy metal concentration includes: Determine whether all intrinsic mode terms contained in the current intrinsic mode function meet the conditions that the difference between the number of extreme value points and the number of zero crossing points is at most 1 and the average value of the upper and lower envelopes is within the set range, retain the intrinsic mode terms that meet the above conditions, and eliminate the heavy metal concentration data corresponding to the intrinsic mode terms that do not meet the above conditions, and obtain the intrinsic mode function that meets the intrinsic conditions; Determine whether the residual term satisfies the following expression:

[0012]

[0013] If it is satisfied, the residual term is determined to meet the intrinsic hypothesis; otherwise, the residual term is determined not to meet the intrinsic hypothesis, and the heavy metal concentration data corresponding to the residual term that does not meet the intrinsic hypothesis are eliminated; For sampling points The residual term of Expressing hope; Indicates the calculation of variance; represents the variogram.

[0014] Furthermore, the expression for calculating the semivariance of the intrinsic part of heavy metal concentration at different distances is:

[0015] in semivariance; is the total number of sampling point pairs in the current distance interval; and The sampling point pairs and The intrinsic concentration of Indicates the representative distance corresponding to the current distance interval. Each distance interval corresponds to a representative distance.

[0016] Furthermore, theoretical variogram models include spherical model, exponential model, Gaussian model, logarithmic model, hole model and linear model; When fitting the variogram, the coefficient of determination and / or RMSE index is used to select the theoretical variogram with the highest fit and the smallest error.

[0017] Furthermore, the expression for calculating the attention weight coefficient between nodes is:

[0018] in For nodes With node The attention weight between LeakyReLU activation function; is the range of variation in the variogram; is a learnable weight matrix; Represents a splicing operation; and Node With node The mapping characteristics of For nodes With node The Euclidean distance of the sample point With sampling point The Euclidean distance of A learnable parameter to control the effect of the variogram on attention; For nodes With node Corresponding variogram; node For nodes The superscript T indicates the transpose of the matrix.

[0019] Furthermore, the specific method of iteratively processing the graph structure through the multi-head attention mechanism and the multi-layer GAT structure to obtain node-level mapping features includes the following steps: The node Its neighboring nodes The attention weights between are softmax normalized to obtain the normalized attention weights, which are expressed as:

[0020] in Representation node With node The normalized attention weights between them; It represents the exponential with the natural constant e as the base; For nodes The number of neighbors; Refers to the node Neighbors; Computational Graph Attention Network Layer output nodes The mapping feature of is expressed as:

[0021] in For the graph attention network Layer output nodes The mapping characteristics of is the activation function; For the graph attention network Nodes in the layer With node The normalized attention weights between ; For the graph attention network The learnable weight matrix in the layer, ; For the graph attention network Layer output nodes The mapping characteristics of hour, .

[0022] The beneficial effects of the present invention are: 1. By integrating the variogram and distance into the GAT model, we enhance our understanding of soil heavy metal pollution patterns, adapt to different geographical conditions and pollution scenarios, and provide more accurate and reliable prediction results. This not only increases the sensitivity to severely polluted areas, but also improves the prediction accuracy of heavy metal concentrations at unsampled points, providing more accurate data support for fields such as earth science, environmental science, and resource management.

[0023] 2. The variogram is used to assess the degree of dispersion of soil heavy metal content between different sampling points to help identify pollution hotspots. The diffusion paths of pollutants are simulated through geographic Euclidean distance and characteristic distances based on factors such as soil type and land use, thereby enhancing the adaptability and generalization ability of the model under different geographical conditions and pollution scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of the process of this method. DETAILED DESCRIPTION

[0025] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0026] like Figure 1 As shown in Figure 2, the soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism includes the following steps: S1. Collect soil samples from the target area, record the geographic coordinates and heavy metal concentrations of each sampling point, calculate the Euclidean distance between each sampling point, and dynamically divide the distance intervals based on the distribution characteristics of the Euclidean distance between each sampling point. S2. Perform empirical mode decomposition on heavy metal concentrations to obtain intrinsic mode functions and residual terms; S3. Verify the intrinsic properties of each intrinsic mode function and residual term, and select the intrinsic mode function and residual term that meet the intrinsic properties condition as the intrinsic part of heavy metal concentration; S4. Calculate the semivariogram of the intrinsic part of heavy metal concentration at different distances based on the experimental semivariogram, and generate data pairs with the structure of (distance, semivariogram), which are recorded as experimental semivariogram data pairs; S5. Fit the experimental semivariogram data to the theoretical variogram model, and select the theoretical variogram with the highest fitting degree and the smallest error as the variogram of the intrinsic part of the heavy metal concentration; S6. Build a graph structure: take the sampling points as nodes of the graph attention network, the heavy metal concentrations of the sampling points as node attributes, and the variogram and distance of the intrinsic part of the heavy metal concentration between two nodes as edge attributes; S7. Input the graph structure into the graph attention network, integrate the heavy metal concentration, the variogram of the intrinsic part of the metal concentration, and the distance into the attention mechanism, and calculate the attention weight coefficient between nodes; S8, iteratively process the graph structure through the multi-head attention mechanism and multi-layer GAT structure to obtain node-level mapping features; S9. The node-level mapping features output by the last layer of the GAT structure are used as the input of the pre-trained multi-layer perceptron to output the interpolation prediction results of soil heavy metal pollution.

[0027] In this embodiment, the heavy metals include cadmium, arsenic, chromium, copper, mercury, manganese, lead, zinc, and nickel (Cd, As, Cr, Cu, Hg, Mn, Pb, Zn, and Ni).

[0028] The calculation expression of the Euclidean distance between two sampling points in step S1 is:

[0029] in and For sampling points With sampling point geographical coordinates; For sampling points With sampling point The Euclidean distance of .

[0030] In step S1, the specific method of dynamically dividing the distance intervals according to the distribution characteristics of the Euclidean distance between two sampling points is as follows: Based on the distribution characteristics of the Euclidean distance between two sampling points, several equal distance intervals are divided to ensure that each distance interval contains enough sampling points for stable calculation of the semi-variance value. The distance range that covers the main spatial correlation is selected as the range for variation calculation.

[0031] The expression for the empirical mode decomposition of heavy metal concentration in step S2 is:

[0032] in Indicates sampling point Heavy metal concentrations; is the eigenmode term, and all eigenmode terms constitute the eigenmode function; For sampling points All residual values ​​constitute the residual term.

[0033] In step S3, the method of verifying the intrinsic properties of each intrinsic mode function and residual term and selecting the intrinsic mode function and residual term that meet the intrinsic property condition as the intrinsic part of the heavy metal concentration includes: Determine whether all intrinsic mode terms contained in the current intrinsic mode function meet the conditions that the difference between the number of extreme value points and the number of zero crossing points is at most 1 and the average value of the upper and lower envelopes is within the set range, retain the intrinsic mode terms that meet the above conditions, and eliminate the heavy metal concentration data corresponding to the intrinsic mode terms that do not meet the above conditions, and obtain the intrinsic mode function that meets the intrinsic conditions; Determine whether the residual term satisfies the following expression:

[0034]

[0035] If it is satisfied, the residual term is determined to meet the intrinsic hypothesis; otherwise, the residual term is determined not to meet the intrinsic hypothesis, and the heavy metal concentration data corresponding to the residual term that does not meet the intrinsic hypothesis are eliminated; For sampling points The residual term of Expressing hope; Indicates the calculation of variance; Represents the variogram, which is the semivariogram value with spatial distance d The overall function model or curve of the change, the semivariogram is the difference between the variogram and the specific spatial distance. d The calculated value below.

[0036] In step S4, the expression for calculating the semi-variance value of the intrinsic part of heavy metal concentration at different distances is:

[0037] in semivariance; is the total number of sampling point pairs in the current distance interval; and The sampling point pairs and The intrinsic concentration of Indicates the representative distance corresponding to the current distance interval. Each distance interval corresponds to a representative distance.

[0038] In this embodiment, the theoretical variogram model includes a spherical model, an exponential model, a Gaussian model, a logarithmic model, a hole model, and a linear model; wherein: The expression of the spherical model is:

[0039] The expression of the exponential model is:

[0040] The expression of Gaussian model is:

[0041] The expression of the logarithmic model is:

[0042] The expression of the hole model is:

[0043] The linear model is expressed as:

[0044] in The nugget effect represents random variation or measurement error at a very small scale; is the structural variance, which represents the part after deducting the nugget effect from the sill value, which means that the total sill value is ; is the range, which defines the maximum distance of spatial autocorrelation.

[0045] When fitting the variogram, the experimental semi-variogram is fitted to the theoretical model using the (weighted) least squares method to estimate the 、 and Three parameters are selected, and then the coefficient of determination and / or RMSE index are used to select the theoretical variogram with the highest fit and the smallest error.

[0046] The expression for calculating the attention weight coefficient between nodes in step S7 is:

[0047] in For nodes With node The attention weight between LeakyReLU activation function; is the range of variation in the variogram; is a learnable weight matrix; Represents a splicing operation; and Node With node The mapping characteristics of For nodes With node The Euclidean distance of the sample point With sampling point The Euclidean distance of A learnable parameter to control the effect of the variogram on attention; For nodes With node Corresponding variogram; node For nodes The superscript T indicates the transpose of the matrix.

[0048] In step S8, the specific method of iteratively processing the graph structure through the multi-head attention mechanism and the multi-layer GAT structure to obtain the node-level mapping features includes the following steps: S8-1, the node Its neighboring nodes The attention weights between are softmax normalized to obtain the normalized attention weights, which are expressed as:

[0049] in Representation node With node The normalized attention weights between them; It represents the exponential with the natural constant e as the base; For nodes The number of neighbors; Refers to the node Neighbors; In a single-layer GAT structure, the node The features of are obtained from its neighbors by using normalized attention weights Aggregate information to perform single-head GAT feature updates:

[0050] in is the mapping feature; In order to enhance the expressiveness of the model, this method adopts a multi-head attention mechanism and combines the outputs of single heads. Finally, a multi-layer GAT structure is stacked, based on this: S8-2. Computational Graph Attention Network Layer output nodes The mapping feature of is expressed as:

[0051] in For the graph attention network Layer output nodes The mapping characteristics of is the activation function; For the graph attention network Nodes in the layer With node The normalized attention weights between ; For the graph attention network The learnable weight matrix in the layer, ; For the graph attention network Layer output nodes Mapping features.

[0052] In GAT (Graph Attention Network), traditional attention mechanisms typically only consider node features to calculate attention weights. However, in this method, we improve the attention mechanism so that it not only relies on node features but also incorporates the variogram and inter-node distance to better model spatial dependencies.

[0053] Pre-trained multi-layer perceptron (MLP) uses the mapping features from the last layer of the graph attention network , which is nonlinearly mapped to predict the value of the unknown node. Specifically, It is input into the MLP, processed through multiple fully connected layers and nonlinear activation functions, and converted into the predicted value of the node. In this way, the MLP provides interpolation for each unknown node, making full use of the rich, high-order local features captured by the multi-layer structure of GAT.

[0054] In one embodiment of the present invention, the graph attention network that integrates the variogram and distance information in this method is named GAT-V&D, and a comparative experiment is designed to verify the performance of this method. The specific comparison models include: This GAT-V&D: In the GAT attention mechanism, the variogram value and distance information are combined to calculate the attention weight. Combining the variogram and spatial distance improves the interpolation prediction accuracy. The corresponding attention weight expression is:

[0055] GAT+Variation (VGAT): When GAT calculates attention, only the variogram value is concatenated. The corresponding attention weight expression is:

[0056] GAT+Distance (DGAT): When GAT calculates attention, only the Euclidean distance information is concatenated. The corresponding attention weight expression is:

[0057] Traditional GAT (GAT): Attention is calculated only based on node features. As a baseline model, its performance difference with the improved model is compared. The corresponding attention weight expression is: .

[0058] The root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²) are used to evaluate the model's predictive performance. The root mean square error (RMSE) is a commonly used metric to measure the difference between the predicted value and the true value.

[0059]

[0060] The smaller the RMSE value, the better the model performs. Similarly, the mean absolute error measures the average absolute difference between the predicted value and the true value, where the formula is:

[0061] The smaller the MAE value, the better the model performance. The R² score is then calculated using the coefficient of determination formula to measure how well the model fits the data. The coefficient of determination formula is:

[0062] The closer the R² value is to 1, the better the model explains the variation in the data.

[0063] Table 1

[0064] Table 1 shows the comparative results for the soil heavy metal prediction task. As can be seen from this table, the GAT-V&D model proposed in this method performs best. Comparing the performance metrics (R², RMSE%, and MAE%) of the different models shows that GAT-V&D achieves the highest R² values ​​and the lowest RMSE% and MAE% for all tested heavy metal elements, indicating that this model has higher prediction accuracy and smaller error. Using only the variogram (V-GAT) or only considering distance information (D-GAT) can also improve the model's predictive ability, but the results are not as good as the GAT-V&D model that combines both. Traditional GAT, due to its lack of consideration of spatial correlation, performs the worst in interpolation prediction tasks. Ordinary Kriging (OK), a classic geostatistical method, underperforms GAT and its variants in most cases, particularly in terms of R² and MAE. This comparative experiment validates the importance of incorporating both variogram and distance information into GAT for heavy metal pollution interpolation prediction.

[0065] In summary, by integrating the variogram and distance into the GAT model, the present invention enhances the understanding of soil heavy metal pollution patterns, can adapt to different geographical conditions and pollution scenarios, and provide more accurate and reliable prediction results. It not only increases the sensitivity to severely polluted areas, but also improves the prediction accuracy of heavy metal concentrations at unsampled points, and can provide more accurate data support for fields such as earth science, environmental science, and resource management.

Claims

1. A soil heavy metal pollution interpolation prediction method based on a variational coupled graph attention mechanism, characterized by: The following steps are involved: Collect soil samples from the target area, record the geographic coordinates and heavy metal concentrations of each sampling point, and calculate the Euclidean distance between any two sampling points. Dynamically divide the distance intervals based on the distribution characteristics of the Euclidean distance between any two sampling points. Perform empirical mode decomposition on heavy metal concentrations to obtain intrinsic mode functions and residual terms; Verify the intrinsic properties of each intrinsic mode function and residual term, and select the intrinsic mode function and residual term that meet the intrinsic conditions as the intrinsic part of heavy metal concentration; According to the experimental semivariogram, the semivariogram of the intrinsic part of heavy metal concentration at different distances is calculated to generate data pairs with the structure of (distance, semivariogram), which are recorded as experimental semivariogram data pairs; Fit the experimental semivariogram data pairs to the theoretical variogram model, and select the theoretical variogram with the highest fitting degree and the smallest error as the variogram of the intrinsic part of heavy metal concentration; Constructing a graph structure: taking sampling points as nodes of the graph attention network, the heavy metal concentrations of the sampling points as node attributes, and the variogram and distance of the intrinsic part of the heavy metal concentration between two nodes as edge attributes; The graph structure is input into the graph attention network, the heavy metal concentration, the variogram of the intrinsic part of the metal concentration, and the distance are integrated into the attention mechanism to calculate the attention weight coefficient between nodes; The graph structure is iteratively processed through a multi-head attention mechanism and a multi-layer GAT structure to obtain node-level mapping features; The node-level mapping features output by the last layer of the GAT structure are used as the input of the pre-trained multi-layer perceptron to output the interpolation prediction results of soil heavy metal pollution.

2. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: Heavy metals include cadmium, arsenic, chromium, copper, mercury, manganese, lead, zinc and nickel.

3. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: The calculation expression of the Euclidean distance between two sampling points is: in and For sampling points With sampling point geographical coordinates; For sampling points With sampling point The Euclidean distance of .

4. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: According to the distribution characteristics of the Euclidean distance between two sampling points, the specific method of dynamically dividing the distance interval is as follows: According to the distribution characteristics of the Euclidean distance between two sampling points, several equal distance intervals are divided to ensure that each distance interval contains enough sampling points for stable calculation of the semi-variance value.

5. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: The expression for empirical mode decomposition of heavy metal concentration is: in Indicates sampling point Heavy metal concentrations; is the eigenmode term, and all eigenmode terms constitute the eigenmode function; For sampling points All residual values ​​constitute the residual term.

6. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: Methods for verifying the intrinsic properties of each intrinsic mode function and residual term and selecting the intrinsic mode function and residual term that meet the intrinsic properties conditions as the intrinsic part of heavy metal concentration include: Determine whether all intrinsic mode terms contained in the current intrinsic mode function meet the conditions that the difference between the number of extreme value points and the number of zero crossing points is at most 1 and the average value of the upper and lower envelopes is within the set range, retain the intrinsic mode terms that meet the above conditions, and eliminate the heavy metal concentration data corresponding to the intrinsic mode terms that do not meet the above conditions, and obtain the intrinsic mode function that meets the intrinsic conditions; Determine whether the residual term satisfies the following expression: If it is satisfied, the residual term is determined to meet the intrinsic hypothesis; otherwise, the residual term is determined not to meet the intrinsic hypothesis, and the heavy metal concentration data corresponding to the residual term that does not meet the intrinsic hypothesis are eliminated; For sampling points The residual term of Expressing hope; Indicates the calculation of variance; represents the variogram.

7. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: The expression for calculating the semivariance of the intrinsic part of heavy metal concentration at different distances is: in semivariance; is the total number of sampling point pairs in the current distance interval; and The sampling point pairs and The intrinsic concentration of Indicates the representative distance corresponding to the current distance interval. Each distance interval corresponds to a representative distance.

8. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: Theoretical variogram models include spherical model, exponential model, Gaussian model, logarithmic model, hole model and linear model; When fitting the variogram, the coefficient of determination and / or RMSE index is used to select the theoretical variogram with the highest fit and the smallest error.

9. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: The expression for calculating the attention weight coefficient between nodes is: in For nodes With node The attention weight between is the LeakyReLU activation function; is the range of variation in the variogram; is a learnable weight matrix; Represents a splicing operation; and Node With node The mapping characteristics of For nodes With node The Euclidean distance of the sample point With sampling point The Euclidean distance of A learnable parameter to control the effect of the variogram on attention; For nodes With node Corresponding variogram; node For nodes The superscript T indicates the transpose of the matrix.

10. The soil heavy metal pollution interpolation prediction method based on the variational coupled graph attention mechanism according to claim 1 is characterized in that: The specific method of obtaining node-level mapping features by iteratively processing the graph structure through the multi-head attention mechanism and the multi-layer GAT structure includes the following steps: The node Its neighboring nodes The attention weight between Perform softmax normalization to obtain the normalized attention weight, which is expressed as: in Representation node With node The normalized attention weights between them; It represents the exponential with the natural constant e as the base; For nodes The number of neighbors; Refers to the node Neighbors; Computational Graph Attention Network Layer output nodes The mapping feature of is expressed as: in For the graph attention network Layer output nodes The mapping characteristics of is the activation function; For the graph attention network Nodes in the layer With node The normalized attention weights between ; For the graph attention network The learnable weight matrix in the layer, ; For the graph attention network Layer output nodes The mapping characteristics of hour, .