Ecological risk management system based on soil heavy metal data

By performing data binding, neighborhood relationship analysis, and hotspot cluster derivation on the ecological risk management system for soil heavy metal data, the problems of unclear regional scope support and data consistency in the existing system have been solved. This has enabled the reliability of risk identification and the executability of management decisions, and supports review and rolling updates.

CN122115180APending Publication Date: 2026-05-29CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
Filing Date
2026-03-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing ecological risk management systems based on soil heavy metal data suffer from several drawbacks in practical applications. These include a lack of structured description of the relationship between the regional scope and the sampling points, inconsistent data processing standards, insufficient comparability of results, and a tendency to over-expose or underestimate the true anomalies when the number of sampling points is limited or unevenly distributed. Consequently, these systems fail to meet the requirements of robustness and traceability.

Method used

The data acquisition module binds and stores the sampling point data into the database, performs access control and robust labeling, constructs spatial neighborhood relationships of sampling points, performs local spatial cluster significance testing, forms hotspot clusters and derives pollution cores, outputs risk level and uncertainty labeling, the system automatically selects the appropriate path and records the triggering reason and parameter version, and supports review and auditing.

Benefits of technology

It improves the reliability of regional risk identification and the feasibility of management decisions, ensures that the conclusions are consistent with the measured evidence at the sites, reduces the unstable output caused by subjective human intervention, and generates uncertainty representations when evidence is insufficient, supporting subsequent review and rolling updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115180A_ABST
    Figure CN122115180A_ABST
Patent Text Reader

Abstract

The application discloses an ecological risk management system based on soil heavy metal data and belongs to the field of heavy metal data analysis, and comprises the following steps: a data acquisition module, which is used for binding and storing the sampling point number, coordinates, field properties, laboratory element content, detection limit and abnormal element data, and performing access control, tail replacement and robust marking to form a traceable original data set; a pollution core analysis module, which is used for constructing the spatial neighborhood relationship of sampling points with heavy metal elements as the object, performing local spatial cluster significance test according to the neighborhood relationship, and performing connected aggregation on high-value points meeting the significance level to form a hot spot cluster, and then deducing the pollution core from the hot spot cluster; and a risk management module, which is used for outputting the risk level, evidence level and uncertainty marking inside and outside the core, and storing the process parameter version and the result to the ecological risk management system based on soil heavy metal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heavy metal data analysis technology, and in particular to an ecological risk management system based on soil heavy metal data. Background Technology

[0002] Heavy metal pollution in soil is characterized by its high degree of concealment, complex migration and transformation mechanisms, and highly non-uniform spatial distribution. Pollutants can enter and accumulate in the soil through agricultural inputs, wastewater irrigation, and industrial emissions, thereby impacting agricultural ecological security and environmental management. In actual supervision and remediation, management departments often need to develop risk zoning, core pollution area identification, key verification, and follow-up sampling strategies for the monitored area based on the test results of a limited number of sampling points. This is to support risk control, determination of remediation boundaries, and continuous updates.

[0003] Existing ecological risk management systems based on soil heavy metal data typically consist of four core components: data acquisition and detection, data storage and management, risk assessment and spatial analysis, and result output and application. First, sampling points are set up and samples are collected according to a pre-defined sampling plan, recording attributes such as point coordinates, depth, and land cover type. The laboratory completes sample pretreatment and tests the content of each heavy metal element, generating test results and related metadata. The system binds the test results with the point information and stores them in the database to form a basic dataset. Based on this, the system performs threshold or standard comparisons to identify points exceeding the standard and outputs risk levels or alarms. Simultaneously, it performs spatial inference between sampling points to generate regional distribution results or risk maps. If necessary, it further forms risk zoning on the assessment grid and outputs the boundaries of remediation or restoration areas. Finally, the risk level, zoning results, boundary results, and alerts are published to the management terminal for display and decision-making.

[0004] For example, Chinese invention patent CN118469314B discloses a method and system for assessing the ecological risk of heavy metals in soil. The method includes: obtaining soil testing information of a preset area; obtaining the content value of heavy metals in the soil based on the soil testing information; if the content value of heavy metals in the soil is greater than or equal to a preset content threshold of the corresponding heavy metal, then the corresponding soil testing point is considered to have heavy metal pollution and is designated as a pollution point; constructing a pollution area based on the pollution point and extracting the area of ​​the corresponding pollution area; calculating the threat index of heavy metals to soil ecology based on the area of ​​the pollution area, the content value of heavy metals in the area, and the name of the corresponding heavy metal; generating ecological risk warning information when the threat index of heavy metals to soil ecology is greater than a preset threat index threshold; and sending the ecological risk warning information to a preset management terminal for notification.

[0005] For example, Chinese invention patent CN114167031B discloses a method for predicting the bioavailability of heavy metals in soil, which includes: S1, determining the properties of the target soil and obtaining the total content of heavy metals corresponding to the target soil through testing; S2, substituting the properties of the target soil and the corresponding total content of heavy metals into a pre-constructed prediction model for the bioavailability of heavy metals in soil for calculation, and outputting the predicted bioavailability data of heavy metals.

[0006] However, in the process of implementing the technical solutions of the embodiments of this application, it was found that the above-mentioned technology has at least the following technical problems: Although the existing ecological risk management system based on soil heavy metal data can realize functions such as sampling point threshold discrimination, risk index calculation, risk map generation or risk zoning output, it still generally has the following problems in practical applications: On the one hand, the existing systems mostly form risk areas or governance boundaries based on the judgment of point exceedance or the interpolation, simulation, and model prediction results between sampling points. The supporting relationship between the area and the specific sampling points often remains at the result level display, lacking a structured explanation of the basis for the formation of the area, making it difficult for the management end to quickly determine whether the area belongs to actual measurement support or inferred extension, thus affecting the review and accountability. On the one hand, the handling of detection data in cases of detection limits, quality control anomalies, and batch differences is often simplified in existing systems to rules outside the process or manual agreements, leading to insufficient comparability of results for the same area under different batches or by different personnel. In addition, when the number of sampling points is limited, the spatial distribution is uneven, or there are blank zones, the risk maps or zoning results output by existing systems are prone to over-extrapolation of local high values ​​or underestimation of true anomalies. The system usually only outputs risk levels or graphical conclusions, lacking synchronous indications of the strength of evidence and expression of uncertainty. Consequently, subsequent verification priorities, encrypted sampling arrangements, and differentiated handling strategies lack sufficient basis and are difficult to meet the requirements of robustness and traceability in actual regulatory scenarios. Summary of the Invention

[0007] To address the technical problems existing in the prior art, embodiments of the present invention provide an ecological risk management system based on soil heavy metal data. Specifically, the module includes: a data acquisition module, used to bind and store sampling point numbers, coordinates, field attributes, laboratory element content, detection limits, and abnormal metadata into a database, and perform access control, truncation substitution, and robust labeling to form a traceable original dataset; a pollution core analysis module, used to construct spatial neighborhood relationships of sampling points based on heavy metal elements, and perform local spatial cluster significance tests based on neighborhood relationship analysis, and connect and aggregate high-value points that meet the significance level to form hotspot clusters, and then deduce the pollution core from the hotspot clusters. At the same time, a method discrimination switching unit is set up to diagnose the number of effective points, neighborhood connectivity and stability, significance results, and cluster stability. When the diagnostic results do not meet the core boundary verification conditions, it automatically switches to the continuous surface estimation process, interpolates between sampling points to form surfaces and simultaneously generates uncertainty representations, and extracts pollution cores under uncertainty constraints; and a risk management module, used to overlay pollution cores, sampling points, and management units, output risk levels inside and outside the core, evidence levels, and uncertainty labels, and uniformly store process parameter versions and results in the ecological risk management system database.

[0008] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. The ecological risk management system based on soil heavy metal data provided by this invention, without changing the constraints of existing sampling point layout, integrates sampling and evidence collection, detection and storage, spatial analysis, risk zoning, and result storage into a closed-loop process. This ensures that risk conclusions cover the needs of comprehensive management and maintain a traceability link consistent with measured evidence at specific points, thereby improving the reliability of regional risk identification, the interpretability of conclusions, and the executability of management decisions. It also supports consistency and comparability of results during subsequent verification and rolling updates. The system introduces a verifiable process diversion mechanism during element-level analysis: before performing spatial identification, it determines diagnostic quantities such as the scale of effective points, spatial connectivity, coverage of significant evidence, and availability of cluster results. These diagnostic quantities are then compared with fixed thresholds for consistency, automatically selecting the appropriate path and recording the triggering reason and parameter version. This mechanism allows the system to automatically select methods based on evidence conditions, reducing unstable output caused by subjective human intervention. It also ensures that each switch has a clear and traceable basis, avoiding incomparable conclusions due to different tasks or personnel operations.

[0009] 2. When the diagnostic data meets the conditions and enters the hotspot cluster identification path, the system performs statistical tests on high-value clusters under preset neighborhood relationship constraints. The cluster type and significance level of the points are used as direct evidence for cluster formation and boundary derivation. This allows the generation process of the core range to be traced back to specific points, neighborhood construction, and test parameters, facilitating verification and audit tracking. Simultaneously, cluster connectivity and aggregation allow the same element to form multiple independent clusters, enabling the system to maintain natural expression even under multiple pollution sources or multiple cores coexisting, without forcibly merging spatially independent risk hotspots into a single result.

[0010] 3. When hotspot cluster paths fail to produce stable conclusions due to sparse points, spatial fragmentation, or insufficient statistical evidence, this invention switches to a continuous surface estimation path. Based on quality gating and detection limit truncation, it spatially completes the information between sampling points and simultaneously generates uncertainty representations and evidence classifications. Then, it extracts suspected core areas using evidence constraints and automatically adds strong prompts to low-evidence regions. This output method ensures comprehensive risk management results even in scenarios with insufficient evidence, avoids misinterpreting the completion results as equivalent to definitive conclusions from actual measurements, and can provide follow-up action suggestions such as encrypted sampling, retesting, or key verification, providing a practical closed-loop support for risk management and sampling optimization. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram of the structure of the ecological risk management system based on soil heavy metal data provided in this application embodiment; Figure 2 A schematic diagram of the spatial aggregation statistical derivation method based on measured evidence from sampling points provided in this application embodiment; Figure 3 This is a schematic flowchart of a spatial information completion method based on continuous surface estimation provided in an embodiment of this application. Figure 4 This is a schematic diagram of the pollution core analysis and identification path switching process provided in the embodiments of this application. Detailed Implementation

[0013] The following provides explanations for some of the terms used in this application. It should be noted that these explanations are for the convenience of those skilled in the art and do not constitute a limitation on the scope of protection claimed in this application.

[0014] The embodiments of this application involve at least one, including one or more; where "multiple" means two or more. Furthermore, it should be understood that in the description of this specification, terms such as "first," "second," and "third" are used only for descriptive purposes and should not be construed as indicating relative importance or order. For example, "first device" and "second device" do not represent the degree of importance of the two or their order, but are merely for descriptive distinction. In the embodiments of this application, "and / or" merely describes an association relationship, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0015] The directional terms mentioned in the embodiments of this application, such as "up", "down", "left", "right", "inner", and "outer", are only for reference to the directions in the accompanying drawings. Therefore, the directional terms used are for better and clearer explanation and understanding of the embodiments of this application, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.

[0016] References to "one embodiment," "in some examples," or "some embodiments" as described in the embodiments of this application mean that one or more embodiments of this specification include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in some examples," "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0017] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0018] like Figure 1The diagram shows the structure of an ecological risk management system based on soil heavy metal data provided in this application embodiment. It includes: a data acquisition module, used to bind and store sampling point numbers, coordinates, field attributes, laboratory element content, detection limits, and abnormal metadata into a database, and to perform access control, truncation substitution, and robust labeling to form a traceable original dataset; and a pollution core analysis module, used to construct spatial neighborhood relationships of sampling points based on heavy metal elements, and to perform local spatial clustering significance tests based on neighborhood relationship analysis, and to connect and aggregate high-value points that meet the significance level to form hotspot clusters, which are then analyzed by... The cluster derivation of the pollution core is accompanied by a method discrimination switching unit, which is used to diagnose the number of effective points, neighborhood connectivity and stability, significance results, and cluster stability. When the diagnostic results do not meet the verifiable condition of the core boundary, it automatically switches to the continuous surface estimation process, interpolates between sampling points to form surfaces and generates uncertainty representations simultaneously, and extracts the pollution core under uncertainty constraints. The risk management module is used to overlay the pollution core with sampling points and management units, output the risk level inside and outside the core, the evidence level, and the uncertainty label, and uniformly store the process parameter version and results in the ecological risk management system database.

[0019] It should be noted that the "system" described in the embodiments of this application is an ecological risk management system based on soil heavy metal data.

[0020] Example 1: This example provides an ecological risk management system based on soil heavy metal data. For example... Figure 2 As shown, without altering the existing sampling point layout, the system uses measured data from the sampling points as primary evidence to first generate a traceable original dataset of heavy metals in the soil. Then, using heavy metal elements as the analysis object, it performs significance tests on high-value clusters under spatial neighborhood constraints to identify one or more hotspot clusters, and derives the pollution core range from these hotspot clusters. After obtaining the core range, the system jointly analyzes the core range with the original sampling point coordinates and detection results, outputting risk levels, evidence strength levels, and uncertainty markers for the core area and its surrounding regions. The process data and result data are uniformly stored in the ecological risk management system database to support subsequent review, rolling updates, and management decisions.

[0021] Specifically, during the data acquisition phase, the system deploys several sampling points within the monitored area according to the project's pre-set sampling plan, assigning a unique location number to each sampling point. It acquires the location coordinates, sampling date, sampling depth, and land cover type (e.g., farmland, forest, or wetland) for each sampling point and collects corresponding soil samples. Samples are packaged, retained, and transferred in batches. The transfer record includes at least the location number, sample batch number, and packaging time, ensuring a traceable one-to-one correspondence between samples and location information. The sampling plan, location numbering rules, and set of field attribute fields are fixed to the system configuration during project initialization to maintain consistent sampling standards and prevent incomparability of subsequent statistical results due to changes in standards.

[0022] During the testing phase, the laboratory performs sample pretreatment, which includes at least air drying, grinding, sieving, and digestion. Detection methods such as atomic absorption spectrometry, inductively coupled plasma mass spectrometry, or inductively coupled plasma atomic emission spectrometry are then used to determine the content of heavy metals such as cadmium, mercury, arsenic, lead, and chromium. In addition to outputting elemental content, the laboratory also simultaneously outputs metadata associated with the test results. This metadata includes at least detection limit information, method or instrument batch information, a quality control conclusion summary, and anomaly marker information. The quality control conclusion summary characterizes whether the batch test passed quality control checks such as blank, spiked recovery, parallel sample consistency, and instrument stability. The anomaly marker information identifies abnormal situations that may affect reliability, such as sample contamination risk, instrument drift risk, or results exceeding the acceptable range. The rules for selecting testing methods and the quality control judgment rules can be formalized in the laboratory's method documentation. Detection limits and quality control conclusions are written into and retained with each test batch for subsequent retrospective analysis.

[0023] During the data entry phase, the system binds the laboratory's output test results with the sampling point number and coordinate information, and then writes them into the ecological risk management system database, forming a raw soil heavy metal dataset. This raw soil heavy metal dataset includes at least the following fields: sampling point spatial information, heavy metal element types and content, detection metadata, quality markers, sample batches, and circulation indexes. To support subsequent review and auditing, the system generates a data fingerprint for each entry record and records the entry time and data version number, ensuring that any risk conclusion can be traced back to the corresponding original record, the corresponding test batch, and the corresponding quality control evidence.

[0024] like Figure 3As shown, before entering spatial statistical analysis, the system performs access control and numerical usability processing. The system determines whether a record should participate in the current round of spatial statistical calculations based on detection metadata and quality markers. For records that do not meet quality control requirements, exhibit abnormal flow, or are explicitly marked as abnormal, the system retains their original values ​​and reasons for the abnormality in the ecological risk management system database, but marks them as unusable for spatial statistics and excludes them from the statistical calculation set. Simultaneously, the system records the basis for removal and an evidence index, thereby preventing unqualified data from misleading hotspot identification. For data items below the detection limit, the system generates alternative values ​​for calculation according to preset truncation rules and retains the truncation marker to express the source of uncertainty, ensuring that spatial statistics can handle continuous numerical inputs without masking the incomplete information caused by values ​​below the detection limit. The aforementioned truncation rules can be calibrated offline using historical samples and laboratory practices and then embedded into the configuration. These rules are also recorded in the audit log with each version update. Furthermore, to reduce the accidental impact of extreme values ​​or suspected outliers on subsequent local clustering tests, the system generates robustness markers for records that pass the access control but exhibit abnormal risk warnings or significantly deviate from the overall distribution. These robustness markers are used to generate robustness weights or robustness substitutes for statistical inputs without altering the original value archive. The reasons for the robustness markers, trigger thresholds, and processing versions are written into the audit log to ensure the results are verifiable. Specifically, the system performs a robust label generation process for valid sampling point records according to the element dimension: First, the system reads the set of valid point concentration values ​​of the element and its associated detection metadata and anomaly label information to form a dataset to be diagnosed; Second, the system calculates the distribution diagnosis threshold for the dataset to be diagnosed and solidifies the diagnostic criteria for this round. The distribution diagnosis threshold includes at least a high quantile threshold for identifying abnormally high values, a low quantile threshold for identifying abnormally low values ​​when necessary, and a deviation threshold based on robust statistics for identifying significant deviations relative to the overall distribution; Third, the system performs metadata risk trigger judgment and numerical deviation trigger judgment for each record: When there are abnormal risk prompts in the detection metadata such as quality control failure, instrument drift risk, contamination risk, or abnormal circulation, it is recorded as a metadata risk trigger; When the concentration value of the record exceeds the high quantile threshold or the deviation of the relative robust statistics exceeds the preset deviation threshold, it is recorded as a numerical deviation trigger; The system determines whether to assign a robust label to the record based on the combination rules configured in the project, such as satisfying any one trigger or needing to satisfy two types of triggers simultaneously.For records assigned robustness tags, the system does not modify the original concentration values ​​archived in the ecological risk management system database. Instead, it generates corresponding robustness inputs in the data view used for spatial statistical calculations: First, it generates robustness weights to reduce the record's contribution to local statistical calculations by a preset ratio; second, it generates robustness substitute values ​​to truncate the record's value used in calculations to the upper diagnostic threshold or compress it according to preset robustness transformation rules; third, when a record is accompanied by strong anomaly risk warnings and has an extremely high degree of deviation, the system can mark it as a strongly robust state that is only archived and does not participate in statistics, but still retains its original value and records the reason for exclusion. The system writes the robustness tagging results into a robustness processing record table. The record table includes at least the point number, element identifier, robustness tagging status, trigger type metadata, trigger threshold reference, type of robustness strategy used, robustness weight or substitute value summary, processing version number, and generation timestamp. It also enables subsequent point aggregation results, hotspot cluster results, and core range results to be associated with this robustness processing record to support review and auditing.

[0025] Subsequently, the system breaks down the analysis objects by element. For each heavy metal element, the system extracts the point concentration value of the element and its associated land type, depth, and other attributes from the effective sampling point records, and generates a high-value indicator for spatial clustering verification. The high-value indicator is used to identify the high-value objects of interest in this round of clustering verification, avoiding indiscriminate verification at all points and reducing interpretability. The generation method of the high-value indicator is specified by the project configuration and maintains a stable caliber, supporting at least one of the following methods or a combination thereof: First, generating high-value indicators based on the exceedance status of management thresholds or background reference thresholds, where the threshold table can be stratified and fixed in the configuration according to land type, soil type, or sampling depth; Second, generating high-value indicators using quantile thresholds under the sample distribution of the region to adapt to the background differences of different regions, where the quantile thresholds can be calibrated offline by historical samples and updated periodically; Third, to reduce the accidental impact of single-point extreme values ​​on the formation of hotspot clusters, the system can introduce robust processing labels for abnormal high values, reducing their impact in statistical input, while still retaining the original value archive, robust processing reasons, and label records to ensure that the results are verifiable.

[0026] In the spatial relationship modeling phase, the system constructs spatial neighborhood relationships based on the coordinates of sampling points according to preset rules, and generates an adjacency list and neighborhood weights to constrain subsequent cluster significance tests to be conducted only within a reasonable range of spatial interactions. The spatial neighborhood relationships define the set of neighboring points for each sampling point, and the neighborhood weights characterize the contribution of neighboring points to the spatial statistics of that point, thus avoiding including points that are far away or unreasonably crossed into the same local statistical range. The construction of neighborhood relationships can employ one or more combinations of distance threshold neighborhoods, fixed nearest neighbor neighborhoods, or adaptive neighborhoods. The distance threshold can be determined and fixed during the installation and commissioning phase based on sampling density and terrain scale; the fixed nearest neighbor number can be configured offline based on historical project experience; and the adaptive neighborhood can set a maximum search radius while prioritizing the minimum nearest neighbor number to avoid crossing unreasonable spatial obstacles or forming excessively distant adjacencies. The system verifies neighborhood connectivity. When isolated points or neighborhood connectivity risks are detected, the system adjusts the neighborhoods according to a fixed completion strategy and records the differences, reasons, and parameters before and after the adjustment to ensure spatial statistical stability and result traceability.

[0027] During the hotspot identification phase, the system performs a spatial clustering significance test on each element under neighborhood relation constraints to determine whether high values ​​present statistical evidence of spatial clustering. This significance test can be implemented using existing hotspot analysis or local spatial autocorrelation analysis, such as Getis-Ord type hotspot statistics or local Moran type autocorrelation statistics. The system calculates local statistics for each sampling point and constructs a null hypothesis distribution through permutation tests. The permutation test involves randomly rearranging the observations multiple times while maintaining the spatial structure of the points, comparing the relative positions of the observed statistics within the null hypothesis distribution to obtain the significance level and significance grade. The significance test parameters include at least the number of permutations, the significance level, and the significance grade threshold. The number of permutations and the significance level can be fixed after offline calibration using historical samples. The significance grade threshold is used to distinguish between strongly significant, moderately significant, and insignificant results, so that the strength of evidence can be clearly defined during subsequent core range derivation. To ensure the reproducibility of significance test results, the system packages the algorithm type used in this round of significance testing (such as Getis-Ord or local Moran type), number of permutations, significance level, significance classification threshold, and necessary random seeds and rearrangement strategies into a test parameter set, and generates a unique test parameter version for the test parameter set. The test parameter version is used to associate the point cluster result table, hot spot cluster result table, and core range result table, so that any cluster type and significance level can be traced back to the corresponding test algorithm and parameter caliber. The system outputs the clustering type and significance level of each sampling point under the given element. The clustering type includes at least high-value clusters, low-value clusters, high-value outliers, low-value outliers, and insignificant points. The clustering type, significance level, and statistical evidence summary are written into the point clustering result table and associated with the point number, data version number, and test parameter version number to avoid unexplained differences in the same data under different parameters that would prevent verification. Based on the point clustering result table, the system further generates a point-level significance evidence diagnostic summary: it counts the number of sampling points that reach the preset significance level and have a high-value clustering type, and calculates their proportion of all valid sampling points; simultaneously, it uses the induced subgraph formed by high-value clusters under the current spatial neighborhood relationship as the object to count its spatial connectivity characteristics. The spatial connectivity characteristics include at least the number of connected components, the proportion of the largest connected component, the proportion of isolated points, or the average adjacency, used to characterize whether high-value significant points exhibit a clustered distribution; the point-level diagnostic summary is written into the diagnostic record and associated with the data version number, neighborhood parameter version number, and test parameter version number for subsequent hotspot cluster availability determination and method switching auditing.

[0028] During the hotspot cluster formation stage, the system connects and aggregates adjacent sampling points that belong to the same high-value cluster type and reach the preset significance level, based on the point aggregation result table, to form one or more hotspot clusters. To eliminate unstable small clusters caused by random fluctuations, the system sets minimum point count, minimum spatial scale, and intra-cluster consistency constraints for hotspot clusters. The maximum cluster point count characterizes the upper limit of the current hotspot cluster size or the dominant cluster size; the system counts the maximum cluster point count when multiple hotspot clusters are formed. The intra-cluster consistency satisfaction rate characterizes the degree of consistency in the high-value judgment criteria within the hotspot cluster. The system verifies the sampling points within the cluster according to preset consistency rules, such as consistency in exceeding standards, consistency in quantile standards, or consistency in robust label constraints. The consistency satisfaction rate is calculated as the proportion of points within the cluster that meet the consistency conditions to the total number of points in the cluster, or the proportion of hotspot clusters that meet the consistency constraints to the total number of hotspot clusters. The cluster spatial scale characterizes the spatial extension range of the hotspot cluster. The system can calculate this based on the cluster coverage area, the area of ​​the outer envelope boundary, the equivalent radius, or the intra-cluster point spacing statistics, and compare it with the minimum spatial scale threshold fixed in the project configuration to determine whether the hotspot cluster has reached the preset minimum spatial scale. The above cluster-level statistics and judgment results are written into the hotspot cluster result table and associated with the parameter version number for core boundary verifiable condition determination. The above constraints are fixed by the project configuration. Intra-cluster consistency constraints are used to limit the consistency of points within the same cluster under the high-value determination criterion, avoiding the mis-aggregation of scattered high-value points into a core. If multiple hotspot clusters exist, the system allows them to exist in parallel without forcibly merging them into a single core, thus adapting to multi-core distribution scenarios caused by multiple pollution sources. The system assigns a cluster number to each hotspot cluster, outputs the cluster point set, cluster-level statistical summary, and cluster-level confidence marker, and writes these to the hotspot cluster result table. Before deriving the contamination core boundary from the hotspot cluster, the system performs a core boundary verifiability condition check on the hotspot cluster results to ensure that the subsequent core boundary has an interpretable, traceable, and repeatable chain of evidence. The core boundary verifiability conditions include at least: the number of effective sampling points meets the minimum statistical requirements; the spatial neighborhood relationship forms a stable connected structure and the proportion of isolated points does not exceed the upper limit; the coverage rate of high-value clusters reaching the preset significance level reaches the lower limit; the hotspot cluster meets the minimum number of points and minimum spatial scale constraints; the hotspot cluster has a consistent lower limit for limited perturbations of key neighborhood parameters; and the core boundary generation rules and parameters are fixed in the project initialization and can be traced back from the boundary results to the supporting hotspot cluster point set and significance evidence. When any condition is not met, the system does not output a highly deterministic contamination core boundary and records the trigger item and threshold version number for auditing and verification.

[0029] In engineering scenarios for soil heavy metal ecological risk management, the management side typically needs to quickly identify concentrated areas potentially affected by pollution sources based on the measured elemental content of a limited number of discrete sampling points, under constraints such as no adjustment to the existing sampling point layout or difficulty in increasing sampling density. Furthermore, it needs to provide core zoning boundaries and zoning risk conclusions that can be used for remediation, verification, and rolling updates. Because sampling point data are spatially discrete, and background levels, land use attributes, and detection limits and quality control conclusions may differ across regions, directly delineating the remediation scope based solely on single-point exceedances or simple interpolation results can easily lead to problems such as uninterpretable boundaries, excessively long chains of evidence, or difficulty in verification. Therefore, this embodiment introduces the concepts of pollution core and pollution core range at the system output layer to form traceable, verifiable, and management-unit-adaptable core zoning results based on measured evidence from sampling points. In this embodiment, the pollution core refers to a high-value contiguous concentrated area of ​​influence supported by significant evidence of spatial clustering under a specific heavy metal element dimension. It is expressed as one or more spatial boundary polygons. Each pollution core boundary is associated with at least a set of hotspot cluster numbers, a summary of evidence from the hotspot cluster point set, a neighborhood scale summary, and the boundary generation rule and parameter version number, enabling the pollution core to be traced back from the boundary results to the supporting sampling points and significance test evidence. During the core range derivation stage, the system derives the pollution core range based on hotspot clusters. The core range is expressed as a spatial boundary, which can be one or more core boundary polygons. To ensure the core range is verifiable and the evidence chain is short, the core boundary generation follows traceability rules: the core boundary is derived from the hotspot cluster point set and its neighborhood scale. The neighborhood scale preferentially uses the typical point spacing statistics in the current neighborhood relationship to characterize the spatial influence range of the point clusters. Boundary generation can be based on the outer envelope boundary of the point clusters, the connected region boundary formed by the buffer fusion of point clusters, or the concave shell boundary generation rule. The boundary generation rule used is fixed and its version is recorded during project initialization. To achieve a more spatially continuous representation and ensure the reproducibility of core boundary results, the system packages the boundary generation rule types and key parameters used in this round of core boundary generation into a boundary parameter set, and generates a unique version of the boundary parameters for each set. The boundary parameter set includes at least: boundary generation method identifiers such as outer envelope, neighborhood scale reference caliber and values, buffer radius or fusion distance, concave shell parameters or shape constraint parameters, boundary normalization parameters such as gap filling threshold, small patch filtering threshold, and coordinate system / unit caliber. The system records the boundary parameter version in the core range result table, ensuring that any core boundary can be traced back to the corresponding boundary parameter set record and generation timestamp. The system can perform boundary normalization and small fragment filtering on sharp corners and gaps, but the normalization and filtering parameters must be recorded with the results, and the core boundary must be traceable back to the hotspot cluster set and saliency evidence supporting that boundary, avoiding the direct substitution of sampling evidence with uninterpretable surface model products.The system writes the core range boundary, the set of supporting hotspot cluster numbers, the boundary generation parameter summary, and the boundary parameter version of each element into the core range result table.

[0030] After obtaining the core area, the system spatially compares the core area boundary with the original coordinates of each sampling point. In addition to this spatial comparison, the system also performs a spatial intersection operation between the pollution core area boundary and the set of management unit boundaries fixed during project initialization, to obtain the set of management units within and outside the core area. Specifically, the system determines the intersection of the boundary polygon of each management unit with the boundary polygon of the pollution core area: when a management unit boundary intersects or contains any pollution core area boundary, the system marks the management unit as an internal management unit and can further calculate its coverage ratio with the pollution core area to express the degree of core impact; when a management unit boundary does not intersect with any pollution core area boundary, the system marks the management unit as an external management unit. The system associates the marking results of management units within and outside the core with the management unit identifier, core scope version number, and boundary parameter version number, writing them into the management unit overlay result table. This ensures that the core affiliation of any management unit can be traced back to the corresponding core boundary results and overlay criteria, triggering subsequent output of differentiated risk levels within and outside the core, assignment of evidence levels, and addition of uncertainty markers. A set of sampling points within the core scope is obtained, triggering differentiated analysis within and outside the core. If sampling points exist within the core scope, the system has the capability to extrapolate using measured points as core anchor evidence. The system defines the sampling points falling within the core scope as the core anchor point set and marks them as areas of sufficient evidence. Within the core area, the system directly generates risk levels and exceedance levels based on measured content and corresponding threshold criteria, marking the evidence source as measured in the results. Subsequently, the system performs zoning risk extension on the area surrounding the core by combining the core boundary and the spatial relationship of the neighborhood: the system constructs several risk extension zones outside the core boundary, the width of which is determined and fixed by both the neighborhood scale and the management unit scale; for each management unit within an extension zone, the system generates a hierarchy of evidence strength based on its spatial distance from the core anchor point, neighborhood connectivity, and the salience of evidence from surrounding points. Areas with close proximity and strong spatial connectivity are marked as areas with stronger evidence, while areas with greater distance or poor connectivity are marked as areas with weaker evidence. For all core extrapolation areas, the system simultaneously attaches uncertainty markers and verifiable evidence summaries to indicate that the risk conclusion for this area is an extended conclusion inferred from spatial evidence, rather than a direct measured conclusion from within the core area.

[0031] If no sampling points exist within the core area, the system considers it lacking in core anchoring evidence and does not directly output a high-certainty, high-risk conclusion. In this case, the system restricts the certainty output of high-risk cases to only be effective in the vicinity of the measured points or in areas where the evidence sufficiency condition is met, and uniformly adds a high uncertainty warning to the core candidate area and its extrapolated area. The system also generates follow-up action suggestions triggered by insufficient evidence. These suggestions include at least encrypted sampling, retesting, or key verification, and the action suggestions and triggering reasons are written into a suggestion list table. This ensures that without changing the existing sampling point layout constraints, the system can still output risk management results for the complete area and avoid misjudgments caused by insufficient evidence.

[0032] During the results aggregation phase, the system uniformly summarizes the core scope, core anchor point set, risk levels inside and outside the core, evidence strength levels, and uncertainty markers for each element, forming ecological risk management results oriented towards management units. Management units can be administrative grids, land parcel boundaries, or user-configured assessment grids. Their boundary system and resolution are fixed during project initialization to ensure comparability between different batches of results. The system outputs at least the following fields for each management unit: comprehensive risk level, dominant element, whether it is located within the core scope of any element, evidence level, uncertainty marker, summary of recommended treatment strategies, and recommended review actions. The fusion rules for the comprehensive risk level are fixed by configuration, and can adopt the worst-case scenario principle or the element weight fusion principle, recording the version of the adopted rule. Simultaneously, when the evidence level is low or the uncertainty is high, the system adds a verification mark to the output to prevent the management side from mistakenly believing that inferred risks have been confirmed by actual measurements.

[0033] Finally, the system uniformly writes the analysis process and results of this round into the ecological risk management system for storage, forming a complete and traceable closed loop. The stored content includes at least a point clustering result table, a hotspot cluster result table, a core area result table, a risk management result table, and an audit and traceability table. The audit and traceability table records the input data version, elimination and truncation replacement records, robustness marking records, neighborhood rule and test parameter versions, boundary generation rule versions, result generation timestamps, and task numbers, ensuring that the risk conclusions of any management unit can be traced back to the corresponding sampling point, the corresponding testing batch, the corresponding quality control evidence, and the corresponding spatial statistical configuration. Through the above process, this embodiment, based on the measured evidence from sampling points, outputs the clustering type and significance level under the constraint of neighborhood relationships, forming hotspot clusters. Then, the pollution core area is deduced from the significant clustering results. This approach relies on fewer assumptions, has a shorter evidence chain, and provides stronger result verifiability, naturally adapting to multi-core distribution scenarios caused by multiple pollution sources.

[0034] Example 2: Under the condition that other conditions remain unchanged in Example 1, when the method for deriving the core range of hotspot clusters in Example 1 fails to form a verifiable core range due to insufficient effective points, disconnected spatial relationships between sampling points, invalid spatial cluster significance, or instability of hotspot clusters, the system, without changing the existing sampling point layout, uses an interpolation method based on continuous surface estimation to complete the spatial information between sampling points, and then outputs a suspected pollution core range with uncertainty marking. Based on this, the system generates the overall ecological risk management results and subsequent action suggestions, and finally stores the process data and result data in the ecological risk management system database to support verification, rolling updates, and management and disposal.

[0035] During the method selection phase, the system prioritizes executing the hotspot cluster derivation core range processing flow and performs stability and verifiability diagnosis on the hotspot cluster results. When the system detects that a preset condition for non-core boundary verifiability is met, it switches to the interpolation surface completion flow of this embodiment. The system writes the triggering cause, diagnostic quantity summary, and method switching record into the audit information to explain the necessity and applicable boundaries of using the surface completion method, avoiding confusion between the two methods in terms of the strength of the evidence chain.

[0036] During the modeling admission and numerical usability phase, the system performs admission control on the raw data based on detection metadata and quality markers to determine the valid sampling point records for surface modeling. For records that fail quality control, exhibit abnormal flow, or have their credibility affected by anomaly marking, the system retains their original values ​​and the reasons for the anomalies for traceability, but marks them as unusable for this round of surface modeling and excludes them from the modeling set. Simultaneously, the system records the basis for removal and the evidence index to prevent anomalies from driving the surface model to generate false cores or unreasonably high-value surfaces. For data items below the detection limit, the system generates alternative values ​​for calculation according to preset truncation rules and retains the truncation marker, enabling the input to be used for continuous surface estimation, while reflecting the information loss caused by this type of data in the uncertainty output. The aforementioned truncation rules, admission control rules, and anomaly removal rules are either fixed by project configuration or fixed after offline calibration of historical samples and archived with the version number to ensure consistency and verifiability of analysis standards across similar projects and different batches.

[0037] During the element-based data splitting and modeling phase, the system constructs modeling subsets based on heavy metal elements. Each modeling subset includes at least the coordinates of the sampling points, the concentration value used for calculation, truncation markers, quality markers, and field attribute information associated with the threshold caliber, such as land type, soil type, and sampling depth, ensuring that subsequent risk assessments are consistent with management standards. To reduce the unreasonable drag of extreme values ​​on interpolation surfaces, the system can set robustness markers for points that significantly deviate from the overall distribution and also have abnormal risk warnings. This reduces their influence weight during modeling or allows them to participate in estimation only as reference points. However, the system still retains the original value of the point, the robustness marker, and the reason for the robustness marker to ensure that the conclusions are interpretable and traceable. The robustness strategy is fixed in the project configuration to avoid incomparable results due to arbitrary adjustments during task execution.

[0038] In this embodiment, interpolation surface formation refers to spatially completing the heavy metal concentration between sampling points using a continuous surface estimation method without changing the existing sampling point layout. It uses the coordinates and concentration values ​​of the effective sampling points after entry control as input, and the project-initialized output units—management grids, evaluation grids, and plot units—as spatial carriers. A predicted value surface for the element is generated on each output unit. Simultaneously, the prediction error and variance output by the algorithm, or the fluctuation statistics based on resampling, are combined to generate an uncertainty characterization. This uncertainty, along with spatial diagnostic quantities such as point density and distance to the nearest sampling point, forms an evidence level, which is used as a constraint condition for subsequent extraction of suspected core areas. In the continuous surface estimation stage, the surface formation output uses the user-defined management unit grid or evaluation grid as the spatial carrier unit. The management unit grid or evaluation grid is configured and its coordinate system and grid size are fixed by the user during project initialization to ensure that output results from different batches and tasks can be directly compared and superimposed under the same spatial reference. Specifically, the system must solidify at least the following coordinate system information: the identifier of the coordinate reference system used, such as the identifier code of the geographic coordinate system or the projected coordinate system; the coordinate units and axis definitions; the projection parameters and transformation methods when necessary; and the coordinate accuracy. When the original coordinates of the sampling points are inconsistent with the grid coordinate system, the system will uniformly transform the coordinates of the sampling points and the project boundary to the grid coordinate system according to the solidified transformation method, and write the transformation method and version number in the audit record to ensure that the subsequent calculation process can be reproduced. For regular grid output, the system further solidifies the grid type, such as square grid or hexagonal grid; the grid resolution, such as the cell side length or cell area diameter; the grid alignment reference, such as the grid origin coordinates; the row and column directions; the alignment rules; and the cell numbering rules, such as encoding by row and column number or globally unique encoding, and archives the above information as the output grid configuration version.First, the system reads the boundary of the monitoring area analysis range fixed during project initialization and converts it to a grid coordinate system. Second, based on the user-defined grid type, resolution, and alignment reference, the system generates a covering regular grid within the bounding rectangle of the analysis range boundary, forming a candidate grid cell set. Then, the system performs spatial pruning on the candidate grid cell set, retaining grid cells that intersect with or fall within the analysis range boundary, and removing grid cells that fall completely outside the analysis range. Furthermore, if the project configuration defines unestimateable masks or invalid areas such as water bodies, hardened surfaces, no-mining zones, soil-free areas, or user-specified exclusion zones, the system also converts the masks to a grid coordinate system and performs mask removal on the pruned grid cells, obtaining the final set of cells to be estimated. Finally, the system generates a unique cell identifier for each cell in the set of cells to be estimated and records its geometric boundary, center point coordinates, row and column index, management level identifier, and grid configuration version number, writing it into the output cell index table so that any subsequent predicted value, uncertainty value, and evidence level can be located, traced back, and verified through the cell identifier. The system determines the spatial carrying unit and resolution of the surface output based on the project configuration. The surface output unit is used to carry predicted values ​​and uncertain results, and can be a management grid unit, an administrative division unit, or an evaluation grid configured by the user. Its boundary and resolution are fixed during project initialization to ensure that results from different batches can be compared and superimposed. The system performs continuous surface estimation on the output unit based on valid sampling point records. In this embodiment, the element concentration prediction surface refers to the set of spatial prediction results obtained by continuously estimating the spatial distribution of a certain heavy metal element between sampling points on a unified coordinate system and fixed output units, such as management grids, evaluation grids, or management unit polygons. In engineering implementation, the element concentration prediction surface is preferably expressed in a cell-value manner. This means that a predicted concentration value is generated for each output cell in the set of cells to be estimated, creating a continuous spatial trend for the predicted values ​​of all output cells, thus forming the element concentration prediction surface. When the output cells are regular grids, the concentration prediction surface can be represented as a raster prediction surface; when the output cells are polygons of land parcels or administrative units, it can be represented as a vector cell attribute surface. Both are indexed and stored using a unified cell identifier. When generating the element concentration prediction surface, the system takes the coordinates of valid sampling points after access control and the element concentration values ​​used for calculation as input. It uses a fixed, continuous surface estimation algorithm and its parameter caliber to output the predicted values ​​for each cell. The predicted values ​​characterize the estimated level of element concentration at the cell location and do not directly replace the measured concentration values ​​of the sampling points. The system distinguishes between measured and predicted values ​​in the results to avoid evidence confusion.To ensure reproducibility and traceability of results, the system records and stores prediction surface version information for elemental concentration prediction surfaces. This version information includes at least: the selected algorithm type and version, model, parameter version number, output unit configuration version number, input data version number, coordinate system and transformation method version number, and generation timestamp. The predicted concentration value for each output unit is then associated with the aforementioned version information and written into the prediction surface result table. Continuous surface estimation is used to construct a continuous spatial variation surface between sampling points, thereby supplementing the estimates for unsampled locations. Continuous surface estimation can be implemented using existing methods such as inverse distance weighted interpolation, radial basis function interpolation, or kriging interpolation. The system permanently stores the selected algorithm type and its key parameter methods in the configuration and records the algorithm and parameter versions to ensure reproducibility of the conclusions. To avoid the surface modeling results merely presenting a smooth appearance without reliable evidence, the system simultaneously generates a modeling quality diagnostic summary. This summary includes at least: a point coverage uniformity diagnosis, used to characterize whether there are large blank areas in the spatial distribution of sampling points; a prediction range reasonableness diagnosis, used to identify anomalies where predicted values ​​exceed acceptable ranges; and an error statistical summary for cross-validation or leave-one-out validation, used to evaluate the model's fit on known points and its generalization level on unknown points. The error statistical summary serves as the basis for subsequent uncertainty grading and credibility constraints on suspected core ranges, and is archived along with the results.

[0039] In this embodiment, uncertainty constraints refer to a set of constraints used to limit the inclusion of interpolated surface inference results into the suspected contamination core range. These constraints prevent high predicted values ​​from being directly solidified as deterministic cores in areas with insufficient evidence. After generating the uncertainty characterization, the system generates uncertainty constraint judgment results for each output unit based on constraint thresholds that are either initially solidified during project initialization or solidified after offline calibration using historical samples. The constraint thresholds include at least an upper uncertainty threshold, a lower evidence level threshold, and optional spatial diagnostic constraint thresholds such as an upper distance limit to the nearest sampling point, a lower point density limit, or a lower neighborhood coverage limit. The system marks output units that meet the above constraints as valid constraint units that can be included in the suspected core extraction and forms an uncertainty constraint mask. For output units that do not meet the constraints but have high predicted values, the system marks them as high-risk candidates with insufficient evidence and adds a strong uncertainty warning. The uncertainty constraint thresholds and judgment criteria generate uncertainty constraint version numbers and write them into the audit log to ensure that the subsequent suspected core range extraction process is verifiable. In the uncertainty assessment and evidence grading stage, the system generates uncertainty characterizations for continuous surface estimation results and maps the uncertainty to levels of evidence strength. Uncertainty characterization is used to quantify the inference risk brought about by model completion, avoiding the misinterpretation of surface predictions as empirical evidence. For surface algorithms that can directly output prediction variance or error estimates, the system uses their output as the basis for uncertainty. For surface algorithms that do not directly output uncertainty, the system uses a resampling error estimation method to generate uncertainty characterization. This resampling error estimation method includes multiple subset extractions or perturbation modeling of the sampling point set, and statistical analysis of the fluctuation degree of the predicted values ​​of each output unit to characterize prediction stability. The system further combines the uncertainty characterization with spatial diagnostic quantities such as point density, distance to the nearest sampling point, and neighborhood coverage to form an evidence level field, and outputs the evidence level on each output unit to distinguish between areas with sufficient evidence and areas with insufficient evidence, ensuring that subsequent risk levels in areas with weak evidence are not mistakenly considered as high-certainty conclusions.

[0040] During the suspected core area extraction stage, the system uses a suspected expression for the pollution core area based on the surface formation results, that is, it outputs the suspected core area without directly equating it with the measured core. For each element, the system extracts high-value regions on the surface formation results according to a preset high-value judgment criterion and forms suspected core candidate areas. The high-value judgment criterion is fixed by configuration and supports at least one of the following methods: determining the exceeding area based on the control threshold or background reference threshold, or determining the relatively high-value area based on the high quantile of the regional distribution. At the same time, the system uses uncertainty characterization as a constraint condition in suspected core extraction. Only when the high-value area meets the preset evidence level requirements or meets the uncertainty upper limit constraint is it included in the suspected core area; for areas that do not meet the evidence constraints but still show high values, the system marks them as high-risk candidate areas and adds a higher uncertainty prompt to avoid directly fixing the model product as a deterministic core. The system performs boundary regularization and small patch filtering on the extracted suspected core areas. Boundary regularization is used to eliminate jagged boundaries caused by meshing output, and small patch filtering is used to remove fragmented areas that are too small and lack supporting evidence. The regularization and filtering rules are fixed by configuration, and the suspected core boundary, extraction caliber version number, and uncertainty constraint version number are written into the results table to ensure that they can be verified.

[0041] In the joint analysis phase of the core area and sampling points, the system spatially compares the boundary of the suspected core area with the original coordinates of each sampling point, outputting the set of sampling points and anchoring status within the suspected core area, and adopting a differentiated output strategy accordingly. If sampling points exist within the suspected core area, the system uses these sampling points as anchoring evidence for the suspected core area. Anchoring evidence means that the area inferred from the surface can be supported by measured points in a local area, and therefore its risk level can be given at a higher confidence level. The system distinguishes between measured evidence and area inference evidence in its output: the risk level within the suspected core area is primarily confirmed by the measured content and threshold caliber of the anchored sampling points, while the surface results are used to supplement spatial continuity and boundary morphology; the risk extension in the area near the boundary of the suspected core area is jointly determined based on the gradient of the surface prediction value, the distance to the sampling point, and the evidence level, and the uncertainty level and evidence level summary are output simultaneously. If no sampling points exist within the suspected core area, the system will not treat that area as a definitive core, but rather as an unanchored core candidate area. The system restricts the definitive output of high-risk areas to only take effect in the neighborhood of the measured points or in areas where the evidence level meets the conditions. It also uniformly adds a high uncertainty warning to the core candidate area and its extrapolated area, and generates follow-up action suggestions triggered by insufficient evidence, including at least encrypted sampling, retesting, or key verification. The suggestions and triggering reasons are written into the suggestion list to guide the next round of sampling optimization without changing the existing sampling point layout constraints of this round.

[0042] During the ecological risk management results generation phase, the system summarizes the suspected core range of each element, the output unit-level predicted value, the evidence level, and the uncertainty marker, forming ecological risk management results oriented towards the management unit. The system outputs at least the comprehensive risk level, dominant element, whether it falls within the suspected core range of any element, evidence level, uncertainty alert, and a summary of remediation recommendations for each management unit. The fusion rules for the comprehensive risk level are fixed by configuration and can adopt the worst-case scenario principle or the weighted fusion principle. However, when the evidence level is low or the uncertainty is high, the system adds a verification mark to the output and adopts a strategy of prioritizing verification and encrypted sampling for remediation recommendations to prevent the management side from mistakenly identifying inferred risks as risks that have been confirmed by actual measurements.

[0043] Finally, the system stores the interpolation process and results of this round in the ecological risk management system database to form a complete and traceable closed loop. The database contents include at least: surface model configuration and parameter version, a list of valid sampling points and removal records, truncation substitution and truncation marker records, robustness processing markers and reason records, cross-validation or error statistics summary, output unit-level predicted surface results, uncertain surface results, suspected core area boundary results, management unit-level risk management results, and audit traceability records. Through the above process, this embodiment, when hotspot clusters cannot form a stable chain of evidence, uses quality gating, truncation substitution, and continuous surface estimation to complete the spatial information between sampling points, and uses uncertainty and evidence level to constrain the output of suspected core areas and risk conclusions. Thus, even with insufficient evidence, it can still output complete regional ecological risk management results, while maintaining traceability, hierarchical interpretation, and guidance for subsequent sampling and verification.

[0044] As can be seen from the above, Example 1 embodies a spatial clustering statistical derivation method based on measured evidence from sampling points. Under the constraint of preset spatial neighborhood relationships, it performs significance testing on high-value clusters of each element to form hotspot clusters, and directly derives the core pollution range from these hotspot clusters. This is a core range identification path with a short and verifiable evidence chain. Example 2 embodies a spatial information completion method based on continuous surface estimation. Based on quality gating and truncated substitution, it interpolates between sampling points to form surfaces and outputs uncertainty and evidence classification. Under uncertainty constraints, it extracts the suspected core range, representing a fallback identification path under conditions of insufficient evidence.

[0045] like Figure 4The diagram illustrates the pollution core analysis and identification path switching process provided in this application. The system prioritizes executing the core range derivation path based on hotspot clusters in Embodiment 1. When this path fails to generate a verifiable conclusion, it switches to the fallback path based on interpolation-formed surfaces in Embodiment 2. Embodiment 1 uses actual observations of sampling points as the primary source of evidence. Under the constraint of preset spatial neighborhood relationships, it performs spatial cluster significance tests on the data of each element, outputs the clustering type and significance level of each sampling point, and forms one or more hotspot clusters accordingly. The pollution core range is then derived from the significant clustering results. Since this path mainly relies on actual measurement points and their spatial statistical evidence, it requires fewer prior assumptions, has a shorter evidence chain, and the conclusions are more verifiable. It can also naturally generate multi-core distribution identification results in the presence of multiple pollution sources.

[0046] When the path in Example 1 fails to form a hotspot cluster that satisfies stability constraints due to insufficient effective sampling points, difficulty in forming a stable connectivity structure in spatial neighborhood relationships, or failure to reach the preset significance level in the cluster significance test, the system switches to the fallback path based on interpolation-based surface formation in Example 2. Example 2, based on quality gating and truncation of data below the detection limit, completes the spatial information between sampling points through continuous surface estimation and simultaneously outputs uncertainty characterization and evidence classification. Under uncertainty constraints, it extracts the suspected contamination core range marked with uncertainty as the fallback result. Through the above switching strategy, the system can prioritize providing verifiable core range conclusions supported by actual measurements when evidence is sufficient, and avoid directly equating the surface completion result with sampling evidence when evidence is insufficient. This achieves robust output of full-domain risk management results without changing the existing sampling point layout.

[0047] For example, when sufficient valid sampling points and connected neighborhoods are met, the spatial clustering statistical derivation method based on measured evidence from sampling points is preferentially adopted. Taking a certain area to be monitored as an example, after sampling points are deployed according to the established sampling plan, the number of valid sampling point records for cadmium meets the minimum point requirement, and the sampling points are relatively evenly distributed in space. The neighborhood relationships constructed based on distance threshold neighborhoods or fixed nearest neighbor neighborhoods can form a stable connected structure. After the system completes the entry control and truncation processing, it generates a high-value indication for cadmium and performs a spatial clustering significance test under the constraint of neighborhood relationships. The results show that multiple sampling points exhibit high-value clustering and the significance level reaches the preset threshold. Furthermore, the adjacent and significant high-value clusters are connected and aggregated to form two hotspot clusters, corresponding to the two suspected pollution source influence areas, respectively. Since hotspot clusters are directly obtained from statistical evidence of spatially clustered high values ​​measured at sampling points, and their core range is derived from the results of significant clustering, the system can output two pollution core ranges with a relatively short chain of evidence. It then labels the sampling points within the core range as a set of core anchor points, and subsequently performs risk expansion by zoning inside and outside the core and outputs tiered evidence strength. In this scenario, the system does not need to continuously complete the data between sampling points to obtain stable hotspot clusters; therefore, Implementation Example 1 is preferred to achieve stronger verifiability and natural adaptability to the distribution of multiple pollution sources and multiple cores.

[0048] For example, when insufficient valid points or disconnected neighborhoods lead to unstable hotspot clusters, a spatial information completion method based on continuous surface estimation is adopted. Taking another monitored area as an example, due to terrain obstruction or on-site accessibility, sampling points exhibit obvious blank zones or are divided into several discrete areas in space. Simultaneously, the detection results for mercury contain many data items below the detection limit, and some points are excluded due to quality control failures, resulting in the number of valid sampling points participating in the statistics being close to or below the minimum point requirement. When the system attempts to construct neighborhood relationships according to Example 1, it finds that neighborhood relationships are not easy to form stable connected structures, with isolated points and broken areas. Furthermore, when performing spatial cluster significance testing, the significance results are highly sensitive to neighborhood parameters, failing to form stable hotspot clusters that satisfy the minimum number of points and minimum spatial scale constraints. Alternatively, although small clusters are formed, the consistency within the clusters is insufficient, and the verification is weak. At this point, the system triggers a failure to meet the core boundary verification condition and transitions to the second implementation process: First, for data below the detection limit, substitute values ​​are generated according to the truncation rule, and the truncation mark is retained. Then, continuous surface estimation is performed on the fixed management grid resolution to obtain the predicted value surface of each grid cell. Simultaneously, uncertainty characterization and evidence level are generated through cross-validation error statistics and diagnostic quantities such as point density and distance to the nearest sampling point. To facilitate subsequent boundary generation and determination of risk expansion band width, after constructing spatial neighborhood relationships, the system obtains neighborhood scale statistics based on adjacency list statistics. Neighborhood scale statistics are used to characterize the typical spatial spacing of sampling points under the current neighborhood relationship. Preferably, it is the median distance from each sampling point to its nearest neighbor, or the quantile statistics of the adjacent edge length distribution. The system writes the neighborhood scale statistics, along with neighborhood rule type, key parameters such as distance threshold, number of nearest neighbors, and maximum search radius, into the neighborhood configuration record and associates them with the version number to ensure that subsequent results are reproducible. When extracting suspected core areas, the system links high-value regions with uncertainty constraints, including only contiguous high-value areas that meet the evidence level criteria within the suspected core area. Areas with insufficient evidence but high predicted values ​​are marked as high-risk candidate areas and given strong uncertainty warnings. Simultaneously, the system outputs action suggestions for encrypted sampling or retesting. In this scenario, the hotspot cluster method cannot form a stable chain of evidence. Therefore, the system employs Implementation Example 2 as a fallback to ensure that comprehensive risk management results can still be output even with insufficient evidence. Furthermore, uncertainty and evidence grading prevent surface-level products from being mistakenly treated as experimental conclusions.

[0049] In the switching control between Embodiment 1 and Embodiment 2, before performing the core range deduction of hot spot clusters for each heavy metal element, the system first generates a verifiable diagnostic quantity based on the effective sampling point set after the entry control, and determines whether to continue executing Method 1 or switch to Method 2 accordingly. The diagnostic metrics include at least: the number of valid sampling points, used to characterize whether the scale of points participating in spatial statistics meets the minimum statistical requirements; neighborhood connectivity diagnostic results, used to characterize whether the spatial neighborhood relationships constructed according to preset neighborhood rules form a stable connected structure, wherein the connectivity diagnostic results include at least the number of connected components and the proportion of isolated points, wherein the number of connected components is used to determine whether the points are divided into too many discrete regions, and the proportion of isolated points is used to determine whether the proportion of points lacking neighbor support is too high; significance evidence coverage rate, used to characterize whether the proportion of high-value cluster points that pass the cluster significance test and reach the preset significance level among all valid points meets the lower limit requirement; hotspot cluster availability constraint results, used to characterize whether the hotspot clusters meet the minimum number of points constraint and minimum spatial scale constraint after aggregation, wherein the minimum spatial scale constraint can be characterized by the cluster coverage area, equivalent scale, or intra-cluster point spacing statistics; and hotspot cluster stability diagnostic results, used to characterize the sensitivity of hotspot clusters to neighborhood parameter perturbations, wherein stability diagnosis can be obtained by performing a finite number of perturbation calculations on the key parameters of neighborhood construction without changing the input data and comparing the consistency of hotspot clusters, wherein consistency can be characterized by the degree of cluster-level overlap or the degree of core boundary overlap. Furthermore, the neighborhood connectivity diagnosis also includes connectivity structure consistency under perturbations of different neighborhood construction parameters, which characterizes the sensitivity of spatial neighborhood relationships to key neighborhood construction parameters. Specifically, without changing the set of effective sampling points, the system performs a finite number of perturbations on the key neighborhood construction parameters and constructs corresponding adjacency lists for each, and calculates the connected component partitions and isolated point sets for each perturbation result. Based on the degree of overlap of connected component partitions, maximum matching coverage, or consistency score among the perturbation results, the system obtains a connectivity structure consistency diagnostic value to determine whether neighborhood relationships undergo structural changes due to minor parameter variations. When the connectivity structure consistency is below a preset lower limit or the number of connected components / isolated point ratio fluctuates drastically under perturbation, the system determines that the neighborhood connectivity structure is unstable, records it as an unusable trigger item, and uses it for method switching auditing. The system compares the above diagnostic values ​​with the fixed thresholds. The fixed thresholds include at least the minimum number of valid points threshold, the maximum number of connected components threshold, the maximum proportion of isolated points threshold, the lower limit threshold for the proportion of significant high-value points, the minimum number of points in hotspot clusters threshold, the minimum spatial scale threshold for hotspot clusters, and the lower limit threshold for consistency of hotspot clusters. Among them, the neighborhood rules and their key parameters, the minimum number of points in hotspot clusters threshold, and the minimum spatial scale threshold are fixed by the initial configuration of the project. The significance test related thresholds and the consistency lower limit threshold can be fixed by offline calibration of historical samples. The neighborhood scale thresholds can be determined and fixed during the installation and commissioning stage in combination with the sampling density and terrain scale.When any diagnostic quantity triggers a condition that does not meet the core boundary verifiability criteria, the system determines that Method 1 cannot form a hotspot cluster and core range that meet the requirements of stability and verifiability. It then automatically switches to Method 2 and writes the trigger item, threshold version number, and diagnostic summary into the audit log. When all diagnostic quantities meet the usability criteria, the system continues to execute Method 1 and outputs the contaminated core range derived from significant clustered evidence. This achieves a hierarchical output mechanism that prioritizes empirically supported conclusions when evidence is sufficient and uses surface completion with uncertainty constraints when evidence is insufficient.

[0050] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope and intent of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. An ecological risk management system based on soil heavy metal data, characterized in that, include:: The data acquisition module is used to bind and store the sampling point number, coordinates, field attributes, laboratory element content and detection limit and abnormal metadata into the database, and perform access control, truncation replacement and robust labeling to form a traceable original dataset. The pollution core analysis module is used to construct the spatial neighborhood relationship of sampling points based on heavy metal elements. It performs a local spatial cluster significance test based on the neighborhood relationship analysis and connects and aggregates high-value points that meet the significance level to form hot spot clusters. Then, the pollution core is derived from the hot spot clusters. At the same time, a method discrimination switching unit is set to diagnose the number of effective points, neighborhood connectivity and stability, significance results and cluster stability. When the diagnostic results do not meet the core boundary verification condition, it automatically switches to the continuous surface estimation process, interpolates the sampling points to form surfaces and generates uncertainty characterization simultaneously, and extracts the pollution core under uncertainty constraints. The risk management module is used to overlay the pollution core with sampling points and management units, output the risk level, evidence level and uncertainty marker inside and outside the core, and uniformly store the process parameter version and results in the ecological risk management system database.

2. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The process of performing admission control, truncation replacement, and robust labeling to form a traceable original dataset is as follows: Based on the detection limit and abnormal metadata, determine the set of valid records to participate in the calculation, and retain the original value and the reason for the abnormality for records that have not passed the access control, and mark them as unusable for calculation; For data items below the detection limit, a replacement value for calculation is generated according to a preset truncation replacement rule, and the truncation mark is retained; The weight of the robustly labeled points is reduced in subsequent statistics or modeling while retaining the reason for the robust labeling, thereby forming the traceable original dataset.

3. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The process of performing local spatial cluster significance testing based on neighborhood relationship analysis and connecting and aggregating high-value points that meet the significance level to form hotspot clusters is as follows: When the pollution core analysis module performs the derivation of the pollution core from the hotspot cluster, it splits the traceable original dataset according to the heavy metal elements, generates high value point judgment results based on threshold or quantile, and constructs the spatial neighborhood relationship and neighborhood weight of the sampling points accordingly. Under the constraint of spatial neighborhood relationship of sampling points, perform local spatial clustering significance test on high value points and use permutation test to obtain significance level, and output the clustering type and significance level of sampling points; Connect adjacent sampling points that belong to the same high-value cluster and reach the preset significance level to form hotspot clusters, and apply minimum point number, minimum spatial scale and intra-cluster consistency constraints to the hotspot clusters. The pollution core boundary is generated based on the hot spot cluster point set and neighborhood-scale statistics, and the pollution core boundary is associated and stored with the hot spot cluster number set and test parameter version that support the pollution core boundary.

4. The ecological risk management system based on soil heavy metal data as described in claim 3, characterized in that: The process of deriving the contamination core from hotspot clusters includes generating the contamination core boundary according to the outer envelope boundary, the point cluster buffer fusion boundary, or the concave shell boundary, performing boundary regularization and small patch filtering on the contamination core boundary, and associating and storing the contamination core boundary with the set of hotspot cluster numbers supporting the contamination core boundary and the boundary parameter version.

5. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The diagnosis of effective point count, neighborhood connectivity and stability, significance results, and cluster stability is specifically as follows: The number of valid sampling points recorded after the access control is counted to obtain the number of valid points, and then compared with the preset minimum number of points threshold. Based on the effective sampling point coordinates, candidate sampling point spatial neighborhood relationships are generated according to preset neighborhood construction rules. Connectivity verification and outlier identification are performed on the candidate sampling point spatial neighborhood relationships, and neighborhood stability diagnostic quantity is calculated. The neighborhood stability diagnostic quantity includes at least the average number of nearest neighbors, the proportion of outliers, and the consistency of connectivity structure under different neighborhood construction parameter perturbations. Under the constraint of the spatial neighborhood relationship of the sampling points, the local spatial cluster significance test is performed, and the number, proportion and spatial connectivity characteristics of high-value cluster points that reach the preset significance level are counted as the diagnostic quantity of significance results. After connecting and aggregating high-value points that meet the significance level to form hotspot clusters, the hotspot cluster stability diagnostic quantity is calculated. The hotspot cluster stability diagnostic quantity includes at least the number of hotspot clusters, the maximum number of cluster points, the intra-cluster consistency satisfaction rate, and whether the cluster space scale reaches the preset minimum space scale. When any diagnostic quantity does not meet the preset core boundary verifiable condition, an unusable trigger conclusion is generated and a combination of trigger reasons is output to drive the automatic switch to the continuous surface estimation process. The number of effective points, the neighborhood stability diagnostic quantity, the significance result diagnostic quantity, the hot spot cluster stability diagnostic quantity, and the corresponding test parameter version or boundary parameter version are written into the audit record.

6. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The continuous surface estimation process specifically includes: Based on the traceable original dataset, a set of valid points for participating in the surface formation is determined, and the coordinates of the valid points and element concentration values ​​are mapped to the user-defined management unit grid or evaluation grid coordinate system. Read the boundary and resolution of the management unit grid or evaluation grid and generate a set of cells to be estimated; On the set of cells to be estimated, inverse distance weighted interpolation, radial basis function interpolation, or kriging interpolation are selected according to the configuration. For each cell to be estimated, the element concentration prediction value is calculated based on its distance relationship with the effective point or the spatial correlation structure to form the element concentration prediction surface. A modeling quality diagnostic summary generation process is performed on the element concentration prediction surface. The modeling quality diagnostic summary includes at least a diagnosis of the coverage uniformity of effective points on the management unit grid or evaluation grid, a diagnosis of the reasonableness of the predicted value range, and an error statistical summary obtained by performing leave-one-out cross-validation on effective points. The error statistical summary includes the algorithm type, key parameters, diagnostic summary, and corresponding test parameter version or boundary parameter version stored together.

7. The ecological risk management system based on soil heavy metal data as described in claim 6, characterized in that: The synchronous generation of uncertainty characterization includes: when the continuous surface estimation process can output the prediction variance, the prediction variance is directly used as the uncertainty characterization; When the continuous surface estimation process does not output the prediction variance, resampling modeling is used to generate the prediction fluctuation statistic as an uncertainty characterization. The step of extracting contamination cores under uncertainty constraints includes identifying high-value regions as contamination cores only when they meet the preset evidence level or uncertainty upper limit constraints; otherwise, they are marked as unanchored candidate regions and encrypted sampling or retesting suggestions are output.

8. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The superposition of the pollution core, sampling points, and management unit includes: Spatial intersection calculations are performed between the pollution core and the management unit boundary to obtain the set of management units within the core and the set of management units outside the core. Sampling points are then mapped to the corresponding management units according to their location relationships to form a management unit-sampling point association table. At least one layer of risk expansion zone is generated outside the pollution core based on the neighborhood scale statistics of the spatial neighborhood relationship of the sampling points. The risk expansion zone is obtained by buffering and layering the outer side of the pollution core boundary. For each risk expansion zone, its number, bandwidth parameter version and spatial coverage relationship with the management unit are recorded. The overlay results, extended band results, and the association table of the sampling points of the management unit are written into the intermediate result table for subsequent risk classification.

9. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The output of risk levels, evidence levels, and uncertainty markers inside and outside the core includes: for management units inside the core, extracting the measured values ​​of the associated sampling point elements and determining the risk level according to the threshold caliber, while marking the evidence level as measured evidence and writing it into the corresponding abnormal metadata summary and truncation marker. For core external management units, based on their respective risk extension zones, distances to the pollution core boundary, neighborhood connectivity with the core anchor point, and uncertainty characterization, an evidence level is generated and constraints are imposed on the risk level output according to the evidence level. The imposition of constraints includes at least reducing the deterministic expression level and outputting a verification mark when the evidence level is low or the uncertainty mark is high. The risk level, evidence level, uncertainty marker, and summary of the calculation basis for each management unit will be included in the database to support review and rolling updates.

10. The ecological risk management system based on soil heavy metal data as described in claim 1, characterized in that: The unified storage of process parameter versions and results includes at least writing the following into the ecological risk management system database: entry control and exclusion criteria, truncation substitution and truncation markers, robustness markers and robustness reasons, sampling point spatial neighborhood relationship parameter versions, local spatial cluster significance test parameter versions, hotspot cluster and pollution core boundary results, continuous surface estimation process algorithm and parameter versions, uncertainty characterization results, and management unit-level risk management results. The task number and timestamp are then associated to support rolling updates and audit traceability.