Intelligent auxiliary identification method and system for high-risk pollution area of industrial gathering area site

By integrating multi-source data and using dynamic model simulation, combined with graph neural networks and risk assessment, the problem of low accuracy and poor adaptability in identifying high-risk pollution areas in industrial clusters has been solved, enabling the provision of efficient and accurate pollution area identification and remediation solutions.

CN121352484APending Publication Date: 2026-01-16BEIJING MUNICIPAL RES INST OF ENVIRONMENT PROTECTION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511487152.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy, poor adaptability, weak reusability, and disconnect from remediation needs in identifying high-risk pollution areas in industrial clusters, making it difficult to achieve accurate, efficient, and dynamic pollution area identification.

Method used

By employing multi-source data integration, dynamic model simulation, and goal-oriented analysis, and through spatial grid construction, multi-source data integration, three-dimensional correlation embedding graph neural network simulation, and risk assessment, combined with dynamic topology adjustment and cross-park transfer learning, we can achieve accurate identification of high-risk pollution areas and provide remediation solutions.

Benefits of technology

It improved recognition accuracy and efficiency, enhanced the model's adaptability to complex scenarios, provided clear governance directions, and enabled rapid deployment across parks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121352484A_ABST
    Figure CN121352484A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent auxiliary identification method and system for a high-risk pollution area of an industrial gathering area site, belongs to the field of pollution identification of industrial gathering areas, and is used for solving the problems of low identification precision, poor adaptability and difficulty in cross-area reuse of the high-risk pollution area of the industrial gathering area in related technologies. According to the method, through the steps of spatial gridding construction, multi-source data integration, three-dimensional binding graph neural network simulation, risk assessment and the like, digital twinborn rehearsal, target-oriented attribution and federal transfer learning are combined; the system comprises a data acquisition subsystem and a server, can realize accurate identification of a high-risk area, dynamic scene adaptation and cross-area rapid deployment, and provides targeted support for pollution abatement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pollution identification in industrial clusters, and more particularly to an intelligent auxiliary identification method and system for high-risk pollution areas in industrial clusters. Background Technology

[0002] Industrial clusters, as concentrated areas of industrial production, contain various types of pollution facilities (such as storage tanks and wastewater ponds). Their site pollution is characterized by its concealment and complexity. Accurate identification of high-risk pollution areas is a key link in ensuring regional environmental safety and human health, and has become a major focus in the field of environmental governance.

[0003] In existing technologies, the identification of high-risk pollution areas in industrial clusters often relies on manual on-site sampling combined with laboratory analysis, or on simple models driven by a single data source (such as traditional statistical models or basic machine learning models). Manual sampling methods have limited coverage and are time-consuming, making it difficult to achieve rapid identification over a large area. Single models often only utilize monitoring data, failing to fully integrate multi-source information such as spatial, facility, and historical data, and do not consider the migration patterns of pollutants across multiple media, resulting in low identification accuracy.

[0004] Existing technologies still have significant shortcomings: First, the fixed grid division cannot adapt to the refined identification needs of high-risk and sensitive areas, resulting in large deviations in the location of small-scale pollution sources; second, the risk assessment weight determination method is singular, using subjective or objective weights alone, making it difficult to reflect the comprehensive impact of multiple factors; third, the fixed model topology cannot dynamically respond to changes in environmental parameters and human intervention, resulting in poor adaptability; fourth, there is a lack of goal-oriented attribution of pollution causes, leading to a disconnect between identification results and governance needs; fifth, there is poor reusability across industrial parks, requiring new parks to train models from scratch, resulting in low deployment efficiency; at the same time, there is a lack of hardware systems adapted to the identification methods, making it difficult to achieve linkage between data collection and identification processes. These shortcomings prevent existing technologies from meeting the actual needs of accurate, efficient, and dynamic identification of high-risk pollution areas in industrial clusters. Summary of the Invention

[0005] This application provides an intelligent auxiliary identification method and system for high-risk pollution areas in industrial clusters. It can solve the problems of low identification accuracy, poor adaptability, weak reusability, and disconnection from governance needs of existing technologies by integrating multi-source data, dynamic model simulation, and target-oriented analysis.

[0006] Firstly, this application provides an intelligent auxiliary identification method for high-risk pollution zones in industrial clusters. The process includes the following steps: A spatial grid construction step, where the industrial cluster is divided into standard grids based on a preset coordinate system. Each grid is assigned a unique identifier and its spatial boundary information is recorded. The grids at the boundaries of high-risk facility areas and environmentally sensitive areas are further refined. A multi-source data integration step, collecting spatial basic data, enterprise facility layout data, pollution monitoring data, and geological and hydrological data of the industrial cluster. After standardizing the collected data, it is mapped to the corresponding grids using spatial overlay or interpolation methods, forming a relationship between the grids and the data. A three-dimensional association embedding graph neural network simulation step, constructing a graph neural network model with grids as nodes and pollutant migration relationships between grids as edges. Three-dimensional features of grid identifiers, pollutant facility types and proportions, and pollutant concentrations are embedded into the model nodes. The model is trained to simulate the spatiotemporal diffusion of pollutants in the soil-groundwater medium. A risk assessment step, constructing an assessment system based on four preset indicators: pollution source intensity, pollution migration channels, environmental vulnerability, and historical records. The comprehensive risk value of each grid is calculated, and high, medium, and low risk levels are divided according to preset risk value ranges. The results of identifying high-risk pollution areas are output.

[0007] By adopting the above technical solutions, a unified spatial benchmark for pollution identification is achieved through spatial grid construction, multi-source data integration solves the problem of data heterogeneity, three-dimensional correlation embedding graph neural network simulation fully integrates spatial, facility, and pollutant information and reflects migration patterns, risk assessment integrates multiple factors to achieve classification, and four steps work together to achieve accurate identification of high-risk pollution areas. Compared with manual sampling, efficiency is significantly improved, and compared with a single model, accuracy is greatly improved.

[0008] Furthermore, in the spatial gridding construction step, the grid refinement process adopts a partitioned hierarchical control method. In general industrial areas, the standard grid size is maintained and the grid boundary coordinate deviation is controlled within a preset range. The grid size of high-risk concentrated areas is adjusted to a second grid size smaller than the standard grid size to adapt to the needs of small-scale pollution source positioning. The grid size of the boundary of environmentally sensitive areas is further adjusted according to the protection level of the sensitive areas to adapt to the accuracy requirements of pollutant diffusion simulation boundary.

[0009] By adopting the above technical solutions, the hierarchical grid division avoids a "one-size-fits-all" approach. While ensuring global computational efficiency, it improves the spatial resolution of high-risk and sensitive areas, makes the location of small-scale pollution sources more accurate, and reduces the simulation deviation of sensitive area boundaries, thereby further improving the identification accuracy.

[0010] Furthermore, in the risk assessment step, the weights of the four types of indicators are determined by a combination of subjective and objective weight determination methods. The secondary indicators under each indicator are quantitatively graded according to a preset scoring standard. The comprehensive risk value is calculated by weighted summation of the weights of each indicator and the corresponding scores.

[0011] By adopting the above technical solutions, the combination of subjective and objective weights balances expert experience and data objectivity, avoids bias caused by a single weight, and makes risk values ​​more comparable through quantitative grading and weighted calculation, making risk level classification more scientific and improving the reliability of risk assessment.

[0012] Furthermore, the three-dimensional association embedding graph neural network simulation step also includes a dynamic topology adjustment process. Every preset time step, it detects whether the environmental parameters corresponding to the grid have changed by a preset amount or whether there is a human intervention instruction. If any condition is met, topology adjustment is triggered. Based on the detection results, the connection weights between grids are adjusted and the adjacency connections of the pollution diffusion front grid are expanded from the original 4 directions to 8 directions.

[0013] By adopting the above technical solutions, dynamic topology adjustment enables the model to respond in real time to environmental changes (such as sudden changes in soil permeability coefficient) and human intervention, avoiding simulation lag caused by fixed topology. The 8-way connectivity of the diffusion front is more in line with the actual diffusion direction, improving the model's adaptability to dynamic scenarios and simulation accuracy.

[0014] Furthermore, the three-dimensional associated embedded graph neural network simulation step also includes a multi-media coupling simulation process, which calculates the pollutant migration flux from the atmosphere to the soil and from the soil to the groundwater using a preset migration flux calculation method, uses the migration flux data as new feature embedded graph neural network model nodes, updates the node feature vectors and retrains the model to realize the multi-media coupling of air-soil-groundwater pollutant migration simulation.

[0015] By adopting the above technical solutions, multi-media coupling simulation makes up for the limitations of single-media simulation, fully reflects the cross-media migration path of pollutants, and the migration flux embedding makes the model more in line with the actual pollution diffusion law, further improving the realism and accuracy of pollutant spatiotemporal diffusion simulation.

[0016] Furthermore, the multi-source data integration step also includes a cross-modal data fusion process, which extracts features from unstructured data such as UAV aerial images and inspection videos to obtain image pollution index and facility risk index. The two types of indices and standardized pollution monitoring data are then combined using a weighted fusion method with preset weights to construct a cross-modal intermediary. The intermediary is then mapped to the corresponding grid to supplement the grid's attribute information.

[0017] By adopting the above technical solutions, cross-modal data fusion transforms unstructured data into effective information, solves the problem of insufficient data in sparse areas (such as remote factory areas), supplements grid attributes to make the data more comprehensive, provides more sufficient data support for subsequent simulation and evaluation, and improves the reliability of recognition results.

[0018] Furthermore, it also includes a digital twin simulation step, which involves constructing a digital twin image of the industrial cluster based on the results of multi-media coupling simulation, loading parameters corresponding to extreme scenarios into the image to correct environmental parameters and pollution source strength data, simulating multiple intervention schemes in the image and outputting pollution control effect data corresponding to each scheme.

[0019] By adopting the above technical solutions, digital twin simulations can achieve virtual simulations of extreme scenarios (such as rainstorms and tank leaks), obtaining the effects of intervention plans without actually conducting risk experiments. This provides a reference for actual governance, reduces trial and error costs, and improves the scientific nature and efficiency of governance decisions.

[0020] Furthermore, it also includes a goal-oriented pollution attribution step, which calculates the basic contribution value of different pollution causes to high-risk grids using a graph neural network backpropagation algorithm, defines the priority of attribution dimensions according to preset dynamic target labels, adjusts the weight of each pollution cause in combination with the priority, and calculates the final contribution ratio of each cause to high-risk grids, and outputs the attribution results in the form of a text report containing the contribution ratio and key evidence.

[0021] By adopting the above technical solutions, the target-oriented attribution clearly identifies the core pollution causes in high-risk areas (such as facility leaks and atmospheric deposition), and can adjust the attribution focus according to different objectives (such as health risks and ecological risks). The attribution results are combined with key evidence to provide a clear direction for targeted governance, avoid indiscriminate control, and improve the pertinence of governance.

[0022] Furthermore, it also includes federated transfer and reinforcement learning closed-loop steps, constructing a cross-park parameter server to store the general graph neural network parameters of 3D association embedding, loading the general parameters into the new park and then fine-tuning the parameters using a small amount of local measured data. The fine-tuning loss function includes local association embedding error and general parameter deviation constraints. The knowledge of the mature park model is transferred to the new park model through knowledge distillation. The model parameters and intervention scheme parameters are the adjustment objects, and the achievement of dynamic goals is the evaluation basis. The parameters are adjusted through iterative optimization algorithms to form a closed loop from target input to parameter iteration.

[0023] By adopting the above technical solutions, federated transfer learning significantly reduces the data requirements and time for training models in newly built parks, enabling rapid cross-regional deployment; the reinforcement learning closed loop enables the model to continuously iterate and optimize, adapt to long-term scenario changes, and improve the model's reusability and long-term effectiveness.

[0024] Secondly, this application provides an intelligent auxiliary identification system for high-risk contaminated areas in industrial clusters. It includes a data acquisition subsystem and a server. The data acquisition subsystem is used to collect remote sensing images of the industrial cluster, surface change data, soil and groundwater environmental parameters, pollutant concentration data, and real-time operational status data of contamination facilities. The server is used to receive data transmitted by the data acquisition subsystem and execute the intelligent auxiliary identification method for high-risk contaminated areas in industrial clusters as described in any one of the first aspects above. Data transmission between the data acquisition subsystem and the server is achieved through a preset communication method.

[0025] By adopting the above technical solutions, the data acquisition subsystem realizes real-time and comprehensive acquisition of multiple types of data, providing data input for the identification method; the server, as the method execution carrier, ensures the efficient operation of the identification process; the two work together through a preset communication method to form a complete "data acquisition-method execution" link, providing hardware support for intelligent identification of high-risk pollution areas in industrial clusters, and ensuring the continuity and stability of the identification process.

[0026] In summary, this application has at least the following beneficial effects:

[0027] It provides a precise and efficient solution for identifying high-risk pollution areas in industrial clusters, solving the problems of low accuracy and poor efficiency of existing technologies;

[0028] By using dynamic topology and multi-media coupling, the model's adaptability to complex scenarios is improved;

[0029] By leveraging goal-oriented attribution and digital twin simulations, we can provide clear directions and solutions for governance.

[0030] Federated migration learning enables rapid deployment across campuses, lowering the application threshold.

[0031] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0032] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0033] Figure 1 The diagram illustrates the principle of an intelligent auxiliary identification system for high-risk pollution zones in industrial clusters, as described in an embodiment of this application.

[0034] Figure 2A flowchart of an intelligent auxiliary identification method for high-risk pollution zones in industrial clusters is shown in an embodiment of this application. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0037] This application provides an intelligent auxiliary identification method and system for high-risk pollution areas in industrial clusters. It can accurately and efficiently identify high-risk areas, adapt to dynamic scenarios, clarify the causes of pollution to support governance decisions, and can be quickly deployed across industrial parks, effectively solving the problems of low accuracy and poor adaptability of existing technologies.

[0038] Figure 1 The diagram illustrates the principle of an intelligent auxiliary identification system for high-risk pollution zones in industrial clusters, as described in an embodiment of this application.

[0039] Reference Figure 1 The system includes a data acquisition subsystem and a server. As a whole, it provides a hardware operating environment for intelligent auxiliary identification methods for high-risk pollution areas in industrial clusters, realizes efficient execution of multi-source data acquisition, transmission and identification algorithms, and ensures the continuity and stability of the identification process.

[0040] The data acquisition subsystem, serving as the data input terminal, is used to collect various types of data required by the identification method. Its hardware components include remote sensing data acquisition equipment (such as drones and remote sensing satellite receiving terminals), ground monitoring equipment (such as soil sensors, groundwater monitors, and pollutant concentration detectors), and facility status acquisition equipment (such as tank level sensors and pipeline pressure sensors). These devices respectively acquire remote sensing images of the industrial cluster, surface change data, soil and groundwater environmental parameters, pollutant concentration data, and real-time operational status data of pollution facilities. Simultaneously, the data acquisition subsystem is also equipped with communication equipment to transmit the collected data to the server in real time according to preset communication protocols (such as 5G and industrial Ethernet).

[0041] The server, as the core of data processing and algorithm execution, receives data transmitted by the data acquisition subsystem and runs identification methods. Its hardware components include a storage unit, a computing unit, and a data interaction unit. The storage unit is equipped with a spatial database to store the raw data collected, the grid data generated during the identification process, model parameters, and intermediate calculation results. The computing unit is equipped with computing acceleration components (such as GPUs) to run identification algorithms such as spatial grid construction, multi-source data integration, and graph neural network simulation to complete the calculation and identification of high-risk pollution areas. The data interaction unit establishes a data link with the communication equipment of the data acquisition subsystem to receive the collected data, and can also connect to external display devices (such as industrial touch screens) to output the high-risk area map, attribution report, and intervention plan obtained from the identification.

[0042] The data acquisition subsystem and the server achieve two-way data interaction through a preset communication method. The data acquisition subsystem continuously transmits real-time data to the server, and the server runs the identification method based on the data and can provide the data acquisition subsystem with data acquisition needs adjustment instructions. The two work together to form a complete hardware support link of "data acquisition-processing-identification-result output", ensuring the stable implementation and operation of the intelligent auxiliary identification method for high-risk pollution areas in industrial clusters.

[0043] This application specifically discloses an intelligent auxiliary identification method for high-risk pollution zones in industrial clusters.

[0044] Figure 2 A flowchart of an intelligent auxiliary identification method for high-risk pollution zones in industrial clusters is shown in an embodiment of this application.

[0045] Reference Figure 2 The method specifically includes the following steps:

[0046] S1: Spatial grid construction steps: Based on the preset coordinate system, the industrial cluster area is divided into standard grids, each grid is assigned a unique identifier and spatial boundary information is recorded, and the grids of high-risk facility areas and environmentally sensitive areas are refined.

[0047] The preset coordinate system mentioned here preferentially adopts the National Geodetic Coordinate System 2000 (CGCS2000), which is the legal benchmark for geospatial data in my country and can ensure the national uniformity of grid spatial location and the compatibility of subsequent data fusion. The standard grid size is set to 1m×1m by default. This size can balance the positioning accuracy of pollution facilities in conventional industrial clusters (such as storage tanks with a diameter of 1.5-3m and sewage outlets with a diameter of 0.5-2m) with the overall calculation efficiency. The unique identifier of the grid adopts the encoding rule of "GridID = G[row number]_[column number]", where the row number and column number are generated based on the top left corner of the smallest bounding rectangle of the industrial cluster in the preset coordinate system as the origin, and incremented along the x-axis (eastward) and y-axis (northward). For example, the grid number located in the 30th row and 15th column is G_30_15. The spatial boundary information is recorded in Well-KnownText (WKT) format, such as "POLYGON((x1y1,x2y2,x3y3,x4y4,x1y1))", where x1 / y1 to x4 / y4 are the planar coordinates of the four vertices of the grid in the preset coordinate system.

[0048] The specific methods in this step include: Mesh refinement is performed using a zoned hierarchical control approach. In general industrial areas, standard mesh sizes are maintained, and mesh boundary coordinate deviations are controlled within a preset range. Here, the mesh boundary coordinate deviation is determined using a formula... The calculation is performed, where ΔP is the coordinate deviation of a vertex on the grid boundary, x_actual and y_actual are the actual planar coordinates of that vertex measured by GNSS (source: industrial cluster survey report), and x_grid and y_grid are the theoretical planar coordinates of that vertex based on the grid generation algorithm (source: grid partitioning program under the preset coordinate system). The preset upper limit for ΔP in general industrial areas is 0.5m to meet the location requirements of conventional pollution facilities. The grid size in high-risk concentrated areas is adjusted to a second grid size smaller than the standard grid size to adapt to the location requirements of small-scale pollution sources. High-risk concentrated areas are determined by the density of pollution facilities, and the determination formula is... ρ represents the density of contaminated facilities within the grid (unit: facilities / m²). 2 ), N fac S represents the number of pollution facilities included within the grid (sourced from CAD drawings of facility layout provided by the company). gridarea The grid area (1m² under standard grid) 2 When ρ≥0.2 units / m 2When an area is identified as a high-risk concentration zone, the second grid size is preferentially set to 0.5m × 0.5m. At this point, the upper limit of the coordinate deviation ΔP for the high-risk concentration zone is adjusted to 0.2m to ensure that the positioning error of small-scale pollution sources (such as valve interfaces with a diameter of 0.8m) is controllable. The grid size of the boundary of the environmentally sensitive area is further adjusted according to the protection level of the sensitive area to adapt to the accuracy requirements of the pollutant diffusion simulation boundary. The protection level of the environmentally sensitive area is divided according to the "Technical Guidelines for Environmental Impact Assessment - Ecological Impact" (HJ19-2022). For example, a primary water source protection area is the highest level, and a secondary water source protection area is the medium level, with corresponding grid sizes adjusted to 0.2m × 0.2m and 0.3m × 0.3m respectively. The accuracy of the pollutant diffusion simulation boundary is determined by the formula ΔD = |L 模拟 -L 实际 |Evaluation, where ΔD is the distance deviation between the diffusion simulation boundary and the actual boundary of the sensitive area, L 模拟 L represents the distance of pollutant diffusion boundary based on the grid (sourced from subsequent GNN simulation results). L is actually the legal boundary distance of the sensitive area (sourced from the red line map of sensitive areas released by the ecological and environmental departments). The preset upper limit of ΔD for environmentally sensitive areas is 2m. This deviation can be made to meet the requirements by adjusting the grid size. For example, when a 0.2m×0.2m grid is used in the primary water source protection area, ΔD can be controlled within 1m.

[0049] In addition, the grid data is ultimately stored in a PostGIS spatial database. The database table structure includes fields such as GridID (character type), WKT_Boundary (text type), Area_Type (character type, identifying "general area / high-risk area / sensitive area"), and Adjacent_Grids (array type, storing the GridIDs of adjacent grids). Among them, Adjacent_Grids is generated through a topological relationship algorithm. The algorithm logic is as follows: if two grids share an edge with a length ≥ 0.8 times the grid side length, they are determined to be adjacent. For example, the adjacent determination condition for 0.5m × 0.5m grids is that the shared edge length is ≥ 0.4m. This topological relationship provides the basis for the edge weight calculation of the subsequent 3D associative embedded graph neural network simulation.

[0050] S2: Multi-source data integration step, collecting spatial basic data of industrial clusters, enterprise facility layout data, pollution monitoring data and geological and hydrological data, and after standardizing the collected data, mapping it to the corresponding grid through spatial overlay or interpolation methods to form the relationship between grid and data.

[0051] The spatial foundation data specifically includes vector boundary maps and land use type maps of the industrial cluster (sourced from GIS data provided by the local natural resources department), in Shapefile format; enterprise facility layout data includes location coordinates, size parameters, and media types of pollution facilities (storage tanks, wastewater treatment ponds, sewage outlets, etc.) (sourced from enterprise environmental impact assessment reports and CAD drawings from on-site surveys); pollution monitoring data includes soil heavy metal concentrations, groundwater VOCs concentrations, and atmospheric particulate matter concentrations (sourced from data from automatic monitoring stations deployed in the cluster and laboratory sampling analysis data), with a monitoring frequency of once per day; geological and hydrological data includes soil texture (sandy / clay), soil permeability coefficient, groundwater depth, and groundwater flow direction (sourced from detailed exploration reports issued by geological survey units). The collected data undergoes standardization processing, the core of which is to eliminate differences between different data dimensions, using the Min-Max normalization method, with the formula:

[0052]

[0053] In the formula, x std Here, x represents the standardized data value (range [0,1]), and x represents the original data value (e.g., the original soil cadmium concentration is 0.8 mg / kg). min x represents the minimum value of this type of data (sourced from historical monitoring datasets of the same region, such as the minimum soil cadmium concentration of 0.1 mg / kg). max The maximum value of this type of data (source as above, such as the maximum soil cadmium concentration of 5.0 mg / kg); for qualitative data (such as land use type "industrial land" "green space"), a unique thermal coding is used to convert it into numerical data (such as industrial land coded as [1,0], green space coded as [0,1]).

[0054] When mapping to a grid using spatial overlay or interpolation methods, spatial overlay is suitable for vector data (such as enterprise facility layout data): The facility vector surface and the grid vector surface are spatially intersected. If the overlapping area of ​​the facility and the grid accounts for ≥30% of the total area of ​​the facility, then the facility is determined to belong to the corresponding grid, and the facility parameters (such as tank volume, media type) are attached to the grid attributes. Interpolation methods are suitable for discrete point monitoring data (such as pollution monitoring data), using inverse distance weighted (IDW) interpolation, with the following formula:

[0055]

[0056] In the formula, C i Let C be the interpolated pollutant concentration (e.g., mg / kg) for grid i, where n is the number of monitoring points participating in the interpolation (usually the 5 nearest surrounding monitoring points), and C is the interpolated concentration for grid i. k The measured concentration at monitoring point k (source: monitoring data), d ikis the straight-line distance (in meters) between the center point of grid i and monitoring point k, calculated using coordinates in a preset coordinate system. p is the distance attenuation coefficient (valued at 2, derived from the Environmental Monitoring Interpolation Technical Specification, ensuring that closer monitoring points have a greater impact on grid concentration). After mapping, the attribute table of each grid contains the associated fields "GridID - Spatial Basic Data - Facility Data - Monitoring Data - Geological and Hydrological Data", forming a complete grid-data association relationship.

[0057] The specific methods in this step include: a cross-modal data fusion process, feature extraction of unstructured data such as UAV aerial images and inspection videos to obtain image pollution index and facility risk index, construction of cross-modal mediators by weighted fusion of the two indices and standardized pollution monitoring data with preset weights, and mapping of the mediators to the corresponding grids to supplement the attribute information of the grids.

[0058] The drone aerial imagery has a resolution of 0.1m / pixel (sourced from a DJI M300RTK drone, flight altitude 100m). Semantic segmentation algorithms (such as the U-Net model) were used to extract soil anomaly areas (e.g., color anomalies, compacted areas) and vegetation withering areas. The Image Pollution Index (IPI) is calculated using the following formula:

[0059]

[0060] In the formula, IPI i S is the image contamination index of grid i (range [0,1]). abn,i Area of ​​soil anomaly region within grid i (unit: m²) 2 (Source: semantic segmentation results), S grid The area of ​​grid i (the standard grid is 1m). 2 ), R wilt,i The vegetation withering rate within grid i (value range [0,1], i.e., the ratio of withered vegetation area to total vegetation area, sourced from semantic segmentation results). The inspection video is in 4K resolution (recorded by a handheld inspection device at 25fps). Rust areas and leakage traces (such as oil stains and water stains) are identified using object detection algorithms (such as the YOLOv8 model). The Facility Risk Index (FRI) is calculated using the following formula:

[0061]

[0062] In the formula, FRI i S represents the risk index (range [0,1]) of the facilities within grid i. rus,i The corroded area of ​​the facility within grid i (unit: m²) 2 (Source: Target Detection Results), S fac,iThe total surface area of ​​the facilities within grid i (unit: m²) 2 (Source: Enterprise facility parameters), I(·) is an indicator function; if a leak sign is detected (LeakSign) i If the value is 1, then take 1; if there is no trace of leakage, then take 0.

[0063] Cross-modal mediator (M) i The weighted fusion formula for ) is:

[0064]

[0065] In the formula, M i Let be the cross-modal mediator of grid i (value range [0,1]), ω1, ω2, and ω3 are weighting coefficients (satisfying ω1+ω2+ω3=1). In the sparse data region (monitoring point spacing > 50m), ω1=0.45, ω2=0.45, and ω3=0.1 (due to the scarcity of monitoring data, unstructured data is relied upon). In the non-sparse data region, ω1=0.3, ω2=0.3, and ω3=0.4 (due to the abundance of monitoring data, measured data is prioritized). C meas,i C represents the pollutant monitoring concentration after standardization for grid i (source: the standardization results described above). std This corresponds to the environmental quality standard limit for the pollutant (e.g., the screening value for cadmium in soil is 0.6 mg / kg, sourced from the "Soil Environmental Quality Standard for Construction Land Soil Pollution Risk Control" (GB36600-2018)). M i As a new attribute, it can be attached to the grid to supplement the grid information in sparse areas of data, thereby improving the reliability of subsequent simulations and evaluations.

[0066] S3: The three-dimensional correlation embedding graph neural network simulation step involves constructing a graph neural network model with grids as nodes and pollutant migration relationships between grids as edges. The three-dimensional features of grid identification, pollution facility type and proportion, and pollutant concentration are embedded into the model nodes to train the model and realize the spatiotemporal diffusion simulation of pollutants in soil-groundwater media.

[0067] The graph neural network model constructed here is an endogenous GNN with object perception. First, it is necessary to clarify the specific composition of the node feature vector: the grid identifier is converted into numerical features through one-hot encoding (e.g., GridID = G_30_15 corresponds to the encoding vector [0,...,1,...,0], with the dimension being the total number of grids), the facility type and proportion are represented through multi-hot encoding + normalization (e.g., if the grid contains storage tanks and wastewater ponds, the type encoding is [1,1,0], and the proportion is storage tank area / grid area = 0.3, wastewater pond area / grid area = 0.2, combined into [1,1,0,0.3,0.2]), and the pollutant concentration is a standardized value (e.g., the cadmium concentration in the soil is standardized to 0.6). The three are concatenated to form the initial node feature vector. (d is the feature dimension, which is determined by the sum of each sub-feature dimension, such as grid identification dimension 1000 + facility feature dimension 5 + concentration dimension 1 = 1006).

[0068] The core of the model is the target-aware message passing mechanism, and the convolutional layer calculation follows the formula:

[0069]

[0070] In the formula, Let N(i) be the node characteristics of the (i+1)th layer grid i, and N(i) be the set of adjacent nodes of grid i (initially 4-way adjacency, i.e., top, bottom, left, and right grids). Let be the connection weight matrix between grid i and j in the I-th layer (obtained through gradient descent optimization during model training). The first layer bias vector (trained together with the weight matrix); GoalAttn(·) is the target attention function, used to amplify the node connection weights that match dynamic targets (such as health risk management), and its calculation formula is... in , where is the feature vector of the dynamic target label (e.g., the label vector corresponding to a health risk target is determined by the risk factor weights labeled by experts), and ⊙ represents the element-wise product operation. The loss function used for model training is the mean squared error (MSE), and the formula is: in The model predicts the pollutant concentration in grid i at time t. Let N be the measured pollutant concentration of grid i at time t (source: monitoring data), N be the total number of grids, and T be the total number of time steps. The training is iterated until the loss function converges (e.g., loss value < 0.001).

[0071] The method in this step specifically includes: a dynamic topology adjustment process, where the environmental parameters corresponding to the grid are checked at preset time steps to see if changes exceeding a preset range occur or if there are any human intervention commands. If any condition is met, topology adjustment is triggered, adjusting the connection weights between grids based on the detection results and expanding the adjacency connections of the pollution diffusion front grid from the original 4 directions to 8 directions. Here, the preset time step is set to 6 hours (determined based on the migration rate of pollutants in the soil-groundwater medium, referring to the pollutant migration timescale in the "Technical Guidelines for Risk Assessment of Contaminated Sites" (HJ25.3-2014)). A significant change is defined as a change of 10% (i.e., the ratio of the change in environmental parameters to the initial value is ≥10%). Specifically, environmental parameters refer to the soil permeability coefficient K (unit: m / d, sourced from soil test data in the geological survey report) and groundwater flow velocity v (unit: m / d, calculated from the groundwater level difference and permeability coefficient using Darcy's law, formula: v = K·J, where J is the hydraulic gradient, sourced from groundwater monitoring well level data). Human intervention commands include setting up barriers and activating pumping wells (input by management personnel through the system, with the command format "Intervention Type - Affected Grid Range - Intervention Intensity").

[0072] The formula for adjusting the connection weights between meshes is:

[0073]

[0074] In the formula, W i,j,t Let W be the connection weight of grid ij at time t. i,j,init K represents the initial connection weights (the weight values ​​at training convergence). i,t v i,t Let K be the soil permeability coefficient and groundwater flow velocity of grid i at time t. i,init v i,init Flag is the corresponding parameter at the initial time (t=0). i,inter The intervention indicator is set to 1 for grid i affected by the intervention, and 0 otherwise. 0.99 is the attenuation coefficient of the intervention on the weight, determined based on experimental data on the blocking efficiency of the barrier on pollutant migration. The determination of the pollution diffusion front grid is based on the concentration gradient threshold, i.e., calculating the concentration difference ΔC between grid i and its adjacent grids. i,j =|C i,t -C j,t |, if ΔC exists i,j ≥ΔC th (ΔC th The concentration gradient threshold is set to 0.1 × C. std C stdIf the limit value of the environmental quality standard for pollutants is used, then grid i is determined to be the diffusion front grid, and its adjacent connections are expanded from the original 4 directions (up, down, left, right) to 8 directions (adding upper left, upper right, lower left, and lower right) to better match the actual diffusion direction of pollutants.

[0075] It also includes a multi-media coupling simulation process, which calculates the pollutant migration flux from the atmosphere to the soil and from the soil to the groundwater using a preset migration flux calculation method. The migration flux data is then embedded as a new feature into the nodes of the graph neural network model. After updating the node feature vectors, the model is retrained to realize the multi-media coupling simulation of pollutant migration between the atmosphere, soil, and groundwater.

[0076] The formula for calculating the atmospheric-to-soil pollutant migration flux (taking VOCs as an example) is as follows:

[0077] F atm→soil,i =C voc,i ×v wind,i ×K sa,i ×δ rain,i

[0078] In the formula, F atm→soil,i The atmospheric VOCs deposition flux into the soil for grid i (unit: mg / (m³) 2 ·d)), C voc,i VOCs concentration in the atmosphere above grid i (unit: mg / m³) 3 (Source: data from atmospheric monitoring stations), v wind,i Let K be the average wind speed of grid i (unit: m / s, source: data from the park's weather station). sa,i δ represents the partition coefficient of VOCs between the atmosphere and soil (unit: m, source: compound property data from the Handbook of Environmental Chemistry). rain,i This is a rainfall correction factor (1.3 when there is rainfall and 1.0 when there is no rainfall, determined based on experimental data on the enhancing effect of rainfall on the wet deposition of air pollutants).

[0079] The formula for calculating the pollutant migration flux from soil to groundwater is:

[0080]

[0081] In the formula, F soil→gw,i The infiltration flux of soil pollutants into groundwater in grid i (unit: mg / (m³)). 2 ·d)), C soil,i K represents the soil pollutant concentration (mg / kg, source: soil monitoring data) for grid i. sg,i K represents the partition coefficient of pollutants between soil and groundwater (unitless, sourced from compound property literature). soil,iSoil permeability coefficient (unit: m / d, source: geological survey report), μ gw The dynamic viscosity of groundwater (unit: Pa·s, taken as 1.002 × 10⁻⁶ at room temperature) is given. -3 Pa·s), d gw,i The groundwater depth for grid i (unit: m, source: groundwater monitoring well data) This is the flux attenuation term due to burial depth (the greater the burial depth, the smaller the flux, which is consistent with the actual migration law).

[0082] After standardizing the two flux data points mentioned above (using Min-Max normalization, the formula is the same as x in S2), std (Calculation), concatenated with the original three-dimensional node feature vector, to form a four-dimensional node feature vector of "grid identifier - facility feature - pollutant concentration - multi-media flux". (d+2 is the original dimension d plus 2 flux feature dimensions), and then re-substitute the aforementioned GNN convolution formula and loss function for training. The training data is supplemented with multi-media monitoring data (such as groundwater VOCs concentration and atmospheric VOCs concentration) to ensure that the model can accurately simulate the cross-media pollutant migration process.

[0083] S4: Risk assessment step. Based on four preset indicators, including pollution source intensity, pollution migration channels, environmental vulnerability, and historical records, an assessment system is constructed. The comprehensive risk value of each grid is calculated, and high, medium, and low risk levels are divided according to the preset risk value range. The results of high-risk pollution area identification are output.

[0084] The composition and data sources of the secondary indicators for the four primary indicators need further clarification: The pollution source intensity indicator includes "enterprise pollution intensity," "pollution facility density," and "number of characteristic pollutant types." Enterprise pollution intensity is represented by the enterprise's annual permitted discharge volume (unit: t / a, source: discharge permit issued by the ecological and environmental department). Pollution facility density is the ratio of the area occupied by pollution facilities within the grid to the grid area (source: CAD drawings of enterprise facility layout). The number of characteristic pollutant types is the count of pollutant types in the "Priority Controlled Chemicals List" within the grid (source: laboratory test reports). The pollution migration channel indicator includes "soil permeability," "groundwater depth," and "pipeline aging probability." Soil permeability is represented by the soil permeability coefficient (unit: m / d, source: geological survey report). Groundwater depth is the measured value from monitoring wells (unit: m, source: groundwater monitoring data). The pipeline aging probability is based on the pipeline's service life and design life. The ratio calculation (source: enterprise pipeline maintenance archives) includes "land use type", "aquifer type", and "ecological sensitivity score". Land use type is classified and assigned a value according to "industrial land / residential land / green space" (source: land use map of natural resources department). Aquifer type is distinguished by "unconfined aquifer / confined aquifer" (source: geological survey report). Ecological sensitivity score is determined with reference to the "Technical Specification for Ecological Function Zoning" (HJ19-2022) (source: zoning data of ecological and environmental departments). Historical record indicators include "frequency of monitoring exceedance", "pollution incident record", and "environmental protection penalty record". The frequency of monitoring exceedance is the number of times the monitoring data in the grid exceeds the standard in the past 3 years (source: monitoring database). The pollution incident record is assigned a value according to whether a pollution accident has occurred in the past 5 years (source: archives of emergency management department). The environmental protection penalty record is assigned a value according to whether the enterprise has been punished for environmental protection in the past 3 years (source: penalty publicity of ecological and environmental departments).

[0085] The specific methods in this step include: the weights of the four types of indicators are determined by a combination of subjective and objective weight determination methods; the secondary indicators under each indicator are quantitatively graded according to preset scoring standards; and the comprehensive risk value is calculated by weighted sum of the weights of each indicator and the corresponding scores.

[0086] Subjective weights are determined using the Analytic Hierarchy Process (AHP). First, a first-level indicator judgment matrix A = (a ij ) 4×4 , where a ij This indicates the importance of the i-th primary indicator relative to the j-th primary indicator, using a 1-9 scale (1 indicates equal importance, 3 indicates slightly important, 5 indicates significantly important, 7 indicates very important, 9 indicates extremely important, and 2 / 4 / 6 / 8 are intermediate values). ji =1 / a ijThe judgment matrix was constructed based on scores from 10 experts in the fields of environmental engineering and pollution control (sourced from expert consultation questionnaires); subsequently, the maximum eigenvalue λ of the judgment matrix was calculated. max The corresponding feature vectors are then normalized to obtain the subjective weight vector W. sub =[W sub1 W sub2 W sub3 W sub4 (These correspond to the subjective weights of pollution source intensity, pollution migration pathways, environmental vulnerability, and historical records, respectively), and pass the consistency test (consistency index CI = (λ)). max -4) / (4-1), the random consistency ratio CR = CI / RI, when CR < 0.1, the matrix consistency is considered qualified, RI is the random consistency index, and for a 4th order matrix RI = 0.90).

[0087] The objective weights are determined using the entropy weight method. First, the raw data of the secondary indicators are standardized (positive indicators). negative indicators Where x ij This represents the original value of the j-th secondary indicator in the i-th grid. (These are the minimum and maximum values ​​of the j-th secondary indicator, respectively); then, the entropy value of the j-th secondary indicator is calculated. (n is the total number of grid cells, If p ij =0 then lnp ij (Treat it as 0); then calculate the information utility value d. j =1-e j Objective weight (m represents the total number of secondary indicators), and the objective weight vector W of the primary indicators is obtained by classifying and summarizing them according to the primary indicators. obj =[W obj1 W obj2 W obj3 W obj4 ].

[0088] The weights of the primary indicator combinations are calculated using a weighted average method, with the formula W. k =αW subk +(1-α)W objk(k=1,2,3,4), where α=0.5 (the value is based on the principle of balancing subjective experience and data objectivity, and the risk assessment results fluctuate by less than 5% in the range of 0.4-0.6 after sensitivity analysis verification). The final determined combination weights usually satisfy W1 (pollution source intensity)∈[0.35,0.45], W2 (pollution migration channel)∈[0.20,0.30], W3 (environmental vulnerability)∈[0.15,0.25], and W4 (historical record)∈[0.10,0.20].

[0089] Each secondary indicator is quantitatively graded according to a preset scoring standard (0-5 points, with 5 points being the highest risk). For example, the grading standard for "density of contaminated facilities" is as follows: 1 point for density <20%, 2 points for 20%-40%, 3 points for 40%-60%, 4 points for 60%-80%, and 5 points for >80% (the grading is based on the correlation study between facility density and pollution risk in the "Technical Guidelines for Risk Control and Remediation of Soil Pollution in Industrial Enterprises" (HJ25.4-2019)); the grading standard for "soil permeability" is as follows: 2 points for permeability coefficient <0.1m / d (clay soil), 3 points for 0.1-1m / d (loam soil), and 5 points for >1m / d (sandy soil) (because pollutants migrate faster in sandy soil, resulting in higher risk).

[0090] The formula for calculating the comprehensive risk value is:

[0091]

[0092] In the formula, Risk_Score i Let W be the overall risk value of the i-th grid. k Let m be the combined weight of the k-th primary indicator. k w represents the number of secondary indicators under the k-th primary indicator. kj The normalized weight of the j-th secondary indicator under the k-th primary indicator s kij The score (0-5 points) is given to the j-th secondary indicator under the k-th primary indicator of the i-th grid.

[0093] The preset risk value range was determined with reference to the "Standard for Risk Control of Soil Pollution in Construction Land" (GB36600-2018) and the characteristics of regional industrial pollution: Risk_Score i A score of ≥4.0 indicates high risk (remediation and remediation should be prioritized), while a score of 2.0 ≤ Risk_Score indicates high risk. i A score below 4.0 indicates medium risk (monitoring frequency needs to be increased). Risk_Score iA score of <2.0 indicates low risk (routine monitoring is sufficient); when outputting the identification results of high-risk pollution areas, the set of high-risk grids is exported in vector surface format, and the GridID, comprehensive risk value and main contribution index of each high-risk grid are labeled (e.g., "G_30_15, Risk=4.8, main contribution: pollution source intensity (weight 0.42, score 4.5)"), providing precise guidance for subsequent governance decisions.

[0094] S5: Digital twin simulation step, constructing a digital twin image of the industrial cluster based on the multi-media coupling simulation results, loading parameters corresponding to extreme scenarios into the image to correct environmental parameters and pollution source strength data, simulating multiple intervention schemes in the image and outputting pollution control effect data corresponding to each scheme.

[0095] The basic data for constructing the digital twin mirror here is entirely derived from the simulation results of the multi-media coupled GNN in S3, including the spatial coordinates of grid nodes, four-dimensional feature vectors (grid identifier - facility feature - pollutant concentration - multi-media flux), spatiotemporal diffusion coefficients of pollutants in atmospheric, soil and groundwater media, and other core data. At the same time, it integrates the grid topology relationship of S1 and the multi-source attribute data of S2 (such as land use type, geological and hydrological parameters) to form a four-layer mirror architecture of "geospatial layer - facility model layer - pollutant migration layer - environmental parameter layer". The geospatial layer uses a preset coordinate system as a reference to reproduce the topography and landforms of the industrial cluster (such as buildings, roads, and green belts) at a 1:1 scale. The data accuracy matches the S1 grid size (standard grid corresponds to a mirror space resolution of 1m, and refined grid corresponds to 0.2-0.5m). The facility model layer uses BIM (Building Information Modeling) technology to construct a three-dimensional solid model of the pollution facilities (such as storage tanks modeled according to actual diameter, height, and material parameters). The model parameters are sourced from the facility design drawings provided by the enterprise. The pollutant migration layer embeds the S3 multi-media coupled GNN model to drive the dynamic update of the pollutant concentration field in real time. The environmental parameter layer is associated with real-time monitoring equipment (such as soil sensors and weather stations) to ensure that the mirror image is synchronized with the environmental status of the actual site (the data update frequency is consistent with the S3 dynamic topology adjustment time step, which is 6 hours / time).

[0096] When loading parameters corresponding to extreme scenarios into the image to correct environmental parameters and pollution source intensity data, the extreme scenarios mainly include rainstorm scenarios and storage tank leakage scenarios. The parameter correction for both types of scenarios is based on quantitative formulas determined by industry standards and practical engineering experience. Specifically, the parameter correction for rainstorm scenarios focuses on soil moisture content and soil permeability coefficient. The formula for calculating the increase in soil moisture content is:

[0097]

[0098] In the formula, Δθ i,RLet θ represent the soil moisture content increment (in %) of grid i under the heavy rain scenario, R be the 24-hour rainfall (in mm, sourced from extreme rainfall data released by the local meteorological bureau, such as 150 mm for a 50-year return period and 200 mm for a 100-year return period), and θ be the soil moisture content increment (in %). i,init Let θ be the initial soil moisture content of grid i (unit: %, sourced from measured values ​​in S2 geological and hydrological data; for example, the initial moisture content of sandy soil is typically 10%-15%). max,i The maximum soil moisture content (unit: %) is given by grid i, derived from soil water-holding capacity experimental data in the geological survey report; for example, the maximum moisture content of sandy soil is 25%-30%. Based on the corrected soil moisture content, the soil permeability coefficient is further corrected using the following formula:

[0099]

[0100] In the formula, K i,R K represents the soil permeability coefficient (unit: m / d) for grid i under the heavy rain scenario. i,init The initial soil permeability coefficient is given by S2 geological and hydrological data, such as an initial K value of 1-5 m / d for sandy soil. 0.5 is the sensitivity coefficient of permeability to water content (from the experimental research conclusions on the permeability characteristics of sandy soil in "Soil Physics").

[0101] The core of parameter correction in tank leakage scenarios is adjusting the pollution source strength. A small-hole leakage model is used to calculate the leakage rate, which is then used to correct the pollution source strength. The formula is as follows:

[0102]

[0103] In the formula, v leak Leakage rate of the storage tank (unit: m) 3 / h), C d Let A be the leakage coefficient (taken as 0.6, derived from the typical coefficient for small-hole leakage in metal pipes in the "Technical Guidelines for Handling Hazardous Chemical Leakage Accidents"), and let A be the leakage outlet area (unit: m²). 2 Based on a common fault leak diameter of 5mm, A = π × (0.0025) 2 ≈1.96×10 -5 m 2 P is the pressure of the medium inside the storage tank (unit: Pa, sourced from the company's storage tank operation records, e.g., P is taken as 101325 Pa for an atmospheric pressure storage tank), P0 is the atmospheric pressure (unit: Pa, taken as 101325 Pa), and ρ is the density of the medium inside the storage tank (unit: kg / m³). 3 For example, the density of benzene is 876 kg / m³. 3 (Source: Media Safety Technical Specification), g is the acceleration due to gravity (taken as 9.81 m / s²). 2h is the height of the liquid column at the leak point (unit: m, sourced from tank level gauge data; for example, h is taken as 3m when the tank height is 5m and the liquid level is 3m). The pollution source strength is corrected based on the leakage rate using the following formula:

[0104] Q i,leak =Q i,init +v leak ×C tank ×Δt

[0105] In the formula, Q i,leak Let Q be the pollution source intensity (unit: mg / h) of grid i in the leakage scenario. i,init The initial pollution source strength (sourced from the normal discharge volume in the facility layout data of enterprise S2), C tank Concentration of medium in storage tank (unit: mg / m³) 3 For example, pure benzene medium C tank Take 876000 mg / m 3 ), where Δt is the duration of the leak (unit: h, taken as 24h based on the common emergency response time).

[0106] When simulating multiple intervention schemes in a mirror environment, common intervention schemes include deploying barrier systems, activating pumping wells, and enhancing ventilation. The parameter settings and effect calculation logic for each scheme are as follows: In the barrier deployment scheme, the barrier parameters are set to a thickness of 0.5m and a permeability coefficient K. barrier =1×10 -7 m / d (source: geomembrane material technical parameters, such as the permeability coefficient of HDPE geomembrane), the soil permeability coefficient of the modified barrier cover grid is K. barrier The performance indicator is "the reduction rate of grid pollutant concentration within 100 days after barrier deployment"; in the pumping well activation plan, the pumping well parameters are set to a pumping volume of 5m³ / h. 3 / h, well depth 10m (source: commonly used parameters for groundwater remediation projects), set the pumping well location in the mirror (preferably in the upstream grid of the pollution diffusion path), simulate the impact of groundwater flow field changes on pollutant migration using a GNN model, and the effect index is "area reduction in pollution plume range after pumping well operation"; in the enhanced ventilation scheme, the ventilation equipment parameters are set to an air volume of 10000m³ / h. 3 / h, wind speed 3m / s (source: Industrial Ventilation Equipment Technical Manual), corrected atmospheric-to-soil pollutant migration flux (δ in the formula) wind =0.7, meaning ventilation reduces atmospheric deposition flux by 30%, and the effectiveness indicator is "the time it takes for soil pollutant concentrations to meet standards within 72 hours after ventilation begins." Pollution control effectiveness data is calculated using the built-in effectiveness evaluation module of the mirror platform, with core indicators including the time to meet standards (T). reach(The time when the simulated pollutant concentration first falls below the screening value of the "Soil Environmental Quality Standard for Construction Land Soil Pollution Risk Control" (GB36600-2018)) and the maximum concentration reduction rate. (C max,init To determine the maximum pollutant concentration in the grid before intervention, C max,inter The maximum concentration after intervention and the cost of the solution (including equipment purchase, operating energy consumption, and labor maintenance costs, sourced from market quotations and engineering quota standards) are output in tabular form, showing the correspondence between "Solution Name - Completion Time - Maximum Concentration Reduction Rate - Solution Cost", providing data support for the selection of actual treatment solutions.

[0107] S6: Goal-oriented pollution attribution step. The basic contribution value of different pollution causes to high-risk grids is calculated by backpropagation algorithm of graph neural network. Attribution dimension priority is defined according to preset dynamic target labels. The weight of each pollution cause is adjusted according to the priority and the final contribution ratio of each cause to high-risk grids is calculated. The attribution results, including contribution ratio and key evidence, are output in the form of text report.

[0108] Here, "different pollution causes" are specifically defined as three core causes: facility leakage, atmospheric deposition, and historical residues. Their data foundations correspond to key data from previous steps: facility leakage is linked to facility status data (e.g., tank level changes, pipeline pressure fluctuations) in S2 and pollutant concentration simulation results in S3; atmospheric deposition is linked to atmospheric-soil migration flux data in S3; and historical residues are linked to historical record indicators in S4 (e.g., monitoring exceedance frequency, pollution event records). When calculating the basic contribution value using the graph neural network backpropagation algorithm, the core logic is to use the comprehensive risk value of the high-risk grid as the output of the loss function, and then deduce the gradient contribution of the node features corresponding to each pollution cause to the loss value. The specific formula is as follows:

[0109]

[0110] In the formula, BaseContrib c,i L represents the baseline contribution of pollution type c (c=1 represents facility leakage, c=2 represents atmospheric deposition, and c=3 represents historical residue) to high-risk grid i. i The risk loss function for grid i (i.e., L) i = |Risk_Scoree i -Risk_Target∣,

[0111] Risk_Target is the high-risk threshold, set to 4.0 (sourced from the S4 risk grading standard). c,iThis is the feature vector corresponding to the c-th cause for grid i (e.g., the feature vector for facility leakage is "tank level change - pipeline pressure deviation - soil pollutant concentration increase", sourced from the fusion data of S2 and S3). W is the partial derivative of the loss function with respect to that feature vector (reflecting the sensitivity of the risk value to changes in the feature, calculated using the chain rule of GNN backpropagation). base,c The basic weight for cause c (facility leakage W) base,1 =0.4, atmospheric deposition W base,2 =0.3, historical remnant W base,3 =0.3, sourced from statistical studies on the causes of pollution in industrial clusters in the field of environmental engineering, such as the weighting recommendations in the "Technical Guidelines for the Analysis of the Causes of Pollution in Industrial Sites".

[0112] When defining attribution dimension priorities based on preset dynamic target labels, the dynamic target labels are mainly divided into two categories: "health risk labels" and "ecological risk labels." Each category corresponds to different attribution dimensions and priority weights. Health risk labels focus on the "exposure path dimension" (such as soil-human direct contact, groundwater-drinking water intake), and their priority weights are derived using the exposure dose formula: for the soil-human contact path, the exposure dose... (C soil,i The soil pollutant concentration in grid i is S3; IR is the ingestion rate, taken as 100 mg / d, sourced from the "Technical Guidelines for Soil Pollution Risk Assessment of Construction Land" (HJ25.3-2014); EF is the exposure frequency, taken as 350 d / a; ED is the exposure duration, taken as 25 a; BW is body weight, taken as 60 kg; AT is the average exposure time, taken as 9125 d. Pathways with higher exposure doses have greater priority weights. Under the health risk label, the priority weight for the exposure path dimension is set to 0.7, and the priority weight for the migration path dimension (e.g., soil-groundwater) is set to 0.3. The ecological risk label focuses on the "migration path dimension" (e.g., soil-groundwater-aquatic organisms, atmosphere-vegetation-insects), and its priority weight is calculated based on the ecological impact range: migration impact range S... eco =π×(v mig ×t) 2 (v mig The pollutant migration velocity is represented by a multi-media coupling simulation originating from S3; t represents the migration time, taken as 1a). Pathways with larger impact areas have higher priority weights. Under the ecological risk label, the priority weight for the migration path dimension is set to 0.6, and the priority weight for the exposure path dimension is set to 0.4. Ultimately, the c-th type of cause is determined under the target label S. tag The dimension weights GoalDimW(S) tagc) Determined by the correlation between causes and dimensions. For example, facility leakage has the highest correlation with the "soil-human contact" exposure path, with GoalDimW(health,1)=0.8 under the health risk label. Atmospheric deposition has the highest correlation with the "atmosphere-vegetation" migration path, with GoalDimW(ecology,2)=0.7 under the ecological risk label.

[0113] When adjusting the weights of each pollution cause based on priority and calculating the final contribution percentage, a normalization formula is used to combine the basic contribution value with the dimensional weights:

[0114]

[0115] In the formula, Contrib c,i The final contribution percentage (in percentage form) of cause c to grid i is calculated, with the denominator being the sum of the "base contribution value × dimension weight" of all causes, ensuring that the total contribution percentage of all causes is 100%. For example, in a high-risk grid i under the health risk label, if the BaseContrib for facility leakage is 0.6 and GoalDimW is 0.8, the BaseContrib for atmospheric deposition is 0.3 and GoalDimW is 0.5, and the BaseContrib for historical remnants is 0.2 and GoalDimW is 0.4, then the final contribution percentage for facility leakage is... Atmospheric deposition Historical remnants

[0116] When outputting attribution results in written report form, the report must include four parts: "Grid Basic Information - Contribution Ratio of Each Cause - Key Evidence - Attribution Conclusion". Grid basic information includes GridID and risk value (e.g., G_30_15, Risk = 4.8); the contribution ratio of each cause is presented in a table; key evidence must be linked to measured or simulated data from previous steps. For example, evidence of facility leakage could be "tank level dropped by 0.5m in 24 hours (source: S2 facility status acquisition module), soil benzene concentration exceeded screening value by 3 times (source: S3 monitoring data)", and evidence of atmospheric deposition could be "daily average atmospheric VOCs concentration of 1.2mg / m³". 3 (Source: S2 atmospheric monitoring data), Atmospheric-soil migration flux: 0.8 mg / (m³) 2 ·d)(Source: S3 Multimedia Simulation)”; The attribution conclusion clearly identifies the core causes (e.g., “The high risk in this grid is mainly caused by facility leakage, accounting for 65%, and the sealing status of the storage tanks should be investigated first”).

[0117] S7: Federated transfer and reinforcement learning closed-loop steps. Construct a cross-park parameter server to store the general graph neural network parameters of 3D association embedding. After loading the general parameters into the new park, fine-tune the parameters using a small amount of local measured data. The fine-tuning loss function includes local association embedding error and general parameter deviation constraints. Transfer the knowledge of the mature park model to the new park model through knowledge distillation. Use the model parameters and intervention plan parameters as the adjustment objects and the achievement of dynamic goals as the evaluation basis. Adjust the parameters through iterative optimization algorithms to form a closed loop from target input to parameter iteration.

[0118] The cross-campus parameter server constructed here adopts a distributed architecture, deployed on a cloud platform (such as Alibaba Cloud ECS server), and supports concurrent access from at least 10 campus nodes. The "3D association embedding general graph neural network parameters" it stores specifically include: the node feature weight matrix W of the GNN model. node (Dimension is d×d, where d is the node feature dimension, sourced from the mean parameters of GNN models trained and converged from more than 3 mature parks (such as industrial clusters that have been operating for more than 3 years and have ≥100,000 data entries), edge weight initialization matrix W edge,init (Dimensions are n×n, where n is the maximum number of grid cells, derived from the statistical regularities of mature park topology), Multi-media flux calculation coefficient K flux (For example, the average values ​​of the atmosphere-soil distribution coefficient and the soil-groundwater permeability coefficient are derived from historical parameters of the S3 multi-media simulation in mature industrial parks). The parameter server adopts a "center-edge" synchronization mechanism. Edge nodes (servers in each industrial park) only upload local model parameter updates, while the central server aggregates and updates general parameters, avoiding the transmission of raw data across industrial parks and ensuring data privacy (complying with the sensitive environmental data protection requirements of the Data Security Law).

[0119] When fine-tuning parameters after loading general parameters in a newly built park, the amount of locally measured data collected is 5%-10% of that in a mature park (e.g., a mature park requires monitoring data from 1000 grids, while a newly built park only requires data from 50-100 grids). The data types are consistent with S2 (spatial foundation data, facility data, monitoring data, etc.). The fine-tuning loss function uses a bi-objective constraint, and the formula is:

[0120]

[0121] In the formula, L federated For federal fine-tuning of total losses, L local,bind The mean squared error is used to represent the local association embedding error (i.e., the prediction loss of the GNN model for local data). (C i,local,pred -C i,local,true ) 2 N localC represents the number of local grid cells. i,local,pred For the predicted concentration of local grid i, C i,local,true For the predicted concentration of local grid i, C i,local,pred λ is the measured concentration of local grid i (source: monitoring data from the newly built park S2); λ is the regularization coefficient (set to 0.1, source: model hyperparameter tuning experiment, to ensure that local parameters do not deviate too much from general parameters); θ local,bind For local 3D association embedding parameters (such as node feature weights, edge weights), θ base,bind Generic 3D associative embedded parameters stored on the parameter server. It is the L2 norm, used to constrain the deviation between local parameters and general parameters.

[0122] When transferring knowledge from mature campus models via knowledge distillation, a "teacher-student" model architecture is adopted: the teacher model is a convergent GNN model trained on a mature campus (accuracy ≥90%, derived from validation results of the mature S3 model), and the student model is a fine-tuned GNN model from a newly established campus. The distillation loss function includes hard-label loss and soft-label loss, and the formula is:

[0123] L distill =α·L CE (C hard C student,pred )+(1-α)·L CE (C soft C student,pred )

[0124] In the formula, L distill The total distillation loss is represented by α, which is a weighting coefficient (0.7, derived from distillation experiments to balance the effects of hard and soft labels); L CE C is the cross-entropy loss function; hard Hard labels (i.e., classification labels for measured pollutant concentrations, such as "high concentration / medium concentration / low concentration," sourced from S4 risk classification); C soft For soft labels (the probability distribution output by the teacher model, such as the teacher model predicting a high concentration probability of 0.8, a medium concentration probability of 0.15, and a low concentration probability of 0.05 for a certain grid); C student,pred This is the probability distribution output by the student model. Through this loss function, the student model can learn the generalization ability of the teacher model, improving model performance with small datasets.

[0125] When adjusting model parameters and intervention parameters, the model parameters specifically refer to the node feature weights and edge weight update step sizes of the GNN (derived from the training parameters of the S3 model), while the intervention parameters refer to the key parameters of the intervention scheme in S5 (such as the thickness of the barrier and the pumping volume of the pumping well, derived from the design parameters of the S5 scheme). Together, they constitute the action space A = {Δθ}. gnn,Δp inter}(Δθ gnn Δp is the adjustment amount for the model parameters. inter (This refers to the adjustment amount of the intervention program parameters). The reward function formula, which is based on the achievement of dynamic goals, is as follows:

[0126]

[0127] In the formula, R is the reward value (range [0,1]), β and γ are weighting coefficients (both are 0.5, derived from target priority settings); C pred C represents the pollutant concentration predicted by the model (source: S3 simulation results). target,tag For the concentration threshold corresponding to the dynamic target label (e.g., C under the health risk target) target,tag The screening values ​​are taken from GB36600-2018 (source: S4 risk standard); T pred T represents the time to achieve the target predicted by the model (source: S5 simulation results). target,tag This refers to the time requirement for achieving the target corresponding to the dynamic goal label (e.g., the user sets a target of 30 days, sourced from governance requirement input). The higher the reward value, the more the parameter adjustment direction aligns with the dynamic goal.

[0128] When adjusting parameters using an iterative optimization algorithm (emphasizing the proximal strategy to optimize the PPO algorithm), the iterative process is as follows: First, input the dynamic objective (e.g., "reduce the concentration of high-risk grid cells below the screening value within 30 days"). Based on the current model parameters and intervention plan parameters, perform S3 simulation and S5 rehearsal, and calculate the reward function R. If R < 0.8 (a preset reward threshold derived from engineering acceptance standards), update the action space A using the PPO algorithm (adjusting model parameters and intervention plan parameters), and repeat the simulation and rehearsal. If R ≥ 0.8, the parameter adjustment converges, and the optimal parameters are output. This process forms a closed loop of "dynamic objective input → S3 simulation / S5 rehearsal → reward calculation → parameter iteration → objective verification," ensuring that the model and intervention plan continuously adapt to changes in pollution within the park.

[0129] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0130] This solution employs a hierarchical technical approach encompassing "spatial benchmark unification, data integration for efficiency enhancement, accurate model simulation, scientific risk assessment, functional expansion and adaptation, and cross-regional reuse optimization," which objectively derives the complete technical effect. The specific derivation logic is as follows:

[0131] First, by constructing a hierarchical grid in a partitioned area (preset coordinate system as a spatial reference, and refinement of the grid in high-risk / sensitive areas), the accuracy imbalance caused by the "one-size-fits-all" approach of traditional fixed grids can be solved. This approach ensures the consistency of the overall spatial reference of industrial clusters through a standard grid, and improves the spatial accuracy of small-scale pollution source location and diffusion boundary simulation by adjusting the size of the second grid in high-risk areas and the gradient in sensitive areas, providing a precise spatial carrier for subsequent data mapping and simulation. Combined with S2, multi-source data integration (Min-Max normalization to eliminate dimensional differences, IDW interpolation to complete discrete monitoring data, and cross-modal fusion to supplement unstructured information), the heterogeneity and sparsity of multi-source data can be solved. Normalization ensures the comparability of different types of data (such as concentration and permeability coefficient), interpolation fills the data gaps in monitoring blind spots, and cross-modal mediators supplement pollution correlation information in images / videos, ultimately improving data integrity and consistency and providing high-quality data support for model input.

[0132] Secondly, by using the S3 "3D correlation embedded GNN simulation" (target perception message passing fusion grid-facility-concentration features, dynamic topology adjustment to respond to environmental changes / human intervention, and multi-media coupling to calculate cross-media flux), the shortcomings of traditional models such as "single features, fixed topology, and fragmented media" can be solved. The feature fusion capability of GNN enables the model to capture spatial correlation and pollution attributes at the same time. Dynamic topology adjustment avoids simulation lag after sudden changes in environmental parameters or intervention. Multi-media coupling fully reflects the migration path of atmosphere-soil-groundwater. The three work together to improve the accuracy and dynamic adaptability of pollutant spatiotemporal diffusion simulation, providing a reliable simulation basis for risk assessment.

[0133] Furthermore, by using the S4 "Risk Assessment" (AHP-entropy weight method to balance subjective experience and objective data, quantitative grading and weighted calculation to quantify risk), the problems of "single weight and ambiguous grading" in traditional assessments can be solved. The combined weight avoids weight distortion caused by subjective bias or data noise, the quantitative grading makes the risk value comparable, and the weighted calculation ensures that the comprehensive impact of multiple indicators is reasonably quantified. Ultimately, the scientific nature of the risk assessment and the accuracy of high-risk area identification are derived, providing a clear direction for the identification of governance target areas.

[0134] Subsequent functional expansion methods further extend the technical effects: S5 "Digital Twin Pre-Drill" (building a mirror based on simulation results, loading parameters to correct extreme scenarios, and simulating the output effect of intervention schemes) can replace actual risk experiments with virtual mirrors, deriving the reduction of trial and error costs of intervention schemes and the improvement of governance decision-making efficiency; S6 "Goal-Oriented Pollution Attribution" (backpropagation to calculate basic contributions and priority adjustment of target dimensions) can derive the clarification of core pollution causes and the improvement of targeted governance measures through the correlation and matching of causes and targets, avoiding indiscriminate control; S7 "Federated Migration and Reinforcement Learning Closed Loop" (parameter server stores general parameters, fine-tuning + distillation to achieve cross-regional knowledge transfer, and PPO iterative optimization of parameters) can solve the problem of "limited data and difficult training" in newly built parks, deriving the ability of models to be quickly deployed across regions and dynamically optimized in the long term, and lowering the threshold for technical application.

[0135] In summary, from basic spatial and data preparation to core simulation and evaluation, and then to functional expansion and cross-regional reuse, the various levels of technical means form a logical closed loop. Ultimately, this allows for the objective derivation of the comprehensive technical effects of identifying high-risk pollution areas in industrial clusters: improved accuracy, optimized efficiency, enhanced adaptability, targeted governance, and convenient application. This effectively addresses the shortcomings of traditional technologies in terms of accuracy, dynamism, and reusability. The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-mentioned technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-mentioned technical features or their equivalents without departing from the aforementioned disclosed concept. For example, technical solutions formed by substituting the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. An intelligent auxiliary identification method for high-risk pollution areas in an industrial cluster site, characterized in that, The method comprises the following steps: a spatial gridding construction step, in which the industrial agglomeration area is divided into standard grids based on a preset coordinate system, each grid is assigned a unique identifier and spatial boundary information is recorded, and the grids at the boundaries of high-risk facility areas and environmentally sensitive areas are subjected to refinement processing; a multi-source data integration step, in which spatial basic data, enterprise facility layout data, pollution monitoring data and geological and hydrological data of the industrial agglomeration area are collected, the collected data are subjected to standardization processing, and then are mapped to the corresponding grids through spatial superposition or interpolation methods to form an association between the grids and the data; a three-dimensional correlation embedded graph neural network simulation step, in which a graph neural network model is constructed with the grids as nodes and the pollutant migration relationship between the grids as edges, three-dimensional features of the grid identifiers, pollution facility types and proportions, and pollutant concentrations are embedded into the model nodes, and the model is trained to realize the simulation of the temporal and spatial diffusion of pollutants in the soil-groundwater medium; a risk assessment step, in which an evaluation system is constructed based on four types of indexes of preset pollution source intensity, pollution migration channel, environmental vulnerability and historical record, the comprehensive risk value of each grid is calculated, the high, medium and low risk levels are divided according to the preset risk value interval, and the identification result of the high-risk pollution area is output.

2. The intelligent auxiliary identification method for the high-risk pollution area of the industrial agglomeration area site according to claim 1, characterized in that in the spatial gridding construction step, the grid refinement processing adopts a zoned hierarchical control mode, the standard grid size is maintained in the general industrial area and the grid boundary coordinate deviation is controlled within a preset range, the grid size of the high-risk concentration area is adjusted to be smaller than the standard grid size to adapt to the small-scale pollution source positioning requirement, and the grid size of the boundary of the environmentally sensitive area is further adjusted according to the protection level of the sensitive area to adapt to the boundary precision requirement of the pollutant diffusion simulation.

3. The intelligent auxiliary identification method for the high-risk pollution area of the industrial agglomeration area site according to claim 1, characterized in that in the risk assessment step, the weights of the four types of indexes are determined by a combination of a subjective weight determination method and an objective weight determination method, the secondary indexes under each index are quantitatively graded according to a preset grading standard, and the comprehensive risk value is calculated by weighted addition of the weights of each index and the corresponding score.

4. The intelligent auxiliary identification method for the high-risk pollution area of the industrial agglomeration area site according to claim 1, characterized in that the three-dimensional correlation embedded graph neural network simulation step further comprises a dynamic topology adjustment process, in which every interval of a preset time step, it is detected whether the environmental parameters corresponding to the grid have changed by more than a preset amplitude or whether there is a human intervention instruction, if any of the conditions is met, the topology adjustment is triggered, the connection weight between the grids is adjusted according to the detection result, and the adjacent connection of the pollution diffusion front grid is expanded from the original 4-way to 8-way.

5. The intelligent auxiliary identification method for the high-risk pollution area of the industrial agglomeration area site according to claim 4, characterized in that The three-dimensional correlation embedded graph neural network simulation step further comprises a multi-medium coupling simulation process, a preset migration flux calculation method is used to calculate the pollutant migration flux from the atmosphere to the soil and from the soil to the groundwater, the migration flux data is embedded into the graph neural network model node as a new feature, and the model is retrained after updating the node feature vector to realize the multi-medium coupling pollutant migration simulation of the atmosphere-soil-groundwater.

6. The intelligent auxiliary identification method for high-risk pollution areas in industrial agglomeration area sites according to claim 1, wherein, The multi-source data integration step further comprises a cross-modal data fusion process, feature extraction is performed on unstructured data such as unmanned aerial vehicle aerial image and inspection video to obtain image pollution index and facility risk index, and the two types of indexes and standardized pollution monitoring data are fused by a preset weighted fusion method to construct a cross-modal intermediate material, and the intermediate material is mapped to the corresponding grid to supplement the attribute information of the grid.

7. The intelligent auxiliary identification method for high-risk pollution areas in industrial agglomeration area sites according to claim 5, wherein, It further comprises a digital twin pre-play step, based on the multi-medium coupling simulation result, an industrial agglomeration area digital twin mirror image is constructed, parameters corresponding to extreme scenarios are loaded in the mirror image to correct the environmental parameters and pollution source intensity data, and multiple intervention schemes are simulated in the mirror image and the pollution control effect data corresponding to each scheme is output.

8. The intelligent auxiliary identification method for high-risk pollution areas in industrial agglomeration area sites according to claim 1, wherein, It further comprises a target-oriented pollution attribution step, the basic contribution value of different pollution causes to the high-risk grid is calculated by a graph neural network back propagation algorithm, the attribution dimension priority is defined according to a preset dynamic target label, the weight of each pollution cause is adjusted according to the priority, and the final contribution proportion of each cause to the high-risk grid is calculated, and the attribution result including the contribution proportion and key evidence is output in the form of a written report.

9. The intelligent auxiliary identification method for high-risk pollution areas in industrial agglomeration area sites according to claim 7 or 8, wherein, It further comprises a federal transfer and reinforcement learning closed loop step, a cross-park parameter server is constructed to store the parameters of the three-dimensional correlation embedded general graph neural network, the general parameters are loaded in the newly built park, and the parameters are fine-tuned through a small amount of local measured data, the fine-tuning loss function includes the local correlation embedding error and the general parameter deviation constraint, the knowledge of the mature park model is transferred to the new park model through knowledge distillation, the model parameters and intervention scheme parameters are adjusted, the dynamic target achievement is evaluated, and the parameters are adjusted through an iterative optimization algorithm to form a closed loop of target input to parameter iteration.

10. An intelligent auxiliary identification system for high-risk pollution areas in industrial agglomeration area sites, comprising a data acquisition subsystem and a server, wherein the data acquisition subsystem is used to acquire remote sensing images, surface change data, soil and groundwater environmental parameters, pollutant concentration data, and real-time operation state data of pollution facilities in the industrial agglomeration area. ​ The server is used for receiving data transmitted by the data acquisition subsystem and performing the intelligent auxiliary identification method of the high-risk pollution area of the industrial cluster area site according to any one of claims 1 to 9; The data transmission is realized between the data acquisition subsystem and the server through a preset communication mode.

Citation Information

Cited By

  • Soil heavy metal pollution ecological risk identification method and system

    CN121920685A