Pollutant tracing method, system and equipment based on multiple media and multiple links and media
By preprocessing and feature extraction of pollutant monitoring data, a transmission coupling relationship model is constructed and causal verification is performed. This solves the problem of pollutant transmission relationships in multiple media and multiple stages, and realizes the quantification, causation, spatialization and dynamism of pollutant source tracing, thereby improving the accuracy and efficiency of pollution control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies are insufficient to effectively analyze the pollutant transfer relationships across multiple media and stages, especially lacking quantitative means for indirect coupling effects across media and stages. Furthermore, traditional methods cannot adapt to the dynamic changes in causal relationships in complex systems, leading to unreliable verification results.
By preprocessing pollutant monitoring data, the main characteristic components representing the distribution and evolution of pollutants are extracted, a transmission coupling relationship model is constructed and jointly verified, including filling missing values using a spatiotemporal attention graph convolutional network, dimensionality reduction based on manifold learning algorithm and principal component analysis, training the transmission coupling relationship model with graph attention neural network, establishing a causal network skeleton and performing dynamic verification.
It achieves the quantification, causation, spatialization, and dynamization of pollution source tracing in target areas, improving the accuracy and efficiency of pollution control. It can accurately simulate the transmission path and coupling strength of pollutants, ensuring the scientific validity and reliability of the source tracing results.
Smart Images

Figure CN121725933A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pollution control technology for target areas, specifically to a method, system, equipment, and media for tracing pollutant sources based on multiple media and multiple stages. Background Technology
[0002] Environmental pollution in the target area is a typical complex system problem, with pollution sources encompassing multiple stages such as mining, industrial production, storage and transportation, solid waste stockpiling, and agricultural fertilization. Pollutants migrate and transform in various media, including groundwater, soil, and gases, exhibiting nonlinear coupling effects. For example, fluoride in mine wastewater can seep into groundwater and migrate to other areas, while lead-containing dust generated during storage and transportation may diffuse through the atmosphere and settle into distant soils, forming a complex cross-media, cross-regional transmission network. This multi-stage, multi-media, and spatiotemporally dynamic pollution characteristic makes precise source tracing and quantitative analysis extremely challenging.
[0003] To address this challenge, existing technologies typically employ two types of methods: one is statistical feature extraction methods, such as principal component analysis (PCA), used to identify correlations between pollutants; the other is methods based on causal inference and time-series testing, such as Granger causality tests and ADF unit root tests, used to determine causal relationships and time-series stationarity among variables. These methods can, to some extent, analyze local or static pollution conditions.
[0004] However, these existing technologies still have significant limitations when dealing with complex pollution systems in target areas: First, traditional methods such as PCA are linear models, making it difficult to capture complex nonlinear synergistic relationships between pollutants (such as a sudden drop in adsorption capacity due to concentration abrupt changes), and they cannot effectively identify the dominant transport pathways in multi-regional networks. Second, for indirect coupling effects across media and stages, traditional models lack effective quantification methods, often leading to an underestimation of the strength of interactions within the system. Finally, traditional causal and temporal verification methods are mostly static or global models, unable to adapt to the dynamic changes in transport pathways and causal strength over time, resulting in unreliable verification results for the behavior of complex systems. Therefore, there is an urgent need for a new method and system capable of systematically analyzing and dynamically verifying the transport coupling relationships of pollutants in multiple regions and media. Summary of the Invention
[0005] In view of this, it is necessary to provide a pollutant source tracing method, system, equipment and medium based on multiple media and multiple links, so as to solve the technical problem that it is difficult to effectively analyze the transmission relationship of pollutants in multiple media in the existing technology.
[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides a pollutant source tracing method based on multiple media and multiple stages, comprising: Pollutant monitoring data is preprocessed to obtain a standardized dataset; the pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized data; Based on the standardized dataset, the main characteristic components representing the distribution and evolution of pollutants are extracted; A transmission coupling relationship model is constructed based on the main characteristic components, and the transmission path and coupling strength coefficient of pollutants are obtained according to the transmission coupling relationship model. Based on the transmission path and coupling strength coefficient, a causal model is constructed and a joint verification is performed; when the verification passes, the pollution source tracing result is output.
[0007] In one possible implementation, the preprocessing of pollutant monitoring data to obtain a standardized dataset includes: Pollutant concentration data from multiple preset stages and multiple media are collected, and the corresponding stage attributes, media type, spatial coordinates and sampling time are labeled for each sampling point to obtain the pollutant monitoring data; Using a spatiotemporal attention graph convolutional network, combined with pollution contribution weights determined by the entropy weight method, missing values in the pollutant monitoring data are filled in; For the dimensional differences in pollutant concentrations in different media, standardization is performed based on the corresponding environmental quality standard limits to ensure that the standardized data for all media are within a preset mean range and a preset standard deviation range.
[0008] In one possible implementation, the extraction of key feature components characterizing the distribution and evolution of pollutants based on the standardized dataset includes: A spatial adjacency graph is constructed based on the spatial coordinates of the sampling points, and the shortest path distance between every two sampling points is calculated based on the spatial adjacency graph to obtain a spatial distance matrix characterizing the migration potential of the pollutants. Based on the spatial distance matrix, the standardized data is reduced in dimensionality using a manifold learning algorithm to obtain a manifold coordinate matrix mapped in the structural space. Based on the manifold coordinate matrix and the pollution contribution weights determined by the entropy weight method, a weighted manifold kernel function is constructed; A time warping algorithm is used to perform time alignment processing on the standardized data corresponding to each sampling point in the standardized dataset. Based on the aligned data, principal component analysis is performed according to the weighted manifold kernel function to extract principal components that meet the preset contribution rate threshold, and the extracted principal components are used as the main feature components.
[0009] One possible implementation also includes: The spatial distance matrix is analyzed to obtain the data manifold structure characteristics, and the corresponding manifold learning algorithm is selected based on the data manifold structure characteristics; the manifold learning algorithm includes the isometric feature mapping algorithm and the Laplacian feature mapping algorithm. When the spatial distance matrix satisfies the global isometry property, the manifold coordinate matrix is obtained by the isometry feature mapping algorithm. When the spatial distance matrix exhibits local nonlinear characteristics, the Laplacian eigenmap algorithm is used to obtain the manifold coordinate matrix.
[0010] In one possible implementation, the step of constructing a transport coupling relationship model based on the main characteristic components, and obtaining the pollutant transport path and coupling strength coefficient according to the transport coupling relationship model, includes: Based on the main feature components, a multi-media coupling transfer graph containing multiple nodes is constructed, wherein each node corresponds to a combination of one of the links, one of the media, and one sampling point; Calculate the pollutant migration probability between any two nodes, and use the pollutant migration probability as the initial weight of the connecting edge between the two nodes; A graph attention neural network is constructed based on the multi-media coupling transfer graph, and the transfer coupling relationship model is obtained by training the graph attention neural network; during the training process of the graph attention neural network, the pollutant migration probability and mass conservation loss are used as constraints. Based on the aforementioned transmission coupling relationship model, the transmission path of pollutants between the nodes is obtained according to the standardized dataset. Based on the aforementioned transmission coupling relationship model, the state change of the node after the control variable is changed is obtained, and the coupling strength coefficient is calculated based on the state change.
[0011] In one possible implementation, the step of constructing a causal model and performing joint verification based on the transmission path and coupling strength coefficient includes: The transmission path is used as a causal relationship to construct a causal network skeleton, and the coupling strength coefficient is set as the initial weight of the causal network skeleton. A structural equation model is established based on the causal network skeleton after weight setting; the time-varying coefficient matrix of the structural equation model is obtained by estimating the standardized dataset; Calculate the stability index of each candidate causal path in the structural equation model; The overall validation loss function is constructed based on the stability index, the mass conservation loss of the trained transitive coupling relationship model, and the data reconstruction error of the extracted main feature components, so as to calculate the overall validation loss value. When the overall verification loss value is lower than the judgment threshold set according to the historical verification set, the verification is determined to be successful.
[0012] One possible implementation also includes: Calculate the degree of agreement between the pollution source tracing results and the actual investigation results; Using a Bayesian optimization algorithm, with the goal of maximizing the fit, the optimal combination of weight coefficients in the overall validation loss function is calculated. The overall validation loss function is updated based on the optimal combination of weight coefficients.
[0013] Secondly, the present invention also provides a pollutant tracing system based on multiple media and multiple stages, comprising: The data acquisition and preprocessing module is used to preprocess pollutant monitoring data to obtain a standardized dataset; the pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized data; The feature extraction module is used to extract the main feature components characterizing the distribution and evolution of pollutants based on the standardized dataset; The transmission and coupling modeling module is used to construct a transmission and coupling relationship model based on the main characteristic components, and to obtain the transmission path and coupling strength coefficient of pollutants according to the transmission and coupling relationship model; The causal verification and source tracing module is used to construct a causal model and perform joint verification based on the transmission path and coupling strength coefficient; when the verification passes, the pollution source tracing result is output.
[0014] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the pollutant tracing method based on multiple media and multiple links described in any of the above implementations.
[0015] Fourthly, the present invention also provides a computer-readable storage medium for storing a computer-readable program or instructions, which, when executed by a processor, can implement the steps in the pollutant tracing method based on multiple media and multiple links described in any of the above implementations.
[0016] The beneficial effects of this invention are as follows: The pollutant source tracing method based on multiple media and multiple stages provided by this invention first achieves quantification, causality, spatialization, and dynamism of pollution source tracing in target areas through a four-step closed loop: precise data preprocessing, in-depth feature extraction, physicalization of transmission modeling, and dynamic causal verification. This improves the accuracy and efficiency of pollution control, resulting in significant environmental, social, and economic benefits. Furthermore, by preprocessing pollutant monitoring data, a high-quality, dimensionless, and complete standardized dataset is provided for subsequent analysis. Differentiated standardization for different media and / or different stages eliminates dimensional interference, ensuring that pollutant data from different media and / or different stages are at the same comparable level, enabling cross-media comprehensive analysis. Moreover, by extracting the main features characterizing the distribution and evolution of pollutants, the inherent nonlinear structure and spatial continuity in the data can be discovered and extracted, thereby capturing complex pollution patterns that traditional methods cannot identify. Furthermore, by constructing a transmission coupling relationship model, the migration process of pollutants in complex environments can be simulated, thereby obtaining accurate pollutant transmission paths and coupling strength coefficients. This ensures that the transmission paths learned by the model conform to environmental migration laws and avoids spurious correlations that may arise from purely data-driven approaches. Moreover, traditional causal verification methods are mostly static or global, unable to adapt to the dynamic evolution of causal relationships over time in complex systems, resulting in unreliable verification results. This application uses joint verification of causal models to conduct dynamic and multi-dimensional rigorous verification of discovered causal relationships, ultimately ensuring the scientific validity of the output source tracing results. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A schematic flowchart of an embodiment of the pollutant source tracing method based on multiple media and multiple links provided by the present invention; Figure 2 For the present invention Figure 1 A schematic diagram of an embodiment of S101; Figure 3 For the present invention Figure 1 A schematic diagram of an embodiment of S102; Figure 4 For the present invention Figure 1 A schematic diagram of an embodiment of S103; Figure 5 For the present invention Figure 1A schematic diagram of an embodiment of S104; Figure 6 A schematic flowchart of another embodiment of the pollutant tracing method based on multiple media and multiple links provided by the present invention; Figure 7 A schematic diagram of an embodiment of the pollutant tracing system based on multiple media and multiple links provided by the present invention; Figure 8 A schematic diagram of an embodiment of the electronic device provided by the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0020] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0021] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.
[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0023] Before demonstrating the embodiments, the following terms will be explained.
[0024] DTW (Dynamic Time Warping).
[0025] KPCA (Kernel Principal Component Analysis).
[0026] LIME (Local Interpretable Model-agnostic Explanations).
[0027] SHAP (SHapley Additive exPlanations).
[0028] WDM-PCA (Weighted Dynamic Manifold PCA).
[0029] Isomap (Isometric Mapping).
[0030] PI-ST-GT (Physics-Informed Spatio-Temporal Graph Network).
[0031] SCM-JV (Structural Causal Model Joint Validation).
[0032] ST-AGCN (Spatio-Temporal Attention Graph Convolutional Network).
[0033] Delaunay (Delaunay Triangulation).
[0034] RMSE (Root Mean Square Error).
[0035] MMCT-Graph (Multi-Medium Coupled Transfer Graph).
[0036] TVC-SI (Time-Varying Causal Stability Index).
[0037] This invention provides a method, system, equipment, and medium for tracing pollutants from multiple media and multiple stages, which will be described below.
[0038] Figure 1 The following is a schematic flowchart of an embodiment of the pollutant source tracing method based on multiple media and multiple stages provided by the present invention, as shown below. Figure 1As shown, the pollutant source tracing method based on multiple media and multiple stages includes: S101. Preprocess the pollutant monitoring data to obtain a standardized dataset; the pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized data.
[0039] It should be noted that the target area can include areas such as mines, chemical industrial parks, landfills, and farmland contaminated with heavy metals. Mining areas can include mining faces, waste rock dumps, tailings ponds, spoil heaps, ore dressing plants, and transportation roads. A multi-point, multi-media monitoring network is deployed within the target area. Pollutant monitoring data is collected through this network, and each pollutant monitoring data point undergoes preprocessing to obtain corresponding standardized data. Preprocessing may include data cleaning, time alignment, spatial registration, normalization, and standardization. The pollutant monitoring data from all stages and all media are then aggregated and corresponding to the standardized data to obtain a standardized dataset.
[0040] S102. Based on the standardized dataset, extract the main characteristic components that characterize the distribution and evolution of pollutants.
[0041] It should be noted that the algorithm described in the following examples can be used to extract key feature components characterizing the dominant spatial distribution patterns and temporal evolution trends or laws of pollutants from massive, high-dimensional standardized data. These key feature components include a few feature information elements that best represent the pollution status and evolution patterns of the entire target area.
[0042] S103. Construct a transmission coupling relationship model based on the main characteristic components, and obtain the transmission path and coupling strength coefficient of pollutants according to the transmission coupling relationship model.
[0043] It's important to note that a transport path refers to the physical migration route of pollutants from their source, traversing different media (water, soil, air) and crossing various human-induced stages (mining, storage, transportation, production, etc.). Transport paths are not fixed. For example, during the rainy season, groundwater flow accelerates, and pollutants may primarily migrate via groundwater; while during the dry season with strong winds, they may mainly be transported via atmospheric particulate matter. Transport coupling models can capture these dynamically changing pollutant transport paths over time. A complete transport path typically involves multiple media and stages. Example path: Mining (stage) → Slag leaching into groundwater (media) → Flowing with groundwater to downstream farmland (space) → Absorbed by crops into the soil (media) → Transported to storage and transportation stages via agricultural products (stage).
[0044] The coupling strength coefficient is a quantitative indicator used to measure the degree of mutual influence and interaction between different pollution sources, processes, or media. It describes not the migration of the pollutants themselves, but rather the intensity of the "side effects" or "interactions" generated during the migration process. The coupling strength coefficient can accurately quantify indirect, non-linear effects. For example, fluoride (source A) in mine wastewater seeps into the soil, altering its pH. This change significantly enhances the activity and toxicity of lead (source B), which was originally stable in the soil, making it easier for plants to absorb or for it to seep into groundwater. Here, there is a strong coupling effect between fluoride and lead.
[0045] The main characteristic components are input into a preset transmission coupling relationship model to simulate the transmission and coupling process of pollutants in multiple links and multiple media, driving it to learn the migration and interaction laws of pollutants in different links and different media, thereby resolving the transmission path and coupling strength coefficient of pollutants.
[0046] S104. Based on the transmission path and coupling strength coefficient, construct a causal model and perform joint verification; when the verification passes, output the pollution source tracing result.
[0047] It should be noted that the dynamic transmission path output by S103 serves as the structural framework of the causal relationship, and the coupling strength coefficient is used as important prior information to measure the strength of the causal relationship. Based on this, time-varying structural equation modeling and other causal models are established. The core of these models lies in acknowledging and characterizing the dynamic evolution of causal relationships in the real environment; that is, the impact path and intensity of pollutants change over time (e.g., seasonal variations, production cycles), rather than remaining static. Joint verification is a multi-dimensional and comprehensive statistical verification process that cross-validates the robustness and rationality of pollutant transmission paths from multiple perspectives. Specifically, it verifies whether a real, stable, and physically consistent causal driving relationship exists between the two nodes connected by a transmission path (i.e., the combination of two "link-medium-sampling point" pairs).
[0048] In summary, the pollutant source tracing method based on multiple media and stages provided in this invention first achieves quantification, causality, spatialization, and dynamism in pollution source tracing of target areas through a four-step closed loop (S101 to S104): precise data preprocessing, in-depth feature extraction, physical transmission modeling, and dynamic causal verification. This improves the accuracy and efficiency of water pollution control, resulting in significant environmental, social, and economic benefits. Furthermore, by preprocessing pollutant monitoring data, a high-quality, dimensionless, and complete standardized dataset is provided for subsequent analysis. Differentiated standardization for different media and / or stages eliminates dimensional interference, ensuring that pollutant data from different media and / or stages are at the same comparable scale, enabling cross-media comprehensive analysis. Moreover, by extracting the main features characterizing the distribution and evolution of pollutants, the inherent nonlinear structure and spatial continuity in the data can be discovered and extracted, thereby capturing complex pollution patterns that traditional methods cannot identify. Furthermore, by constructing a transmission coupling relationship model, the migration process of pollutants in complex environments can be simulated, thereby obtaining accurate pollutant transmission paths and coupling strength coefficients. This ensures that the transmission paths learned by the model conform to environmental migration laws and avoids spurious correlations that may arise from purely data-driven approaches. Moreover, traditional causal verification methods are mostly static or global, unable to adapt to the dynamic evolution of causal relationships over time in complex systems, resulting in unreliable verification results. This application uses joint verification of causal models to conduct dynamic and multi-dimensional rigorous verification of discovered causal relationships, ultimately ensuring the scientific validity of the output source tracing results.
[0049] In order to integrate multi-source, heterogeneous, and incomplete raw environmental monitoring data into a unified, complete, and comparable standardized dataset.
[0050] In some embodiments of the present invention, such as Figure 2 As shown, step S101 includes: S201. Collect pollutant concentration data from multiple preset stages and multiple media, and label each sampling point with the corresponding stage attribute, media type, spatial coordinates and sampling time to obtain the pollutant monitoring data.
[0051] It should be noted that, for ease of explanation, the target area will be assumed to be a mining area for illustrative purposes throughout the following description. The five key stages within a mining area include mining operations (e.g., mine drainage, slag heaps), production and processing (e.g., ore dressing plants, smelting workshops), storage and transportation (e.g., ore stockpiles, transport roads), solid waste disposal (e.g., tailings ponds, waste rock dumps), and agricultural fertilization (e.g., surrounding farmland, orchards). The three key media within a mining area include groundwater, soil, and gas. Pollutant concentration data for multiple stages and media include: groundwater pollutant concentrations (fluoride, sulfate, nitrate, ammonia nitrogen, total phosphorus, lead); soil pollutant concentrations (lead, cadmium); and gaseous pollutant concentrations (sulfur dioxide, nitrogen oxides).
[0052] Of course, the medium can also be set according to different scenarios, including surface water (such as open water bodies such as rivers, lakes, ponds, and reservoirs), sediments (such as silt and sediments at the bottom of rivers, lakes, oceans, and irrigation canals), and organisms (such as crops such as rice and vegetables, as well as indicator plants such as moss and lichen, and aquatic organisms such as fish and shellfish, and terrestrial indicator organisms such as earthworms).
[0053] After obtaining the pollutant concentration data, record the spatial coordinates (latitude and longitude) and timestamp (year, month, day, hour, minute) of each sampling point. Construct the original dataset according to a unified format to ensure that each data point can be traced back to a specific process, medium, and spatiotemporal location. For example, process number-medium type-longitude-latitude-timestamp-pollutant concentration.
[0054] S202. Using a spatiotemporal attention graph convolutional network, combined with the pollution contribution weight determined by the entropy weight method, missing values in the pollutant monitoring data are filled in.
[0055] It should be noted that the entropy weight method is an objective weighting method that determines weights based on the degree of variation in pollutant data at each stage. Due to factors such as sensor malfunction and severe weather, the proportion of missing pollutant monitoring data may be between 5% and 10%. The pollution contribution weights determined by the entropy weight method specifically involve standardizing the pollution contribution data (such as pollutant concentration and emissions) at each stage. The entropy value of each stage is calculated using the information entropy formula. The pollution contribution weight of each stage is then calculated based on the entropy values.
[0056] S203. Based on the dimensional differences in pollutant concentrations in different media, standardization processing is performed to ensure that the standardized data of all media are within a preset mean range and a preset standard deviation range.
[0057] It should be noted that because the environmental quality standard limits differ for different media (groundwater, soil, and gas), and the background values of pollutants also vary, different standardization methods are required for different media. For each media, the standard limit is obtained according to its corresponding environmental quality standard (e.g., GB / T14848 for groundwater, GB15618 for soil, and GB3095 for gas). Additionally, for groundwater, regional background values also need to be considered.
[0058] Each pollutant in each medium is standardized to convert it into a dimensionless value, and then adjusted to a preset mean and standard deviation. Different standardization methods are used for different media with varying data characteristics. For example, taking a mining area as an example again: Groundwater pollutant standardization involves standardizing each pollutant in the groundwater: Standardized value = (Measured concentration - Background value) / (Standard limit - Background value). The background value adopts the regional geochemical background value, which can be determined based on local groundwater background value survey data. The standard limit adopts the Class III water standard limit in the "Groundwater Quality Standard" (GB / T14848).
[0059] Soil pollutant standardization involves standardizing each pollutant in the soil: Standardized value = Measured content / Standard limit, where the standard limit is the screening value in the "Soil Environmental Quality Construction Land Soil Pollution Risk Control Standard" (GB36600).
[0060] Standardization of gaseous pollutants begins with converting their concentration units to the corresponding national standard units (e.g., mg / m³). For example, converting ppm to mg / m³ is calculated as: mg / m³ = (ppm × molecular weight) / 22.4 × (273 / (273 + T)) × (P / 101.325), where T is temperature (degrees Celsius) and P is pressure (kPa). Standard conditions (25℃, 101.325 kPa) are typically used for simplified calculation. Each pollutant in the gas is then standardized: Standardized value = Measured concentration / Standard limit. The standard limit is the 24-hour average concentration limit from the "Ambient Air Quality Standard" (GB3095).
[0061] After the standardization process described above, the contaminant concentrations for each medium have been converted to dimensionless values, but their distributions may differ. Further adjustments are needed to ensure that the standardized data for all media have the same preset mean range μ (e.g., 0.5) and preset standard deviation range σ (e.g., 0.2). The adjustment method is as follows: Let the original standardized value be x, and the adjusted value be y. Here, μ and σ are the mean and standard deviation of the original standardized values of all pollutants at all sampling points of the medium, respectively. Note that when calculating μ and σ, it is necessary to calculate them separately for each type of medium, i.e., groundwater, soil, and gas, and then adjust them separately.
[0062] Of course, for some pollutants, there may be no standard limits or background values, requiring handling based on the actual situation (e.g., using other standards or reference values). If the measured concentration of a pollutant is lower than the background value, the standardized value may be negative, which is permissible because subsequent adjustments will map it to the new range. If a standard limit for a pollutant is missing, other relevant standards or expert empirical values can be considered. Ultimately, the standardized data for all media should have a mean of μ and a standard deviation of σ to eliminate dimensional interference.
[0063] In this embodiment, by integrating spatiotemporal features and pollution contribution weights, the accuracy of missing value imputation is improved. Media-specific standardization solves the problem of inconsistent dimensions between different media. Differential processing based on national environmental quality standards ensures the standardization process is scientifically sound and can eliminate dimensional differences in pollutant concentrations across different media, allowing for direct comparison and computation of data from different media. For example, if the migration probability of pollutants is high between certain stages, the attention weights between nodes in these stages should be greater. Furthermore, the data is unified to the same statistical magnitude (mean 0.5, standard deviation 0.2), laying the foundation for subsequent multi-media collaborative analysis. Furthermore, each sampling point is labeled with its corresponding stage attribute, media type, spatial coordinates, and sampling time, enabling precise location and tracing of pollution sources. Furthermore, the pollution contribution weights objectively determined by the entropy weight method reflect the actual pollution contribution of each stage, are insensitive to missing patterns, and can handle random and continuous missing data, improving the accuracy of processing missing data from major pollution stages. Furthermore, by using a spatiotemporal attention graph convolutional network, both spatial and temporal correlations are considered simultaneously. The attention mechanism focuses on important neighboring nodes, resulting in higher filling accuracy than traditional methods (such as KNN and spatiotemporal kriging). This allows the model to be constrained by physical laws during learning, avoiding physical inconsistencies caused by purely data-driven approaches. It can adaptively learn the relationships between nodes without needing to predefine fixed spatial neighborhoods or time windows.
[0064] To accurately extract key characteristic components reflecting the spatial migration patterns and driving forces of pollutants from complex spatiotemporal pollution data, in some embodiments of the present invention, such as... Figure 3 As shown, step S102 includes: S301. Construct a spatial adjacency graph based on the spatial coordinates of the sampling points, and calculate the shortest path distance between every two sampling points based on the spatial adjacency graph to obtain a spatial distance matrix characterizing the migration potential of the pollutants.
[0065] It's important to note that Delaunay triangulation is used to construct the spatial adjacency graph. Delaunay triangulation triangulates a set of planar points such that the circumcircle of any triangle contains no other points. This allows each sampled point to be treated as a node, with the edges between points determined by the Delaunay triangulation results. On the constructed adjacency graph, Dijkstra's algorithm (or Floyd-Warshall algorithm) is used to calculate the shortest path distance between every two nodes. Assuming the edge weights are equal to the Euclidean distance between the two points, the shortest path distances are stored in a matrix, which represents the spatial distance matrix characterizing the pollutant migration potential. This spatial distance matrix effectively characterizes the pollutant migration potential because pollutant migration in the environment is often not a straight line but follows paths such as topography and hydrology. Delaunay triangulation can simulate this path to some extent (because adjacent points are connected by triangle edges, and the shortest path follows these edges).
[0066] S302. Based on the spatial distance matrix, the standardized data is reduced in dimensionality using a manifold learning algorithm to obtain a manifold coordinate matrix mapped in the structural space.
[0067] It should be noted that, having already obtained the spatial distance matrix, the next step is to use manifold learning algorithms to reduce the dimensionality of the high-dimensional standardized data (standardized data with multiple pollutant concentrations at each sampling point) to a low-dimensional manifold space, obtaining the manifold coordinate matrix. The choice of manifold learning algorithm depends on the geometric properties of the spatial distance matrix. The main manifold learning algorithms include isomaps and Laplacian eigenmaps. The choice of manifold learning algorithm type is based on whether the spatial distance matrix maintains global isometry. According to the selected manifold learning algorithm, the standardized data (pollutant concentration vectors at each sampling point) is reduced in dimensionality to obtain a low-dimensional manifold coordinate matrix mapped onto the structural space.
[0068] S303. Construct a weighted manifold kernel function based on the manifold coordinate matrix and the pollution contribution weight determined by the entropy weight method.
[0069] It should be noted that the manifold coordinate matrix represents the coordinates of each data point in a low-dimensional manifold space. The pollution contribution weights are weight vectors for each stage determined using the entropy weighting method; the length of the pollution contribution weights is equal to the number of stages. Each node corresponds to a combination of stage-medium-sampling point. After calculating each pollution contribution weight using the improved entropy weighting method, the weighted manifold kernel function constructed based on the manifold coordinate matrix is: ; in, To calculate the weighted manifold similarity between two data points X and Y, and The coordinates of data points X and Y after being mapped to a low-dimensional manifold space through manifold learning (such as Isomap). The pollution contribution weights determined by the entropy weight method, This indicates element-wise multiplication. The bandwidth parameter of the Gaussian kernel function controls the rate at which similarity decays with distance. This is a constant term used to adjust the numerical stability of the kernel function. The degree is polynomial, controlling the nonlinearity of the weights' influence. This weighted manifold kernel function can simultaneously measure the similarity of data points in the manifold space and the importance of their respective components.
[0070] S304. Time warping algorithm is used to perform time alignment processing on the standardized data corresponding to each sampling point in the standardized dataset.
[0071] It should be noted that the pollutant concentration time series from different sampling points are not synchronized on the timeline due to factors such as monitoring start time, sampling frequency, or data transmission delays. Their sequence lengths may differ, and the specific time points of key events (such as pollution peaks) may also be misaligned. From the standardized dataset, the concentration time series of a single pollutant is extracted for each sampling point. For example: the groundwater fluoride concentration sequence for sampling point A = [C1, C2, C3, ..., C...]. m For example: the groundwater fluoride concentration sequence at sampling point B = [D1, D2, D3, ..., D...]. n Note that alignment is typically required for each contaminant separately. The core idea of the DTW algorithm is to find the optimal alignment path between two sequences, even if their lengths and speeds differ. Based on DTW, the normalized data corresponding to each sampling point in the normalized dataset is time-aligned. The average correlation coefficient between the aligned sequences can be calculated, and the aligned sequences can be visualized to ensure that key features such as contaminant peaks are aligned on the time axis.
[0072] S305. Based on the aligned data, perform principal component analysis according to the weighted manifold kernel function, extract principal components that meet the preset contribution rate threshold, and use the extracted principal components as the main feature components.
[0073] It should be noted that weighted dynamic manifold principal component analysis (WDM-PCA) is used to perform principal component analysis, which involves calculating the kernel values of all sample pairs using a weighted manifold kernel function to form a kernel matrix. To perform principal component analysis in the feature space, the data in the feature space needs to be centered, which can be achieved by centering the kernel matrix. Eigenvalue decomposition is performed on the centered kernel matrix, selecting the eigenvectors corresponding to the first k principal components. These eigenvectors are the principal component directions in the feature space. The centered kernel matrix is then projected onto these principal component directions to obtain the principal components that satisfy the preset contribution rate threshold. In other words, PCA (or kernel KPCA) is performed on the standardized data to output a series of principal components and their corresponding eigenvalues. The variance contribution rate = eigenvalue of the principal component / sum of the eigenvalues of all principal components. The contribution rates of the first k principal components are accumulated sequentially. Accumulation begins with the first principal component and continues until the cumulative variance contribution rate first exceeds the preset contribution rate threshold (e.g., 90%). The corresponding k value at this point is the number of principal components to be extracted. These k principal components are the main feature components that can represent most of the information in the original data.
[0074] In this embodiment, a combination of weighted manifold kernel function and principal component analysis achieves efficient feature extraction from complex pollutant data, providing high-quality feature input for subsequent transmission coupling modeling and pollution source tracing. The shortest path distance based on the spatial adjacency graph accurately reflects the actual migration path of pollutants, improving the accuracy of pollution diffusion pattern recognition under complex terrain conditions compared to traditional Euclidean distance. Furthermore, the manifold learning algorithm effectively captures the essential low-dimensional structure within high-dimensional data, enhancing the ability to handle nonlinear relationships and reducing computational complexity. Furthermore, the low-dimensional manifold coordinate matrix effectively characterizes the spatial distribution pattern of pollutants. The weighted manifold kernel function integrates spatial structure and pollution contribution weight information, and the adaptive adjustment mechanism of the kernel function ensures adaptability to different pollution scenarios. Multi-dimensional information fusion makes feature extraction more comprehensive and accurate. Furthermore, by constructing a spatial adjacency graph and using the shortest path calculation method, a spatial distance matrix that accurately characterizes the migration potential of pollutants is obtained, reflecting the actual migration path of pollutants in the environment more accurately than a simple straight-line distance. By considering spatial adjacency, the impact of natural conditions such as topography and hydrology on pollutant migration is captured, providing higher quality and more meaningful input features, thereby significantly improving the accuracy of pollution source tracing.
[0075] In order to adaptively select the optimal manifold learning algorithm based on the inherent spatial structure characteristics of the data, so as to accurately capture the nonlinear features of pollutant distribution.
[0076] In some embodiments of the present invention, it further includes: The spatial distance matrix is analyzed to obtain the data manifold structure characteristics, and the corresponding manifold learning algorithm is selected based on the data manifold structure characteristics; the manifold learning algorithm includes the isometric feature mapping algorithm and the Laplacian feature mapping algorithm. When the spatial distance matrix satisfies the global isometry property, the manifold coordinate matrix is obtained by the isometry feature mapping algorithm. When the spatial distance matrix exhibits local nonlinear characteristics, the Laplacian eigenmap algorithm is used to obtain the manifold coordinate matrix.
[0077] It should be noted that the stress index of the spatial distance matrix is calculated to evaluate the global isometry. ,in, d ij The original distances in the spatial distance matrix. δ ij The Euclidean distance is the distance after low-dimensional embedding. Multidimensional scaling (MDS) is used to evaluate global isometry; that is, MDS is applied to the spatial distance matrix to obtain the embedded coordinates. Stress is calculated based on the Euclidean distance of the embedded coordinates and the original spatial distance matrix. If Stress is less than a target threshold (e.g., 0.15), it indicates that the data is approximately isometric globally, and the Isomap algorithm is a suitable manifold learning algorithm. If Stress is greater than or equal to the target threshold (e.g., 0.15), the Laplacian eigenmap algorithm is selected. Alternatively, the local curvature of the data can be calculated. When the local curvature is greater than a set threshold, it indicates significant local nonlinearity, making Laplacian eigenmap suitable.
[0078] The Isomap algorithm execution flow is as follows: Based on the spatial distance matrix, find k nearest neighbors (usually k=8-12) for each sampling point to construct an unweighted neighborhood graph. Use Dijkstra's algorithm on the unweighted neighborhood graph to calculate the shortest path distance between all pairs of nodes; this distance is the geodesic distance, which forms a new N×N geodesic distance matrix D_G. Calculate the centering matrix based on the geodesic distance matrix D_G. ,in Perform eigenvalue decomposition on the centered matrix B, and take the first d largest eigenvalues ( The manifold coordinate matrix is derived from the first d largest eigenvalues and their corresponding eigenvectors v1, v2, ..., vd. Each row represents the coordinates of a sampling point in a d-dimensional low-dimensional structure space.
[0079] The Laplacian eigenmap algorithm execution flow is as follows: Based on the spatial distance matrix, an adjacency graph is constructed using the k-nearest neighbor or ε-ball method, and the weights of the edges connecting each node in the adjacency graph are calculated using a hot kernel function. Calculate the degree matrix D (a diagonal matrix). The Laplacian matrix L is calculated based on the degree matrix D. The Laplacian matrix L is solved, and the eigenvectors corresponding to the 2nd to (d+1)th smallest non-zero eigenvalues are taken from the solution structure to form the manifold coordinate matrix.
[0080] In this embodiment, the spatial distance matrix contains geospatial information and potential pollutant migration path information. Based on the characteristics of the data manifold structure, a corresponding manifold learning algorithm is selected to transform the spatial distance matrix into a pure manifold coordinate matrix representing relative positions and structural relationships. This allows subsequent kernel function calculations to be based on the intrinsic structural similarity of the data, rather than apparent numerical similarity, ensuring optimal feature extraction results under different data characteristics. This provides high-quality feature input for subsequent pollution source tracing analysis, significantly improving the performance and reliability of the entire system. Furthermore, the dimensionality-reduced manifold coordinate matrix has extremely low dimensions (typically <10 dimensions), and each dimension represents a macroscopic pattern of spatial distribution of pollutants (e.g., "distribution along rivers," "diffusion pattern from mining center"), reducing data processing volume and improving the efficiency of pollutant source tracing.
[0081] To construct a dynamic graph network constrained by physical laws, and to simultaneously identify the migration paths of pollutants in multiple stages and media, and quantify their indirect coupling effects. In some embodiments of the present invention, such as Figure 4 As shown, step S103 includes: S401. Construct a multi-media coupling transfer graph containing multiple nodes based on the main feature components, wherein each node corresponds to a combination of one of the links, one of the media, and one sampling point.
[0082] It should be noted that each node represents a specific "stage-medium-sampling point" combination. For example, a node could be "mine-groundwater-sampling point 1", "production-soil-sampling point 2", etc. A multi-medium coupled transport graph G=(V, E) is constructed, where V is the set of nodes and E is the set of directed edges. The direction of the edges indicates the possible migration direction of the pollutants (e.g., from the mining stage to the production stage).
[0083] S402. Calculate the pollutant migration probability between any two nodes, and use the pollutant migration probability as the initial weight of the connecting edge between the two nodes.
[0084] It should be noted that the probability of pollutant migration between any two nodes is calculated based on physical models (such as Darcy's law and diffusion equations). For example, for groundwater, Darcy's law is used to calculate the water flow velocity and direction, and then the probability of pollutant migration with the water flow is estimated; for soil and gas, diffusion equations are used to consider the concentration gradient and diffusion coefficient, and then the probability of pollutant migration with the soil or airflow is estimated. The calculated pollutant migration probability is used as the initial weight of the corresponding edge in the directed graph, and the initial weight reflects the physical probability of pollutant migration between nodes.
[0085] S403. Construct a graph attention neural network based on the multi-media coupling transfer graph, and train the graph attention neural network to obtain the transfer coupling relationship model; during the training of the graph attention neural network, the pollutant migration probability and mass conservation loss are used as constraints.
[0086] It should be noted that a Graph Attention Neural Network (PI-ST-GT) is constructed, which contains multiple graph attention layers. The initial feature vector of each node is the principal component extracted in steps S301-S305. During training, the contaminant migration probability calculated in step S402 is used as a constraint on the attention mechanism. Specifically, when calculating the attention coefficient, in addition to the similarity between node features, the contaminant migration probability is also incorporated as prior knowledge. Simultaneously, a mass conservation loss is introduced as part of the loss function to ensure that the transmission relationships learned by the graph attention neural network conform to mass conservation. The graph attention neural network is trained using a standardized dataset, and its parameters are optimized through backpropagation, enabling the graph attention neural network to learn the contaminant transmission relationships between nodes.
[0087] A multi-media coupled transfer graph is input into a graph attention neural network (PI-ST-GT). A physically constrained attention mechanism is used, where the attention coefficients between nodes i and j are... The calculation is as follows: ; in, It is the attention coefficient of node i to node j. It is an activation function. It is a learnable attention weight vector. It is a learnable linear transformation matrix used to generate queries. It is a learnable linear transformation matrix used to generate keys. It is the feature vector of the node. It is the feature vector of node j. It's a vector concatenation operation. It is a feasible option (i.e., the probability of pollutant migration) calculated based on a physical model. It is a hyperparameter that controls the strength of physical constraints. It is a constant.
[0088] In model training, a mass conservation loss function is introduced. As a constraint: ; in, It is the sum of the mass flow rates of all pollutants entering a certain node at time t. yes, It is the sum of the mass flow rates of all pollutants flowing out of a certain node at time t. It represents the change in the mass of pollutants stored within this node at time t.
[0089] S404. Based on the aforementioned transmission coupling relationship model, the transmission path of pollutants between the nodes is obtained according to the standardized dataset.
[0090] It should be noted that: A standardized dataset is input into a trained graph attention neural network (i.e., a trained graph attention neural network) to obtain the output features of each node. By analyzing the attention weights in the graph attention neural network, edges with weights exceeding a preset threshold are extracted to form dynamic transport paths for pollutants. These transport paths represent the main migration routes of pollutants between various stages, media, and sampling points. The process of obtaining transport paths from the standardized dataset is as follows: the standardized dataset is input into the trained graph attention neural network to calculate the attention weight matrix between nodes. Based on the attention weight matrix, edges connecting nodes with weights exceeding a preset threshold are extracted to form dynamically changing transport paths for pollutants between nodes.
[0091] S405. Based on the transmission coupling relationship model, obtain the state change of the node after changing the control variable, and calculate the coupling strength coefficient according to the state change.
[0092] It should be noted that: After sequentially shielding a specific link or medium (setting its node characteristics to zero or changing them), the transitive coupling model is rerun, and the changes in the states of other nodes are observed. The changes in the node states (e.g., pollutant concentration characteristics) under normal and shielded conditions are recorded. The coupling strength coefficient is calculated based on these state changes. Specifically, the complete standardized dataset is input into the transitive coupling model to obtain the baseline pollution state feature vector for each node. After shielding all node data related to a specific link or medium, the standardized dataset is re-inputted into the transitive coupling model to obtain the post-shielding pollution state feature vector for each node. The degree of difference between the baseline pollution state feature vector and the post-shielding pollution state feature vector is compared. Based on this difference, the coupling strength coefficient of the shielded link or medium on other links or media is calculated. The coupling strength coefficient can be calculated using the following formula: Coupling strength coefficient = (Normal state value - Post-shielding state value) / Normal state value. The coupling strength coefficient reflects the degree of influence of the shielded link or medium on other links or media.
[0093] In this embodiment, by combining a physical model (transfer probability, mass conservation) with a graph attention neural network, both physical laws are considered, and complex nonlinear relationships can be learned from data, improving the model's accuracy and interpretability. Furthermore, by identifying important transmission paths through graph attention weights, the migration patterns of pollutants in complex environments can be dynamically reflected. Furthermore, through shielding experiments, the interactions between different stages and media can be precisely quantified, providing targeted decision support for pollution control. Furthermore, physical constraints during training ensure that the transmission relationships learned by the model conform to physical laws, improving the model's generalization ability.
[0094] To ensure the scientific validity and reliability of source tracing conclusions, multi-dimensional joint verification of dynamic and complex candidate pollution causal relationships is conducted. In some embodiments of this invention, such as... Figure 5 As shown, step S104 includes: S501. Construct a causal network skeleton using the transmission path as a causal relationship, and set the coupling strength coefficient as the initial weight of the causal network skeleton.
[0095] It should be noted that the dynamic propagation paths output from the PI-ST-GT model are directly used as candidate causal relationships to construct a directed causal graph. Nodes in the directed causal graph represent a specific combination of link-medium-sampling point (e.g., mine 1-groundwater-sampling point P), and directed edges represent candidate causal paths from pollution source (cause) to pollution receptor (effect). Each coupling strength coefficient output from the PI-ST-GT model is assigned to the corresponding directed edge in the causal network skeleton as the initial weight for that candidate causal relationship.
[0096] S502. A structural equation model is established based on the causal network skeleton after weight setting; the time-varying coefficient matrix of the structural equation model is obtained by estimating the standardized dataset.
[0097] It should be noted that the structural equation model can be expressed as: ; Where X(t) is the vector of observed values of all variables at time t (i.e., the pollutant concentration at each sampling point), Φ(t) is the time-varying coefficient matrix, which reflects the dynamic causal relationship between variable X(t-1) and variable X(t), and ε(t) is the error term with a mean of 0.
[0098] The time-varying vector autoregressive (TV-VAR) model is used to estimate the time-varying coefficient matrix Φ(t). TV-VAR allows the time-varying coefficient matrix Φ(t) to change over time, and common estimation methods include state-space models, recursive least squares, and rolling window regression.
[0099] S503. Calculate the stability index of each candidate causal path in the structural equation model.
[0100] It should be noted that for each transmission path i->j, the stability index is calculated at time t: ; in, Let be the time-varying causal stability index of the transmission path from node i to node j at time t. The closer to 1, the higher the stability. The width of the confidence interval for causal strength. The p-value is for the time-varying unit root test. The causal lag time predicted by the model. This is the physical lag time estimated based on physical migration laws (such as water flow velocity and diffusion rate). The closer the predicted lag is to the physical lag, the larger this value is and the higher the stability.
[0101] S504. Construct an overall verification loss function based on the stability index, the mass conservation loss of the trained transmission coupling relationship model, and the data reconstruction error of the extracted main feature components, so as to calculate the overall verification loss value.
[0102] It should be noted that: the overall validation loss function The analysis combines three components: data reconstruction error, stability index, and mass conservation loss. The data reconstruction error originates from the WDM-PCA stage and represents the error in reconstructing the original data using principal component analysis. The mass conservation loss originates from the PI-ST-GT stage and represents the degree of violation of mass conservation constraints. The overall validation loss function for the entire analysis workflow is calculated as follows: ; in, For the overall validation loss function, For data reconstruction error, For loss due to conservation of mass, As a stability index, , , It's the weighting coefficient, the weighting coefficient. , , The Bayesian optimization method determines that the optimization objective is to maximize the consistency between the pollution source tracing results output by the model and the actual investigation results.
[0103] S505. When the overall verification loss value is lower than the judgment threshold set according to the historical verification set, the verification is determined to be successful.
[0104] It should be noted that: when If the result is less than a preset threshold θ, the verification is considered successful. The threshold θ is set based on the performance of the historical validation set. Specifically, it is determined through Bayesian optimization on the historical validation set (which includes known actual exploration results). , , Then, calculate a series of The values are then analyzed, and the distribution of these loss values is determined. A quantile threshold (e.g., the 95th quantile) is selected as the preset threshold θ, so that when new data... When the data is below a preset threshold θ, the analysis results are considered reliable. Passing the validation means that the output of the entire analysis process (from data preprocessing to causal verification), such as pollution source tracing results and remediation recommendations, is considered accurate and reliable and can be used for decision support. If the error is greater than or equal to a preset threshold θ, the verification is deemed a failure, and data reprocessing or model adjustment may be necessary. Specifically, the problem is precisely located based on the source of error in each component of the overall verification loss function. If the data reconstruction error (Lrecon) is large, it indicates that the feature extraction stage failed to effectively capture the main patterns of the data, or that the preprocessed data quality is insufficient; in this case, the feature extraction or data preprocessing parameters are re-optimized. If the mass conservation loss (Lphysics) is high, it indicates that the path learned by the transitive coupling relationship model (PI-ST-GT) violates the laws of physical conservation; in this case, the physical constraints or network structure of the transitive coupling relationship model are adjusted to ensure that it conforms to basic physical laws such as mass conservation. If the causal stability error (1-TVS-SI) is prominent, it indicates that the initially discovered causal relationship is statistically unrobust or inconsistent in time series; in this case, the parameters of the structural causal model (SCM-JV) are optimized, or the alignment of the time series is recalibrated. Based on the above diagnosis, the model or parameters of the corresponding stages are adjusted and retrained to gradually adapt the model to the current data distribution and patterns until the verification passes.
[0105] To automatically optimize the key weight coefficients in the model and maximize the accuracy of pollution source tracing results, in some embodiments of the present invention, such as... Figure 6 As shown, it also includes: S601. Calculate the degree of agreement between the pollution source tracing results and the actual investigation results.
[0106] It should be noted that the consistency degree is defined as a measure of the consistency between the pollution source tracing results output by the model (e.g., pollution source location, contribution rate, etc.) and the actual investigation results (e.g., actual pollution sources determined through on-site exploration and monitoring). Assume the pollution source tracing results output a series of candidate pollution source locations and their contribution rates, while the actual investigation results are a set of determined pollution source locations. For each actual pollution source, the nearest predicted pollution source (in spatial location) in the transitive coupling relationship model is found, and the location error (e.g., Euclidean distance) is calculated. The similarity of contribution rates is calculated; for example, for matched pollution source pairs, the correlation coefficient or relative error of contribution rates is calculated. The overall consistency degree is obtained by combining the location error and contribution rate similarity. The consistency degree process is as follows: Assume the actual investigation results have m pollution sources, and the transitive coupling relationship model predicts n pollution sources. First, a bipartite graph matching is constructed to match the actual pollution sources with the predicted pollution sources, minimizing the total location error. Then, for successfully matched pollution source pairs, the similarity of contribution rates is calculated. The similarity score is calculated as follows: w1 × (1 - mean normalized position error) + w2 × contribution rate similarity, where w1 and w2 are weights, and w1 + w2 = 1. The mean normalized position error is the average of the position errors of the matched pairs divided by the maximum permissible error (e.g., the maximum diagonal distance of the study area).
[0107] S602. Using the Bayesian optimization algorithm, with the goal of maximizing the degree of fit, the optimal combination of weight coefficients in the overall verification loss function is calculated.
[0108] It should be noted that: due to the three weight coefficients in the overall validation loss function , , These are relative weights, which can be constrained to the interval [0, 1] and satisfy the following conditions: (Or, without constraints, normalization can be used). Since Bayesian optimization maximizes the objective function, the objective function is the goodness of fit calculated by S601. Gaussian process regression is used to model the objective function, and a sampling function (such as EI, Expected Improvement) is used to guide the next sampling point. In each iteration, the parameter combination that maximizes the sampling function is selected. Calculate the degree of fit for this combination, update the Gaussian process model until the maximum number of iterations is reached or convergence is achieved. After the iteration ends, select the parameter combination with the highest degree of fit in history as the optimal weight coefficient combination.
[0109] S603. Update the overall verification loss function according to the optimal weight coefficient combination.
[0110] It should be noted that the obtained optimal weighting coefficients Substitute this into the updated overall validation loss function for subsequent joint validation.
[0111] In this embodiment, Bayesian optimization automatically determines the weights of each component in the loss function, avoiding the subjectivity and tediousness of manual parameter tuning. With consistency as the objective, it ensures that the optimization direction aligns with the actual survey results, thereby improving the accuracy of the source tracing results. Furthermore, Bayesian optimization can handle non-convex, high-dimensional optimization problems, is suitable for complex weight coefficient adjustments, and typically finds a better solution within a fewer iterations, resulting in high computational efficiency. Moreover, the optimization process is repeatable, facilitating consistent optimization results across different datasets. By optimizing the weights, the overall validation loss function more reasonably balances feature reconstruction error, mass conservation loss, and causal stability, thereby improving the model's reliability.
[0112] For example, the overall technical process includes the following first stage: extracting dominant feature patterns (i.e., main feature components) from complex data using weighted dynamic manifold principal component analysis (WDM-PCA). Specifically, this involves receiving time-stamped pollutant monitoring data from five stages and three media. Subsequently, media-specific standardization is performed, differentiating the dimensions and standard limits for different media to ensure all data are at comparable magnitudes. Next, manifold structure learning is performed to capture the inherent nonlinear correlation and spatial continuity of pollutant concentrations across multiple regions. A manifold learning algorithm is used to map high-dimensional data to a low-dimensional manifold space, obtaining manifold coordinates. This involves integrating the complex relationships of geospatial data into the data representation. Then, a weighted manifold kernel function is constructed, which combines the manifold structure and the weights of each element to create a... Its specific form is as follows: Next, dynamic time alignment and kernel PCA parsing are performed, that is, the Dynamic Time Warping (DTW) algorithm is used to align all unequal time sequences. Then, a kernel function is used... Kernel Principal Component Analysis (KPCA) is performed to extract principal components whose cumulative variance contribution rate exceeds a preset contribution rate threshold (e.g., 90%). These principal components represent the dominant synergistic patterns of pollutant distribution and evolution throughout the region. SHAP and LIME models can be combined to quantify the contribution of each link and each medium to each principal component, accurately identifying key pollutants and core responsible links. The output of this stage will serve as important input features for constructing the graph attention neural network in the next stage. Second Stage: The entire mining area is abstracted as a dynamic graph structure to simulate the flow and interaction of pollutants in a complex network. Specifically, a multi-media coupled transfer graph (MMCT-Graph) is constructed: using "link-medium-data point" as nodes and reachability calculated based on physical migration laws (such as Darcy's law and diffusion equations) as edges, an initial graph structure (i.e., graph attention neural network) is built. The MMCT-Graph is input into the graph attention neural network (PI-ST-GT). Based on a physically constrained attention mechanism, the attention coefficient between nodes i and j is calculated. This mechanism ensures that the model prioritizes learning physically plausible transport paths. Feasibility rights are used to construct the initial graph structure based on reachability calculated using physical models such as Darcy's law and diffusion equations. A quantitative value representing the ease of pollutant transport from node i to node j is calculated based on the spatial location and environmental attributes (such as hydrogeological parameters and meteorological data) of data points using classic physical transport models (e.g., Darcy's law for groundwater, diffusion equations for soil and gases). A mass conservation loss function is introduced during model training. As a constraint: The mass conservation loss function ensures that the transmission laws learned by the model conform to the law of mass conservation. Dynamic transmission strength is obtained through network learning, and coupling effects are quantified through shielded experiments. The output transmission paths and coupling strength coefficients in this stage constitute a set of candidate causal relationships, which are then sent to the final verification stage. The third stage dynamically verifies the causal laws in the complex system through joint causal model verification (SCM-JV). This stage rigorously verifies the discovered relationships statistically to ensure their authenticity, stability, and conformity to physical time-series laws. Specifically, the dynamic paths output by PI-ST-GT are used as the causal network skeleton to establish a time-varying structural equation model. And calculate the time-varying causal stability index (i.e., the stability index TVC-SI): ; The system finally calculates the overall verification loss function. Perform global optimization and judgment: ;in For WDM-PCA reconstruction error, The average causal stability index is denoted by α, β, and γ. The coefficients α, β, and γ are determined through Bayesian optimization, with the optimization objective being to maximize the agreement between the final analysis results and the actual exploration results. When the values are less than a preset threshold θ, the entire analysis process is considered valid. This application addresses multi-media pollutant data (groundwater, fluoride, sulfate, nitrate, ammonia nitrogen, total phosphorus, lead), soil, and gases (sulfur dioxide, ammonia) across five stages: mining, production, storage and transportation, solid waste, and fertilization. In scenarios lacking environmental factor data, a workflow of "Weighted Dynamic Manifold Principal Component Analysis (WDM-PCA) → Physical Information Spatiotemporal Graph Network (PI-ST-GT) → Structural Causal Model Joint Verification (SCM-JV)" is constructed. By integrating manifold structure and stage weights in WDM-PCA, the dominant pollution patterns and key driving stages can be clearly decoupled from the mixed signals interwoven across multiple regions, overcoming the shortcomings of traditional linear methods in analyzing nonlinear synergistic effects. The PI-ST-GT model, through the introduction of a physically constrained attention mechanism, ensures automatic focusing on physically feasible real migration routes in complex path networks. Combined with shielding experiments, it can accurately quantify indirect coupling effects that are difficult for traditional models to capture. The Time-Varying Causal Stability Index (TVC-SI) proposed in the SCM-JV stage provides a comprehensive quantitative assessment from three dimensions: causal strength, time-series stationarity, and lag effects. Its dynamic verification capability closely matches the characteristics of causal relationships evolving over time in complex systems, resulting in more reliable results. Through an overall verification loss function, feature fidelity (i.e., data reconstruction error), physical rationality (i.e., mass conservation loss), and causal robustness (i.e., stability index) are incorporated into a unified framework for synergistic optimization, ensuring that the entire process from data preprocessing to final verification points to the same optimal goal, significantly improving the overall scientific rigor and decision support value of the source tracing conclusions. This application does not rely on environmental factors that are difficult to define uniformly; it is directly driven by spatiotemporal data of pollutants and exhibits good adaptability and robustness to the complex transmission and coupling problems commonly found in various mining areas. It can effectively improve the accuracy and efficiency of pollution control and reduce prevention and control costs. In summary, WDM-PCA, which incorporates stage weights and manifold structure, accurately extracts nonlinear principal components. A graph attention mechanism constrained by physical laws ensures that the transmission path conforms to the transfer law. Based on the established Time-Varying Causal Stability Index (TVC-SI), a unified quantitative verification of causality and temporal stability is achieved. By organically linking each stage through a single overall verification loss function, the entire process from feature extraction and transmission modeling to causal verification is optimized, significantly improving the scientific rigor and accuracy of mine pollution source tracing.
[0113] To better implement the pollutant source tracing method based on multiple media and multiple stages in the embodiments of the present invention, based on the pollutant source tracing method based on multiple media and multiple stages, correspondingly, as follows: Figure 7 As shown, this embodiment of the invention also provides a pollutant source tracing system 700 based on multiple media and multiple stages. The pollutant source tracing system 700 based on multiple media and multiple stages includes: The data acquisition and preprocessing module 701 is used to preprocess pollutant monitoring data to obtain a standardized dataset; the pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized datasets; Feature extraction module 702 is used to extract the main feature components characterizing the distribution and evolution of pollutants based on the standardized dataset; The transmission and coupling modeling module 703 is used to construct a transmission and coupling relationship model based on the main characteristic components, and to obtain the transmission path and coupling strength coefficient of pollutants according to the transmission and coupling relationship model; The causal verification and source tracing module 704 is used to construct a causal model and perform joint verification based on the transmission path and coupling strength coefficient; when the verification passes, the pollution source tracing result is output.
[0114] The pollutant source tracing system 700 based on multiple media and multiple links provided in the above embodiments can realize the technical solutions described in the above embodiments of the pollutant source tracing method based on multiple media and multiple links. The specific implementation principles of each module or unit can be found in the corresponding content in the above embodiments of the pollutant source tracing method based on multiple media and multiple links, and will not be repeated here.
[0115] like Figure 8 As shown, the present invention also provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802, and a display 803. Figure 8 Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead.
[0116] In some embodiments, processor 801 may be a central processing unit (CPU), microprocessor, or other data processing chip, used to run program code stored in memory 802 or process data, such as the pollutant tracing method based on multiple media and multiple links in this invention.
[0117] In some embodiments, processor 801 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 801 may be local or remote. In some embodiments, processor 801 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, intranet, multi-cloud, etc., or any combination thereof.
[0118] In some embodiments, memory 802 may be an internal storage unit of electronic device 800, such as a hard disk or memory of electronic device 800. In other embodiments, memory 802 may also be an external storage device of electronic device 800, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 800.
[0119] Furthermore, the memory 802 may include both internal storage units of the electronic device 800 and external storage devices. The memory 802 is used to store application software and various types of data installed on the electronic device 800.
[0120] In some embodiments, display 803 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 803 is used to display information from electronic device 800 and to display a visual user interface. Components 801-803 of electronic device 800 communicate with each other via a system bus.
[0121] In one embodiment, when processor 801 executes the pollutant tracing program based on multiple media and multiple stages in memory 802, the following steps can be implemented: Pollutant monitoring data is preprocessed to obtain a standardized dataset; the pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized data; Based on the standardized dataset, the main characteristic components representing the distribution and evolution of pollutants are extracted; A transmission coupling relationship model is constructed based on the main characteristic components, and the transmission path and coupling strength coefficient of pollutants are obtained according to the transmission coupling relationship model. Based on the transmission path and coupling strength coefficient, a causal model is constructed and a joint verification is performed; when the verification passes, the pollution source tracing result is output.
[0122] It should be understood that when the processor 801 executes the pollutant tracing program based on multiple media and multiple links in the memory 802, in addition to the functions mentioned above, it can also perform other functions, as can be found in the description of the corresponding method embodiments above.
[0123] Furthermore, this embodiment of the invention does not specifically limit the type of electronic device 800 mentioned. Electronic device 800 can be a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, laptop computer, or other portable electronic device. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic device can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the invention, electronic device 800 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0124] Accordingly, this application also provides a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, they can implement the steps or functions of the pollutant tracing method based on multiple media and multiple links provided in the above-described method embodiments.
[0125] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0126] The above provides a detailed description of the pollutant tracing method, system, equipment, and media based on multiple media and multiple stages provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A pollutant source tracing method based on multiple media and multiple stages, characterized in that, include: Pollutant monitoring data are preprocessed to obtain a standardized dataset; The pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized datasets. Based on the standardized dataset, the main characteristic components representing the distribution and evolution of pollutants are extracted; A transmission coupling relationship model is constructed based on the main characteristic components, and the transmission path and coupling strength coefficient of pollutants are obtained according to the transmission coupling relationship model. Based on the transmission path and coupling strength coefficient, a causal model is constructed and joint verification is performed; When the verification passes, the pollution source tracing result is output.
2. The pollutant source tracing method based on multiple media and multiple stages according to claim 1, characterized in that, The process of preprocessing pollutant monitoring data to obtain a standardized dataset includes: Pollutant concentration data from multiple preset stages and multiple media are collected, and the corresponding stage attributes, media type, spatial coordinates and sampling time are labeled for each sampling point to obtain the pollutant monitoring data; Using a spatiotemporal attention graph convolutional network, combined with pollution contribution weights determined by the entropy weight method, missing values in the pollutant monitoring data are filled in; For the dimensional differences in pollutant concentrations in different media, standardization is performed based on the corresponding environmental quality standard limits to ensure that the standardized data for all media are within a preset mean range and a preset standard deviation range.
3. The pollutant source tracing method based on multiple media and multiple stages according to claim 1, characterized in that, Based on the standardized dataset, the main feature components characterizing the distribution and evolution of pollutants are extracted, including: A spatial adjacency graph is constructed based on the spatial coordinates of the sampling points, and the shortest path distance between every two sampling points is calculated based on the spatial adjacency graph to obtain a spatial distance matrix characterizing the migration potential of the pollutants. Based on the spatial distance matrix, the standardized data is reduced in dimensionality using a manifold learning algorithm to obtain a manifold coordinate matrix mapped in the structural space. Based on the manifold coordinate matrix and the pollution contribution weights determined by the entropy weight method, a weighted manifold kernel function is constructed; A time warping algorithm is used to perform time alignment processing on the standardized data corresponding to each sampling point in the standardized dataset. Based on the aligned data, principal component analysis is performed according to the weighted manifold kernel function to extract principal components that meet the preset contribution rate threshold, and the extracted principal components are used as the main feature components.
4. The pollutant source tracing method based on multiple media and multiple stages according to claim 3, characterized in that, Also includes: The spatial distance matrix is analyzed to obtain the data manifold structure characteristics, and the corresponding manifold learning algorithm is selected based on the data manifold structure characteristics; the manifold learning algorithm includes the isometric feature mapping algorithm and the Laplacian feature mapping algorithm. When the spatial distance matrix satisfies the global isometry property, the manifold coordinate matrix is obtained by the isometry feature mapping algorithm. When the spatial distance matrix exhibits local nonlinear characteristics, the Laplacian eigenmap algorithm is used to obtain the manifold coordinate matrix.
5. The pollutant source tracing method based on multiple media and multiple stages according to claim 1, characterized in that, The construction of a transport coupling relationship model based on the main characteristic components, and the determination of the pollutant transport path and coupling strength coefficient based on the transport coupling relationship model, include: Based on the main feature components, a multi-media coupling transfer graph containing multiple nodes is constructed, wherein each node corresponds to a combination of one of the links, one of the media, and one sampling point; Calculate the pollutant migration probability between any two nodes, and use the pollutant migration probability as the initial weight of the connecting edge between the two nodes; A graph attention neural network is constructed based on the multi-media coupling transfer graph, and the transfer coupling relationship model is obtained by training the graph attention neural network; during the training process of the graph attention neural network, the pollutant migration probability and mass conservation loss are used as constraints. Based on the aforementioned transmission coupling relationship model, the transmission path of pollutants between the nodes is obtained according to the standardized dataset. Based on the aforementioned transmission coupling relationship model, the state change of the node after the control variable is changed is obtained, and the coupling strength coefficient is calculated based on the state change.
6. The pollutant source tracing method based on multiple media and multiple stages according to claim 1, characterized in that, The step of constructing a causal model and performing joint verification based on the transmission path and coupling strength coefficient includes: The transmission path is used as a causal relationship to construct a causal network skeleton, and the coupling strength coefficient is set as the initial weight of the causal network skeleton. A structural equation model is established based on the causal network skeleton after weight setting; the time-varying coefficient matrix of the structural equation model is obtained by estimating the standardized dataset; Calculate the stability index of each candidate causal path in the structural equation model; The overall validation loss function is constructed based on the stability index, the mass conservation loss of the trained transitive coupling relationship model, and the data reconstruction error of the extracted main feature components, so as to calculate the overall validation loss value. When the overall verification loss value is lower than the judgment threshold set according to the historical verification set, the verification is determined to be successful.
7. The pollutant source tracing method based on multiple media and multiple stages according to claim 6, characterized in that, Also includes: Calculate the degree of agreement between the pollution source tracing results and the actual investigation results; Using a Bayesian optimization algorithm, with the goal of maximizing the fit, the optimal combination of weight coefficients in the overall validation loss function is calculated. The overall validation loss function is updated based on the optimal combination of weight coefficients.
8. A pollutant source tracing system based on multiple media and multiple stages, characterized in that, include: The data acquisition and preprocessing module is used to preprocess pollutant monitoring data to obtain a standardized dataset; The pollutant monitoring data includes data on pollutants in multiple stages and multiple media within the target area, and the standardized dataset includes multiple standardized datasets. The feature extraction module is used to extract the main feature components characterizing the distribution and evolution of pollutants based on the standardized dataset; The transmission and coupling modeling module is used to construct a transmission and coupling relationship model based on the main characteristic components, and to obtain the transmission path and coupling strength coefficient of pollutants according to the transmission and coupling relationship model; The causality verification and tracing module is used to construct a causal model and perform joint verification based on the transmission path and coupling strength coefficient. When the verification passes, the pollution source tracing result is output.
9. An electronic device, characterized in that, Including memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the pollutant tracing method based on multiple media and multiple links as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the pollutant tracing method based on multiple media and multiple links as described in any one of claims 1 to 7.
Citation Information
Cited By
Mine heavy metal pollution source positioning method and device based on space-time inversion, equipment and medium
CN122201493A
A machine learning-based intelligent decision system for soil pollution remediation of a dump
CN122222425A
Groundwater pollution plume source identification method and system based on three-dimensional dynamic monitoring
CN122262595A