A soil pollution source tracing method, device and medium based on causality
Through the combination of PMF, emission list method and GCCM model, the problem of insufficient causal inference ability in the existing technology is solved, accurate traceability and targeted control of soil pollution are achieved, and scientific basis for pollution control is provided.
Patent Information
- Application Number
- CN202510805548.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing soil pollution traceability methods have insufficient causal inference ability, making it difficult to accurately trace the source in complex environments. Traditional methods cannot effectively detect nonlinear causal relationships and the reliability of spatial data when they are scarce, resulting in inaccurate pollution control measures.
The soil pollution traceability method based on causal relationship is adopted to identify hidden pollution sources through positive matrix factor decomposition (PMF), and a high-resolution emission distribution map is generated based on emission list method and nuclear density method. The causal relationship is verified by GCCM model, and a framework for soil pollution traceability and control are constructed.
It realizes accurate traceability of soil pollution in complex environments, identify key pollution sources and their contribution areas, provides a scientific basis for targeted governance, avoids false correlations and data dependence restrictions, and improves the detection ability of the causal relationship between pollutant distribution and emission sources.
Smart Images

Figure CN120317537B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of pollution detection and control, and in particular to a soil pollution source tracing method, equipment, and medium based on causality. Background Art
[0002] With the accelerated development of industrialization and urbanization, soil pollution is becoming increasingly serious in many regions. Soil pollution source tracing involves collecting and analyzing harmful substances in soil samples, combined with relevant information such as environmental surveys, to trace the source of soil pollution, providing a reasonable basis for pollution control.
[0003] Current methods for tracing the source of soil contaminants typically involve performing PMF (Positive Matrix Factorization) on soil contaminant content data to identify hidden soil pollution sources, then using correlation analysis or geographically weighted regression to verify causal relationships. This method lacks causal inference capabilities. For example, correlation analysis only provides statistical associations, and causal relationships can be obscured by other collinear pollution sources, easily leading to "spurious correlations." Geographically weighted regression relies on linear assumptions and cannot detect nonlinear causal chains in weakly coupled systems. Furthermore, it requires high spatiotemporal data, which reduces its reliability in real-world scenarios when data is scarce. Consequently, none of these methods can accurately trace the source of soil contamination in complex environments.
[0004] In the prior art, patent application CN116559148A discloses a soil pollutant source tracing method, and patent application CN118171060A discloses a soil pollutant source tracing analysis method, neither of which can solve the above-mentioned problem. Summary of the Invention
[0005] The embodiments of the present invention provide a soil pollution source tracing method, device and medium based on causality to solve the above problems.
[0006] In a first aspect, an embodiment of the present invention provides a soil pollution source tracing method based on causality, comprising:
[0007] Obtain soil pollutant concentration detection data in the area to be traced, as well as data on emission sources from enterprises and transportation;
[0008] Performing positive matrix factorization on the detection data to determine the contribution of each hidden pollution source to each pollutant; determining at least one pollutant that is ranked at the top of the hidden pollution source contribution and is associated with enterprises and / or transportation as at least one target pollutant to be traced;
[0009] For each target pollutant, perform the following operations:
[0010] S1-1. interpolating a first two-dimensional distribution map of target pollutant concentration based on the detection data;
[0011] S1-2. Execute the emission inventory method based on the data of the two emission sources to obtain the target total pollutant emissions of each enterprise and the target total pollutant emissions of transportation;
[0012] S1-3. Using the kernel density method, allocate the total target pollutant emissions of each enterprise to each grid in the region to obtain a second two-dimensional distribution map of enterprise emissions; using the road network density weight, allocate the total target pollutant emissions of the traffic to each grid in the region to obtain a third two-dimensional distribution map of traffic emissions;
[0013] S1-4. Performing GCCM causal analysis on the second two-dimensional distribution map and the third two-dimensional distribution map, respectively, with the first two-dimensional distribution map to verify the causal relationship and significance of the target pollutants between enterprises and transportation.
[0014] Based on the verification results, auxiliary decisions are made on the blocks that are given priority for control within the area, as well as the types of pollutants and emission sources that are given priority for control in each block.
[0015] In a second aspect, an embodiment of the present invention provides an electronic device, comprising:
[0016] one or more processors;
[0017] a memory for storing one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the soil pollution source tracing method based on causality as described in any embodiment.
[0019] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the soil pollution source tracing method based on causality as described in any embodiment is implemented.
[0020] In summary, the embodiments of the present invention provide a soil pollution source tracing method based on causal relationships. For areas with industry and transportation as the main emission sources, first, the soil detection data is decomposed by PMF to determine the hidden pollution sources, and the target pollutants with large contributions from the hidden pollution sources and related to the emission sources are extracted from dozens of pollutants as the tracing objects, and the two-dimensional spatial distribution of the target pollutant concentration is obtained by interpolation; then, the emissions of each industrial enterprise and transportation are spatialized by the emission inventory method, and according to the characteristics of enterprise emissions and traffic emissions, the emission inventory is spatialized by the kernel density function and road density respectively, to obtain a more accurate and detailed two-dimensional spatial distribution of emissions; finally, the causal relationship between each emission source and the spatial distribution of pollutant concentration is verified by the GCCM model, and the nonlinear causal chain in the two weakly coupled systems of PMF and emission inventory is detected; and different control strategies are recommended for different blocks based on the detection results to assist in the precise targeted control of soil pollution.
[0021] Specifically, compared with the prior art, the embodiments of the present invention have the following advantages:
[0022] 1. Existing technologies lack the ability to infer causality in soil pollution. Traditional correlation analysis (such as the Pearson correlation coefficient) and MGWR cannot effectively detect nonlinear causal relationships in weakly coupled systems, making it difficult to reveal the complex interaction chains between pollutant concentration distributions and emission sources. These methods are often based on linear assumptions and have high data requirements (such as time series data). They perform poorly in complex environmental systems, especially when spatiotemporal data is scarce. This limitation is particularly prominent.
[0023] The embodiment of the present invention adopts the GCCM model, which is based on dynamic system theory and generalized embedding theorem, and reconstructs the state space of the system through spatial cross-sectional data. It overcomes the data dependence limitation of traditional methods and can accurately identify hidden one-way causal relationships. That is, it can clarify the causal direction of enterprises, traffic emission sources and pollutant concentration distribution, avoid "false correlation" interference, and provide a scientific basis for precise governance.
[0024] 2. Existing emission inventories lack spatialization. While existing emission inventory methods can estimate total pollutant emissions, they lack detailed characterization of their spatial distribution, making it difficult to accurately locate high-emission hotspots.
[0025] The embodiment of the present invention uses kernel density allocation and road network weight allocation methods to accurately allocate emissions from industrial emission sources and transportation emission sources to a spatial grid, generating a high-resolution emission inventory map and providing high-quality input data for subsequent causal inference.
[0026] 3. Existing technologies struggle to analyze complex, multi-source systems. In complex, multi-source, nonlinearly coupled environments, existing technologies struggle to fully analyze the sources of pollutants and their interactions, potentially leading to the neglect of minor pollution sources or the masking of the contributions of major pollution sources.
[0027] The embodiments of the present invention, through multi-model coupling (PMF+emission inventory+GCCM), combined with spatial analysis and dynamic causal inference, can comprehensively analyze the pollution characteristics of multi-source complex systems and reveal hidden pollution sources and interaction mechanisms.
[0028] 4. Inaccurate pollution control in existing technologies. Existing technology solutions lack a deep understanding of the causal relationship between visible emission sources and pollutant concentration distribution. As a result, pollution control measures are often overly broad and fail to target key pollution sources.
[0029] The embodiment of the present invention integrates PMF (source identification), emission inventory (source calculation) and GCCM (source verification) to build a complete soil pollution traceability and control framework. It can establish a correspondence between the hidden pollution sources in the theoretical model and the explicit emission sources in reality, accurately lock in key pollution sources and their contributing areas, and provide a scientific basis for formulating feasible targeted governance strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 This is a flow chart of a soil pollution source tracing method based on causality provided by an embodiment of the present invention;
[0032] Figure 2 is a two-dimensional distribution diagram of As concentration obtained by interpolation within an exemplary region provided by an embodiment of the present invention;
[0033] Figure 3 This is a two-dimensional distribution diagram of As emissions from enterprises in an exemplary region provided by an embodiment of the present invention;
[0034] Figure 4 This is a schematic diagram of the causal relationship verification results provided by an embodiment of the present invention, with Gejiu City, Yunnan Province as the area to be traced, smelting and mining as the two industries to be controlled, and As and Pb as the two target pollutants; wherein, Figure 4 (a) is the verification result of smelting enterprises and As concentration, Figure 4(b) is the verification result of industrial and mining enterprises and As concentration, Figure 4 (c) is the verification result of traffic emission sources and As concentration, Figure 4 (d) is the verification result of smelting enterprises and Pb concentration, Figure 4 (e) is the verification result of industrial and mining enterprises and Pb concentration, Figure 4 (f) is the verification result of traffic emission sources and Pb concentration;
[0035] Figure 5 This is a two-dimensional distribution diagram showing a mixed display of concentrations of multiple target pollutants within an exemplary area provided by an embodiment of the present invention;
[0036] Figure 6 This is a technical framework diagram of a soil pollution source tracing method based on causality provided by an embodiment of the present invention;
[0037] Figure 7 This is a schematic diagram of the results of causal relationship verification provided by an embodiment of the present invention, with Gejiu City, Yunnan Province as the area to be traced, and the density of smelting enterprises, the density of industrial and mining enterprises, the density of road networks, and the detection concentration of As as input data; wherein, Figure 7 (a) is the verification result of the causal relationship between smelting enterprise density and As concentration, Figure 7 (b) is the verification result of the causal relationship between road network density and As concentration. Figure 7 (c) is the verification result of the causal relationship between industrial and mining enterprise density and As concentration;
[0038] Figure 8 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.
[0040] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0041] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0042] Figure 1 This is a flow chart of a soil pollution source tracing method based on causal relationships provided by an embodiment of the present invention. This method is applicable to areas where industry and transportation are the main sources of soil pollutants, and is executed by electronic equipment. Figure 1 As shown, the method specifically includes:
[0043] S110. Obtain soil pollutant concentration detection data in the area to be traced, as well as data from two emission sources: enterprises and transportation.
[0044] This embodiment first determines a certain area as the target for subsequent soil pollutant tracing and control. The area can be a province or a city, and this embodiment does not impose specific restrictions. Then, basic data of the area is collected as the data source for the entire method. Optionally, the basic data includes detection data of soil pollutant concentrations, enterprise data, and traffic data. Enterprises and transportation are the main sources of soil pollutants in the area and are also the targets of control in subsequent control measures.
[0045] In one specific embodiment, representative sampling points (e.g., soil pollutant content monitoring sites) can be selected within the region using a grid or random sampling method. These sampling points cover different land use types (industrial areas, agricultural areas, traffic-intensive areas, etc.) to ensure balanced sampling. Soil pollutant concentration data are tested at each sampling point using standard methods. Generally speaking, testing stations detect dozens of types of soil pollutants (for example, according to GB36600, soil testing for construction land requires 45 items and 40 other items), including heavy metals such as arsenic (As), lead (Pb), cadmium, mercury, and nickel; and organic substances such as carbon tetrachloride, chloroform, and methyl chloride.
[0046] Enterprise data includes the enterprise's location, activity data over a period of time (such as output and process parameters), emission factors, and the removal efficiency and technical implementation coefficients of the pollution control technologies used by the enterprise. This data can be obtained from public online platforms such as online POIs (Point of Interest) maps (such as OpenStreetMap), Qichacha.com (http: / / www.qichacha.com / ), Tianyancha (http: / / www.tianyancha.com / ), and the National Pollutant Discharge Permit Management Information Platform (https: / / permit.mee.gov.cn / ). The National Enterprise Credit Information Publicity System (http: / / www.gsxt.gov.cn / ) is used as a supplement. Data is scraped from various websites and then further cleaned for use. The enterprise's location can be determined based on geographic information such as industrial layout maps. Emission factors, as well as the removal efficiency and technical implementation coefficients of pollution control technologies, can be obtained from the publicly available "Manual on Pollutant Discharge Accounting Methods and Coefficients for Emission Source Statistical Surveys."
[0047] Traffic data includes the number of motor vehicles (categorized by type: passenger cars, trucks, motorcycles, etc.) and their corresponding emission factors, as well as road network density maps. Vehicle data can be sourced from provincial statistical yearbooks. For example, if the source area is Yunnan Province or a city in Yunnan Province, vehicle data can be sourced from the Yunnan Provincial Statistical Yearbook. Road traffic information can be sourced from OpenStreetMap (https: / / master.apis.dev.openstreetmap.org). Emission factors can be sourced from the publicly available Manual on Pollutant Emission Accounting Methods and Coefficients for Emission Source Statistical Surveys.
[0048] S120. Perform positive matrix factor decomposition on the detection data to determine the contribution of each hidden pollution source to each pollutant; and determine at least one pollutant that is ranked at the top in terms of contribution by each hidden pollution source and is associated with enterprises and / or transportation as at least one pollutant to be traced.
[0049] This embodiment uses the PMF receptor model to select one or more pollutants from the numerous pollutants in the detection data as the targets for subsequent traceability and control. For ease of distinction and description, this embodiment refers to these pollutants to be traced and controlled as target pollutants.
[0050] Specifically, PMF is a relatively excellent receptor model for source attribution of air pollutants. This method uses a mathematical formula to decompose the concentration of pollutants in a sample into the contributions of multiple potential pollution sources, and then identifies the characteristic components of each pollution source and their relative contribution. This example uses this method to decompose pollutant concentration data measured at sampling points, aiming to identify target pollutants that are both highly polluting and controllable.
[0051] In a specific embodiment, the soil pollutant concentration detection data is first standardized to remove outliers. Then, a pollutant concentration matrix is constructed based on the standardized data. The matrix dimension is the number of sampling points × type of pollutant. Each row of the matrix is the concentration of each pollutant at the same sampling point, and each column is the concentration of the same pollutant at each sampling point.
[0052] By performing PMF on this matrix, we can obtain the number of potential pollution sources and the contribution of each pollution source to each pollutant. It should be noted that the pollution sources obtained here are implicit pollution sources analyzed from the pollutant concentration detection data. They are not explicit objects such as a certain block, a certain geographical grid or a certain enterprise, but may be multiple blocks, multiple grids or multiple enterprises, or even the comprehensive result of multiple factors. Therefore, this embodiment refers to this pollution source as an implicit pollution source, and its specific corresponding explicit object will be gradually clarified in the subsequent operations of this embodiment. Specifically, the PMF model formula is as follows:
[0053] (1)
[0054] in, represents the concentration of pollutant j at sampling point i, represents the contribution ratio of the kth hidden pollution source in sampling point i, represents the characteristic pollutant concentration of the kth hidden pollution source, represents the residual corresponding to pollutant j in sampling point i, and q represents the number of hidden pollution sources.
[0055] By applying the PMF to the pollutant concentration matrix using the above model, we can determine the number of hidden pollution sources within the region and the proportion of each source's contribution to each pollutant. For each hidden pollution source, we rank each pollutant based on its contribution ratio and select the top pollutants. Ultimately, we obtain the top pollutants for each hidden pollution source. These pollutants are then combined to form a set of pollutants with the highest impact.
[0056] At the same time, based on the data of the two emission sources, at least one pollutant associated with the two emission sources of enterprises or transportation is determined to form a set of associated pollutants. Optionally, pollutants associated with each enterprise can be determined from the enterprise data, and pollutants associated with transportation emission sources can be extracted from the traffic data, such as pollutants targeted by each emission factor. These associated pollutants can improve their pollution situation through the control of enterprises and transportation, and the preliminary range of target pollutants can be limited. For the sake of ease of distinction and description, this embodiment refers to the determined associated pollutant set here as the first pollutant set, and the pollutant set with serious impact constructed in the previous paragraph as the second pollutant combination.
[0057] The intersection of the first and second pollutant sets is taken, and each pollutant in the intersection is a target pollutant to be traced. For example, using Gejiu City in Yunnan Province as an example, the target pollutants ultimately screened out include arsenic and lead.
[0058] S130. For each target pollutant, perform the following operations S1-1 to S1-4 respectively to determine the causal relationship between each type of emission source and the current target pollutant and its significance level.
[0059] S1-1. Based on the detection data, a two-dimensional distribution diagram of the target pollutant concentration is obtained by interpolation.
[0060] For the target pollutant currently being processed, the detection data includes the concentration of the target pollutant at multiple sampling points. These discrete concentration values can be interpolated using the inverse distance interpolation method to form a spatially continuous two-dimensional distribution map of the target pollutant concentration. Figure 2 An example is shown of a two-dimensional distribution diagram of As concentration obtained by interpolation within a rectangular target area, where mg / kg (milligrams per kilogram) is the concentration unit, indicating the As content per kilogram of soil; SV represents the standard value, i.e., the second-category land use screening value.
[0061] S1-2. Execute the emission inventory method based on the data of the two emission sources to obtain the target total pollutant emissions of each enterprise and the target total pollutant emissions of transportation.
[0062] The emissions inventory method systematically quantifies the total amount of pollutants emitted by various pollution sources (such as industry, transportation, and daily life) over a specific time span and spatial scope. This example uses the emissions inventory method to quantify the actual emissions from each enterprise and transportation under control for the target pollutant, providing more accurate data support for source control.
[0063] In one specific implementation, the emissions inventory method is applied to each enterprise based on its data. Specifically, the production emissions of each enterprise can be determined based on its activity data during the production process and its emission coefficient for the target pollutant. The emission reductions of each enterprise can then be determined based on its production emissions and the removal efficiency of its pollution control technology. The total target pollutant emissions of each enterprise can be obtained by subtracting the emission reductions from the production emissions. Optionally, the total pollutant emissions of each enterprise can be calculated using the following formula:
[0064] (2)
[0065] in, (kg, kilogram) represents the total amount of target pollutant emissions of the enterprise; Indicates that the company Activity data of annual production processes; (kg / unit activity data) is the emission factor of the target pollutant, which is used to characterize the content of the target pollutant emitted per unit activity data; (%) represents the removal efficiency of the enterprise's pollution control technology, (Value range: 0-1) represents the technical implementation coefficient of the pollution control technology.
[0066] At the same time, the entire region's traffic is considered as a single emission source, and the emission inventory method is applied to the entire traffic source based on the traffic data. Optionally, the total target pollutant emissions of the entire traffic source are calculated based on the number of vehicles of each type and the emission factors for the target pollutants. The formula is as follows:
[0067] (3)
[0068] in, (kg) represents the total amount of target pollutants emitted by traffic emission sources, (kg / year / vehicle) indicates the Emission factors of target pollutants for each vehicle type, Indicates the Vehicle type in the Years of holdings.
[0069] S1-3. Using the kernel density method, the total target pollutant emissions of each enterprise are allocated to each grid in the region to obtain a two-dimensional distribution map of enterprise emissions; using the road network density weight, the total target pollutant emissions of the traffic are allocated to each grid in the region to obtain a two-dimensional distribution map of traffic emissions.
[0070] This embodiment uses different methods to spatially expand the emissions calculated by enterprises and transportation through emission inventories. Optionally, the area to be traced can be rasterized into 1 km x 1 km grid cells, and continuous spatial expansion can be achieved by calculating the emissions within each grid cell.
[0071] In one specific implementation, taking into account the pollutant diffusion characteristics of fixed emission sources, this embodiment uses the kernel density method to sequentially allocate the total target pollutant emissions of each enterprise to each grid in the area to be traced, thereby achieving spatial expansion of enterprise emissions. The kernel density estimation method is used to estimate the probability density function of point-like geographic data. This method scientifically quantifies the spatial distance attenuation effect of geographic phenomena, assigning greater weight to objects with closer geographic distances. Its core function is defined as follows:
[0072] (4)
[0073] in, Indicates estimated point The kernel density value of represents the search radius (bandwidth), Indicates the number of point features within the search radius. Indicates a point Arrive distance, is the default spatial weight function;
[0074] (5)
[0075] in, represents the standard distance, Represents the median distance, with the highest kernel density of the point feature as the upper limit and 0 as the lower limit. The higher the kernel density, the higher the emission intensity of the pollution source.
[0076] Based on the above core function, each grid in this embodiment can be used as an estimation point , the total amount of target pollutant emissions of each enterprise is taken as a point element and the kernel density method is implemented. Specifically, first, according to the spatial weight function , the total target pollutant emissions of each enterprise are sequentially allocated to multiple grids around the grid where each enterprise is located (i.e., the grids within the search radius mentioned above). As shown in formula (5), the kernel density method automatically determines a reasonable search radius based on the overall size of the area to be expanded and the grid size. Ultimately, the sum of the emissions allocated to each grid by the same enterprise is equal to the total emissions of the same enterprise. The farther the distance between each grid and the same enterprise, the greater the corresponding spatial weight and the greater the total emissions allocated. Then, the emissions allocated to each grid from each enterprise are added together to obtain the total target pollutant emissions of each grid. By arranging these total emission values in sequence according to the geographical location of the grids, a two-dimensional distribution map of enterprise emissions can be obtained.
[0077] In practice, the kernel density method in ArcGIS Pro can also be used to implement the above process. The input is the point features of each enterprise's emission inventory (i.e., emissions with coordinates). Kernel density analysis in the Spatial Analyst toolbox is used to set the specified output area unit factor (i.e., unit, whether it is square meters or square kilometers), and the rest are set to default; the output grid is the emissions of each unit (grid). Furthermore, the calculation results can be visualized using GIS (Geographical Information System) tools to generate high-resolution emission inventory maps. Figure 3 The example shows a two-dimensional distribution map of the As emissions of enterprises within a rectangular area. This embodiment also refers to this process as spatialization of the enterprise emission inventory.
[0078] The main emission carriers of traffic emission sources are moving vehicles. The corresponding pollutant diffusion characteristics are not exactly the same as those of fixed emission sources, and the specific diffusion laws are still unclear. Therefore, this embodiment distributes the total emissions of traffic emission sources to each road section based on road network density to ensure that the spatial distribution of emissions is consistent with the actual traffic pattern.
[0079] In a specific embodiment, first, the road network density weight of each grid is determined based on the ratio of the total road length of each grid in the area to be traced to the total road length of the area:
[0080] (6)
[0081] in, Representation Grid The road network density weight (or road network density), and Represents the grid and grid The total length of roads, Indicates the number of grid cells in the area.
[0082] Then, the target total pollutant emissions from the traffic emission sources are allocated to each grid according to the road network density weight of each grid:
[0083] (7)
[0084] in, Indicates allocation to the grid target pollutant emissions.
[0085] For ease of distinction and description, this embodiment refers to the two-dimensional distribution map of the target pollutant concentration obtained by interpolation of detection data in S1-1 as the first two-dimensional distribution map; the two-dimensional distribution map obtained by spatialization of the enterprise emission inventory in S1-3 is called the second two-dimensional distribution map, and the two-dimensional distribution map obtained by spatialization of the traffic emission inventory is called the third two-dimensional distribution map.
[0086] S1-4. Performing GCCM causal analysis on the second two-dimensional distribution map and the third two-dimensional distribution map with the first two-dimensional distribution map, respectively, to verify the causal relationship and significance of the target pollutants between enterprises and transportation.
[0087] This example uses the GCCM (Geographical Convergent Cross Mapping) model to verify the causal relationship between enterprise emissions and target pollutant concentrations, as well as the causal relationship between traffic emissions and target pollutant concentrations, and quantify the strength of these causal relationships. This process effectively couples pollution source information obtained from both PMF and emission inventories to comprehensively reflect the interactions among multiple pollution sources.
[0088] Specifically, GCCM is a spatial causal inference method based on dynamic systems theory and the generalized embedding theorem. It reconstructs the state space of a system using spatial cross-sectional data and identifies causal relationships between geographical factors. Unlike traditional correlation analysis, causal inference aims to reveal the independent influence of one phenomenon (cause) on another phenomenon (effect). Specifically, it aims to reveal how the outcome variable responds to changes in the cause variable, distinguishing true causal relationships from spurious correlations.
[0089] Furthermore, for the two variables X and Y whose causal relationship is to be verified, the core steps of the GCCM model include:
[0090] a) Embedding construction: Determine the spatial lag of each unit (e.g., pixel or polygon) in a 2D distribution map and organize it into vectors and matrices.
[0091] b) Predicting Y from X: Predict Y from X for a series of window sizes and calculate the prediction skill ρ for each window.
[0092] c) Predict X from Y: Use the same method to predict X from Y, and obtain a series of ρ values.
[0093] d) Causal interpretation: Plot the predicted output based on the convergence, significance, and confidence interval of the ρ value at the maximum window size to determine the presence, direction, and strength of causality. For more details, see Gao, B., Yang, J., Chen, Z., Sugihara, G., Li, M., Stein, A., Kwan, M.-P., Wang, J., 2023. Causal inference from cross-sectional earth system data with geographical convergent cross mapping. Nat. Commun. 14, 5875. https: / / doi.org / 10.1038 / s41467-023-41619-6.
[0094] In this embodiment, the causal relationship judgment results may include the following:
[0095] Unidirectional causal relationship (X→Y or Y→X): If the cross-mapping prediction skill (ρ value) of predicting Y through X (Y xmap X) is significantly higher at the maximum library size and shows a clear upward trend and statistical significance (p<0.05) as the library size (i.e., window size) increases, while the ρ value of the reverse prediction (X xmap Y) is low and insignificant (or much weaker than Y xmap X), then a unidirectional causal relationship (e.g., X→Y) can be inferred. Conversely, if the ρ value of X xmap Y is significant and shows a steady upward trend with increasing library size, while the ρ value of Y xmap X is low and insignificant, then it can be inferred that Y is the cause of X (Y→X).
[0096] Bidirectional asymmetric causality: When the ρ values for two cross-mapping predictions (Y xmap X and X xmap Y) are both significant and statistically significant (p < 0.05), but their magnitudes differ significantly, this indicates a bidirectional causal relationship between the two variables, but with asymmetric causal strength. For example, if the ρ value for Y xmap X is significantly higher than that for X xmap Y, it can be inferred that X primarily drives Y, with weak feedback from Y to X. This asymmetry may result from a strong correlation effect in the dominant causal direction, rather than completely independent bidirectional effects.
[0097] No causal relationship: If neither of the two cross-mapping prediction methods (Y xmap X and X xmap Y) shows a significant upward trend in ρ values, or if their prediction skills (ρ values) are statistically insignificant (p ≥ 0.05), a causal relationship between the two variables cannot be inferred. This may indicate a lack of dynamic correlation between the variables or noise in the data that interferes with the validity of causal inference.
[0098] Based on the above core steps, this embodiment completes the causal relationship verification between the second two-dimensional distribution map (enterprise emissions) and the first two-dimensional distribution map (pollutant concentrations), and the causal relationship verification between the third two-dimensional distribution map (traffic emissions) and the first two-dimensional distribution map (pollutant concentrations). Taking the verification process corresponding to traffic emissions as an example, the specific steps include the following:
[0099] Step 1: Input the third two-dimensional distribution graph and the first two-dimensional distribution graph into the GCCM model, and let the model predict the target pollutant concentration based on a series of traffic emissions of window sizes, calculate the prediction skill under each window, and draw a line graph showing the prediction skill changing with the window size; at the same time, predict traffic emissions based on a series of target pollutant concentrations of window sizes, calculate the prediction skill under each window, and draw a line graph showing the prediction skill changing with the window size. For ease of distinction and description, this embodiment refers to the broken lines in the two broken line graphs mentioned above as the first broken line and the second broken line, respectively. For example, Figure 4 (c) The result of causal verification between traffic emission sources and the detection concentration of target pollutant As is taken from Gejiu City, Yunnan Province as the source area to be traced. The horizontal axis L is the window size, the vertical axis ρ is the prediction skill, As xmapTraffic represents the prediction of As detection concentration based on As emissions from traffic sources, and its corresponding red broken line is the first broken line. Traffic xmap As represents the prediction of As emissions from traffic sources based on As detection concentration, and its corresponding blue broken line is the second broken line.
[0100] Step 2: Determine the causal relationship and significance based on the two broken lines. Optionally, if the prediction skill of the first broken line increases with the increase of the window size, and the prediction skill of the first broken line under the same window size is significantly higher than that of the second broken line, then it is determined that traffic affects the concentration of the target pollutant; and the greater the distance between the two broken lines, the more significant the causal relationship. Figure 4 Taking (c) as an example, the prediction skills of the two broken lines did not show a significant upward trend, and the prediction skills in the same window were almost the same, indicating that traffic emissions had little effect on As concentration. Figure 4(f) is the result of causal verification between traffic emission sources and the detection concentration of target pollutant Pb, with Gejiu City, Yunnan Province as the traceable area. Similarly, Pb xmapTraffic represents the prediction of Pb detection concentration based on Pb emissions from traffic sources, and its corresponding red broken line is the first broken line. Traffic xmap Pb represents the prediction of Pb emissions from traffic sources based on Pb detection concentration, and its corresponding blue broken line is the second broken line. It can be seen that the ρ value of the first broken line is higher and increases with the increase of window size, while the ρ value of the second broken line is significantly lower than that of the first broken line, indicating that traffic emission sources affect Pb concentration. The distance between the two broken lines under the maximum window represents the significance of the causal relationship.
[0101] As for corporate emission sources, when there are a large number of enterprises and they involve different industries, the enterprises can be classified according to industry in advance. Accordingly, when using the kernel density method to generate the second two-dimensional distribution map in S1-3, the total target pollutant emissions of each enterprise in the industry can be allocated to each grid in the area to be traced for each industry, and the two-dimensional distribution map corresponding to the corporate emissions of the industry can be obtained. Ultimately, each industry corresponds to a second two-dimensional distribution map. Accordingly, when verifying the causal relationship in this step, the causal relationship between the second two-dimensional distribution map and the first two-dimensional distribution map of each industry is also verified separately, and the causal relationship between the corporate emission sources of each industry and the target pollutants and their significance are obtained.
[0102] After performing operations S1-1 to S1-4 for each target pollutant, the verification results of each industry's enterprises and each target pollutant, as well as the causal inference results of traffic and each target pollutant, can be obtained. For example, Figure 4 The verification results are based on Gejiu City, Yunnan Province as the area to be traced, non-ferrous metal smelting and rolling processing industry (hereinafter referred to as "smelting") and non-ferrous metal mining and dressing industry (hereinafter referred to as "mining") as two industries to be controlled, and As and Pb as two target pollutants. Among them, Figure 4 (a) is the verification result of smelting enterprises (i.e. non-ferrous metal smelting and rolling processing enterprises) and As concentration. Asxmap NFM Smelting represents the prediction of As detection concentration based on As emission of smelting enterprises, and NFM Smelting xmapAs represents the prediction of As emission of smelting enterprises based on As detection concentration. Figure 4 (b) is the verification result of industrial and mining enterprises (i.e., non-ferrous metal mining and dressing enterprises) and As concentration. As xmap NFM Mining represents the predicted As concentration based on the As emissions of industrial and mining enterprises, and NFM Mining xmap As represents the predicted As emissions of industrial and mining enterprises based on the As concentration. Figure 4(d) is the verification result of the smelting enterprise and Pb concentration. Pb xmap NFM Smelting represents the prediction of Pb detection concentration based on the Pb emission of the smelting enterprise, and NFM Smelting xmap Pb represents the prediction of Pb emission of the smelting enterprise based on the Pb detection concentration. Figure 4 (e) is the verification result of industrial and mining enterprises and Pb concentration. Pb xmap NFM Mining represents the prediction of Pb detection concentration based on Pb emissions from industrial and mining enterprises, and NFM Mining xmap Pb represents the prediction of Pb emissions from industrial and mining enterprises based on Pb detection concentration. It can be seen that in the causal relationship results of each figure, the change trend and distance of the two broken lines are different. Figure 4 (a) The causal relationship of Chinese metallurgical enterprises to As is most significant.
[0103] S140. Based on the verification results, assist in deciding the blocks within the area that are to be prioritized for control, as well as the types of pollutants and emission sources that are to be prioritized for control in each block.
[0104] This embodiment recommends different control strategies for different blocks based on detection data and causal verification results, assisting in achieving precise targeted control of soil pollution.
[0105] In one embodiment, the first two-dimensional distribution diagrams of all target pollutants can be mixed and displayed. Figure 5 This is a two-dimensional distribution map showing a mixed display of multiple target pollutant concentrations, provided by an embodiment of the present invention. Different colors in the map represent different target pollutant concentrations, with three representative colors labeled 1, 2, and 3, respectively. From this mixed concentration map, at least one area where the target pollutant concentration exceeds the standard can be extracted as the at least one target area to be controlled.
[0106] For each target block, a control strategy can be recommended based on the following two situations:
[0107] Case 1: If there is only one target pollutant exceeding the standard in a target area, the one with the most significant causal relationship with the target pollutant can be extracted from the emission sources of various industries, enterprises and transportation, and recommended as the priority control object for the target area. Figure 5 The target block is formed by the aggregation of grids with color 1. The target pollutants exceeding the standard only include As, so As is the pollutant to be controlled. Figure 4 It can be seen that the causal relationship between various emission sources on As concentration is: smelting enterprises > industrial and mining enterprises > transportation. Therefore, the smelting enterprises in the target block are the priority control objects of the target block, and then the industrial and mining enterprises are controlled. Transportation has almost no effect on As concentration and can be excluded from control.
[0108] Case 2: If there are multiple target pollutants exceeding the standard in a target area, the most significant one will be extracted from all the causal relationships verified for all the pollutants exceeding the standard, and the target pollutant corresponding to the most significant one will be recommended as the priority control object for the target area. Figure 5 In the block formed by the grids with color 3, both As and Pb exceed the standard. Figure 4 Among the six causal relationships verified for these two target pollutants, select the one with the most significant causal relationship. Figure 4 (a) The corresponding target pollutant, As, is prioritized for control. This is because the more significant the causal relationship, the more pronounced the expected control effect, and therefore should be implemented first. For specific control of As, the same approach as in Case 1 is adopted, with smelters first and then mining enterprises. After As control measures are implemented, Pb control can be implemented, still following the approach described in Case 1, with limited control targets determined in the order of mining enterprises, transportation, and smelters.
[0109] Further, Figure 6 The process from data collection to causal inference in the above method is shown in the form of a technical framework diagram. Figure 6 As shown, the technical framework of this embodiment includes the following stages:
[0110] Source Identification: Based on soil testing data and PMF receptor models in specific areas, we can identify and quantify pollution sources, which are hidden pollution sources and cannot be directly linked to actual controllable objects.
[0111] Source Inventory Calculation: This method estimates total pollutant emissions by systematically collecting emission factors and activity data for sources such as industrial activities and traffic flows within a specific area. This method directly links explicit emission sources and identifies high-emission hotspots.
[0112] Source Verification: Source verification is a spatial causal inference method based on dynamic system theory and the generalized embedding theorem, used to verify the causal relationship between emission sources and pollutant distribution. In this embodiment, this process also makes implicit pollution sources explicit. Specifically, implicit pollution sources are derived from actual detection data and are relatively accurate pollution sources, but they are difficult to correspond to explicit locations or controllable factors. This embodiment uses GCCM causal inference to correspond implicit pollution sources to explicit emission sources such as enterprises, industries, and transportation, achieving explicit tracing of implicit pollution sources and providing a basis for formulating implementable fine-grained control measures.
[0113] To further demonstrate the advantages of this example's method, this example uses Gejiu City, Yunnan Province, as an example. Using the Pearson correlation coefficient method and the MGWR (Multiscale Geographically Weighted Regression) method, we conducted correlation analyses between As and Pb, two pollutants, and three types of data: smelting emissions, industrial and mining emissions, and transportation emissions. Specifically, the correlation analysis used the pollutant concentration and emissions (if available) for each grid as data points. The final verification results obtained by the three methods are shown in the following table:
[0114]
[0115] In the table, ** indicates significant correlation at the 0.01 level, and * indicates significant correlation at the 0.05 level; r indicates the Pearson correlation coefficient, ARC indicates the average regression coefficient in MGWR, and r / ARC / ρ indicate the corresponding parameters in different methods; p indicates the Pearson correlation coefficient and the p-value of the significance test in GCCM; SRR indicates the proportion of significant regions, that is, the proportion of spatial regions with p-values less than the significance level in the MGWR analysis to the total area of the study area.
[0116] The table above demonstrates that traditional methods lack causal inference capabilities. Pearson correlation analysis only provides statistical correlations, and causal relationships can be obscured by other sources of collinearity, leading to "spurious correlations" (for example, two data sets that appear correlated may not actually have a causal relationship). The MGWR method, on the other hand, relies on linear assumptions and cannot detect nonlinear causal chains in weakly coupled systems. Furthermore, it requires high spatiotemporal data, which reduces its reliability in real-world scenarios when data is scarce. Therefore, in this example, the results of the two methods above are insignificant or low (no asterisk or less than 0.01), and it is impossible to determine the cause and effect. The GCCM method, on the other hand, achieves significantly higher significance results and, by following the principle that "the effect of X predicting Y is significant and converges, while the effect of Y predicting X is insignificant or approaches zero," can confirm a unidirectional causal relationship from X to Y (X is the cause, Y is the effect).
[0117] In addition, this example also takes Gejiu City, Yunnan Province as an example, and uses the density of industrial and mining enterprises, the density of smelting enterprises, and the road network density as variables to be verified, and verifies their causal relationship with the concentration of target pollutants. Taking the target pollutant As as an example, the verification results are as follows: Figure 7 Specifically, the density of smelting enterprises in each grid is arranged into a two-dimensional distribution map and the two-dimensional distribution map of As detection concentration is used for GCCM causal inference to obtain Figure 7(a) shows the causal relationship verification results, where xxmap y represents the predicted As detection concentration based on the density of smelting enterprises, and y xmap x represents the predicted smelting enterprise density based on the As detection concentration. The GCCM causal inference is performed on the two-dimensional distribution map of the density of industrial and mining enterprises in each grid and the two-dimensional distribution map of the As detection concentration, and the following is obtained: Figure 7 (b) shows the causal relationship verification results, where x xmap y represents the predicted As detection concentration based on the density of industrial and mining enterprises, and y xmap x represents the predicted industrial and mining enterprise density based on the As detection concentration. The GCCM causal inference is performed on the two-dimensional distribution map of the road network density of each grid and the two-dimensional distribution map of the As detection concentration, and the following is obtained: Figure 7 (c) shows the causal relationship verification results, where x xmap y represents the predicted As concentration based on road network density, and y xmap x represents the predicted road network density based on As concentration. It can be seen that for areas where industry and transportation are the primary sources of pollutant emissions, there is no significant causal relationship between enterprise density, road network density, and pollutant concentrations. This is because, despite the large number of enterprises, the impact of different enterprise sizes and emissions on soil pollution varies; and despite the large number of roads, the impact of emissions from different roads on soil pollution also varies. Therefore, this embodiment uses an emissions inventory and spatial allocation, using the spatial distribution of enterprise emissions from a kernel density model and the spatial distribution of traffic emissions from a road network model as the verification targets for causal inference. This not only establishes a coupling relationship between implicit and explicit pollution sources, but also conforms to the laws governing the effects of industry and transportation on the surrounding environment (e.g., stationary emission sources diffuse pollutants in a point-like manner, while mobile emission sources emit pollutants in a fuzzy manner), achieving more accurate and manageable explicit pollutant source tracing.
[0118] In summary, this embodiment provides a soil pollution source tracing method based on causal relationships. For areas where industry and transportation are the main emission sources, first, the soil detection data is decomposed using the PMF method to determine the hidden pollution sources. Target pollutants with large contributions from hidden pollution sources and related to the emission sources are extracted from dozens of pollutants as tracing objects, and the two-dimensional spatial distribution of the target pollutant concentrations is obtained by interpolation. Then, the emissions of various industrial enterprises and transportation are spatialized using the kernel density function and road density according to the characteristics of enterprise emissions and traffic emissions, respectively, to obtain a more accurate and detailed two-dimensional spatial distribution of emissions. Finally, the GCCM model is used to verify the causal relationship between various emission sources and the spatial distribution of pollutant concentrations, and to identify the nonlinear causal chain in the two weakly coupled systems of PMF and emission inventory. Different control strategies are recommended for different blocks based on the identification results to assist in achieving precise targeted control of soil pollution.
[0119] Specifically, compared with the prior art, this embodiment has the following advantages:
[0120] 1. Existing technologies lack the ability to infer causality in soil pollution. Traditional correlation analysis (such as the Pearson correlation coefficient) and MGWR cannot effectively detect nonlinear causal relationships in weakly coupled systems, making it difficult to reveal the complex interaction chains between pollutant concentration distributions and emission sources. These methods are often based on linear assumptions and have high data requirements (such as time series data). They perform poorly in complex environmental systems, especially when spatiotemporal data is scarce. This limitation is particularly prominent.
[0121] This embodiment adopts the GCCM model, which is based on dynamic system theory and generalized embedding theorem. It reconstructs the state space of the system through spatial cross-sectional data, overcomes the data dependence limitation of traditional methods, and can accurately identify hidden one-way causal relationships, that is, it can clarify the causal direction of enterprises, traffic emission sources and pollutant concentration distribution, avoid "false correlation" interference, and provide a scientific basis for precise governance.
[0122] 2. Existing emission inventories lack spatialization. While existing emission inventory methods can estimate total pollutant emissions, they lack detailed characterization of their spatial distribution, making it difficult to accurately locate high-emission hotspots.
[0123] This embodiment uses kernel density allocation and road network weight allocation methods to accurately allocate emissions from industrial and transportation sources to a spatial grid, generating a high-resolution emission inventory map and providing high-quality input data for subsequent causal inference.
[0124] 3. Existing technologies struggle to analyze complex, multi-source systems. In complex, multi-source, nonlinearly coupled environments, existing technologies struggle to fully analyze the sources of pollutants and their interactions, potentially leading to the neglect of minor pollution sources or the masking of the contributions of major pollution sources.
[0125] This embodiment, through multi-model coupling (PMF+emission inventory+GCCM), combined with spatial analysis and dynamic causal inference, can comprehensively analyze the pollution characteristics of multi-source complex systems and reveal hidden pollution sources and interaction mechanisms.
[0126] 4. Inaccurate pollution control in existing technologies. Existing technology solutions lack a deep understanding of the causal relationship between visible emission sources and pollutant concentration distribution. As a result, pollution control measures are often overly broad and fail to target key pollution sources.
[0127] The embodiment of the present invention integrates PMF (source identification), emission inventory (source calculation) and GCCM (source verification) to build a complete soil pollution traceability and control framework. It can establish a correspondence between the hidden pollution sources in the theoretical model and the explicit emission sources in reality, accurately lock in key pollution sources and their contributing areas, and provide a scientific basis for formulating feasible targeted governance strategies.
[0128] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 8 As shown, the device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the device can be one or more. Figure 8 In the embodiment, a processor 60 is used as an example; the processor 60, the memory 61, the input device 62 and the output device 63 in the device can be connected by a bus or other means. Figure 8 The bus connection is taken as an example.
[0129] Memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the causal relationship-based soil pollution source tracing method in the embodiments of the present invention. Processor 60 executes the software programs, instructions, and modules stored in memory 61 to perform various functional applications and data processing of the device, thereby implementing the causal relationship-based soil pollution source tracing method described above.
[0130] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0131] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 63 may include a display device such as a display screen.
[0132] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the soil pollution source tracing method based on causality of any embodiment.
[0133] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device.
[0134] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0135] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0136] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A soil pollution source tracing method based on causality, characterized in that: include: Obtain soil pollutant concentration detection data in the area to be traced, as well as data on emission sources from enterprises and transportation; Performing positive matrix factorization on the detection data to determine the contribution of each hidden pollution source to each pollutant; Identify at least one pollutant that ranks first in contribution from each hidden pollution source and is related to enterprises and / or transportation as at least one target pollutant to be traced; For each target pollutant, perform the following operations: S1-1. interpolating a first two-dimensional distribution map of target pollutant concentration based on the detection data; S1-2. Execute the emission inventory method based on the data of the two emission sources to obtain the target total pollutant emissions of each enterprise and the target total pollutant emissions of transportation; S1-3. Using the kernel density method, allocate the total target pollutant emissions of each enterprise to each grid in the region to obtain a second two-dimensional distribution map of enterprise emissions; using the road network density weight, allocate the total target pollutant emissions of the traffic to each grid in the region to obtain a third two-dimensional distribution map of traffic emissions; S1-4. Performing GCCM causal analysis on the second two-dimensional distribution map and the third two-dimensional distribution map, respectively, with the first two-dimensional distribution map to verify the causal relationship and significance of the target pollutants between enterprises and transportation. Based on the verification results, auxiliary decisions are made on the blocks that are given priority for control within the area, as well as the types of pollutants and emission sources that are given priority for control in each block.
2. The method according to claim 1, characterized in that The at least one pollutant that is ranked as the top contributor to each hidden pollution source and is associated with enterprises and / or transportation is determined as the at least one target pollutant to be traced, including: Determining, based on the data of the two emission sources, at least one pollutant associated with the enterprise or traffic to form a first pollutant set; For each hidden pollution source: sort the pollutants according to the contribution ratio of the hidden pollution source to each pollutant, and select the pollutants that rank at the top; The union of the top pollutants corresponding to each hidden pollution source is taken to form the second pollutant set; The intersection of the first pollutant set and the second pollutant set is taken, and at least one pollutant in the intersection is determined as at least one target pollutant to be traced.
3. The method according to claim 1, characterized in that The emission inventory method is implemented based on the data of the two emission sources to obtain the target total pollutant emissions of each enterprise and the target total pollutant emissions of transportation, including: Determine the production emissions of each enterprise based on the activity data of each enterprise in the production process and the emission coefficient of the target pollutants; The target total pollutant emissions for each enterprise shall be determined based on the production emissions and removal efficiency of pollution control technologies of each enterprise.
4. The method according to claim 1, wherein The kernel density method is used to allocate the total target pollutant emissions of each enterprise to each grid in the region, including: According to a preset spatial weight function, the total target pollutant emissions of each enterprise are sequentially allocated to multiple grids centered on each enterprise within the region, wherein the sum of the emissions of the same enterprise allocated to each grid is equal to the total emissions of the same enterprise, and the farther the distance between each grid and the same enterprise, the greater the corresponding spatial weight; The emission amounts allocated to each grid from each enterprise are added together to obtain the total target pollutant emissions for each grid.
5. The method according to claim 1, wherein The method of allocating the target total amount of traffic pollutant emissions to each grid in the region using the road network density weight includes: Determining a road network density weight for each grid according to a ratio of the total road length of each grid in the region to the total road length of the region; The target total amount of pollutant emissions from the traffic is allocated to each grid according to the road network density weight of each grid.
6. The method according to claim 1, characterized in that The second two-dimensional distribution map and the third two-dimensional distribution map are subjected to GCCM causal analysis with the first two-dimensional distribution map to verify the causal relationship and significance of the target pollutants between enterprises and transportation, respectively, including: Inputting the third two-dimensional distribution graph and the first two-dimensional distribution graph into the GCCM causal analysis model, the model predicting the target pollutant concentration based on traffic emissions for a series of window sizes, calculating the prediction skill in each window, and plotting a first broken line showing how the prediction skill varies with the window size; concurrently, the model predicting traffic emissions based on the target pollutant concentration for a series of window sizes, calculating the prediction skill in each window, and plotting a second broken line showing how the prediction skill varies with the window size; If the prediction skill in the first broken line increases with the increase of the window size, and the prediction skill of the first broken line under the maximum window is significantly higher than that of the second broken line, it is judged that traffic affects the concentration of the target pollutant, and the greater the distance between the two broken lines, the more significant the causal relationship.
7. The method according to claim 1, characterized in that The method of allocating the total target pollutant emissions of each enterprise to each grid in the region using the kernel density method to obtain a second two-dimensional distribution map of enterprise emissions includes: classifying each enterprise according to industry; and allocating the total target pollutant emissions of each enterprise in the same industry to each grid in the region using the kernel density method to obtain a second two-dimensional distribution map of emissions of enterprises in the same industry. Accordingly, the second two-dimensional distribution map and the third two-dimensional distribution map are respectively subjected to GCCM causal analysis with the first two-dimensional distribution map to respectively verify the causal relationship between enterprises and transportation and the target pollutants and the degree of significance thereof, including: performing GCCM causal analysis on the second two-dimensional distribution map of emissions from enterprises in various industries and the first two-dimensional distribution map to respectively verify the causal relationship between enterprises in various industries and the target pollutants and the degree of significance thereof.
8. The method according to claim 7, characterized in that Based on the verification results, auxiliary decision-making is made on the priority control blocks in the region, as well as the types of pollutants and emission sources that are prioritized for control in each block, including: Mixing and displaying the first two-dimensional distribution map of the at least one target pollutant, and extracting at least one target block where the concentration of the target pollutant exceeds the standard; For target areas where only one target pollutant exceeds the standard, the emission source type with the most significant causal relationship with the target pollutant is extracted from various industrial enterprises and transportation emission source types, and recommended as the priority control target for the target area; For target blocks where multiple target pollutants exceed the standards, the most significant one is extracted from all causal relationships verified for all pollutants exceeding the standards, and the target pollutant corresponding to the most significant one is recommended as the priority control object for the target block.
9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the soil pollution source tracing method based on causality as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the soil pollution source tracing method based on causal relationship described in any one of claims 1 to 8.
Citation Information
Patent Citations
Soil pollutant tracing method
CN116559148A
Soil pollutant traceability analysis method, apparatus and device, and storage medium
CN118171060A
Soil heavy metal risk identification method based on space interaction relationship
CN112288247A
Heavy metal source analysis method based on PMF and geographic detector model
CN115598204A