Air pollution dynamic tracing method based on multi-source data association analysis
By using multi-dimensional data cleaning and multi-source data correlation analysis, the problems of data quality and pollution source location in air pollution source tracing were solved, enabling accurate identification of pollution source types and assessment of emission entities, and improving the efficiency of pollution control.
Patent Information
- Application Number
- CN202610040928.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-02-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing air pollution source tracing technologies have shortcomings in data quality control, multi-source data correlation, and terrain and industry adaptation, making it difficult to achieve accurate and dynamic pollution source location and assessment.
By employing a multi-dimensional data cleaning mechanism, a dynamic sliding window algorithm, a terrain-corrected wind field trajectory inversion model, and a multi-source data association rule base, combined with Pearson correlation coefficient and mutual information entropy index, an air quality dataset is constructed to optimize pollution source identification and emission subject assessment.
It improves the accuracy and timeliness of air pollution source tracing, ensures data integrity, accurately identifies pollution source types and emission entities, and enhances pollution control efficiency.
Smart Images

Figure CN121504701A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of environmental monitoring and data analysis, specifically relating to a method for dynamic source tracing of air pollution based on multi-source data correlation analysis. Background Technology
[0002] With the acceleration of industrialization and urbanization, air pollution has become a critical issue affecting human health, the ecological environment, and sustainable social development. Airborne pollutants, including particulate matter and gaseous pollutants, originate from various sources, including stationary sources (such as industrial enterprises), mobile sources (such as motor vehicles), and area sources (such as straw burning and dust). Furthermore, influenced by meteorological conditions, topographical features, and diffusion patterns, they exhibit uneven spatial and temporal distribution and complex causes. Therefore, accurately locating pollution sources and identifying the main contributors to pollution are core prerequisites for developing targeted pollution control strategies and effectively reducing pollution concentrations, and are also a core requirement in the current environmental management field.
[0003] While current air pollution source tracing technologies have developed into various research directions, they still have limitations in practical applications, making it difficult to meet the needs for accurate, dynamic, and comprehensive source tracing. Traditional source tracing methods rely on monitoring station data. However, existing monitoring networks have coverage blind spots. Due to terrain limitations (such as mountainous areas and remote suburbs) or economic constraints, monitoring stations are sparsely deployed in some areas, resulting in missing pollution concentration information in these blind spots. At the same time, monitoring data is susceptible to delays caused by communication link interference during transmission, and sensor failures and extreme weather interference can lead to extreme outliers. Existing data processing mechanisms mostly use simple interpolation or threshold removal methods, without forming a multi-dimensional cleaning system that addresses "transmission delay, extreme anomalies, and blind spot completion." As a result, the data quality is insufficient to support accurate source tracing analysis.
[0004] Existing technologies for analyzing pollutants are often limited to a single dimension or a small amount of data correlation. Some methods rely solely on the time series changes of a single pollutant to determine the pollution source, ignoring the synergistic relationship between different pollutants. Even when using multi-site data, they often aggregate sites through simple correlation analysis without considering the diffusion characteristics of pollutants (such as the difference between linear diffusion and nonlinear reactive pollution) to optimize correlation indicators. This leads to biases in the division of site clusters, and consequently, misjudgments of the location and diffusion range of pollution sources.
[0005] The adaptability of wind field trajectory inversion to pollution source type identification is insufficient. Traditional wind field trajectory models are mostly constructed based on standard atmospheric conditions and do not fully consider the disturbance of wind field by regional topography (such as mountain bypass, river valley funneling effect, and urban building cluster obstruction), resulting in deviation in the location of upwind pollution source areas. In terms of pollution source type identification, they mostly rely on single industry attributes or pollutant concentration data and have not established a deep mapping relationship between "pollutant component fingerprints and industrial emission characteristics", making it difficult to distinguish the pollution contribution of different types of industries in the same region and easily leading to misidentification of pollution source types (such as misidentifying emissions from chemical enterprises as emissions from steel enterprises).
[0006] Weakness in multi-source data integration and pollution source identification: Although current technologies have begun to incorporate multi-source information such as remote sensing fire points, vehicle positioning, and enterprise pollution control facility data, they mostly remain at the level of data overlay and fail to explore deep coupling relationships (such as the temporal synergy between enterprise emission concentrations and the power load of pollution control facilities, and the spatial correlation between vehicle trajectories and pollution peaks). At the same time, source tracing of stationary sources, mobile sources, and area sources is mostly carried out independently, lacking cross-source collaborative assessment mechanisms. It is difficult to quantify the pollution contribution of different emission entities, resulting in the inability to accurately identify specific abnormal emission entities, ultimately affecting the efficiency of pollution control measures.
[0007] Furthermore, with the development of monitoring technology, the types of air quality data are becoming increasingly diverse (such as minute-level online monitoring data, high-resolution remote sensing data, and real-time vehicle trajectory data). However, existing technologies lack the ability to construct spatiotemporal alignment and association rules for multi-source heterogeneous data, resulting in a large amount of data resources not being effectively utilized. This not only wastes data but also makes it difficult to balance dynamism and accuracy in the source tracing results.
[0008] In summary, current air pollution source tracing technologies have significant shortcomings in areas such as data quality control, multi-source data correlation, terrain and industry adaptation, and identification of specific emission entities. A dynamic source tracing method that can integrate multi-source data, optimize data processing mechanisms, and strengthen multi-dimensional correlation analysis is needed to improve the accuracy and timeliness of air pollution source tracing and provide scientific and reliable technical support for pollution control. Summary of the Invention
[0009] To address the aforementioned problems in the existing technology, this invention provides a method for dynamic source tracing of air pollution based on multi-source data correlation analysis.
[0010] The objective of this invention can be achieved through the following technical solutions: A dynamic source tracing method for air pollution based on multi-source data correlation analysis includes: S1: Delineate the air pollution impact area, generate air quality data based on the obtained pollution factors, construct a multi-dimensional data cleaning mechanism, generate and identify data transmission delay characteristics and extreme outlier distribution patterns, supplement the pollution concentration information in blind areas, and generate an air quality dataset; S2: Time series analysis is performed on the pollution factors respectively. The pollution concentration change sequence is segmented using a dynamic sliding window algorithm. The degree of coordination of pollution concentration changes between different stations is calculated by combining the Pearson correlation coefficient and mutual information entropy. Stations with coordination reaching a preset threshold are aggregated to form a station cluster. Based on the spatial distribution characteristics and concentration change synchronicity of the station cluster, a pollution source judgment is generated. S3: Based on the pollution source determination, obtain real-time wind direction and speed data, input the terrain correction wind field trajectory inversion model to locate the upwind pollution source area; collect the pollutant component characteristic data and regional industry attribute data of the upwind pollution source area, establish a mapping relationship, compare with the preset regional pollutant characteristic database, generate the component fingerprint difference of the pollution source, and generate the type of the pollution source by combining the emission characteristics of the regional industry. S4: Retrieve multi-source correlation data within the upwind pollution source area, construct a multi-source data correlation rule base, mine the coupling relationship between enterprise emission data and power consumption data of treatment facilities, generate the temporal synchronization and spatial correlation between vehicle driving trajectory and emission activities, combine remote sensing fire point data with spatial distance threshold verification of surrounding pollution sources, evaluate the emission level of emission subjects in multiple dimensions, and generate emission anomaly ranking results.
[0011] Specifically, the implementation process of the multi-dimensional data cleaning mechanism is as follows: For the data transmission delay feature, the concentration changes of adjacent monitoring stations within the delay period are first split into trend terms, periodic terms and residual terms using a time series decomposition algorithm. Then, based on the linear continuity of the trend terms and the repetitive pattern of the periodic terms, the trend extrapolation method is selected to complete the data missing due to the delay.
[0012] Specifically, the process of generating air quality data by fusing pollutants is as follows: a weighted fusion algorithm is used to integrate the pollutant data. First, the analytic hierarchy process (AHP) is used to assign subjective weights based on the importance of the pollutants to human health and the environment. Then, the entropy weight method is used to assign objective weights based on the dispersion of the pollutant data. The final weight is the normalized weighted result of the two. For pollutant data from different sources, standardization is first used to eliminate dimensional differences, and then spatiotemporal consistency verification is performed to finally generate a spatiotemporally unified air quality dataset.
[0013] Specifically, when the dynamic sliding window algorithm segments the pollution concentration change sequence, it adopts an adaptive window adjustment strategy: first, it calculates the pollution concentration change rate through the first derivative, sets a change rate threshold to distinguish between smooth segments and fluctuating segments, uses the least squares method to fit the trend of the segmented sequence, and calculates the sum of squared residuals between the fitted curve and the segmented sequence.
[0014] Specifically, when calculating the degree of synergy between pollution concentration changes at different stations using the Pearson correlation coefficient and mutual information entropy dual indicators, a pollution concentration dataset based on time series is first constructed. For the pollutant concentration data collected from the monitoring stations, the influence of dimensions is eliminated through standardization processing. The similarity of pollution concentration change trends is quantified by calculating the ratio of the product of the covariance and the standard deviation of the data. The mutual information entropy captures nonlinear dependencies by calculating the relative entropy of the joint probability distribution and the marginal probability distribution. A two-dimensional screening model is constructed by setting thresholds for the Pearson correlation coefficient and mutual information entropy.
[0015] Specifically, when generating pollution source judgment based on the spatial distribution characteristics and concentration change synchronicity of the site cluster, optimization is performed by combining regional underlying surface characteristics: first, the regional underlying surface type is divided, then the spatial correlation between the site cluster and the underlying surface is analyzed, and the temporal difference of the concentration peak of each site in the site cluster is calculated through cross-correlation analysis, and the location of the pollution source is determined according to the direction of the difference.
[0016] Specifically, the specific correction process of the terrain-corrected wind field trajectory inversion model is as follows: after inputting the real-time wind direction and speed data, the regional terrain data is first imported, and the influence of terrain on the wind field is calculated based on the atmospheric boundary layer model in fluid dynamics; then, the Lagrange trajectory inversion algorithm is used to simulate the pollutant diffusion path, and the upwind trajectory obtained by inversion is combined with the spatial location of terrain obstacles to make an effectiveness judgment, generate effective terrain and wind direction trajectories, and locate the upwind pollution source area.
[0017] Specifically, the process of generating the component fingerprint difference of the pollution source is as follows: from the pollutant component feature data, feature components are screened based on preset principles, the relative content of the feature components is calculated using a normalization algorithm, and the ratio between feature components is calculated at the same time. The component fingerprint generated by fusing the relative content and the feature ratio is compared with the standard component fingerprint of different types of pollution sources in the preset regional pollutant feature database. The difference between the two is quantified using Euclidean distance, and the matching degree of fingerprint morphology is judged by combining cosine similarity.
[0018] Specifically, the acquired regional industry attribute data includes industry type, production process, raw material consumption types, main products, and historical emission records. When establishing the mapping relationship between the pollutant component characteristic data and the regional industry attribute data, an association rule mining algorithm is used: first, the component characteristics and industry attributes are treated as frequent itemsets, and the support and confidence between the itemsets are calculated; then, strong association rules with support and confidence reaching preset standards are selected to form an industry attribute-component characteristic mapping table.
[0019] Specifically, a multi-dimensional spatiotemporal alignment optimization process is performed on the multi-source associated data. This process is implemented as follows: In the time dimension, considering the differences in sampling frequency of data from different sources, a combination of linear interpolation and sliding window averaging is used to standardize all data to the same time granularity. At the same time, based on the timestamp deviation records of data collection, time offset correction is performed on each data source. In the spatial dimension, the spatial coordinates of enterprise locations, vehicle trajectories, remote sensing fire points, and monitoring stations are uniformly transformed to the same coordinate system. After spatiotemporal alignment, a three-dimensional index of time-space-data features is constructed.
[0020] Specifically, the retrieved multi-source associated data includes fixed source data, mobile source data, and area source data, employing a classification-adaptive preprocessing mechanism: fixed source data covers emission concentration data monitored by enterprises, electricity load data of treatment facilities, and emission limit parameters approved by environmental protection authorities. During preprocessing, timestamps are aligned, and instantaneous spikes and anomalies in the electricity data are removed; mobile source data includes freight vehicle positioning trajectory data, exhaust gas remote sensing monitoring data, and traffic flow statistics. During preprocessing, trajectory drift is corrected by combining the road network topology, and vehicle identity is bound to emission data through license plate association; area source data includes remote sensing data of straw burning fire points, agricultural fertilization records, and dust monitoring data. During preprocessing, regional attribution labels are generated by matching the latitude and longitude of fire points with administrative divisions.
[0021] Specifically, the multi-dimensional assessment of the emission level of the emission subject constructs a three-layer indicator system of emission intensity, facility operation, and correlation verification: the emission intensity indicator adopts the weighted result of emission per unit output value and the frequency of exceeding the standard, and the weight is based on the preset industrial pollution intensity level; the facility operation indicator includes the commissioning rate of treatment facilities, the compliance rate of electricity load, and the completeness score of maintenance records; the correlation verification indicator includes the time-series correlation coefficient between emission data and fire point data, and the spatial distance deviation between vehicle trajectory and concentration peak. The three-layer indicators are weighted by the analytic hierarchy process and then the comprehensive evaluation value is calculated.
[0022] The beneficial effects of this invention are as follows: To address the issues of transmission delays, extreme outliers, and spatiotemporal inconsistencies in traditional source tracing using multi-source data (ground monitoring, mobile monitoring, and remote sensing inversion), this invention achieves a breakthrough through a multi-dimensional data cleaning mechanism and a weighted fusion strategy. On the one hand, it completes delayed data based on time series decomposition and trend extrapolation, and corrects outliers by combining meteorological factors and spatial correlation to ensure data integrity. On the other hand, it uses a weighted fusion method combining the analytic hierarchy process (AHP) and entropy weighting to balance the influence weight of pollution factors and data dispersion, and simultaneously performs spatiotemporal consistency verification to eliminate dimensional differences and spatiotemporal misalignments. The resulting air quality dataset has significantly improved accuracy, avoiding source tracing deviations caused by data errors.
[0023] Traditional methods often rely on single-site data or simple concentration comparisons to determine pollution sources, which are susceptible to diffusion interference and misjudgment of the scope. This invention uses a dynamic sliding window algorithm to adaptively segment the pollution concentration sequence, accurately capturing details of concentration fluctuations. By combining Pearson correlation coefficient (quantifying linear synergy) and mutual information entropy (capturing nonlinear dependence) as dual indicators, it achieves accurate quantification of the degree of synergy among different sites, and the aggregated site clusters better reflect the pollution diffusion pattern. At the same time, it incorporates regional underlying surface characteristics (road network, industrial park, agricultural area) and the temporal differences in peak concentrations of site clusters, which can initially distinguish between mobile sources (linearly distributed along the road network with delayed peaks) and stationary sources (area-like distribution around a specific area with synchronous fluctuations), and pinpoint the approximate location of the pollution source, narrowing the subsequent source tracing scope to the target area and improving source tracing efficiency.
[0024] Traditional wind field trajectory inversion often ignores the influence of topography, leading to inaccuracies in locating upwind pollution source areas. Pollution source type identification relies heavily on experience, easily confusing industries with similar emission characteristics. This invention's topography-corrected wind field trajectory inversion model corrects for the deflection and attenuation effects of mountains, valleys, and building clusters on the wind field using an atmospheric boundary layer model. Combined with Lagrange trajectory inversion to eliminate invalid trajectories, the positioning accuracy is significantly improved compared to traditional models. Simultaneously, by selecting identifiable pollutant components to generate "component fingerprints," which are compared with a pre-set feature library (using Euclidean distance and cosine similarity), and combined with strong correlation rules between industry attributes and component features (mined using the Apriori algorithm), it can match pollution source types (such as steel mills, oil refineries, and straw burning), avoiding misclassification due to "different industries with the same component," and providing a clear direction for subsequently identifying the main emitter.
[0025] Traditional source tracing methods often focus on "regions" rather than "specific entities," and their singular assessment dimensions can easily lead to distorted anomaly ranking. This invention addresses this by preprocessing multi-source correlated data (stationary source electricity consumption-emission coupling, mobile source trajectory-emission binding, area source fire point-regional attribution) and optimizing spatiotemporal alignment to construct a three-dimensional index of "time-space-data characteristics," ensuring the effectiveness of data correlation. Based on a three-layer indicator system of "emission intensity-facility operation-correlation verification" (emissions per unit output value, facility operation rate, fire point temporal correlation, etc.), combined with weighted calculation of comprehensive evaluation values using the analytic hierarchy process, the generated emission anomaly ranking can accurately identify high-risk entities (such as enterprises falsely operating treatment facilities, vehicles with strong correlations between trajectory and concentration peaks). Simultaneously, cross-validation of anomaly results (resource consumption correlation, vehicle maintenance records, ground inspection traces) further ensures credibility, providing regulatory authorities with a basis for "precise positioning and tiered handling," avoiding blind inspections and improving pollution control efficiency. Attached Figure Description
[0026] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0027] Figure 1 This is a flowchart illustrating a dynamic source tracing method for air pollution based on multi-source data correlation analysis according to the present invention. Figure 2 This is a diagram illustrating the multi-dimensional data cleaning and fusion mechanism in this invention; Figure 3 This is a diagram illustrating the site collaborative analysis and preliminary pollution source identification in this invention. Detailed Implementation
[0028] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.
[0029] Please see Figure 1-3 A dynamic source tracing method for air pollution based on multi-source data correlation analysis includes: S1: Delineate the air pollution impact area, generate air quality data based on the obtained pollution factors, construct a multi-dimensional data cleaning mechanism, generate and identify data transmission delay characteristics and extreme outlier distribution patterns, supplement the pollution concentration information in blind areas, and generate an air quality dataset; S2: Time series analysis is performed on the pollution factors respectively. The pollution concentration change sequence is segmented using a dynamic sliding window algorithm. The degree of coordination of pollution concentration changes between different stations is calculated by combining the Pearson correlation coefficient and mutual information entropy. Stations with coordination reaching a preset threshold are aggregated to form a station cluster. Based on the spatial distribution characteristics and concentration change synchronicity of the station cluster, a pollution source judgment is generated. S3: Based on the pollution source determination, obtain real-time wind direction and speed data, input the terrain correction wind field trajectory inversion model to locate the upwind pollution source area; collect the pollutant component characteristic data and regional industry attribute data of the upwind pollution source area, establish a mapping relationship, compare with the preset regional pollutant characteristic database, generate the component fingerprint difference of the pollution source, and generate the type of the pollution source by combining the emission characteristics of the regional industry. S4: Retrieve multi-source correlation data within the upwind pollution source area, construct a multi-source data correlation rule base, mine the coupling relationship between enterprise emission data and power consumption data of treatment facilities, generate the temporal synchronization and spatial correlation between vehicle driving trajectory and emission activities, combine remote sensing fire point data with spatial distance threshold verification of surrounding pollution sources, evaluate the emission level of emission subjects in multiple dimensions, and generate emission anomaly ranking results.
[0030] Specifically, the implementation process of the multi-dimensional data cleaning mechanism is as follows: For the data transmission delay feature, the concentration changes of adjacent monitoring stations within the delay period are first split into trend terms, periodic terms and residual terms using a time series decomposition algorithm. Then, based on the linear continuity of the trend terms and the repetitive pattern of the periodic terms, the trend extrapolation method is selected to complete the data missing due to the delay.
[0031] To address the identified data transmission delay characteristics, a time series decomposition algorithm (such as STL decomposition) is first used to split the concentration changes of adjacent monitoring stations within the delay period into trend, periodic, and residual terms. Then, based on the linear continuity of the trend term and the repetitive pattern of the periodic term, an appropriate trend extrapolation method is selected (linear trend extrapolation is used when the concentration changes smoothly, and exponential trend extrapolation is used when the concentration fluctuates periodically) to complete the missing data due to the delay. At the same time, meteorological factors such as wind speed and humidity are introduced to correct the extrapolation results and reduce the completion bias caused by meteorological interference. For extreme outliers, the deviation of the outlier from the concentration of surrounding stations in the same period is first calculated. Combined with the regional meteorological background (such as whether there is sudden dust or precipitation), the cause of the anomaly is determined. If it is a sensor malfunction (manifested as data deviating from the normal range for multiple consecutive time periods and a sharp drop in correlation with surrounding stations), the concentration change pattern of the station in the same period over the past three months is used to construct a baseline curve. At the same time, the spatial correlation coefficient between the surrounding stations and the station is calculated. The baseline curve is corrected according to the correlation coefficient weight to generate alternative values for the outlier, ensuring the completeness and accuracy of the air quality dataset.
[0032] Specifically, the process of generating air quality data by fusing pollutants is as follows: a weighted fusion algorithm is used to integrate the pollutant data. First, the analytic hierarchy process (AHP) is used to assign subjective weights based on the importance of the pollutants to human health and the environment. Then, the entropy weight method is used to assign objective weights based on the dispersion of the pollutant data. The final weight is the normalized weighted result of the two. For pollutant data from different sources, standardization is first used to eliminate dimensional differences, and then spatiotemporal consistency verification is performed to finally generate a spatiotemporally unified air quality dataset.
[0033] For pollution factor data from different sources (ground station monitoring data, mobile monitoring data, and remote sensing inversion data), the data preprocessing stage uses Z-score standardization to eliminate dimensional differences. The formula is as follows: Where x represents the original data, μ is the data mean, and σ is the standard deviation. Regarding spatiotemporal consistency verification, in the time dimension, for high-frequency data from mobile monitoring, a moving average method is used for downsampling, with a 1-hour sampling frequency from ground station monitoring as the benchmark to ensure data timestamp alignment; for missing values in the time series, cubic spline interpolation is used for imputation. In the spatial dimension, Kriging interpolation is used to map data from different sources to a unified spatial grid, with a grid resolution set to 1km × 1km. The relative deviation of data from different sources within the grid is calculated using the following formula: ,in , These are data values from different sources within the same grid.
[0034] Specifically, when the dynamic sliding window algorithm segments the pollution concentration change sequence, it adopts an adaptive window adjustment strategy: first, it calculates the pollution concentration change rate through the first derivative, sets a change rate threshold to distinguish between smooth segments and fluctuating segments, uses the least squares method to fit the trend of the segmented sequence, and calculates the sum of squared residuals between the fitted curve and the segmented sequence.
[0035] Specifically, when calculating the degree of synergy between pollution concentration changes at different stations using the Pearson correlation coefficient and mutual information entropy dual indicators, a pollution concentration dataset based on time series is first constructed. For the pollutant concentration data collected from the monitoring stations, the influence of dimensions is eliminated through standardization processing. The similarity of pollution concentration change trends is quantified by calculating the ratio of the product of the covariance and the standard deviation of the data. The mutual information entropy captures nonlinear dependencies by calculating the relative entropy of the joint probability distribution and the marginal probability distribution. A two-dimensional screening model is constructed by setting thresholds for the Pearson correlation coefficient and mutual information entropy.
[0036] For highly diffusive pollutants (such as PM2.5 and PM10, whose concentration changes are linearly correlated with airflow diffusion), the weight of the Pearson correlation coefficient is increased. This coefficient quantifies the degree of linear synchronization of concentration changes among stations, while eliminating spurious correlations caused by concentration values close to the detection limit. For highly reactive pollutants with nonlinear concentration changes (such as ozone and nitrogen oxides, whose concentrations fluctuate nonlinearly due to photochemical reactions and the intermittent influence of emission sources), the weight of the mutual information entropy is increased. This indicator measures the nonlinear dependence of concentration changes among stations (without assuming that the data follows a specific distribution). The results of the two indicators are weighted and fused, with the weight values determined according to the reactivity level of the pollutants (the higher the reactivity, the greater the weight of the mutual information entropy). The fused result serves as the final quantitative basis for the degree of coordination among stations.
[0037] Specifically, when generating pollution source judgment based on the spatial distribution characteristics and concentration change synchronicity of the site cluster, optimization is performed by combining regional underlying surface characteristics: first, the regional underlying surface type is divided, then the spatial correlation between the site cluster and the underlying surface is analyzed, and the temporal difference of the concentration peak of each site in the site cluster is calculated through cross-correlation analysis, and the location of the pollution source is determined according to the direction of the difference.
[0038] First, the underlying surface types of the region are classified (urban road network, industrial park, agricultural area, nature reserve). Then, the spatial relationship between the site cluster and the underlying surface is analyzed: if the site cluster is linearly distributed along the urban road network and the concentration change shows intermittent peaks (coinciding with peak traffic hours), and the synchronicity is shown as "gradually delayed along the extension direction of the road network", then the pollution source is initially judged to be a mobile emission source (such as motor vehicles); if the site cluster is distributed in a planar pattern around the industrial park, the concentration change shows a continuous high value, and gradually decreases from the center of the park to the periphery (consistent with the diffusion law of stationary sources), and the synchronicity is shown as "the sites around the park fluctuate almost synchronously", then the pollution source is initially judged to be a stationary emission source in the region; at the same time, the temporal difference of the peak concentration of each site in the site cluster is calculated by cross-correlation analysis, and the direction of the difference (such as the peak gradually appearing from northwest to southeast) is used to help judge the approximate location of the pollution source (the northwest direction is the potential pollution source area).
[0039] Specifically, the specific correction process of the terrain-corrected wind field trajectory inversion model is as follows: after inputting the real-time wind direction and speed data, the regional terrain data is first imported, and the influence of terrain on the wind field is calculated based on the atmospheric boundary layer model in fluid dynamics; then, the Lagrange trajectory inversion algorithm is used to simulate the pollutant diffusion path, and the upwind trajectory obtained by inversion is combined with the spatial location of terrain obstacles to make an effectiveness judgment, generate effective terrain and wind direction trajectories, and locate the upwind pollution source area.
[0040] After inputting real-time wind direction and speed data, high-resolution regional terrain data (including mountain slope, valley direction, building height and distribution) is first imported. Based on the atmospheric boundary layer model in fluid dynamics, the impact of terrain on the wind field is calculated: for mountainous terrain, the wind field deflection angle is corrected according to the slope (the greater the slope, the greater the deflection angle caused by the flow around the mountain), and the wind speed attenuation coefficient is calculated according to the mountain height (the higher the height, the more significant the wind speed attenuation); for valley terrain, the wind speed enhancement and wind direction fixation caused by the funneling effect are considered (the wind direction shifts along the valley direction); for urban building clusters, the roughness length is calculated according to the building density, and the near-surface wind speed is corrected (the greater the density, the more obvious the wind speed attenuation); then, the Lagrange trajectory inversion algorithm is used to simulate the pollutant diffusion path. The validity of the inverted upwind trajectory is judged in combination with the spatial location of terrain obstacles: if the trajectory needs to pass through mountains with an altitude exceeding the pollutant diffusion height, or if the trajectory direction is opposite to the corrected wind direction, it is judged as an invalid trajectory and discarded. Finally, the valid trajectory that matches the terrain and wind direction is retained to accurately locate the upwind pollution source area.
[0041] Specifically, the process of generating the component fingerprint difference of the pollution source is as follows: from the pollutant component feature data, feature components are screened based on preset principles, the relative content of the feature components is calculated using a normalization algorithm, and the ratio between feature components is calculated at the same time. The component fingerprint generated by fusing the relative content and the feature ratio is compared with the standard component fingerprint of different types of pollution sources in the preset regional pollutant feature database. The difference between the two is quantified using Euclidean distance, and the matching degree of fingerprint morphology is judged by combining cosine similarity.
[0042] Specifically, the acquired regional industry attribute data includes industry type, production process, raw material consumption types, main products, and historical emission records. When establishing the mapping relationship between the pollutant component characteristic data and the regional industry attribute data, an association rule mining algorithm is used: first, the component characteristics and industry attributes are treated as frequent itemsets, and the support and confidence between the itemsets are calculated; then, strong association rules with support and confidence reaching preset standards are selected to form an industry attribute-component characteristic mapping table.
[0043] From the collected pollutant component characteristic data, characteristic components are screened based on the principle of "strong identifiability and high stability"—prioritizing components that have a significantly higher proportion in specific pollution sources than the background concentration and are not easily consumed by atmospheric chemical reactions (such as specific heavy metals emitted by the steel industry and characteristic aromatic hydrocarbons emitted by the chemical industry). The relative content of each characteristic component (concentration of a certain component / total concentration of all characteristic components) is calculated using a normalization algorithm. At the same time, the ratios between characteristic components with discriminative power (such as the concentration ratio of different heavy metal elements and the peak area ratio of different organic compounds) are calculated. These ratios can effectively eliminate the interference of diffusion distance on the absolute value of concentration. The "component fingerprint" composed of the relative content and the characteristic ratio is compared with the standard component fingerprints of different types of pollution sources (such as steel plants, oil refineries, and waste incineration plants) in the preset regional pollutant characteristic database. The difference between the two is quantified by Euclidean distance (the smaller the distance, the higher the similarity). At the same time, the matching degree of fingerprint morphology is judged by combining the cosine similarity. The combined result of the two is used as the core basis for subsequent judgment of pollution source type.
[0044] Specifically, a multi-dimensional spatiotemporal alignment optimization process is performed on the multi-source associated data. This process is implemented as follows: In the time dimension, considering the differences in sampling frequency of data from different sources, a combination of linear interpolation and sliding window averaging is used to standardize all data to the same time granularity. At the same time, based on the timestamp deviation records of data collection, time offset correction is performed on each data source. In the spatial dimension, the spatial coordinates of enterprise locations, vehicle trajectories, remote sensing fire points, and monitoring stations are uniformly transformed to the same coordinate system. After spatiotemporal alignment, a three-dimensional index of time-space-data features is constructed.
[0045] Specifically, the retrieved multi-source associated data includes fixed source data, mobile source data, and area source data, employing a classification-adaptive preprocessing mechanism: fixed source data covers emission concentration data monitored by enterprises, electricity load data of treatment facilities, and emission limit parameters approved by environmental protection authorities. During preprocessing, timestamps are aligned, and instantaneous spikes and anomalies in the electricity data are removed; mobile source data includes freight vehicle positioning trajectory data, exhaust gas remote sensing monitoring data, and traffic flow statistics. During preprocessing, trajectory drift is corrected by combining the road network topology, and vehicle identity is bound to emission data through license plate association; area source data includes remote sensing data of straw burning fire points, agricultural fertilization records, and dust monitoring data. During preprocessing, regional attribution labels are generated by matching the latitude and longitude of fire points with administrative divisions.
[0046] Specifically, the multi-dimensional assessment of the emission level of the emission subject constructs a three-layer indicator system of emission intensity, facility operation, and correlation verification: the emission intensity indicator adopts the weighted result of emission per unit output value and the frequency of exceeding the standard, and the weight is based on the preset industrial pollution intensity level; the facility operation indicator includes the commissioning rate of treatment facilities, the compliance rate of electricity load, and the completeness score of maintenance records; the correlation verification indicator includes the time-series correlation coefficient between emission data and fire point data, and the spatial distance deviation between vehicle trajectory and concentration peak. The three-layer indicators are weighted by the analytic hierarchy process and then the comprehensive evaluation value is calculated.
[0047] In this embodiment, dynamic source tracing is carried out for a sudden air pollution event in a certain area.
[0048] Delineation of pollution impact areas and data acquisition Based on the coverage of the regional monitoring network, an air pollution impact area A0 was delineated, and data on multiple types of pollutants within this area were obtained—including particulate matter factor F1 and gaseous pollutant factor F2. Data sources included monitoring data from ground-based fixed stations D. s Mobile monitoring vehicle data D m Satellite remote sensing inversion data Dᵣ.
[0049] Pollutant weighted fusion The subjective weights of each pollutant were determined using the analytic hierarchy process (AHP). (k corresponds to the factor type, such as k=1 for F1 and k=2 for F2), and the weights are assigned based on the level of the factor's impact on the environment and health. The objective weights of each factor are calculated using the entropy weight method. The information entropy is quantified by the degree of dispersion of factor data (such as the range of concentration fluctuations), and then converted into objective weights. Calculate the overall weight (α and β are weighting adjustment coefficients, set according to data reliability), according to the formula The data is merged to generate a unified air quality data set, Q.
[0050] Multi-dimensional data cleaning Identify data transmission delay characteristics: For missing data within the delay period, use a time series decomposition algorithm to split the concentration changes of adjacent stations into a trend term T, a periodic term C, and a residual term R. Based on the linear continuity of T and the repetitive pattern of C, use trend extrapolation to complete the delayed data and obtain the completed data D1. Handling extreme outliers: Calculate the deviation Δ between the outlier and the concentration of surrounding stations during the same period, combine regional meteorological data (such as whether there is sudden dust storm) to determine the cause of the outlier. If it is a sensor failure, call the historical concentration pattern of the station during the same period to construct a baseline curve L, correct L by weighting it according to the spatial correlation coefficient of surrounding stations, generate outlier replacement values, and finally form an air quality dataset D.
[0051] Site cluster aggregation and preliminary assessment of pollution sources Pollution concentration sequence segmentation The concentration change sequences of each pollutant in dataset D are processed using a dynamic sliding window algorithm: Calculate the rate of change of concentration (ΔC is the concentration difference between adjacent time points, and Δt is the time interval), set a threshold for the rate of change. Distinguish between gentle segments ( ) and fluctuation segmentation ( ); Enlarge the window size for smooth segments Reduce the window size by segmenting the fluctuations. The segmented concentration sequence was obtained. .
[0052] Site Coordination Calculation Calculate the Pearson correlation coefficient r: for concentration sequences at any two sites. , Through formula (Cov is the covariance, σ is the standard deviation) Quantifying linear synergy; Calculate mutual information entropy I: by constructing , joint probability distribution With marginal probability distribution , According to the formula Capturing nonlinear synergy; Set threshold (The threshold for determining r) (The decision threshold of I), when and At that time, it was determined that the coordination between the two sites met the standard.
[0053] Site cluster aggregation and pollution source identification Sites that meet the collaboration criteria will be aggregated to form a site cluster. Analyze the spatial distribution of each site cluster: If site cluster The concentrations are linearly distributed along the road network, and the peak concentrations appear progressively later along the direction of the road network. Based on the underlying surface type (urban road network), a preliminary judgment can be made. Corresponding to mobile source pollution; If site cluster The concentrations are distributed in a planar pattern around an industrial park, decreasing from the center to the periphery. Preliminary judgment suggests... Corresponding to stationary source pollution.
[0054] Upwind area location and pollution source type identification Topographically corrected wind field trajectory inversion Acquire real-time wind direction and speed data Wd (wind direction) and Ws (wind speed), and import regional terrain data. (Including mountain slope and river valley orientation); Based on the atmospheric boundary layer model, the influence of topography on the wind field is corrected: the wind direction deflection angle ΔWd is corrected for mountainous areas, and the wind speed enhancement coefficient k is corrected for river valley areas, resulting in corrected wind field data. , ; The Lagrange trajectory inversion algorithm is used to simulate the pollutant diffusion trajectory. Eliminate invalid tracks that conflict with terrain obstacles (such as mountains), and retain valid tracks. The pollution source area is located upwind, in region A1.
[0055] Correlation between pollutant composition characteristics and industry attributes Pollutant component characteristic data C (including characteristic components) of area A1 were collected. (e.g., specific heavy metals, organic compounds), calculate the relative content of each component. and the ratio of characteristic components (Select component pairs with industry-identifiable characteristics) to form component fingerprints Fp; Retrieve the industry attribute data Ind (including industry type) for region A1 Production process Using the raw material type Ma, an association rule mining algorithm is employed to construct an "industry attribute-component characteristic" mapping table M (e.g., process). Corresponding characteristic components ); Comparison of component fingerprints Compared with the standard fingerprint in the preset area pollutant feature library B Calculate Euclidean distance (Quantify numerical differences) Calculate cosine similarity (Quantitative morphological matching degree), when and ( , When determining the threshold, the pollution source type is determined by combining the mapping table M. (such as stationary sources for chemical industries).
[0056] Multi-source data correlation and emission anomaly ranking Multi-source correlation data preprocessing and rule base construction Retrieve multi-source correlation data for area A1: stationary source data Df (enterprise emission concentration E, power load of treatment facilities P), mobile source data Dm (vehicle trajectory) Exhaust gas monitoring value Em), area source data Ds (remote sensing fire point data F, dust monitoring value Du); Perform spatiotemporal alignment: In the time dimension, all data are unified to the same time granularity (low-frequency data is supplemented by linear interpolation), and in the spatial dimension, all coordinates are transformed to the same coordinate system to construct a three-dimensional index Index(t,s,d) (t is time, s is space, and d is data feature). Construct a multi-source data association rule base Rb, such as "Decrease in power load P of treatment facilities → Increase in emission concentration E" and "Vehicle trajectory". Spatially overlapping with the concentration peak Ep → high contribution from mobile sources.
[0057] Key Coupling Relationship Mining Stationary source coupling analysis: For enterprise data, calculate the time-series correlation coefficient between emission concentration E and electricity load P. ,when When the threshold of negative correlation is reached, the operation of the treatment facility and emissions are determined to be synergistic; if Marked as a potential anomaly; Spatiotemporal matching of mobile sources: An improved hidden Markov model is used to match vehicle trajectories. As an observed sequence, the emission peak Ep is used as a hidden state, and the temporal overlap between the trajectory and the peak is calculated. Spatial distance deviation ,when and When the two are determined to be strongly correlated; Area source fire point verification: Calculate the spatial distance Df between the remote sensing fire point F and the surrounding pollution sources. When Df≤Thf (the threshold of the impact range) and the fire point type matches the pollution source type, increase the anomaly assessment weight of the pollution source.
[0058] Multidimensional emission level assessment and ranking Constructing a three-tiered assessment index: emission intensity index (Emissions per unit of output value × Frequency of exceedances, weighted) ), facility operation indicators (Environmental protection facility commissioning rate × Electricity consumption compliance rate × Maintenance completeness, weighted) ), correlation verification metrics (Time-series correlation coefficient × spatial matching degree, weight) ); Calculate the comprehensive evaluation value Generate a ranking result of the degree of emission anomaly by sorting by S in descending order. The entities ranked first (such as enterprises) ,vehicle (This is) the key target for verification.
[0059] Final output: Air pollution impact area source tracing report, including pollution source types. upwind area The spatial scope, the ranking of abnormal emission entities (Result), and key parameters of each stage (such as the synergy coefficient r, fingerprint similarity) The quantitative results of the comprehensive assessment value (S) provide a basis for pollution control.
[0060] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for dynamic source tracing of air pollution based on multi-source data correlation analysis, characterized in that, include: S1: Delineate the air pollution impact area, generate air quality data based on the obtained pollution factors, construct a multi-dimensional data cleaning mechanism, generate and identify data transmission delay characteristics and extreme outlier distribution patterns, supplement the pollution concentration information in blind areas, and generate an air quality dataset; S2: Time series analysis is performed on the pollution factors respectively. The pollution concentration change sequence is segmented using a dynamic sliding window algorithm. The degree of coordination of pollution concentration changes between different stations is calculated by combining the Pearson correlation coefficient and mutual information entropy. Stations with coordination reaching a preset threshold are aggregated to form a station cluster. Based on the spatial distribution characteristics and concentration change synchronicity of the station cluster, a pollution source judgment is generated. S3: Based on the pollution source determination, obtain real-time wind direction and speed data, input the terrain correction wind field trajectory inversion model to locate the upwind pollution source area; collect the pollutant component characteristic data and regional industry attribute data of the upwind pollution source area, establish a mapping relationship, compare with the preset regional pollutant characteristic database, generate the component fingerprint difference of the pollution source, and generate the type of the pollution source by combining the emission characteristics of the regional industry. S4: Retrieve multi-source correlation data within the upwind pollution source area, construct a multi-source data correlation rule base, mine the coupling relationship between enterprise emission data and power consumption data of treatment facilities, generate the temporal synchronization and spatial correlation between vehicle driving trajectory and emission activities, combine remote sensing fire point data with spatial distance threshold verification of surrounding pollution sources, evaluate the emission level of emission subjects in multiple dimensions, and generate emission anomaly ranking results.
2. The method according to claim 1, characterized in that, In S1, the specific implementation process of the multi-dimensional data cleaning mechanism is as follows: For the data transmission delay feature, the concentration changes of adjacent monitoring stations within the delay period are first split into trend terms, periodic terms and residual terms by time series decomposition algorithm. Then, based on the linear continuity of the trend terms and the repetition pattern of the periodic terms, the trend extrapolation method is selected to complete the data missing due to the delay.
3. The method according to claim 1, characterized in that, In S1, the process of generating air quality data by fusing pollutants is as follows: a weighted fusion algorithm is used to integrate the pollutant data. First, the analytic hierarchy process (AHP) is used to assign subjective weights based on the importance of the pollutants to human health and the environment. Then, the entropy weight method is used to assign objective weights based on the dispersion of the pollutant data. The final weight is the normalized weighted result of the two. For pollutant data from different sources, standardization is first used to eliminate dimensional differences. Then, a spatiotemporal consistency check is performed to finally generate the spatiotemporally unified air quality dataset.
4. The method according to claim 1, characterized in that, In S2, when the dynamic sliding window algorithm segments the pollution concentration change sequence, it adopts an adaptive window adjustment strategy: first, it calculates the pollution concentration change rate through the first derivative, sets a change rate threshold to distinguish between smooth segments and fluctuating segments, uses the least squares method to fit the trend of the segmented sequence, and calculates the sum of squared residuals between the fitted curve and the segmented sequence.
5. The method according to claim 1, characterized in that, In S2, when calculating the degree of synergy between pollution concentration changes at different sites using the Pearson correlation coefficient and mutual information entropy dual indicators, a pollution concentration dataset based on time series is first constructed. For the pollutant concentration data collected from the monitoring sites, the influence of dimensions is eliminated through standardization. The similarity of pollution concentration change trends is quantified by calculating the ratio of the product of the covariance and the standard deviation of the data. The mutual information entropy captures nonlinear dependencies by calculating the relative entropy of the joint probability distribution and the marginal probability distribution. A two-dimensional screening model is constructed by setting thresholds for the Pearson correlation coefficient and mutual information entropy.
6. The method according to claim 1, characterized in that, In S2, when generating pollution source judgment based on the spatial distribution characteristics and concentration change synchronicity of the site cluster, optimization is performed by combining regional underlying surface characteristics: first, the regional underlying surface type is divided, then the spatial correlation between the site cluster and the underlying surface is analyzed, and the temporal difference of the concentration peak of each site in the site cluster is calculated through cross-correlation analysis, and the location of the pollution source is determined according to the direction of the difference.
7. The method according to claim 1, characterized in that, In S3, the specific correction process of the terrain-corrected wind field trajectory inversion model is as follows: after inputting the real-time wind direction and speed data, the regional terrain data is first imported, and the influence of terrain on the wind field is calculated based on the atmospheric boundary layer model in fluid mechanics; then, the Lagrange trajectory inversion algorithm is used to simulate the pollutant diffusion path, and the upwind trajectory obtained by inversion is combined with the spatial location of terrain obstacles to make an effectiveness judgment, generate effective terrain and wind direction trajectories, and locate the upwind pollution source area.
8. The method according to claim 1, characterized in that, In S3, the process of generating the component fingerprint difference of the pollution source is as follows: from the pollutant component feature data, feature components are screened based on preset principles, the relative content of the feature components is calculated using a normalization algorithm, and the ratio between feature components is calculated at the same time. The component fingerprint generated by fusing the relative content and the feature ratio is compared with the standard component fingerprints of different types of pollution sources in the preset regional pollutant feature database. The difference between the two is quantified using Euclidean distance, and the matching degree of fingerprint morphology is judged by combining cosine similarity.
9. The method according to claim 1, characterized in that, In S3, the acquired regional industry attribute data includes industry type, production process, raw material consumption type, main products and historical emission records; when establishing the mapping relationship between the pollutant component characteristic data and the regional industry attribute data, an association rule mining algorithm is used: first, the component characteristics and industry attributes are treated as frequent itemsets, and the support and confidence between itemsets are calculated; then, strong association rules with support and confidence reaching the preset standard are selected to form an industry attribute-component characteristic mapping table.
10. The method according to claim 1, characterized in that, In S4, a multi-dimensional spatiotemporal alignment optimization process is performed on the multi-source associated data. Specifically, in the time dimension, for the sampling frequency differences of data from different sources, a combination of linear interpolation and sliding window averaging is used to standardize all data to the same time granularity. At the same time, based on the timestamp deviation record of data collection, time offset correction is performed on each data source. In the spatial dimension, the spatial coordinates of enterprise location, vehicle trajectory, remote sensing fire point and monitoring station are uniformly transformed to the same coordinate system. After spatiotemporal alignment, a three-dimensional index of time-space-data features is constructed.
11. The method according to claim 1, characterized in that, In S4, the retrieved multi-source associated data includes fixed source data, mobile source data, and area source data, employing a classification-adaptive preprocessing mechanism: fixed source data covers emission concentration data monitored by enterprises, electricity load data of treatment facilities, and emission limit parameters approved by environmental protection authorities. During preprocessing, timestamps are aligned, and instantaneous spikes in electricity data are removed; mobile source data includes freight vehicle positioning trajectory data, exhaust gas remote sensing monitoring data, and traffic flow statistics. During preprocessing, trajectory drift is corrected by combining the road network topology, and vehicle identity is bound to emission data through license plate association; area source data includes remote sensing data of straw burning fire points, agricultural fertilization records, and dust monitoring data. During preprocessing, regional attribution labels are generated by matching the latitude and longitude of fire points with administrative divisions.
12. The method according to claim 1, characterized in that, In S4, the multi-dimensional assessment of the emission level of the emission subject constructs a three-layer indicator system of emission intensity, facility operation, and correlation verification: the emission intensity indicator adopts the weighted result of emission per unit output value and the frequency of exceeding the standard, and the weight is based on the preset industrial pollution intensity level; the facility operation indicator includes the commissioning rate of treatment facilities, the compliance rate of electricity load, and the completeness score of maintenance records; the correlation verification indicator includes the time-series correlation coefficient between emission data and fire point data, and the spatial distance deviation between vehicle trajectory and concentration peak. The three-layer indicators are weighted by the analytic hierarchy process and then the comprehensive evaluation value is calculated.