A method for air quality prediction and pollution source tracing based on multi-source data coupling

By simulating pollutant distribution through multi-source data coupling and graph neural network models, combined with LSTM numerical models and path optimization algorithms, the problems of data quality and model coupling in air quality prediction and source tracing are solved, achieving high-precision pollution prediction and source tracing.

CN121075474BActive Publication Date: 2026-01-30南京创蓝科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511624439.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-01-30
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

Existing air quality forecasting and source tracing technologies suffer from data quality defects, weak model coupling, low forecast accuracy, insufficient source tracing accuracy, and difficulty in taking into account the physical mechanisms and long-term temporal dependencies of pollutant transport. This leads to the accumulation of simulation biases, numerous false early warning signals, and inaccurate pollution source location.

Method used

By employing a multi-source data coupling method, data standardization and credibility assessment are performed. A graph neural network model is constructed to simulate pollutant distribution, and an LSTM-numerical model is used for prediction. Pollution concentration thresholds are used to filter out outliers and verify credibility. Finally, a path optimization algorithm is used for pollution source tracing.

Benefits of technology

It improved data quality and model coupling, enhanced the accuracy of pollution prediction and source tracing, reduced false warnings, and enabled precise location of pollution sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075474B_ABST
    Figure CN121075474B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data source tracing technology, and discloses a method and system for air quality prediction and pollution source tracing based on multi-source data coupling. The method includes: standardizing pre-acquired meteorological data, pollution source data, emission source data, and topographic data to generate standardized air quality data; simulating pollutant distribution on the standardized air quality data based on a preset atmospheric chemistry model, and then predicting the pollution concentration of the current monitored atmospheric environment; comparing the obtained pollution prediction results with a preset pollution concentration threshold and verifying them to obtain a final high pollution warning signal; performing refined emission analysis on the pollution prediction results and standardized air quality data to obtain pollution emission data; and using a path optimization algorithm to trace the pollution source on the standardized air quality data to obtain the location of the pollution source. This invention can improve the accuracy of air quality prediction and pollution source tracing based on multi-source data coupling.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data traceability, in particular to an air quality prediction and pollution traceability method and system based on multi-source data coupling. BACKGROUND

[0002] With the acceleration of urbanization, the demand for precise prevention and control of air pollution is increasingly urgent. The existing air quality prediction and traceability technology has the following limitations. First, the data quality is defective. Traditional methods lack reliable evaluation of multi-source heterogeneous data (meteorology, pollution sources, topography, etc.), lack dynamic fusion mechanisms, resulting in low credibility of input data, affecting prediction accuracy, weak model coupling, physical models and machine learning models often run independently, making it difficult to balance the physical mechanism of pollutant transport and long-term time series dependence, leading to simulation bias accumulation, high false alarm rate, abnormal value filtering relying on static threshold, not combining spatial clustering and influence domain analysis, easy to generate false warning signals, and lack of credibility review mechanism, insufficient traceability accuracy, pollution source positioning relying on single concentration gradient analysis, not integrating multi-element over-standard degree and path optimization algorithm, difficult to accurately identify the location of pollution sources in complex environments, therefore, how to improve data quality, enhance model coupling, and improve pollution traceability accuracy has become a problem to be solved. SUMMARY

[0003] The present application provides an air quality prediction and pollution traceability method based on multi-source data coupling, which mainly aims to solve the problem of low accuracy of air quality prediction and pollution traceability based on multi-source data coupling.

[0004] To achieve the above purpose, the present application provides an air quality prediction and pollution traceability method based on multi-source data coupling, which comprises:

[0005] S1, data standardization processing is performed on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data;

[0006] S2, based on a pre-set atmospheric chemical model, simulating the distribution of pollutants based on the air quality standardized data, generating pollutant distribution data containing the content and concentration chemical analysis results of each element;

[0007] S3, based on the pollutant distribution data, predicting the pollution concentration of the current monitoring atmospheric environment, generating a pollution prediction result of the current monitoring atmospheric environment;

[0008] S4, based on a pre-set pollution concentration threshold, filtering abnormal values in the pollution prediction result to obtain an initial high pollution warning signal, and performing credibility review on the initial high pollution warning signal to obtain a final high pollution warning signal;

[0009] S5, performing fine emission analysis on the pollution prediction result and the air quality standardized data to obtain pollution emission data;

[0010] S6, based on the high pollution early warning signal and the pollution emission data, performing pollution tracing on the air quality standardized data by using a path optimization algorithm to obtain a pollution source position of the current monitored atmospheric environment.

[0011] In a preferred embodiment, the data standardization processing of the pre-acquired meteorological data, pollution source data, emission source data and topographic data is to generate air quality standardized data, which comprises:

[0012] constructing an original data set containing pre-acquired meteorological data, pollution source data, emission source data and topographic data;

[0013] performing credibility evaluation on the original data set to obtain a credibility score;

[0014] calculating a dynamic fusion weight based on the credibility score;

[0015] performing weighted fusion on the original data set based on the dynamic fusion weight to obtain air quality standardized data.

[0016] In a preferred embodiment, the credibility evaluation on the original data set to obtain a credibility score comprises:

[0017] based on the original data set, constructing a credibility score model containing power consumption-pollution discharge coupling verification rules, wherein the mathematical formula of the credibility score model is as follows:

[0018]

[0019] wherein, is the initial credibility score of the i-th original data vector in the original data set, is the data acquisition equipment calibration record score of the i-th original data vector in the original data set, is the historical data volatility score of the i-th original data vector in the original data set, is the spatial consistency score of the i-th original data vector in the original data set, , and are respectively the calibration record weight coefficient, the volatility weight coefficient and the spatial consistency weight coefficient;

[0020] ​​​​The initial confidence score is verified using the aforementioned electricity consumption-pollution discharge coupling verification rule to obtain a confidence score. The specific verification rules of the electricity consumption-pollution discharge coupling verification rule are as follows:

[0021] The rate of change in electricity consumption is obtained by calculating the ratio of the current electricity consumption to the historical average electricity consumption.

[0022] The rate of change in pollution discharge is obtained by calculating the ratio of the current discharge volume to the historical average discharge volume.

[0023] When the rate of change in electricity consumption is greater than a preset threshold for change in electricity consumption, and the rate of change in pollution discharge is less than a preset threshold for change in pollution discharge, a downgrade factor is added to the initial credibility score, and a downgrade penalty is imposed on the initial credibility score to obtain a credibility score.

[0024] When the rate of change in electricity consumption is less than or equal to a preset threshold for change in electricity consumption, and the rate of change in sewage discharge is greater than a preset threshold for change in sewage discharge, the initial credibility score is the credibility rating.

[0025] In a preferred embodiment, the step of simulating pollutant distribution on the standardized air quality data based on a preset atmospheric chemical model to generate pollutant distribution data includes:

[0026] Based on the standardized air quality data, a graph neural network with cities as nodes is constructed;

[0027] An atmospheric diffusion equation is embedded in the message passing process of the graph neural network to simulate the generation, transport, and deposition of pollutants, thereby generating the pollutant distribution data. The atmospheric diffusion equation is as follows:

[0028]

[0029] In the formula, For the first Nodes in a layered GNN The characteristic vector of pollutant concentration, For activation function, This is the weight matrix. For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, For nodes arrive exist The average wind speed component in the direction, For nodes arrive exist The average wind speed component in the direction, is a node to a deposition coefficient, is a bias vector, denotes a normalization factor, is a pollutant concentration, denotes a spatial coordinate.

[0030] In a preferred embodiment, the pollution concentration prediction of the current monitoring atmospheric environment based on the pollutant distribution data generates a pollution prediction result of the current monitoring atmospheric environment, comprising:

[0031] initial field construction is performed on the pollutant distribution data;

[0032] the completed initial field is input into an LSTM-numerical model hybrid framework to capture long-term dependence of pollutant concentration change over time to obtain a one-stage prediction result;

[0033] atmospheric physical mechanism numerical analysis is performed on the pollutant distribution data based on the completed initial field to generate a two-stage prediction result;

[0034] adaptive weighted fusion is performed on the one-stage prediction result and a plurality of two-stage prediction results based on a prediction period to obtain a pollution prediction result of the current monitoring atmospheric environment.

[0035] In a preferred embodiment, the pollution prediction result is filtered based on a preset pollution concentration threshold to obtain an initial high pollution early warning signal, comprising:

[0036] dynamic hierarchical filtering is performed on the pollution prediction result based on a preset pollution concentration threshold to obtain a pollution concentration abnormal value label;

[0037] spatial clustering analysis is performed on the pollution concentration abnormal value label to obtain a pollution concentration abnormal area block;

[0038] spatial impact domain evaluation is performed on the pollution concentration abnormal area block to obtain a spatial impact index;

[0039] significance determination is performed on the spatial impact index and a preset pollution concentration threshold to obtain an initial high pollution early warning signal.

[0040] In a preferred embodiment, the initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal, comprising:

[0041] S701: performing credibility review on the initial high-pollution early warning signal by using a credibility review algorithm to obtain a credibility review score, wherein a mathematical expression of the credibility review algorithm is as follows:

[0042]

[0043] is the credibility review score, is a similarity matching score, is a spatial consistency index, and are a similarity weight coefficient and a spatial consistency weight coefficient, respectively;

[0044] S702, based on the credibility review score and a preset credibility threshold, dynamically adjusting an initial pollution level in the initial high-pollution early warning signal to generate a final pollution level;

[0045] S703, based on the credibility review score, performing influence area boundary optimization on an abnormal area position in the initial high-pollution early warning signal to generate final air influence area information;

[0046] S704, based on the final air influence area information, outputting a final high-pollution early warning signal including the final pollution level and the final air influence area information.

[0047] In a preferred embodiment, the fine emission analysis on the pollution prediction result and the air quality standardized data to obtain pollution emission data includes:

[0048] constructing a pollutant concentration distribution matrix containing predicted concentration values of pollution elements at each spatial grid point based on the pollution prediction result;

[0049] constructing a pollution element-threshold mapping table containing standard concentration thresholds of each pollution element based on pollution source type data in the air quality standardized data;

[0050] comparing the pollutant concentration distribution matrix with the standard concentration thresholds in the pollution element-threshold mapping table to obtain a set of pollution concentration exceeding elements;

[0051] calculating the exceeding degree of pollution elements based on the set of pollution concentration exceeding elements and the pollutant concentration distribution matrix by using an exceeding degree calculation formula;

[0052] integrating the set of pollution concentration exceeding elements and the exceeding degree of pollution elements to obtain pollution emission data containing pollution concentration exceeding elements, exceeding degrees, and pollution ranges. ​

[0053] In a preferred embodiment, the over-standard degree calculation formula is as follows:

[0054]

[0055] In the formula, is an over-standard degree index, is a total number of spatial grids, is a spatial grid point identifier, is a dynamic distance weight, is a predicted concentration, is a standard concentration threshold, is a pollution element identifier.

[0056] In a preferred embodiment, the pollution source location of the current monitored atmospheric environment is obtained by using a path optimization algorithm to perform pollution tracing on the air quality standardized data, and the method comprises the following steps:

[0057] Based on the final high-pollution early warning signal, spatial gridding concentration data of the pollution concentration distribution matrix is extracted;

[0058] Based on the spatial gridding concentration data, a concentration change quantitative processing is performed on the pollution concentration over-standard elements, so as to obtain a concentration change direction and intensity quantitative index of the spatial gridding concentration data;

[0059] Based on the concentration change direction and intensity quantitative index, a maximum value is located, so as to obtain a pollution element candidate pollution source coordinate;

[0060] The centroid position of all pollution element candidate pollution source coordinates is calculated, so as to obtain the pollution source location of the current monitored atmospheric environment.

[0061] Compared with the prior art, the present application has the following beneficial effects:

[0062] 1. A multi-source data credibility quantitative evaluation is adopted to construct a credibility score model based on data sources, collection equipment and historical consistency, and to dynamically weight and fuse data, so as to solve the prediction deviation caused by false monitoring data and improve data quality.

[0063] 2. A physical constraint graph neural network transmission model is adopted to embed an atmospheric diffusion equation into a GNN node update mechanism, to constrain the transmission path of pollutants across regions and to enhance the coupling of the model.

[0064] 3. A prediction-tracing two-way error feedback mechanism is adopted to use the tracing result to reversely correct the emission source parameters of the prediction model, to form a closed-loop optimization and to improve the accuracy of pollution tracing. BRIEF DESCRIPTION OF DRAWINGS

[0065] Figure 1A first flowchart of a method for air quality prediction and pollution tracing based on multi-source data coupling is provided in an embodiment of the present application.

[0066] Figure 2 A second flowchart of a method for air quality prediction and pollution tracing based on multi-source data coupling is provided in an embodiment of the present application.

[0067] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0068] It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0069] Embodiments of the present application provide a method for air quality prediction and pollution tracing based on multi-source data coupling. The execution subject of the method for air quality prediction and pollution tracing based on multi-source data coupling includes, but is not limited to, at least one of electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiments of the present application. In other words, the method for air quality prediction and pollution tracing based on multi-source data coupling can be executed by software or hardware installed in a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be a stand-alone server, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms, etc. basic cloud computing services.

[0070] Reference Figure 1 A flowchart of a method for air quality prediction and pollution tracing based on multi-source data coupling is provided in an embodiment of the present application. In this embodiment, the method for air quality prediction and pollution tracing based on multi-source data coupling includes:

[0071] S1, data standardization processing is performed on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data;

[0072] In the embodiment of the present application, the data standardization processing is performed on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data, including:

[0073] An original data set containing pre-acquired meteorological data, pollution source data, emission source data and topographic data is constructed;

[0074] performing credibility evaluation on the original data set to obtain a credibility score;

[0075] calculating a dynamic fusion weight based on the credibility score;

[0076] performing weighted fusion on the original data set based on the dynamic fusion weight to obtain air quality standardized data.

[0077] It should be noted that the meteorological data is a global data set made by using meteorological observation data, which can predict the meteorological data of the next 10 days.

[0078] It should be noted that the credibility evaluation refers to an operation of evaluating the reliability of data by a quantitative model, which means identifying low-quality data and can improve the overall accuracy of the fused data.

[0079] It should be noted that the dynamic fusion weight is obtained by calculating the ratio of the initial credibility score of the i-th original data vector to the sum of the initial credibility scores of all original data vectors.

[0080] Further, the air quality standardized data obtained by the weighted fusion is obtained by calculating the sum of the product of each original data vector and the corresponding dynamic fusion weight.

[0081] In the embodiment of the present application, the credibility evaluation on the original data set to obtain a credibility score comprises:

[0082] Based on the original data set, a credibility score model containing electricity consumption-pollutant discharge coupling verification rules is constructed, wherein the mathematical formula of the credibility score model is as follows:

[0083]

[0084] In the formula, is the initial credibility score of the i-th original data vector in the original data set, is the data acquisition equipment calibration record score of the i-th original data vector in the original data set, is the historical data volatility score of the i-th original data vector in the original data set, is the spatial consistency score of the i-th original data vector in the original data set, , and are calibration record weight coefficient, volatility weight coefficient and spatial consistency weight coefficient, respectively.

[0085] ​​​​​The initial credibility score is checked by using the electricity consumption-pollutant discharge coupling checking rule to obtain a credibility score, wherein the specific checking rule of the electricity consumption-pollutant discharge coupling checking rule is as follows:

[0086] The electricity consumption change rate is obtained by calculating the ratio of the current electricity consumption to the average of the historical electricity consumption.

[0087] The pollutant discharge change rate is obtained by calculating the ratio of the current pollutant discharge to the average of the historical pollutant discharge.

[0088] When the electricity consumption change rate is greater than a preset electricity consumption change threshold, and the pollutant discharge change rate is less than a preset pollutant discharge change threshold, a degradation factor is added to the initial credibility score, and the initial credibility score is degraded and punished to obtain a credibility score.

[0089] When the electricity consumption change rate is less than or equal to a preset electricity consumption change threshold, and the pollutant discharge change rate is greater than a preset pollutant discharge change threshold, the initial credibility score is the credibility score.

[0090] It should be noted that the electricity consumption-pollutant discharge coupling checking refers to the logical judgment of the change of enterprise electricity consumption and pollutant discharge, and detects data anomalies to prevent data falsification and ensure the reliability of the fine emission analysis in step S5.

[0091] The degradation factor is used to reduce the credibility score to punish inconsistent data and ensure data authenticity.

[0092] Further, the calibration record weight coefficient, the volatility weight coefficient and the spatial consistency weight coefficient are all 1 / 3.

[0093] S2, based on a preset atmospheric chemical model, simulating pollutant distribution of the air quality standardized data to generate pollutant distribution data containing chemical analysis results of the content and concentration of each element;

[0094] In the embodiment of the present application, the pollutant distribution data generated by simulating the pollutant distribution of the air quality standardized data based on the preset atmospheric chemical model comprises:

[0095] Based on the air quality standardized data, a graph neural network with cities as nodes is constructed.

[0096] In the message passing process of the graph neural network, an atmospheric diffusion equation is embedded to simulate the generation, transmission and deposition process of pollutants to generate the pollutant distribution data, wherein the atmospheric diffusion equation is as follows:

[0097]

[0098] In the formula, For the first Nodes in a layered GNN The characteristic vector of pollutant concentration, For activation function, This is the weight matrix. For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, For nodes arrive exist The average wind speed component in the direction, For nodes arrive exist The average wind speed component in the direction, For nodes arrive The settlement coefficient, For bias vectors, Represents the normalization factor. For pollutant concentration, Represents spatial coordinates.

[0099] It should be noted that pollutant distribution simulation is a simulation method that simulates the generation, transport, and deposition processes of pollutants;

[0100] The pollutant concentration feature vector is used to quantify the pollution level and represent the current pollution status of urban nodes;

[0101] The settlement coefficient is calculated based on topographic data and represents path-dependent attenuation, which can correct for transmission loss.

[0102] The bias vector represents adjusting the output offset, which can improve the fitting accuracy;

[0103] The normalization factor represents the effect of balancing node degree, which can stabilize the training process.

[0104] Furthermore, the pre-defined atmospheric chemistry model refers to a computational framework built based on physical and chemical principles. It is a mathematical representation that describes the behavior of atmospheric pollutants, provides a basis for simulation, and ensures the physical rationality of predictions. The atmospheric chemistry model includes the WRF model and the CMAQ model.

[0105] Graph neural networks are a deep learning architecture that uses graph structures to process the relationships between nodes and edges. They are used to model the transport of pollutants between cities and can capture spatial dependencies.

[0106] Further, the WRF model is a mesoscale numerical weather prediction system, mainly used for simulating and predicting atmospheric state, and is a meteorological prediction special model for simulating atmospheric physical state by performing operation on coarse resolution data to obtain high resolution data.

[0107] The WRF model has a highly configurable feature, supports multiple nested grids, and can simulate weather of different scales from global to urban blocks, and contains multiple physical process parameterization schemes.

[0108] Further, the CMAQ model is a high simulation model, which divides the ground into grids to obtain the specific amount of pollutants emitted by each grid of the ground.

[0109] Further, the CMAQ model is a high simulation model, which divides the ground into grids to obtain the specific amount of pollutants emitted by each grid of the ground.

[0110] S3, based on the pollutant distribution data, the pollution concentration of the current monitoring atmospheric environment is predicted, and a pollution prediction result of the current monitoring atmospheric environment is generated.

[0111] In the embodiment of the present application, the pollution concentration of the current monitoring atmospheric environment is predicted based on the pollutant distribution data, and the pollution prediction result of the current monitoring atmospheric environment is generated, comprising:

[0112] The initial field construction is performed on the pollutant distribution data;

[0113] The completed initial field is input into the LSTM-numerical model hybrid framework to capture the long-term dependence of the pollutant concentration change with time to obtain a one-stage prediction result;

[0114] Based on the completed initial field, the pollutant distribution data is subjected to atmospheric physical mechanism numerical analysis to generate a two-stage prediction result;

[0115] Based on the prediction period, the one-stage prediction result and a plurality of two-stage prediction results are adaptively weighted and fused to obtain the pollution prediction result of the current monitoring atmospheric environment.

[0116] It should be noted that the long-term dependence of the pollutant concentration change with time is captured by using the LSTM-numerical model, and the operation equation group of the LSTM-numerical model is as follows:

[0117]

[0118] In the formula, is the forgetting gate for controlling the retention degree of historical information, is an activation function, is a weight matrix, is a hidden state storing historical pollution evolution characteristics, is a bias vector, is a pollutant distribution data at time t, is an input gate controlling new information entry, is a temporary storage of new characteristics, is a mapping function, is an output gate controlling prediction result generation, is a hyperbolic tangent function.

[0119] It should be noted that the algorithm used in the numerical analysis of atmospheric physical mechanism is as follows:

[0120]

[0121] wherein, is a pollutant concentration field, is a three-dimensional wind speed field, is a turbulent diffusion coefficient, is a chemical reaction source term, is a dry and wet deposition term, is a gradient operator, is a time dimension of the simulation process.

[0122] Further, the pollutant concentration field is used to describe the spatial distribution of pollutants (such as NO2 or SO2) in the atmosphere;

[0123] The turbulent diffusion coefficient represents the mixing effect caused by atmospheric turbulence, improving spatial accuracy;

[0124] The chemical reaction source term is used to describe the generation / consumption of pollutants, quantifying chemical processes;

[0125] The dry and wet deposition term represents the loss of pollutants due to dry and wet deposition;

[0126] The time dimension of the simulation process is used to capture the dynamic changes of pollutants, ensuring time continuity;

[0127] The three-dimensional wind speed field is used to simulate the carrying effect of wind on pollutants.

[0128] Further, the adaptive weighted fusion is to multiply the LSTM output concentration and the pollutant concentration field by the corresponding weight coefficients and sum them up, wherein the corresponding weight coefficients of the LSTM output concentration and the pollutant concentration field are both 50%.

[0129] S4, filtering abnormal values in the pollution prediction result based on a preset pollution concentration threshold to obtain an initial high pollution early warning signal, and performing credibility review on the initial high pollution early warning signal to obtain a final high pollution early warning signal;

[0130] In the embodiment of the application, filtering abnormal values in the pollution prediction result based on a preset pollution concentration threshold to obtain an initial high pollution early warning signal comprises:

[0131] performing dynamic hierarchical filtering on the pollution prediction result based on a preset pollution concentration threshold to obtain a pollution concentration abnormal value label;

[0132] performing spatial clustering analysis on the pollution concentration abnormal value label to obtain a pollution concentration abnormal area block;

[0133] performing spatial influence domain evaluation on the pollution concentration abnormal area block to obtain a spatial influence index;

[0134] performing significance determination on the spatial influence index and the preset pollution concentration threshold to obtain an initial high pollution early warning signal.

[0135] It should be noted that the initial high pollution early warning signal comprises a pollution concentration abnormal value, an abnormal area position and an initial pollution level.

[0136] It should be noted that the process of dynamic hierarchical filtering is as follows: using a preset pollution concentration threshold, comparing each data point in the pollution prediction result to obtain a pollution comparison result, which is a matrix containing pollution concentration values;

[0137] According to the pollution comparison result, the data points are divided into multiple pollution levels, and if the pollution concentration exceeds the threshold, the abnormal value is marked to generate a pollution concentration abnormal value label.

[0138] It should be noted that the process of spatial clustering analysis is based on the pollution concentration abnormal value label, and adjacent abnormal points in space are aggregated by a spatial clustering algorithm to form a continuous pollution concentration abnormal area block.

[0139] It should be noted that the process of spatial influence domain evaluation is based on the pollution concentration abnormal area block, and the influence degree of each block on the surrounding geographical range is evaluated to calculate the spatial influence index to quantify the influence.

[0140] Further, the spatial influence domain evaluation considers the pollution intensity, spatial range and diffusion potential of the block, and outputs the spatial influence index as a quantitative index.

[0141] It should be noted that the process of significance determination is to compare the spatial influence index with the preset pollution concentration threshold value to determine whether the index is significantly higher than the threshold value, and if the index exceeds the threshold value, it is determined as a significantly high pollution event, thereby generating an initial high pollution early warning signal.

[0142] In the embodiment of the application, the initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal, which comprises:

[0143] S701: The initial high pollution early warning signal is subjected to credibility review by using a credibility review algorithm to obtain a credibility review score, wherein the mathematical expression of the credibility review algorithm is as follows:

[0144]

[0145] In the formula, is the credibility review score, is a similarity matching score, is a spatial consistency index, and are a similarity weight coefficient and a spatial consistency weight coefficient, respectively;

[0146] S702, based on the credibility review score and a preset credibility threshold value, the initial pollution level in the initial high pollution early warning signal is dynamically adjusted to generate a final pollution level;

[0147] S703, based on the credibility review score, the abnormal area position in the initial high pollution early warning signal is subjected to influence area boundary optimization to generate final air influence area information;

[0148] S704, based on the final air influence area information, a final high pollution early warning signal including the final pollution level and the final air influence area information is output.

[0149] It should be noted that the similarity matching score is based on a preset historical pollution event database to perform spatio-temporal similarity matching on the initial high pollution early warning signal to generate a similarity matching score.

[0150] The spatial consistency index is based on the abnormal area position in the initial high pollution early warning signal to perform spatial consistency analysis on the real-time acquired pollution concentration data to obtain a spatial consistency index.

[0151] It should be noted that the process of dynamically adjusting the initial pollution level in the initial high pollution early warning signal is based on the comparison result of the credibility review score and the preset credibility threshold, and the initial pollution level in the initial high pollution early warning signal is dynamically corrected. If the credibility review score is lower than the threshold, the initial pollution level is reduced;

[0152] If the score is higher than the threshold, the initial level is improved or maintained;

[0153] Through the elastic adjustment of the score driven level, the reliability of the final pollution level is ensured to output the final pollution level.

[0154] It should be noted that the process of optimizing the boundary of the affected area is to use the credibility review score to finely correct the spatial boundary of the abnormal area position in the initial high pollution early warning signal. If the score is lower, the abnormal area boundary is contracted to exclude the low confidence area;

[0155] If the score is higher, the boundary is expanded to cover the potential impact range;

[0156] Finally, the accurate final air impact area information is generated to ensure that the early warning range matches the actual risk of pollution diffusion.

[0157] S5, the pollution prediction result and the air quality standardization data are analyzed in detail to obtain pollution emission data;

[0158] In the embodiment of the application, the pollution prediction result and the air quality standardization data are analyzed in detail to obtain pollution emission data, which comprises:

[0159] Based on the pollution prediction result, a pollution concentration distribution matrix containing the predicted concentration value of each pollution element at each spatial grid point is constructed;

[0160] Based on the pollution source type data in the air quality standardization data, a pollution element-threshold mapping table containing the standard concentration threshold of each pollution element is constructed;

[0161] The pollution concentration distribution matrix and the standard concentration threshold in the pollution element-threshold mapping table are compared to obtain a set of pollution concentration exceeding elements;

[0162] Based on the set of pollution concentration exceeding elements and the pollution concentration distribution matrix, the exceeding degree of the pollution element is calculated by using an exceeding degree calculation formula;

[0163] The set of pollution concentration exceeding elements and the exceeding degree of the pollution element are integrated to obtain pollution emission data containing pollution concentration exceeding elements, exceeding degrees and pollution ranges.

[0164] It should be noted that the pollutant concentration distribution matrix refers to the predicted concentration value of structured storage, the matrix row represents the pollution element, and the list represents the spatial grid point, which is convenient for algorithm processing and improves the calculation efficiency.

[0165] The pollution element-threshold mapping table refers to a preset pollution threshold table, which defines the safety limit of each pollutant and provides a standard for exceeding the standard to ensure objectivity.

[0166] In the embodiment of the application, the exceeding degree calculation formula is as follows:

[0167]

[0168] In the formula, is the exceeding degree index, is the total number of spatial grids, is the spatial grid point identifier, is the dynamic distance weight, is the predicted concentration, is the standard concentration threshold, is the pollution element identifier.

[0169] It should be noted that the exceeding degree index represents the overall exceeding severity of the pollution element, the greater the value, the more serious the pollution, the spatial grid point identifier represents the position number after regional division, the total number of spatial grids represents the resolution of the analysis area, the greater the value, the more detailed the analysis, the dynamic distance weight adjusts the contribution weight based on the distance of the pollution source, the closer to the pollution source, the higher the weight, and the predicted concentration represents the predicted concentration value of the pollution element at the grid point.

[0170] S6, based on the high pollution early warning signal and the pollution emission data, using a path optimization algorithm to perform pollution tracing on the air quality standardized data to obtain the pollution source position of the current monitored atmospheric environment.

[0171] In the embodiment of the application, the pollution tracing on the air quality standardized data using a path optimization algorithm to obtain the pollution source position of the current monitored atmospheric environment comprises:

[0172] Based on the final high pollution early warning signal, spatial gridding concentration data extraction is performed on the pollution concentration distribution matrix;

[0173] Based on the spatial gridding concentration data, concentration change quantization processing is performed on the pollution concentration exceeding element to obtain concentration change direction and intensity quantization indexes of the spatial gridding concentration data;

[0174] Based on the concentration change direction and intensity quantization indexes, maximum value positioning is performed to obtain pollution element candidate pollution source coordinates;

[0175] The centroid position of all pollution element candidate pollution source coordinates is calculated to obtain the pollution source position of the current monitored atmospheric environment.

[0176] It should be noted that the concentration change quantification processing is based on the spatial gridding concentration data and the pollution concentration exceeding element, and the spatial concentration change of each exceeding element is quantitatively analyzed, including: calculating the concentration gradient of each pollution element on the spatial grid point to determine the concentration change direction and change strength, and through the quantification processing, the structured concentration change direction and strength quantification index are generated.

[0177] Compared with the prior art, the present application has the following beneficial effects:

[0178] 1. Multi-source data credibility quantification evaluation is adopted to construct a credibility score model based on data sources, collection equipment and historical consistency, and the data is dynamically weighted and fused to solve the prediction deviation caused by false monitoring data and improve the data quality.

[0179] 2. The physical constraint graph neural network transmission model is adopted to embed the atmospheric diffusion equation into the GNN node update mechanism, constrain the pollution cross-region transmission path, and enhance the model coupling.

[0180] 3. The prediction-tracing two-way error feedback mechanism is adopted to use the tracing result to correct the emission source parameters of the prediction model in reverse, form a closed loop optimization, and improve the pollution tracing accuracy.

[0181] In several embodiments provided by the present application, it should be understood that the disclosed method can be implemented by other ways.

[0182] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.

[0183] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limited. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. An air quality prediction and pollution tracing method based on multi-source data coupling, characterized in that, The method comprises: S1, data standardization processing is performed on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data; The data standardization processing on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data comprises: An original data set containing the pre-acquired meteorological data, pollution source data, emission source data and topographic data is constructed; The original data set is subjected to credibility evaluation to obtain a credibility score; A dynamic fusion weight is calculated based on the credibility score; The original data set is subjected to weighted fusion based on the dynamic fusion weight to obtain air quality standardized data; S2, based on a pre-set atmospheric chemical model, the air quality standardized data is subjected to pollution distribution simulation to generate pollution distribution data containing chemical analysis results of the content and concentration of each element; The credibility evaluation on the original data set to obtain a credibility score comprises: Based on the original data set, a credibility score model containing a power consumption-emission coupling verification rule is constructed, wherein the mathematical formula adopted by the credibility score model is as follows: wherein, is an initial trust score for an original data vector in the original data set, is an initial trust score for an original data vector in the original data set, is a data acquisition device calibration record score for an original data vector in the original data set, is a data acquisition device calibration record score for an original data vector in the original data set, is a historical data volatility score for an original data vector in the original data set, is a historical data volatility score for an original data vector in the original data set, is a spatial consistency score for an original data vector in the original data set, is a spatial consistency score for an original data vector in the original data set, , and are calibration record weight coefficient, volatility weight coefficient and spatial consistency weight coefficient, respectively. The initial credibility score is verified by using the power consumption-emission coupling verification rule to obtain a credibility score, wherein the specific verification rule of the power consumption-emission coupling verification rule is as follows: A power consumption change rate is obtained by calculating the ratio of the current power consumption to the historical power consumption average; An emission change rate is obtained by calculating the ratio of the current emission to the historical emission average; When the power consumption change rate is greater than a pre-set power consumption change threshold and the emission change rate is less than a pre-set emission change threshold, a degradation factor is added to the initial credibility score, and the initial credibility score is degraded and punished to obtain a credibility score; When the power consumption change rate is less than or equal to a pre-set power consumption change threshold and the emission change rate is greater than a pre-set emission change threshold, the initial credibility score is the credibility score; S3, based on the pollution distribution data, a pollution concentration prediction of the current monitored atmospheric environment is performed to generate a pollution prediction result of the current monitored atmospheric environment; S4, based on a pre-set pollution concentration threshold, abnormal values in the pollution prediction result are filtered to obtain an initial high pollution early warning signal, and the initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal; S5, the pollution prediction result and the air quality standardized data are subjected to fine emission analysis to obtain pollution emission data; S6, based on the final high pollution early warning signal and the pollution emission data, a path optimization algorithm is used to perform pollution tracing on the air quality standardized data to obtain a pollution source position of the current monitored atmospheric environment.

2. The method of claim 1, wherein, The pollution distribution simulation on the air quality standardized data based on the pre-set atmospheric chemical model to generate pollution distribution data comprises: Based on the air quality standardized data, a graph neural network with cities as nodes is constructed; In the message passing process of the graph neural network, an atmospheric diffusion equation is embedded to simulate the generation, transmission and deposition process of pollutants, and the pollutant distribution data is generated, wherein the atmospheric diffusion equation is as follows: In the formula, For the first Nodes in a layered GNN The characteristic vector of pollutant concentration, For activation function, This is the weight matrix. For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, For nodes arrive exist The average wind speed component in the direction, For nodes arrive exist The average wind speed component in the direction, For nodes arrive The settlement coefficient, For bias vectors, Represents the normalization factor. For pollutant concentration, Represents spatial coordinates.

3. The method of claim 1, wherein, The pollution concentration of the current monitoring atmospheric environment is predicted based on the pollutant distribution data, and a pollution prediction result of the current monitoring atmospheric environment is generated, which includes: An initial field is constructed based on the pollutant distribution data; The constructed initial field is input into an LSTM-numeric model hybrid framework to capture the long-term dependence of the pollutant concentration over time to obtain a one-stage prediction result; Based on the constructed initial field, the pollutant distribution data is subjected to numerical analysis of atmospheric physical mechanisms to generate a two-stage prediction result; Based on the prediction period, the one-stage prediction result and a plurality of two-stage prediction results are adaptively weighted and fused to obtain the pollution prediction result of the current monitoring atmospheric environment.

4. The method of claim 3, wherein, Based on the preset pollution concentration threshold, the pollution prediction result is filtered to obtain an initial high pollution early warning signal, which includes: Based on the preset pollution concentration threshold, the pollution prediction result is dynamically graded and filtered to obtain a pollution concentration abnormal value label; The pollution concentration abnormal value label is subjected to spatial clustering analysis to obtain a pollution concentration abnormal area block; The pollution concentration abnormal area block is subjected to spatial influence domain evaluation to obtain a spatial influence index; The spatial influence index and the preset pollution concentration threshold are subjected to significance determination to obtain an initial high pollution early warning signal.

5. The method of claim 1, wherein, The initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal, which includes: S501: A credibility review algorithm is used to review the credibility of the initial high pollution early warning signal to obtain a credibility review score, wherein the mathematical expression of the credibility review algorithm is as follows: wherein is a plausibility check score, is a similarity match score, is a spatial consistency index, and are a similarity weight coefficient and a spatial consistency weight coefficient, respectively. S502, based on the credibility review score and the preset credibility threshold, the initial pollution level in the initial high pollution early warning signal is dynamically adjusted to generate a final pollution level; S503, based on the credibility review score, the abnormal area position in the initial high pollution early warning signal is subjected to influence area boundary optimization to generate final air influence area information; S504, based on the final air influence area information, the final high pollution early warning signal including the final pollution level and the final air influence area information is output.

6. The method of claim 1, wherein, The pollution prediction result and the air quality standardized data are subjected to fine emission analysis to obtain pollution emission data, which includes: A pollutant concentration distribution matrix containing the predicted concentration values of pollution elements at each spatial grid point is constructed based on the pollution prediction result; Based on the pollution source type data in the air quality standardized data, a pollution element-threshold mapping table containing the standard concentration threshold of each pollution element is constructed; The pollutant concentration distribution matrix is compared with the standard concentration threshold in the pollution element-threshold mapping table to obtain a pollution concentration exceeding element set; Based on the exceeding standard element set and the pollutant concentration distribution matrix, an exceeding standard degree calculation formula is used to calculate the exceeding standard degree of the pollution elements; The exceeding standard element set and the exceeding standard degree of the pollution elements are integrated for pollution emission data to obtain pollution emission data containing pollution concentration exceeding standard elements, exceeding standard degrees and pollution ranges.

7. The method of claim 6, wherein, The exceeding standard degree calculation formula is as follows: wherein, is an over-standard degree index, is a total number of spatial grids, is a spatial grid point identification, is a dynamic distance weight, is a predicted concentration, is a standard concentration threshold value, is a pollution element identification.

8. The method of claim 7, wherein, The pollution source position of the current monitored atmospheric environment is obtained by using a path optimization algorithm to perform pollution tracing on the air quality standardized data, including: Based on the final high-pollution early warning signal, spatial gridding concentration data of the pollution concentration distribution matrix is extracted; Based on the spatial gridding concentration data, a concentration change quantification processing is performed on the exceeding standard element set to obtain a concentration change direction and intensity quantification index of the spatial gridding concentration data; Based on the concentration change direction and intensity quantification index, a maximum value is located to obtain pollution element candidate pollution source coordinates; The centroid position of all pollution element candidate pollution source coordinates is calculated to obtain the pollution source position of the current monitored atmospheric environment.

Citation Information

Patent Citations

  • Atmospheric pollution monitoring method based on unmanned aerial vehicle remote sensing and machine learning

    CN120446400A

  • Traceability analysis method and system for atmospheric pollutants

    CN120895127A