Air quality prediction and pollution tracing method based on multi-source data coupling

By standardizing and assessing the reliability of multi-source data, and combining graph neural networks and LSTM models for air quality prediction and pollution source tracing, the problem of insufficient reliability assessment of multi-source data is solved, and high-precision air quality prediction and pollution source tracing are achieved.

CN121075474AActive Publication Date: 2025-12-05南京创蓝科技有限公司 +1

Patent Information

Application Number
CN202511624439.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2025-12-05
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

In existing air quality prediction and source tracing technologies, the reliability assessment of multi-source heterogeneous data is insufficient, the model coupling is weak, and it is difficult to take into account both the physical mechanism of pollutant transport and the long-term time-series dependence, resulting in low prediction accuracy and insufficient source tracing accuracy.

Method used

By standardizing meteorological, pollution source, and topographic data, a credibility scoring model is constructed. A graph neural network is used to simulate pollutant distribution, and an LSTM-numerical model is combined to predict pollution concentration. A path optimization algorithm is used to trace the source, outliers are dynamically filtered, and credibility is verified to ultimately locate the pollution source.

Benefits of technology

It improved data quality and model coupling, enhanced the accuracy of air quality prediction and pollution source tracing, formed a closed-loop optimization, and reduced the false alarm rate of early warnings and source tracing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075474A_ABST
    Figure CN121075474A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data traceability, and discloses an air quality prediction and pollution traceability method and system based on multi-source data coupling, and the method comprises the steps: carrying out the data standardization processing of pre-obtained meteorological data, pollution source data, emission source data and landform data, and generating air quality standardization data, pollutant distribution simulation is carried out on the air quality standardized data based on a preset atmospheric chemical model, then pollution concentration prediction is carried out on the currently monitored atmospheric environment, an obtained pollution prediction result is compared with a preset pollution concentration threshold value, rechecking is carried out, and a final high-pollution early warning signal is obtained. And carrying out refined emission analysis on the pollution prediction result and the air quality standardized data to obtain pollution emission data, and carrying out pollution tracing on the air quality standardized data by adopting a path optimization algorithm to obtain a pollution source position. According to the invention, the accuracy of air quality prediction and pollution tracing based on multi-source data coupling can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data traceability, in particular to an air quality prediction and pollution traceability method and system based on multi-source data coupling. BACKGROUND

[0002] With the acceleration of urbanization, the demand for precise prevention and control of air pollution is increasingly urgent. The existing air quality prediction and traceability technology has the following limitations. First, the data quality is defective. Traditional methods lack sufficient reliability evaluation of multi-source heterogeneous data (meteorology, pollution sources, topography, etc.), lack dynamic fusion mechanisms, result in low credibility of input data, affect prediction accuracy, weak model coupling, physical models and machine learning models often run independently, it is difficult to balance the physical mechanism of pollutant transport and long-term time series dependence, resulting in simulation deviation accumulation, high false alarm rate, abnormal value filtering relying on static threshold, not combined with spatial clustering and influence domain analysis, easy to generate false warning signals, and lack of credibility review mechanism, insufficient traceability accuracy, pollution source positioning relies on single concentration gradient analysis, not combined with multi-element over-standard degree and path optimization algorithm, difficult to accurately identify the location of pollution sources in complex environment, therefore, how to improve data quality, enhance model coupling, and improve pollution traceability accuracy has become a problem to be solved. SUMMARY

[0003] The present application provides an air quality prediction and pollution traceability method based on multi-source data coupling, which mainly aims to solve the problem of low accuracy of air quality prediction and pollution traceability based on multi-source data coupling.

[0004] To achieve the above purpose, the present application provides an air quality prediction and pollution traceability method based on multi-source data coupling, which comprises: S1, data standardization processing is performed on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data; S2, based on a pre-set atmospheric chemical model, pollutant distribution simulation is performed on the air quality standardized data to generate pollutant distribution data containing chemical analysis results of the content and concentration of each element; S3, based on the pollutant distribution data, pollution concentration prediction is performed on the current monitoring atmospheric environment to generate pollution prediction results of the current monitoring atmospheric environment; S4, based on a pre-set pollution concentration threshold, abnormal values in the pollution prediction results are filtered to obtain initial high pollution warning signals, and credibility review is performed on the initial high pollution warning signals to obtain final high pollution warning signals; S5, fine emission analysis is performed on the pollution prediction results and the air quality standardized data to obtain pollution emission data; S6. Based on the high pollution warning signal and the pollution emission data, a path optimization algorithm is used to trace the pollution source of the standardized air quality data to obtain the location of the pollution source in the current monitored atmospheric environment.

[0005] In a preferred embodiment, the step of standardizing the pre-acquired meteorological data, pollution source data, emission source data, and topographic data to generate standardized air quality data includes: Construct a raw dataset containing pre-acquired meteorological data, pollution source data, emission source data, and topographic data; The original dataset is subjected to a credibility assessment to obtain a credibility score; Calculate dynamic fusion weights based on the aforementioned credibility score; The original dataset is weighted and fused based on the dynamic fusion weights to obtain standardized air quality data.

[0006] In a preferred embodiment, the step of performing a credibility assessment on the original dataset to obtain a credibility score includes: Based on the original dataset, a credibility scoring model is constructed that includes coupled verification rules for electricity consumption and pollution discharge. The mathematical formula used in the credibility scoring model is as follows:

[0007] In the formula, For the first in the original dataset The initial confidence score of each original data vector. For the first in the original dataset Scoring of data acquisition device calibration records for each raw data vector. For the first in the original dataset Historical data volatility score for each original data vector For the first in the original dataset Spatial consistency score of the original data vector , and These are the calibration record weighting coefficient, volatility weighting coefficient, and spatial consistency weighting coefficient, respectively. The initial confidence score is verified using the aforementioned electricity consumption-pollution discharge coupling verification rule to obtain a confidence score. The specific verification rules of the electricity consumption-pollution discharge coupling verification rule are as follows: The rate of change in electricity consumption is obtained by calculating the ratio of the current electricity consumption to the historical average electricity consumption. The rate of change in pollution discharge is obtained by calculating the ratio of the current discharge volume to the historical average discharge volume. When the rate of change in electricity consumption is greater than a preset threshold for change in electricity consumption, and the rate of change in pollution discharge is less than a preset threshold for change in pollution discharge, a downgrade factor is added to the initial credibility score, and a downgrade penalty is imposed on the initial credibility score to obtain a credibility score. When the rate of change in electricity consumption is less than or equal to a preset threshold for change in electricity consumption, and the rate of change in sewage discharge is greater than a preset threshold for change in sewage discharge, the initial credibility score is the credibility rating.

[0008] In a preferred embodiment, the step of simulating pollutant distribution on the standardized air quality data based on a preset atmospheric chemical model to generate pollutant distribution data includes: Based on the standardized air quality data, a graph neural network with cities as nodes is constructed; An atmospheric diffusion equation is embedded in the message passing process of the graph neural network to simulate the generation, transport, and deposition of pollutants, thereby generating the pollutant distribution data. The atmospheric diffusion equation is as follows:

[0009] In the formula, For the first Nodes in a layered GNN The characteristic vector of pollutant concentration, For activation function, This is the weight matrix. For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, For nodes arrive exist The average wind speed component in the direction, For nodes arrive exist The average wind speed component in the direction, For nodes arrive The settlement coefficient, For bias vectors, Represents the normalization factor. For pollutant concentration, Represents spatial coordinates.

[0010] In a preferred embodiment, the step of predicting the pollution concentration of the current monitored atmospheric environment based on the pollutant distribution data and generating the pollution prediction result of the current monitored atmospheric environment includes: An initial field is constructed based on the pollutant distribution data; The completed initial field input LSTM-numeric model hybrid framework is used to capture the long-term dependence of the pollutant concentration over time to obtain a one-stage prediction result; Based on the completed initial field, the pollutant distribution data is subjected to atmospheric physical mechanism numerical analysis to generate a two-stage prediction result; Based on the prediction period, the one-stage prediction result and most two-stage prediction results are adaptively weighted and fused to obtain the pollution prediction result of the current monitoring atmospheric environment.

[0011] In a preferred embodiment, the pollution prediction result is filtered based on a preset pollution concentration threshold to obtain an initial high pollution early warning signal, comprising: Based on a preset pollution concentration threshold, the pollution prediction result is dynamically graded and filtered to obtain a pollution concentration abnormal value label; The pollution concentration abnormal value label is subjected to spatial clustering analysis to obtain a pollution concentration abnormal area block; The pollution concentration abnormal area block is subjected to spatial influence domain evaluation to obtain a spatial influence index; The spatial influence index and the preset pollution concentration threshold are subjected to significance determination to obtain an initial high pollution early warning signal.

[0012] In a preferred embodiment, the initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal, comprising: S701: A credibility review algorithm is used to review the credibility of the initial high pollution early warning signal to obtain a credibility review score, wherein the mathematical expression of the credibility review algorithm is as follows:

[0013] In the formula, is the credibility review score, is the similarity matching score, is the spatial consistency index, and are the similarity weight coefficient and the spatial consistency weight coefficient, respectively; S702, based on the credibility review score and a preset credibility threshold, the initial pollution level in the initial high pollution early warning signal is dynamically adjusted to generate a final pollution level; S703, based on the credibility review score, the abnormal area position in the initial high pollution early warning signal is subjected to influence area boundary optimization to generate final air influence area information; S704, output a final high pollution early warning signal including the final pollution level and the final air influence area information based on the final air influence area information.

[0014] In a preferred embodiment, the fine emission analysis on the pollution prediction result and the air quality standardized data to obtain pollution emission data comprises: constructing a pollution concentration distribution matrix containing predicted concentration values of pollution elements at each spatial grid point based on the pollution prediction result; constructing a pollution element-threshold mapping table containing standard concentration thresholds of each pollution element based on pollution source type data in the air quality standardized data; comparing the pollution concentration distribution matrix with the standard concentration thresholds in the pollution element-threshold mapping table to obtain a set of pollution concentration exceeding elements; calculating the exceeding degree of pollution elements using an exceeding degree calculation formula based on the set of pollution concentration exceeding elements and the pollution concentration distribution matrix; integrating the set of pollution concentration exceeding elements and the exceeding degree of pollution elements to obtain pollution emission data containing pollution concentration exceeding elements, exceeding degrees, and pollution ranges.

[0015] In a preferred embodiment, the exceeding degree calculation formula is as follows:

[0016] In the formula, is the exceeding degree index, is the total number of spatial grids, is the spatial grid point identifier, is the dynamic distance weight, is the predicted concentration, is the standard concentration threshold, is the pollution element identifier.

[0017] In a preferred embodiment, the pollution tracing on the air quality standardized data using a path optimization algorithm to obtain the pollution source location of the current monitored atmospheric environment comprises: spatial gridding concentration data extraction on the pollution concentration distribution matrix based on the final high pollution early warning signal; concentration change quantification processing on the pollution concentration exceeding elements based on the spatial gridding concentration data to obtain concentration change direction and intensity quantification indexes of the spatial gridding concentration data; maximum value positioning based on the concentration change direction and intensity quantification indexes to obtain pollution element candidate pollution source coordinates; The centroid position of all pollution element candidate pollution source coordinates is calculated to obtain the pollution source position of the current monitored atmospheric environment.

[0018] Compared with the prior art, the present application has the following beneficial effects: 1. A multi-source data credibility quantitative evaluation is used to construct a credibility score model based on data sources, collection equipment, and historical consistency, dynamically weighted data fusion, solve the prediction deviation caused by monitoring data fraud, and improve data quality.

[0019] 2. A physically constrained graph neural network transmission model is used to embed the atmospheric diffusion equation into the GNN node update mechanism, constrain the pollution cross-region transmission path, and enhance the model coupling.

[0020] 3. A prediction-tracing two-way error feedback mechanism is used to use the tracing result to correct the emission source parameters of the prediction model in reverse, form a closed loop optimization, and improve the pollution tracing accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A first flowchart of a multi-source data coupled air quality prediction and pollution tracing method according to an embodiment of the present application is provided. Figure 2 A second flowchart of a multi-source data coupled air quality prediction and pollution tracing method according to an embodiment of the present application is provided. The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0022] It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0023] Embodiments of the present application provide a multi-source data coupled air quality prediction and pollution tracing method. The execution subject of the multi-source data coupled air quality prediction and pollution tracing method includes but is not limited to at least one of electronic devices capable of being configured to execute the method provided by the embodiments of the present application, such as a server and a terminal. In other words, the multi-source data coupled air quality prediction and pollution tracing method can be executed by software or hardware installed in a terminal device or a server device. The server includes but is not limited to a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be a stand-alone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms, etc. basic cloud computing services.

[0024] Referring to Figure 1 Fig. 1 shows a flowchart of a method for air quality prediction and pollution tracing based on multi-source data coupling according to an embodiment of the present application. In this embodiment, the method for air quality prediction and pollution tracing based on multi-source data coupling comprises the following steps: S1, performing data standardization processing on pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data; In this embodiment, the data standardization processing on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data comprises: constructing an original data set containing the pre-acquired meteorological data, pollution source data, emission source data and topographic data; performing credibility evaluation on the original data set to obtain a credibility score; calculating a dynamic fusion weight based on the credibility score; performing weighted fusion on the original data set based on the dynamic fusion weight to obtain air quality standardized data.

[0025] It should be noted that the meteorological data is a global data set made of meteorological observation data, which can predict meteorological data for the next 10 days.

[0026] It should be noted that credibility evaluation refers to an operation of evaluating data reliability through a quantitative model, which means identifying low-quality data and improving the overall accuracy of fused data.

[0027] It should be noted that the dynamic fusion weight is obtained by calculating the ratio of the initial credibility score of the i-th original data vector to the sum of the initial credibility scores of all original data vectors.

[0028] Further, the air quality standardized data obtained by weighted fusion is obtained by calculating the sum of the product of each original data vector and the corresponding dynamic fusion weight.

[0029] In this embodiment, the credibility evaluation on the original data set to obtain a credibility score comprises: based on the original data set, constructing a credibility score model containing power consumption-emission coupling verification rules, wherein the mathematical formula of the credibility score model is as follows:

[0030] In the formula, i is the index of the original data vector, and is the initial credibility score of the i-th original data vector. ​The initial confidence score of each original data vector. For the first in the original dataset Scoring of data acquisition device calibration records for each raw data vector. For the first in the original dataset Historical data volatility score for each original data vector For the first in the original dataset Spatial consistency score of the original data vector , and These are the calibration record weighting coefficient, volatility weighting coefficient, and spatial consistency weighting coefficient, respectively. The initial confidence score is verified using the aforementioned electricity consumption-pollution discharge coupling verification rule to obtain a confidence score. The specific verification rules of the electricity consumption-pollution discharge coupling verification rule are as follows: The rate of change in electricity consumption is obtained by calculating the ratio of the current electricity consumption to the historical average electricity consumption. The rate of change in pollution discharge is obtained by calculating the ratio of the current discharge volume to the historical average discharge volume. When the rate of change in electricity consumption is greater than a preset threshold for change in electricity consumption, and the rate of change in pollution discharge is less than a preset threshold for change in pollution discharge, a downgrade factor is added to the initial credibility score, and a downgrade penalty is imposed on the initial credibility score to obtain a credibility score. When the rate of change in electricity consumption is less than or equal to a preset threshold for change in electricity consumption, and the rate of change in sewage discharge is greater than a preset threshold for change in sewage discharge, the initial credibility score is the credibility rating.

[0031] It should be noted that the electricity consumption-pollution discharge coupling verification refers to the logical judgment of comparing changes in enterprise electricity consumption and pollution discharge, detecting data anomalies, and ensuring the reliability of the refined emission analysis in step S5 to prevent data fraud. Degradation factors are used to lower credibility scores to penalize inconsistent data and ensure data authenticity.

[0032] Furthermore, the calibration record weighting coefficient, volatility weighting coefficient, and spatial consistency weighting coefficient are all 1 / 3.

[0033] S2, based on a preset atmospheric chemistry model, simulate the distribution of pollutants on the standardized air quality data to generate pollutant distribution data that includes the content and concentration of each element in the chemical analysis. In this embodiment of the invention, the step of simulating pollutant distribution on the standardized air quality data based on a preset atmospheric chemical model to generate pollutant distribution data includes: Based on the standardized air quality data, a graph neural network with cities as nodes is constructed; An atmospheric diffusion equation is embedded in the message passing process of the graph neural network to simulate the generation, transport, and deposition of pollutants, thereby generating the pollutant distribution data. The atmospheric diffusion equation is as follows:

[0034] In the formula, For the first Nodes in a layered GNN The characteristic vector of pollutant concentration, For activation function, This is the weight matrix. For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, For nodes arrive exist The average wind speed component in the direction, For nodes arrive exist The average wind speed component in the direction, For nodes arrive The settlement coefficient, For bias vectors, Represents the normalization factor. For pollutant concentration, Represents spatial coordinates.

[0035] It should be noted that pollutant distribution simulation is a simulation method that simulates the generation, transport, and deposition processes of pollutants; The pollutant concentration feature vector is used to quantify the pollution level and represent the current pollution status of urban nodes; The settlement coefficient is calculated based on topographic data and represents path-dependent attenuation, which can correct for transmission loss. The bias vector represents adjusting the output offset, which can improve the fitting accuracy; The normalization factor represents the effect of balancing node degree, which can stabilize the training process.

[0036] Furthermore, the pre-defined atmospheric chemistry model refers to a computational framework built based on physical and chemical principles. It is a mathematical representation that describes the behavior of atmospheric pollutants, provides a basis for simulation, and ensures the physical rationality of predictions. The atmospheric chemistry model includes the WRF model and the CMAQ model. Graph neural networks are a deep learning architecture that uses graph structures to process the relationships between nodes and edges. They are used to model the transport of pollutants between cities and can capture spatial dependencies.

[0037] Further, the WRF model is a mesoscale numerical weather prediction system, mainly used for simulating and predicting atmospheric state, obtaining high-resolution data through operation of coarse-resolution data, and simulating meteorological prediction special model of atmospheric physical state; The WRF model has the characteristics of high configurability, supports multiple nested grids, can simulate weather of different scales from global to urban blocks, and contains multiple physical process parameterization schemes.

[0038] Further, the CMAQ model is a high-resolution data, ground air quality observation data and ground emission inventory input into the model, chemical analysis is carried out to obtain the content and concentration of each element in the atmosphere, and the content and concentration of each element in the atmosphere are integrated to output prediction data for the next 10 days. Further, the CMAQ model is a high simulation model, and the ground is divided into grids to obtain the specific amount of pollutants emitted by each grid of the ground.

[0039] S3, based on the pollutant distribution data, the pollution concentration of the current monitoring atmospheric environment is predicted, and the pollution prediction result of the current monitoring atmospheric environment is generated; In the embodiment of the present application, the pollution concentration of the current monitoring atmospheric environment is predicted based on the pollutant distribution data, and the pollution prediction result of the current monitoring atmospheric environment is generated, which comprises: The initial field construction is performed on the pollutant distribution data; The completed initial field is input into the LSTM-numerical model hybrid framework to capture the long-term dependence of the pollutant concentration change with time to obtain a one-stage prediction result; Based on the completed initial field, the pollutant distribution data is subjected to atmospheric physical mechanism numerical analysis to generate a two-stage prediction result; Based on the prediction period, the one-stage prediction result and a plurality of two-stage prediction results are adaptively weighted and fused to obtain the pollution prediction result of the current monitoring atmospheric environment.

[0040] It should be noted that the long-term dependence of the pollutant concentration change with time is captured by using the LSTM-numerical model, and the operation equation group of the LSTM-numerical model is as follows:

[0041] In the formula, is a forgetting gate for controlling the retention degree of historical information, is an activation function, is a weight matrix, is a hidden state for storing historical pollution evolution characteristics, is a bias vector, is the pollutant distribution data at time t, is the input gate for controlling new information entry, is the temporary storage of new features, is the mapping function, is the output gate for controlling the generation of prediction results, is the hyperbolic tangent function.

[0042] It should be noted that the algorithm used in the numerical analysis of atmospheric physical mechanism is as follows:

[0043] In the formula, is the pollutant concentration field, is the three-dimensional wind speed field, is the turbulent diffusion coefficient, is the chemical reaction source term, is the dry and wet deposition term, is the gradient operator, is the time dimension of the simulation process.

[0044] Further, the pollutant concentration field is used to describe the spatial distribution of pollutants (such as NO2 or SO2) in the atmosphere; The turbulent diffusion coefficient represents the mixing effect caused by atmospheric turbulence, which improves the spatial accuracy; The chemical reaction source term is used to describe the generation / consumption of pollutants and quantify chemical processes; The dry and wet deposition term represents the loss of pollutant dry and wet deposition; The time dimension of the simulation process is used to capture the dynamic changes of the pollutant and ensure the time continuity; The three-dimensional wind speed field is used to simulate the carrying effect of wind on pollutants.

[0045] Further, the adaptive weighted fusion is to multiply the LSTM output concentration and the pollutant concentration field by the corresponding weight coefficients and sum them up, wherein the weight coefficients corresponding to the LSTM output concentration and the pollutant concentration field are both 50%.

[0046] S4, based on a preset pollution concentration threshold, filtering the abnormal values in the pollution prediction result to obtain an initial high pollution early warning signal, and performing credibility review on the initial high pollution early warning signal to obtain a final high pollution early warning signal; In the embodiment of the present application, the filtering of abnormal values in the pollution prediction result based on the preset pollution concentration threshold to obtain the initial high pollution early warning signal comprises: Based on the preset pollution concentration threshold, the pollution prediction result is dynamically graded and filtered to obtain a pollution concentration abnormal value mark; perform spatial clustering analysis on the pollution concentration abnormal value label to obtain a pollution concentration abnormal area block; perform spatial influence domain evaluation on the pollution concentration abnormal area block to obtain a spatial influence index; perform significance determination on the spatial influence index and a preset pollution concentration threshold to obtain an initial high pollution early warning signal.

[0047] It should be noted that the initial high pollution early warning signal includes a pollution concentration abnormal value, an abnormal area position and an initial pollution level.

[0048] It should be noted that the process of dynamic grading filtering is as follows: a preset pollution concentration threshold is used to compare each data point in the pollution prediction result to obtain a pollution comparison result, which is a matrix containing pollution concentration values; According to the pollution comparison result, the data points are divided into multiple pollution levels, and if the pollution concentration exceeds the threshold, the abnormal value is marked to generate a pollution concentration abnormal value label.

[0049] It should be noted that the process of spatial clustering analysis is based on the pollution concentration abnormal value label, and the spatial clustering algorithm is used to aggregate the abnormal points adjacent in space to form a continuous pollution concentration abnormal area block.

[0050] It should be noted that the process of spatial influence domain evaluation is based on the pollution concentration abnormal area block, and the influence degree of each block on the surrounding geographical range is evaluated, and the spatial influence index is calculated to quantify the influence.

[0051] Further, the spatial influence domain evaluation considers the pollution intensity, spatial range and diffusion potential of the block, and outputs the spatial influence index as a quantitative indicator.

[0052] It should be noted that the process of significance determination is to compare the spatial influence index with the preset pollution concentration threshold to determine whether the index is significantly higher than the threshold. If the index exceeds the threshold, it is determined to be a significantly high pollution event, thereby generating an initial high pollution early warning signal.

[0053] In the embodiment of the present application, the initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal, which includes: S701: a credibility review algorithm is used to review the credibility of the initial high pollution early warning signal to obtain a credibility review score, wherein the mathematical expression of the credibility review algorithm is as follows:

[0054] In the formula, is the credibility review score, is a similarity matching score, a spatial consistency index, and are similarity weight coefficients and spatial consistency weight coefficients, respectively; S702, based on the credibility review score and the preset credibility threshold, dynamically adjusting the initial pollution level in the initial high-pollution early warning signal to generate a final pollution level; S703, based on the credibility review score, performing influence area boundary optimization on the abnormal area position in the initial high-pollution early warning signal to generate final air influence area information; S704, based on the final air influence area information, outputting a final high-pollution early warning signal including the final pollution level and the final air influence area information.

[0055] It should be noted that the similarity matching score is based on a preset historical pollution event database, and the initial high-pollution early warning signal is matched in space and time to generate a similarity matching score; The spatial consistency index is based on the abnormal area position in the initial high-pollution early warning signal to perform spatial consistency analysis on the real-time acquired pollution concentration data to obtain a spatial consistency index.

[0056] It should be noted that the process of dynamically adjusting the initial pollution level in the initial high-pollution early warning signal is based on the comparison result of the credibility review score and the preset credibility threshold, and the initial pollution level in the initial high-pollution early warning signal is dynamically corrected. If the credibility review score is lower than the threshold, the initial pollution level is reduced; If the score is higher than the threshold, the initial level is improved or maintained; Through the elastic adjustment of the score driven level, the reliability of the final pollution level is ensured to output the final pollution level.

[0057] It should be noted that the process of influence area boundary optimization is to use the credibility review score to perform spatial boundary fine-tuning on the abnormal area position in the initial high-pollution early warning signal. If the score is lower, the abnormal area boundary is contracted to exclude the low-confidence area; If the score is higher, the boundary is expanded to cover the potential impact range; Finally, the accurate final air influence area information is generated to ensure that the early warning range matches the actual risk of pollution diffusion.

[0058] S5, performing fine emission analysis on the pollution prediction result and the air quality standardized data to obtain pollution emission data; In the embodiment of the application, the fine emission analysis on the pollution prediction result and the air quality standardized data to obtain pollution emission data comprises: constructing a pollutant concentration distribution matrix containing predicted concentration values of pollution elements at each spatial grid point based on the pollution prediction result; constructing a pollution element-threshold mapping table containing standard concentration thresholds of each pollution element based on pollution source type data in the air quality standardization data; comparing the pollutant concentration distribution matrix with the standard concentration thresholds in the pollution element-threshold mapping table to obtain a set of pollution concentration exceeding elements; calculating pollution element exceeding degrees based on the set of pollution concentration exceeding elements and the pollutant concentration distribution matrix using an exceeding degree calculation formula; integrating the set of pollution concentration exceeding elements and the pollution element exceeding degrees with pollution emission data to obtain pollution emission data containing pollution concentration exceeding elements, exceeding degrees, and pollution ranges.

[0059] It should be noted that the pollutant concentration distribution matrix refers to structured stored predicted concentration values, the matrix row represents a pollution element, and the list represents a spatial grid point, which facilitates algorithm processing and improves calculation efficiency. The pollution element-threshold mapping table refers to a pre-set pollutant threshold table that defines the safety limits of each pollutant and provides a standard for exceeding judgment to ensure objectivity.

[0060] In the embodiments of the present application, the exceeding degree calculation formula is as follows:

[0061] In the formula, is an exceeding degree index, is the total number of spatial grids, is a spatial grid point identifier, is a dynamic distance weight, is a predicted concentration, is a standard concentration threshold, is a pollution element identifier.

[0062] It should be noted that the exceeding degree index represents the overall exceeding severity of the pollution element, the larger the value, the more serious the pollution, the spatial grid point identifier represents the position number after regional division, the total number of spatial grids represents the resolution of the analysis area, the larger the value, the more detailed the analysis, the dynamic distance weight adjusts the contribution weight based on the distance from the pollution source, the closer to the pollution source, the higher the weight, and the predicted concentration represents the predicted concentration value of the pollution element at the grid point.

[0063] S6, based on the high pollution warning signal and the pollution emission data, a path optimization algorithm is used to perform pollution tracing on the air quality standardization data to obtain the pollution source position of the current monitored atmospheric environment.

[0064] In the embodiment of the present application, the pollution source position of the current monitored atmospheric environment is obtained by using the path optimization algorithm to trace the pollution of the air quality standardized data. Based on the final high-pollution early warning signal, spatial gridding concentration data of the pollution concentration distribution matrix is extracted. Based on the spatial gridding concentration data, concentration change quantitative processing is performed on the pollution concentration exceeding elements to obtain concentration change direction and intensity quantitative indicators of the spatial gridding concentration data. Based on the concentration change direction and intensity quantitative indicators, maximum value positioning is performed to obtain pollution element candidate pollution source coordinates. The centroid position of all pollution element candidate pollution source coordinates is calculated to obtain the pollution source position of the current monitored atmospheric environment.

[0065] It should be noted that the concentration change quantitative processing is based on the spatial gridding concentration data and the pollution concentration exceeding elements, and the spatial concentration change of each exceeding element is quantitatively analyzed, including: calculating the concentration gradient of each pollution element on the spatial grid point to determine the concentration change direction and change intensity, and generating a structured concentration change direction and intensity quantitative indicator through quantitative processing.

[0066] Compared with the prior art, the present application has the following beneficial effects: 1. Multi-source data credibility quantitative evaluation is adopted to construct a credibility score model based on data source, collection equipment and historical consistency, and to dynamically weight and fuse data, so as to solve the prediction deviation caused by false monitoring data and improve data quality.

[0067] 2. A physical constraint graph neural network transmission model is adopted to embed the atmospheric diffusion equation into the GNN node update mechanism, constrain the pollution cross-region transmission path, and enhance the model coupling.

[0068] 3. A prediction-tracing two-way error feedback mechanism is adopted to use the tracing result to correct the emission source parameters of the prediction model in reverse, form a closed loop optimization, and improve the pollution tracing accuracy.

[0069] In the several embodiments provided in the present application, it should be understood that the disclosed method can be implemented in other ways.

[0070] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.

[0071] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. The artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use the knowledge to obtain the best results.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. An air quality prediction and pollution tracing method based on multi-source data coupling, characterized in that, The method comprises: S1, data standardization processing is carried out on pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data; S2, based on a preset atmospheric chemical model, the air quality standardized data is simulated for pollutant distribution to generate pollutant distribution data containing chemical analysis results of content and concentration of each element; S3, based on the pollutant distribution data, the pollution concentration of the current monitoring atmospheric environment is predicted to generate a pollution prediction result of the current monitoring atmospheric environment; S4, based on a preset pollution concentration threshold, abnormal values in the pollution prediction result are filtered to obtain an initial high pollution early warning signal, and the initial high pollution early warning signal is reviewed for credibility to obtain a final high pollution early warning signal; S5, fine emission analysis is carried out on the pollution prediction result and the air quality standardized data to obtain pollution emission data; S6, based on the high pollution early warning signal and the pollution emission data, a path optimization algorithm is used to trace the pollution of the air quality standardized data to obtain the pollution source position of the current monitoring atmospheric environment.

2. The method of claim 1, wherein, The data standardization processing on the pre-acquired meteorological data, pollution source data, emission source data and topographic data to generate air quality standardized data comprises: An original data set containing pre-acquired meteorological data, pollution source data, emission source data and topographic data is constructed; The original data set is evaluated for credibility to obtain a credibility score; A dynamic fusion weight is calculated based on the credibility score; The original data set is weighted and fused based on the dynamic fusion weight to obtain air quality standardized data.

3. The method of claim 2, wherein, The credibility evaluation on the original data set to obtain a credibility score comprises: Based on the original data set, a credibility score model containing a power consumption-emission coupling verification rule is constructed, wherein the mathematical formula used in the credibility score model is as follows: , wherein, is an initial trust score for an original data vector in the original data set, is a data acquisition device calibration record score for an original data vector in the original data set, is a historical data volatility score for an original data vector in the original data set, is a spatial consistency score for an original data vector in the original data set, , and are calibration record weight coefficient, volatility weight coefficient, and spatial consistency weight coefficient, respectively.​​​​ The initial credibility score is verified by using the power consumption-emission coupling verification rule to obtain a credibility score, wherein the specific verification rule of the power consumption-emission coupling verification rule is as follows: The power consumption change rate is obtained by calculating the ratio of the current power consumption to the historical power consumption average; The emission change rate is obtained by calculating the ratio of the current emission to the historical emission average; When the power consumption change rate is greater than a preset power consumption change threshold, and the emission change rate is less than a preset emission change threshold, a degradation factor is added to the initial credibility score, and the initial credibility score is degraded and punished to obtain a credibility score; When the power consumption change rate is less than or equal to a preset power consumption change threshold, and the emission change rate is greater than a preset emission change threshold, the initial credibility score is the credibility score.

4. The method of claim 1, wherein, The air quality standardized data is simulated for pollutant distribution based on the preset atmospheric chemical model to generate pollutant distribution data, comprising: Based on the air quality standardized data, a graph neural network with cities as nodes is constructed; In the message passing process of the graph neural network, an atmospheric diffusion equation is embedded to simulate the generation, transmission and deposition process of pollutants, and the pollutant distribution data is generated, wherein the atmospheric diffusion equation is as follows: , In the formula, For the first Nodes in a layered GNN The characteristic vector of pollutant concentration, For activation function, This is the weight matrix. For nodes The set of neighboring nodes, For nodes The set of neighboring nodes, For nodes arrive exist The average wind speed component in the direction, For nodes arrive exist The average wind speed component in the direction, For nodes arrive The settlement coefficient, For bias vectors, Represents the normalization factor. For pollutant concentration, Represents spatial coordinates.

5. The method of claim 1, wherein, The pollution concentration of the current monitoring atmospheric environment is predicted based on the pollutant distribution data, and a pollution prediction result of the current monitoring atmospheric environment is generated, which includes: An initial field is constructed based on the pollutant distribution data; The constructed initial field is input into an LSTM-numeric model hybrid framework to capture the long-term dependence of the pollutant concentration over time to obtain a one-stage prediction result; Based on the constructed initial field, the pollutant distribution data is subjected to numerical analysis of atmospheric physical mechanisms to generate a two-stage prediction result; Based on the prediction period, the one-stage prediction result and a plurality of two-stage prediction results are adaptively weighted and fused to obtain the pollution prediction result of the current monitoring atmospheric environment.

6. The method of claim 5, wherein, Based on the preset pollution concentration threshold, the pollution prediction result is filtered to obtain an initial high pollution early warning signal, which includes: Based on the preset pollution concentration threshold, the pollution prediction result is dynamically graded and filtered to obtain a pollution concentration abnormal value label; The pollution concentration abnormal value label is subjected to spatial clustering analysis to obtain a pollution concentration abnormal area block; The pollution concentration abnormal area block is subjected to spatial influence domain evaluation to obtain a spatial influence index; The spatial influence index and the preset pollution concentration threshold are subjected to significance determination to obtain an initial high pollution early warning signal.

7. The method of claim 1, wherein, The initial high pollution early warning signal is subjected to credibility review to obtain a final high pollution early warning signal, which includes: S701: A credibility review algorithm is used to review the credibility of the initial high pollution early warning signal to obtain a credibility review score, wherein the mathematical expression of the credibility review algorithm is as follows: , wherein is a plausibility check score, is a similarity match score, is a spatial consistency index, and are a similarity weight coefficient and a spatial consistency weight coefficient, respectively. S702, based on the credibility review score and the preset credibility threshold, the initial pollution level in the initial high pollution early warning signal is dynamically adjusted to generate a final pollution level; S703, based on the credibility review score, the abnormal area position in the initial high pollution early warning signal is subjected to influence area boundary optimization to generate final air influence area information; S704, based on the final air influence area information, the final high pollution early warning signal including the final pollution level and the final air influence area information is output.

8. The method of claim 1, wherein, The pollution prediction result and the air quality standardized data are subjected to fine emission analysis to obtain pollution emission data, which includes: A pollutant concentration distribution matrix containing the predicted concentration values of pollution elements at each spatial grid point is constructed based on the pollution prediction result; Based on the pollution source type data in the air quality standardized data, a pollution element-threshold mapping table containing the standard concentration threshold of each pollution element is constructed; The pollutant concentration distribution matrix is compared with the standard concentration threshold in the pollution element-threshold mapping table to obtain a pollution concentration exceeding element set; Based on the exceeding standard element set and the pollutant concentration distribution matrix, an exceeding standard degree calculation formula is used to calculate the exceeding standard degree of the pollution elements; The exceeding standard element set and the exceeding standard degree of the pollution elements are integrated for pollution emission data to obtain pollution emission data containing pollution concentration exceeding standard elements, exceeding standard degrees and pollution ranges.

9. The method of claim 8, wherein, The exceeding standard degree calculation formula is as follows: , wherein, is an over-standard degree index, is a total number of spatial grids, is a spatial grid point identification, is a dynamic distance weight, is a predicted concentration, is a standard concentration threshold value, is a pollution element identification.

10. The method of claim 8, wherein, The pollution source position of the current monitored atmospheric environment is obtained by using a path optimization algorithm to perform pollution tracing on the air quality standardized data, including: Based on the final high-pollution early warning signal, spatial gridding concentration data of the pollution concentration distribution matrix is extracted; Based on the spatial gridding concentration data, a concentration change quantification processing is performed on the exceeding standard element set to obtain a concentration change direction and intensity quantification index of the spatial gridding concentration data; Based on the concentration change direction and intensity quantification index, a maximum value is located to obtain pollution element candidate pollution source coordinates; The centroid position of all pollution element candidate pollution source coordinates is calculated to obtain the pollution source position of the current monitored atmospheric environment.

Citation Information

Patent Citations

  • Accurate traceability method and system for atmospheric pollution source

    CN120146397A

  • Atmospheric pollution monitoring method based on unmanned aerial vehicle remote sensing and machine learning

    CN120446400A

  • Indoor air quality joint detection system fused with space-time diagram neural network

    CN120847348A

  • Traceability analysis method and system for atmospheric pollutants

    CN120895127A

  • System and method for dynamic management and control of air pollution

    US20230252487A1

Cited By

  • Monitoring and early warning method for rapid and accurate traceability

    CN121481276A

  • Environmental data denoising and feature enhancement method based on physical and data dual drive

    CN121808206A

  • Air quality high-value event identification method, device, equipment, medium and product

    CN122065012A

  • An air quality high-value event identification method, device, equipment, medium and product

    CN122065012B