Atmospheric pollution gridding prediction method and system based on meteorological parameter optimization

By constructing a parameter combination-meteorological result mapping database and causal verification, combining the ST-GCN model and conditional generation adversarial network, the shortcomings in regional adaptability and prediction accuracy of the traditional atmospheric pollutant PM2.5 prediction model are solved, and high-precision pollutant concentration prediction and distribution modeling are achieved.

CN120387379AActive Publication Date: 2025-07-29SICHUAN GUOLAN ZHONGTIAN ENVIRONMENTAL TECH GRP CO LTD

Patent Information

Application Number
CN202510873694.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The traditional atmospheric pollutant PM2.5 prediction model has insufficient regional adaptability and prediction accuracy, and the traditional feature screening method cannot distinguish the pseudo-correlation between variables, and a single loss function is difficult to balance the goal conflict between online learning and integrated optimization.

Method used

Using WRF-based intelligent parameter matching, multi-dimensional pollutant feature screening and dynamic mapping fusion methods, the parameter combination-meteorological result mapping database is constructed, causal verification and correlation variable screening are carried out, and extreme meteorological scenes are generated by combining the ST-GCN model and the conditional generation adversarial network to achieve high-precision prediction of pollutant concentration.

Benefits of technology

It significantly improves the accuracy, efficiency and scenario adaptability of grid prediction of pollutants, realizes high-precision modeling of meteorological prediction and pollution distribution, and provides reliable technical support for the control of air pollution problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387379A_ABST
    Figure CN120387379A_ABST
Patent Text Reader

Abstract

The invention discloses an atmospheric pollution gridding prediction method and system based on meteorological parameter optimization, and the method comprises the steps: constructing a parameter combination-meteorological result mapping database, carrying out the matching of a recommendation parameter combination corresponding to regional meteorological forecast data in the mapping database, and outputting the gridding meteorological forecast data corresponding to the recommendation parameter combination; taking the historical pollutant concentration monitoring data and the historical meteorological data as an observation instance set to screen a preliminary candidate variable set; performing causal verification on the preliminary candidate quantity set to obtain a single variable set, and performing functional coupling to obtain an associated variable set; inputting the single variable set and the associated variable set into the pollutant concentration prediction model, and outputting a pollutant concentration prediction result; according to gridding weather forecast data, a pollutant concentration prediction result is interpolated into each grid to obtain grid distribution of pollutant concentration, and high-precision modeling of weather prediction and pollution distribution is realized based on mapping database matching, multi-dimensional pollutant feature screening and dynamic mapping fusion methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of atmospheric environment monitoring, and particularly relates to an atmospheric pollution grid prediction method and system optimized based on meteorological parameters. Background Art

[0002] Since the generation of pollutants such as PM2.5 in atmospheric pollutants is affected by many complex factors, the changes in space and time make it difficult for traditional single models to achieve sufficient prediction accuracy. In the prior art, the selection of WRF parameterization schemes highly depends on manual experience, resulting in poor regional adaptability and low efficiency; moreover, for PM2.5 prediction, traditional feature screening methods (such as Pearson correlation coefficient, analysis of variance) are relied on, but the pseudo-correlation between variables cannot be distinguished; at the same time, traditional machine learning models adjust weights through a single loss function (such as MSE), and it is difficult to balance the objective conflicts between online learning and integrated optimization. Summary of the Invention

[0003] The purpose of the present invention is to provide an atmospheric pollution grid prediction method and system optimized based on meteorological parameters, and propose a method for intelligent matching of parameters based on WRF, multi-dimensional pollutant feature screening, and dynamic mapping fusion, so as to achieve high-precision modeling of meteorological prediction and pollution distribution, and provide reliable technical support for solving the problem of atmospheric pollution.

[0004] To achieve the above purpose, the present application adopts the following solutions: On the one hand, the present invention provides an atmospheric pollution grid prediction method optimized based on meteorological parameters, specifically including the following steps: S1. Obtain historical pollutant concentration monitoring data, historical meteorological data, and regional meteorological forecast data of micro monitoring stations in the target area; S2. Construct a mapping database of parameter combination - meteorological result according to the physical parameterization scheme of the WRF forecast model, match the recommended parameter combination corresponding to the regional meteorological forecast data in the mapping database, input the recommended parameter combination into the WRF forecast model for operation, and output the grid meteorological forecast data in the target area; S3. Use the historical pollutant concentration monitoring data and historical meteorological data as an observation instance set to screen a preliminary candidate variable set; conduct causal verification on the preliminary candidate variable set to obtain a single variable set, and conduct functional coupling on the single variable set to obtain an associated variable set; S4. Input the single variable set and the associated variable set into the pollutant concentration prediction model, and output the pollutant concentration prediction result; S5. Map the micro monitoring stations into the grids of the target area, and interpolate the pollutant concentration prediction result into each grid according to the grid meteorological forecast data to obtain the grid distribution of the pollutant concentration.

[0005] In some specific embodiments, the specific process of step S2 is as follows: S21. Obtain GFS data within the target area, select multiple combinations of core parameters from the physical parameterization scheme of the WRF prediction model, use the GFS data as the initial input of the WRF prediction model for each combination of core parameters, and output the corresponding meteorological prediction data; S22. Store the combination of core parameters and its corresponding meteorological prediction data in the mapping database in the form of parameter combination - meteorological result combination; S23. Use the classification model with the combination of core parameters as the classification label and the meteorological prediction data of the meteorological station as the feature, and match the parameter combination with the minimum error in the mapping database for the meteorological prediction data of the meteorological station as the recommended parameter combination; S24. Input the recommended parameter combination into the WRF prediction model for operation, and output the gridded meteorological forecast data.

[0006] In some specific embodiments, the specific process of obtaining a single variable set in step S3 is as follows: S31. Use the historical pollutant concentration monitoring data and historical meteorological data as the observation instance set, perform Bayesian neural network modeling on the original variables, where the original variables include geographical factors and meteorological factors. Through Monte Carlo sampling, quantify the prediction variance contribution of each original variable, and select candidate variables from each original variable according to the prediction variance contribution and store them in the preliminary candidate variable set; S32. Conduct causal verification on each variable in the preliminary candidate variable set to obtain a single variable set.

[0007] In some specific embodiments, the specific process in step S32 is as follows: Use the causal forest to calculate the average treatment effect ATE of each variable in the preliminary candidate variable set, verify the causal direction p of each variable through counterfactual analysis, and select variables with ATE > 0.1 and p < 0.05 from the preliminary candidate variable set and save them as a single variable set.

[0008] In some specific embodiments, the process of obtaining the associated variable set is as follows: S01. Perform variable combination on the variables in the single variable set using wrapper feature selection, and conduct co - verification on the variable combinations to obtain multiple significantly associated variable combinations; S02. Divide the target area, merge similar areas according to the variables within each area to obtain several sub - areas, use the sub - areas as the input nodes of the ST - GCN model, and obtain the obstacle coefficient of the terrain to pollutant diffusion through the k - hop neighbor aggregation of the ST - GCN model; S03. Use a conditional generative adversarial network to generate extreme meteorological scenarios, and combine the obstacle coefficient of terrain to the pollutant diffusion to screen the significant associated variable combinations to obtain a candidate variable set; S04. Filter out unreasonable variables in the candidate variable set through physical constraints to obtain an associated variable set.

[0009] In some specific implementation manners, the specific process of step S01 is as follows: Evaluate the prediction gain of variable combinations through model importance scoring, and screen out variable combinations with clear physical meanings from multiple variable combinations to obtain multiple Top-K combined variables; Calculate the conditional average treatment effect of each Top-K combined variable, then conduct counterfactual analysis to simulate the joint intervention of the Top-K combined variables, calculate the synergy between the conditional average treatment effect and the joint intervention of each Top-K combined variable, and output the Top-K combined variables that meet the synergy condition as significant associated variable combinations.

[0010] In some specific implementation manners, the specific process of step S4 is as follows: S41. Use FTRL to screen out variables with high contribution and high stability from a single variable set and an associated variable set to form a basic variable set; S42. Use the basic variable set as the input of the pollutant concentration prediction model, and construct several weak classifiers in the pollutant concentration prediction model to quantify the prediction ability of different variable combinations, and output the pollutant concentration prediction result.

[0011] In some specific implementation manners, the specific process of step S5 is as follows: S51. Correlate the grids of the target area with the grids of the WRF forecast model, and map the longitude and latitude of the micro monitoring sites into the grids of the WRF forecast model as the benchmark points for grid interpolation; S52. Select the corresponding interpolation algorithm according to the meteorological factors and terrain factors of each grid; S53. Insert the pollutant concentration prediction result into each grid according to the benchmark points according to the interpolation algorithm.

[0012] In some specific implementation manners, the interpolation algorithms include integrated Kriging, inverse distance weighting, radial basis function, and modified IDW algorithms. A dynamic selector is constructed to store the corresponding relationships between different meteorological factors and terrain factors and the adopted interpolation algorithms. The terrain deviation in the terrain factors of the grid is corrected through a neural network, and the corresponding interpolation algorithm is selected from the dynamic selector according to the meteorological factors and the corrected terrain factors.

[0013] In a second aspect, the present application provides an atmospheric pollution grid prediction system based on meteorological parameter optimization, including: One or more processors; A storage unit for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement a method for grid prediction of air pollution optimized based on meteorological parameters according to the first aspect.

[0014] Advantages of the present invention: By constructing a parameter combination - meteorological result mapping database, the dynamic optimization and intelligent adaptation of the physical parameterization scheme of numerical weather prediction (WRF) are realized; through the collaborative mechanism of single variable screening and functional coupling of associated variables, the multi-scale dynamic optimization of pollutant concentration prediction is realized, significantly improving the generalization ability of the model in complex meteorological and geographical scenarios. At the same time, through causal - physical joint driving, dynamic modeling of spatial heterogeneity, and incremental weight allocation, a multi-algorithm integration and intelligent decision-making mechanism is adopted, significantly improving the accuracy, efficiency, and scenario adaptability of grid prediction of pollutants. Description of the drawings

[0015] Figure 1 It is a flowchart of a method for grid prediction of air pollution optimized based on meteorological parameters provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the process of grid prediction of air pollution provided by an embodiment of the present invention. Detailed implementation manners

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way constitutes a limitation on the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0017] Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps described in these embodiments do not limit the scope of the present invention.

[0018] Meanwhile, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0019] In addition, for the sake of clarity and conciseness, the description of well-known structures, functions, and configurations may be omitted. Those of ordinary skill in the art will recognize that various changes and modifications can be made to the examples described herein without departing from the spirit and scope of the present disclosure.

[0020] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices shall be considered part of the authorization specification.

[0021] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0022] Embodiment 1 As Figure 1 - Figure 2 shown, this embodiment provides an atmospheric pollution grid prediction method optimized based on meteorological parameters, specifically including the following steps: 1. Data processing S1. Obtain regional geographical data, pollutant emission data, historical pollutant concentration monitoring data of micro monitoring stations, historical meteorological data, and regional meteorological forecast data within the target area, where: Regional static geographical data: terrain, vegetation coverage rate, land use type, etc.; Historical pollutant concentration monitoring data of micro monitoring stations: PM2.5 (data of multiple micro monitoring stations in the target area and hourly data of the monitoring stations) Historical meteorological data: temperature, humidity, air pressure, wind speed, wind direction, dew point temperature, rainfall (data of micro meteorological stations in the target area and hourly data of meteorological monitoring stations) Regional meteorological forecast data from the meteorological data network: wind speed, rainfall, temperature, wind direction, air pressure; GFS meteorological data; input data required by the WRF forecast model; Perform data processing on the data obtained above, including the following processes: (1) Data missing: There are missing values in the original dataset, and linear interpolation is used to complete them.

[0023] (2) Data anomaly: Anomalous values are regarded as missing values and processed using the data missing processing method.

[0024] (3) Data normalization: Use the maximum minimum normalization (Min - Max Normaliation) to normalize the data.

[0025] 2. Dynamic optimization of WRF parameters S2. Construct a mapping database of parameter combinations - meteorological results according to the physical parameterization scheme of the WRF forecast model, match the recommended parameter combination corresponding to the regional meteorological forecast data in the mapping database, input the recommended parameter combination into the WRF forecast model for operation, and output the grid meteorological forecast data within the target area; The specific process of step S2 is: S21. Obtain GFS data within the target area, select multiple core parameter combinations from the physical parameterization schemes of the WRF prediction model, use the GFS data as the initial input of the WRF prediction model for each core parameter combination, and output the corresponding meteorological prediction data; S22. Store the core parameter combinations and their corresponding meteorological prediction data in the mapping database in the form of parameter combination - meteorological result combination; S23. Use the classification model with the core parameter combination as the classification label and the meteorological station prediction data as the feature, and match the parameter combination with the minimum error in the mapping database for the meteorological station prediction data as the recommended parameter combination; S24. Input the recommended parameter combination into the WRF prediction model and run it in combination with the regional static geographical data to output the grid - based meteorological forecast data.

[0026] Among them, the physical parameterization schemes include, but are not limited to, cumulus schemes: KF, GR, BMJ; planetary boundary layer schemes: MYJ, ACM2, Boulac, etc. When selecting multiple core parameter combinations, the contribution degree of parameters to the prediction error can be quantified through a machine learning model (such as a random forest model) to screen out the core parameter combinations. The random forest model can train the non - linear relationship between the parameter combinations and the meteorological results. The input is the parameter scheme (such as cumulus scheme = KF, boundary layer scheme = MYJ), and the output is the meteorological prediction error (such as temperature error ΔT, precipitation error ΔP). Key parameters are screened according to the feature importance (Gini index) (for example, the cumulus scheme has the greatest impact on precipitation error (weight 0.35), and the boundary layer scheme has a significant impact on wind speed error (weight 0.28); the positive correlation between the KF cumulus scheme and the precipitation data of the meteorological station (r = 0.72), and the negative correlation between the MYJ boundary layer scheme and the temperature error (r = - 0.65)). Use a classification model such as (XGBoost), with the parameter combination as the classification label and the meteorological station prediction data as the feature, train the classification model, and substitute the parameter combination with the minimum error under similar meteorological conditions in the historical database into the WRF to run and output the grid - based meteorological forecast data.

[0027] 3. Pollutant Concentration Prediction (Taking PM2.5 as an Example) S3. Single - variable and associated - variable screening: Use the historical pollutant concentration monitoring data and historical meteorological data as the observation instance set to screen the preliminary candidate variable set; conduct causal verification on the preliminary candidates to obtain the single - variable set, and perform functional coupling on the single variables to generate the associated - variable set. The specific process of step S3 is as follows: Obtain the single - variable set: S31. Data Preparation and Variable Screening: Use historical pollutant concentration monitoring data and historical meteorological data as the observation instance set. Build a Bayesian neural network model for the observation instance set. Through Monte Carlo sampling, quantify the predictive variance contribution of each variable, and screen out candidate variables from each variable according to the predictive variance contribution and store them in the preliminary candidate variable set. Use the collected and processed historical daily average pollutant concentration data set and meteorological data set as the observation instance set (such as temperature, wind speed, vegetation coverage, etc.). Build a Bayesian neural network model for the observation instance set. Through Monte Carlo Dropout sampling 1000 times, quantify the predictive variance contribution of each variable (such as wind speed variance = 3.2%, vegetation coverage variance = 8.5%). Then, by setting a threshold (such as variance < 5%), screen out highly stable variables (wind speed passes, vegetation coverage is excluded) as the preliminary candidate variable set.

[0028] S32. Causal Effect Verification: Conduct causal verification on each variable in the preliminary candidate variable set to obtain a single variable set.

[0029] The specific process in step S32 is as follows: Use the causal forest to calculate the average treatment effect ATE of each variable in the preliminary candidate variable set. Through counterfactual analysis, verify the causal direction p of each variable. Screen out variables with ATE > 0.1 and p < 0.05 from the preliminary candidate variable set and output them as a single variable set.

[0030] For the variables that passed the stability verification in step S31 above, use the causal forest to calculate the average treatment effect (ATE). For example, when the wind speed increases by 1 m / s, PM2.5 decreases by 0.25 μg / m³. Through counterfactual analysis, verify the causal direction (if the wind speed is artificially reduced by 10%, whether the PM2.5 concentration significantly increases (p < 0.05)), and exclude variables with insignificant causal effects (such as the ATE of land use type = 0.08, p = 0.12). Store variables with ATE > 0.1 and p < 0.05 in the single variable set; The process of obtaining the associated variable set is as follows: S01. For the variables in the single variable set, perform wrapper feature selection for variable combination, and conduct collaborative verification on the variable combination to obtain multiple significant associated variable combinations; The specific process of step S01 is as follows: Evaluate the prediction gain of variable combinations through model importance scoring, and screen out variable combinations with clear physical meanings from multiple variable combinations to obtain multiple Top-K combined variables; Calculate the conditional average treatment effect of each Top-K combination variable, then conduct counterfactual analysis to simulate the joint intervention of the Top-K combination variables, calculate the synergy between the conditional average treatment effect of each Top-K combination variable and the joint intervention, and output the Top-K combination variables that meet the synergy condition as the significant association variable combination.

[0031] For the variables selected based on single variable screening, use the Wrapper Methods for feature selection. Evaluate the prediction gain of variable combinations after pairwise or triple combinations of variables through model importance scoring (such as the Gini importance of random forest and the gain contribution of XGBoost), and preferentially generate combinations with clear physical meanings (such as wind speed × inversion layer thickness, etc., not all-permutation combinations to avoid the curse of dimensionality), and output the Top-K combination variables.

[0032] Calculate the conditional average treatment effect (CATE) of the Top-K combination variables. For example, evaluate the sensitivity of the PM2.5 concentration change when the wind speed increases by 1 m / s and the inversion layer thickness > 100 m. Then conduct counterfactual analysis to simulate the joint intervention of the variable combination (such as simultaneously reducing the wind speed by 10% and closing a certain industrial area), verify whether the synergy is significant (p < 0.05), and output the significant association variable set.

[0033] S02, Spatial relationship modeling: Divide the target area into regions, merge similar regions according to the variables within each region to obtain a sub-region grid, use the sub-region grid as the input node of the ST-GCN model, and through the k-hop neighbor aggregation of the ST-GCN model, obtain the obstacle coefficient of the terrain to pollutant diffusion; Use the GeoDetector to quantify the spatial heterogeneity (only for preprocessing), output the heterogeneity index according to the regional characteristics, dynamically adjust the variance contribution threshold of the previous variable screening, merge sub-regions with similar characteristics (such as grids with vegetation coverage > 60%), reduce the number of calculation nodes, and retain the internal heterogeneity. The merged sub-regions are used as the input nodes of the ST-GCN, and through the k-hop neighbor aggregation of the ST-GCN (meeting the boundary conditions of the diffusion equation), capture the spatio-temporal coupling effect between variables, quantify the obstacle coefficient of the terrain to pollutant diffusion, and assign weights to the terrain obstacles (such as when the mountain elevation > 500 m, the weight +0.2).

[0034] S03, Extreme scenario generation: Use the conditional generative adversarial network to generate extreme meteorological scenarios, and screen the significant association variable combinations in combination with the obstacle coefficient of the terrain to pollutant diffusion to obtain the candidate variable set; Use conditional generative adversarial network (cGAN) to generate extreme meteorological scenarios (such as stable weather + high industrial emissions); filter out unreasonable variables in the associated variable set through physical constraints (mass conservation equation), and conduct distribution similarity tests (KL divergence < 0.1) between the generated data and the observed values of historical heavy pollution events (such as during red alerts for smog), ensuring the physical interpretability of the generated data, and finally obtaining the candidate variable set.

[0035] S04. Feature optimization: Reduce the number of candidate variables in the candidate variable set through dynamic causal screening, reducing the subsequent computational complexity. Based on the high-weight variables (higher ATE) in a single variable set, preferentially generate their combined features (such as wind speed × terrain undulation degree) rather than all permutations and combinations, reducing the feature dimension. Use the screened candidate variable set as the associated variable set.

[0036] It can be understood that the above process realizes the multi-scale dynamic optimization of PM2.5 concentration prediction through the collaborative mechanism of single variable screening and the functional coupling of associated variables, significantly improving the generalization ability of the model in complex meteorological and geographical scenarios. At the same time, through the full-chain technical closed-loop of "causal-driven variable screening → spatial heterogeneity modeling → dynamic weight allocation", the PM2.5 concentration prediction realizes the leap from "experience-driven" to "mechanism-data fusion".

[0037] S4. Input the single variable set and the associated variable set into the pollutant concentration prediction model, and output the pollutant concentration prediction result; the specific process of step S4 is as follows: S41. Use FTRL to screen out variables with high contribution and high stability from the single variable set and the associated variable set to form a basic variable set; S42. Use the basic variable set as the input of the pollutant concentration prediction model, and construct several weak classifiers in the pollutant concentration prediction model to quantify the prediction ability of different variable combinations, and output the pollutant concentration prediction result.

[0038] Dynamic weight allocation Based on the online learning framework, high - contribution variables and high - stability variables are selected through the L1 regularization (sparsity) and L2 regularization (stability) of FTRL (Follow - the - Regularized - Leader) to form a secondary candidate set. The high - stability variables output by FTRL are input into the Adaboost model to construct weak classifiers (decision stumps), and each weak classifier focuses on different variable combinations (such as wind speed × terrain undulation). The classifier weights are optimized through Boosting iteration: in each round of training, a weak classifier is generated based on the current sample weights (initially uniformly distributed); the classification error is calculated and the classifier weights are refined. The lower the error, the greater the weight, thus dynamically strengthening the key combinations; the sample weights are fixed (here, the sample refers to each independent observation point in the training dataset, and each sample contains multiple variables and a target value). The model realizes the fusion of single - variable associated variables through weighted voting of the classifier, and thus trains the Adaboost model. The new air quality data is input into the trained Adaboost model. The model combines the high - stability variables screened by FTRL with the classifier weights learned by Adaboost to dynamically calculate the predicted values (for example, in a static and stable weather scenario, the reduction in wind speed leads to a significant increase in the weight of the "wind speed × inversion layer thickness" combination. The final prediction result is fused by the weighted voting of multiple classifiers, and the optimization effect is evaluated by combining indicators such as the mean squared error (MSE) and the coefficient of determination (R²), forming a closed - loop feedback of "data input → weight update → prediction optimization").

[0039] 4. Dynamic Fusion of Gridded Data S5. Map the micro - monitoring stations into the grids of the target area, and interpolate the pollutant concentration prediction results into each grid according to the gridded meteorological forecast data to obtain the grid distribution of the pollutant concentration.

[0040] The specific process of step S5 is as follows: S51. Correlate the grids of the target area with the grids of the WRF forecast model, and map the longitude and latitude of the micro - monitoring stations into the grids of the WRF forecast model as the benchmark points for grid interpolation; Map the longitude and latitude of the PM2.5 prediction stations (i.e., micro - monitoring stations) into the WRF grids as the benchmark points for grid interpolation; S52. Select the corresponding interpolation algorithm according to the meteorological and topographical factors of each grid; Integrate multiple interpolation algorithms such as Kriging, inverse distance weighting, and radial basis function (covering different meteorological and topographical condition requirements), and design a dynamic selector to construct a multi - interpolation algorithm library. Based on the grid meteorological conditions and topographical features, an online selection of interpolation algorithms is made through a lightweight reinforcement learning model (e.g., when the wind direction changes suddenly and the terrain undulates greatly, automatically switch to the modified IDW algorithm considering obstacles; use Kriging interpolation in static and stable weather).

[0041] S53. Insert the pollutant concentration prediction results into each grid according to the interpolation algorithm based on the reference points. Dynamically adjust the influence weights of the single variable set (key variables) and the associated variable set (variable combination) in model prediction and data fusion by using the parameter correlation rules (the parameter correlation rules include causal effects, spatial heterogeneity, combined feature sensitivity, and the feedback mechanism of the dynamic learning framework) mentioned in the pollutant concentration prediction in step 3. Then correct the terrain deviation through a BP neural network (such as reducing the concentration in the valley area by 12%), and finally output a 1 km × 1 km high-resolution PM2.5 grid distribution map.

[0042] Model update: When the WRF meteorological grid data is updated (such as a new forecast every 6 hours) or new PM2.5 monitoring data is added, the following process is automatically triggered: ① Rematch the station and grid data; ② Dynamically adjust the weight coefficients (such as updating the weights after the wind speed changes); ③ Run the hybrid interpolation algorithm to generate a new distribution map, using a lightweight neural network model (such as MobileNet) to ensure that the 1 km × 1 km grid update is completed within 10 s.

[0043] Embodiment 2 This embodiment provides an atmospheric pollution grid prediction system optimized based on meteorological parameters, including: One or more processors; A storage unit for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement an atmospheric pollution grid prediction method optimized based on meteorological parameters according to the first aspect.

[0044] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Based on the technical essence of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments within the spirit and principles of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An atmospheric pollution grid prediction method optimized based on meteorological parameters, characterized in that, Specifically, it includes the following steps: S1. Obtain the historical pollutant concentration monitoring data, historical meteorological data, and regional meteorological forecast data of the micro-monitoring stations in the target area; S2. Construct a mapping database of parameter combination - meteorological result according to the physical parameterization scheme of the WRF forecast model. Match the recommended parameter combination corresponding to the regional meteorological forecast data in the mapping database, and input the recommended parameter combination into the WRF forecast model to run, and output the gridded meteorological forecast data in the target area; S3. Use the historical pollutant concentration monitoring data and historical meteorological data as the observation instance set to screen the preliminary candidate variable set; conduct causal verification on the preliminary candidate variable set to obtain a single variable set, and conduct functional coupling on the single variable set to obtain an associated variable set; S4. Input the single variable set and the associated variable set into the pollutant concentration prediction model, and output the pollutant concentration prediction result; S5. Map the micro-monitoring stations into the grids of the target area, and interpolate the pollutant concentration prediction result into each grid according to the gridded meteorological forecast data to obtain the grid distribution of the pollutant concentration.

2. The air pollution grid prediction method optimized based on meteorological parameters according to claim 1, wherein The specific process of step S2 is as follows: S21. Obtain the GFS data in the target area, select multiple core parameter combinations from the physical parameterization scheme of the WRF forecast model, and use the GFS data as the initial input of the WRF forecast model for each core parameter combination, and output the corresponding meteorological prediction data; S22. Store the core parameter combination and its corresponding meteorological prediction data in the mapping database in the form of parameter combination - meteorological result combination; S23. Use the classification model with the core parameter combination as the classification label and the meteorological station prediction data as the feature to match the parameter combination with the minimum error in the mapping database for the meteorological station prediction data as the recommended parameter combination; S24. Input the recommended parameter combination into the WRF forecast model to run, and output the gridded meteorological forecast data.

3. A method for grid prediction of air pollution optimized based on meteorological parameters according to claim 1, characterized in that, The specific process of obtaining the single variable set in step S3 is as follows: S31. Use the historical pollutant concentration monitoring data and historical meteorological data as the observation instance set to conduct Bayesian neural network modeling on the original variables. The original variables include geographical factors and meteorological factors. Through Monte Carlo sampling, quantify the prediction variance contribution degree of each original variable, and screen out the candidate variables from each original variable according to the prediction variance contribution degree and store them in the preliminary candidate variable set; S32. Conduct causal verification on each variable in the preliminary candidate variable set to obtain a single variable set.

4. The air pollution grid prediction method optimized based on meteorological parameters according to claim 3, characterized in that The specific process in step S32 is as follows: Use the causal forest to calculate the average treatment effect ATE of each variable in the preliminary candidate variable set, and verify the causal direction p of each variable through counterfactual analysis. Screen out the variables with ATE > 0.1 and p < 0.05 from the preliminary candidate variable set and save them as a single variable set.

5. A method for grid prediction of air pollution optimized based on meteorological parameters according to claim 1, characterized in that The process of obtaining the associated variable set is as follows: S01. Conduct variable combination on the variables in the single variable set using wrapper feature selection, and conduct collaborative verification on the variable combination to obtain multiple significant associated variable combinations; S02. Divide the target area, merge similar areas according to the variables in each area to obtain several sub-areas, use the sub-areas as the input nodes of the ST-GCN model, and obtain the obstacle coefficient of the terrain to pollutant diffusion through the k-hop neighbor aggregation of the ST-GCN model; S03. Use a conditional generative adversarial network to generate extreme meteorological scenarios, and jointly screen the significant associated variable combinations with the obstacle coefficient of the terrain to pollutant diffusion to obtain a candidate variable set; S04. Filter out unreasonable variables in the candidate variable set through physical constraints to obtain an associated variable set.

6. The method for grid prediction of air pollution optimized based on meteorological parameters according to claim 5, characterized in that, The specific process of step S01 is as follows: Evaluate the prediction gain of variable combinations through model importance scoring, and screen out variable combinations with clear physical meanings from multiple variable combinations to obtain multiple Top-K combined variables; Calculate the conditional average treatment effect of each Top-K combined variable, then conduct counterfactual analysis to simulate the joint intervention of the Top-K combined variables, calculate the synergy between the conditional average treatment effect of each Top-K combined variable and the joint intervention, and output the Top-K combined variables that meet the synergy condition as significant associated variable combinations.

7. A method for grid-based prediction of air pollution optimized based on meteorological parameters according to claim 1, characterized in that, The specific process of step S4 is as follows: S41. Use FTRL to screen out variables with high contribution and high stability from the single variable set and the associated variable set to form a basic variable set; S42. Use the basic variable set as the input of the pollutant concentration prediction model, and construct several weak classifiers in the pollutant concentration prediction model to quantify the prediction ability of different variable combinations, and output the pollutant concentration prediction result.

8. A method for grid prediction of air pollution optimized based on meteorological parameters according to claim 1, characterized in that The specific process of step S5 is as follows: S51. Corresponding the grid of the target area with the grid of the WRF forecast model, and mapping the longitude and latitude of the micro monitoring site into the grid of the WRF forecast model as the benchmark point for grid interpolation; S52. Select the corresponding interpolation algorithm according to the meteorological factors and terrain factors of each grid; S53. Insert the pollutant concentration prediction result into each grid according to the benchmark point according to the interpolation algorithm.

9. The method for grid prediction of air pollution optimized based on meteorological parameters according to claim 8, characterized in that, The interpolation algorithms include integrated Kriging, inverse distance weighting, radial basis function and modified IDW algorithms. A dynamic selector is constructed to store the corresponding relationship between different meteorological factors and terrain factors and the interpolation algorithms used. The terrain deviation in the terrain factors of the grid is corrected through a neural network, and the corresponding interpolation algorithm is selected from the dynamic selector according to the meteorological factors and the corrected terrain factors.

10. An air pollution grid prediction system optimized based on meteorological parameters, characterized in that, It includes: One or more processors; A storage unit for storing one or more programs, which when executed by the one or more processors can enable the one or more processors to implement an air pollution grid prediction method based on meteorological parameter optimization as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Consignment smoke-related prediction method fusing space-time and network topology characteristics, medium and equipment

    CN114580499A

  • Meteorological data downscaling method and device based on space-time diagram neural network

    CN118915192A

  • Associated variable determination method and device, equipment, storage medium and program product

    CN119067243A

  • Multi-model fusion pollutant emission distribution rapid correction method and system

    CN119207615A

Cited By

  • Method and system for evaluating sewage treatment quality of mixing station

    CN121639045A

  • Runoff pollutant concentration prediction method, system, equipment and medium

    CN121960905A