Air quality management effectiveness evaluation method based on machine learning algorithm
Through machine learning algorithms and meteorological standardization technology, the problem of evaluating data of pollutant concentration data in traditional methods is solved, and simple evaluation and decision-making support for air quality management results is achieved.
Patent Information
- Application Number
- CN202510164102.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The traditional air quality management effectiveness evaluation method relies on the pollution source emission inventory data. Affected by the uncertainty of the initial meteorological field and the imperfect chemical mechanism, it is difficult to accurately evaluate the changes in pollutant concentrations, and there is a lack of effective strong indicator variables for pollution source emission sources.
The machine learning algorithm is used to process pollutant concentration data through meteorological standardization technology, and a pollutant prediction model is established using a random forest regression algorithm, and a statistical model based on pollutant meteorological standardization concentration is constructed to evaluate the contribution of pollution emissions and meteorological conditions to air quality.
It realizes an objective assessment of the effectiveness of localized air quality management on a hourly scale, simplifies the operation process, provides decision-making support for environmental air quality management measures, and breaks through the technical difficulty limitations of traditional methods.
Smart Images

Figure CN120105362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air pollution prevention and control management, and in particular to an air quality management effectiveness evaluation method based on a machine learning algorithm. Background Art
[0002] Evaluating the effectiveness of air quality management involves answering the question of how much changes in regional or urban air quality are due to pollution emission controls ("human effort") and changes in meteorological conditions ("weather's help"). Traditionally, this evaluation primarily relies on scenario-based simulations using air quality models, such as the Weather Research and Forecasting Model-Community Multiscale Air Quality Model (WRF-CMAQ). These simulations use atmospheric pollution source emission inventory data to simulate ambient air pollutant concentrations. This method uses a switch-off simulation experiment, using actual baseline reference year pollution emission inventory data and the actual year's meteorological conditions. Simulated pollutant concentrations are obtained under the baseline year's emissions scenario, and the simulation results are evaluated by comparing the predicted values with the actual values. Once the model achieves the expected accuracy, further scenario-based simulations are conducted using pollution source emission inventory data from the target year to simulate ambient air pollutant concentrations under the baseline year's meteorological conditions. Under the same meteorological conditions (i.e., the base year's meteorological conditions), the difference between the model's simulated average annual ambient concentration of pollutants emitted in the assessment year and the base year's average observed concentration represents the portion of the change in ambient air quality during the assessment period attributable to pollution reduction (the "human effort" contribution). The meteorological contribution ("weather's help") is calculated by subtracting the model-simulated ambient concentration of pollutants under the base year's meteorological conditions and the assessment year's pollution emission inventory from the average observed concentration in the target year.
[0003] The air quality model establishes a mathematical model based on the partial differential equations of atmospheric physical and chemical mechanisms to describe the relationship between pollution source emissions and pollutant concentrations. It can reflect the impact of pollution source emission reduction and weather and meteorological changes on air quality. However, the model results are affected by the uncertainty of the initial meteorological field, the imperfection of the chemical mechanism, and the accuracy and timeliness of the pollution emission inventory.
[0004] Given that changes in air quality are jointly influenced by pollution source emissions and meteorological factors, the temporal changes in pollutants are considered to be a function of pollution source intensity and meteorological factors. Therefore, a statistical model can be established using meteorological monitoring data and explanatory variables such as pollution source intensity and pollutant concentration. The difficulty of this method lies in the difficulty in obtaining indicator variables that can fully characterize the intensity of pollution emission sources. In recent years, meteorological standardization technology based on random forest regression algorithms has been widely used. This method establishes a pollutant concentration prediction model based on historical observation data. For any given pollutant at a given time, its ambient concentration under different meteorological conditions is predicted, and the average of multiple concentrations is taken to eliminate the disturbance of meteorological factors. The temporal changes in the meteorologically standardized concentration of pollutants block the interference of different meteorological conditions, thereby reflecting the changes in the intensity of local pollution emissions. Summary of the Invention
[0005] The present invention aims to use meteorologically standardized pollutant time series to characterize the changes in local emission source intensity, incorporate them into statistical models together with meteorological parameters, and predict the environmental concentrations of pollutants in the evaluation period under baseline meteorological conditions or baseline emission intensity through the control variable method. The differences between them are compared to evaluate the contributions of local pollution emission control ("human efforts") and changes in meteorological conditions ("help from nature"), and construct a data-driven air quality management effectiveness evaluation method.
[0006] To achieve the above objectives, the present invention provides an air quality management effectiveness evaluation method based on a machine learning algorithm, comprising the following steps:
[0007] S1, obtain the indicator variables of local source emission intensity of pollutants, including:
[0008] S101, prepare a model training dataset by collecting and selecting time series data on the concentrations of six air pollutants, SO2, NO2, CO, O3, PM10, and PM2.5, as well as environmental meteorological parameters and time trend variable data, which are monitored hourly over two years in the target area. The time trend variable is used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes.
[0009] S102, training the prediction model, for a selected air pollutant, using the regression algorithm to predict the concentration of the pollutant ( ) and environmental meteorological parameters ( ) and time trend variables ( ) modeling, using the full hourly data set of two full calendar years to train a pollutant prediction model, comparing the consistency and difference between the pollutant concentrations predicted by the model and the actual monitored concentrations, and calculating the correlation coefficient and root mean square error to evaluate the model fitting effect;
[0010] in Is the model residual term, for any time t, the model ( ) The predicted pollutant concentration is , then:
[0011]
[0012] ,
[0013] Where i represents the meteorological parameter to be modeled; j represents the time trend variable;
[0014] S103, meteorological standardization processing, normalizes the interference of meteorological conditions that change with time series on pollutant concentrations to the average meteorological conditions, uses the trained random forest pollutant prediction model f, and predicts the concentration of pollutants at time t under N randomly selected historical meteorological conditions, taking the algebraic mean of the predicted pollutant concentrations ( ), that is, the meteorological standardized concentration of pollutants at time t is:
[0015]
[0016] Where k is the kth sample of the selected fixed N groups of historical meteorological data, is the pollutant concentration predicted by the model under the kth historical meteorological data of the selected fixed N groups;
[0017] S2: Build a machine learning model for pollutants. Use the meteorologically standardized concentration of pollutants as an indicator variable for local pollution emission intensity and incorporate it into the machine learning training of pollutant regression prediction models, including:
[0018] S201: Prepare the model training dataset, which includes the hourly monitoring data of six air pollutants, SO2, NO2, CO, O3, PM10, and PM2.5, as well as environmental meteorological parameters and time trend variable data, over the past two years in the target area.
[0019] S202, training the prediction model, using regression algorithm to compare the concentration of pollutants with the ambient meteorological parameters ( ) and meteorological normalized concentration ( ) for modeling, is the model residual term, and a pollutant prediction model is retrained with all hourly data sets for two full natural years. For any time t, the model ( ) The predicted pollutant concentration is :
[0020]
[0021] , i represents the listed modeled meteorological parameters;
[0022] S3. Evaluate the contribution of changes in pollution emissions and meteorological conditions to changes in environmental concentrations, determine the baseline reference period for changes in pollutant concentrations in the proposed evaluation period, use the pollutant machine learning model constructed in step S2, simulate pollutant concentrations under different emission and meteorological scenarios based on the control variable method, and obtain the environmental concentration of pollutants under the meteorological conditions of the baseline reference year under the emission intensity of the year to be evaluated. Take the difference between the predicted environmental concentration and the actual monitored concentration in the evaluation year as the contribution, and obtain a quantitative assessment of the change in pollutant concentration caused by pollution emission reduction in the proposed evaluation year compared with the baseline year.
[0023] Preferably, the environmental meteorological parameters include ground temperature T, relative humidity RH, wind speed WS, wind direction WD, air pressure SP, radiation intensity SSR, mixing layer height BLH, total cloud cover TCC, precipitation Prep, trajectory category Air cluster and trajectory length Air length .
[0024] Preferably, the time trend variables include a timestamp Unix time, a lunar day number LunarDay, a solar day number Day-of-Year, a day of the week Day-of-Week, and a daily hour sequence Hour.
[0025] Preferably, the regression algorithm is a random forest algorithm, a neural network algorithm, a gradient regression tree algorithm or an extreme gradient boosting tree algorithm.
[0026] Preferably, the environmental meteorological parameters include a variety of easily accessible meteorological observation data or meteorological reanalysis data, and the variables used to characterize the emission source intensity are any air-related data related to pollution emission activities.
[0027] Preferably, the assessment of the contribution of changes in pollutant emissions and meteorological conditions to changes in ambient concentrations includes:
[0028] The second year or the second month is used as the evaluation year (eval), the first year or the first month is used as the reference year (base), and the model constructed in step S2 is used. Predict the ambient concentration of pollutants under the meteorological conditions of the first year or the first month of the reference year and the emission intensity of the second year or the second month , then the difference between the ambient concentration predicted by the model and the actual monitored concentration in the first year or the first month is the contribution of the change in pollution emission intensity. for:
[0029]
[0030]
[0031] Replace the meteorological data of the second year or the second month with the meteorological data of the first year or the first month, and use the model constructed in step S2 Predict the environmental concentration of pollutants under the meteorological conditions of the first year or the first month and the emission intensity of the second year or the second month , then the difference between the ambient concentration predicted by the model and the actual monitored concentration in the second year or the second month is the contribution of the change in meteorological conditions. for:
[0032]
[0033] Where, is the average annual ambient observation concentration of pollutants in the base year, To assess the annual average environmental monitoring concentration of annual pollutants.
[0034] Based on the above technical solution, the advantages of the present invention are:
[0035] Compared with traditional air quality model assessment methods that rely on pollution source emission inventory data and have high technical difficulty in operation, the technical method constructed by the present invention has simple principles, convenient operation, and is easy to implement. It only requires historical monitoring data of environmental meteorological parameters and air pollutants. It can provide a simple technical method for objectively evaluating the effectiveness of localized air quality management on an hourly time scale, and further provide support technology for decision-making required for the evaluation of environmental air quality standards and the effectiveness of urban and regional air quality management measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0037] Figure 1 A schematic diagram of the steps of the air quality management effectiveness evaluation method of the present invention;
[0038] Figure 2 Schematic diagram of the steps to obtain indicator variables for local source emission intensity of pollutants;
[0039] Figure 3 These are the changes in PM2.5 concentration and the control effectiveness evaluation results from the embodiments of the present invention between 2015 and 2023. DETAILED DESCRIPTION
[0040] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments.
[0041] The present invention provides an air quality management effectiveness evaluation method based on a machine learning algorithm, taking the evaluation of the changes in air pollutant concentrations caused by pollution reduction and meteorological factors over two natural years as an example, which mainly includes the following three steps: Figure 1 As shown, the following steps are included:
[0042] S1. Obtain indicator variables for local source pollutant emission intensity. This method uses the principle of inversely calculating emission intensity from monitored air pollutant concentrations. Air pollutant concentration information includes both emission intensity and meteorological information. By using meteorological standardization techniques to "strip" meteorological information from changes in air pollutant concentrations, information on changes in pollution emission intensity can be obtained.
[0043] Specifically, if Figure 2 As shown, including:
[0044] S101, prepare a model training data set, collect and select the time series data of the six air pollutant concentrations of SO2, NO2, CO, O3, PM10, and PM2.5 monitored hourly in the target area over two years, as well as environmental meteorological parameters and time trend variable data. The time trend variables are used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes.
[0045] Specifically, the study is based on hourly concentration data of six pollutants (SO2, NO2, CO, O3, PM10, and PM2.5) monitored continuously at air quality monitoring stations within the study area, along with synchronized ambient meteorological parameters and time trend variables. Ambient meteorological parameters include ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, trajectory type, and trajectory length. Time trend variables are used to characterize pollutant emissions or atmospheric physical and chemical processes with periodic variations. For example, the day of the Gregorian calendar, the day of the week, and the hourly time series of each day are used to indicate pollution source emissions with seasonal, weekly, and daily variations on an annual time scale. Long-term, interdecadal trends in pollutant emissions are indicated using hourly linearly increasing variables, such as Unix time (a timestamp representing the number of seconds since midnight, January 1, 1970, Greenwich Mean Time). During the evaluation process, the evaluation period can use a whole year's data or a month's hourly average data. In principle, the evaluation period and the benchmark period have the same time span and the amount of data should be as small as possible to meet the training requirements of the machine learning model.
[0046] S102, training the prediction model, for a selected air pollutant, using the regression algorithm to predict the concentration of the pollutant ( ) and environmental meteorological parameters ( ) and time trend variables ( ) was used for modeling, and a pollutant prediction model was trained with the entire hourly data set of two full natural years. The consistency and difference between the pollutant concentrations predicted by the model and the actual monitored concentrations were compared, and the correlation coefficient and root mean square error were calculated to evaluate the model fitting effect.
[0047] in Is the model residual term, for any time t, the model ( ) The predicted pollutant concentration is , then:
[0048]
[0049] ,
[0050] Where i represents the meteorological parameter to be modeled; j represents the time trend variable.
[0051] S103, meteorological standardization. After the pollutant prediction model is trained, meteorological standardization is performed to normalize the interference of meteorological conditions that change over time on pollutant concentrations to the average meteorological conditions, thereby eliminating the disturbance of meteorological conditions that reflects changes in pollution emission intensity. Using the trained random forest pollutant prediction model f, for the pollutant at time t, the concentration of the pollutant is predicted under N randomly selected historical meteorological conditions, and the algebraic mean of the predicted pollutant concentration is taken ( The N randomly selected sets of meteorological data are drawn from historical environmental meteorological data in the original training dataset. According to the law of large numbers and the central limit theorem, when N is large enough, the meteorological perturbation in pollutant concentration at that moment will approach zero. In theory, the time-series changes in the meteorologically standardized concentration of pollutants can reflect changes in pollution emission intensity.
[0052] The existing technology typically uses 300-1000 sets of random meteorological samples. The larger the sample size, the longer the calculation time. To meet timeliness requirements and eliminate meteorological temporal disturbances with as small a historical meteorological sample as possible, this paper proposes using a trained pollutant random forest model to predict the ambient concentration of pollutants at time t under a fixed set of historical meteorological conditions (non-random), such as 24 sets of hourly meteorological conditions on a specific historical day. The algebraic mean of these values is then used as a proxy for emission intensity.
[0053] The present invention normalizes the interference of meteorological conditions that change with time series on pollutant concentrations to the average meteorological conditions, and uses the trained random forest pollutant prediction model f to predict the concentration of pollutants at time t under N randomly selected historical meteorological conditions, and takes the algebraic mean of the predicted pollutant concentrations ( ), that is, the meteorological standardized concentration of pollutants at time t is:
[0054]
[0055] Where k is the kth sample of the selected fixed N groups of historical meteorological data, It is the pollutant concentration predicted by the model under the kth historical meteorological data of the selected fixed N groups.
[0056] S2: Build a machine learning model for pollutants. Use the meteorologically standardized concentration of pollutants as an indicator variable for local pollution emission intensity and incorporate it into the machine learning training of pollutant regression prediction models, including:
[0057] S201, prepare the model training data set, the time series data of the six air pollutants SO2, NO2, CO, O3, PM10, and PM2.5 concentrations monitored hourly in the target area for two years, as well as the environmental meteorological parameters and time trend variable data. The time trend variable ( ) is replaced by the meteorological standardized concentration, i.e. .
[0058] S202, training the prediction model, using regression algorithm to compare the concentration of pollutants with the ambient meteorological parameters ( ) and meteorological normalized concentration ( ) for modeling, is the model residual term, and a pollutant prediction model is retrained with all hourly data sets for two full natural years. For any time t, the model ( ) The predicted pollutant concentration is :
[0059]
[0060] , i represents the listed modeled meteorological parameters;
[0061] S3. Evaluate the contribution of changes in pollution emissions and meteorological conditions to changes in environmental concentrations, determine the baseline reference period for changes in pollutant concentrations in the proposed evaluation period, use the pollutant machine learning model constructed in step S2, simulate pollutant concentrations under different emission and meteorological scenarios based on the control variable method, and obtain the environmental concentration of pollutants under the meteorological conditions of the baseline reference year under the emission intensity of the year to be evaluated. Take the difference between the predicted environmental concentration and the actual monitored concentration in the evaluation year as the contribution, and obtain a quantitative assessment of the change in pollutant concentration caused by pollution emission reduction in the proposed evaluation year compared with the baseline year.
[0062] Preferably, the assessment of the contribution of changes in pollutant emissions and meteorological conditions to changes in ambient concentrations includes:
[0063] The second year or the second month is used as the evaluation year (eval), the first year or the first month is used as the reference year (base), and the model constructed in step S2 is used. Predict the ambient concentration of pollutants under the meteorological conditions of the first year or the first month of the reference year and the emission intensity of the second year or the second month , then the difference between the ambient concentration predicted by the model and the actual monitored concentration in the first year or the first month is the contribution of the change in pollution emission intensity. for:
[0064]
[0065]
[0066] Replace the meteorological data of the second year or the second month with the meteorological data of the first year or the first month, and use the model constructed in step S2 Predict the ambient concentration of pollutants under the meteorological conditions of the first year or the first month and the emission intensity of the second year or the second month , then the difference between the ambient concentration predicted by the model and the actual monitored concentration in the second year or the second month is the contribution of the change in meteorological conditions. for:
[0067]
[0068] Where, is the average annual ambient observation concentration of pollutants in the base year, To assess the annual average environmental monitoring concentration of annual pollutants.
[0069] Preferably, the regression algorithm is a random forest algorithm, a neural network algorithm, a gradient regression tree algorithm or an extreme gradient boosting tree algorithm.
[0070] Preferably, the environmental meteorological parameters include a variety of easily accessible meteorological observation data or meteorological reanalysis data, and the variables used to characterize the emission source strength are any air-related data related to pollution emission activities, including but not limited to motor vehicle traffic, key source pollutant emission monitoring data, etc.
[0071] Current statistical modeling lacks an effective indicator variable for characterizing changes in pollution emission intensity, limiting the application of machine learning algorithms for evaluating pollution control effectiveness. This paper uses the meteorologically standardized concentration of pollutants as an indicator variable for "emission source intensity" and proposes an air quality management effectiveness evaluation method based on a machine learning algorithm, using the control variable approach.
[0072] Taking the PM2.5 continuous monitoring data from an air quality monitoring station in Tianjin from 2015 to 2023 as an example, the specific implementation method includes the following steps:
[0073] Step 101, obtain the meteorologically standardized concentration of PM2.5. Collect and organize the hourly monitoring concentration data of PM2.5 at the site from 2015 to 2023, as well as synchronized environmental meteorological parameters (ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, air mass trajectory category and length) and time trend variables (timestamp, number of lunar days, number of solar days, number of days of the week, and daily hourly time series). The selected environmental meteorological variables may include a variety of easily accessible meteorological observation data or meteorological reanalysis data. The variables used to characterize the intensity of emission sources may be air-related data related to pollution emission activities, including but not limited to motor vehicle traffic and key source pollutant emission monitoring data. After the data collection is completed, the deviation caused by the variation of meteorological covariates in the pollutant concentration time series is adjusted to obtain the meteorologically standardized concentration of the pollutant as the indicator variable for the subsequent "emission intensity" modeling.
[0074] Step 102, replace the time variable contained in the pollutant meteorological standardized modeling data set with the "emission variable", that is, the meteorological standardized concentration of the pollutant. The base year data is selected as a new modeling data set, and a new machine learning random forest modeling is performed to construct the response relationship between emissions, meteorology and pollutant observation concentrations under the base year scenario. The above operations can be implemented in the Python environment through the scikit-learn library machine learning random forest modeling code (or neural network, lightweight gradient boosting machine (LightGBM), extreme gradient boosting tree (XGBoost) and other regression algorithms), based on the root mean square error (RMSE) and correlation coefficient (R 2 ) to judge the quality of model fitting.
[0075] Step 103, use the machine learning model constructed in step 102 to predict the environmental concentration of pollutant emission intensity under the meteorological conditions of the base year 2015 and other years. The emission intensity of the base reference year is replaced by the meteorological standardized concentration (emission intensity) of the corresponding pollutant in each year to be evaluated, and the remaining modeling variables are retained as the base reference year. Based on this, the reconstruction of the data set in the virtual scenario where the emission intensity is based on the evaluation year and the meteorological conditions are based on the base reference year is completed. Based on the reconstructed data, the environmental concentration of the pollutant emission intensity in the year to be evaluated under the meteorological conditions of the base reference year is predicted by the machine learning model constructed above. Under the comparable meteorological conditions of the base year, a quantitative assessment of the change in PM2.5 concentration caused by pollution reduction in the proposed evaluation year compared to the base year is achieved, and the contribution of the "sky-helping" meteorological conditions is also calculated. The final result is as follows. Figure 3 shown.
[0076] exist Figure 3In the study, PM2.5 concentrations decreased by 26.6 micrograms per cubic meter between 2015 and 2023. Human efforts contributed to a 20.3 microgram per cubic meter reduction, accounting for 76% of the decrease, demonstrating significant progress in emission reduction and pollution control. Meteorological factors reduced PM2.5 concentrations by 6.3 micrograms per cubic meter, and weather factors contributed 24% of the decrease. Based on this methodology, a quantitative contribution to the evaluation of air quality management effectiveness was obtained.
[0077] The method presented in this paper is simple in principle and computationally efficient, overcoming the limitations of traditional air quality models, which are limited by the accuracy and timeliness of pollution source emission inventories and the difficulty of computational operations. This method can be directly applied to diverse atmospheric environmental management scenarios, including evaluating the effectiveness of implementing ambient air quality standards, assessing emergency responses to heavy pollution events, and assessing air quality assurance for major national events, thus possessing significant application and promotional value.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or some technical features can be replaced by equivalents without departing from the spirit of the technical solution of the present invention. They should all be included in the scope of the technical solution for protection of the present invention.
Claims
1. A method for evaluating air quality management effectiveness based on a machine learning algorithm, characterized by: The steps include: S1, obtain the indicator variables of local source emission intensity of pollutants, including: S101, prepare a model training dataset by collecting and selecting time series data on the concentrations of six air pollutants, SO2, NO2, CO, O3, PM10, and PM2.5, as well as environmental meteorological parameters and time trend variable data, which are monitored hourly in the target area over a two-year period. The time trend variable is used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes. S102, training the prediction model, for a selected air pollutant, using the regression algorithm to predict the concentration of the pollutant at the time t under study. and environmental meteorological parameters and time trend variables Modeling was conducted to train a pollutant prediction model using the full hourly data set of two full calendar years. The consistency and difference between the pollutant concentrations predicted by the model and the actual monitored concentrations were compared, and the correlation coefficient and root mean square error were calculated to evaluate the model fitting effect. in is the model residual term. For any time t, the model The predicted pollutant concentration is , then: ; ; In the formula represents the modeled ambient meteorological parameters; represents the time trend variable; T represents the ground temperature, RH represents the relative humidity, WS represents the wind speed, WD represents the wind direction, SP represents the air pressure, SSR represents the radiation intensity, BLH represents the height of the mixing layer, TCC represents the total cloud cover, Prep represents the precipitation, Air cluster Indicates the trajectory category, Air length Indicates the trajectory length; Hour indicates the daily hour sequence, Day-of-Week indicates the day of the week, Day-of-Year indicates the day of the year in the Gregorian calendar, LunarDay indicates the day of the year in the lunar calendar, and Unix time indicates the timestamp; S103, meteorological standardization processing, normalizes the interference of meteorological conditions that change with time series on pollutant concentrations to the average meteorological conditions, uses the trained random forest pollutant prediction model f, and predicts the concentration of pollutants at time t under N randomly selected historical meteorological conditions, taking the algebraic mean of the predicted pollutant concentrations , that is, the meteorological standardized concentration of pollutants at time t is: Where k is the kth sample of the selected fixed N groups of historical meteorological data, is the pollutant concentration predicted by the model under the kth historical meteorological data of the selected fixed N groups; S2: Build a machine learning model for pollutants. Use the meteorologically standardized concentration of pollutants as an indicator variable for local pollution emission intensity and incorporate it into the machine learning training of pollutant regression prediction models, including: S201: Prepare the model training dataset, which includes the hourly monitoring data of six air pollutants, SO2, NO2, CO, O3, PM10, and PM2.5, as well as environmental meteorological parameters and time trend variable data, over the past two years in the target area. S202, training the prediction model, using regression algorithm to compare the concentration of pollutants with the ambient meteorological parameters for a selected air pollutant and meteorological normalized concentration Modeling, is the model residual term. A pollutant prediction model is retrained with all hourly data sets for two full natural years. For any time t, the model The predicted pollutant concentration is : ; ; In the formula Represents the modeled environmental meteorological parameters, T represents ground temperature, RH represents relative humidity, WS represents wind speed, WD represents wind direction, SP represents air pressure, SSR represents radiation intensity, BLH represents mixed layer height, TCC represents total cloud cover, Prep represents precipitation, Air cluster Indicates the trajectory category, Air length represents the trajectory length; S3. Evaluate the contribution of changes in pollution emissions and meteorological conditions to changes in environmental concentrations, determine the baseline reference period for changes in pollutant concentrations in the proposed evaluation period, use the pollutant machine learning model constructed in step S2, simulate pollutant concentrations under different emission and meteorological scenarios based on the control variable method, and obtain the environmental concentration of pollutants under the meteorological conditions of the baseline reference year under the emission intensity of the year to be evaluated. Take the difference between the predicted environmental concentration and the actual monitored concentration in the evaluation year as the contribution, and obtain a quantitative assessment of the change in pollutant concentration caused by pollution emission reduction in the proposed evaluation year compared with the baseline year.
2. The air quality management effectiveness evaluation method according to claim 1, characterized in that: The regression algorithm is a random forest algorithm, a neural network algorithm, a gradient regression tree algorithm or an extreme gradient boosting tree algorithm.
3. The air quality management effectiveness evaluation method according to claim 1, characterized in that: The environmental meteorological parameters include a variety of easily accessible meteorological observation data or meteorological reanalysis data, and the variables used to characterize the emission source intensity are any air-related data related to pollution emission activities.
4. The air quality management effectiveness evaluation method according to claim 1, characterized in that: Assessment of the contribution of changes in pollutant emissions and meteorological conditions to changes in ambient concentrations includes: The second year or the second month is used as the evaluation year (eval), the first year or the first month is used as the reference year (base), and the model constructed in step S2 is used. Predict the ambient concentration of pollutants under the meteorological conditions of the first year or the first month of the reference year and the emission intensity of the second year or the second month , then the difference between the ambient concentration predicted by the model and the actual monitored concentration in the first year or the first month is the contribution of the change in pollution emission intensity. for: ; Replace the meteorological data of the second year or the second month with the meteorological data of the first year or the first month, and use the model constructed in step S2 Predict the ambient concentration of pollutants under the meteorological conditions of the first year or the first month and the emission intensity of the second year or the second month , then the difference between the ambient concentration predicted by the model and the actual monitored concentration in the second year or the second month is the contribution of the change in meteorological conditions. for: ; Where, is the average annual ambient observation concentration of pollutants in the base year, To assess the annual average environmental monitoring concentration of annual pollutants.
Citation Information
Patent Citations
Pollutant concentration influence assessment method and device, storage medium and electronic equipment
CN117517581A