Pollution source strong dynamic inversion method based on air quality machine learning rolling prediction model
Through the strong dynamic inversion method of pollution source based on the air quality machine learning rolling prediction model, the shortcomings of the existing technology under the demand for high-time management are solved, real-time rolling prediction and dynamic inversion of pollutants are achieved, and the accuracy and timeliness of heavy pollution emergency assessment are improved.
Patent Information
- Application Number
- CN202510163961.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-14
Smart Images

Figure CN120104983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to air pollution prevention and control technology, and in particular to a pollution source strong dynamic inversion method based on an air quality machine learning rolling prediction model. Background Art
[0003] The traditional method of assessing the emission intensity of pollution sources is to analyze the ambient air quality monitoring data, and to adjust the deviation caused by the variation of meteorological covariates in the pollutant concentration time series by selecting appropriate mathematical statistics methods, so as to reveal the contribution of pollution sources from the changes in air quality. There are technical difficulties in the practical application of the traditional mathematical statistics assessment method: 1) Classical mathematical statistics has relatively strict parameter test conditions, and the data of atmospheric environment monitoring usually does not fully meet the requirements of Gaussian normal distribution and homogeneity of variance; 2) Meteorological factors and air pollution have complex nonlinear effects, and the fitting performance of classical statistical models for complex nonlinear problems is insufficient.
[0004] With the rapid development of artificial intelligence machine learning algorithms, a series of new algorithms suitable for modeling complex problems have been continuously proposed. David K. Carslaw and other scholars from York University in the UK have developed a meteorological standardization technology based on random forest algorithm modeling to separate meteorological and emission information in air quality changes. This technical method assumes that the emission of air pollution has certain periodic characteristics. It uses machine learning algorithms such as random forests with better fitting performance to fit and train historical data such as meteorological parameters and pollution source intensity periodic characterization variables to train its regression model with the monitoring concentration of atmospheric pollutants. In order to reduce the interference of meteorological factors on the monitoring concentration of pollutants at different times, at any time, several groups of meteorological data are randomly selected from the historical data, and the trained model is used to predict the environmental concentration of pollutants under these historical meteorological conditions. The concentration of pollutants under these historical meteorological conditions at each moment in the study period is calculated. The time series changes of the arithmetic mean of the multiple predicted concentrations can reflect the changes in emission intensity, and the pollutants at each moment are standardized under average meteorological conditions. This technical method makes up for the shortcomings and limitations of the classical statistical evaluation method. At present, this technology has been widely used in the trend research of air pollution emission changes in hundreds of cities around the world.
[0005] However, in addition to normalized control measures, air pollution prevention and control also includes temporary and sudden control measures, such as emergency response to heavy pollution weather and air quality assurance for major events. These temporary controls have greatly increased the timeliness requirements, and the current meteorological standardization technology based on machine learning modeling is mainly used for post-analysis of scenarios such as policy intervention. The analysis is time-consuming and difficult to automate, and the analysis results are delayed and cannot effectively support high-timeliness decision-making and management demand scenarios. Summary of the invention
[0006] In order to meet the above-mentioned high-timeliness management needs, the present invention proposes a dynamic inversion method for pollution source strength based on an air quality machine learning rolling prediction model. By designing a machine learning pollutant rolling prediction model and designing an algorithm to effectively control the disturbance of changes in meteorological conditions, it is possible to quickly invert the changes in local pollutant emission source strength from real-time air quality monitoring data, so as to provide timely and efficient automated feedback evaluation for environmental management scenarios such as heavy pollution emergency control effects.
[0007] To achieve the above object, the present invention provides a pollution source strong dynamic inversion method based on an air quality machine learning rolling prediction model, and the pollution source strong dynamic inversion method comprises the following steps:
[0008] Step S1, acquisition and processing of air quality modeling dataset: based on the hourly continuous monitoring of SO2 by air quality monitoring stations in the study area, 2 、NO 2 ,CO,O 3 , PM10 and PM2.5 concentration data, as well as synchronous environmental meteorological parameters and time trend variables, the time trend variables are used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes;
[0009] Step S2, construction of air quality machine learning prediction model: for any air pollutant, at the studied time t, use machine learning regression algorithm to compare pollutant concentration C with environmental meteorological parameters M i and the time trend variable T j Modeling was performed to compare the consistency and difference between the pollutant concentrations predicted by the model and the actual monitored concentrations, and the correlation coefficient and root mean square error were calculated to evaluate the model fitting effect;
[0010] where ∈ t is the model residual term, and the pollutant prediction model trained at the current time t is f t :
[0011] C t =f t (M i,t ,T j,t )+∈ t ,
[0012] i∈(T,RH,BLH,…,I); j∈(hour,day,…,Unix time),
[0013] Where i represents the meteorological parameters of the model, including temperature T, relative humidity RH, boundary layer height BLH, etc.; j represents the time trend variable, including the daily hourly time series hour, the day of the Gregorian calendar date day, the linear trend variable Unix time, etc.;
[0014] Step S3, meteorological standardization assessment of air quality: the trained model is used to predict the pollutants at time t under a fixed set of historical meteorological conditions, replacing the mean value as the representation of emission intensity, that is, the meteorological standardization concentration C of pollutants at time t t,norm for:
[0015]
[0016] In the formula, k is the kth sample of the selected fixed N groups of historical meteorological data, is the pollutant concentration predicted by the model under the kth historical meteorological data of the selected fixed N groups;
[0017] Step S4, air quality rolling forecast and meteorological standardization iterative calculation:
[0018] For the pollutant monitoring concentration at the next time t+1, after the monitoring data collection is completed, the pollutant, meteorological parameters and time variable data at time t+1 are included as a new sample in the original training data set of the same length, and steps S2 and S3 are repeated using the new data set. The new training model f at time t+1 is used. t+1 The meteorological standardized concentrations of pollutants at time t and t+1 are calculated respectively, and the N groups of historical meteorological conditions remain unchanged; as the new monitoring data continues to synchronize with time, the steps of t+1 are repeated for time t+2, and the new data at time t+2 are included to retrain a pollutant prediction model f t+2 , using the model f t+2 Update the meteorological standardized concentration of pollutants at time t, t+1, and t+2, and the newly trained model f t+2 It is used to predict the concentration of various pollutants under fixed N groups of weather conditions, and discard the previous round of f t+1 The calculation results are repeated until the rolling iterative calculation is performed at each moment.
[0019] Preferably, the environmental meteorological parameters include ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, trajectory category and trajectory length.
[0020] Preferably, the time trend variables include a timestamp, a lunar day, a solar day, a day of the week, and a daily hour sequence.
[0021] Preferably, the regression algorithm is a random forest algorithm, a neural network algorithm, a gradient regression tree algorithm or an extreme gradient boosting tree algorithm.
[0022] Preferably, the environmental meteorological parameter M i It includes a variety of easily accessible meteorological observation data or meteorological reanalysis data, and the variables used to characterize the emission source intensity are any air-related data related to pollution emission activities.
[0023] Preferably, in step S4, the fixed meteorological conditions selected for the fixed meteorological standardization are N groups of meteorological conditions fixed within any period of time.
[0024] Based on the above technical solution, the advantages of the present invention are:
[0025] The pollution source strong dynamic inversion method based on the air quality machine learning rolling prediction model of the present invention realizes the rolling prediction of pollutants through hourly iterative modeling, and uses each round of newly trained pollutant prediction model to normalize the pollutants in the study period (such as heavy pollution process) to fixed preset meteorological conditions, thereby realizing dynamic inversion under the premise of ensuring that the meteorological normalized pollutant concentrations are comparable in time series.
[0026] The pollution source intensity dynamic inversion method of the present invention not only solves the problem that the existing meteorological normalization technology is limited by insufficient timeliness and is mainly used in post-analysis of policy interventions, but also reduces the biased estimation caused by random meteorological normalization, thereby improving the accuracy of the results. It is particularly suitable for environmental management application scenarios with high timeliness requirements such as emergency assessment of heavy pollution weather, and has important application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0028] Figure 1 It is a schematic diagram of the process flow of the strong dynamic inversion method of pollution sources of the present invention;
[0029] Figure 2 is SO in Example 1 2 The root mean square error and correlation coefficient of the concentration output model fitting data;
[0030] Figure 3 The SO after establishing the infinite loop modeling and prediction system in Example 1 2 Hourly concentration data. DETAILED DESCRIPTION
[0031] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments.
[0032] The present invention provides a method for strong dynamic inversion of pollution sources based on an air quality machine learning rolling prediction model, such as Figure 1 As shown, the pollution source strong dynamic inversion method includes the following steps:
[0033] Step S1, acquisition and processing of air quality modeling dataset: based on the hourly continuous monitoring of SO2 by air quality monitoring stations in the study area, 2 、NO 2 ,CO,O 3 , PM10 and PM2.5, as well as synchronized environmental meteorological parameters and time trend variables. The time trend variables are used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes.
[0034] Specifically, based on the six pollutants (SO 2 、NO 2 ,CO,O 3 , PM10, PM2.5) concentration data, as well as synchronized environmental meteorological parameters and time trend variables. Environmental meteorological parameters include ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, trajectory category and trajectory length. Time trend variables are used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes. For example, the number of days in the Gregorian calendar, the number of days of the week, the hourly time sequence of each day, etc., indicate the pollution source emission activities with seasonal, weekly-daily and daily variation patterns within the annual time scale. The long-term interdecadal pollution emission change trend is indicated by a linearly increasing variable hour by hour, such as Unix time (timestamp, that is, the accumulated seconds since midnight on January 1, 1970, Greenwich Mean Time). The data set used to train the pollutant model is based on the basic requirement of no less than 1 month of continuous monitoring data, and the longer the coverage time, the better.
[0035] Step S2, construction of air quality machine learning prediction model: for the six pollutants, at the studied time t, the concentration C of the pollutant is calculated one by one using regression algorithm t and environmental meteorological parameters M i and the time trend variable T j Modeling is performed to compare the consistency and difference between the pollutant concentration predicted by the model and the actual monitored concentration, and the correlation coefficient and root mean square error are calculated to evaluate the model fitting effect; preferably, the regression algorithm is a random forest algorithm, a neural network algorithm, a gradient regression tree algorithm or an extreme gradient boosting tree algorithm.
[0036] where ∈ t is the model residual term, and the pollutant prediction model trained at the current time t is f t:
[0037] C t =f t (M i,t ,T j,t )+∈ t ,
[0038] i∈(T,RH,BLH,…,I); j∈(hour,day,…,Unix time),
[0039] Where i represents the meteorological parameters of the model, including temperature T, relative humidity RH, boundary layer height BLH, etc.; j represents the time trend variable, including the daily hourly time series hour, the day of the Gregorian calendar date day, the linear trend variable Unix time, etc.
[0040] Taking a full year of monitoring data as an example, the pollutant prediction model is trained with all hourly data sets. Considering the timeliness requirements of practical applications, when the training sample is large enough (for example, more than a full year), it is not required to report the generalization ability of the model on the test set, but the root mean square error and correlation coefficient of the model fitting all the data are required.
[0041] The environmental meteorological parameters used in the modeling may include a variety of easily accessible meteorological observation data or meteorological reanalysis data, such as ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, trajectory category and trajectory length, etc. The variables used to characterize the emission source intensity can be any air-related data related to pollution emission activities, including but not limited to motor vehicle flow, key source pollutant emission monitoring data, etc.
[0042] Step S3, meteorological standardization assessment of air quality: the trained model is used to predict the pollutants at time t under a fixed set of historical meteorological conditions, replacing the numerical mean as the representation of emission intensity, that is, the meteorological standardization concentration of pollutants at time t is:
[0043]
[0044] In the formula, k is the kth sample of the selected fixed N groups of historical meteorological data, It is the pollutant concentration predicted by the model under the kth historical meteorological data of the selected fixed N groups.
[0045] Specifically, the purpose of meteorological standardization is to standardize the interference of meteorological conditions that change with time to pollutant concentrations to average meteorological conditions, so as to eliminate the disturbance of meteorological conditions that reflects the change of pollution emission intensity. The existing technology uses the trained random forest pollutant prediction model f to predict the concentration of pollutants at time t under N randomly selected historical meteorological conditions, and takes the algebraic mean C of the predicted pollutant concentration. t,norm . The N randomly selected groups of meteorological data are extracted from the historical environmental meteorological data in the original training data set. According to the law of large numbers and the central limit theorem, when N is large enough, the meteorological disturbance in the pollutant concentration at that moment will tend to 0. In theory, the change in the meteorological standardized concentration of pollutants in time series can reflect the change in the intensity of pollution emissions. The more common number of random meteorological samples is 300 to 1000 groups. The larger the meteorological samples extracted, the longer the calculation time. At present, this meteorological standardization method still has shortcomings, that is, the randomly extracted historical meteorological data introduces random errors, resulting in the risk of inconsistency of the randomly extracted "average meteorological conditions" at different times; when the extracted meteorological samples are large enough to eliminate random risks, the cost of computing resources is greatly increased.
[0046] In order to meet the requirements of timeliness and eliminate meteorological disturbances in time series with as small historical meteorological samples as possible, the present invention proposes to use the trained pollutant model to predict the environmental concentration of pollutants at time t under a fixed set of historical meteorological conditions (non-random), such as 24 sets of hourly meteorological conditions fixed on a certain historical day, and take the algebraic mean as the representation of emission intensity.
[0047] Step S4, air quality rolling forecast and meteorological standardization iterative calculation:
[0048] For the pollutant monitoring concentration at the next moment t+1, after the monitoring data collection is completed, the pollutant, meteorological parameters and time variable data at moment t+1 are included as a new sample in the original training data set of the same length (such as a complete natural year) (remove the first sample of the original data set).
[0049] Repeat steps S2 and S3 using the new data set, and use the new training model f at time t+1 t+1 The meteorological standardized concentrations of pollutants at time t and t+1 are calculated respectively, and the historical meteorological conditions of group K remain unchanged; then the model f used to describe the relationship between pollutant concentration, emission and meteorological mapping at time t+1 is t+1 for:
[0050] C t+1 =f t+1 (M i,t+1 ,t j,t+1 )+∈ t+1 ,
[0051] i∈(T,RH,BLH,…,I); j∈(hour,day,…,Unix time)
[0052] Using the new training model f at time t+1 t+1 The meteorological standardized concentrations of pollutants at time t and time t+1 are calculated respectively, and the historical meteorological conditions of group K can remain unchanged:
[0053]
[0054] As new monitoring data continues to be synchronized over time, the steps of t+1 are repeated for time t+2, and a pollutant prediction model f is retrained by incorporating the new data at time t+2. t+2 , using the model f t+2 Update the meteorological standardized concentration of pollutants at time t, t+1, and t+2, and the newly trained model f t+2 It is used to predict the concentration of various pollutants under fixed N groups of meteorological conditions. Therefore, the meteorological standardization results of the new prediction model constructed in each round are comparable in time series. At this time, the previous round of f t+1 The calculation results are repeated until the rolling iterative calculation is performed at each moment.
[0055] Preferably, in the fixed meteorological standardization, the selected fixed meteorological conditions are N groups of meteorological conditions that are fixed within any period of time.
[0056] The pollution source strong dynamic inversion method based on the air quality machine learning rolling prediction model of the present invention realizes the rolling prediction of pollutants through hourly iterative modeling, and uses each round of newly trained pollutant prediction model to normalize the pollutants in the study period (such as heavy pollution process) to fixed preset meteorological conditions, thereby realizing dynamic inversion under the premise of ensuring that the meteorological normalized pollutant concentrations are comparable in time series.
[0057] The pollution source intensity dynamic inversion method of the present invention not only solves the problem that the existing meteorological normalization technology is limited by insufficient timeliness and is mainly used in post-analysis of policy interventions, but also reduces the biased estimation caused by random meteorological normalization, thereby improving the accuracy of the results. It is particularly suitable for environmental management application scenarios with high timeliness requirements such as emergency assessment of heavy pollution weather, and has important application value.
[0058] Example 1
[0059] The present invention can be implemented through program software operation. Taking the rolling prediction of PM2.5 concentration hourly monitoring data in Tianjin as an example to invert emission intensity, the specific implementation method includes the following steps:
[0060] Step 101, based on the hourly SO2 concentration observation data of Tianjin from 2015 to date, as well as the synchronized environmental meteorological parameters (ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, trajectory category, trajectory length) and time trend variables (timestamp, number of days in the lunar calendar, number of days in the solar calendar, number of days in the week, hourly time series of each day), Python code is used to complete data preprocessing (missing values, abnormal values, etc.) and loading. The selected environmental meteorological variables may include a variety of easily accessible meteorological observation data or meteorological reanalysis data, and the variables used to characterize the emission source strength may be air-related data related to pollution emission activities, including but not limited to motor vehicle flow, key source pollutant emission monitoring data, etc.
[0061] Step 102, for SO 2 hourly concentration, based on the python code, the random forest algorithm is used to analyze the SO 2 The concentration is modeled with environmental meteorological parameters and time trend variables, and the root mean square error (RMSE) and correlation coefficient (R2) of the model fitting data are output, such as Figure 2 The algorithm used in the modeling process is random forest, and it can also be a regression algorithm such as neural network, gradient regression tree, extreme gradient boosting tree (XGBoost), etc.
[0062] Step 103, based on the Python code, filter out all dates with complete 24-hour data for one day, randomly select 24 groups (hours) of complete meteorological data for a certain day in history, and predict SO at time t based on the random forest model. 2 24 sets of ambient concentrations under the fixed meteorological conditions are used to characterize the emission intensity at that moment by replacing the average value.
[0063] Step 104, establish an infinite loop modeling and prediction system, such as completing the processing of steps 101 to 103 at time t+1, and outputting the corresponding data and image files of the corresponding pollutant observation concentration and emission intensity in the last 7 days, and iterating every hour, such as Figure 3 As shown, the output at 19:00 on August 31, 2024 is the SO from 19:00 on August 24, 2024 to 18:00 on August 31, 2024. 2 Hourly concentration data, solid line is the monitored concentration (Observed SO 2 ), the dotted line is the predicted concentration after fixed meteorological standardization (Predicted SO 2 ).
[0064] like Figure 3 As shown in the figure, the solid line is the monitored environmental concentration, which is greatly affected by the weather and has a large fluctuation. After the meteorological standardization of the present invention, the dotted line SO2 The volatility is significantly reduced, and it shows a stable change every day, which reflects the stable daily cycle change of pollution emission intensity. From August 27 to August 29, SO 2 Taking the phenomenon of continuous increase as an example, it can be found that SO 2 This increase in environmental concentration is not caused by an increase in emissions from pollution sources, but rather by unfavorable meteorological conditions causing SO 2 Pollution accumulates. This verifies the feasibility of the method of the present invention.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or some technical features can be replaced by equivalents without departing from the spirit of the technical solution of the present invention, which should be included in the scope of the technical solution for protection of the present invention.
Claims
1. A strong dynamic inversion method for pollution sources based on an air quality machine learning rolling prediction model, characterized by: The pollution source strong dynamic inversion method comprises the following steps: Step S1, acquisition and processing of air quality modeling data set: based on the six pollutant concentration data of SO2, NO2, CO, O3, PM10 and PM2.5 continuously monitored hourly by air quality monitoring stations in the study area, as well as synchronized environmental meteorological parameters and time trend variables, the time trend variables are used to characterize pollution emissions or atmospheric physical and chemical processes with periodic changes; Step S2, construction of air quality machine learning prediction model: for the six pollutants, at the studied time t, the concentration of pollutants (C t ) and environmental meteorological parameters (M i ) and time trend variables (T j ) to build a model, compare the consistency and difference between the pollutant concentrations predicted by the model and the actual monitored concentrations, and calculate the correlation coefficient and root mean square error to evaluate the model fitting effect; where ∈ t is the model residual term, and the pollutant prediction model trained at the current time t is (f t ): C t =f t (M i ,t,T j ,t)+∈ t , i∈(T,RH,BLH,…,I); j∈(hour,day,…,Unix time), Where i represents the meteorological parameter to be modeled; j represents the time trend variable; Step S3, meteorological standardization assessment of air quality: the trained model is used to predict the pollutants at time t under a fixed set of historical meteorological conditions, replacing the numerical mean as the representation of emission intensity, that is, the meteorological standardization concentration of pollutants at time t is: In the formula, k is the kth sample of the selected fixed N groups of historical meteorological data, It is the pollutant concentration predicted by the model under the kth historical meteorological data of the selected fixed N groups. Step S4, air quality rolling forecast and meteorological standardization iterative calculation: For the pollutant monitoring concentration at the next time t+1, after the monitoring data collection is completed, the pollutant, meteorological parameters and time variable data at time t+1 are included as a new sample in the original training data set of the same length, and steps S2 and S3 are repeated using the new data set. The new training model f at time t+1 is used. t+1 The meteorological standardized concentrations of pollutants at time t and t+1 are calculated respectively, and the historical meteorological conditions of group K remain unchanged; as the new monitoring data continues to synchronize with time, the steps of t+1 are repeated for time t+2, and the new data at time t+2 are included to retrain a pollutant prediction model f t+2 , using the model f t+2 Update the meteorological standardized concentration of pollutants at time t, t+1, and t+2, and the newly trained model f t+2 It is used to predict the concentration of various pollutants under fixed N groups of weather conditions, and discard the previous round of f t+1 The calculation results are repeated until the rolling iterative calculation is performed at each moment.
2. The pollution source strong dynamic inversion method according to claim 1 is characterized in that: The environmental meteorological parameters include ground temperature, relative humidity, wind speed, wind direction, air pressure, radiation intensity, mixing layer height, total cloud cover, precipitation, track category and track length.
3. The pollution source strong dynamic inversion method according to claim 1 is characterized in that: The time trend variables include a timestamp, a lunar day, a solar day, a day of the week, and a daily hour sequence.
4. The pollution source strong dynamic inversion method according to claim 1 is characterized in that: The regression algorithm is a random forest algorithm, a neural network algorithm, a gradient regression tree algorithm or an extreme gradient boosting tree algorithm.
5. The pollution source strong dynamic inversion method according to claim 2 is characterized in that: The environmental meteorological parameters include a variety of easily accessible meteorological observation data or meteorological reanalysis data, and the variables used to characterize the emission source intensity are any air-related data related to pollution emission activities.
6. The pollution source strong dynamic inversion method according to claim 1 is characterized by: In step S4, the fixed meteorological conditions selected for fixed meteorological standardization are N groups of meteorological conditions fixed within any period of time.
Citation Information
Patent Citations
Ozone concentration prediction method and system based on spatio-temporal data and statistical learning
CN107943928A
Air quality forecasting method based on dynamic inversion of emission source data
CN110824107A
Atmospheric pollution emission inversion method, system and equipment based on machine learning
CN115481558A
Small-scale near-real-time atmospheric pollution tracing method based on multi-method fusion
CN116011317A
VOCs emission source list dynamic inversion method based on four-dimensional variation assimilation
CN116415408A
Cited By
Atmospheric pollutant formation mechanism analysis method and device
CN120708767A
Air quality data screening method based on data consistency check
CN120724174A
Air quality multi-model forecast result screening method based on sliding optimal matching
CN120724175A