Air pollutant concentration dynamic hourly prediction method based on machine learning model

By using machine learning models and a wide range of meteorological factors, the problems of high data demand and low accuracy in the air pollutant concentration prediction method are solved, and dynamic time-by-time prediction is achieved, which reduces costs and improves prediction accuracy. It is suitable for dynamic pollution control and heavy pollution warning.

CN120277354APending Publication Date: 2025-07-08CHANGZHOU ENVIRONMENTAL MONITORING CENT
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510319499.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing air pollutant concentration prediction methods require a large amount of air pollutant emission source data, which is costly to implement and incomplete selection of meteorological factors, resulting in low prediction accuracy and inability to achieve dynamic time-by-time prediction.

Method used

Using a machine learning model-based method, a wide range of meteorological factors such as near-ground wind field, temperature, radiation, etc., supplement data loss through exponential smoothing method, randomly divide training and test sample sets, and train using XGBoost or neural network model to achieve dynamic time-by-time prediction of air pollutant concentration.

Benefits of technology

It realizes high-precision dynamic time-by-time prediction without air pollutant emission source data, reduces implementation costs, is suitable for personal computer operations, is suitable for dynamic pollution control and heavy pollution warning, and improves the accuracy and efficiency of air quality prediction.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses an air pollutant concentration dynamic hourly prediction method based on a machine learning model, and the method comprises the steps: a, extracting historical data, report starting data and prediction data: the historical data comprises meteorological mode historical hour data and air pollutant concentration historical hour data in N months before the report starting time of a certain year, a certain month and a certain day in a certain region, the report starting data comprises air pollutant concentration data at a report starting moment, the prediction data comprises meteorological mode prediction hour data at x moments after the report starting moment, N is greater than or equal to 6, and x is greater than or equal to 2; b, processing the historical data and the prediction data; c, determining the trained machine learning model as an air pollutant concentration prediction model; and d, inputting meteorological mode prediction hour data at x moments after the report starting moment into the air pollutant concentration prediction model to obtain dynamic hourly prediction data of the air pollutant concentration. According to the method, air pollutant emission source data is not needed, the used meteorological factors are comprehensive and reasonable, the prediction precision is high, and dynamic hourly prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting pollutant concentration, and particularly to a dynamic hourly prediction method for air pollutant concentration based on a machine learning model, belonging to the technical field of environmental monitoring. Background Technique

[0002] In recent years, with the in-depth promotion of air pollution control, the requirements for pollution control have become gradually refined, and short-term air quality forecasting has received increasing attention. Short-term forecasting of air pollutant concentration has great application prospects in supporting dynamic pollution control, starting and lifting of heavy pollution warnings, and ensuring air quality for major events.

[0003] With the development of artificial intelligence technology, machine learning algorithms based on statistical theory such as machine learning models have been gradually applied to the prediction of ambient air quality. The advantages of machine learning models in prediction models of multi-dimensional data have been continuously verified, showing the advantages of being simple and easy to implement, saving manpower and material resources, and having a short operation time. However, the existing methods for predicting air pollutant concentration usually have the following three disadvantages in practical applications: on the one hand, they have high requirements for computing resources, need to obtain a large amount of air pollutant emission source data, and have a high business implementation cost and a long cycle; on the other hand, the existing prediction methods mainly use traditional meteorological factors such as temperature, wind speed, wind direction, humidity, and precipitation, and do not use meteorological factors such as radiation amount and boundary layer height, resulting in low prediction accuracy and still room for improvement; on the other hand, they cannot achieve dynamic hourly prediction of air pollutant concentration. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a dynamic hourly prediction method for air pollutant concentration based on a machine learning model that does not require air pollutant emission source data, uses comprehensive and reasonable meteorological factors, has a low implementation cost, high prediction accuracy, and can achieve dynamic hourly prediction.

[0005] To solve the above technical problem, the present invention adopts such a dynamic hourly prediction method for air pollutant concentration based on a machine learning model to obtain dynamic hourly prediction data of air pollutant concentration at x moments after the starting time, including the following steps:

[0006] a. Extract historical data, initial reporting data, and prediction data: The historical data includes historical hourly meteorological pattern data and historical hourly air pollutant concentration data for the N months before the initial reporting time on a certain date, month, and year in a certain region. The initial reporting data includes the air pollutant concentration data at the initial reporting time. The prediction data includes the predicted hourly meteorological pattern data for x moments after the initial reporting time, where N is a natural number ≥ 6 and x is a natural number ≥ 2. The historical hourly meteorological pattern data includes the 10-meter u-wind field near the ground, the 10-meter v-wind field near the ground, the 2-meter temperature near the ground, the boundary layer height, the shortwave radiation amount, the relative humidity, the precipitation, the air pressure value, the specific humidity value, the u-wind field at 925 hPa, the v-wind field at 925 hPa, the u-wind field at 850 hPa, the v-wind field at 850 hPa, the u-wind field at 700 hPa, the v-wind field at 700 hPa, the temperature difference between 975 hPa and 950 hPa, the temperature difference between 950 hPa and 925 hPa, the temperature difference between 925 hPa and 900 hPa, the temperature difference between 900 hPa and 850 hPa, the average cloud amount, the 10-meter wind speed near the ground, and the 10-meter wind direction near the ground;

[0007] b. Process the historical data and prediction data: First, check whether the historical hourly meteorological pattern data and historical hourly air pollutant concentration data for the N months before the initial reporting time are complete. If there are incomplete historical missing data, use the exponential smoothing method ES to supplement and process the historical missing data to obtain complete historical hourly meteorological pattern data and historical hourly air pollutant concentration data, and check whether the predicted hourly meteorological pattern data for x moments after the initial reporting time is complete. If there are incomplete predicted missing data, use the exponential smoothing method ES to supplement and process the predicted missing data to obtain complete predicted hourly meteorological pattern data. Then, perform corresponding matching and association processing on the complete historical hourly meteorological pattern data and historical hourly air pollutant concentration data according to the same date and the same moment to obtain the historical hourly matching data of the meteorological pattern and air pollutant concentration;

[0008] c. Randomly divide the historical hourly matching data of the meteorological pattern and air pollutant concentration into a training sample set and a test sample set according to a preset ratio. Use the training sample set to train a machine learning model to obtain a trained machine learning model. Input the test sample set into the trained machine learning model for detection. When the difference between each historical hourly air pollutant concentration data input from the test sample set and the predicted hourly air pollutant concentration data output by the trained machine learning model is within a preset range, determine the trained machine learning model as the air pollutant concentration prediction model;

[0009] d. Input the meteorological model prediction hourly data at x moments after the starting time of step a into the air pollutant concentration prediction model to obtain the dynamic hourly prediction data of air pollutant concentration at x moments after the starting time. Specifically: T0 represents the starting time, and T i (i = 1, 2,..., x) represents the i-th moment after the starting time, where T1 represents the 1st moment after the starting time, T2 represents the 2nd moment after the starting time, and T x represents the x-th moment after the starting time. The prediction method for the air pollutant concentration at T1 moment is to input the meteorological model prediction hourly data at the 1st moment T1 after the starting time and the air pollutant concentration data at the starting time T0 into the air pollutant concentration prediction model to obtain the predicted hourly data of air pollutant concentration at T1 moment; the prediction method for the air pollutant concentration at T2 moment is to input the meteorological model prediction hourly data at the 2nd moment T2 after the starting time and the predicted hourly data of air pollutant concentration at T1 moment into the air pollutant concentration prediction model to obtain the predicted hourly data of air pollutant concentration at T2 moment; the prediction method for the air pollutant concentration at T x moment is to input the meteorological model prediction hourly data at the x-th moment T x after the starting time and the predicted hourly data of air pollutant concentration at T x-1 moment into the air pollutant concentration prediction model to obtain the predicted hourly data of air pollutant concentration at T x moment; and so on, to obtain the dynamic hourly prediction data of air pollutant concentration at x moments after the starting time.

[0010] As a preferred implementation of the present invention, in step a, the historical data and the starting data are extracted from the database of a regional multi-parameter atmospheric station, and the prediction data is extracted from the mesoscale atmospheric model WRF; the historical data includes the meteorological model historical hourly data and the air pollutant concentration historical hourly data in the 8 - 12 months before the starting time on a certain date in a certain region, and the prediction data includes the meteorological model prediction hourly data at 72 - 168 moments after the starting time; the air pollutants include PM 2.5 , O3, PM 10 , NO2, SO2, and VOCs.

[0011] As a preferred embodiment of the present invention, in step c, the meteorological pattern and the historical hourly matching data of air pollutant concentration are randomly divided into a training sample set and a test sample set according to a ratio of 8:2 or 7:3. The training sample set and the test sample set are subjected to dimensionless processing. The machine learning model is trained using the dimensionless processed training sample set to obtain a trained machine learning model. The dimensionless processed test sample set is input into the trained machine learning model for detection. Until the difference between the historical hourly data of air pollutant concentration input from each dimensionless processed test sample set and the predicted hourly data of air pollutant concentration output by the trained machine learning model is within a preset range, the trained machine learning model is determined as the air pollutant concentration prediction model.

[0012] As a preferred embodiment of the present invention, the machine learning model includes XGBoost, random forest, linear regression, support vector machine or neural network.

[0013] As a preferred embodiment of the present invention, in step c, the XGBoost model is used for the machine learning model. The machine learning model is trained using the dimensionless processed training sample set to obtain a trained machine learning model. Specifically:

[0014] (1) Feature variables and target variables are extracted from the dimensionless processed training sample set. The feature variables are the historical hourly data of the meteorological pattern, and the target variables are the historical hourly data of air pollutant concentration corresponding to the same date and the same time as the historical hourly data of the meteorological pattern. Features are extracted from the historical hourly data of the meteorological pattern, including the near-surface 10-meter u-wind field, the near-surface 10-meter v-wind field, the near-surface 2-meter temperature, the boundary layer height, the short-wave radiation amount, the relative humidity, the precipitation, the air pressure value, the specific humidity value, the u-wind field at 925 hPa, the v-wind field at 925 hPa, the u-wind field at 850 hPa, the v-wind field at 850 hPa, the u-wind field at 700 hPa, the v-wind field at 700 hPa, the temperature difference between 975 hPa and 950 hPa, the temperature difference between 950 hPa and 925 hPa, the temperature difference between 925 hPa and 900 hPa, the temperature difference between 900 hPa and 850 hPa, the average cloud amount, the near-surface 10-meter wind speed, and the near-surface 10-meter wind direction as the input features of the XGBoost model; the dimensionless processed training sample set is further randomly divided into a training set and a validation set according to a ratio of 8:2;

[0015] (2) Set the initial parameters of the XGBoost model, including the objective function, the learning rate, the maximum depth of the tree, and the number of trees, as well as the evaluation index;

[0016] (3) Convert the training set data and validation set data into the DMatrix format of XGBoost. Use the xgboost.train() function to input the training set data, validation set data, initial parameters, and features extracted from the historical hourly meteorological model data for XGBoost model training. Use the validation set data to evaluate the performance of the XGBoost model by calculating the mean squared error and root mean squared error. Repeatedly train the XGBoost model by adjusting the initial parameters of the XGBoost model until the difference between the historical hourly air pollutant concentration data input from the validation set and the predicted hourly air pollutant concentration data output by the XGBoost model is within a preset range, and then obtain the trained XGBoost model.

[0017] After adopting the above method, the present invention has the following beneficial effects:

[0018] The meteorological factors used in the present invention include 22 meteorological factors such as the 10-meter u-wind field near the ground, the 10-meter v-wind field near the ground, the 2-meter temperature near the ground, the boundary layer height, the short-wave radiation amount, the relative humidity, the precipitation, the air pressure value, the specific humidity value, the u-wind field at 925 hPa, the v-wind field at 925 hPa, the u-wind field at 850 hPa, the v-wind field at 850 hPa, the u-wind field at 700 hPa, the v-wind field at 700 hPa, the temperature difference between 975 hPa and 950 hPa, the temperature difference between 950 hPa and 925 hPa, the temperature difference between 925 hPa and 900 hPa, the temperature difference between 900 hPa and 850 hPa, the average cloud amount, the 10-meter wind speed near the ground, and the 10-meter wind direction near the ground. The meteorological factors used are comprehensive and reasonable, avoiding the disadvantages of low prediction accuracy caused by fewer meteorological factors and missing important factors in the prior art. The present invention greatly improves the prediction accuracy; the historical hourly meteorological model data, the historical hourly air pollutant concentration data, the air pollutant concentration data at the starting time, and the predicted hourly meteorological model data at x moments after the starting time in the present invention can all be extracted from the database, and the extraction is simple and convenient, and no air pollutant emission source data is required, so the present invention greatly reduces the implementation cost; the air pollutant concentration prediction method at time T1 in the prediction method of the present invention is to use the predicted hourly meteorological model data at the first moment T1 after the starting time and the air pollutant concentration data at the starting time T0 as input quantities to input into the air pollutant concentration prediction model to obtain the predicted hourly air pollutant concentration data at time T1. The air pollutant concentration prediction method at time T2 is to use the predicted hourly meteorological model data at the second moment T2 after the starting time and the predicted hourly air pollutant concentration data at time T1 as input quantities to input into the air pollutant concentration prediction model to obtain the predicted hourly air pollutant concentration data at time T2, and so on, to obtain the dynamic hourly prediction data of the air pollutant concentration at x moments after the starting time. That is to say, the present invention realizes the dynamic hourly prediction of the air pollutant concentration.

[0019] The present invention can well solve the hourly update problem in air quality prediction, can achieve short-term accurate prediction and medium- and long-term trend prediction of air quality pollutants, and improves the prediction accuracy.

[0020] The present invention performs dimensionless processing on the training sample set and the test sample set. This processing can make the conversion values of all data be between 0 and 10, preventing the influence of too large a value of a certain data on the prediction accuracy, and further improving the prediction accuracy of the air pollutant concentration.

[0021] The present invention does not require a large amount of computing resources, nor does it require a dedicated server. The present invention has extremely low requirements for the computing platform, and a personal computer can be used for operation with a fast response speed.

[0022] The present invention is simple to implement, fast in calculation, low in economic cost, and easy to promote, and has great economic value and social benefits in aspects such as dynamic pollution control, start and cancellation of heavy pollution warnings, and air quality guarantee for major activities. Specific implementation manner

[0023] The present invention provides a dynamic hourly prediction method for air pollutant concentration based on a machine learning model, which is used to obtain dynamic hourly prediction data of air pollutant concentration at x moments after the starting moment, and includes the following steps:

[0024] a. Extract historical data, starting data, and prediction data: The historical data includes historical hourly data of meteorological patterns and historical hourly data of air pollutant concentrations in the N months before the starting moment on a certain date, a certain month, and a certain year in a certain area. The starting data includes air pollutant concentration data at the starting moment. The prediction data includes predicted hourly data of meteorological patterns at x moments after the starting moment, where N is a natural number greater than or equal to 6, that is, more than 6 months, and x is a natural number greater than or equal to 2, that is, more than 2 moments / hours. The historical hourly data of meteorological patterns includes the 10-meter u-wind field near the ground, the 10-meter v-wind field near the ground, the 2-meter temperature near the ground, the boundary layer height, the short-wave radiation amount, the relative humidity, the precipitation, the air pressure value, the specific humidity value, the u-wind field at 925 hPa, the v-wind field at 925 hPa, the u-wind field at 850 hPa, the v-wind field at 850 hPa, the u-wind field at 700 hPa, the v-wind field at 700 hPa, the temperature difference between 975 hPa and 950 hPa, the temperature difference between 950 hPa and 925 hPa, the temperature difference between 925 hPa and 900 hPa, the temperature difference between 900 hPa and 850 hPa, the average cloud amount, the 10-meter wind speed near the ground, and the 10-meter wind direction near the ground;

[0025] b. Process the historical data and predicted data: First, check whether the historical hourly data of the meteorological pattern and the historical hourly data of the air pollutant concentration in the N months before the reporting time are complete. If there are incomplete historical missing data, use the well-known exponential smoothing method ES to supplement and process the historical missing data to obtain complete historical hourly data of the meteorological pattern and the air pollutant concentration, and check whether the predicted hourly data of the meteorological pattern at x moments after the reporting time are complete. If there are incomplete predicted missing data, use the well-known exponential smoothing method ES to supplement and process the predicted missing data to obtain complete predicted hourly data of the meteorological pattern. Then, perform corresponding matching and correlation processing on the complete historical hourly data of the meteorological pattern and the historical hourly data of the air pollutant concentration according to the same date and the same time to obtain the historical hourly matching data of the meteorological pattern and the air pollutant concentration;

[0026] c. Randomly divide the historical hourly matching data of the meteorological pattern and the air pollutant concentration into a training sample set and a test sample set according to a preset ratio. Use the training sample set to train the machine learning model to obtain a trained machine learning model. Input the test sample set into the trained machine learning model for detection. Until the difference between each historical hourly data of the air pollutant concentration input from the test sample set and the predicted hourly data of the air pollutant concentration output by the trained machine learning model is within the preset range, determine the trained machine learning model as the air pollutant concentration prediction model;

[0027] d. Input the predicted hourly data of the meteorological pattern at x moments after the reporting time in step a into the air pollutant concentration prediction model to obtain the dynamic hourly prediction data of the air pollutant concentration at x moments after the reporting time. Specifically: T0 represents the reporting time, and T i (i = 1, 2,..., x) represents the i-th moment after the reporting time, where T1 represents the first moment after the reporting time, T2 represents the second moment after the reporting time, and T x represents the x-th moment after the reporting time. The prediction method for the air pollutant concentration at the T1 moment is to input the predicted hourly data of the meteorological pattern at the first moment T1 after the reporting time and the air pollutant concentration data at the reporting time T0 into the air pollutant concentration prediction model to obtain the predicted hourly data of the air pollutant concentration at the T1 moment; The prediction method for the air pollutant concentration at the T2 moment is to input the predicted hourly data of the meteorological pattern at the second moment T2 after the reporting time and the predicted hourly data of the air pollutant concentration at the T1 moment into the air pollutant concentration prediction model to obtain the predicted hourly data of the air pollutant concentration at the T2 moment; The prediction method for the air pollutant concentration at the T x moment is to input the predicted hourly data of the meteorological pattern at the x-th moment T xThe hourly data of the meteorological model prediction and the hourly data of the predicted air pollutant concentration at time T x-1 are input into the air pollutant concentration prediction model as input quantities to obtain the hourly data of the predicted air pollutant concentration at time T x ; and so on, to obtain the dynamic hourly prediction data of the air pollutant concentration at x moments after the starting time.

[0028] As a preferred embodiment of the present invention, in step a, the historical data and the starting data are extracted from the database of a regional multi-parameter atmospheric station, and the prediction data is extracted from the mesoscale atmospheric model WRF; the historical data includes the historical hourly data of the meteorological model and the historical hourly data of the air pollutant concentration in the 8 - 12 months before the starting time on a certain date in a certain region. Of course, only the historical hourly data of the meteorological model and the historical hourly data of the air pollutant concentration in the 6 months before the starting time or a longer time such as more than 12 months can also be used. The prediction data includes the hourly data of the meteorological model prediction at 72 - 168 moments after the starting time. Of course, the hourly data of the meteorological model prediction at a shorter time such as 24 moments / hour or a longer time such as 192 moments / hour can also be used; the air pollutants include PM 2.5 , O3, PM 10 , NO2, SO2, VOCs, etc.

[0029] As a preferred embodiment of the present invention, in step c, the historical hourly matching data of the meteorological model and the air pollutant concentration is randomly divided into a training sample set and a test sample set according to a ratio of 8:2 or 7:3. The training sample set and the test sample set are subjected to a well-known dimensionless processing. By eliminating the influence of the dimension, the data is made more comparable. The machine learning model is trained using the dimensionless training sample set to obtain a trained machine learning model. The dimensionless test sample set is input into the trained machine learning model for detection until the difference between the historical hourly data of the air pollutant concentration input from each dimensionless test sample set and the predicted hourly data of the air pollutant concentration output by the trained machine learning model is within a preset range, and the trained machine learning model is determined as the air pollutant concentration prediction model.

[0030] As a preferred embodiment of the present invention, the machine learning model includes XGBoost, random forest, linear regression, support vector machine or neural network.

[0031] As a preferred embodiment of the present invention, in step c, the machine learning model adopts the XGBoost (Extreme Gradient Boosting) model, and the trained machine learning model is obtained by training the machine learning model with the dimensionless processed training sample set. Specifically:

[0032] (1) Extract feature variables and target variables from the dimensionless processed training sample set. The feature variables are historical hourly data of meteorological patterns, and the target variables are historical hourly data of air pollutant concentrations corresponding and matched to the historical hourly data of meteorological patterns at the same date and the same time. Extract features from the historical hourly data of meteorological patterns, including near-surface 10-meter u-wind field, near-surface 10-meter v-wind field, near-surface 2-meter temperature, boundary layer height, short-wave radiation amount, relative humidity, precipitation, air pressure value, specific humidity value, u-wind field at 925 hPa, v-wind field at 925 hPa, u-wind field at 850 hPa, v-wind field at 850 hPa, u-wind field at 700 hPa, v-wind field at 700 hPa, temperature difference between 975 hPa and 950 hPa, temperature difference between 950 hPa and 925 hPa, temperature difference between 925 hPa and 900 hPa, temperature difference between 900 hPa and 850 hPa, average cloud amount, near-surface 10-meter wind speed, and near-surface 10-meter wind direction as input features of the XGBoost model; further randomly divide the dimensionless processed training sample set into a training set and a validation set according to a ratio of 8:2.

[0033] (2) Set the initial parameters of the XGBoost model, including the objective function, learning rate, maximum depth of the tree, number of trees, and evaluation metrics, etc.

[0034] (3) Convert the training set data and validation set data into the DMatrix format of XGBoost, use the xgboost.train() function, input the training set data, validation set data, initial parameters, and features extracted from the historical hourly data of meteorological patterns for XGBoost model training, use the validation set data to evaluate the performance of the XGBoost model by calculating the mean squared error and root mean squared error, and repeatedly train the XGBoost model by adjusting the initial parameters of the XGBoost model until the difference between each historical hourly data of air pollutant concentration input from the validation set and the predicted hourly data of air pollutant concentration output by the XGBoost model is within a preset range. The preset range is, for example, 0 to 0.2 micrograms per cubic meter, to obtain the trained XGBoost model.

[0035] As another preferred embodiment of the present invention, in step c, the machine learning model adopts a neural network model, and the trained machine learning model is obtained by training the machine learning model with the dimensionless processed training sample set. Specifically:

[0036] Set the CNN network structure. Based on the VGG16 network model, adjust the parameters of the VGG16 network model. Set the region of interest pooling of the VGG16 network model as ROIAlign, introduce the Inception network structure, and establish the FasterR-CNN network model. Use the FasterR-CNN network model as the neural network model. The specific steps for training the FasterR-CNN network model with the dimensionless processed training sample set are as follows: Randomly divide the dimensionless processed training sample set into a training set, a test set, and a validation set according to a preset ratio. Train using the training set through a deep learning algorithm, and update and iterate the weight parameters of the convolutional neural network according to the results of the test set and the validation set to train the FasterR-CNN network model until the FasterR-CNN network model is trained optimally, obtaining the trained FasterR-CNN neural network model.

[0037] As a preferred embodiment of the present invention, taking Changzhou City, Jiangsu Province as an example, in order to obtain the dynamic hourly prediction data of air pollutant concentrations at 72 moments / hours after 20:00 on August 31, 2023, the following steps are adopted:

[0038] a. Extract historical data from the database of the Changzhou Atmospheric Multi-parameter Station (31.7586°E, 119.9523°N), initial data, and extract prediction data from the mesoscale atmospheric model WRF (Weather Research and Forecasting): The historical data includes the historical hourly data of the meteorological model and the historical hourly data of air pollutant concentrations in Changzhou from 20:00 on January 1, 2023 to 19:00 on August 31, 2023, which is the 8 months before 20:00 on August 31, 2023. The initial data includes the air pollutant concentration data at 20:00 on August 31, 2023. The prediction data includes the predicted hourly data of the meteorological model at 72 moments / hours after 20:00 on August 31, 2023. The historical hourly data of the meteorological model includes the near-surface 10-meter u wind field, near-surface 10-meter v wind field, near-surface 2-meter temperature, boundary layer height, shortwave radiation, relative humidity, precipitation, air pressure value, specific humidity value, u wind field at 925 hPa, v wind field at 925 hPa, u wind field at 850 hPa, v wind field at 850 hPa, u wind field at 700 hPa, v wind field at 700 hPa, temperature difference between 975 hPa and 950 hPa, temperature difference between 950 hPa and 925 hPa, temperature difference between 925 hPa and 900 hPa, temperature difference between 900 hPa and 850 hPa, average cloud amount, near-surface 10-meter wind speed, and near-surface 10-meter wind direction;

[0039] b. Process the historical data and prediction data: First, check whether the historical hourly data of the meteorological model and the historical hourly data of air pollutant concentrations in the 8 months before 20:00 on August 31, 2023 are complete. If there are incomplete historical missing data, use the exponential smoothing method ES to supplement and process the historical missing data to obtain complete historical hourly data of the meteorological model and air pollutant concentrations, and check whether the predicted hourly data of the meteorological model at 72 moments / hours after 20:00 on August 31, 2023 are complete. If there are incomplete predicted missing data, use the exponential smoothing method ES to supplement and process the predicted missing data to obtain complete predicted hourly data of the meteorological model. Then, perform corresponding matching and association processing on the complete historical hourly data of the meteorological model and the historical hourly data of air pollutant concentrations according to the same date and the same time to obtain the historical hourly matching data of the meteorological model and air pollutant concentrations;

[0040] c. Randomly divide the meteorological pattern and the historical hourly matching data of air pollutant concentrations into a training sample set and a test sample set according to the ratio of 8:2. Perform dimensionless processing on the training sample set and the test sample set, and use the dimensionless training sample set to train the XGBoost model to obtain the trained XGBoost model. Specifically: (1) Extract the feature variables and target variables from the dimensionless training sample set. The feature variables are the historical hourly meteorological pattern data for the 8 months before 20:00 on August 31, 2023, and the target variable is the historical hourly air pollutant concentration data corresponding and matched to the meteorological pattern historical hourly data at the same date and the same time. Extract features from the meteorological pattern historical hourly data, including the near-surface 10-meter u-wind field, near-surface 10-meter v-wind field, near-surface 2-meter temperature, boundary layer height, short-wave radiation amount, relative humidity, precipitation, air pressure value, specific humidity value, u-wind field at 925 hPa, v-wind field at 925 hPa, u-wind field at 850 hPa, v-wind field at 850 hPa, u-wind field at 700 hPa, v-wind field at 700 hPa, temperature difference between 975 hPa and 950 hPa, temperature difference between 950 hPa and 925 hPa, temperature difference between 925 hPa and 900 hPa, temperature difference between 900 hPa and 850 hPa, average cloud amount, near-surface 10-meter wind speed, and near-surface 10-meter wind direction as the input features of the XGBoost model; further randomly divide the dimensionless training sample set into a training set and a validation set according to the ratio of 8:2; (2) Set the initial parameters of the XGBoost model, including the objective function, learning rate, maximum depth of the tree, and number of trees, as well as the evaluation index; the objective function is usually reg:squarederror for regression tasks, the learning rate controls the step size of each iteration, the maximum depth of the tree and the number of trees control the complexity of the model, the maximum depth of the tree is preferably 10, the number of trees is preferably 100, and the evaluation index is such as rmse (root mean square error); (3) Convert the training set data and the validation set data into the DMatrix format of XGBoost, use the xgboost.train() function, input the training set data, validation set data, initial parameters, and the features extracted from the historical hourly meteorological pattern data for the 8 months before 20:00 on August 31, 2023 to train the XGBoost model, use the validation set data to evaluate the performance of the XGBoost model by calculating the mean squared error and root mean square error, and repeatedly train the XGBoost model by adjusting the initial parameters of the XGBoost model until the difference between each historical hourly air pollutant concentration data input from the validation set and the predicted hourly air pollutant concentration data output by the XGBoost model is within the preset range. The preset range is, for example, 0 to 0.2 μg / m³ to obtain the trained XGBoost model; input the test sample set into the trained XGBoost model for detection until the difference between the historical hourly data of air pollutant concentration input from the test sample set and the predicted hourly data of air pollutant concentration output by the trained XGBoost model is within the preset range. For example, the preset range is 0 - 0.2 μg / m³. Then, determine this trained XGBoost model as the air pollutant concentration prediction model.

[0041] d. Input the predicted hourly data of the meteorological model at 72 moments / hours after 20:00 on August 31, 2023 into the air pollutant concentration prediction model to obtain the dynamic hourly prediction data of air pollutant concentration at 72 moments / hours after 20:00 on August 31, 2023. Specifically, T0 represents 20:00 on August 31, 2023, and T i (i = 1, 2, …, 72) represents the i-th moment after 20:00 on August 31, 2023. Among them, T1 represents the 1st moment after 20:00 on August 31, 2023, that is, 21:00, T2 represents the 2nd moment after 20:00 on August 31, 2023, that is, 22:00, and T 72 represents the 72nd moment after 20:00 on August 31, 2023. The prediction method for the air pollutant concentration at T1 moment is to take the predicted hourly data of the meteorological model at the 1st moment T1 after 20:00 on August 31, 2023 and the air pollutant concentration data at T0 (20:00 on August 31, 2023) as input quantities and input them into the air pollutant concentration prediction model to obtain the predicted hourly data of the air pollutant concentration at T1 moment; the prediction method for the air pollutant concentration at T2 moment is to take the predicted hourly data of the meteorological model at the 2nd moment T2 after 20:00 on August 31, 2023 and the predicted hourly data of the air pollutant concentration at T1 moment as input quantities and input them into the air pollutant concentration prediction model to obtain the predicted hourly data of the air pollutant concentration at T2 moment; the prediction method for the air pollutant concentration at T 72 moment is to take the predicted hourly data of the meteorological model at the 72nd moment T 72 after 20:00 on August 31, 2023 and the predicted hourly data of the air pollutant concentration at T 71 moment as input quantities and input them into the air pollutant concentration prediction model to obtain the predicted hourly data of the air pollutant concentration at T 72 moment; and so on, to obtain the dynamic hourly prediction data of air pollutant concentration at 72 moments / hours after 20:00 on August 31, 2023.

[0042] After trial, the implementation cost of the present invention is low, the prediction accuracy is high, the dynamic hourly prediction of air pollutant concentration is well achieved, and there is no need for a dedicated server, and it can be operated on a personal computer, achieving good results.

Claims

1. A dynamic hourly prediction method for air pollutant concentration based on a machine learning model, used to obtain dynamic hourly prediction data of air pollutant concentration at x moments after the starting time, characterized in that, Including the following steps: a. Extract historical data, starting data, and prediction data: The historical data includes the historical hourly data of the meteorological pattern and the historical hourly data of air pollutant concentration for the N months before the starting time on a certain date, month, and year in a certain region. The starting data includes the air pollutant concentration data at the starting time. The prediction data includes the predicted hourly data of the meteorological pattern for x moments after the starting time, where N is a natural number ≥ 6, and x is a natural number ≥ 2. The historical hourly data of the meteorological pattern includes the 10-meter u-wind field near the ground, the 10-meter v-wind field near the ground, the 2-meter temperature near the ground, the boundary layer height, the shortwave radiation amount, the relative humidity, the precipitation, the air pressure value, the specific humidity value, the u-wind field at 925 hPa, the v-wind field at 925 hPa, the u-wind field at 850 hPa, the v-wind field at 850 hPa, the u-wind field at 700 hPa, the v-wind field at 700 hPa, the temperature difference between 975 hPa and 950 hPa, the temperature difference between 950 hPa and 925 hPa, the temperature difference between 925 hPa and 900 hPa, the temperature difference between 900 hPa and 850 hPa, the average cloud amount, the 10-meter wind speed near the ground, and the 10-meter wind direction near the ground; b. Process the historical data and prediction data: First, check whether the historical hourly data of the meteorological pattern and the historical hourly data of air pollutant concentration for the N months before the starting time are complete. If there are incomplete historical missing data, use the exponential smoothing method (ES) to supplement and process the historical missing data to obtain complete historical hourly data of the meteorological pattern and air pollutant concentration, and check whether the predicted hourly data of the meteorological pattern for x moments after the starting time are complete. If there are incomplete prediction missing data, use the exponential smoothing method (ES) to supplement and process the prediction missing data to obtain complete predicted hourly data of the meteorological pattern. Then, perform corresponding matching and association processing on the complete historical hourly data of the meteorological pattern and air pollutant concentration according to the same date and the same moment to obtain the historical hourly matching data of the meteorological pattern and air pollutant concentration; c. Randomly divide the historical hourly matching data of the meteorological pattern and air pollutant concentration into a training sample set and a test sample set according to a preset ratio. Use the training sample set to train a machine learning model to obtain a trained machine learning model. Input the test sample set into the trained machine learning model for detection. Until the difference between each historical hourly data of air pollutant concentration input from the test sample set and the predicted hourly data of air pollutant concentration output by the trained machine learning model is within the preset range, determine this trained machine learning model as the air pollutant concentration prediction model; d. Input the meteorological model prediction hourly data at x time instances after the starting reporting time in step a into the air pollutant concentration prediction model to obtain the dynamic hourly prediction data of air pollutant concentrations at x time instances after the starting reporting time. Specifically, T0 represents the starting reporting time, and T i (i = 1, 2, …, x) represents the i-th time instance after the starting reporting time, where T1 represents the 1st time instance after the starting reporting time, T2 represents the 2nd time instance after the starting reporting time, and T x represents the x-th time instance after the starting reporting time. The air pollutant concentration prediction method for T1 time instance is to input the meteorological model prediction hourly data at the 1st time instance T1 after the starting reporting time and the air pollutant concentration data at the starting reporting time T0 into the air pollutant concentration prediction model to obtain the air pollutant concentration prediction hourly data at T1 time instance; the air pollutant concentration prediction method for T2 time instance is to input the meteorological model prediction hourly data at the 2nd time instance T2 after the starting reporting time and the air pollutant concentration prediction hourly data at T1 time instance into the air pollutant concentration prediction model to obtain the air pollutant concentration prediction hourly data at T2 time instance; the air pollutant concentration prediction method for T x time instance is to input the meteorological model prediction hourly data at the x-th time instance T x after the starting reporting time and the air pollutant concentration prediction hourly data at T x-1 time instance into the air pollutant concentration prediction model to obtain the air pollutant concentration prediction hourly data at T x time instance; and so on, to obtain the dynamic hourly prediction data of air pollutant concentrations at x time instances after the starting reporting time.

2. The dynamic hourly prediction method for air pollutant concentration based on a machine learning model according to claim 1, wherein: In step a, the historical data and the initial reporting data are extracted from the database of the atmospheric multi-parameter station in a certain area, and the prediction data is extracted from the mesoscale atmospheric model WRF; the historical data includes the historical hourly data of the meteorological model and the historical hourly data of the air pollutant concentration in the 8 - 12 months before the initial reporting time on a certain date in a certain area, and the prediction data includes the predicted hourly data of the meteorological model at 72 - 168 moments after the initial reporting time; the air pollutants include PM 2.5 , O3, PM 10 , NO2, SO2 and VOCs.

3. The dynamic hourly prediction method for air pollutant concentration based on a machine learning model according to claim 1, characterized in that: In step c, the meteorological pattern and the historical hourly matching data of air pollutant concentrations are randomly divided into a training sample set and a test sample set according to a ratio of 8:2 or 7:

3. The training sample set and the test sample set are subjected to dimensionless processing. The machine learning model is trained using the dimensionless training sample set to obtain a trained machine learning model. The dimensionless test sample set is input into the trained machine learning model for detection. Until the difference between each historical hourly data of air pollutant concentrations input from the dimensionless test sample set and the predicted hourly data of air pollutant concentrations output by the trained machine learning model is within a preset range, the trained machine learning model is determined as the air pollutant concentration prediction model.

4. The dynamic hourly prediction method for air pollutant concentration based on a machine learning model according to claim 1 or 2 or 3, characterized in that: The machine learning model includes XGBoost, random forest, linear regression, support vector machine, or neural network.

5. The dynamic hourly prediction method for air pollutant concentration based on a machine learning model according to claim 3, characterized in that: In step c, the XGBoost model is used for the machine learning model. The machine learning model is trained using the dimensionless training sample set to obtain a trained machine learning model, specifically: (1) Extract the feature variables and target variables from the dimensionless training sample set. The feature variables are the historical hourly data of the meteorological pattern, and the target variables are the historical hourly data of air pollutant concentrations corresponding to the same date and the same time as the historical hourly data of the meteorological pattern. Extract features from the historical hourly data of the meteorological pattern, including the near-surface 10-meter u-wind field, the near-surface 10-meter v-wind field, the near-surface 2-meter temperature, the boundary layer height, the shortwave radiation amount, the relative humidity, the precipitation, the air pressure value, the specific humidity value, the u-wind field at 925 hPa, the v-wind field at 925 hPa, the u-wind field at 850 hPa, the v-wind field at 850 hPa, the u-wind field at 700 hPa, the v-wind field at 700 hPa, the temperature difference between 975 hPa and 950 hPa, the temperature difference between 950 hPa and 925 hPa, the temperature difference between 925 hPa and 900 hPa, the temperature difference between 900 hPa and 850 hPa, the average cloud amount, the near-surface 10-meter wind speed, and the near-surface 10-meter wind direction as the input features of the XGBoost model; the dimensionless training sample set is further randomly divided into a training set and a validation set according to a ratio of 8:2; (2) Set the initial parameters of the XGBoost model, including the objective function, the learning rate, the maximum depth of the tree, and the number of trees, as well as the evaluation metrics; (3) Convert the training set data and the validation set data into the DMatrix format of XGBoost. Use the xgboost.train() function to input the training set data, the validation set data, the initial parameters, and the features extracted from the historical hourly data of the meteorological model for XGBoost model training. Use the validation set data to evaluate the performance of the XGBoost model by calculating the mean squared error and the root mean squared error. Repeatedly train the XGBoost model by adjusting the initial parameters of the XGBoost model until the difference between the historical hourly data of the air pollutant concentration input from the validation set and the predicted hourly data of the air pollutant concentration output by the XGBoost model is within a preset range, and then obtain the trained XGBoost model.

Citation Information

Cited By

  • Method for predicting air quality ranking by using ascending and descending channels

    CN121299043A