Air pollutant concentration prediction method
By transforming pollutant data and meteorological data into spatiotemporal data matrix and using specific deep learning models for prediction, the problem of low prediction accuracy in the prior art is solved, and higher prediction accuracy and ability to adapt to complex changing laws are achieved.
Patent Information
- Application Number
- CN202510130140.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems with low accuracy in pollutant concentration prediction, especially when faced with larger data sets and more complex nonlinear relationships.
An air pollutant concentration prediction method is adopted, including the acquisition of pollutant data and meteorological data, the data is converted into a spatiotemporal data matrix, and the prediction is made using a predetermined prediction model of the time convolution network segment, the extrusion excitation mechanism segment, the bidirectional gated circulation unit segment and the convolution block attention module segment.
This method significantly improves the accuracy of air pollutant concentration prediction and can better adapt to the complex air pollutant concentration changes.
Smart Images

Figure CN120069302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pollutant concentration prediction. Background Art
[0002] Accurately predicting pollutant concentrations can provide timely warnings for residents and governments, so that they can take countermeasures. At present, the prediction of pollutants mainly adopts statistical model methods, simple machine learning methods and deep learning methods. Statistical models involve a lot of mathematical calculations and the prediction efficiency is not high. Simple machine learning methods are still weak when faced with larger data sets and more complex nonlinear relationships. Most machine learning methods rely on manual feature engineering to extract effective features from the data, which makes the machine learning model not have good generalization ability. Deep learning methods have been widely used in many fields due to their strong advantages in data processing capabilities and nonlinear modeling, but at present, deep learning methods still have the problem of low accuracy in pollutant prediction. Summary of the invention
[0003] The present invention is made in view of the above problems. According to one aspect of the present invention, a method for predicting air pollutant concentration includes the following steps: a pollutant data and meteorological data acquisition step, obtaining pollutant data and meteorological data within a predetermined time interval before a time period to be predicted; a transformation step, transforming the collected data into a spatiotemporal data matrix; and a prediction step, using the pollutant data and meteorological data collected within a predetermined time interval, using a predetermined prediction model to predict the air pollutant concentration at a predetermined time, wherein the predetermined prediction model includes a temporal convolutional network segment, a squeeze excitation mechanism segment, a bidirectional gated recurrent unit segment, and a convolutional block attention module segment connected in series. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] The present invention can be better understood with reference to the accompanying drawings, which are merely illustrative and are not intended to limit the scope of protection of the present invention.
[0005] Figure 1 A schematic flow chart of a method for predicting air pollutant concentration according to an embodiment of the present invention is shown.
[0006] Figure 2 A schematic diagram of a prediction model according to an embodiment of the present invention is shown.
[0007] Figure 3 A schematic diagram of a squeeze excitation network (SENet) segment 30 is shown according to one embodiment.
[0008] Figure 4 A schematic diagram of a squeeze-excitation mechanism BiGRU segment according to an embodiment is shown.
[0009] Figure 5 FIG. shows a schematic diagram of a channel attention module according to an embodiment of the present invention.
[0010] Figure 6 FIG. shows a schematic diagram of a spatial attention module according to an embodiment of the present invention. Detailed Embodiments
[0011] The following describes the detailed embodiments of the present invention with reference to the accompanying drawings. These descriptions are exemplary and are intended to enable those skilled in the art to implement the embodiments of the present invention, rather than limiting the scope of protection of the present invention. The descriptions also do not describe content that is essential for actual implementation but is irrelevant to understanding the present invention.
[0012] Figure 1 FIG. shows a schematic flowchart of an air pollutant concentration prediction method according to an embodiment of the present invention. As Figure 1 shown, according to an embodiment of the present invention, an air pollutant concentration prediction method includes a pollutant data and meteorological data acquisition step S100, a transformation step S200, a training step 300, and a prediction step 400.
[0013] The pollutant data and meteorological data acquisition step S100 acquires pollutant data and meteorological data within a predetermined time period before the time period to be predicted. According to an embodiment, the pollutant data includes the pollutant data to be predicted and the pollutant data related to the pollutant data to be predicted. For example, if the pollutant data to be predicted is the PM 2.5 concentration tomorrow, the pollutant data may include the PM 2.5 concentration for the week before today and the pollutant data related to the PM 2.5 concentration, such as the hourly concentration data of PM 10 , SO 2 , NO 2 , O 3 , and CO in the prediction area. Since the pollutant concentration changes continuously over time, only the historical data for a previous adjacent period of time is needed to predict the pollutant concentration tomorrow. The meteorological data may include the meteorological data related to the pollutant to be predicted, such as temperature, ground pressure, relative humidity, instantaneous wind direction, instantaneous wind speed, maximum wind speed in 1 hour, precipitation in 1 hour, and average visibility in 10 minutes.
[0014] According to one embodiment, not only the pollutant data and meteorological data of the location to be predicted are collected, but also the pollutant data and meteorological data of the locations associated with the location to be predicted are collected. For example, if the location to be predicted is Beijing, then not only the pollutant data and meteorological data of Beijing are collected, but also the pollutant data and meteorological data of six adjacent cities such as Tianjin, Langfang, Baoding, Zhangjiakou, Chengde, and Tangshan are collected. The cities associated with the location to be predicted are selected according to the distance from the location to be predicted, the spreadability of the pollutant to be predicted, the population base of the city, and the activity level of social activities. Administrative divisions below the prefecture-level city are generally ignored.
[0015] Then, in the transformation step S200, the collected data is transformed into a spatio-temporal data matrix. The collected data is in the form of time series data, and in the data processing module of the prediction system, the time series data will be processed into a three-dimensional vector. According to one embodiment, this three-dimensional vector is a spatio-temporal matrix. According to one embodiment, its three dimensions (height, width, and channel) are respectively: the time step (i.e., the number of time points to look back), the city (in the above example, 7 cities), and the characteristic air pollutant data and meteorological data.
[0016] Next, in step S300, the spatio-temporal data matrix is used to train the prediction model. Various methods known to those skilled in the art or to be known in the future can be used to train the prediction model. It should be noted that in practice, it is not necessary to train the prediction model every time a prediction is made.
[0017] Figure 2 A schematic diagram of a prediction model according to an embodiment of the present invention is shown. As Figure 2 shown, the prediction model includes a time convolutional network (TCN) segment 20, a squeeze-and-excitation network (SENet) segment 30, a bidirectional gated recurrent unit (BiGRU) segment 40, and a convolutional block attention module (CBAM) segment 50 connected in series.
[0018] The TCN segment 20 includes a first branch (dilated convolution route) and a second branch (residual connection route). The first branch includes a dilated convolutional layer, a ReLU activation function layer, and a padding operation layer, and the result of the padding operation is the output of the first branch. The second branch includes a 1x1 convolutional layer (a convolutional operation with a kernel size of 1x1). The input spatio-temporal data matrix is input into both the first branch and the second branch at the same time. The output of the first branch and the output of the second branch are added and output to the SENet segment 30.
[0019] The processing flow in TCN is as follows: After the spatio-temporal data matrix enters the network, it will pass through the dilated convolution route. The dilated convolution route contains one or more dilated convolution layers. In each dilated convolution layer, dilated convolution kernels are used to capture dependencies at different time scales. The result after the convolution operation will pass through the ReLU activation function layer to introduce non-linearity and enhance the model's expressive power. To keep the size of the output features consistent with the input, padding operations are used, that is, adding zero values at the beginning of the sequence to ensure that the output at each time step only depends on the current and previous time steps, rather than future time steps. Here, the sequence can be the original input data or the output of the previous convolution layer. At the same time, the input data will also pass through the residual connection route. The residual connection route contains a 1×1 convolution layer, which is used to adjust the number of input channels to match the output channels of the dilated convolution route, so as to ensure that the input and output can be added element-wise in terms of dimensions. Finally, the output of the dilated convolution layer and the input adjusted by the 1×1 convolution are added element-wise at each time step to form a new feature representation.
[0020] One of the major features of TCN is that it has two routes. It obtains the output through the dilated convolution route and retains the input through the residual connection route, and then adds the input and output element-wise to form the final output, which solves the problem of gradient disappearance in deep networks and enables TCN to better extract the long-term dependence features of data. The SENet segment 30 includes a squeezing operation and an excitation operation to process the output information (feature map) of TCN. Figure 3 FIG. shows a schematic diagram of the squeeze-and-excitation mechanism (SENet) segment 30 according to an embodiment.
[0021] The processing flow in SENet is as follows: For the input features, that is, the features after summation in the TCN segment, a feature transformation (such as a convolution operation) is performed to generate an intermediate feature map. This feature map will undergo a squeezing operation. In this step, SENet uses global average pooling technology to average all the values within each channel, thus compressing the rich feature information into a single value, that is, the channel descriptor. The advantage of this is that it enables SENet to capture the global information of each channel and provides a concise and effective data representation for subsequent processing. Immediately afterwards, these channel descriptors will enter the excitation operation stage. In this stage, the descriptors will first be transformed through two fully connected layers to ensure the generation of meaningful weights. The first fully connected layer plays a role in dimensionality reduction. It reduces the number of descriptors, enabling SENet to process information more efficiently and reduce the computational burden. The second fully connected layer restores the data to the original number of channels, so that the generated weights can correspond to each channel of the original feature map one by one. Finally, each channel descriptor is converted into a weight value between 0 and 1 through the Sigmoid activation function.
[0022] The weights generated by the SENet segment 30 are used to weight the original features, emphasizing the channels considered important by SENet while suppressing the relatively unimportant channels. In this way, SENet can adaptively adjust the channel responses of the feature map, enabling the model to focus more on the features beneficial to the current task, thereby enhancing the overall performance of the model.
[0023] Figure 4 The schematic diagram of the squeeze-and-excitation mechanism BiGRU segment according to an embodiment is shown.
[0024] The BiGRU segment 40 includes forward and backward GRU layers to process the information in both directions of the input sequence, forming a bidirectional feature representation for each time step. After the data enters the BiGRU, the forward GRU layer processes the input data sequentially from the start position to the end position of the sequence to capture the temporal dependence of the sequence; at the same time, the backward GRU layer processes the input data from the end position to the start position of the sequence to capture the reverse dependence relationship of the sequence. Finally, the outputs of the forward GRU and the backward GRU are combined to generate an output that fuses bidirectional information to enhance the global context understanding of the sequence.
[0025] Inside each direction (forward or backward), feature extraction is completed through the internal mechanism of a unidirectional GRU (gated recurrent unit). The GRU processes sequence data through three gating mechanisms: the reset gate, the update gate, and the candidate hidden state gate. The specific process is as follows: First, the reset gate determines whether to ignore the information of the previous hidden state when generating the candidate hidden state. It generates a value between 0 and 1 through the sigmoid function to control the influence of the previous hidden state. Next, the update gate determines how much information from the previous time step and how much new information from the current time step to retain in the current hidden state. Similarly, it generates a value between 0 and 1 through the sigmoid function to control the update ratio of the information. Finally, the candidate hidden state gate generates a new candidate hidden state through the tanh function based on the current input and the previous hidden state processed by the reset gate. This state combines the information of the current input and part of the previous hidden state. Finally, the hidden state is obtained by combining the previous hidden state and the candidate hidden state through the update gate in a certain proportion, gradually accumulating temporal dependence information.
[0026] This bidirectional processing mechanism enables BiGRU to better capture the complex dependence relationships in the time series, thereby improving the performance of the model.
[0027] The CBAM segment 50 includes a channel attention module and a spatial attention module. CBAM dynamically adjusts each part of the feature by gradually applying the channel attention and spatial attention mechanisms to achieve the localization and focusing on the key regions in the sequence data.
[0028] Figure 5 FIG. 2 shows a schematic diagram of a channel attention module according to an embodiment of the present invention.
[0029] As Figure 5 shown, the channel attention module first extracts the channel information of the feature map through global average pooling and global max pooling operations, respectively extracting the channel information of the feature map through average pooling and max pooling. These global pooling operations generate two different feature representations, capturing the global and local information of the feature map in the channel dimension respectively. Then, these two representations are processed through a shared multi-layer perceptron (MLP), the results of the processing are added together, and finally a sigmoid function is used to generate the channel attention weights. These weights are used to re-weight each channel of the original feature map, emphasizing important channel information and suppressing unimportant channels.
[0030] Figure 6 FIG. 3 shows a schematic diagram of a spatial attention module according to an embodiment of the present invention. As Figure 6 shown, according to an embodiment of the present invention, after generating the channel attention weights, the spatial attention mechanism is applied using the spatial attention module to adjust the spatial distribution of the feature map. The input of the spatial attention module is the output of the channel attention mechanism, that is, the feature map reshaped using the channel attention weights. Specifically, CBAM generates two different representations of spatial information by performing global average pooling and global max pooling operations on the spatial dimension of the feature map. These representations are processed through a convolutional layer and a sigmoid function to generate the spatial attention weights. These weights are used to re-weight each spatial position of the original feature map, emphasizing important spatial regions and suppressing unimportant regions.
[0031] CBAM combines the results of the channel attention and spatial attention mechanisms by sequentially performing channel attention and spatial attention, achieving a joint adjustment of the feature map. The channel attention and spatial attention are adjusted in a sequential manner. As mentioned above, first, channel attention weighting is performed on the feature map, and then the feature map reshaped by channel attention is input into the spatial attention, and spatial attention weighting is performed again using the generated spatial attention weights. This process dynamically adjusts the weights of each part in the feature map, achieving the localization and focusing of key regions in the sequence data, thereby improving the expressiveness and information utilization efficiency of the model.
[0032] After receiving the output of CBAM, the fully connected layer converts the input feature map into a specific prediction result through matrix multiplication and activation function. The prediction result is a vector composed of the predicted values for each time step
[0033] Then, in step S400, air pollutant and meteorological data for a predetermined time interval (e.g., the recent week) are input into the trained prediction model for prediction. It should be noted that in practice, the model can be divided into a training stage and a prediction stage. The training stage is to enable the model to learn the variation law of pollutants over a period of time, which can be saved as a.pth file for example. When making predictions later, only the data of the previous week needs to be input and the.pth file is loaded to make predictions and output the predicted values. It is not necessarily to train every time before making predictions. Sometimes, only the trained model (saved as a.pth file) needs to be loaded. The.pth file generated by training is time-sensitive. If the recent weather changes are stable and there are no extreme weather events, the model only needs to be retrained once a week or every 15 days to update the model; if there are extreme weather or other weather events with large changes, it needs to be retrained on the same day to learn the latest features.
[0034] During training, the final prediction results, that is, the specific predicted values, will also be output. For each round of training, the model will generate the learned weights and save them as a.pth file. It is precisely by comparing the fitting degree between the "output predicted value" and the "input true value" that it can finally be determined whether these weights are optimal. Therefore, the training process will output the predicted values to compare with the true values to judge the quality of the training results of this round.
[0035] The predicted values generated in each round of training are only used to evaluate the quality of the model training results, that is, whether the optimal weights have been learned. After all rounds of training are completed, the selected optimal weights will be saved as a.pth file for use in the prediction stage.
[0036] According to the embodiments of the present invention, by connecting the above four modules in series, the multi-level time and space features can be fully utilized to improve the accuracy of air pollutant concentration prediction. Specifically: TCN captures the long-term dependence relationship of the time series; SENet enhances the expression of important features through the channel attention mechanism and further enhances the output result of TCN; BiGRU fuses the bidirectional information of the past and the future to further enhance the extraction of time features; CBAM optimizes the model's understanding and expression of the time and space dimensions through the combination of spatial and channel attention. This series structure can enhance the model's feature extraction ability and prediction performance from multiple dimensions, enabling the model to better adapt to the complex variation law of air pollutant concentrations.
[0037] Furthermore, the sorting of the four modules forms a feature extraction chain with complementary functions and logical progression. Specifically: TCN is responsible for extracting temporal features but does not pay attention to the importance of each feature; SENet generates weights for features through channel attention, enabling the model to focus on important features and suppress unimportant features; BiGRU combines past and future bidirectional information, can extract features from longer-term data, and further enriches the temporal features but does not pay attention to the spatial distribution; CBAM makes up for the deficiency of BiGRU in spatial dimension feature extraction ability through spatial and channel attention.
[0038] Through this arrangement, the smoothness of information transmission is ensured, and problems such as gradient disappearance and information loss commonly found in deep models are avoided. Specifically: TCN has a residual connection that can effectively solve the problem of gradient disappearance in deep models, so TCN is placed at the first stage of the model; SENet performs feature recalibration to suppress the problem of feature redundancy, enabling information to be transmitted more efficiently in subsequent modules; the parallel computing ability and memory mechanism of BiGRU make the transmission of features in the temporal dimension smoother, avoiding the problem of information loss in deep networks; CBAM is a lightweight design, and placing it at the end of the model minimizes the delay of information transmission.
[0039] The data in the subscript illustrates the advantages of the implementation manner of the present invention that uses the above four segments in series compared to the cases of only selecting three segments or two segments.
[0040]
[0041] Root mean square error (RMSE) and mean absolute error (MAE) are model accuracy metrics used to calculate the difference between predicted values and true values. The smaller the two values, the higher the prediction accuracy of the model. It can be seen that our model has the highest accuracy.
[0042] The above description is only illustrative and not a limitation on the protection scope of the present invention. Any changes and substitutions within the concept of the present invention are within the protection scope of the present invention.
Claims
1. A method for predicting air pollutant concentration, characterized in that: The steps include: A pollutant data and meteorological data acquisition step, acquiring pollutant data and meteorological data within a predetermined time period before a time period to be predicted; The transformation step transforms the collected data into a spatiotemporal data matrix; as well as The prediction step uses the pollutant data and meteorological data collected within a predetermined time interval and a predetermined prediction model to predict the concentration of air pollutants at a predetermined time. The predetermined prediction model includes a temporal convolutional network segment, a squeeze excitation mechanism segment, a bidirectional gated recurrent unit segment and a convolutional block attention module segment connected in series.
2. The method for predicting air pollutant concentration according to claim 1, characterized in that: The temporal convolutional network segment includes a first branch, a second branch, and an adder. The first branch includes an expanded convolution layer, an activation function layer, and a padding operation layer. The second branch includes a 1x1 convolution layer. The input spatiotemporal data matrix is simultaneously input into the first branch and the second branch. The adder adds the output of the first branch and the output of the second branch.
3. The method for predicting air pollutant concentration according to claim 2, characterized in that: The first branch includes one or more dilated convolution layers. In each dilated convolution layer, a dilated convolution kernel is used to capture dependencies at different time scales. The result after the convolution operation is activated by the activation function layer to introduce nonlinear characteristics and enhance the expressiveness of the model. The padding operation layer adds zero values at the beginning of the sequence input to the current padding operation layer to ensure that the output of each time step only depends on the current and previous time steps, but not on future time steps.
4. The method for predicting air pollutant concentration according to claim 2, characterized in that: The squeeze-excitation mechanism stage performs feature transformation on the input features to generate an intermediate feature map, and the intermediate feature map then undergoes a squeeze operation and an excitation operation. The squeezing operation uses a global average pooling technique to average all values in each channel, thereby compressing the rich feature information into a single value, namely, a channel descriptor; The excitation operation transforms the channel descriptor through two fully connected layers. The first fully connected layer plays a role of dimensionality reduction, reducing the number of descriptors, and the second fully connected layer restores the data to the original number of channels. Then, the descriptor of each channel is converted into a weight value between 0 and 1 through an activation function.
5. The method for predicting air pollutant concentration according to claim 1, characterized in that: The pollutant data and the weather data within a predetermined time interval before the time period to be predicted include the pollutant data and the weather data of the location associated with the location to be predicted.
6. The method for predicting air pollutant concentration according to claim 1, characterized in that: The bidirectional gated recurrent unit segment includes a forward gated recurrent unit layer, a backward gated recurrent unit layer and a summation layer. The forward gated recurrent unit layer processes the input data from the starting position to the end position of the output of the temporal convolutional network segment in sequence to capture the temporal dependency of the sequence. At the same time, the backward gated recurrent unit layer processes the input data from the end position to the starting position of the sequence to capture the inverse dependency of the sequence. The output of the forward gated recurrent unit and the output of the backward gated recurrent unit are merged by the summation layer.
7. The method for predicting air pollutant concentration according to claim 1, characterized in that: The convolutional block attention module segment first extracts the channel information of the feature map through average pooling and maximum pooling, respectively capturing the global and local information of the feature map in the channel dimension. Then, these two representations are processed through a shared multi-layer perceptron, and finally the channel attention weight is generated through the activation function.
8. The method according to claim 1, characterized in that: The method further comprises a training step, wherein the training step utilizes the collected pollutant data and meteorological data to retrain the predetermined model at predetermined intervals in the absence of extreme weather.
9. The method according to claim 1, characterized in that: The pollutant data and meteorological data include pollutant data and meteorological data collected from cities near the location to be predicted, and the city is selected based on the distance from the location to be predicted, the transmissibility of the pollutant to be predicted, the population base of the city and the level of social activity.