PM2.5 concentration prediction method based on LSTM and attention mechanism
By combining empirical modal decomposition and LSTM+ attention mechanism PM2.5 concentration prediction model, the problems of high computational costs and inaccurate prediction in the existing technology are solved, and high-precision and stable PM2.5 concentration prediction is achieved, supporting environmental monitoring and pollution warning.
Patent Information
- Application Number
- CN202510368130.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems such as high calculation cost, poor practicality and inaccurate prediction of PM2.5 concentration, especially in real-time monitoring systems.
Combining the empirical modal decomposition method with LSTM and attention mechanism, a PM2.5 concentration prediction model is constructed. By preprocessing and decomposing historical data, the attention mechanism is used to automatically select important information in the time series to improve the prediction accuracy and stability of the model.
It improves the accuracy and stability of PM2.5 concentration prediction, provides more effective environmental monitoring and pollution warning support, and is suitable for real-time monitoring systems.
Smart Images

Figure CN120280027A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a PM concentration prediction method based on LSTM and attention mechanism, belonging to the technical field of environmental monitoring and prediction. 2.5 Background Technique
[0002] PM 2.5 refers to particulate matter with a diameter less than or equal to 2.5 micrometers in the air. Because it is extremely easy to enter the human respiratory system and pose a serious threat to human health, it has become an important environmental problem globally. With the acceleration of industrialization and urbanization, the concentration of PM 2.5 is increasing day by day, causing the deterioration of air quality and having a profound impact on public health and the environment. Therefore, accurately predicting the PM 2.5 concentration has become an important task for air pollution control and environmental management.
[0003] PM 2.5 formation and influencing factors are complex. At present, the concentration prediction of PM is roughly divided into two categories: simulation-driven models and data-driven prediction models. In terms of simulation-driven models, it mainly relies on analyzing and mining the relationship between the chemical and physical components of the atmosphere and the PM 2.5 concentration, by analyzing the source of PM 2.5 , simulating the generation mechanism and atmospheric diffusion mode of PM 2.5 , and simulating the formation, propagation and conversion process of PM 2.5 particles in the environment from the perspective of the occurrence mechanism. Traditional physical models may have certain limitations in complex spatio-temporal relationships and environmental changes. Because of this, data-driven prediction models become particularly important. Data-driven methods mainly rely on relevant theories such as mathematics and statistics, and mine the change laws of pollutants through historical PM 2.5 data and related data characteristics, and model and predict the PM 2.5 concentration from a non-mechanistic perspective.
[0004] The simulation-driven method for prediction has the characteristics of good real-time performance and high accuracy. However, the physicochemical reactions of the models involved in atmospheric science are relatively complex, resulting in the need for complex algorithms and computing technologies for the coordination and integration of models, and highly relying on a large amount of real-time meteorological data, emission data and other data resources, which require high computer hardware and software, resulting in high cost and poor practicability of this method, and it is not suitable for large-scale applications.
[0005] Data-driven models (such as machine learning and deep learning) can capture PM 2.5The non - linear and complex relationship between concentration and various variables, without relying on precise physical models, is more adaptable. However, it requires a large amount of computing resources during training and operation, which may pose challenges to some real - time monitoring systems. Moreover, data - driven models highly depend on the quality of input data, and missing, noisy, or inconsistent data may lead to inaccurate prediction results. Summary of the Invention
[0006] To solve the technical problems existing in the prior art, the present invention proposes a prediction method combining the empirical mode decomposition method with the LSTM + attention mechanism.
[0007] To achieve the above - mentioned purpose, the technical solution proposed by the present invention is as follows:
[0008] A PM 2.5 concentration prediction method based on LSTM and attention mechanism, comprising the following steps:
[0009] Step 1, obtain PM 2.5 historical data and pre - process the historical data;
[0010] Step 2, perform empirical decomposition on the processed PM 2.5 historical data to form a data set;
[0011] Step 3, construct an LSTM model based on the attention mechanism, and the model includes an input layer, an LSTM layer, an attention mechanism layer, a merging layer, a fully - connected layer, and an output layer;
[0012] Step 4, use the data set to train the constructed model with the Adam optimizer and the mean - square error loss function, and the optimization objective of training is to minimize the error between the predicted value and the true value to obtain a trained model;
[0013] Step 5, input the PM 2.5 data collected in real - time into the trained model for PM 2.5 concentration prediction.
[0014] For the further design of the above - mentioned technical solution: The attention mechanism layer takes the output of the LSTM layer as input and assigns different weights to each time step.
[0015] The output of the LSTM layer and the output of the attention layer are merged through the merging layer, and the merged result enters the fully - connected layer. The fully - connected layer contains 32 neurons and uses the ReLU activation function for non - linear transformation.
[0016] A Dropout layer is provided after the fully - connected layer.
[0017] The PM 2.5 historical data includes PM2.5 Concentration data, as well as temperature, humidity, wind speed, and boundary layer height data related to air quality.
[0018] The preprocessing is to normalize the features and target values of the input historical data, construct a time series dataset based on the time step, and use the time point data as input features.
[0019] The beneficial effects of the present invention are as follows:
[0020] The present invention uses empirical mode decomposition technology to denoise and decompose the PM 2.5 historical data, thereby improving the prediction accuracy and stability of the model. By combining the advantages of empirical mode decomposition and the LSTM + attention mechanism model, the most relevant information is automatically selected in the time series, improving the model's ability to understand the time series, and thus improving the accuracy and stability of PM 2.5 concentration prediction, providing more effective technical support for environmental monitoring and pollution warning systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the present invention;
[0022] Figure 2 is a prediction scatter plot corresponding to three deep learning models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] Embodiment
[0025] The PM concentration prediction method based on LSTM and attention mechanism in this embodiment, as 2.5 shown, includes the following steps: Figure 1 shown, includes the following steps:
[0026] Step 1: Data collection and preprocessing
[0027] Collect the PM historical data of the SGP C1 site in Oklahoma from 2021 to 2023, including PM 2.5 concentration data and auxiliary data related to air quality (such as temperature, humidity, wind speed, boundary layer height, etc.), and extract the features in the data and define the target column. Perform missing value processing on the data, and use normalization to remove noise and seasonal factors to obtain a stable and unified time series data. Normalization helps to improve the training efficiency of the LSTM model because the LSTM model is sensitive to the scale of the data. And construct a time series dataset based on the time step, use the time point data as input features, and predict the PM 2.5 value at the next time point. 2.5 value.
[0028] Step 2: Perform empirical decomposition on the processed PM 2.5 historical data to form a dataset;
[0029] Perform empirical decomposition on the PM 2.5 historical concentration data to remove the long-term trend, seasonal variations, and noise, obtaining the residuals and periodic components. Further analyze the decomposed data to ensure its stationarity, which is suitable for training the deep learning model.
[0030] The empirical mode decomposition technique (EMD, Empirical Mode Decomposition) is mainly used to process non-linear and non-stationary signals. This technique decomposes a complex signal into a series of intrinsic mode functions (IMFs), where each IMF represents different time-scale components in the signal. When performing empirical mode decomposition on the PM 2.5 data, first decompose the original PM 2.5 time series data. The EMD method adaptively decomposes the signal into multiple intrinsic mode functions (IMFs), and each IMF represents the fluctuating characteristics of different frequencies in the data, providing a clearer perspective for subsequent analysis and prediction.
[0031] When processing time series data, it is usually necessary to use past time points (historical data) to predict future values, and construct feature and label data through a sliding window method to ensure that the model can predict future values based on past information. In this embodiment, the window starts from the first time point, selects the first 12 time points as input features, and predicts the PM 2.5 value at the 13th time point. Then, the window slides one time step, selects the data from time point 2 to 13 as input, and predicts the value at the 14th time point, and so on.
[0032] Divide the dataset into a training set (80%) and a test set (20%), and split it in chronological order, that is, without shuffling the data order, to preserve the continuity of the time series.
[0033] Step 3: Build a deep learning model. In this embodiment, a model with LSTM + attention mechanism is adopted.
[0034] The structure of the model includes an input layer, an LSTM layer, an attention mechanism layer, a merging layer, a fully connected layer (Dense layer), and an output layer.
[0035] The input of the LSTM layer is data with a certain number of time points and features. This LSTM layer uses 32 units and is set to output only the hidden state at the final moment. The output layer is used to output the predicted PM 2.5 value.
[0036] The attention mechanism is implemented through the Attention() layer. This layer takes the output of the LSTM layer as input and assigns weights to different time steps by calculating scores related to each time step. The attention mechanism focuses on the key features of the input sequence through these weights, thereby affecting the final output so that the model can focus on the important parts of the input sequence. In this way, the attention mechanism enables the model to automatically select the most relevant information in the time series, thus improving the model's ability to understand time series.
[0037] The output of the LSTM layer and the output of the attention layer are connected through a merging layer to form a new feature representation. Next, the merged result enters a fully connected layer, which contains 32 neurons and uses the ReLU activation function for non-linear transformation to enhance the model's expressive ability. To prevent overfitting, a Dropout layer is added after the fully connected layer, whose role is to randomly discard neuron connections, thereby improving the model's generalization ability. The model is also trained using the Adam optimizer and the mean squared error (MSE) loss function, with the learning rate set to 0.01, and the optimization goal is to minimize the error between the predicted value and the true value.
[0038] Step 4, Training and Evaluation. The model is trained for 50 epochs using the training set and validated using the test set. During the training process, the LSTM will learn the mapping relationship between the input sequence and the target.
[0039] Step 5, Apply the trained optimal prediction model to the actual PM 2.5 concentration prediction system. Input the real-time collected PM 2.5 concentration data into the trained model for concentration prediction, providing hourly forecasts to prevent and control air pollution.
[0040] Comparative Example
[0041] In this example, the LSTM model, the LSTM+Transformer model, and the LSTM+attention mechanism model described in the above embodiment are respectively used for PM 2.5 concentration prediction; explore the impact of different LSTM variants on the prediction accuracy.
[0042] Among them, the LSTM model adopts the LSTM architecture, and the input of the LSTM layer is data with a certain number of time points and features. This LSTM layer uses 32 units and is set to output only the hidden state at the final moment. The LSTM layer is followed by a fully connected layer with 32 units and the ReLU activation function, as well as an output layer for predicting the PM 2.5 value. And the Adam optimizer is used, with the learning rate set to 0.01 and the loss function being the mean squared error.
[0043] The construction of the LSTM+Transformer model is similar to that of the above LSTM model. The difference is that the model structure includes an input layer, an LSTM layer, a Transformer layer, a merging layer, a fully connected layer (Dense layer), and an output layer. The Transformer layer is introduced to enhance the self-attention mechanism of the model. This layer calculates attention in different subspaces through multiple attention heads, capturing the dependencies between different time steps in the input data. This mechanism enables the model to focus on the more important parts of the input data, helping to improve the modeling effect of long time series data. In addition, [layer name] is used to prevent overfitting, while [layer name] performs normalization processing to help stabilize the training process. Then, [layer name] flattens the multi-dimensional output into one dimension for passing to the fully connected layer. The model is compiled using the Adam optimizer, with a learning rate set to 0.01 and the mean squared error used as the loss function.
[0044] The above three models are trained using the training set data, and the model parameters are adjusted according to the test set to optimize the model performance. Techniques such as cross-validation and early stopping are used to avoid overfitting. Evaluate the performance of different models in the hourly prediction task of PM 2.5 concentration, and determine the best model through comparative analysis.
[0045] Specifically: make predictions on the test data set and evaluate the prediction accuracy of the model (such as root mean square error RMSE, mean absolute error MAE, coefficient of determination R 2 etc.). Further adjust the model based on the evaluation results to improve the prediction accuracy. Among them, the coefficient of determination R 2 , RMSE, and MSE values of the LSTM model with the attention mechanism are 0.835, 1.709 μg / m3, and 2.92 μg / m3 respectively. Compared with other models, the LSTM with the attention mechanism performs better. The scatter plots are as Figure 2 shown. Figure (a) is the LSTM model, (b) is the LSTM model with the attention mechanism, and (c) is the LSTM model with the Transformer.
[0046] From the prediction results, it can be seen that the prediction results of the LSTM+attention mechanism model in the above embodiments are the most accurate.
[0047] The technical solutions of the present invention are not limited to the above embodiments, and all technical solutions obtained by equivalent replacement fall within the scope of protection required by the present invention.
Claims
1. A PM concentration prediction method based on LSTM and attention mechanism, characterized in that, 2.5 It includes the following steps: Step 1, obtain PM 2.5 historical data and preprocess the historical data; Step 2. Perform empirical decomposition on the processed PM 2.5 historical data to form a dataset; Step 3: Construct an LSTM model based on the attention mechanism, where the model includes an input layer, an LSTM layer, an attention mechanism layer, a merging layer, a fully connected layer, and an output layer; Step 4: Use the dataset to train the constructed model with the Adam optimizer and the mean squared error loss function. The optimization objective of the training is to minimize the error between the predicted value and the true value, and obtain a trained model; Step 5: Input the PM 2.5 data collected in real time into the trained model for PM 2.5 concentration prediction.
2. The PM concentration prediction method based on LSTM and attention mechanism according to claim 1 2.5 , characterized in that: The attention mechanism layer takes the output of the LSTM layer as input and assigns different weights to each time step.
3. The PM concentration prediction method based on LSTM and attention mechanism according to claim 2 2.5 is characterized in that: The output of the LSTM layer and the output of the attention layer are merged through the merging layer, and the merged result enters the fully connected layer. The fully connected layer contains 32 neurons and uses the ReLU activation function for non-linear transformation.
4. The PM concentration prediction method based on LSTM and attention mechanism according to claim 3 2.5 , characterized in that: A Dropout layer is provided after the fully connected layer.
5. The PM concentration prediction method based on LSTM and attention mechanism according to claim 1, characterized in that: 2.5 PM 2.5 The historical data includes PM 2.5 concentration data, as well as temperature, humidity, wind speed, and boundary layer height data related to air quality.
6. The PM concentration prediction method based on LSTM and attention mechanism according to claim 5 2.5 , characterized in that: The preprocessing is to normalize the features and target values of the input historical data, construct a time series dataset based on the time step length, and use the time point data as input features.