Deep learning-based pollutant prediction method
By improving the GRU gating mechanism to C-GRUCell and combining convolutional networks and KAN networks for feature fusion, the problem of low prediction accuracy in existing models when dealing with data anomalies and extremes is solved, and higher pollutant prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510233944.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
Existing deep learning models are difficult to deal with data anomalies and extremes in pollutant prediction, resulting in low prediction accuracy and poor robustness.
By improving the GRU gated unit to C-GRUCell, and combining convolutional networks and residual networks to enhance the model expression ability, combining Encoder to extract spatial features, and feature fusion is performed through KAN networks to improve the prediction accuracy and robustness of the model.
It improves the accuracy and robustness of pollutant prediction, can process data under complex conditions more effectively, and improves the model's expression and generalization ability in extreme cases.
Smart Images

Figure CN120183534A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sequence prediction, and specifically relates to a method for pollutant prediction based on deep learning. Background Art
[0002] The sequence prediction task is a very classic task with wide application requirements, and it is widely used in fields such as urban traffic flow, meteorological data prediction, and population movement. For example, in the urban traffic flow prediction based on monitoring station data, the model analyzes the past traffic flow data to predict the future traffic conditions and helps optimize traffic signals and road network planning; in meteorological data prediction, the historical meteorological data of meteorological stations are modeled to predict the future change trends of temperature, humidity, etc. However, these models also have certain drawbacks.
[0003] In applications, the quality of monitoring station data may be affected by equipment failures, external interferences, or data transmission problems, resulting in abnormal acquisition of pollutant data; in addition, sudden environmental changes or extreme weather events may also lead to serious deviations in prediction results. Traditional deep learning models often have difficulty providing accurate pollutant predictions in such cases. Therefore, although deep learning models such as RNN, GRU, LSTM, and Transformer perform well in dealing with complex sequence prediction tasks, when using these models for pollutant prediction, there are still certain limitations in the face of data anomalies and extreme situations, resulting in low accuracy and poor robustness in predicting environmental data at future times by monitoring stations. Summary of the Invention
[0004] In order to improve the accuracy of model prediction and give full play to the role of deep learning in the field of sequence prediction, the present invention proposes a method for pollutant prediction based on deep deep learning. By utilizing the self-advantages of GRU, Encoder, and Decoder, and improving the gating unit of GRU, namely the C-GRUCell proposed in the present invention, the model expression ability is enhanced through convolutional networks and residual networks. And better spatial features are obtained by combining with Encoder, feature fusion is performed through the KAN network, and the final prediction output is obtained through Decoder. The present invention improves the model structure and the ability to process abnormal data to improve the accuracy and robustness of prediction.
[0005] The technical solution adopted by the present invention to achieve the above object is: a method for pollutant prediction based on deep learning, comprising the following steps:
[0006] 1) Collect data from each monitoring station and perform preprocessing;
[0007] 2) Based on the convolution-improved C-GRUCell network, capture the temporal features of the preprocessed sequence data, and at the same time capture the spatial features through the encoder;
[0008] 3) Fuse the temporal features and spatial features obtained in step 2) to get the input of the decoder, and make predictions through the decoder to obtain the prediction results of each monitoring site within the set future time period.
[0009] Pre-construct a sequence task dataset, construct a prediction model and train it, including the following steps:
[0010] 1.1) Construct a sequence task dataset:
[0011] Obtain the historical data of each monitoring site and preprocess it;
[0012] Then normalize the preprocessed historical data and perform preprocessing on time feature segmentation, that is, split the historical data according to year, month, and day for time features;
[0013] Embed the segmented preprocessed time features into the historical data through word embedding to construct a sequence task dataset;
[0014] 1.2) For the sequence data in the sequence task dataset, based on the convolution-improved C-GRUCell network, capture the temporal features, and at the same time capture the spatial features through the encoder;
[0015] 1.3) Fuse the temporal features and spatial features obtained in step 1.2) to get the input of the decoder, and obtain the prediction results of each monitoring site within the set future time period through the decoder; among them, the prediction model is obtained through steps 1.2) to 1.3).
[0016] The convolution-improved C-GRUCell network captures temporal features as follows:
[0017] Enhance the data features of the input data through a feature enhancement network;
[0018] For the sequence data of each enhanced monitoring site, after embedding the time features through the embedding layer encoding, based on the convolution-improved C-GRUCell network, capture the temporal feature vectors;
[0019] Fuse the temporal feature vectors of all detection sites to obtain the temporal features.
[0020] Capturing spatial features through the encoder is as follows:
[0021] After encoding the input data through the embedding layer in the encoder and embedding the time features, spatial feature learning between stations is performed through the encoder, and then through the multi-layer perceptron layer MLP to obtain spatial features.
[0022] The working process of the convolution-improved C-GRUCell network is as follows:
[0023] Input the current time step input sequence x t After performing a 1*1 convolution operation and a residual connection operation, a new sequence x of the same size as x is obtained t ;
[0024] The sequence x t is combined with the hidden state H of the previous time step of the hidden state t-1 to obtain the input of the gating mechanism, which respectively enters the update gate and the reset gate; the working processes of the update gate and the reset gate are as follows:
[0025] First, the input sequence data x t passes through the first convolutional layer to obtain the output of the update gate after convolutional transformation, and through the first sigmod activation function to obtain the updated output z of the update gate, so as to update the hidden state by representing the input of the current time step and the state of the previous time step;
[0026] The input sequence data passes through the second convolutional layer to obtain the output of the reset gate after convolutional transformation, and through the second sigmod activation function to obtain the updated output of the update gate. The output obtained by the reset gate is combined with the hidden state to obtain the updated output r of the reset gate, which is used to determine the retention of the state information of the previous time step;
[0027] Finally, the output r of the reset gate and the sequence x t are combined, and a new hidden state H is obtained through tanh. Finally, through the hidden state calculation formula H of the gating mechanism t =(1 - z)*H + z*H t-1 the final output H of the convolution-improved C-GRUCell network is obtained t , where t represents the time step; the shape of the output of the convolution-improved C-GRUCell network is the same as the input shape.
[0028] In step 3), the fusion of the temporal feature and the spatial feature obtained in step 2) to obtain the input of the decoder is specifically:
[0029] Using the temporal feature and the spatial feature obtained in step 2), through fusion, an input sequence with spatio-temporal features is obtained, which is used as the input to further enhance the data through the KAN network for the decoder, as the input of the decoder.
[0030] The data of each monitoring station includes meteorological characteristic data, time characteristics, and pollutant concentrations; the meteorological data characteristics include temperature, humidity, air pressure, wind speed, and wind direction; the time characteristics are time points; the pollutant concentrations include PM2.5, sulfides, and nitrides.
[0031] A sequence prediction system based on deep learning includes:
[0032] A time acquisition module for collecting data from each monitoring station and performing preprocessing;
[0033] A spatio-temporal feature capture module for capturing temporal features of the preprocessed sequence data based on the convolution-improved C-GRUCell network, and simultaneously capturing spatial features through an encoder;
[0034] A prediction output module for fusing the temporal features and spatial features to obtain the input of the decoder, and performing prediction through the decoder to obtain the prediction results for each monitoring station in a future set time period.
[0035] A pollutant prediction method based on deep learning includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the described pollutant prediction method based on deep learning when executing the computer program.
[0036] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the described pollutant prediction method based on deep learning is implemented.
[0037] The present invention has the following beneficial effects and advantages:
[0038] 1. Based on deep learning, the present invention uses the existing basic algorithms of GRU, Encoder, and Decoder, and uses convolution operations to improve GRU, which can effectively enhance the expression ability of the model, improve the prediction accuracy of the model, and improve the prediction accuracy of pollutants under complex conditions by enhancing the expression and generalization ability of the model under complex conditions.
[0039] 2. The present invention improves the GRU gating mechanism through a 1*1 convolutional layer, enhances the input data, and introduces a residual connection. Modifying the linear transformation of the traditional GRU gating mechanism can enhance the expression ability and robustness of the model.
[0040] 3. The KAN network is used to fuse the spatial features extracted by the Encoder and the temporal features extracted by the C-GRUCell, improving the prediction accuracy and efficiency of future environmental information. Description of the Drawings
[0041] Figure 1It is the method flow chart of the present invention;
[0042] Figure 2 It is the schematic diagram of the C-GRUCell structure improved based on convolution;
[0043] Figure 3 It is the schematic diagram of the deep learning model structure. Specific embodiments
[0044] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0045] As Figure 1 shown, a pollutant prediction method based on deep learning includes the following steps:
[0046] Step 1: Collect sequence prediction data.
[0047] Step 2: Preprocess the data, including processing abnormal data, for example: filling empty data to obtain a data set for training and prediction.
[0048] Step 3: Enhance the features of the input sequence. Taking the prediction of meteorological data as an example, the input shape is a matrix composed of stations, time intervals, feature dimensions, and pollutant compositions; the rows represent the input time intervals, and the columns represent the feature dimensions, where the features include temperature, humidity, air pressure, wind speed, and wind direction.
[0049] Step 4: For the spatial feature extraction part, perform word embedding processing on the data of each station, embed the time dimension and input features, and enhance the interpretability of the data.
[0050] Step 5: For the time series feature extraction part, first perform stationarity processing (data standardization processing) on the data.
[0051] Step 6: Use time feature + spatial feature fusion (the fusion method here is the dot product method) to obtain input data with spatio-temporal features.
[0052] Step 7: Obtain the final prediction output through the decoder; that is, if it is necessary to predict the meteorological data of n monitoring stations, then the obtained prediction output data is the prediction results of n stations for a future period of time (such as 24 hours).
[0053] The collected data is the data of air monitoring stations, and at the same time, data from multiple stations is received,
[0054] The data preprocessing is to separate the time features and normalize the collected data.
[0055] The data enhancement is to use the data enhancement KAN network to train and enhance the features of the input sequence.
[0056] The described time features are processed using the convolution-improved C-GRUCell in the invention. The spatial features are implemented using the encoder of the Transformer, and other KAN networks are used for feature enhancement.
[0057] The final predicted output obtained by the decoder is the output obtained by training the new features that fuse temporal and spatial features using the decoder.
[0058] As Figure 2 shown, the specific process of the convolution-improved C-GRUCell network is as follows:
[0059] Step 1: Input the current time-step input sequence x t After passing through a 1*1 convolution operation and performing a residual connection operation, a sequence x t of the same size as x t is obtained;
[0060] Step 2: Combine the sequence x t with the hidden state H t-1 to obtain the input of the gating mechanism, which respectively enters the update gate and the reset gate. The specific processes of the update gate and the reset gate are as follows: First, input the sequence x t through the Conv1d layer ( Figure 2 the first conv1d from right to left in
[0061] ), obtain the output of the updated update gate after convolution transformation, and obtain the updated output z of the update gate through the sigmod activation function; t-1 Step 3: Input the sequence (hidden state H t and x Figure 2 ) through the Conv1d layer (
[0062] the second conv1d from right to left in t ), obtain the output of the reset gate after convolution transformation, and obtain the updated output of the update gate through the sigmod activation function. Combine the output obtained by the reset gate with the hidden state to obtain the updated output r of the reset gate; t Step 4: Finally, combine the output r of the reset gate with the input sequence x t-1 and obtain the new hidden state H through tanh. Finally, obtain the final output of the C-GRUCell through the hidden state calculation formula H
[0063] During the training phase of the C-GRUCell network, if it is not the last time step, the input is given to the next time step to continue the calculation of the current convolution-improved C-GRUCell network. If it is the last time step, the calculation ends and then enters Figure 2 the next convolution-improved C-GRUCell network for training.
[0064] As Figure 3 shown, the sequence prediction method based on deep learning:
[0065] Step 1: Collect the meteorological data collected to form a training data set.
[0066] The data in the data set comes from the monitoring data of each monitoring site, such as meteorological data: wind speed, wind direction, etc., pollutant concentrations: PM2.5, sulfides, nitrides, etc. The data set is a matrix composed of site numbers, time intervals, feature dimensions, and pollutant concentration values; the rows represent the input time intervals, feature values, and pollutant values corresponding to the sites; the columns represent the feature dimensions; among them, the features include temperature, humidity, air pressure, wind speed, and wind direction; the time interval is a set time period to represent the corresponding monitoring data collected within this time interval; the longitude and latitude information of the monitoring site can also be added to the feature dimensions.
[0067] Step 2: Perform data set preprocessing, separate the time and longitude and latitude information, and form an input sequence x (corresponding to monitoring site 1, monitoring site 2, monitoring site 3... monitoring site n in the figure). The process is as follows:
[0068] 1). Obtain the historical data of each monitoring site and perform preprocessing;
[0069] 2). Normalize the preprocessed historical data and perform time feature segmentation preprocessing, that is, split the historical data according to years, months, and days for time features;
[0070] 3). Embed the segmented preprocessed time features into the historical data through word embedding to construct a sequence task data set.
[0071] Step 3: Preprocess the training data and perform stationarity processing on the data in Figure 3 the left part (time series feature extraction part), and the specific implementation is: calculate the mean and variance ( where x ′ is the processed data, x represents the input data, represents the mean, and δ represents the variance).
[0072] Step 4: Embed the time information and location information into the input sequence through word embedding ( Figure 3 the embedding part in to obtain input data that is more suitable for model training.
[0073] Step 5: Extract spatial features through the improved C-GRUCell and use the encoder to extract spatial features. The specific steps are as Figure 2 described.
[0074] Step 6: After passing through the C-GRUCell model, perform detrending on the output of the temporal features and perform feature enhancement through the KAN network. Then, splice the results of each monitoring site to obtain a multi-site temporal feature (number of sites, time interval, feature). Among them, x ′ is the processed data, x represents the input data, represents the mean, and δ represents the variance.
[0075] Step 7: In Figure 3 the left spatial feature extraction part, pass the obtained input data (monitoring site 1, monitoring site 2, monitoring site 3... monitoring site n) through word embedding processing (embed time and location information into the input data to obtain more suitable input data for training).
[0076] Step 8: The training data passes through the encoder Encoder and the multi-layer perceptron mlp to obtain an output sequence with spatial features (here is the mutual influence relationship between sites, and the specific shape is (time interval, number of sites, number of sites).
[0077] Step 9: Fuse the obtained temporal and spatial features (here is the fusion through dot product operation) to obtain the input of the final prediction output module. At this time, the shape is (number of sites, time interval, feature).
[0078] Step 10: For the input data obtained in Step 9, perform data enhancement on the data of different sites through the corresponding KAN network respectively, and send the obtained results to the decoder Decoder.
[0079] Step 11: Combine the prediction results obtained through each decoder Decoder and each KAN network (here the KAN network is a new KAN network, not the same as the network in Step 6) to obtain the final prediction output result out; at this time, the shape is the same as the shape of the input data, both are (number of sites, time interval, feature, pollutant concentration. Here the prediction result is the meteorological data of each monitoring site, such as meteorological data (wind speed, wind direction, etc.) and pollutant concentration (PM2.5, sulfide, nitride, etc.).
[0080] Step 12: Obtain a prediction model for pollutant prediction through Steps 1-11;
[0081] Step 13: Collect data from each monitoring point in real time. After preprocessing through Step 3, perform prediction using the obtained prediction model to obtain the pollutant sequence for a future period of time.
[0082] Example:
[0083] The modeling steps of the present invention are as follows:
[0084] Step 1: The data of the present invention is sourced from meteorological data monitoring stations. The data includes the collected meteorological information, including temperature, humidity, air pressure, wind speed, and wind direction, etc.
[0085] Step 2: Data preprocessing. Preprocess the collected data. The specific operations are as follows: fill in missing values, separate time and location information, and perform time feature segmentation processing by separating time (in the subsequent word embedding stage, the separated time features will be embedded into the task sequence through word embedding encoding):
[0086] The information contained in the collected data is: meteorological data features (such as wind speed and wind direction) + time features (time points) + pollutant concentration information (such as PM2.5, sulfides, nitrides, etc.). Split the time points by year, month, and day to better obtain its periodicity (such as seasons, months, weeks, etc.).
[0087] Then embed the processed time features into the task sequence through the word embedding (embedding encoding, which is achieved by constructing an embedding matrix for the time features) method to construct a sequence task data set; among them, for the embedding achieved through word embedding encoding, the main operation is to obtain a matrix regarding time expression through embedding encoding and splice this matrix into the feature matrix.
[0088] Step 3: Divide the preprocessed data set into a training set, a validation set, and a test set according to a ratio of 4:1:1.
[0089] Step 4: Enhance the features of the processed data set through the KAN network.
[0090] Step 5: Embed time and location information into the input sequence through word embedding processing to facilitate the capture of periodic features.
[0091] Step 6: Capture time features through the improved C-GRUCell.
[0092] Enhance the data features of the sequence data through the KAN network;
[0093] For the sequence data of each enhanced monitoring site, after embedding the time features through the embedding layer encoding, then based on the convolutional improved C-GRUCell network, capture the time series feature vectors;
[0094] Fuse the time series feature vectors of all detection sites to obtain time series features.
[0095] Step 7: Perform spatial feature capture through an encoder.
[0096] Step 8: Fuse the captured temporal and spatial features as the input sequence and feed it into the decoder for the final prediction output.
Claims
1. A pollutant prediction method based on deep learning, characterized in that: The following steps are involved: 1) Collect data from each monitoring site and pre-process it; 2) The preprocessed sequence data is used to capture the temporal features based on the convolution-improved C-GRUCell network, and the spatial features are captured through the encoder; 3) The temporal features and spatial features obtained in step 2) are integrated to obtain the input of the decoder, and the decoder is used to perform prediction to obtain the prediction results of each monitoring site within a set time period in the future.
2. The pollutant prediction method based on deep learning according to claim 1, characterized in that: Pre-build a sequence task dataset, build a prediction model and train it, including the following steps: 1.1) Constructing sequence task dataset: Obtain historical data from each monitoring site and perform preprocessing; Then the preprocessed historical data is normalized and preprocessed by time feature segmentation, that is, the historical data is split into time features according to year, month, and day; The time features after segment preprocessing are embedded into historical data through word embedding to construct a sequence task dataset; 1.2) For the sequence data in the sequence task dataset, the C-GRUCell network improved based on convolution captures the temporal features, and the encoder captures the spatial features; 1.3) The temporal features and spatial features obtained in step 1.2) are integrated to obtain the input of the decoder, and the prediction results of each monitoring station within a set time period in the future are obtained through the decoder; wherein, the prediction model is obtained through steps 1.2) to 1.3).
3. A pollutant prediction method based on deep learning according to claim 1 or 2, characterized in that: The C-GRUCell network based on convolution improvement captures the timing features, as follows: The input data is enhanced through the feature enhancement network; For the enhanced sequence data of each monitoring site, the temporal features are embedded through the embedding layer, and the convolution-improved C-GRUCell network is used to capture the temporal feature vector. The time series feature vectors of all detection sites are fused to obtain the time series features.
4. A pollutant prediction method based on deep learning according to claim 1 or 2, characterized in that: The encoder captures spatial features as follows: After the input data is embedded through the embedding layer in the encoder, the temporal features are embedded, and then the spatial features are learned between sites through the encoder, and then the spatial features are obtained through the multi-layer perceptron layer MLP.
5. A pollutant prediction method based on deep learning according to claim 1 or 2, characterized in that: The convolution-improved C-GRUCell network workflow is as follows: Enter the current time step into the sequence x t After a 1*1 convolution operation and a residual connection operation, a new sequence x with the same size as x is obtained. t ; The sequence x t The hidden state H of the previous time step t-1 The input of the gating mechanism is combined and enters the update gate and reset gate respectively; the workflow of the update gate and reset gate is as follows: First, the input sequence data x t After the first convolutional layer, the update gate output after the convolution transformation is obtained, and the updated update gate output z is obtained through the first sigmoid activation function to represent the input of the current time step and the state of the previous time step to update the hidden state; The input sequence data is passed through the second convolutional layer to obtain the output of the reset gate after the convolution transformation, and the updated update gate output is obtained through the second sigmoid activation function. The output obtained by the reset gate and the hidden state are combined to obtain the updated reset gate output r, which is used to determine the state information retention of the previous time step; Finally, the output r of the reset gate and the sequence x t Combined, and the new hidden state H is obtained through tanh, and finally the hidden state calculation formula H of the gating mechanism is used t =(1-z)*H+z*H t-1 Get the final convolution-improved C-GRUCell network output H t , where t represents the time step; the shape of the output of the convolution-improved C-GRUCell network is the same as the input shape.
6. The pollutant prediction method based on deep learning according to claim 1, characterized in that: In step 3), the temporal features and spatial features obtained in step 2) are fused to obtain the input of the decoder, specifically: Using the temporal features and spatial features obtained in step 2), the input sequence with temporal and spatial features is fused and further data enhanced through the KAN network for the decoder as the input of the decoder.
7. The pollutant prediction method based on deep learning according to claim 1, characterized in that: The data of each monitoring station include meteorological characteristic data, time characteristics and pollutant concentration; the meteorological data characteristics include temperature, humidity, air pressure, wind speed and wind direction; the time characteristics are time points; the pollutant concentrations include PM2.5, sulfides and nitrogen compounds.
8. A sequence prediction system based on deep learning, characterized in that: include: Time collection module, used to collect data from each monitoring site and perform preprocessing; The spatiotemporal feature capture module is used to capture the temporal features of the preprocessed sequence data based on the convolution-improved C-GRUCell network, and capture the spatial features through the encoder; The prediction output module is used to fuse the temporal features and spatial features to obtain the input of the decoder, and to perform prediction through the decoder to obtain the prediction results of each monitoring station for the set time period in the future.
9. A pollutant prediction method based on deep learning, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement a pollutant prediction method based on deep learning as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, a pollutant prediction method based on deep learning as described in any one of claims 1 to 6 is implemented.