Air pollutant prediction method and system based on multi-modal deep learning
By combining multimodal deep learning methods with multi-layer perceptron, bidirectional long and short-term memory network and multi-head attention mechanism, the problems of unstable data quality and low prediction accuracy in air pollutant prediction are solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510060647.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems such as unstable data quality, complex and changeable environmental factors, and undesirable prediction accuracy in the prediction of air pollutant concentrations, and has failed to fully utilize the advantages of multimodal deep learning.
A hybrid neural network model is constructed for preprocessing and analyzing air quality data and dynamically adjusting the learning rate through the improved Adam algorithm.
It improves the accuracy and reliability of air pollutant prediction, can effectively capture complex dependencies in time series data, enhances the robustness and stability of the model, and shows high accuracy in the prediction results.
Smart Images

Figure CN120069170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for predicting air pollutants, and particularly to a method and system for predicting air pollutants based on multi-modal deep learning, belonging to the technical field of regional air pollutant prediction. Background Art
[0002] Air pollution is an important public health issue globally, and air pollutants have a significant impact on human health, especially the respiratory system. With the rapid development of science and technology, accurately predicting the changing trend of air pollutant concentration has become a key task in environmental monitoring and pollution prevention and control.
[0003] In recent years, deep learning technology has made remarkable progress in the field of regional air pollutant prediction. Among them, the multi-layer perceptron (MLP) can effectively process and extract static features; the bidirectional LSTM (BiLSTM) can more comprehensively understand the time-dependent features of time series data by considering both the forward and backward information of the sequence; in addition, the introduction of the attention mechanism provides a new research direction for deep learning models. The multi-head attention mechanism can capture complex patterns in the sequence from different representation subspaces by parallel processing multiple attention calculations, further improving the feature extraction ability of the model.
[0004] Most existing studies focus on the application of single technology or adopt a simple model combination method, failing to fully utilize the advantages of multi-modal deep learning. At the same time, air pollutant concentration prediction also faces challenges such as unstable data quality, complex and variable environmental factors, and less than ideal prediction accuracy. Therefore, there is an urgent need to study a multi-modal deep learning prediction method that can combine the multi-layer perceptron, bidirectional LSTM, and multi-head attention mechanism to improve the accuracy and reliability of air pollutant prediction. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to provide a method and system for predicting air pollutants based on multi-modal deep learning that can improve prediction accuracy.
[0006] Technical Solution: A method for predicting air pollutants based on multi-modal deep learning according to the present invention includes:
[0007] (1) Obtain historical air quality data and make a data set;
[0008] (2) Preprocess the data set;
[0009] (3) Based on the bidirectional long short-term memory network BiLSTM, use the multi-layer perceptron MLP and the multi-head attention mechanism Multi-Head Attention to construct a hybrid neural network model;
[0010] (4) Divide the preprocessed dataset into a training set, a validation set, and a test set according to a preset ratio;
[0011] (5) Use the training set to train the hybrid neural network model;
[0012] (6) Evaluate the trained hybrid neural network model;
[0013] (7) Repeat the training and evaluation steps in steps (5) to (6) until the evaluation result is greater than the preset threshold, output the MLP-BiLSTM-MHAT prediction model, and use the MLP-BiLSTM-MHAT prediction model to predict the air to be measured.
[0014] Further, the step (2) includes:
[0015] (21) Use the mean filling method to process continuous missing values and the backward filling method to process non-continuous missing values; the calculation formula of the mean filling method is as follows:
[0016]
[0017] where, mean j represents the mean of the j-th column of the dataset, x ij represents the i-th non-missing value in the j-th column, n j represents the number of non-missing values in the j-th column;
[0018] (22) For the dataset after processing missing values, use the interquartile range method to process abnormal data:
[0019] sup = Q3 + (θ * iqr)
[0020] inf = Q1 - (θ * iqr)
[0021] where, sup represents the maximum value allowed in the data, and points exceeding this value are regarded as abnormal, inf represents the minimum value allowed in the data, and points below this value are regarded as abnormal, Q1 and Q3 respectively represent the first quartile and the third quartile, and θ represents the threshold;
[0022] (23) Normalize the result after processing by the interquartile range method:
[0023]
[0024] where, X 1 represents the normalized data value, X 0 represents the data value to be processed, X maxx represents the maximum value in the dataset, X min represents the minimum value in the dataset.
[0025] Furthermore, step (3) includes:
[0026] (31) Extract static features from the preprocessed dataset, process the static features using an MLP, and let the static features be X static ∈R ds , where X static is a data point containing ds features, and ds is the dimension of the static features, with the expression:
[0027] h (n) = f(W (n) h (n-1) + b (n) ), n = 1, 2,..., N
[0028] where h (n) represents the output feature vector of the n-th layer of the MLP, W (n) ∈R dn×dn-1 represents the weight matrix of the first layer, b (n) ∈R dn is the bias term of the n-th layer, f() represents the activation function, and N is the total number of layers of the MLP;
[0029] After being processed by the MLP, the static feature representation vector is output:
[0030]
[0031] where d mlp represents the dimension of the output features of the MLP;
[0032] (32) Extract temporal features from the preprocessed dataset, process the temporal features using a BiLSTM, and let the temporal features be X seq ∈R T×dt , where T represents the number of time steps and dt is the feature dimension of the time step, with the expression:
[0033]
[0034] where represents the hidden state of the forward LSTM cell at time step t, with the dimension of represents the hidden state of the backward LSTM cell at time step t;
[0035] By concatenating the forward and backward hidden states, the feature vector for each time step is obtained:
[0036]
[0037] Finally, the feature matrix of the time series is:
[0038]
[0039] (33) Enhance the interaction between static features and temporal features using the multi-head attention mechanism. Calculate the attention scores through dot product and normalize them into weights using softmax. Aggregate the features according to the weight matrix, and the expression is:
[0040]
[0041] H′ = AV
[0042] Among them, A represents the similarity between each time step; Q represents the query matrix, K represents the key matrix, V represents the value matrix, and H′ is the weighted representation of each time step;
[0043] (34) Use the concatenation operation Concat to fuse static features and temporal features, and the formula is as follows:
[0044] Z fusion = Concat(Z static , Flatten(Z seq ))
[0045] Among them, Z fusion represents the fused feature vector, Z static represents the static feature vector, Z seq represents the temporal feature vector, and Flatten() represents flattening the feature matrix into a one-dimensional vector.
[0046] Furthermore, the step (5) includes:
[0047] (51) Use He initialization for the weights, and the initial values of the weights for each layer are randomly drawn from the following normal distribution:
[0048]
[0049] Among them, represents a normal distribution with a mean of 0 and a variance of , and b represents the number of input units;
[0050] (52) Use the mean squared error MSE as the loss function, and the formula is:
[0051]
[0052] Among them, L p represents the loss value of the p-th task, P is the number of tasks, C represents the number of samples, y i represents the true value, represents the predicted value, Loss represents the total loss value, and ω p represents the weight of the task;
[0053] (53) Dynamically adjust the learning rate using the improved Adam algorithm;
[0054] (54) Set the number of training epochs Epoch.
[0055] Further, the step (53) includes:
[0056] Calculate the gradient and introduce a memory unit. For each parameter θ t , calculate its gradient g t , and for each θ t there is a corresponding memory unit M t ;
[0057] Use the forgetting factor and the update factor to update the memory unit. The formula is:
[0058] M t = f t M t-1 + u t g t
[0059]
[0060] where g t is the gradient at the current time step t, M t is the memory state at the current time step, f t is the forgetting factor, u t is the update factor, a and b are hyperparameters, and μ is the reference value;
[0061] Update the momentum m t and the second moment v t , and add the memory information. The formula is:
[0062] m t = β 1 m t-1 + (1 - β 1 )g t + γM t
[0063]
[0064] where β 1 represents the decay rate controlling the momentum, β 2 represents the decay rate controlling the squared gradient, and γ is a hyperparameter;
[0065] Correct the bias of the momentum and the second moment. The formula is:
[0066]
[0067] where, is the corrected momentum estimate, is the corrected second-order matrix estimate;
[0068] Update the gradient using the corrected parameters, specifically:
[0069]
[0070] where θ t represents the updated parameter value, represents the learning rate, and ε is a small constant.
[0071] Further, step (6) includes:
[0072] (61) Select the mean absolute error MAE and the coefficient of determination R 2 as the model evaluation metrics;
[0073] (62) Compare the evaluation results predicted by the constructed model with a preset threshold.
[0074] Based on the same inventive concept, the present invention also provides an air pollutant prediction system based on multi-modal deep learning, including:
[0075] A collection module for obtaining historical air quality data and making a data set;
[0076] A preprocessing module for preprocessing the data set;
[0077] A model construction module for constructing a hybrid neural network model based on a bidirectional long short-term memory network BiLSTM, using a multi-layer perceptron MLP and a multi-head attention mechanism Multi-Head Attention;
[0078] A classification module for dividing the preprocessed data set into a training set, a validation set, and a test set according to a preset ratio;
[0079] A training module for training the hybrid neural network model using the training set;
[0080] An evaluation module for evaluating the trained hybrid neural network model;
[0081] A prediction module for repeating the training module and the evaluation module until the evaluation result is greater than a preset threshold, outputting an MLP-BiLSTM-MHAT prediction model, and using the MLP-BiLSTM-MHAT prediction model to predict the air to be measured.
[0082] Based on the same inventive concept, the present invention also provides a computer program product, including computer programs / instructions, which when executed by a processor implement the steps of the air pollutant prediction method based on multi-modal deep learning according to any one of the above.
[0083] Based on the same inventive concept, the present invention also provides a computing device, including: one or more processors, one or more memories, and one or more programs, where the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the air pollutant prediction method based on multi-modal deep learning according to any one of the above are implemented.
[0084] Based on the same inventive concept, the present invention also provides a storage medium, where the storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the steps of the air pollutant prediction method based on multi-modal deep learning according to any one of the above.
[0085] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: 1. The present invention combines a Multilayer Perceptron (MLP), a Bidirectional Long Short-Term Memory Network (BiLSTM), and a Multi-Head Attention Mechanism (MHAT) and applies them to the field of regional air pollutant prediction; through the combination of MLP, BiLSTM, and MHAT, the model can not only effectively capture the complex dependencies in time series data, but also further enhance the robustness and stability of model prediction through the input of static features; 2. The present invention improves the Adam algorithm by introducing a memory mechanism and uses it to dynamically adjust the learning rate, making the optimizer more adaptable to complex gradient patterns and non-stationary optimization processes, thereby achieving more efficient and stable global optimization; 3. The MLP-BiLSTM-MHAT prediction model constructed by the present invention has a much higher prediction accuracy compared with traditional methods, and can be extended to the field of visibility prediction to solve similar problems, having universality. Description of the Drawings
[0086] Figure 1 It is a flowchart of the method according to an embodiment of the present invention;
[0087] Figure 2 It is a comparison chart of PM2.5 indicators of the prediction results according to an embodiment of the present invention;
[0088] Figure 3 It is a comparison chart of PM10 indicators of the prediction results according to an embodiment of the present invention;
[0089] Figure 4 It is a comparison chart of SO2 indicators of the prediction results according to an embodiment of the present invention;
[0090] Figure 5 It is a comparison chart of the NO2 index for the prediction results of the embodiments of the present invention;
[0091] Figure 6 It is a comparison chart of the CO index for the prediction results of the embodiments of the present invention;
[0092] Figure 7 It is a comparison chart of the O3 index for the prediction results of the embodiments of the present invention. Detailed implementation manners
[0093] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.
[0094] As shown in the attached Figure 1 figure, the air pollutant prediction method based on multi-modal deep learning in this embodiment includes:
[0095] Step 1: Obtain historical air quality data and make a data set;
[0096] Step 2: Preprocess the data set;
[0097] Step 3: Based on the bidirectional long short-term memory network BiLSTM, use the multi-layer perceptron MLP and the multi-head attention mechanism Multi-Head Attention to construct a hybrid neural network model;
[0098] Step 4: Divide the preprocessed data set into a training set, a validation set, and a test set according to a preset ratio;
[0099] Step 5: Use the training set to train the hybrid neural network model;
[0100] Step 6: Evaluate the trained hybrid neural network model;
[0101] Step 7: Repeat steps 5 to 6 for the training and evaluation steps until the evaluation result is greater than a preset threshold, output the MLP-BiLSTM-MHAT prediction model, and use the MLP-BiLSTM-MHAT prediction model to predict the air to be measured.
[0102] Specifically, in step 1, obtaining historical air quality data specifically includes indicators such as temperature, air pressure, dew point, rainfall, wind direction, wind speed, PM2.5, PM10, SO 2 , NO 2 , CO, O 3 and other data indicators. Among them, the units of temperature and dew point are degrees Celsius (°C), the unit of air pressure is Pascal (Pa), the unit of rainfall is millimeter (mm), the unit of wind speed is m / s, PM2.5, PM10, SO2 , NO 2 , O 3 Unit: μg / m 3 , and for CO, the unit is mg / m 3 ;
[0103] Step 2, Preprocess the air quality data. The specific process includes:
[0104] 2.1 Use the mean filling and backward filling methods to handle missing data. The mean filling method is used to handle continuous missing values, and the backward filling method is used to handle non - continuous missing values. The calculation formula of the mean filling method is as follows:
[0105]
[0106] where mean j represents the mean of the j - th column of the data set, x ij represents the i - th non - missing value in the j - th column, and n j represents the number of non - missing values in the j - th column;
[0107] 2.2 Use the interquartile range (IQR) method to handle abnormal data. The specific formula is as follows:
[0108] sup = Q3+(θ * iqr)
[0109] inf = Q1-(θ * iqr)
[0110] where sup represents the upper bound, inf represents the lower bound, Q1 and Q3 represent the first quartile and the third quartile respectively, and θ is set to 1.5;
[0111] 2.3 Normalize the imputed and corrected data set. The specific formula is:
[0112]
[0113] where X 1 represents the normalized data value, X 0 represents the data value before normalization, X max represents the maximum value in the data set, and X min represents the minimum value in the data set;
[0114] Step 3, Construct a hybrid neural network: This network is used to extract static features and time - series features. Use a multi - layer perceptron (MLP), a bidirectional long short - term memory network (BiLSTM), and a multi - head attention mechanism (Multi - Head Attention) to construct a hybrid neural network. Specifically, it includes:
[0115] 3.1 Use MLP to process static features: Let the static feature be Xstatic ∈R ds , X static is a data point containing ds features, where ds is the dimension of static features, and the specific expression is:
[0116] h (n) = f(W (n) h (n-1) + b (n) ), n = 1, 2, 3,..., N
[0117] Among them, h (n) represents the output feature vector of the nth layer, W (n) ∈R dn×dn-1 represents the weight matrix of the lth layer, b (n) ∈R dn is the bias term of the nth layer, f represents the activation function, and N represents the total number of layers of the MLP;
[0118] After being processed by the MLP, the output is a static feature representation vector:
[0119]
[0120] Among them, d mlp represents the dimension of the output features of the MLP;
[0121] 3.2 Process the time series features using BiLSTM: Let the time series feature be X seq ∈R T×dt , T represents the number of time steps, and dt is the feature dimension of the time step. The specific expression is:
[0122]
[0123] Among them represents the hidden state of the forward LSTM cell at time step t, and the dimension is represents the hidden state of the backward LSTM cell at time step t;
[0124] By concatenating the forward and backward hidden states, the feature vector of each time step is obtained:
[0125]
[0126] The final feature matrix of the time series output is:
[0127]
[0128] 3.3 Use the multi-head attention mechanism to enhance features: Multiple attention heads calculate attention scores through dot product and normalize them to weights by softmax, and then aggregate features according to the weight matrix. The specific expression is:
[0129]
[0130] H' = AV
[0131] Among them, A represents the similarity between each time step; Q represents the query matrix, K represents the key matrix, V represents the value matrix, and H′ is the weighted representation of each time step;
[0132] 3.4 Feature fusion: Use the concatenation operation (Concat) to fuse static features and temporal features. The specific formula is:
[0133] Z fusion = Concat(z static , Flatten(H seq ))
[0134] Among them, Z fusion represents the fused feature vector, z static represents the static feature vector, H seq represents the temporal feature vector, and Flatten() represents flattening the feature matrix into a one-dimensional vector;
[0135] Step 4: Divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1;
[0136] Step 5: Model training. The specific process includes:
[0137] 5.1 Use He initialization for weights. The initial value of each layer of weights is randomly drawn from the following normal distribution:
[0138]
[0139] Among them, represents a normal distribution with a mean of 0 and a variance of , and b represents the number of input units;
[0140] 5.2 Use the mean squared error (MSE) as the loss function. The formula is:
[0141]
[0142] Among them, L p represents the loss value of the p-th task, P is the number of tasks, C represents the number of samples, y i represents the true value, represents the predicted value, Loss represents the total loss value, and ω p represents the weight of the task;
[0143] 5.3 Use the improved Adam algorithm to dynamically adjust the learning rate, including:
[0144] 5.3.1 Calculate the gradient and introduce the memory unit: For each parameter θ t , calculate its gradient g t , and for each θ t there is a corresponding memory unit M t ;
[0145] 5.3.2 Update the memory unit: Use the forgetting factor and the update factor to update the memory unit. The specific formula is:
[0146] M t = f t M t-1 + u t g t
[0147]
[0148] where g t is the gradient at the current time step t, and M t is the memory state at the current time step;
[0149] f t represents the forgetting factor, adjusted by the magnitude of the gradient, and is used to control the decay of memory;
[0150] u t represents the update factor and is used to determine the impact of the current gradient on memory;
[0151] a and b represent hyperparameters, which are used to control the forgetting speed and the update amplitude respectively;
[0152] μ is a reference value used to judge the magnitude of the current gradient;
[0153] 5.3.3 Update the momentum and the second moment: After introducing the memory mechanism, when updating the momentum m t and the second moment v t , memory information needs to be added. The specific formula is:
[0154] m t = β 1 m t-1 + (1 - β 1 )g t + γM t
[0155]
[0156] where β 1 represents the decay rate of the momentum, β 2 represents the decay rate of the squared gradient, and γ is a hyperparameter used to control the contribution of memory information to the momentum update;
[0157] 5.3.4 Correct the deviations of momentum and second moment, and the specific formula is as follows:
[0158]
[0159] Where, is the corrected momentum estimate, is the corrected second-order matrix estimate;
[0160] 5.3.5 Update the gradient using the corrected parameters, and the specific formula is as follows:
[0161]
[0162] Where, θ t represents the updated parameter value, represents the learning rate, and ε is a small constant;
[0163] 5.4 Set the number of training epochs;
[0164] 5.5 Set the early stopping mechanism to prevent overfitting;
[0165] Step 6, Model evaluation, including:
[0166] 6.1 Select the mean absolute error (MAE) and the coefficient of determination (R 2 ) as the model evaluation metrics;
[0167] 6.2 Compare and analyze the prediction results of the constructed model with those of other models;
[0168] Step 7, Repeat the above steps 5 and 6 until the optimal MLP-BiLSTM-MHAT prediction model is output.
[0169] In this embodiment, the experimental results are as Figures 2 to 7 shown. It can be seen from the figure that in the prediction of PM2.5 and PM10, the MLP-BiLSTM-MHAT model can accurately capture the peak and valley change trends, especially the fitting effect in the mutation region is significantly better than that of other models; in the prediction of SO 2 and NO 2 , the curve of the MLP-BiLSTM-MHAT model almost coincides with the actual value, showing a high-precision modeling ability for the fluctuations of low-concentration gases; in the prediction of CO, the MLP-BiLSTM-MHAT successfully predicts the shape and amplitude of the peak, and performs better than the obvious underestimation of GRU; in the prediction of O 3 , the MLP-BiLSTM-MHAT is smoother and more accurate in the trend fitting of continuous fluctuations, while other models show certain noise; therefore, this method has higher robustness and accuracy in the prediction of air pollutant concentrations.
[0170] Based on the same inventive concept, this embodiment also provides an air pollutant prediction system based on multimodal deep learning, including:
[0171] A collection module, configured to obtain historical air quality data and create a data set;
[0172] A preprocessing module, configured to preprocess the data set;
[0173] A model construction module, configured to construct a hybrid neural network model based on the bidirectional long short-term memory network BiLSTM, using a multi-layer perceptron MLP and a multi-head attention mechanism Multi-Head Attention;
[0174] A classification module, configured to divide the preprocessed data set into a training set, a validation set, and a test set according to a preset ratio;
[0175] A training module, configured to train the hybrid neural network model using the training set;
[0176] An evaluation module, configured to evaluate the trained hybrid neural network model;
[0177] A prediction module, configured to repeat the training module and the evaluation module until the evaluation result is greater than a preset threshold, output the obtained MLP-BiLSTM-MHAT prediction model, and use the MLP-BiLSTM-MHAT prediction model to predict the air to be measured.
[0178] Based on the same inventive concept, this embodiment also provides a computer program product, including a computer program / instructions, which when executed by a processor implement the steps of the air pollutant prediction method based on multimodal deep learning according to any one of the above.
[0179] Based on the same inventive concept, this embodiment also provides a computing device, including: one or more processors, one or more memories, and one or more programs, where the programs are stored in the memory and are configured to be executed by the processor, and when the programs are loaded into the processor, they implement the steps of the air pollutant prediction method based on multimodal deep learning according to any one of the above.
[0180] Based on the same inventive concept, this embodiment also provides a storage medium, where the storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the steps of the air pollutant prediction method based on multimodal deep learning according to any one of the above.
[0181] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit of the present invention should be covered by the present invention.
Claims
1. A method for predicting air pollutants based on multimodal deep learning, characterized in that: include: (1) Obtain historical air quality data and create a data set; (2) Preprocess the data set; (3) Based on the bidirectional long short-term memory network BiLSTM, a hybrid neural network model is constructed using the multi-layer perceptron MLP and the multi-head attention mechanism Multi-Head Attention; (4) Divide the preprocessed data set into a training set, a validation set, and a test set according to a preset ratio; (5) Use the training set to train the hybrid neural network model; (6) Evaluate the trained hybrid neural network model; (7) Repeat the training and evaluation steps from (5) to (6) until the evaluation result is greater than a preset threshold, and output the MLP-BiLSTM-MHAT prediction model, and use the MLP-BiLSTM-MHAT prediction model to predict the air to be tested.
2. The air pollutant prediction method based on multimodal deep learning according to claim 1, characterized in that: The step (2) comprises: (21) The mean filling method is used to process continuous missing values, and the backward filling method is used to process non-continuous missing values; the calculation formula of the mean filling method is as follows: Among them, mean j represents the mean of the jth column of the data set, x ij represents the i-th non-missing value in the j-th column, n j represents the number of non-missing values in the jth column; (22) For the data set after processing missing values, the interquartile range method is used to process abnormal data: sup=Q3+(θ*iqr) inf=Q1-(θ*iqr) Among them, sup represents the maximum value allowed in the data, and points exceeding this value are considered abnormal. Inf represents the minimum value allowed in the data, and points below this value are considered abnormal. Q1 and Q3 represent the first quartile and the third quartile respectively, and θ represents the threshold. (23) The results after the interquartile range method are normalized: Among them, X1 represents the normalized data value, X0 represents the data value to be processed, and X maxx Represents the maximum value in the data set, X min Represents the minimum value in the data set.
3. The air pollutant prediction method based on multimodal deep learning according to claim 1, characterized in that: The step (3) comprises: (31) Extract static features from the preprocessed data set and use MLP to process the static features. Let the static features be X static ∈R ds , X static is a data point containing ds features, where ds is the dimension of the static feature, expressed as: h (n) =f(W (n) h (n-1) +b (n) ),n=1,2,…,N Among them, h (n) represents the output feature vector of the nth layer of MLP, W (n) ∈R dn×dn-1 represents the weight matrix of the lth layer, b (n) ∈R dn is the bias term of the nth layer, f() represents the activation function, and N is the total number of layers of the MLP; After MLP processing, the static feature representation vector is output: Among them, d mlp Represents the dimension of the MLP output features; (32) Extract time series features from the preprocessed data set and use BiLSTM to process the time series features. Let the time series features be X seq ∈R T×dt , T represents the number of time steps, dt is the feature dimension of the time step, and the expression is: in, represents the hidden state of the forward LSTM unit at time step t, with dimension represents the hidden state of the reverse LSTM unit at time step t; By concatenating the forward and reverse hidden states, we get the feature vector for each time step: The feature matrix of the final output time series is: (33) A multi-head attention mechanism is used to enhance the interaction between static features and temporal features. The attention score is calculated by dot product and normalized to weight by softmax. The features are aggregated according to the weight matrix. The expression is: H′=AV Among them, A represents the similarity between each time step; Q represents the query matrix, K represents the key matrix, V represents the value matrix, and H′ is the weighted representation of each time step; (34) The concatenation operation Concat is used to fuse static features and temporal features. The formula is as follows: Z fusion =Concat(Z static ,Flatten(Z seq )) Among them, Z fusion represents the fused feature vector, Z static represents the static eigenvector, Z seq Represents the time series feature vector, and Flatten() means flattening the feature matrix into a one-dimensional vector.
4. The method for predicting air pollutants based on multimodal deep learning according to claim 1, characterized in that: The step (5) comprises: (51) Use He to initialize the weights, and the initial value of each layer's weight is randomly drawn from the following normal distribution: in, It means the mean is 0 and the variance is The normal distribution of , b represents the number of input units; (52) Using mean square error MSE as the loss function, the formula is: Among them, L p represents the loss value of the pth task, P is the number of tasks, C is the number of samples, and y i represents the true value, Represents the predicted value, Loss represents the total loss value, ω p Indicates the weight of the task; (53) Use the improved Adam algorithm to dynamically adjust the learning rate; (54) Set the training epoch.
5. The method for predicting air pollutants based on multimodal deep learning according to claim 4, characterized in that: The step (53) comprises: Calculate the gradient and introduce the memory unit, for each parameter θ t , calculate its gradient g t , and for each θ t There is a corresponding memory unit M t ; Use the forgetting factor and the updating factor to update the memory unit. The formula is: M t =f t M t-1 +u t g t Among them, g t is the gradient of the current time step t, M t is the memory state of the current time step, f t is the forgetting factor, u t is the update factor, a and b are hyperparameters, and μ is the reference value; Update momentum m t and the second moment v t , add memory information, the formula is: m t =β1m t-1 +(1-β1)g t +γM t Among them, β1 represents the decay rate of the control momentum, β2 represents the decay rate of the control square gradient, and γ is a hyperparameter; Correct the deviation of momentum and second-order moment, the formula is: in, is the revised momentum estimate, is the modified second-order matrix estimate; Use the corrected parameters to update the gradient, specifically: Among them, θ t represents the updated parameter value, represents the learning rate and ε is a small constant.
6. The method for predicting air pollutants based on multimodal deep learning according to claim 1, characterized in that: The step (6) comprises: (61) Select mean absolute error MAE and coefficient of determination R 2 As a model evaluation metric; (62) The evaluation results predicted by the constructed model are compared with the preset threshold.
7. An air pollutant prediction system based on multimodal deep learning, characterized in that: include: The acquisition module is used to obtain historical air quality data and create a data set; Preprocessing module, used to preprocess the data set; Model building module, which is used to build a hybrid neural network model based on the bidirectional long short-term memory network BiLSTM, using the multi-layer perceptron MLP and the multi-head attention mechanism Multi-Head Attention; The classification module is used to divide the preprocessed data set into training set, validation set and test set according to the preset ratio; A training module, used to train the hybrid neural network model using a training set; An evaluation module, used to evaluate the trained hybrid neural network model; The prediction module is used to repeat the training module and the evaluation module until the evaluation result is greater than a preset threshold, and the MLP-BiLSTM-MHAT prediction model is output, and the MLP-BiLSTM-MHAT prediction model is used to predict the air to be tested.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the air pollutant prediction method based on multimodal deep learning according to any one of claims 1 to 6 are implemented.
9. A computing device, characterized in that include: One or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the air pollutant prediction method based on multimodal deep learning according to any one of claims 1 to 6 are implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, which includes program instructions, and when the program instructions are executed by a processor, the processor executes the steps of the air pollutant prediction method based on multimodal deep learning according to any one of claims 1 to 6.
Citation Information
Cited By
Regression prediction and abnormal state detection classification system and method for gas dissolved in oil based on multi-expert fusion learning and medium
CN121479727A