LSTM urban traffic network CO and NOx emission prediction method and system based on Attention mechanism
By introducing an Attention mechanism into the LSTM model, using urban traffic data to predict CO and NOx emissions, the problems of slow response and high cost of traditional pollutant monitoring methods are solved, and more efficient and accurate pollutant emission prediction is achieved.
Patent Information
- Application Number
- CN202510149006.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional pollutant monitoring methods are slow to respond and costly, making it difficult to capture the complex and changeable pollutant emissions in urban transportation networks in real time and comprehensively.
Using the LSTM model based on the Attention mechanism, a Attention-LSTM model including the Attention mechanism layer is constructed by collecting road monitoring data and vehicle feature data, conducting feature analysis and screening, and training is performed to predict CO and NOx emissions.
It improves the accuracy and stability of pollutant emission forecasts, can better capture the complex relationship between traffic flow and emissions, and meet the needs of overall urban traffic management and pollution control.
Smart Images

Figure CN120069207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air pollutant prediction, and particularly to a method and system for predicting CO and NOx emissions in urban traffic road networks based on the Attention mechanism and LSTM. Background Art
[0002] Due to the slow response speed and high cost of traditional pollutant monitoring methods, it is difficult to comprehensively and real-time capture the complex and variable pollutant emissions in urban traffic networks. Therefore, finding efficient and accurate pollutant emission prediction methods has become an urgent problem to be solved.
[0003] With the continuous progress of artificial intelligence technology, deep learning, especially prediction models based on Long Short-Term Memory (LSTM) networks, has shown significant advantages in processing time series data. LSTM can effectively capture long-term dependencies in sequence data, thereby improving prediction accuracy. However, simply relying on the LSTM network may still face the problem of insufficient information extraction when dealing with large-scale complex traffic data. To this end, the attention mechanism can be introduced into the LSTM model. By assigning different weights to different traffic input features, key features can be better captured, further enhancing the prediction ability and stability of the road network emission model. Summary of the Invention
[0004] In view of some or all of the problems in the prior art, the present invention provides a method for predicting CO and NOx emissions in urban traffic road networks based on the Attention mechanism and LSTM. The method includes the following steps:
[0005] Collect road monitoring data and vehicle feature data as input data;
[0006] Use principal component analysis and correlation coefficient method to perform feature analysis on the input data, and screen out feature variables related to CO and NOx emissions as input feature variables;
[0007] Perform data cleaning and normalization processing on the dataset of the input feature variables to obtain a normalized dataset of feature variables;
[0008] Divide the normalized dataset of feature variables into a training set and a test set;
[0009] Construct an Attention-LSTM model including an Attention mechanism layer, where the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer;
[0010] Construct the loss function of the Attention-LSTM model, and configure the hyperparameters of the Attention-LSTM model, where the hyperparameters include the learning rate, batch size, number of LSTM layers, and number of neurons; and
[0011] Input the normalized feature variable dataset into the Attention-LSTM model with set hyperparameters for training to obtain an emission prediction model for CO and NOx in the traffic road network, and use evaluation metrics to evaluate the emission prediction models of CO and NOx.
[0012] Furthermore, use principal component analysis and correlation coefficient method to perform feature analysis on the input data, and screen out the feature variables related to CO and NOx emissions as input feature variables, including:
[0013] Use principal component analysis to map the input data to a low-dimensional space and screen the feature variable data; and
[0014] Calculate the Pearson correlation coefficient of the feature variable data, and screen out the feature variable data related to CO and NOx emissions as input feature variable data;
[0015] The Pearson correlation coefficient rho(m,n) is
[0016]
[0017] where X m and X n are the m-th and n-th feature variable data respectively, are the averages of the m-th and n-th feature variable data respectively, and N is the number of feature variable data;
[0018] The input feature variables include feedback speed, stage distance, DPF operating speed, filtration speed, pre-filter temperature, filter face velocity, and battery temperature.
[0019] Furthermore, the normalization formula is
[0020]
[0021] where X i is the input feature variable data, X min and X max are the maximum and minimum values of the input feature variable data respectively, and yi is the normalized feature variable data;
[0022] After normalization, a normalized feature variable dataset is obtained, and the normalized feature variable dataset is scaled to the range of [0,1].
[0023] Further, dividing the normalized feature variable data set into a training set and a test set includes:
[0024] Dividing the normalized feature variable data set into a training set and a test set according to a ratio of 8:2.
[0025] Further, constructing an Attention-LSTM model including an Attention mechanism layer, where the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer, includes:
[0026] The input data set enters the LSTM layer in the form of a time series. The input of the LSTM layer is X = [x 1 , x 2 ,..., x T , the sequence length is T, and the hidden state h t at time step t is updated according to the formula:
[0027] h t = f(W x x t + W h h t-1 + b)
[0028] where f is the activation function, W x is the input weight matrix, W h is the hidden state weight matrix, b is the bias vector, and x t is the input at time step t;
[0029] Input the output residual r t at time step t of the LSTM layer into the attention mechanism layer;
[0030] In the Squeeze layer of the attention mechanism layer, compress the output residual r t at time step t of the LSTM layer into a global description feature vector v;
[0031] Input the global description feature vector v into the extract layer;
[0032] Input the LSTM weights and residuals into the fully connected layer, and the fully connected layer outputs a vector [H t , V t ;
[0033] Input the output vector [H t , V t into the ReLU layer to obtain the final Attention weight V t '; and
[0034] The final Attention weight V t ' is used to weight the output of the LSTM to generate the final prediction y of the Attention-LSTM model final ;
[0035] Among them, H t is the reference vector, and V t is the attention score;
[0036] The calculation formulas for the reference vector and the attention score are
[0037] H t = W e h t + b e
[0038] V t = softmax(H t )
[0039] Among them, W e and b e are the weights and biases of the fully connected layer, and softmax is used to normalize the scores into a probability distribution;
[0040] The final Attention weight V t ' is
[0041] V t ' = ReLU(V t )
[0042] The final prediction y of the Attention-LSTM model final is
[0043]
[0044] Among them, h t is the hidden state of the LSTM, and T is the sequence length.
[0045] Further, construct the loss function of the Attention-LSTM model and configure the hyperparameters of the Attention-LSTM model. The hyperparameters include the learning rate, batch size, number of LSTM layers, and number of neurons, including:
[0046] The loss function Loss of the Attention-LSTM model is
[0047]
[0048] Among them, Train_Y is the training set, Predict_Y is the prediction set, and n is the number of samples in the training set;
[0049] The learning rate is 0.001, the batch size is 30, the number of LSTM layers is single layer, and the number of neurons is 128.
[0050] Furthermore, the evaluation metrics include root mean square error, mean absolute error, coefficient of determination, and mean absolute percentage error.
[0051] The present invention also provides a system for the above-mentioned LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism, and this system includes the following modules:
[0052] A data collection module, configured to collect road monitoring data and vehicle characteristic data as input data;
[0053] A feature variable screening module, configured to perform feature analysis on the input data using principal component analysis and correlation coefficient method, and screen out the feature variables related to CO and NOx emissions as input feature variables;
[0054] A data preprocessing module, configured to perform data cleaning and normalization processing on the dataset of the input feature variables to obtain a normalized dataset of feature variables;
[0055] A dataset division module, configured to divide the normalized dataset of feature variables into a training set and a test set;
[0056] An Attention-LSTM model construction module, configured to construct an Attention-LSTM model including an Attention mechanism layer, and the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer;
[0057] A hyperparameter configuration module, configured to construct a loss function of the Attention-LSTM model and configure the hyperparameters of the Attention-LSTM model, and the hyperparameters include learning rate, batch size, number of LSTM layers, and number of neurons; and
[0058] An emission prediction and evaluation module, configured to input the normalized dataset of feature variables into the Attention-LSTM model with set hyperparameters for training to obtain an emission prediction model for traffic road network CO and NOx, and evaluate the CO and NOx emission prediction models using evaluation metrics.
[0059] The present invention also provides a computer system, including:
[0060] A processor, configured to execute machine-readable instructions;
[0061] A graphics card with an artificial intelligence chip, which is configured to train the LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism; and
[0062] A memory, which is configured to store machine-readable instructions that, when executed by a processor and / or a graphics card, perform the steps of the LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism.
[0063] The present invention also provides a computer-readable storage medium, on which machine-readable instructions are stored that, when executed by a processor, perform the steps of the LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism.
[0064] The technical solution provided by the present invention has the following advantages:
[0065] 1. Traditional prediction models rely on experts' experience and subjective judgment and are difficult to process complex traffic data. The LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism proposed by the present invention can automatically extract and learn key features in traffic data through deep learning, thereby achieving more intelligent and accurate emission prediction.
[0066] 2. The LSTM model is good at processing time series data, but when traffic data involves complex spatial relationships, a simple LSTM model is insufficient to cope with it. The LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism proposed by the present invention can effectively focus on important time points and spatial positions, capture the complex relationship between traffic flow and emissions, and improve the accuracy and stability of prediction.
[0067] 3. The LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism proposed by the present invention realizes a global prediction mechanism by modeling and analyzing the global data of the entire traffic road network, and can better meet the needs of urban overall traffic management and pollution control.
[0068] 4. The LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism proposed by the present invention can continuously absorb and learn new traffic data, dynamically update the prediction model, ensure a high degree of agreement with the actual situation at all times, and provide real-time and effective management suggestions.
[0069] 5. The LSTM urban traffic network CO and NOx emission prediction method based on the Attention mechanism proposed by the present invention allows for flexible adjustment of the objective function to meet different management requirements and environmental policies, providing greater flexibility and choice space for the formulation of traffic control strategies.
[0070] 6. The LSTM urban traffic network CO and NOx emission prediction method based on the Attention mechanism proposed by the present invention is not only applicable to CO and NOx emission prediction, but can also be extended to the prediction and management of other types of traffic pollutants, providing an advanced and efficient solution for the accurate prediction and control strategies of urban traffic pollutant emissions, and having broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] To further clarify the above and other advantages and features of the embodiments of the present invention, a more specific description of the embodiments of the present invention will be presented with reference to the accompanying drawings. It can be understood that these drawings only depict typical embodiments of the present invention and will not be considered as limiting its scope. In the drawings, for clarity, the same or corresponding components will be denoted by the same or similar reference numerals.
[0072] Figure 1 FIG. shows a schematic flow chart of the LSTM urban traffic network CO and NOx emission prediction method based on the Attention mechanism according to an embodiment of the present invention;
[0073] Figure 2 FIG. shows a schematic diagram of the basic framework of the LSTM urban traffic network CO and NOx emission prediction model based on the Attention mechanism according to an embodiment of the present invention;
[0074] Figure 3 FIG. shows a schematic diagram of the Pearson correlation coefficient matrix of the characteristic variables according to an embodiment of the present invention;
[0075] Figure 4 FIG. shows a schematic diagram of the Attention-LSTM model structure according to an embodiment of the present invention;
[0076] Figure 5 FIG. shows a schematic diagram of the comparison of the CO emission prediction results of the LSTM urban traffic network based on the Attention mechanism according to an embodiment of the present invention;
[0077] Figure 6 FIG. shows a schematic diagram of the comparison of the NOx emission prediction results of the LSTM urban traffic network based on the Attention mechanism according to an embodiment of the present invention; and
[0078] Figure 7Schematic diagram of an LSTM urban traffic road network CO and NOx emission prediction system based on the Attention mechanism according to an embodiment of the present invention. Detailed implementation manners
[0079] In the following description, the present invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments can be implemented without one or more specific details or in combination with other alternative and / or additional methods or components. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring the inventive points of the present invention. Similarly, for purposes of explanation, specific numbers and configurations are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, the present invention is not limited to these specific details.
[0080] In this specification, the reference to "an embodiment" or "the embodiment" means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. The phrase "in an embodiment" appearing throughout this specification does not necessarily all refer to the same embodiment.
[0081] It should be noted that the embodiments of the present invention describe the method steps in a specific order. However, this is only for the purpose of explaining the specific embodiment and does not limit the order of the steps. On the contrary, in different embodiments of the present invention, the order of the steps can be adjusted according to the actual requirements.
[0082] In the present invention, each module of the system according to the present invention can be implemented using software, hardware, firmware, or a combination thereof. When a module is implemented using software, the functions of the module can be realized through a computer program process. For example, the module can be implemented by a code segment (such as a code segment in languages like C, C++) stored in a storage device (such as a hard disk, memory, etc.). When the code segment is executed by a processor, the corresponding functions of the module can be realized. When a module is implemented using hardware, the functions of the module can be realized by setting up corresponding hardware structures. For example, the functions of the module can be realized by hardware programming of a programmable device such as a Field Programmable Gate Array (FPGA), or by designing an Application Specific Integrated Circuit (ASIC) including multiple electronic devices such as transistors, resistors, and capacitors. When a module is implemented using firmware, the functions of the module can be written in a read-only memory such as an EPROM or EEPROM of the device in the form of program code, and when the program code is executed by a processor, the corresponding functions of the module can be realized. Additionally, certain functions of the module may need to be realized by separate hardware or in cooperation with the hardware. For example, the detection function is realized through corresponding sensors (such as proximity sensors, acceleration sensors, gyroscopes, etc.), the signal emission function is realized through corresponding communication devices (such as Bluetooth devices, infrared communication devices, baseband communication devices, Wi-Fi communication devices, etc.), the output function is realized through corresponding output devices (such as displays, speakers, etc.), and so on.
[0083] The present invention proposes an LSTM-based prediction method for CO and NOx emissions in urban traffic road networks based on the Attention mechanism for the prediction of CO and NOx emissions in urban traffic road networks. Through the Attention mechanism, the model can dynamically adjust the focus of attention, more accurately extract key features related to pollutant emissions, and combine the time series modeling ability of LSTM to achieve efficient prediction of traffic pollutant emissions. This method not only improves the accuracy of prediction but also provides a new technical means for the real-time monitoring and prediction of urban traffic pollutant emissions, helps to formulate more scientific and effective traffic control strategies, thereby improving urban air quality and enhancing life satisfaction.
[0084] Figure 1 The flowchart shows the LSTM-based prediction method for CO and NOx emissions in urban traffic road networks based on the Attention mechanism according to an embodiment of the present invention. Figure 2 The schematic diagram shows the basic framework of the LSTM-based prediction model for CO and NOx emissions in urban traffic road networks based on the Attention mechanism according to an embodiment of the present invention. Figure 2Among them, the purpose of dilution is to make the data more suitable for the training of the learning model to improve the generalization ability of the model. The relationship between the diluted CO and NOx and those before dilution can be expressed by the dilution ratio, which is the proportional relationship between the concentration of the original pollutant and the concentration after dilution. The concentration dilution ratio of CO is 10 times, and the dilution ratio of NOx is 1. In a multi-input single-output regression model, the diluted data can be used to train the model to reflect the pollutant diffusion characteristics in the real environment and their impact on urban traffic emissions.
[0085] Next, in combination with Figure 1 and Figure 2 , the Attention mechanism-based LSTM urban traffic road network CO and NOx emission prediction method proposed by the present invention will be described. In an embodiment of the present invention, the Attention mechanism-based LSTM urban traffic road network CO and NOx emission prediction method can be executed by a computer. As Figure 1 shown, the Attention mechanism-based LSTM urban traffic road network CO and NOx emission prediction method includes the following steps:
[0086] First, collect road monitoring data and vehicle characteristic data as input data. Urban traffic pollutant emission indicators include carbon monoxide, carbon dioxide, nitrogen oxides, hydrocarbons, etc. The present invention selects carbon monoxide (CO) and nitrogen oxides (NOx) as the prediction targets of the Attention-LSTM emission model. Vehicle characteristics mainly include vehicle speed, stage driving distance, Diesel Particulate Filter (DPF) operation rate, filter flow rate (L / min), filtration temperature, filter windward surface speed, battery temperature, etc. at each time series. In an embodiment of the present invention, the road monitoring data is traffic data sampled every hour from January 1, 2016 to April 9, 2016.
[0087] Next, use the principal component analysis and correlation coefficient method to perform feature analysis on the input data, and screen out the feature variables related to the CO and NOx emissions as input feature variables.
[0088] Perform principal component analysis (PCA) on the data set and map its features to a low-dimensional space. First, select the feature variable data with complete records and value in the data set, and exclude some feature variable data with a large amount of redundant residuals and wrong judgments. Secondly, calculate the correlation coefficients of each feature variable and draw a Pearson correlation analysis coefficient matrix.
[0089] The Pearson correlation coefficient rho(m,n) is
[0090]
[0091] Among them, X m and X n are the m-th and n-th characteristic variable data respectively, are the average values of the m-th and n-th characteristic variable data respectively, and N is the number of characteristic variable data.
[0092] Figure 3 shows a schematic diagram of the Pearson correlation coefficient matrix of the characteristic variables in an embodiment of the present invention. According to the scores between the characteristic variables, there is a weak or even extremely weak correlation between the feedback speed, stage distance, battery temperature and the DPF operating speed, filtration speed, pre-filter temperature, and filter windward surface speed. However, there is a relatively strong correlation among the DPF operating speed, filtration speed, pre-filter temperature, and filter windward surface speed, and there is also a relatively strong correlation among the feedback speed, stage distance, and battery temperature. Therefore, a total of seven variables, namely the feedback speed, stage distance, DPF operating speed, filtration speed, pre-filter temperature, filter windward surface speed, and battery temperature, are selected as the input characteristic variables.
[0093] In principal component analysis, the principal components are calculated from the variance-covariance matrix between the characteristics and are used to capture the direction with the largest variance in the data. Features with low correlation can provide independent directions of variation for the principal components in PCA, enabling the principal components to better reflect the diversity of the input data. Therefore, the influence of these features on the principal components is to make them contain more dimensions of unique information, thereby enriching the predictive ability of the model.
[0094] Next, clean and normalize the dataset of the input feature variables to obtain a normalized feature variable dataset. Data cleaning and normalization are the preprocessing of data. Select appropriate preprocessing methods according to the accuracy of different datasets. Since there are situations such as missing values and outliers in the data, in order to ensure data quality and detect and correct errors, missing values, redundancy, etc. in the data, data cleaning is required. In the present invention, first, it is necessary to count the data information under each feature in the dataset. Secondly, judge whether the data is within a reasonable range according to variables such as the total amount, average value, maximum value, and minimum value of the data under each feature. For the null values in the data, convert them to 0. Delete features with a large number of missing values, excessive data dispersion, etc., and retain the feature data with high data quality and predictive value. Taking the Euro 6 emission standard vehicle dataset as an example, round all features to two decimal places to avoid the model being partial to the easy part and losing its due accuracy due to too many decimal places in all data under the feature. At the same time, since the feedback speed recorded in this dataset is recorded with a negative sign for the reverse direction, and the negative values account for less than 5% in the overall data. Therefore, all negative signs are deleted here, and only the absolute value of the speed is retained to avoid the existence of a small number of negative numbers from reducing the regression performance and generalization ability of the model. The standard normalization formula is as follows:
[0095]
[0096] Y min and Y max represent the minimum and maximum values of the interval to be normalized respectively. X min and X max represent the minimum and maximum values in the input feature variable data respectively. Normalize each data X i of the input feature to between Y min and Y max to unify the sample data of each parameter.
[0097] This method normalizes all selected input feature variable data to [0, 1]. According to the requirements, take Y max as 1, take Y min as 0, and the actually used standardized operation formula is,
[0098]
[0099] where, X i is the input feature variable data, X min and X max are the maximum and minimum values of the input feature variable data respectively, and yi is the normalized feature variable data.
[0100] After normalization, a normalized feature variable data set is obtained, and the normalized feature variable data set is scaled to the range of [0, 1].
[0101] Next, the normalized feature variable data set is divided into a training set and a test set. In an embodiment of the present invention, the normalized feature variable data set is divided into a training set and a test set according to a ratio of 8:2. Taking the Euro 6 emission standard vehicle data set as an example, there are a total of 2,381 data, which are used to predict CO and NOx. The first 1,881 data are taken as the training set, the last 500 data are taken as the test set, and at the same time, the last 200 data in the training set are taken as the validation set.
[0102] Next, an Attention-LSTM model including an Attention mechanism layer is constructed. The Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer. For the regression target of multiple inputs and single output, since each time step of the input sequence is equally treated by each neuron, therefore, the present invention introduces an attention mechanism to give higher weights to more valuable variables in the input sequence. In the Attention-LSTM model, the attention mechanism layer usually receives the output of the LSTM layer and dynamically adjusts the weights according to these outputs to better capture the key features in the time series data. In the present invention, the constructed attention mechanism layer mainly includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer. Under this model, the parameters to be learned mainly include: LSTM weights, Attention weights, and the residuals of the LSTM output.
[0103] Figure 4 FIG. shows a schematic structural diagram of the Attention-LSTM model according to an embodiment of the present invention. As Figure 4 shown, the structure of the Attention-LSTM model is that after substituting the data into the LSTM model, the difference between the fitted value and the true value, that is, the residual, will be obtained; the output residual r of the LSTM layer at time step t t is input into the attention mechanism layer; in the Squeeze layer of the attention mechanism layer, the output residual r of the LSTM layer at time step t t is compressed into a global description feature vector v; this global description feature vector v is used as a new output and input into the Extract layer to extract effective information and perform calculations in the fully connected layer; the LSTM weights and residuals are input into the fully connected layer, and the fully connected layer outputs a vector [H t , V t ; the output vector [H t , V t is input into the ReLU layer to obtain the final Attention weight Vt '; and the final Attention weight V t ' is used to weight the output of the LSTM to generate the final prediction yfinal of the Attention-LSTM model.
[0104] The input dataset enters the LSTM layer in the form of a time series. The input to the LSTM layer is X = [x 1 , x 2 ,..., x T , the sequence length is T, and the hidden state h t of the LSTM layer at time step t is updated according to the formula
[0105] h t = f(W x x t + W h h t-1 + b)
[0106] where f is the activation function, W x is the input weight matrix, W h is the hidden state weight matrix, b is the bias vector, and x t is the input at time step t.
[0107] The output of the LSTM layer passes through the attention mechanism layer to dynamically adjust the weights of the output at each time step, enabling the model to better focus on the time steps that are more important for the regression target.
[0108] The residual r t represents the difference between the true value ytrue,t and the fitted value ypred,t of the LSTM layer.
[0109] r t = y true,t - y pred,t .
[0110] The compression layer can aggregate the residuals at different time steps into a global vector to extract the overall features of the sequence. The compressed feature vector is input into the Extract layer to extract more meaningful information. The Extract layer can enhance the expressive power of the features through pooling operations and output a feature vector.
[0111] where H t is the reference vector, i.e., recording the state of the model training at the t-th moment, and V t is the attention score, i.e., the importance evaluation score of the feature data at this moment;
[0112] The calculation formulas for the reference vector and the attention score are
[0113] H t= W e h t + b e
[0114] V t = softmax(H t )
[0115] where W e and b e are the weights and biases of the fully connected layer, and softmax is used to normalize the scores into a probability distribution;
[0116] The final Attention weight V t ' is
[0117] V t ' = ReLU(V t )
[0118] The final prediction yfinal of the Attention-LSTM model is
[0119]
[0120] where h t is the hidden state of the LSTM, and T is the sequence length.
[0121] Next, construct the loss function of the Attention-LSTM model and configure the hyperparameters of the Attention-LSTM model. The hyperparameters include the learning rate, batch size, number of LSTM layers, and number of neurons.
[0122] In the present invention, the Attention-LSTM model inputs seven feature variables, aims to reduce the loss function, and continuously iteratively adjusts the Attention weight according to the time step, and finally outputs a traffic road network emission prediction regression model.
[0123] The loss function Loss of the Attention-LSTM model is
[0124]
[0125] where Train_Y is the training set, Predict_Y is the prediction set, and n is the number of samples in the training set, which is used to calculate the loss function.
[0126] The LSTM is a single-layer LSTM, that is, the number of LSTM layers is one. The number of neurons is 128. The feature dimension of the data input is 7, and the dimension of the data output is 1. In terms of hyperparameters, the number of hidden neurons in the LSTM layer is default set to 128, the state activation function selects tanh, the gate activation function selects the Sigmoid function, and the lag length is 24. Set the size of the fully connected layer to the number of output responses, which is 1. The Adam gradient descent algorithm is selected, the batch size is set to 30, the number of iterations is 1200 times, the learning rate is set to 0.001, and the number of training rounds is set to 15. Among them, the learning rate decay period is set to 800, and the learning rate decay factor is 0.15, that is, after the 800th training iteration, the learning rate drops to 0.001 multiplied by 0.15. At the same time, add the L2 regularization parameter 0.001 to prevent the model from overfitting.
[0127] Update the Attention weights in each training iteration, continuously use the validation set to verify the training effect, and reduce the loss function. Import the initialization parameters of the LSTM layer, the Attention layer, and the fully connected layer, including (i.e., the reference vector, used to record the initial state and the initial score), the LSTM initial residual, the LSTM weights, the recurrent weights, and the Attention weights, etc. Sending all the feature data into the Attention-LSTM emission model for training once is called an epoch. This traffic emission model is trained for 15 epochs, and the batch size is updated once per round, that is, how many batches of data to train per round, denoted as batch.
[0128] In each epoch, the parameters of the model are updated by gradient descent (such as the Adam optimizer). The update formula is,
[0129]
[0130] where η is the learning rate, and ω are the parameters of the model, i.e., the weights. During the training process, the model updates the weights to minimize the loss function, L total is the total loss function value, representing the error between the predicted value and the true value output by the model. After obtaining the output h t in the LSTM layer, the Attention layer calculates the attention weights V t according to the LSTM output and the reference vector H t :
[0131]
[0132] In each epoch, the data is divided and trained in batches, and the batch size is dynamically adjusted during training to enhance the stability of the model. In each batch, the model calculates the loss function using forward propagation and then updates the model parameters through backpropagation. After each epoch, the model is evaluated using the validation set to ensure that the model has not overfitted, and the model parameters are fine-tuned according to the performance of the validation set. As the number of training epochs increases, the loss value continuously decreases, and the prediction curve of the model gradually approaches the true value, indicating that the model is approaching the optimal fitting state.
[0133] Finally, the normalized feature variable dataset is input into the Attention-LSTM model with set hyperparameters for training to obtain an emission prediction model for CO and NOx in the traffic road network, and the CO and NOx emission prediction models are evaluated using evaluation metrics. By comparing the predicted values with the actual values and using the selected evaluation metrics, the effectiveness and accuracy of the model are verified to ensure its good generalization ability in practical applications. The evaluation metrics include root mean square error, mean absolute error, coefficient of determination, and mean absolute percentage error. The root mean square error RMSE is used to measure the deviation between the model's predicted values and the true values, the mean absolute error MAE represents the average absolute deviation between the predicted values and the true values, the coefficient of determination R 2 is used to measure the degree of fit of the model to the data, and the mean absolute percentage error MAPE reflects the relative error between the model's predicted values and the true values, expressed as a percentage, and the smaller the better.
[0134] Table 1 shows the evaluation results of the CO and NOx emission prediction models based on the Attention-LSTM model. Relying solely on the squared term or MAPE error for evaluation may lead to a one-sided judgment of the model's performance. The overall prediction ability of the model should also be considered comprehensively. Taking the CO emission model as an example, although its squared term is -14.9609 and the MAPE reaches about 45%, this does not mean that the regression performance of the model is poor. In fact, the result graph of the predicted values and the true values of this model shows a certain emission prediction ability, and its emission prediction curve gradually approaches the true value, as Figure 5 shown.
[0135] Table 1 Evaluation results of the CO and NOx emission prediction models based on the Attention-LSTM model
[0136] Pollutant type RMSE MAE <![CDATA[R 2 > MAPE CO emission model 4.695 2.6764 -14.9609 45.1719% NOx emission model 0.1782 0.0413 0.4991 76.8083%
[0137] Figure 5 and Figure 6Among them, the abscissa represents the prediction samples, which are the serial numbers of the samples used for model prediction and represent each prediction point in the test set. The ordinate represents the prediction results of CO and NOx emission concentrations, with the unit of ppm (parts per million concentration).
[0138] Figure 5 It shows the comparison between the prediction results of the CO emission concentration by the LSTM urban traffic road network CO emission prediction model based on the Attention mechanism and the actual observed values. The red curve represents the real CO emission concentration, and the blue curve represents the predicted value of the model. It can be seen from the figure that the model can follow the actual CO emission trend well most of the time, indicating that the model has a good ability to capture the temporal characteristics of CO emissions.
[0139] Figure 6 It shows the comparison between the prediction results of the NOx emission concentration by the LSTM urban traffic road network NOx emission prediction model based on the Attention mechanism and the actual observed values. Red represents the real NOx emission concentration, and blue represents the predicted value of the model. Generally speaking, the model can also track the trend of NOx emissions relatively accurately, especially the prediction at multiple peak positions is relatively close to the actual values.
[0140] The method for predicting CO and NOx emissions in the urban traffic road network based on the Attention mechanism proposed by the present invention can automatically extract and learn the key features in traffic data through deep learning, so as to achieve more intelligent and accurate emission prediction; it can effectively focus on important time points and spatial positions, capture the complex relationship between traffic flow and emissions, and improve the accuracy and stability of prediction; by modeling and analyzing the global data of the entire traffic road network, a global prediction mechanism is realized, which can better meet the needs of urban overall traffic management and pollution control; it is not only applicable to CO and NOx emission prediction, but also can be extended to the prediction and management of other types of traffic pollutants, providing an advanced and efficient solution for the accurate prediction and control strategy of urban traffic pollutant emissions, and having a broad application prospect.
[0141] In an embodiment of the present invention, the present invention also provides a system for predicting CO and NOx emissions in the urban traffic road network based on the Attention mechanism. Figure 7 It shows a schematic diagram of the system for predicting CO and NOx emissions in the urban traffic road network based on the Attention mechanism according to an embodiment of the present invention. As Figure 7 shown, the system includes the following modules:
[0142] A data collection module, configured to collect road monitoring data and vehicle characteristic data as input data;
[0143] A feature variable screening module, configured to perform feature analysis on the input data using principal component analysis and the correlation coefficient method, and screen out the feature variables related to the CO and NOx emissions as input feature variables;
[0144] A data preprocessing module, configured to perform data cleaning and normalization processing on the dataset of the input feature variables to obtain a normalized dataset of feature variables;
[0145] A dataset partitioning module, configured to partition the normalized dataset of feature variables into a training set and a test set;
[0146] An Attention-LSTM model construction module, configured to construct an Attention-LSTM model including an Attention mechanism layer, where the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer;
[0147] A hyperparameter configuration module, configured to construct a loss function of the Attention-LSTM model and configure the hyperparameters of the Attention-LSTM model, where the hyperparameters include a learning rate, a batch size, the number of LSTM layers, and the number of neurons; and
[0148] An emission prediction and evaluation module, configured to input the normalized dataset of feature variables into the Attention-LSTM model with the hyperparameters set for training to obtain an emission prediction model for CO and NOx in the traffic road network, and evaluate the CO and NOx emission prediction models using evaluation metrics.
[0149] In one embodiment of the present invention, the present invention further provides a computer system, which includes a processor, a graphics card with an artificial intelligence chip, and a memory. The memory is configured to store machine-readable instructions, the graphics card is configured to train the LSTM urban traffic road network CO and NOx emission prediction method based on the Attention mechanism, and the processor is configured to execute the machine-readable instructions. When the processor and / or the graphics card execute the machine-readable instructions, the following processing steps are implemented: collecting road monitoring data and vehicle feature data as input data; using the principal component analysis and the correlation coefficient method to perform feature analysis on the input data, and screening out the feature variables related to the CO and NOx emissions as input feature variables; performing data cleaning and normalization processing on the dataset of the input feature variables to obtain a normalized dataset of feature variables; dividing the normalized dataset of feature variables into a training set and a test set; constructing an Attention-LSTM model including an Attention mechanism layer, where the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer; constructing a loss function of the Attention-LSTM model, and configuring the hyperparameters of the Attention-LSTM model, where the hyperparameters include a learning rate, a batch size, the number of LSTM layers, and the number of neurons; inputting the normalized dataset of feature variables into the Attention-LSTM model with the set hyperparameters for training to obtain an emission prediction model for traffic road network CO and NOx, and using evaluation metrics to evaluate the CO and NOx emission prediction model.
[0150] The graphics card preferably can be a graphics card with a GPU computing power higher than model 5.0. Since the amount of data to be trained is large, providing the graphics card configuration can significantly improve the training speed.
[0151] The memory includes various media that can store machine-readable instructions, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc.
[0152] It can be understood that in addition to the memory and the processor described above, the above computer system further includes other software and hardware components not listed in this specification. Specifically, it can be determined according to the model of the specific data processing device in different application scenarios, and this specification will not list and elaborate one by one.
[0153] In one embodiment of the present invention, the present invention further provides a computer-readable storage medium, on which machine-readable instructions are stored, and when the machine-readable instructions are executed by a processor, the following processing steps are implemented: collecting road monitoring data and vehicle feature data as input data; using principal component analysis and correlation coefficient method to perform feature analysis on the input data, and screening out feature variables related to CO and NOx emissions as input feature variables; performing data cleaning and normalization processing on the dataset of the input feature variables to obtain a normalized dataset of feature variables; dividing the normalized dataset of feature variables into a training set and a test set; constructing an Attention-LSTM model including an Attention mechanism layer, and the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer; constructing a loss function of the Attention-LSTM model, configuring hyperparameters of the Attention-LSTM model, and the hyperparameters include a learning rate, a batch size, the number of LSTM layers, and the number of neurons; inputting the normalized dataset of feature variables into the Attention-LSTM model with set hyperparameters for training to obtain an emission prediction model for CO and NOx in a traffic road network, and using evaluation metrics to evaluate the emission prediction model for CO and NOx.
[0154] Although the embodiments of the present invention have been described above, it should be understood that they are presented only as examples and not as limitations. It will be obvious to those skilled in the relevant art that various combinations, variations, and changes can be made to them without departing from the spirit and scope of the present invention. Therefore, the width and scope of the present invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined according to the technical solution of the present invention and its equivalent replacements.
Claims
1. A LSTM urban traffic network CO and NOx emission prediction method based on Attention mechanism, characterized in that: The steps include: Collect road monitoring data and vehicle characteristic data as input data; Performing feature analysis on the input data using principal component analysis and correlation coefficient method, and screening out feature variables related to CO and NOx emissions as input feature variables; Performing data cleaning and normalization processing on the data set of the input feature variables to obtain a normalized feature variable data set; Dividing the normalized feature variable data set into a training set and a test set; Construct an Attention-LSTM model including an Attention mechanism layer, wherein the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer; Construct the loss function of the Attention-LSTM model and configure the hyperparameters of the Attention-LSTM model, including the learning rate, batch size, number of LSTM layers, and number of neurons; as well as The normalized feature variable data set is input into the Attention-LSTM model with set hyperparameters for training to obtain an emission prediction model for CO and NOx in a traffic network, and the CO and NOx emission prediction model is evaluated using an evaluation indicator.
2. According to claim 1, the LSTM urban traffic network CO and NOx emission prediction method based on the Attention mechanism is characterized in that: The input data is analyzed using principal component analysis and correlation coefficient method, and characteristic variables related to CO and NOx emissions are screened out as input characteristic variables including: Mapping the input data to a low-dimensional space using principal component analysis to filter feature variable data; and Calculating the Pearson correlation coefficient of the characteristic variable data, and selecting the characteristic variable data related to CO and NOx emissions as input characteristic variable data; The Pearson correlation coefficient rho(m,n) is, Among them, X m , X n are the mth and nth characteristic variable data respectively, are the average values of the mth and nth characteristic variable data, respectively, and N is the number of characteristic variable data; The input characteristic variables include feedback speed, stage distance, DPF operation speed, filtration speed, pre-filtration temperature, filter windward surface speed and battery temperature.
3. According to claim 1, the LSTM urban traffic network CO and NOx emission prediction method based on the Attention mechanism is characterized in that: The normalized formula is: Among them, X i is the input feature variable data, X min and X max are the maximum and minimum values of the input feature variable data, y i is the normalized characteristic variable data; After normalization, a normalized feature variable data set is obtained, and the normalized feature variable data set is scaled to the range of [0, 1].
4. The LSTM urban traffic network CO and NOx emission prediction method based on Attention mechanism according to claim 1 is characterized in that: Dividing the normalized feature variable data set into a training set and a test set includes: The normalized feature variable data set is divided into a training set and a test set in a ratio of 8:
2.
5. According to claim 1, the LSTM urban traffic network CO and NOx emission prediction method based on Attention mechanism is characterized in that: Construct an Attention-LSTM model including an Attention mechanism layer, wherein the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer including: The input data set enters the LSTM layer in the form of a time series. The input of the LSTM layer is X = [x1, x2, ..., x T ], the sequence length is T, and the hidden state of the LSTM layer is h t The update formula at time step t is, h t =f(W x x t +W h h t-1 +b) Among them, f is the activation function, W x is the input weight matrix, W h is the hidden state weight matrix, b is the bias vector, x t is the input at time step t; The output residual r of the LSTM layer at time step t is t Input attention mechanism layer; In the Squeeze layer of the attention mechanism layer, the output residual r of the LSTM layer at time step t is t Compress into a global description feature vector v; Input the global description feature vector v into the extract layer; The LSTM weights and residuals are input into the fully connected layer, and the fully connected layer outputs the vector [H t ,V t ]; The output vector [H t ,V t ] Input the ReLU layer to get the final Attention weight V t ';as well as The final Attention weight V t ′ is used to weight the output of LSTM to generate the final prediction y of the Attention-LSTM model final ; Among them, H t is the reference vector, V t Score for attention; The calculation formula of reference vector and attention score is, H t =W e h t +b e V t =softmax(H t ) Among them, W e and b e are the weights and biases of the fully connected layer, and softmax is used to normalize the scores into a probability distribution; The final Attention weight V t 'for, V t ′=ReLU(V t ) The final prediction y of the Attention-LSTM model final for, Among them, h t is the hidden state of LSTM, and T is the sequence length.
6. The LSTM urban traffic network CO and NOx emission prediction method based on Attention mechanism according to claim 1 is characterized in that: Construct the loss function of the Attention-LSTM model and configure the hyperparameters of the Attention-LSTM model, including the learning rate, batch size, number of LSTM layers, and number of neurons: The loss function of the Attention-LSTM model is: Among them, Train_Y is the training set, Predict_Y is the prediction set, and n is the number of samples in the training set; The learning rate is 0.001, the batch size is 30, the number of LSTM layers is single, and the number of neurons is 128.
7. The LSTM urban traffic network CO and NOx emission prediction method based on Attention mechanism according to claim 1 is characterized in that: Evaluation metrics include root mean square error, mean absolute error, coefficient of determination, and mean absolute percentage error.
8. A system for the LSTM urban traffic network CO and NOx emission prediction method based on the Attention mechanism according to any one of claims 1 to 7, characterized in that: Includes the following modules: A data collection module is configured to collect road monitoring data and vehicle characteristic data as input data; A characteristic variable screening module is configured to perform characteristic analysis on the input data using principal component analysis and correlation coefficient method, and screen out characteristic variables related to CO and NOx emissions as input characteristic variables; A data preprocessing module is configured to perform data cleaning and normalization processing on the data set of the input feature variables to obtain a normalized feature variable data set; A data set division module is configured to divide the normalized feature variable data set into a training set and a test set; An Attention-LSTM model building module is configured to build an Attention-LSTM model including an Attention mechanism layer, wherein the Attention mechanism layer includes a squeeze layer, an extract layer, a fully connected layer, and a ReLU layer; A hyperparameter configuration module, configured to construct a loss function of the Attention-LSTM model and configure hyperparameters of the Attention-LSTM model, wherein the hyperparameters include a learning rate, a batch size, a number of LSTM layers, and a number of neurons; and The emission prediction and evaluation module is configured to input the normalized feature variable data set into the Attention-LSTM model with set hyperparameters for training, obtain the emission prediction model for CO and NOx of the traffic network, and use the evaluation index to evaluate the emission prediction model of CO and NOx.
9. A computer system, characterized in that: include: a processor configured to execute machine-readable instructions; A graphics card with an AI chip configured to train an Attention-based LSTM urban traffic network CO and NOx emission prediction method; as well as A memory configured to store machine-readable instructions, wherein the machine-readable instructions, when executed by a processor and / or a graphics card, perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Machine-readable instructions are stored thereon, and when the machine-readable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.