Traffic flow hybrid prediction model based on self-attention-time convolutional network-generative adversarial network
By introducing a hybrid prediction model of self-attention-time convolutional network-generating adversarial network in traffic flow prediction, the problem that existing methods are difficult to capture nonlinear fluctuations and dynamics of traffic flow data is solved, and more efficient and accurate traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510191771.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
Existing traffic flow prediction methods are difficult to accurately capture the nonlinear fluctuations and dynamics of traffic flow data, resulting in limited prediction accuracy and difficulty in automatically extracting nonlinear features in the data.
A traffic flow hybrid prediction model (SA-TGAN) based on self-attention-time convolution network-generating adversarial network is proposed, which captures long-term dependencies through self-attention mechanism and time convolution network, and automatically extracts nonlinear features in data through the generation adversarial network.
It improves the accuracy and generalization ability of traffic flow prediction, significantly improves the prediction accuracy, can effectively capture the dynamic characteristics and nonlinear relationships of traffic flow, and shows strong anti-interference ability when facing external interference.
Smart Images

Figure CN120048112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and integrates a self-attention mechanism, a temporal convolutional network (TCN) and a generative adversarial network (GAN). The present invention is specifically a traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network. Background Art
[0002] Traffic flow prediction plays a vital role in the development of intelligent transportation systems. It not only provides a scientific basis for traffic planning and management, but also has a far-reaching impact on safety emergency response, optimization of travel services, etc. However, the complexity and dynamics of traffic flow data pose a huge challenge to traditional prediction methods. Traditional traffic flow prediction methods, such as ARIMA, VAR models, and linear regression models, are mainly based on the assumption of constant statistical characteristics. However, traffic flow data has significant nonlinear fluctuations and dynamics, and these traditional methods are difficult to accurately capture these characteristics, resulting in limited prediction accuracy. In addition, these methods also have obvious deficiencies in dealing with data integrity and time dependence, which further limits their application in intelligent transportation systems.
[0003] To overcome these challenges, prediction models based on intelligent methods have gradually emerged. Intelligent methods such as neural networks and machine learning models perform well in processing complex time series data, but they still face some pain points. For example, neural network models are prone to gradient vanishing or gradient exploding problems when dealing with long-term dependencies, making model training difficult; while machine learning models often rely on manual feature engineering and have difficulty automatically extracting nonlinear features from data.
[0004] Therefore, there is an urgent need for a more efficient and accurate traffic flow prediction algorithm to cope with the complexity and dynamics of traffic flow data. This algorithm needs to be able to automatically extract nonlinear features in the data, capture long-term dependencies, and have strong generalization capabilities. Based on such requirements, the present invention proposes a traffic flow hybrid prediction model (SA-TGAN) based on self-attention-temporal convolutional network-generative adversarial network, which aims to improve the accuracy and generalization ability of traffic flow prediction and provide more reliable data support for intelligent transportation systems. Summary of the invention
[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network.
[0006] The GAN network structure combines the characteristics of unsupervised learning and supervised learning. Its core consists of two basic networks: the generator (G) and the discriminator (D).
[0007] This framework is designed to handle complex data structures with a static part s and a time-varying part, whose latent variables can come from the encoding of real data or distributed random noise. In the temporal GAN, the generator G is responsible for generating data from a given random noise distribution p(z), and its goal is to make the generated data distribution G(z) as close to the real data distribution as possible. The discriminator D is responsible for distinguishing whether the input data comes from the real data distribution or is generated by the generator G. In order to train the network, a special combination of loss functions is used. On the one hand, the latent variables encoded from the real data are used to restore the original data through the decoder to construct a recovery loss function to ensure the accuracy of the encoding and decoding process. On the other hand, the error between the output of the discriminator D and the real label constitutes the loss function of unsupervised learning, which is used to guide the generator G to generate more realistic data. During the training process, a noise distribution p(z) is first randomly initialized and used as the input of the generator G to generate the corresponding output G(z). Then, the discriminator D will evaluate these outputs to determine whether they are realistic enough to fool the discriminator. Through continuous iteration and optimization, the generator G will gradually learn to generate data that is closer to the real data distribution, and the discriminator D will become more sensitive.
[0008] Temporal Convolutional Neural Network is a new time series prediction model based on Convolutional Neural Network, which can use convolution to extract features of longer time spans. Therefore, TCN has shown better performance than RNN, LSTM and other machine learning methods in many time series prediction tasks. TCN consists of three parts: causal convolution, dilated convolution and residual connection.
[0009] The attention mechanism is a method that mimics the human visual and cognitive system, which allows neural networks to focus on relevant parts when processing input data. By integrating the attention mechanism, the neural network can automatically learn and selectively focus on important information in the input, improving the performance and generalization ability of the model. The basic idea of the self-attention mechanism is that when processing sequence data, each element can be associated with other elements in the sequence, rather than just relying on elements in adjacent positions. It adaptively captures long-range dependencies between elements by calculating the relative importance between elements. Specifically, for each element in the sequence, the self-attention mechanism calculates the similarity between it and other elements and normalizes these similarities into attention weights. Then, the output of the self-attention mechanism can be obtained by weighted summing each element with the corresponding attention weight.
[0010] The present invention provides the following technical solutions:
[0011] The present invention provides a traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network. The model framework consists of two modules: the self-attention layer and the temporal convolutional layer of the generator, and the convolutional neural network of the discriminator. The working steps are as follows:
[0012] Data preprocessing: Collect historical traffic flow data and normalize the data to ensure that all data is on a standardized scale to improve the stability and efficiency of model training;
[0013] Generator module processing: The normalized traffic flow data is input into the temporal convolutional network (TCN) for processing. The receptive field of the model is expanded through causal convolution and dilated convolution structures to capture longer time dependencies, and the residual link technology is used to retain the traffic flow characteristics of historical moments.
[0014] Self-attention mechanism processing: The high-dimensional features extracted by the temporal convolutional network (TCN) are input into the self-attention mechanism for processing. By calculating the importance of each data point and weighted summing them, a more accurate and representative feature representation is generated;
[0015] Discriminator module processing: A discriminator with a convolutional neural network (CNN) structure is used to distinguish between the fake data and real data generated by the generator, extract the feature information of the data and output the discrimination result;
[0016] Adversarial training: The generator and the discriminator compete with each other in the generative adversarial network (GAN) model. The generator learns the distribution probability of the original data and generates realistic fake data to challenge the discriminator. The discriminator continuously improves its discrimination ability to accurately distinguish fake data from real data.
[0017] Model optimization and prediction result output: Optimize the parameters of the generator and discriminator through iterative training, use the trained model to predict traffic flow, and output the prediction results;
[0018] Result evaluation and optimization: Compare the predicted results with the true values, evaluate the prediction performance of the model, and optimize and adjust the model based on the evaluation results to improve the prediction accuracy and generalization ability of the model.
[0019] As a preferred technical solution of the present invention, the data preprocessing step includes: aggregating the collected data into summary data at 5-minute intervals according to the records of each detector, and using the first 80% of the data of each site for model training, and the remaining 20% of the data for model verification.
[0020] As a preferred technical solution of the present invention, in the processing step of the generator module, the temporal convolutional network (TCN) uses causal convolution to ensure that the model strictly relies on historical data at and before a certain moment in the future when predicting traffic flow, expands the receptive field of the network model by dilated convolution, and uses residual link technology to alleviate the gradient vanishing problem in deep networks.
[0021] As a preferred technical solution of the present invention, in the self-attention mechanism processing step, a series of attention weights are obtained by calculating the dot product or similarity measurement between the query vector and all key vectors, and the value vectors are weighted and summed according to these weights to generate a new feature representation.
[0022] As a preferred technical solution of the present invention, in the processing step of the discriminator module, the discriminator uses convolution operations and pooling operations to extract feature information of the data, and outputs a scalar value or probability distribution to determine whether the input data is real data or false data generated by the generator.
[0023] As a preferred technical solution of the present invention, in the adversarial training step, the parameters of the generator and the discriminator are continuously optimized through iterative training, so that the generator can generate more realistic false data, and at the same time the discriminator can improve its discrimination ability.
[0024] As a preferred technical solution of the present invention, in the model optimization and prediction result output step, indicators such as mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE) and mean absolute percentage error (MAPE) are used to evaluate the performance of the model, and the trained model is used to predict traffic flow.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. Improved prediction accuracy and generalization ability. Through comparative experiments, the SA-TGAN model showed a significant improvement in prediction accuracy in the short-term traffic flow prediction task. Compared with the baseline algorithm, the model has significant improvements in evaluation indicators such as mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE) and mean absolute percentage error (MAPE), which are increased by about 26.8%, 21.8%, 49.4% and 29.2% respectively. The generalization ability of the model has also been enhanced, and it can maintain stable prediction performance in different traffic flow scenarios.
[0027] 2. Effectively capture the dynamic characteristics and nonlinear relationships of traffic flow. The SA-TGAN model uses the self-attention mechanism to enhance the capabilities of the temporal convolutional network (TCN), which can efficiently mine the time-domain nonlinear characteristics and global dynamic dependencies in traffic flow data; the self-attention mechanism enables the model to capture the dependency between any two positions in the sequence, which is crucial for processing traffic flow data with nonlinear fluctuations.
[0028] 3. By combining the advantages of TCN and generative adversarial network (GAN), the SA-TGAN model can process local and global features at the same time and continuously optimize the prediction results through generative adversarial training; the convolution operation mechanism of TCN automatically learns and extracts nonlinear features in traffic flow data, and the characteristics of dilated convolution enable the model to expand the temporal receptive field and capture dependencies over a longer time range; the generator of GAN can simulate the dynamic changes of traffic flow and mine the dynamic B features of the input traffic flow data, while the discriminator optimizes the performance of the generator by comparing the differences between real data and predicted data.
[0029] 4. In the disturbance experiment, the SA-TGAN model showed strong anti-interference ability when facing external interferences such as Gaussian noise and Poisson noise, and the prediction performance of the model was less affected by the interference.
[0030] 5. Accurate traffic flow prediction is of great significance for traffic planning management, safety emergency response and optimization of travel services; the application of the SA-TGAN model can provide more accurate and reliable prediction data support for urban traffic management, which helps to improve the overall operating efficiency and safety of the urban transportation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0032] Figure 1 This is the GAN principle diagram;
[0033] Figure 2 This is a GAN model diagram with TCN as the core of the generator;
[0034] Figure 3 It is the structure diagram of SA-TGAN model;
[0035] Figure 4 3D surface plot of RMSE of prediction accuracy under different parameters of SA-TGAN model;
[0036] Figure 5 The curves of the predicted results and actual values of the SA-TGAN model;
[0037] Figure 6 The predicted results and actual value curves of the SA-TGAN model and the selected baseline model;
[0038] Figure 7 The curves of the predicted results and actual values of the SA-TGAN model and the selected baseline model for the same period on the second day. DETAILED DESCRIPTION
[0039] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0040] Example 1
[0041] like Figure 1-7 As shown, the present invention provides a traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network (SA-TGAN), and the prediction includes the following steps:
[0042] Step 1: Data Preprocessing
[0043] In order to fully evaluate the performance of the model of the present invention, the present invention conducted a case study using detailed traffic volume data from Caltrans PeMS. The Caltrans PeMS system deploys more than 15,000 detectors that collect traffic data every 30 seconds. In this study, the present invention focuses on traffic flow data on non-working days because the traffic patterns on these days tend to have specific regularity and predictability. Therefore, non-working day volume data within two months was selected as the basis for the study. In order to make more effective use of this data, the present invention performed data preprocessing and aggregated the collected data into summary data at 5-minute intervals according to the records of each detector. This helps to capture the changing trends of traffic flow in a short period of time while reducing the complexity of data processing. A total of 5284 data points were accumulated at each station in two months, providing a rich sample of traffic flow data. The present invention uses the first 80% of the data of each station for model training to ensure that the model can learn the main characteristics and laws of traffic flow. The remaining 20% of the data is used for model validation to evaluate the prediction performance of the model on unseen data.
[0044] Step 2: Generator module processing
[0045] The generator module is mainly composed of a temporal convolutional network. The temporal convolutional network (TCN) plays a vital role in processing traffic flow data. It first uses the characteristics of causal convolution to ensure that the model strictly relies on historical data at and before a certain moment in the future when predicting traffic flow, thereby avoiding the leakage of future information and ensuring the fairness and accuracy of the prediction. In order to further expand the receptive field of the model, that is, to enable the model to capture traffic flow changes over a longer time range, TCN introduces dilated convolution. Dilated convolution inserts holes in the convolution kernel so that the convolution operation can span multiple time steps, thereby effectively increasing the model's ability to model long-distance dependencies without increasing computational complexity. In addition, TCN also uses residual link technology, which enables the model to retain feature information from previous layers by adding direct connections between convolutional layers, which helps alleviate the gradient vanishing problem in deep networks and improves the stability and training efficiency of the model. Therefore, through the synergy of causal convolution, dilated convolution and residual link, TCN can efficiently extract the time-dependent features in traffic flow data, providing strong support for the subsequent self-attention mechanism and feedforward neural network processing.
[0046] Causal convolution means that the data at the first layer only depends on the convolution results of the data before the first layer. It is a strict time constraint model. To obtain correlation information over a longer time span, just add more hidden layers.
[0047] For the traffic flow data series X=(x 1 ,x 2 ,…,x k ,), mapping F=(f 1 ,f 2 ,f 3 ,…,f k ), in x t The causal convolution at is:
[0048]
[0049] If causal convolution wants to obtain longer dependencies, it needs to stack more hidden layers. However, adding too many hidden layers will make the model calculation more complex and time-consuming. To this end, the receptive field of the network model is expanded through dilated convolution to enable causal convolution to obtain longer dependencies. For a convolution layer, if you want to increase the receptive field of the output unit, it can generally be achieved in three ways: increase the size of the convolution kernel, increase the number of layers (for example, two layers of 3×3 convolution can approximate the effect of one layer of 5×5 convolution), and perform pooling operations before convolution. As a variant of causal convolution, dilated convolution uses the dilation factor d to dynamically set the interval size between each sampling neuron so that it can accept longer input sequences.
[0050] For the traffic flow data series X=(x 1 ,x 2 ,…,x k ,), mapping F=(f 1 ,f 2 ,f 3 ,…,f k ), in x t The dilated convolution with a dilation factor of d is:
[0051]
[0052] When the number of network layers becomes very deep, the gradient will gradually decrease, making network training more difficult and performance limited. Residual connection is a technology that uses direct connections across layers in neural networks to solve the problem of gradient disappearance and model training. Residual connection adds a jump connection between the output of the previous layer and the output of the current layer, which can learn nonlinear mapping and residual at the same time.
[0053] For the traffic flow data series X=(x 1 ,x 2 ,…,x k ,) and nonlinear mapping F, the output of the residual link is:
[0054] O(X)=F(X)+X#(3)
[0055] Step 3: Self-attention mechanism processing
[0056] The self-attention mechanism further refines and enhances the high-dimensional features extracted from the temporal convolutional network (TCN). Specifically, this processing stage first inputs the feature sequence output by the TCN into the self-attention mechanism, the core of which is the ability to dynamically calculate the correlation or association between each position (or time step) in the sequence and other positions. This calculation is usually implemented through an attention mechanism called "query-key-value", in which the features of each position are mapped to three different vectors: query, key, and value. Then, by calculating the dot product or some similarity measure between the query vector and all key vectors, a series of attention weights are obtained, which reflect the importance of different positions in the sequence to the current position. Finally, the value vector is weighted and summed according to these weights to generate a new feature representation. This process not only enhances the model's ability to capture key information, but also enables the model to flexibly adjust the degree of attention to information at different time steps, thereby more accurately understanding the dynamic changes and internal laws of traffic flow. In this way, the self-attention mechanism significantly improves the model's expressiveness and prediction accuracy when processing complex time series data.
[0057] Attention imitates the way people perceive, selectively filtering out a small amount of important information and focusing on it, while ignoring most of the less important information. The core of the attention information selection process lies in the calculation of the information weight coefficient. The self-attention mechanism calculates the attention weight by itself. Suppose the input vector sequence of a self-attention layer of the network is X i (i=1,2,…,n) dimensions are all d, and the output vector sequence of this layer is a i (i=1,2,…,n).
[0058] First, X i Multiply by the same matrix W k , W q , W v , get the key vector K i 、query vector Q i and the value vector V i Then use X 1 Output to a 1 Take the process of Q 1 After doing dot products with all the keys in turn, we get the initial attention value α 1,i .
[0059] Q i =W q ·X i #(4)
[0060] K i =W k ·X i #(5)
[0061] V i =W v ·X i #(6)
[0062] X i is the one-dimensional vector of the input data after word embedding, W k , W q , W v is a two-dimensional matrix with learnable parameters
[0063] The self-attention mechanism is divided into three steps: First, the similarity or correlation between the query and the key is calculated:
[0064] Sim(Q i ,K i )=Q i ·K i #(7)
[0065] Then use softmax to normalize the result calculated in step 1 to get the weight coefficient:
[0066] a i =Softmax(Sim i )#(8)
[0067] The weighting factor is the weighted sum of the Value.
[0068]
[0069] Step 5: Discriminator module processing
[0070] The discriminator module first receives as input the fake data from the generator and the real data directly obtained from the data source. These data are usually presented in the form of time series and contain the historical records of traffic flow. Subsequently, the discriminator processes these data using the powerful feature extraction capabilities of convolutional neural networks (CNNs). CNN gradually extracts high-level feature representations from the input data through multiple layers of convolution and pooling operations. These features not only contain the local patterns of the data, but also reflect the global structure and time dependency of the data. In the process of feature extraction, the discriminator gradually learns the distribution law of the real data and forms a deep understanding of the real data. Finally, based on these feature representations, the discriminator outputs a scalar value or probability distribution to determine whether the input data is real data or fake data generated by the generator. Through continuous adversarial training with the generator, the discriminator continuously improves its discriminative ability, while also prompting the generator to generate more realistic and indistinguishable fake data, thereby promoting the performance of the entire SA-TGAN model in the task of traffic flow prediction.
[0071] Step 6: Adversarial Training
[0072] During the training process, the generative adversarial network first randomly initializes the G model parameters θ g and D model parameters θ d , and then start iterative training. Each round of iteration is divided into two steps:
[0073] First, based on the generative ability of G, we train D with stronger discriminative ability, that is, we randomly sample l samples before time t from the real data. Then randomly sample l samples from the initial noise distribution Then generate samples through G And train and update the parameters θ of the discriminant model D d Maximize the objective function:
[0074]
[0075] where θ d The value range of is as follows:
[0076]
[0077] The discriminant model D needs to learn to better distinguish between real samples and generated samples. The higher it is, the stronger the ability to distinguish real samples is. The higher it is, the stronger the resolution of the generated samples.
[0078] The second step is to train G with stronger generation ability based on the discriminative ability of D. First, randomly sample l samples from the random distribution. Training updates the parameters θ of the generative model G g Maximize the objective function:
[0079]
[0080] where θ g The value range of is as follows:
[0081]
[0082] The generative model G must learn to generate more realistic samples so that D cannot distinguish between real samples and generated samples.
[0083] The objective function in the model includes the generator loss and the discriminator loss. The goal of the generator is to generate data that follows the real data distribution, which can be formed as shown in equation (3):
[0084] L G =-D(G(X t ))+α·L cons #(14)
[0085] The consistency loss L cons Used to reduce the reconstruction error of sensor data records, α is the coefficient of consistency loss.
[0086] The discriminator receives real data or generated data and outputs a scalar. The goal of the discriminator is to maximize and The expected difference between , which can be expressed as equation (4):
[0087]
[0088] in is the real historical data of the detector at time t, Represents the traffic flow data generated at time t.
[0089] Step 7: Model optimization and prediction result output
[0090] The parameter setting of the model uses the grid search method to determine the selection of key parameters. Test-rate determines the amount of data used for model testing. If the Test-rate is set too high, the amount of data used for training will be reduced, which may cause the model to underfit; if it is set too low, the performance of the model on the test set may not be accurate enough because the amount of data used for evaluation is too small. Disc-hidden determines the capacity of the discriminator network. A larger Disc-hidden value may enable the discriminator to learn more complex patterns, but it may also cause overfitting; a smaller value may make it difficult for the discriminator to learn the subtle features of the data, resulting in underfitting. Pretrain-learning-rate determines the update step size of the model parameters during training. A larger learning rate may cause the model to converge quickly in the early stages of training, but it may also cause oscillations near the optimal solution; a smaller learning rate may make the model converge more stably, but it may also cause the training speed to be too slow. G-learning-rate and D-learning-rate are similar to Pretrain-learning-rate, but are specifically used for training generators and discriminators. Appropriate learning rates can help the generator generate more and more realistic data and help the discriminator quickly and stably learn to distinguish between real data and generated data. Gan-epoches determines the number of rounds of model training. More cycles may enable the model to achieve better performance, but may also lead to overfitting or too long training time; fewer cycles may make the model insufficiently trained. Num_channels defines the number of convolution kernels in the TCN layer, affecting the complexity and learning ability of the model; Kernel_size defines the size of the convolution kernel and determines the range of time dependencies that the model can capture; Dropout is used to prevent overfitting and enhance the generalization ability of the model by randomly discarding some neurons. The selection of these parameters is based on task requirements, data set characteristics, and model complexity. For example, the selection of Kernel_size and Num_channels will affect the model's ability to capture time series and needs to be adjusted according to the time dependencies of the data. The selection of these parameters will directly affect the performance and results of the model.
[0091] The prediction accuracy of the model under different iterations and learning rates. Through the grid search strategy, the dynamic changes of the root mean square error (RMSE) under different combinations of generator learning rate and iteration number are observed. It can be found that when the number of iterations reaches 400, the SA-TGAN model shows excellent prediction performance, and the computational efficiency is relatively high at this time. This result is mainly attributed to the appropriate choice of the number of iterations. As the number of iterations increases, although the GAN model may perform better and better on the training data, too high a number of iterations may cause the model to overfit. Overfitting means that the model is too sensitive to the training data, making it difficult to maintain good performance on unseen test data. For GAN, overfitting may cause the generated samples to lack diversity or have significant differences from the distribution of real samples. In addition, the training process of GAN itself is unstable. Especially at a high number of iterations, the "competition" between the generator and the discriminator may be too fierce, resulting in fluctuations or oscillations during the training process, which in turn affects the prediction performance of the model.
[0092] In addition, compared with the learning rate of 0.0007, a learning rate that is too low will also lead to poor prediction results. The learning rate determines the step size of the model's parameter update at each iteration. When the learning rate is too low, the amplitude of each parameter update of the model is very small, resulting in a slower convergence speed. Within a limited number of iterations, the model may not be able to fully learn the intrinsic characteristics of the data, making it difficult to produce high-quality prediction results. Although lowering the learning rate can improve the stability of training, in some cases, a learning rate that is too small may also cause the training process to be too stable and lack the necessary "exploration" to drive the model to further learn. This may cause the model to stagnate in the later stages of training and fail to further improve the prediction performance.
[0093] Step 8: Result evaluation and optimization
[0094] To evaluate the accuracy of our method, we compared SA-TGAN with five baseline algorithms: convolutional neural network, artificial neural network (RNN), long short-term memory network, gated recurrent unit (GRU), and temporal convolutional network.
[0095] CNN is a feedforward neural network inspired by the natural visual cognition mechanism of organisms. It avoids complex pre-processing of images and can directly input the original image, so it is widely used in image classification, target recognition and other fields. The basic structure of CNN consists of convolutional layers, pooling layers and fully connected layers, but the specific formula varies depending on the specific network structure. Generally, the convolution operation can be expressed as:
[0096] y=conv(x,w)+b#(20)
[0097] Where x is the input, w is the convolution kernel, b is the bias term, and y is the convolution result.
[0098] RNN is a neural network that processes sequence data. It captures the temporal dependencies in the sequence through cyclic connections. RNN is suitable for data with time series or sequence structures, such as language models, machine translation, etc. The hidden state update formula of RNN is:
[0099] h t =f(W x h x +W h h t-1 +b)#(21)
[0100] where h t represents the hidden state of the current time step, h x represents the input of the current time step, h t-1 represents the hidden state of the previous time step, W x and W h is the weight matrix, b is the bias term, and f is the activation function.
[0101] LSTM is a special RNN that can learn long-term dependencies. It uses a gating mechanism to solve the gradient vanishing and gradient exploding problems in long sequence training. LSTM includes gating mechanisms such as forget gate, input gate, and output gate, and its calculation formula is relatively complex. Taking the forget gate as an example, its formula is:
[0102] f t =σ(W f *[h t-1 ,x t ]+b f )#(twenty two)
[0103] Where σ is the sigmoid function, W f is the weight matrix of the forget gate, [h t-1 ,x t ] is the concatenation of the hidden state of the previous time step and the input of the current time step, b f is the bias term.
[0104] GRU is a variant of RNN. Like LSTM, it is also proposed to solve problems such as long-term memory and gradient in back propagation. GRU is simpler than LSTM and easier to train or calculate. GRU includes two gating mechanisms: reset gate and update gate. Taking the reset gate as an example, its formula is:
[0105] r t =σ(W r *[h t-1 ,x t ]+br )#(twenty three)
[0106] Where σ is the sigmoid function, W r is the weight matrix of the forget gate, [h t-1 ,x t ] is the concatenation of the hidden state of the previous time step and the input of the current time step, b r is the bias term.
[0107] TCN is a deep learning model for processing time series data. It is based on the idea of CNN and uses convolution operations to extract and learn features in time series data. The convolution operation of TCN is similar to that of traditional CNN, but it uses specific convolution kernels (such as causal convolution, dilated convolution, etc.) to capture long-term dependencies in time series data.
[0108] By comparing the degree of fit between the actual value and the predicted value, it can be clearly seen that this experiment has achieved relatively ideal results in the prediction results. Compared with the baseline algorithm, this experimental model has shown significant advantages in multiple indicators. Specifically, the performance indicators of this experimental model have been improved. Among them, MAE has increased by about 26.8% compared with the baseline algorithm that produces the maximum result, and MAPE, MSE, and RMSE have increased by about 21.8%, 49.4%, and 29.2%, respectively. These improvements all indicate that this experimental model has achieved significant improvements in prediction accuracy.
[0109] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network, characterized by: The model framework consists of two modules: the self-attention layer and temporal convolution layer of the generator, and the convolutional neural network of the discriminator. The construction steps are as follows: Data preprocessing: Collect historical traffic flow data and normalize the data to ensure that all data is on a standardized scale to improve the stability and efficiency of model training; Generator module processing: The normalized traffic flow data is input into the temporal convolutional network (TCN) for processing. The receptive field of the model is expanded through causal convolution and dilated convolution structures to capture longer time dependencies, and the residual link technology is used to retain the traffic flow characteristics of historical moments. Self-attention mechanism processing: The high-dimensional features extracted by the temporal convolutional network (TCN) are input into the self-attention mechanism for processing. By calculating the importance of each data point and weighted summing them, a more accurate and representative feature representation is generated; Discriminator module processing: A discriminator with a convolutional neural network (CNN) structure is used to distinguish between the fake data and real data generated by the generator, extract the feature information of the data and output the discrimination result; Adversarial training: The generator and the discriminator compete with each other in the generative adversarial network (GAN) model. The generator learns the distribution probability of the original data and generates realistic fake data to challenge the discriminator. The discriminator continuously improves its discrimination ability to accurately distinguish fake data from real data. Model optimization and prediction result output: Optimize the parameters of the generator and discriminator through iterative training, use the trained model to predict traffic flow, and output the prediction results; Result evaluation and optimization: Compare the predicted results with the true values, evaluate the prediction performance of the model, and optimize and adjust the model based on the evaluation results to improve the prediction accuracy and generalization ability of the model.
2. According to claim 1, a traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network is characterized in that: The data preprocessing step includes: aggregating the collected data into summary data at 5-minute intervals according to the records of each detector, and using the first 80% of the data of each site for model training and the remaining 20% of the data for model verification.
3. According to claim 1, a traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network is characterized in that: In the processing step of the generator module, the temporal convolutional network (TCN) uses causal convolution to ensure that the model strictly relies on historical data at and before a certain moment in the future when predicting traffic flow. The receptive field of the network model is expanded by dilated convolution, and the residual link technology is used to alleviate the gradient vanishing problem in the deep network.
4. The traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network according to claim 1 is characterized in that: In the self-attention mechanism processing step, a series of attention weights are obtained by calculating the dot product or similarity measure between the query vector and all key vectors, and the value vectors are weighted and summed according to these weights to generate a new feature representation.
5. The traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network according to claim 1 is characterized in that: In the processing step of the discriminator module, the discriminator extracts feature information of the data using convolution operations and pooling operations, and outputs a scalar value or probability distribution for determining whether the input data is real data or false data generated by the generator.
6. The traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network according to claim 1 is characterized in that: In the adversarial training step, the parameters of the generator and the discriminator are continuously optimized through iterative training, so that the generator can generate more realistic false data, and the discriminator can improve its discrimination ability.
7. The traffic flow hybrid prediction model based on self-attention-temporal convolutional network-generative adversarial network according to claim 1 is characterized in that: In the model optimization and prediction result output step, indicators such as mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE) and mean absolute percentage error (MAPE) are used to evaluate the performance of the model, and the trained model is used to predict traffic flow.
Citation Information
Cited By
Network traffic intrusion detection method based on knowledge tracking model
CN120811776A
Data-mechanism hybrid-driven mixed-flow manufacturing system performance evaluation method
CN121187237A
Data center non-intrusive load monitoring method based on generative adversarial network
CN122084969A