Traffic prediction method based on long-term and short-term space-time memory
By constructing a timing dependency graph and a memory adaptive graph, combining the input traffic sequence for multi-level feature extraction and adaptive integration, the problem that existing traffic data prediction methods are difficult to capture space-time dependence and dynamic changes is solved, and higher prediction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510184418.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Existing traffic data prediction methods are difficult to fully capture the spatial and temporal dependence and dynamic changes of traffic data, resulting in low prediction accuracy.
The traffic prediction method based on long and short-term space-time memory is adopted, and the timing dependence graph and memory adaptive graph are constructed, and multi-level feature extraction and adaptive integration are combined with the input traffic sequence to reconstruct the long-term space-time information and generate the prediction results of traffic data.
It effectively alleviates the problem that traditional models are difficult to deal with long-term and short-term spatial dependence at the same time, improves the accuracy and adaptability of traffic data prediction, and can more comprehensively identify complex spatio-temporal patterns in traffic data.
Smart Images

Figure CN120183176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic data prediction, and specifically, to a traffic prediction method based on long short-term spatio-temporal memory. Background Art
[0002] With the acceleration of the global urbanization process and the rapid growth of the number of vehicles, the problem of traffic congestion has become increasingly serious, making the importance of traffic management and traffic prediction continuously increase. Traffic data (such as traffic speed and flow) are key indicators reflecting the state of the traffic system. These data are collected and transmitted by sensor networks, providing basic support for the operation of the intelligent transportation system (ITS). As the core technology of ITS, traffic data prediction is crucial for urban traffic management and road planning, and at the same time has profound practical significance in alleviating traffic congestion.
[0003] The traffic data prediction task is similar to multivariate time series (MTS) prediction. However, compared with general MTS, traffic data has obvious spatio-temporal heterogeneity and is affected by various dynamic factors such as morning and evening rush hours and traffic accidents. This complex non-stationary characteristic brings great challenges to the prediction task. Therefore, effectively capturing the spatio-temporal dependence of traffic data while coping with dynamic changes has become an important research topic in this field.
[0004] Traditional time series prediction methods only consider a single time series and are not applicable to time series prediction tasks where there is spatial correlation between sequences such as traffic data. The existing models mainly consist of graph neural networks (GCN, GAT) and sequence models (such as RNN, transformer) to extract the spatial and time dependencies in traffic data respectively. In particular, graph neural networks require additional topological structures to represent the spatial relationships of data. Most of the existing research is based on such additional prior knowledge as geographical spatial positions to represent the correlations between sensors, but it is not comprehensive enough. For example, in a city, although a residential area and a working area may not be adjacent geographically, the traffic flow between the two shows strong dependence due to daily commuting, and there are potential spatial relationships. Some studies construct parameter nodes to learn adaptive graphs or calculate the Euclidean distance and cosine similarity between sequences to measure the topological relationships between traffic sequences. However, the adjacency matrices defined in these studies are in a static mode over time, ignoring the time-varying interaction relationships in the traffic network and unable to accurately capture the dynamic characteristics of traffic data that change over time. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a traffic prediction method based on long short-term spatio-temporal memory, which can obtain a comprehensive spatial relationship of traffic sequences, capture the dynamic characteristics of data that change over time, and improve the prediction accuracy.
[0006] The technical solution of the present invention is as follows:
[0007] A traffic prediction method based on long short-term spatio-temporal memory, specifically including the following steps:
[0008] (1) After collecting and preprocessing traffic data, use a sliding window to convert the collected traffic data into a time series form to obtain a traffic sequence;
[0009] (2) Construct a traffic prediction model. Input the traffic sequence into the traffic prediction model. First, construct a temporal dependence graph, and generate a memory adaptive graph through a memory node parameter library to obtain the adjacency matrix of the temporal dependence graph and the adjacency matrix of the memory adaptive graph. Then, the two adjacency matrices are respectively combined with the input traffic sequence to perform multi-level feature extraction through corresponding encoders. The outputs of the two encoders are then subjected to multi-level feature adaptive integration to obtain integrated encoded features. Finally, use the memory node parameter library to reconstruct long-term spatio-temporal information, and after splicing the reconstructed long-term spatio-temporal information with the integrated encoded features, use a decoder to generate the prediction result of traffic data;
[0010] (3) Divide the real traffic sequence into a training set, a validation set, and a test set, thereby training the traffic prediction model, and obtaining the prediction error combined with a contrastive learning loss function for backpropagation training, and finally generating a trained traffic prediction model.
[0011] The specific steps of collecting traffic data, preprocessing it, and then using a sliding window to convert the collected traffic data into a time series form to obtain a traffic sequence are as follows:
[0012] S11. Set the sensors in the urban traffic road network to sample traffic data of vehicle flow or speed every set time interval, and then aggregate the collected traffic data into an increment of a set duration;
[0013] S12. Use a sliding window to divide the sampled traffic data into a time series at a set step size, and perform normalization processing on the data using maximum-minimum normalization;
[0014] S13. Divide the normalized data into a training set, a validation set, and a test set, and input it into the traffic prediction model for model training, validation, and testing; the input sequence of the traffic prediction model is X h ∈R N×P , where N represents that there are N sensors in the urban traffic road network, P is the set step size of the sliding window, the traffic prediction model is used to predict traffic data for the next Q time steps, and the output sequence of the traffic prediction model is X f ∈R N×Q .
[0015] The specific steps for constructing the temporal dependence graph and generating the memory adaptive graph through the memory node parameter library to obtain the adjacency matrix of the temporal dependence graph and the adjacency matrix of the memory adaptive graph are as follows:
[0016] S21. First, input X h ∈R N×P into the gated recurrent unit GRU, retain and output the hidden representations F ∈ R N×h at each time step of the last layer, where h is the hidden feature dimension. Specifically, see the following formula (1):
[0017] F = GRU(X h ) (1);
[0018] Then, pass the hidden representation F through the linear projections of learnable weights W q , W k ∈R B×g×c to calculate the output query value Q and key value K, and then use the self-attention mechanism to calculate the attention score Score between Q and K to obtain the similarity between each traffic sequence. Specifically, see the following formula (2):
[0019]
[0020] To avoid unstable gradients caused by too large inner products, Score is divided by the square root of the linear projection dimension c, and finally the softmax function is used to normalize the inner product to obtain the adjacency matrix A s of the temporal dependence graph. Specifically, see the following formula (3):
[0021]
[0022] S22. First, construct the memory adaptive graph, randomly generate two node embedding vectors E1, E2 ∈
[0023] R N×φ , then construct the node memory parameter library M ∈ R φ×d , where φ represents the number of nodes stored in the node memory parameter library, and d is the feature dimension of the stored nodes. Then, introduce the node memory parameter library into the generation of the memory adaptive graph to enhance the extended learning of the node embedding vectors and provide additional context information for the node embedding vectors at the current moment to obtain the adjacency matrix A m of the memory adaptive graph. The specific processing process is shown in the following formula (4):
[0024] A m = softmax(relu((E1 * M)(E2 * M) T ) (4);
[0025] In formula (4), relu represents the ReLU activation function, softmax represents the exponential normalization function, and T represents the transpose.
[0026] After the two described adjacency matrices are respectively combined with the input traffic sequence, multi-level feature extraction is performed through the corresponding encoders. The specific steps for adaptively integrating the outputs of the two encoders at multiple levels to obtain the integrated encoded features are as follows:
[0027] S23. The encoder is mainly composed of graph convolutional recurrent units, as shown in the following formula (5):
[0028]
[0029] In formula (5), X t ∈R N , H t ∈R N×h represent the input and output at time t, h is the hidden feature dimension, ⊙ represents the dot product operation, u t , r t , C t respectively represent the update gate, reset gate, and candidate gate at the current time t in the graph convolutional recurrent unit, θ u , θ r , θ c respectively represent the kernel parameters required for the convolution of the update gate, reset gate, and candidate gate, * g represents the graph convolution operation, b u , b r , b c respectively represent the offsets corresponding to the update gate, reset gate, and candidate gate, σ represents the Sigmoid activation function, and tanh represents the tanh activation function;
[0030] The specific graph convolution operation is shown in the following formula (6):
[0031]
[0032] In formula (6), X is the input signal, θ is the kernel parameter, θ * g X represents performing a graph convolution operation on X, θ(L)X represents performing a graph convolution operation by introducing the adjacency matrix, L represents the symmetric normalized Laplacian matrix of the adjacency matrix of the time series dependence graph or the memory adaptive graph, T k (L)∈R N×N is the k-order Chebyshev polynomial expansion calculated on L, θ k is the kernel parameter of order k, and K represents the highest order of the Chebyshev polynomial expansion;
[0033] The input sequences X of the two encoders h, The initial state is set to zero, and the two encoders respectively output the hidden state h at the last time step obtained by the graph convolutional recurrent unit s and h m , where the hidden state h s is the short-term spatio-temporal information obtained by combining the temporal dependence graph, and h m is the long-term spatio-temporal information obtained by combining the memory adaptive graph;
[0034] S24. Feature weighting is performed on h s and h m . Then, the weighted h s and h m are integrated to obtain the integrated coding feature h ms , as shown in the following formula (7):
[0035]
[0036] In formula (7), is the parameter space for backpropagation learning of the attention mechanism, and W w ∈R B×N×N represents the weight score of h s relative to h m .
[0037] The specific steps of reconstructing the long-term spatio-temporal information using the memory node parameter library, splicing the reconstructed long-term spatio-temporal information with the integrated coding feature, and then generating the prediction result of traffic data using the decoder are as follows:
[0038] S25. Reconstruct the long-term spatio-temporal information h m to obtain the reconstructed long-term spatio-temporal information, as shown in the following formula (8):
[0039]
[0040] In formula (8), represents the i-th node vector of h m ∈R N×h , M[j] represents the node memory vector in the node memory parameter library, j ∈ (1…φ), and M T [j] represents the transpose of M[j], represents the similarity between the i-th node vector and the node memory vector M[j], represents the reconstructed i-th node vector, which is obtained by similarity weighted calculation,
[0041] S26. Splice the reconstructed long-term spatio-temporal information h ma with the integrated coding feature h ms to obtain [hma , h ms , and then input it into the decoder, which is also composed of graph convolutional recurrent units. Input the daily timestamp embedding vectors of the predicted Q time steps, and the input initial state is [h ma , h ms . Through cyclic decoding, the output of each future time step is obtained, and finally the final prediction result X of Q time steps is obtained by using linear projection f ∈ R N×Q .
[0042] The specific steps for backpropagation training by combining the predicted error with the contrastive learning loss function are as follows:
[0043] S31. First, calculate the loss between the prediction result X f and the corresponding real traffic data Y f to obtain the prediction loss L1, as shown in the following formula (9):
[0044]
[0045] S32. Use the contrastive learning loss function to constrain the adjustment of node memory parameters during training. The calculation process is specifically shown in the following formula (10):
[0046]
[0047] In formula (10), Triplet loss represents the Triplet loss function, and Compact loss represents the Compactness Loss function. The long-term spatio-temporal information of the i-th node vector is used as the anchor point i. At the same time, through the similarity S (i) ∈ R φ , two node memory vectors in the node memory parameter library M that are most similar to the anchor point i are selected as the positive sample pos and the negative sample neg, denoted as M[pos] and M[neg] respectively. Margin is used to control the distance difference between the positive and negative samples, which is a set threshold. λ1 and λ2 are balance factors, and the contrastive learning loss function L2 is used to make the node memory vectors as compact and similar as possible;
[0048] S33. Add the contrastive learning loss function L2 to the prediction loss L1 as the target criterion, and perform backpropagation training to achieve the convergence of training.
[0049] Advantages of the present invention:
[0050] (1) The present invention proposes an adaptive graph construction method based on a node memory parameter library, enabling the graph structure to have long-term spatial memory capabilities. Then, a temporal dependence graph generated through a self-attention mechanism is used to capture short-term spatial relationships, effectively alleviating the problem that traditional traffic data prediction models are difficult to simultaneously handle long-term and short-term spatial dependencies.
[0051] (2) The present invention conducts multi-level feature adaptive integration of the outputs of two encoders, overcomes the problems brought about by short-term volatility, effectively integrates short-term spatio-temporal information with strong volatility into long-term feature information with a stable trend through dynamic weight allocation, and adjusts short-term spatial fluctuations and long-term traffic trends through a dynamic weight allocation mechanism, thereby more comprehensively identifying complex dynamic spatio-temporal patterns in traffic data, better adapting to the complexity and non-stationarity of traffic data, and obtaining more reasonable prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flowchart of the present invention.
[0053] Figure 2 is a schematic diagram of the framework of the traffic prediction model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] See Figure 1 , a traffic prediction method based on long-term and short-term spatio-temporal memory, specifically including the following steps:
[0056] (1) After collecting and preprocessing traffic data, a sliding window is used to convert the collected traffic data into a time series form to obtain a traffic sequence. The specific steps are as follows:
[0057] S11: Set the sensors in the urban traffic road network to sample traffic data of vehicle flow or speed every 30 s, and then aggregate the collected traffic data into an increment every 5 min.
[0058] S12: Use a sliding window to divide the sampled traffic data into a time series every 12 steps, and normalize the data using maximum-minimum normalization.
[0059] S13. Divide the normalized data into a training set, a validation set, and a test set, and input them into the traffic prediction model for model training, validation, and testing; the input sequence of the traffic prediction model is X h ∈R N×P , where N represents that there are N sensors in the urban traffic road network, P = 12 represents the past P time steps, the traffic prediction model is used to predict the traffic data for the next Q time steps, and the output sequence of the traffic prediction model is X f ∈R N×Q ;
[0060] (2). See Figure 2 , construct a traffic prediction model for traffic data prediction, which specifically includes the following steps:
[0061] S21. First, input X h ∈R N×P into the gated recurrent unit GRU, retain and output the hidden representations F ∈ R N×h of each time step in the last layer, where h is the hidden feature dimension, as shown in the following formula (1):
[0062] F = GRU(X h ) (1);
[0063] Then, respectively pass the hidden representation F through the linear projections of the learnable weights W q , W k ∈R B×h×c to calculate the output query value Q and the key value K, and then use the self-attention mechanism to calculate the attention score Score of Q and K to obtain the similarity between each traffic sequence, as shown in the following formula (2):
[0064]
[0065] To avoid unstable gradients caused by too large inner products, Score is divided by the square root of the linear projection dimension c, and finally the softmax function is used to normalize the inner product to obtain the adjacency matrix A s of the temporal dependence graph, as shown in the following formula (3):
[0066]
[0067] S22. First, construct a memory adaptive graph, randomly generate two node embedding vectors E1, E2 ∈
[0068] R N×φ , and then construct a node memory parameter library M ∈ R φ×d, where φ represents the number of nodes stored in the node memory parameter library, d is the feature dimension of the stored nodes. Then, the node memory parameter library is introduced into the memory adaptive graph generation to enhance the extended learning of the node embedding vectors, providing additional context information for the node embedding vectors at the current moment, and obtaining the adjacency matrix A of the memory adaptive graph m , and the specific processing process is shown in the following formula (4):
[0069] A a = softmax(relu((E1 * M)(E2 * M) T )) (4);
[0070] In formula (4), relu represents the ReLU activation function, softmax represents the exponential normalization function, and T represents the transpose;
[0071] Adjacency matrix The i-th row of the matrix represents the spatial association degree between node i and other nodes, and a in the matrix ij represents the unique spatial relationship between node i and node j;
[0072] S23. After the two adjacency matrices are combined with the input traffic sequences respectively, multi-level feature extraction is performed through the corresponding encoders. The encoder is mainly composed of the graph convolutional recurrent unit GCRU. The graph convolutional recurrent unit is improved by using the graph convolution operation to replace the matrix multiplication operation in the gated recurrent unit GRU. The graph convolution can efficiently capture the relationships between nodes in the non-Euclidean space, while the gated recurrent unit can model the time dependence in the sequence data. The combination of the two enables the GCRU to simultaneously process the complex spatio-temporal patterns in the traffic data, as shown in the following formula (5):
[0073]
[0074] In formula (5), X t ∈R N , H t ∈R N×h represent the input and output at time t, h is the hidden feature dimension, ⊙ represents the dot product operation, u t , r t , C t represent the update gate, reset gate, and candidate gate at the current time t in the graph convolutional recurrent unit respectively, θ u , θ r , θ c represent the kernel parameters required for the convolution of the update gate, reset gate, and candidate gate respectively, * g represents the graph convolution operation, b u , b r , b crespectively represent the offsets corresponding to the update gate, the reset gate, and the candidate gate, σ represents the Sigmoid activation function, and tanh represents the tanh activation function;
[0075] The graph convolution operation is specifically shown in the following formula (6):
[0076]
[0077] In formula (6), X is the input signal, θ is the kernel parameter, θ* g X represents performing a graph convolution operation on X, θ(L)X represents implementing the graph convolution operation by introducing the adjacency matrix, L represents the symmetric normalized Laplacian matrix of the adjacency matrix of the temporal dependence graph or the memory adaptive graph, T k (L) ∈ R N×N is the k - order Chebyshev polynomial expansion calculated on L, θ k is the kernel parameter of order k, and K represents the highest order of the Chebyshev polynomial expansion;
[0078] The input sequence has P time steps. If the input to the current GCRU is X i , then k ∈ (t - P +
[0079] 1, …, t - 1, t). According to formula (5), the current input information X i performs a graph convolution operation with the hidden state H i-1 at the previous moment, and can obtain the reset gate r i , which determines the degree of combination of the current input information with the previous memory; from the calculated candidate gate C i viewpoint, the smaller the value of r i , the smaller the value of r i ⊙H i-1 , the more information from the previous moment needs to be discarded. Conversely, more of the hidden state at the previous moment is retained, which helps to capture the short - term relationships in the time series; the current input information X i performs a graph convolution operation with the hidden state H i-1 at the previous moment, and can also obtain the update gate u i used to control the degree to which the state information at the previous moment is brought into the current state. u i ⊙h i-1 represents the retained historical information, and (1 - u i )⊙C i represents the new information. Use u i to calculate and update the current hidden state H i ;
[0080] The two encoder input sequences X h , with the initial state set to zero. The two encoders respectively output the hidden states h at the last time step obtained by the graph convolutional recurrent units With h m , the hidden state h s is the short-term spatio-temporal information obtained by combining the temporal dependence graph, and h m is the long-term spatio-temporal information obtained by combining the memory adaptive graph;
[0081] S24. Perform feature weighting on h s and h m , and then integrate the weighted h s and h m to obtain the integrated coding feature h ms , as shown in the following formula (7):
[0082]
[0083] In formula (7), is the parameter space for backpropagation learning of the attention mechanism, and W w ∈R B×N×N represents the weight score of h s relative to h m ;
[0084] S25. To further enhance the memory and discrimination ability of the node memory parameter library for different nodes and different time points, use the node memory parameter library to reconstruct the long-term spatio-temporal information h m , to obtain the reconstructed long-term spatio-temporal information, as shown in the following formula (8):
[0085]
[0086] In formula (8), represents the i-th node vector of h m ∈R N×h , M[j] represents the node memory vector in the node memory parameter library, j ∈ (1…φ), and M T [j] represents the transpose of M[j], represents the similarity between the i-th node vector and the node memory vector M[j], represents the reconstructed i-th node vector, which is obtained by similarity weighted calculation,
[0087] S26. Concatenate the reconstructed long-term spatio-temporal information h ma and the integrated coding feature h ms to obtain [h ma , h ms , and then input it into the decoder. The decoder is also composed of a graph convolutional recurrent unit, input the daily timestamp embedding vectors of the predicted Q time steps, and the input initial state is [h ma , h ms, perform cyclic decoding to obtain the outputs at each future time step, and finally use linear projection to obtain the prediction results X for the final Q time steps. f ∈R N×Q ;
[0088] (3) Divide the real traffic sequence into a training set, a validation set, and a test set, thereby training the traffic prediction model, and obtaining the prediction error. Combine with the contrastive learning loss function to perform backpropagation training, and finally generate a trained traffic prediction model;
[0089] Among them, the specific steps of combining the prediction error with the contrastive learning loss function for backpropagation training are as follows:
[0090] S31. First, calculate the loss between the prediction result X f and the corresponding real traffic data Y f to obtain the prediction loss L1, as shown in the following formula (9):
[0091]
[0092] S32. Use the contrastive learning loss function to constrain the adjustment of the node memory parameters during training. The calculation process is specifically shown in the following formula (10):
[0093]
[0094] In formula (10), Triplet loss represents the Triplet loss function, Compact loss represents the Compactness Loss function. Take the long-term spatio-temporal information of the i-th node vector as the anchor point i. At the same time, through the similarity S (i) ∈R φ screen out the two node memory vectors in the node memory parameter library M that are most similar to the anchor point i as the positive sample pos and the negative sample neg, denoted as M[pos] and M[neg] respectively. margin is used to control the distance difference between the positive and negative samples, which is a set threshold. λ1 and λ2 are balance factors, and use the contrastive learning loss function L2 to make the node memory vectors as compact and similar as possible;
[0095] S33. Add the contrastive learning loss function L2 to the prediction loss L1 as the target criterion, and perform backpropagation training to achieve the convergence of training.
[0096] Performance evaluation:
[0097] The PEMSBAY traffic speed dataset and the PEMS03 traffic flow dataset are used for training the model, and the traffic prediction model MDGCRN of the present invention is verified to have significantly better performance than other neural network prediction models. The evaluation metrics are the mean absolute error (MAE), the root mean square error (RMSE), and the mean average percentage error (MAPE). The performance evaluation results are shown in Table 1 below:
[0098]
[0099] Table 1
[0100]
[0101] In Table 1, DCRNN is a diffusion convolutional recurrent neural network that proposes graph convolution for the diffusion process and combines GCN with a recurrent neural network; GWNet is the first to use an adaptive graph for graph convolution operations and combines dilated causal convolution to extract temporal features; StemGNN combines graph Fourier transform and discrete Fourier transform to jointly capture sequence spatial correlation and temporal dependence in the spectral domain; PDFormer proposes a Transformer that can perceive traffic delay propagation; HimNet proposes spatio-temporal meta-parameter pooling to learn spatio-temporal heterogeneity information of traffic sequences.
[0102] The smaller the three evaluation metrics are, the better the model prediction performance. From the performance evaluation results in Table 1, it can be seen that DCRNN and PDFormer only use traffic graphs defined based on geographical location to build models, GWNet and HimNet only use adaptive graphs, and StemGNN only uses instantaneous graphs. The average evaluation metrics of the MDGCRN model of the present invention for predicting 12 time steps on the PEMSBAY and PEMS03 datasets are all smaller than those of these neural network prediction models, indicating that the memory adaptive graph and temporal sequence dependence graph constructed by the MDGCRN model of the present invention can simultaneously capture long-term trend features and short-term dynamic changes, and through multi-level feature adaptive integration and a dynamic weight allocation mechanism, it can effectively balance the information contributions of different graph structures, improve the modeling ability for complex traffic scenarios, and achieve more accurate predictions.
[0103] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A traffic prediction method based on long-term and short-term spatiotemporal memory, characterized by: The specific steps include: (1) After collecting and preprocessing the traffic data, the collected traffic data is converted into a time series form using a sliding window to obtain a traffic sequence; (2) Construct a traffic prediction model and input the traffic sequence into the traffic prediction model. First, construct a time-dependent graph and generate a memory adaptive graph through the memory node parameter library to obtain the adjacency matrix of the time-dependent graph and the adjacency matrix of the memory adaptive graph. Then, the two adjacency matrices are combined with the input traffic sequence and the corresponding encoders are used to extract multi-level features. The outputs of the two encoders are then adaptively integrated to obtain integrated coding features. Finally, the memory node parameter library is used to reconstruct the long-term spatiotemporal information. After the reconstructed long-term spatiotemporal information is spliced with the integrated coding features, the decoder is used to generate the prediction results of the traffic data. (3) The real traffic sequence is divided into a training set, a validation set, and a test set to train the traffic prediction model. The prediction error is combined with the contrast learning loss function to perform back propagation training, and finally a trained traffic prediction model is generated.
2. The traffic prediction method based on long-term and short-term spatiotemporal memory according to claim 1, characterized in that: After the traffic data is collected and preprocessed, the collected traffic data is converted into a time series form using a sliding window. The specific steps for obtaining the traffic series are: S11, setting the sensors in the urban traffic network to sample traffic data of vehicle flow or speed once at a set time interval, and then aggregating the collected traffic data into increments of a set time length; S12, using a sliding window to divide the sampled traffic data into a time series according to a set step size, and using maximum and minimum standardization to normalize the data; S13, divide the normalized data into training set, validation set and test set, input them into the traffic prediction model, and perform model training, validation and testing; the input sequence of the traffic prediction model is X h ∈R N×P , where N represents the number of sensors in the urban traffic network, P is the set step size of the sliding window, and the traffic prediction model is used to predict the traffic data for the next Q time steps. The output sequence of the traffic prediction model is X f ∈R N×Q .
3. The traffic prediction method based on long-term and short-term spatiotemporal memory according to claim 2 is characterized by: The specific steps of constructing the timing dependency graph, generating the memory adaptive graph through the memory node parameter library, and obtaining the adjacency matrix of the timing dependency graph and the adjacency matrix of the memory adaptive graph are as follows: S21, first X h ∈R N×P Input into the gated recurrent unit GRU, retain and output the hidden representation F∈R of each time step in the last layer N×h , h is the hidden feature dimension, see the following formula (1): F=GRU(X h ) (1); Then the hidden representation F is passed through the learnable weights W q ,W k ∈R B×h×c The linear projection of is calculated to obtain the output query value Q and key value K, and then the self-attention mechanism is used to calculate the attention score Score of Q and K to obtain the similarity between each traffic sequence, as shown in the following formula (2): In order to avoid gradient instability caused by too large an inner product, Score is divided by the square root of the linear projection dimension c, and finally the inner product is normalized using the softmax function to obtain the adjacency matrix A of the time-series dependency graph. s , see the following formula (3) for details: S22, first build a memory adaptive graph and randomly generate two node embedding vectors E1, E2∈R N×φ , and then construct the node memory parameter library M∈R φ×d , φ represents the number of nodes stored in the node memory parameter library, d is the feature dimension of the stored node, and then the node memory parameter library is introduced into the memory adaptive graph generation to enhance the extended learning of the node embedding vector and provide additional context information for the node embedding vector at the current moment, and obtain the adjacency matrix A of the memory adaptive graph m The specific processing process is shown in the following formula (4): YOUR m =softmax(clock((E1*M)(E2*M) T ) (4)? In formula (4), relu represents the ReLU activation function, softmax represents the exponential normalization function, and T represents transpose.
4. The traffic prediction method based on long-term and short-term spatiotemporal memory according to claim 3 is characterized by: The two adjacency matrices are respectively combined with the input traffic sequence and then the corresponding encoders are used to extract multi-level features. The outputs of the two encoders are then adaptively integrated with multi-level features to obtain the integrated coding features. The specific steps are as follows: S23, the encoder is mainly composed of a graph convolution cycle unit, as shown in the following formula (5): In formula (5), X t ∈R N ,H t ∈R N×h represents the input and output at time t, h is the hidden feature dimension, ⊙ represents the dot product operation, u t 、r t , C t They represent the update gate, reset gate, and candidate gate at the current time t in the graph convolutional recurrent unit, θ u ,θ r ,θ c Respectively represent the kernel parameters required for the update gate, reset gate and candidate gate convolution, * g represents the graph convolution operation, b u 、b r 、b c They represent the offsets corresponding to the update gate, reset gate, and candidate gate, respectively. σ represents the Sigmoid activation function, and tanh represents the tanh activation function. The graph convolution operation is specifically shown in the following formula (6): In formula (6), X is the input signal, θ is the kernel parameter, and θ* g X represents the graph convolution operation on X, θ(L)X represents the graph convolution operation by introducing the adjacency matrix, L represents the symmetric normalized Laplacian matrix of the adjacency matrix of the time-dependent graph or the adjacency matrix of the memory adaptive graph, T k (L)∈R N×N is the k-order Chebyshev polynomial expansion calculated on L, θ k is the kernel parameter of order k, K represents the highest order of Chebyshev polynomial expansion; Two encoders input sequence X h , the initial state is set to zero, and the two encoders output the hidden state h of the last time step obtained by the graph convolutional recurrent unit respectively. s With h m , hidden state h s is the short-term spatiotemporal information obtained by combining the temporal dependency graph, h m To combine the long-term spatiotemporal information obtained by memorizing adaptive graphs; S24, for h s and h m Perform feature weighting and then perform weighted h s and h m Perform integration to obtain the integrated coding feature h ms , see the following formula (7) for details: In formula (7), is the parameter space of attention mechanism back propagation learning, W w ∈R B×N×N Indicates h s Relative to h m The weight score of .
5. The traffic prediction method based on long-term and short-term spatiotemporal memory according to claim 4 is characterized by: The specific steps of reconstructing the long-term spatiotemporal information by using the memory node parameter library, splicing the reconstructed long-term spatiotemporal information with the integrated coding features, and using the decoder to generate the prediction results of the traffic data are as follows: S25. Reconstructing long-term spatiotemporal information using node memory parameter library m , and obtain the reconstructed long-term spatiotemporal information, as shown in the following formula (8): In formula (8), Indicates h m ∈R N×h The i-th node vector of M[j] represents the node memory vector in the node memory parameter library, j∈(1…φ), M T [j] represents the transpose of M[j], Represents the i-th node vector The similarity with the node memory vector M[j], Represents the reconstructed i-th node vector, which is calculated by similarity weighting. S26, reconstruct the long-term spatiotemporal information h ma With the integrated encoding feature h ms After splicing, we get [h ma ,h ms ], and then input to the decoder, which is also composed of graph convolutional recurrent units, input the predicted Q time steps of daily timestamp embedding vector, and input the initial state [h ma ,h ms ], loop decoding to get the output of each future time step, and finally use linear projection to get the final prediction result X of Q time steps f ∈R N×Q .
6. The traffic prediction method based on long-term and short-term spatiotemporal memory according to claim 2 is characterized by: The specific steps of back propagation training by combining the prediction error with the contrastive learning loss function are as follows: S31, firstly, the prediction result X f And the corresponding real traffic data Y f The loss is calculated to obtain the predicted loss L1, as shown in the following formula (9): S32. Use contrastive learning loss function to constrain the adjustment of node memory parameters during training. The calculation process is specifically shown in the following formula (10): In formula (10), Triplet loss represents the Triplet loss function, Compact loss represents the Compactness Loss loss function, and the long-term spatiotemporal information of the i-th node vector is converted into As anchor point i, through the similarity S (i) ∈R φ The two node memory vectors that are most similar to the anchor point i in the node memory parameter library M are selected as the positive sample pos and the negative sample neg, respectively expressed as M[pos] and M[neg]. Margin is used to control the distance difference between positive and negative samples, and is used to set the threshold. λ1 and λ2 are balance factors. The contrastive learning loss function L2 is used to make the node memory vectors as compact as possible. S33, adding the contrastive learning loss function L2 to the prediction loss L1 as the target criterion, performing back propagation training, and achieving training convergence.
Citation Information
Patent Citations
Traffic speed prediction method and device, electronic equipment and computer readable medium
CN112863180A
Traffic flow prediction model construction method and prediction method based on adaptive dynamic graph
CN116187555A
Traffic flow prediction method based on fusion of space-time adaptive graph learning and dynamic graph convolution
CN117392846A
Generative adversarial network-based urban local carbon emission hotspot prediction and regulation method
CN118982155A
Flow prediction method based on dynamic graph space-time correlation and adaptive adversarial training
CN119211044A