Traffic flow time sequence prediction system based on point block cross attention
Through the traffic flow timing prediction system of point-block cross attention, the parallel design of point encoder and block encoder is used, combined with the multi-head cross attention mechanism, the problem of feature fragmentation in the traditional method is solved, and efficient and accurate prediction of traffic flow is achieved, which is suitable for complex traffic data of cities and highways.
Patent Information
- Application Number
- CN202510582646.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional traffic flow prediction methods have significant limitations in capturing the macro trend of traffic flow and accurately retaining key details such as short-term fluctuations and mutation points. They cannot effectively integrate point-level and block-level features, resulting in blurred details or insufficient dependence on the prediction results.
The traffic flow timing prediction system based on point-block cross attention is adopted. The point encoder and block encoder are designed in parallel, combined with the multi-head cross attention mechanism to achieve the fusion of point-level and block-level features. The local detailed features are extracted using the convolutional network, and the Transformer encoding captures long-term trends, and the cross-granular feature fusion is achieved through the cross-granular attention mechanism.
It improves the accuracy and generalization ability of traffic flow forecasts, can capture short-term fluctuations and long-term trends at the same time, is suitable for complex traffic flow data in cities and highways, and supports traffic management and travel planning.
Smart Images

Figure CN120299252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic flow prediction, and particularly to a traffic flow time series prediction system based on point-block cross attention. Background Art
[0002] In the fields of artificial intelligence and time series analysis, long-term time series prediction of traffic flow aims to model and predict the traffic flow changes in a relatively long future period (such as hourly, daily, weekly) based on historical traffic condition data. The core challenge lies in how to capture the macroscopic trends of traffic flow (such as the seasonal periodicity of urban commuting, the long-term traffic flow growth trend of arterial roads) while accurately retaining key detailed information such as short-term fluctuations and mutation points (such as a sudden drop in traffic flow on a section caused by a traffic accident, a regional traffic surge caused by the dispersal of a large event). And the macroscopic trend is an important basis for traffic management departments to allocate resources across cycles (such as long-term road network planning, public transport capacity deployment), while the high-frequency details directly affect the accuracy of real-time traffic control (such as emergency lane scheduling, temporary signal timing).
[0003] However, the traditional long-term time series prediction method PatchTST (patch time series transformer) based on Transformer and chunking strategy exposes significant limitations in the traffic flow estimation scenario:
[0004] Detail loss of the chunking strategy: Existing methods divide the time series with a fixed window (such as dividing 1-day data into 24 1-hour chunks), and perform dimensionality transformation through operations such as linear projection. Essentially, it is a dimensionality reduction and compression of timestamp-level details. This processing causes key detailed information in traffic flow (such as a 10-minute abnormal increase in traffic flow caused by a traffic accident at a certain intersection) to be averaged or filtered during the chunking process, resulting in the problem of "trends can be distinguished, details are blurred" in model prediction, and it is impossible to accurately capture the instantaneous impact of sudden situations on traffic flow.
[0005] One-sidedness of single-view feature extraction: Traditional methods either use a recurrent neural network model to process point by point (focusing on short-term fluctuations of a single timestamp), or use a chunked Transformer to process (focusing on chunk-level trends). However, the complex characteristics of traffic flow require the model to simultaneously consider local details and global dependencies. For example, the overall rising trend of traffic flow during the morning rush hour (chunk-level feature) needs to be co-modeled with the real-time congestion changes at each intersection (point-level feature). Existing methods lack a co-processing mechanism for these two types of features from the same data source, resulting in prediction results that either lose short-term abnormal fluctuations or cannot capture long-term periodic patterns.
[0006] Lack of cross-granularity information fusion: The point-level features in traffic flow data (such as the second-level flow values of individual sensors) and block-level features belong to different granularity representations. Traditional methods do not design explicit interaction mechanisms, making it impossible for block-level features to perceive point-level anomalies, and it is also difficult for point-level features to participate in long-term dependence modeling. This fragmentation is particularly prominent in traffic flow prediction - for example, the overall traffic flow trend (block-level) during holidays needs to be comprehensively modeled by combining abnormal activities at different times of the day (such as sudden increases in midday traffic around scenic spots - point-level features), and existing methods cannot achieve effective fusion of cross-granularity information from the same data source. Summary of the Invention
[0007] In view of the deficiencies in the prior art, the present invention provides a traffic flow time series prediction system based on point-block cross-attention.
[0008] The present invention achieves the above technical objectives through the following technical means.
[0009] A traffic flow time series prediction system based on point-block cross-attention includes:
[0010] Step 1, preprocess the single-modal multi-channel time series data collected by urban traffic sensors;
[0011] Step 2, extract feature vectors of point level for the single-channel data at each timestamp in the single-modal multi-channel time series data preprocessed in Step 1 through a point encoder;
[0012] Step 3, perform block processing on the single-modal multi-channel time series data preprocessed in Step 1 through a block encoder, and extract block-level feature vectors;
[0013] Step 4, use a multi-head cross-attention mechanism, take the block-level feature vector as the query, and the point-level feature vectors as both the key and the value at the same time to obtain fused feature vectors;
[0014] Step 5: Input the fused feature vectors into the micro-training prediction module, and output the multi-channel time series prediction results after inverse normalization processing.
[0015] Furthermore, the point encoder includes two convolutional layers and a ReLU activation function.
[0016] Even further, the formula for the point-level feature vector is expressed as:
[0017]
[0018] point encoder(H) = conv2(ReLU(conv1(H)))
[0019] Where, is the point-level feature vector, H ∈ R C×S is the preprocessed single-modal multi-channel time series data, point encoder is the point encoding function, conv1 and conv2 are convolutional layers, ReLU is the activation function, C is the number of channels, S is the sequence length, D is the feature dimension, and W is the format conversion matrix.
[0020] Furthermore, the steps for the block encoder to extract the block-level feature vector include partitioning the single-modal multi-channel time series data through a sliding window, linear projection and position embedding, and Transformer encoding.
[0021] Even further, the formula for the block-level feature vector is expressed as:
[0022] H p = Patch(H)
[0023]
[0024] where Patch is the partitioning function, H ∈ R C×S is the preprocessed single-modal multi-channel time series data, H p ∈ R C×P×L is the partitioning result, Droupout is a regularization technique, Linear is the linear layer, W pos is the position embedding, is the output of the long-term feature, LN is the layer normalization, MHSA is the multi-head self-attention mechanism, is the block-level feature vector, C is the number of channels, S is the sequence length, P is the number of partitions, L is the partition dimension, and D is the feature dimension.
[0025] Furthermore, the formula for the fused feature vector is expressed as:
[0026]
[0027] Z = MHCA(Q, K, V) = Concat(head1, … head i , … head h )W O
[0028]
[0029] where W1, W2, and W3 are weight matrices, b1, b2, and b3 are bias vectors, Q, K, and V are the query, key, and value respectively, is the block-level feature vector, is the point-level feature vector, Z ∈ R C×P×Dis the fused feature vector, MHCA is the multi-head cross-attention mechanism, Concat is the fusion function, head i is the head module of the multi-head mechanism, W O is the linear transformation matrix of the output of the multi-head cross-attention, CrossAttention is the cross-attention function in each head module, softmax is an activation function, K T is the transposed matrix of K, is the scaling factor, C is the number of channels, P is the number of blocks, and D is the feature dimension.
[0030] Furthermore, the formula of the micro-training prediction module is expressed as:
[0031]
[0032] Among them, is the fused feature vector after flattening, F ∈ R C×Y is the output of the micro-training prediction module, Droupout is a regularization technique, Linear is a linear layer, W f is the weight matrix, b f is the bias vector, and Y is the length of the prediction task for each channel.
[0033] Furthermore, the formula of the inverse normalization is expressed as:
[0034]
[0035] Among them, F out ∈ R C×Y is the multi-channel time series prediction result, Denormlization is the inverse normalization function, F i,j is the time series timestamp value in the output F of the micro-training prediction module, μ i and σ i are the mean and standard deviation of the data of the i-th channel respectively, and i and j correspond to the i-th channel and the j-th timestamp value respectively.
[0036] The beneficial effects of the present invention are:
[0037] (1) Dual - perspective encoding structure: Through the parallel design of the point encoder and the block encoder, a complementary multi - scale feature representation system is constructed for the same traffic flow data. The point encoder focuses on extracting local detailed features at the timestamp level, and can accurately capture the fluctuations and outliers in traffic flow in a short time, such as the short - term congestion changes caused by sudden traffic accidents; the block encoder, with the help of the block - splitting and Transformer mechanism, effectively captures long - term dependencies across blocks, and can identify seasonal trends and periodic patterns of traffic flow, such as the morning and evening rush hours on weekdays and special traffic flow patterns during holidays, thus solving the problems of detail loss or insufficient capture of long - term dependencies in single - perspective modeling of existing methods in traffic flow prediction.
[0038] (2) Innovation of cross - attention mechanism: Pioneeringly propose a cross - granularity attention mechanism with block - level features as queries and point - level features as keys and values, breaking through the limitation of detail loss caused by feature dimensionality reduction in traditional block - splitting methods. Based on the same traffic flow data, this mechanism dynamically adjusts the attention weights, enabling the model to accurately focus on the detailed features of key timestamps when modeling the global trend of traffic flow, such as the mutation points of road section traffic flow when large - scale events disperse and abnormal traffic flow values caused by special weather, realizing the deep fusion of multi - scale features and significantly enhancing the model's representation ability for complex traffic flow patterns.
[0039] (3) Efficient long - term prediction ability: The Transformer structure of the block encoder has a natural advantage in processing long sequences. Combined with the sliding window block - splitting strategy, while reducing the computational complexity, it can still maintain the ability to capture the long - term trend of traffic flow. This feature is especially suitable for complex scenarios such as traffic flow prediction that require cross - cycle decision - making. Whether it is predicting the traffic flow differences between weekdays and weekends within a week or estimating the long - term trend of traffic flow during holidays, it can be completed efficiently and accurately.
[0040] (4) Improvement of generalization ability: Through the explicit interaction between point - level and block - level features of the same traffic flow data, the model has strong generalization ability and can effectively process different types of time - series data. In traffic flow prediction, whether facing the traffic flow data of urban roads with high - frequency changes during morning and evening rush hours or the traffic flow data of highways with different holiday cycle characteristics, the model can achieve accurate modeling through an adaptive feature fusion strategy, providing reliable data support for traffic management, travel planning, etc. Brief Description of the Drawings
[0041] Figure 1 It is a structural diagram of the traffic flow time - series prediction system based on point - block cross - attention according to the present invention. Detailed Embodiment
[0042] The present invention will be further described below in conjunction with the drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.
[0043] This embodiment proposes a traffic flow time series prediction system based on dot-block cross-attention, as Figure 1 shown, which includes the following steps:
[0044] Step 1, data preprocessing. Represent the single-modal multi-channel time series data as X ∈ R C×S , where C is the number of channels and S is the sequence length, and normalize X to obtain the processed single-modal multi-channel time series data H ∈ R C×S .
[0045]
[0046] Among them, μ i and σ i are the mean and standard deviation of the data of the i-th channel respectively, and i and j are the i-th channel and the j-th timestamp value respectively. The normalized data is input into the subsequent module for feature extraction.
[0047] Step 2, point-level detailed feature extraction. Based on the single-modal multi-channel time series data H in Step 1, the point encoding module extracts local detailed features for the single-channel data of each timestamp using a two-layer convolutional network structure to obtain the point-level feature vector The formula is expressed as:
[0048]
[0049] point encoder(H) = conv2(ReLU(conv1(H)))
[0050] Among them, point encoder is the point encoding function, which maps from 1 dimension to multiple dimensions for each timestamp value and extracts detailed features at one time. Point encoder contains two convolutional layers and an activation function. Conv1 and conv2 are convolutional layers, which perform sliding convolutional operations on the data to extract local features of the data. ReLU is the activation function, which introduces non-linearity to the output of the convolutional layer. D is the feature dimension and W is the format conversion matrix.
[0051] The first convolutional layer conv1: The number of input channels is 1 (single timestamp single channel value), the number of output channels corresponds to the feature dimension, the number of output channels is D, and the convolutional kernel size is 3×3, which is used to capture local context information near the timestamp. The activation function uses ReLU to introduce non-linear transformation.
[0052] The second convolutional layer conv2: The number of input and output channels is both D, and the convolutional kernel size is 3×3, which further refines the features and enhances the expression of detailed information.
[0053] After two layers of convolutional processing, a point-level feature vector is obtained through format transformation.
[0054] Step 3, Block-level long-term feature extraction. Similarly based on the single-modal multi-channel time series data H in Step 1, the block encoding module processes the time series data using a block strategy to capture the global long-term trend and obtain the block-level feature vector H~. The general formula is as follows:
[0055]
[0056] Among them, patch encoder is the block encoding function, which performs a block operation on H, and then performs feature mapping on each block of each channel to extract long-term features.
[0057] The following are the specific steps:
[0058] Step 3.1, Block operation: For each channel of the input single-modal multi-channel time series data H, use a sliding window to divide it into P blocks, with a window length of L and a step size of L / 2 (or other adjustable parameters), ensuring overlap between blocks to retain continuity. After division, the dimension of each block is L, and finally the block result H p ∈R C×P×L :
[0059] H p = Patch(H)
[0060] Among them, Patch is the block function, and H p is the block result. At this time, each channel stores a two-dimensional array containing the number of blocks and the size of the blocks.
[0061] Step 3.2, Linear projection and position embedding: Through the linear layer Linear, each block H p inside H p n ∈R L is mapped from the L-dimensional block length to a D-dimensional feature vector, where n = {0, 1, 2,..., P - 1}, corresponding to the specific position of the block, used to extract the detailed information of each block and obtain the long-term features of the time series. Then add the position embedding W pos (such as sine position encoding) to retain the time sequence order information, and finally obtain the output of the long-term features
[0062]
[0063] Among them, Droupout is a regularization technique used to prevent the model from overfitting during training.
[0064] Step 3.3, Transformer encoding: Input the Transformer encoding layer, which includes multi-head self-attention and layer normalization, to obtain the block-level feature vector The formula is expressed as:
[0065]
[0066] Among them, LN (Layer Normalization) is layer normalization, and MHSA (Multi-Head Self-Attention) is the multi-head self-attention mechanism. The long-term dependencies between blocks are captured by the MHSA block encoder, and the block-level feature vector is output Characterize the global trend of the sequence
[0067] Step 4: Multi-head cross-attention feature fusion. Use the multi-head cross-attention mechanism MHCA (Multi-Head Cross-Attention) to interact the point-level and block-level feature vectors extracted from the single-modal multi-channel time series data H in steps 2 and 3 to obtain the fused feature vector. The following are the specific steps:
[0068] Step 4.1, Definition of query (Q, Query), key (K, Key), and value (V, Value): Generate the query Q through linear transformation with the block-level feature vector ~H, and the point-level feature vector Generate the key K and value V through linear transformation. The formula is expressed as:
[0069]
[0070] Among them, W1, W2, and W3 are weight matrices, and b1, b2, and b3 are bias vectors. The bias vectors are used to adjust the adaptability of the feature space
[0071] Step 4.2, Point-block feature fusion calculation: Use multi-head cross-attention to obtain the fused feature vector Z. The formula is expressed as:
[0072] Z = MHCA(Q, K, V) = Concat(head1,…head i ,…head h )W O
[0073]
[0074] Among them, head iIt is the head module of the multi-head mechanism, which calculates the attention of multiple subspaces in parallel through the multi-head mechanism; CrossAttention is the cross-attention function inside each head module, and softmax is an activation function that converts the output of the multi-classification model into a probability distribution; Q, K, and V are the query, key, and value in the attention mechanism; K T is the transposed matrix of K; is the scaling factor, representing the dimension of the key K; Concat is the fusion function that performs weighted summation on the results of multiple head modules; W O is the linear transformation matrix of the output of the multi-head cross-attention, Z ∈ R C×P×D is the fused feature vector. This mechanism enables each block-level feature to adaptively focus on relevant point-level details (such as the key timestamps within the time window corresponding to the block), achieving "detail enhancement under the guidance of the global trend" and avoiding the problem of detail loss in traditional block methods.
[0075] Step 5: Output of the micro-training prediction module and inverse normalization processing
[0076] Input the fused feature vector Z ∈ R C×P×D into the prediction head module, and generate the final prediction result through the following steps:
[0077] Step 5.1, Flattening and dimension transformation: Perform a flattening operation on the fused feature vector Z to convert the three-dimensional tensor into a two-dimensional feature vector for easy processing by the fully connected layer. The formula is expressed as:
[0078]
[0079] where, is the flattened fused feature vector, P is the number of blocks, D is the feature dimension, and the features of each channel after flattening are concatenated into a vector of length P * D.
[0080] Step 5.2, Mapping by the fully connected layer and regularization: Map the flattened fused feature vector to the prediction result vector through the linear transformation layer. During this process, the length of the vector changes from P * D to Y, where Y corresponds to the length of each channel's prediction task, and Dropout regularization is introduced to prevent overfitting. Then, the output F ∈ R C×Y of the micro-training prediction module is obtained. The formula is expressed as:
[0081]
[0082] where, Linear is the linear layer; W f is the weight matrix, and W f ∈ R C×(P×D) ; b f is the bias vector, and bf ∈R Y ; Dropout is a regularization technique.
[0083] Step 5.3, Inverse normalization processing: Perform inverse normalization on the output F of the micro-training prediction module to restore the data scale. Let the mean of each channel calculated in the training phase be μ i and the standard deviation be σ i , the input data is F, and the final output multi-channel time series prediction result is F out ∈R C×Y , Y is the length of the prediction task for each channel, corresponding to the prediction result. The inverse normalization formula is:
[0084]
[0085] where Denormlization is the inverse normalization function, which performs inverse normalization operations on each time series timestamp value F i,j in the output F of the linear layer to restore the data size, μ i and σ i are the mean and standard deviation of the data in the i-th channel respectively, and i and j correspond to the i-th channel and the j-th timestamp value respectively.
[0086] Example:
[0087] Step 1, Traffic data preprocessing:
[0088] In this example, the single-modal multi-channel time series data (urban traffic dataset) collected by urban traffic sensors is used. The urban traffic data is stored in a table form, with each column being a monitoring channel and each row corresponding to the traffic flow value at a time stamp. This dataset contains 862 monitoring channels, and each channel corresponds to 96 time stamps.
[0089] However, urban traffic data often has missing or abnormal values due to sensor failures, transmission delays, etc., and the traffic flow scales at different monitoring points vary significantly. If not processed, it will lead to distorted detail features or deviated trend features. Therefore, it is necessary to preprocess the data. The preprocessing steps are as follows:
[0090] First, in the data cleaning stage, for missing values, the mean of adjacent time points is used for filling. For example, if the data at a certain moment is missing, the average of its previous and next time points is taken for filling to ensure the continuity of the time series; for extreme values exceeding three times the standard deviation, the moving average of the previous and next time points is used for correction to avoid the interference of sudden abnormal values on the overall trend. For example, if the traffic flow suddenly increases due to a traffic accident at a certain minute, it is smoothed by the moving average of the previous and next 5 minutes to restore the true traffic flow fluctuation trend.
[0091] Secondly, the cleaned data is converted into a multi-dimensional matrix. The number of channels corresponds to the number of monitoring points, and the sequence length corresponds to the total number of timestamps. Then, normalization is performed. The mean and standard deviation are calculated separately for the data of each channel. Using the standard normalization method, the data is converted into a distribution with a mean of 0 and a standard deviation of 1, eliminating the differences in the flow scales of different monitoring points. For example, the basic flows of the main road and the secondary road are different. After normalization, the dimensions are unified, laying a foundation for subsequent feature extraction. The normalization formula is as follows:
[0092]
[0093] Among them, H i,j is the urban traffic data after normalization processing, X i,j is the urban traffic data before normalization processing, μ i and σ i are the mean and standard deviation of the data of the i-th channel respectively.
[0094] Step 2, Point-block dual-view feature extraction:
[0095] Urban traffic data has both high-frequency details of short-term fluctuations (such as minute-level flow mutations caused by sudden accidents) and low-frequency trends of long-term evolution (such as the periodic patterns of morning and evening rush hours on weekdays). Traditional single-view processing methods are difficult to capture these two types of features simultaneously. Point-by-point processing can retain details but is difficult to model long-distance dependencies; block processing can capture trends but loses key details due to averaging operations. In this embodiment, a dual-view feature extraction mechanism with point-level and block-level parallelism is designed for the same set of urban traffic data. Point-level processing is based on timestamp-level data, and uses the local perception ability of the convolutional network to accurately extract the detailed features of short-term flow fluctuations (such as the amplitude and direction of flow mutations at specific time points). Block-level processing divides the sequence into overlapping time blocks through a sliding block strategy, and uses the self-attention mechanism of the Transformer to capture cross-block long-term dependencies (such as weekly cyclic flow patterns and seasonal flow baseline changes). The two form differential feature representations based on the same set of urban traffic data. Point-level features focus on micro fluctuations, and block-level features represent macro trends, providing complementary multi-dimensional information for subsequent cross-scale fusion and solving the problem of the disconnection between details and trends in traditional methods.
[0096] Step 2.1, Point-level detailed feature extraction:
[0097] The point encoder captures the details of local traffic changes at each timestamp layer by layer through a two-layer convolutional network. The first convolutional layer sets the input channels to 1, the output channels to 32, and the convolutional kernel size to 3×3, which can capture the short-term traffic changes at the current timestamp and the timestamps before and after it, such as sudden local congestion or instantaneous traffic fluctuations caused by lane changes. After the convolutional operation, the ReLU activation function is used to enhance the non-linear feature expression of the rising or falling trend of traffic, and distinguish different traffic change patterns in different directions. The input and output channels of the second convolutional layer are both 32, and the convolutional kernel size is 3×3, which further refines the features and extracts more complex local patterns, such as the periodic traffic fluctuations every 5 minutes during the morning and evening rush hours. After two-layer convolutional processing, a 32-dimensional feature vector is generated for each timestamp, containing the detailed information of short-term traffic changes, such as the time point of traffic mutation and the fluctuation amplitude, and finally a point-level feature tensor with the dimension of 862×96×32 is formed, completely retaining the microscopic traffic fluctuation features of each timestamp. The formula is as follows:
[0098]
[0099] point encoder(H)=conv2(ReLU(conv1(H)))
[0100] where, is the point-level feature vector, H∈R 862×96 is the normalized urban traffic data, point encoder is the point encoding function, conv1 and conv2 are convolutional layers, which perform sliding convolutional operations on the data to extract the local features of the data, ReLU is the activation function, which introduces non-linearity to the output of the convolutional layer, D is the feature dimension, and W is the format conversion matrix, and the format is adjusted to facilitate the calculation during feature fusion.
[0101] Step 2.2, Block-level long-term feature extraction:
[0102] The block encoder divides the normalized urban traffic data into blocks through a sliding window to capture the long-term trends across time blocks. Specifically, when dividing the blocks, the window length is set to 16 timestamps, and the data of each channel is divided into 12 blocks. Each block is mapped to a 32-dimensional feature space through a linear projection layer, and positional embeddings are added to provide the position information of the time series data. The projected block features are input into the Transformer encoding layer, which contains 4 encoding units. Each unit uses the multi-head self-attention mechanism (through multi-head self-attention, the model can capture the long-term dependencies between different blocks, such as the weekly cyclic traffic patterns, the traffic change trends before and after holidays, etc.) and layer normalization. The finally output block-level feature tensor has a dimension of 862×12×32, representing the global long-term trends of the sequence, such as the traffic differences in different seasons, the regular fluctuations between weekdays and weekends, etc. The formula is as follows:
[0103] H p = Patch(H)
[0104]
[0105] where Patch is the block division function, and H p ∈R 862×12×16 is the result of block division. At this time, each channel stores a two-dimensional array containing the number of blocks and the size of the blocks; Linear is the linear projection layer, which performs feature mapping on several blocks of each channel, maps the block vector of length 16 to a feature vector of length 32 to extract long-term features; Dropout is the regularization function to prevent overfitting; W pos is the positional embedding vector, is the output of the long-term features. The Transformer encoding layer contains the multi-head self-attention mechanism to capture the long-term dependencies between blocks, and finally outputs the block-level feature vector representing the global trend
[0106] Step 2.3, Point-block feature fusion:
[0107] Essentially, the point-level and block-level features are two different perspectives of the same set of data. In order to fuse the point-level detailed features into the block-level long-term features to obtain a set of fused features, this embodiment uses the cross-attention mechanism to complete the fusion of the two types of features. The cross-attention mechanism uses the block-level features as queries, and the point-level features as both keys and values at the same time to achieve cross-granularity feature interaction. In the specific operation, the block-level features generate query vectors through linear transformation, the point-level features generate key and value vectors through linear transformation, and the multi-head cross-attention is used to calculate the weights, so that each block-level feature can adaptively focus on the relevant point-level details. The formula is as follows:
[0108]
[0109] Z = MHCA(Q, K, V) = Concat(head1, … head i ,... head h )W O
[0110]
[0111] where head i is the head module of the multi - head mechanism, which calculates the attention of multiple sub - spaces in parallel through the multi - head mechanism; CrossAttention is the cross - attention function inside each head module; softmax is an activation function that converts the output of a multi - classification model into a probability distribution; Q, K, and V correspond to the query, key, and value in the attention mechanism; Concat is a fusion function that performs weighted summation on the results of multiple head modules (head); W O is the linear transformation matrix of the multi - head cross - attention output, Z ∈ R 862×12×32 is the fused feature vector. When processing the block - level features during the weekday evening rush hour, the attention mechanism will focus on the key time points within the same type of blocks in the historical data, such as the specific time periods when congestion occurred in the past, so as to integrate these detailed information into the block - level feature representation and avoid the loss of details caused by the averaging process in the traditional block - splitting method. After multi - head attention calculation and feature concatenation, the fused feature vector is obtained, which combines both global trends and local detailed information.
[0112] Step 3, micro - training prediction:
[0113] First, flatten the fused features to convert the three - dimensional tensor into a two - dimensional vector for easy processing by the fully - connected layer. Second, the flattened feature vector is mapped into a result vector through the linear transformation layer. The length of the flattened feature vector is 12 * 32, and the length of the result vector is 96, and Dropout regularization is introduced to prevent overfitting. Finally, the results after micro - training are inverse - normalized using the mean and standard deviation of each channel recorded during the training phase to obtain the final multi - channel time - series prediction results (such as the multi - channel traffic flows such as vehicle flow and vehicle speed at each monitoring point in the future time period). The formula is as follows:
[0114] F out = Denormalization(F)
[0115] F=(Dropout(Linear(Flatten(Z)))
[0116] Among them, F is the output of the micro-training prediction module, which aims to perform low-complexity training on the result Z of point-block feature fusion to prevent noise interference after feature fusion. It includes Flatten processing and a Linear layer. Flatten converts the three-dimensional tensor into a two-dimensional feature vector for easy processing by the fully connected layer. The Linear layer maps the flattened fused feature vector to the prediction result vector. During this process, the length of the vector changes from 12 * 32 to 96, and 96 corresponds to the length required for each channel's prediction task. Dropout regularization is introduced to prevent overfitting. Denormalization is the inverse normalization process, which restores the data to its normal size and outputs the final prediction result F out ∈R 862×96 。
[0117] In summary, this embodiment proposes a traffic flow time series prediction system based on point-block cross-attention. For the same input of urban traffic data, it realizes differential processing and deep fusion through a parallel dual-view encoding structure: uses a convolutional network to extract local detailed features of each timestamp to capture short-term traffic fluctuations, and at the same time divides the sequence into overlapping sub-blocks through sliding window partitioning and Transformer encoding to extract cross-block long-term trends. The cross-attention mechanism uses block-level features as queries and point-level features as keys and values, dynamically calculates attention weights to accurately focus on key timestamp details during block-level trend modeling, breaking through the limitation of detail loss in traditional partitioning methods. The fused multi-scale features are processed by the prediction head to output interpretable physical quantity prediction results, directly supporting decision-making such as signal timing optimization and emergency lane scheduling. The system reduces the computational complexity through a sliding window and improves robustness by combining regularization techniques, adapting to real-time processing of urban-level large-scale traffic data. Its "same data dual-view processing - cross-granularity fusion" framework provides a general solution for long-range time series analysis in multiple fields
[0118] The described embodiment is a preferred embodiment of the present invention, but the present invention is not limited to the above embodiment. Without departing from the essential content of the present invention, any obvious improvement, replacement, or variation that those skilled in the art can make belongs to the protection scope of the present invention
Claims
1. A traffic flow time series prediction system based on dot-block cross-attention, characterized in that: Step 1, preprocess the single-modal multi-channel time series data collected by urban traffic sensors; Step 2, use a dot encoder to extract features from the single-channel data at each timestamp in the single-modal multi-channel time series data preprocessed in Step 1 to obtain dot-level feature vectors; Step 3, use a block encoder to perform block processing on the single-modal multi-channel time series data preprocessed in Step 1 to extract block-level feature vectors; Step 4, use a multi-head cross-attention mechanism, take the block-level feature vectors as queries, and the dot-level feature vectors as both keys and values at the same time to obtain fused feature vectors; Step 5: Input the fused feature vectors into a micro-training prediction module, and output multi-channel time series prediction results after inverse normalization processing.
2. The traffic flow time series prediction system based on dot-block cross-attention according to claim 1, wherein The dot encoder includes two convolutional layers and a ReLU activation function.
3. The traffic flow time series prediction system based on dot-block cross-attention according to claim 2, characterized in that, The dot-level feature vector is represented by the formula: point encoder(H) = conv2(ReLU(conv1(H))) Among them, is the point-level feature vector, H ∈ R C×S is the preprocessed single-modal multi-channel time series data, point encoder is the point encoding function, conv1 and conv2 are convolutional layers, ReLU is the activation function, C is the number of channels, S is the sequence length, D is the feature dimension, and W is the format conversion matrix.
4. The traffic flow time series prediction system based on dot-block cross-attention according to claim 1, characterized in that, The steps for the block encoder to extract block-level feature vectors include dividing the single-modal multi-channel time series data into blocks through a sliding window, linear projection and position embedding, and Transformer encoding.
5. The traffic flow time series prediction system based on dot-block cross-attention according to claim 4, wherein The block-level feature vector is represented by the formula: H p = Patch(H) Among them, Patch is the chunking function, and H ∈ R C×S is the preprocessed single-modal multi-channel time series data, and H p ∈ R C×P×L is the chunking result. Dropout is a regularization technique, and Linear is a linear layer. W pos is the positional embedding, is the output of the long-term features. LN is layer normalization, and MHSA is the multi-head self-attention mechanism, is the block-level feature vector. C is the number of channels, S is the sequence length, P is the number of chunks, L is the chunk dimension, and D is the feature dimension.
6. The traffic flow time series prediction system based on dot-block cross-attention according to claim 1, characterized in that The fused feature vector is represented by the formula: Z = MHCA(Q, K, V) = Concat(head1, … head i ,...head h )W O Among them, W1, W2, and W3 are weight matrices, b1, b2, and b3 are bias vectors, Q, K, and V are the query, key, and value respectively, is the block-level feature vector, is the point-level feature vector, Z ∈ R C×P×D is the fused feature vector, MHCA is the multi-head cross-attention mechanism, Concat is the fusion function, head i is the head module of the multi-head mechanism, W O is the linear transformation matrix of the multi-head cross-attention output, CrossAttention is the cross-attention function inside each head module, softmax is an activation function, K T is the transpose matrix of K, is the scaling factor, C is the number of channels, P is the number of blocks, and D is the feature dimension.
7. The traffic flow time series prediction system based on dot-block cross-attention according to claim 6, characterized in that The micro-training prediction module is represented by the formula: Among them, is the flattened fused feature vector, F ∈ R C×Y is the output of the micro-training prediction module. Dropout is a regularization technique, and Linear is a linear layer. W f is the weight matrix, b f is the bias vector, and Y is the length of the prediction task for each channel.
8. The traffic flow time series prediction system based on dot-block cross-attention according to claim 7, characterized in that The inverse normalization is represented by the formula: where, F out ∈R C×Y is the multi-channel time series prediction result, Denormlization is the anti-normalization function, F i,j is the time series timestamp value in the output F of the micro-training prediction module, μ i and σ i are the mean and standard deviation of the data of the i-th channel respectively, and i and j correspond to the i-th channel and the j-th timestamp value respectively.