A Traffic Speed ​​Prediction Method Based on Multi-Spatial-Scale Spatiotemporal Transformer

By using a multi-scale spatiotemporal Transformer model, the problems of complex spatial structures and strict temporal-spatial dependencies in transportation systems are solved, enabling accurate prediction and rapid modeling of traffic speeds.

CN116311921BActive Publication Date: 2025-10-31CHINA UNIV OF MINING & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310182427.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-01
Publication Date
2025-10-31
Estimated Expiration
2043-03-01

AI Technical Summary

Technical Problem

Existing traffic data prediction methods are unable to accurately and quickly extract spatial structural features and spatiotemporal dependencies when processing traffic systems, resulting in inaccurate traffic speed predictions.

Method used

A multi-scale spatiotemporal Transformer model is adopted. Through a multi-scale spatial feature extraction module, a traffic spatiotemporal feature extraction module, and a prediction module, traffic dynamic spatial structure features are extracted from the node, region, and road levels, respectively, and prediction is performed in combination with spatiotemporal dependencies.

Benefits of technology

It achieves accurate modeling and rapid prediction of traffic speed, reduces the computation of useless information, improves prediction accuracy, and solves the problems of noise and information redundancy in existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311921B_ABST
    Figure CN116311921B_ABST
Patent Text Reader

Abstract

This invention discloses a traffic speed prediction method based on a multi-scale spatiotemporal Transformer, belonging to the field of traffic prediction and planning technology. The prediction method includes: sequentially inputting preprocessed road segment sensor speed sequence data into a multi-scale spatial feature extraction module, a traffic spatiotemporal feature extraction module, and a prediction module, gradually realizing feature extraction of multi-scale dynamic spatial structures and static road network structures, accurate modeling of spatiotemporal dependencies, and prediction of traffic speeds over a future period. The multi-scale spatial feature extraction module of this invention can comprehensively and specifically extract spatial features, improving prediction accuracy while reducing a large amount of useless computation. Furthermore, the traffic spatiotemporal feature extraction module selects more valuable historical data based on traffic characteristics and the relative position information of the data for sufficient spatiotemporal feature extraction, solving the problem of losing relative position information when extracting spatiotemporal dependencies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of traffic forecasting and planning technology, and in particular to a traffic speed prediction method based on a multi-spatial-scale spatiotemporal Transformer. Background Technology

[0002] Transportation systems are among the most critical infrastructure components of modern cities, supporting the daily commutes and travel of millions of people. With urbanization and population growth, transportation systems become increasingly complex. Early intervention based on traffic forecasting is considered key to improving transportation system efficiency and mitigating traffic-related problems.

[0003] Current traffic data prediction methods model spatial structure at the scale of the entire traffic network, essentially placing numerous sensor nodes onto a single traffic map for feature extraction. Some models construct spatial graphs based on the distances between sensor nodes and directly apply graph convolutions to extract meaningful patterns and features in the spatial domain. While this approach of treating all sensors as interconnected can effectively extract spatial relationships, it leads to over-extraction, introducing more noise and a large amount of useless information. Some models define the connectivity between sensors using distance thresholds, modeling the sensor network as a weighted directed graph and proposing diffusing convolutions to capture spatial dependencies. However, this method still operates at the scale of the entire traffic network, which is not conducive to accurate and targeted extraction of spatial structural features. With the rapid development of Transformers, some models use spatial Transformers to extract global spatial structural features and calculate the dependencies between each sensor. While achieving good results, this approach still suffers from problems such as a single scale and the computation of a large amount of useless information when extracting spatial structure. Furthermore, existing models do not consider the temporal and spatial relationships and relative location information of traffic speed data when modeling traffic speed data. However, this is crucial for traffic speed prediction because traffic speed data is time-series data, and the influence between each time step is different. For example, the influence of the previous time step on a given time step is greater than the influence of the two time steps before that. In addition, the current time step is not affected by subsequent time steps. Therefore, how to perform accurate spatiotemporal modeling is a major problem that needs to be solved. Summary of the Invention

[0004] Technical Problem: The purpose of this invention is to overcome the shortcomings of the prior art and provide a traffic speed prediction method based on multi-spatial-scale spatiotemporal Transformer, so as to solve the problem that existing traffic data prediction cannot accurately and quickly predict traffic data due to the complex spatial structure of the traffic system and the strict order of spatiotemporal dependencies.

[0005] Technical Solution: This invention relates to a traffic speed prediction method based on a multi-scale spatiotemporal Transformer. It utilizes urban traffic speed data to design a prediction model, achieving comprehensive and targeted extraction of spatial features, accurate modeling of spatiotemporal dependent features, and prediction of traffic speed over a future period. The prediction model includes a multi-scale spatial feature extraction module, a traffic spatiotemporal feature extraction module, and a prediction module; it includes the following steps:

[0006] Step 1: Preprocess the acquired road segment sensor speed sequence data: This includes processing sensor node data and generating a sample set to obtain the preprocessed speed sample set and the weighted adjacency matrix of the road network.

[0007] Step 2: Input the data processed in Step 1 into the multi-scale spatial feature extraction module to obtain speed data that integrates multi-scale dynamic spatial structure and static road network structure features;

[0008] Step 3: The speed data with extracted spatial structural features obtained in Step 2 is processed by the traffic spatiotemporal feature extraction module to construct spatiotemporal dependencies, thereby obtaining speed data with accurate spatiotemporal dependencies;

[0009] Step 4: Take the velocity data X with precise spatiotemporal dependence obtained in Step 3. ST The input prediction module performs multi-step predictions to forecast traffic speeds over a future period; simultaneously, a loss function is used to train the traffic speed prediction model, progressively training and optimizing parameters to achieve accurate predictions of urban traffic speeds.

[0010] The process involves processing sensor node data and generating a sample set;

[0011] The sensor node data refers to the average vehicle speed information over a period of time obtained from road sensors. The method for processing the sensor node data is as follows: the sensor data is aggregated every 5 minutes, missing values ​​are filled using linear interpolation, and finally the sensor data with missing values ​​filled is normalized using the z-score method to obtain the traffic dataset. The linear interpolation method is a method of using a straight line connecting two known quantities to determine the value of an unknown quantity between the two known quantities. The z-score method is a process of dividing the difference between a measured value and the mean by the standard deviation. The z-score method can transform data of different magnitudes into a uniform z-score score.

[0012] The method for generating the sample set is as follows:

[0013] Define a sliding window of length l with a step size of 1; make the sliding window move across the dataset [x1,…,x...]. T Slide the slider up to get the set H = [X1, ..., X...] of all data samples. h,…,X T-l+1 ],in

[0014] All data samples are subjected to feature extraction and prediction in sequence. The feature extraction and prediction processes are consistent, and both go through the multi-scale spatial feature extraction module, the traffic spatiotemporal feature extraction module, and the prediction module in sequence.

[0015] The traffic dataset includes speed data and a weighted adjacency matrix determined by the distances between sensor nodes. The speed data in the traffic dataset is time-series data, represented as follows: in, Let G represent the observations of N sensor nodes at time step t. These observations are represented as a traffic graph G = (V, E, W), where V represents the set of sensor nodes, |V| = N; and E represents the set of edges. Let W represent the weighted adjacency matrix of the traffic graph G; where the adjacency matrix and edge weights are determined by the distance between the locations of the sensors, and the edge weight matrix W is an adjacency matrix constructed based on connectivity. For sensor i and sensor j, w... ij =d ij Among them, w ij d represents the weight between sensor i and sensor j. ij It is the distance between sensor i and sensor j.

[0016] In step 2, the multi-scale spatial feature extraction module includes a node feature extraction layer, a region feature extraction layer, a road feature extraction layer, a static road network feature extraction layer, and a fusion layer. The speed data samples obtained in step 1 are input into each feature extraction layer to extract dynamic spatial structure features and static road network structure features at three scales: node level, region level, and road level. Then, the dynamic features and static feature data at the three scales are input into the fusion layer for fusion to obtain speed data that integrates multi-scale dynamic spatial structure and static road network structure features. The multi-scale spatial feature extraction module can comprehensively and specifically extract spatial features, improving prediction accuracy while reducing the computation of a large amount of useless information.

[0017] The extraction process of the multi-scale spatial feature extraction module is as follows:

[0018] First, take the velocity sample X h For example, let's take the velocity sample X. h First, a 1×1 convolutional layer is used to expand the number of feature channels, resulting in data with expanded feature channels.

[0019]

[0020] Where: Conv represents the convolution operation;

[0021] Will Parallel input sensor node feature extraction layer, region feature extraction layer, road feature extraction layer, and static feature extraction layer are used to obtain node-level features S. node Regional-level characteristics S area Features of the road layer S road and static road network structure characteristics S static The aforementioned features are then input into a fusion layer to fuse the dynamic and static features across the three scales, resulting in speed data X that integrates multi-scale dynamic spatial structure and static road network structure features. S :

[0022] X S =Fusion(S node ,S area ,S road ,S static )

[0023] Wherein: Fusion represents the fusion operation of the fusion layer.

[0024] The node feature extraction layer has its own unique traffic features for each sensor node, and does not need to aggregate the features of other sensor nodes. Therefore, the node feature extraction layer only extracts features from the original input features.

[0025] First, the velocity samples after expanding the number of feature channels. Perform layer normalization (LN) to ensure the stability of features in the data;

[0026] Secondly, a feedforward neural network is used to extract nonlinear features. The feedforward neural network consists of two linear layers and a nonlinear activation function.

[0027] Finally, to prevent gradient vanishing, a layer normalization operation with residual connections was added after feature extraction to obtain the feature values ​​S at the sensor node level. node The extraction process is as follows:

[0028]

[0029] Here, LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; and ReLU is a non-linear activation, used to learn the non-linear characteristics of the data.

[0030] The region feature extraction layer includes a region location embedding unit, a region multi-head self-attention unit, and a feedforward neural network unit;

[0031] First, a learnable spatial location embedding matrix is ​​used to learn the dynamic positional relationships between nodes and incorporate them into the original data;

[0032] Secondly, the data input region after location embedding is used to learn features from multi-head self-attention units;

[0033] Finally, it passes through a feedforward neural network unit to extract deeper features;

[0034] The embedding process of the region location embedding unit is as follows:

[0035] Using a learnable spatial location embedding matrix To learn the dynamic positional relationships between nodes, R area Initialize as a weighted adjacency matrix W to obtain the data after position embedding.

[0036]

[0037] Here, F is a 1×1 convolutional layer used to incorporate dynamic positional information into the input data;

[0038] The feature extraction process for the multi-head self-attention unit in the region is as follows:

[0039] Used in regional multi-head self-attention units Each attention head learns different features, and then the results of each attention head are aggregated; within each attention head, the input data is processed... Spatial feature extraction is performed, where, Parallel computing, the feature extraction process is as follows:

[0040] First, train three latent subspaces for a sequence of N sensor nodes, including the query subspace Q. area Key space K area Sum subspace V area : in, They are Q area ,K area V area The learnable weight matrix;

[0041] Secondly, the attention scores between nodes are calculated. When calculating the regional attention scores, nodes are filtered, and only the attention scores within their respective regions are calculated. The filtering process is as follows:

[0042]

[0043] in, express Query the corresponding value in the subspace. express The corresponding value in the key subspace d represents the transpose of a matrix; k K represents area Dimensions Used to prevent gradient vanishing and problems with excessively large input values; B ij This indicates the selection variable, where B is selected when node j is within the region of node i. ij The value is 0, otherwise it is set to negative infinity:

[0044]

[0045] Where R i This represents the set of all other nodes within the region centered at node i, determined by a given distance threshold K.

[0046] Next, the obtained attention scores are mapped to the range [0,1] using the softmax activation function to ensure that their sum is 1 throughout the entire sequence. Then, they are multiplied and added with the corresponding value subspace to obtain the data M from which the region features have been extracted. area :

[0047]

[0048] Finally, layer normalization with residual connections is used to stabilize the output of the unit, yielding data M′ with extracted regional spatial features. area :

[0049]

[0050] Among them: LN represents the layer standardization operation, which is used to ensure data stability;

[0051] The feature extraction process of the feedforward neural network unit is as follows:

[0052] The feedforward neural network consists of two linear layers and a non-activation function. To prevent gradient vanishing, a normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in data S with extracted regional spatial features. area :

[0053] S area =LN(Linear(ReLU(Linear(M′) area )))+M′ area )

[0054] Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

[0055] The road feature extraction layer first processes the input data. After passing through the linear mapping layer, three subspaces different from those of the region extraction layer are projected, including the query subspace Q. road Key space K road Sum subspace V road :

[0056]

[0057] in: They are Q road ,K road V road The learnable weight matrix;

[0058] Secondly, the correlation scores between nodes are calculated, and sparse self-attention is used to extract spatial features; by Q road With K road First, the attention score set S is obtained by dot product calculation. road :

[0059]

[0060] Among them: (K) road ) T Representation matrix K road The transpose of the result; in the resulting set of attention scores S road The highest attention scores are selected and multiplied by their corresponding value subspaces to obtain the data M from which road features have been extracted. road :

[0061] M road =F top-p (S road V

[0062] Among them, F top-p This represents the filtering function used to select from S road In the middle, select the top-p attention scores according to their numerical values ​​and keep them at their original values, while setting all other attention scores to 0;

[0063] Then, a layer normalization operation with residual connections is used to stabilize the output of the unit, resulting in data M′ with extracted road spatial features. road :

[0064]

[0065] Among them: LN represents the layer standardization operation, which is used to ensure data stability;

[0066] Finally, the data M′ from which road spatial features have been extracted will be used... road A feedforward neural network unit is input to learn the nonlinear features of the data; the feedforward neural network unit consists of two linear layers and a nonlinear activation function; to prevent gradient vanishing, a layer normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in data S with extracted road spatial features. road :

[0067] S road =LN(Linear(ReLU(Linear(M′) road )))+M′ road )

[0068] Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

[0069] The static spatial feature extraction layer uses graph convolution operations to aggregate information from neighboring nodes to extract the static features S of the traffic network. static ;

[0070] The extraction process of the static spatial feature extraction layer is as follows:

[0071]

[0072] in, It is the input data. It is an adjacency matrix with additional self-connections; yes The degree matrix, W static σ is a trainable weight matrix, and σ is the activation function.

[0073] The fusion process of the fusion layer is as follows:

[0074] The fusion layer utilizes a gating mechanism to fuse dynamic and static spatial features across multiple spatial scales. First, a gate g is calculated based on the data to be fused. Then, the gate g is used to calculate a weighted approach to selectively process the input data. The data-calculated gate g is expressed as:

[0075] g = sigmoid(f node (S node )+f area (S area )+f road (S road )+fstatic (S static ))

[0076] Where: f node f area f road and f static They are S node S area S road and S static The linear function is converted into a one-dimensional vector; the gating mechanism uses a transformation gate and a carry gate to represent how much output is generated through the transformation input and carry output, respectively; the speed data X integrates the characteristics of multi-scale dynamic spatial structure and static road network structure. S Represented as:

[0077] X S =g(S node )+g(S area )+g(S road )+(1-g)(S static ).

[0078] In step 3, the traffic spatiotemporal feature extraction module includes a traffic time-location embedding layer, a multi-head self-attention layer, and a feedforward neural network layer; the traffic spatiotemporal feature extraction module first extracts the speed data X with extracted spatial features. S The system inputs a traffic time and location embedding layer to learn the chronological relationship between corresponding times in the data. Then, it passes through a multi-head self-attention layer to learn the spatiotemporal features in the data. Finally, it inputs a feedforward neural network layer to learn the non-linear dependencies between the data.

[0079] The embedding process of the traffic time and location embedding layer is as follows:

[0080] Since absolute position embedding methods lose some positional information, relative positional information is injected into the sequence, and a trainable parameter representing the relative position is added when calculating the attention score later.

[0081] Velocity data with extracted spatial features The data is of length l, and there are 2l-1 relative positional relationships between them. The relative positional relationship table RPR is represented as follows:

[0082] RPR=[-l+1,…,-2,-1,0,1,2,…,l-1];

[0083] in, and The relative positional relationship between them is a ij =ji∈RPR; then generate the corresponding weight matrices for each of the 2l-1 relative positional relationships, where a ijThe corresponding weight vector is W i-j Since traffic data is highly time-fluid, future times will not affect the data flow of previous time steps. Therefore, the weight vector representing future values ​​in RPR is set to 0, i.e., W. i-j =0, ij<0;

[0084] The feature extraction process of the multi-head self-attention layer is as follows:

[0085] Used in multi-head self-attention layer Each attention head is used to learn different features, and then the results of each attention head are aggregated; within each attention head, the input data X is processed. S Spatiotemporal feature extraction is performed. The spatiotemporal feature extraction process is as follows:

[0086] First, for Training query subspace Q T Key space K T Sum subspace V T : in, They are Q T ,K T V T The learnable weight matrix;

[0087] Secondly, the dependencies between computing nodes are determined by... The result obtained after dot product calculation is:

[0088]

[0089] in, Representation matrix transpose; d k K represents T Dimensions Used to prevent gradient vanishing and the problem of excessively large input values; W i-j It is a weight matrix representing the relative positions between sequences;

[0090] Next, multiply by the corresponding value subspace to obtain the data M from which time dependencies have been extracted. T :

[0091] M T =S T V T

[0092] Among them, S T Include This represents the attention score across all time steps in the sequence;

[0093] Finally, layer normalization with residual connections is used to stabilize the output of this unit, yielding the data M′ after the extraction of time-dependent features. T :

[0094] M′ T =LN(X S +M T )

[0095] Among them: LN represents the layer standardization operation, which is used to ensure data stability;

[0096] The feature extraction process of the feedforward neural network unit is as follows:

[0097] The feedforward neural network consists of two linear layers and a non-linear activation function. To prevent gradient vanishing, a normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in traffic speed data X with precise time dependence. ST :

[0098] X ST =LN(Linear(ReLU(Linear(M′) T )))+M′ T )

[0099] Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

[0100] In step 4, the prediction module consists of two classic convolutional layers. The first convolutional layer reduces the dimensionality of the time step, and the second convolutional layer reduces the dimensionality of the feature dimension; ultimately, the future T is obtained. τ Traffic data at each time step, T τ This represents the number of future time steps to predict; the predicted data Y is obtained:

[0101]

[0102] Where: Conv represents the convolution operation;

[0103] After prediction, the Huber loss function is used for optimization:

[0104]

[0105] The Huber loss function is a parameterized loss function used for regression problems. δ represents a parameter that adjusts robustness. When the prediction bias is less than δ, the squared error is used; when the prediction bias is greater than δ, the linear error is used. Y represents the predicted value, and Y represents the actual value.

[0106] Beneficial effects: Due to the adoption of the above technical solution, this invention utilizes three modules: multi-scale spatial feature extraction, traffic spatiotemporal feature extraction, and prediction. In the multi-scale spatial feature extraction module, spatial structure features of traffic dynamics are extracted from three scales: node level, road level, and region level. For the road level and region level extraction, corresponding filters are designed to model the spatial structure in a targeted manner. Furthermore, graph convolutional neural networks are used in the road level and region level modules to extract static road network features. Finally, a fusion layer is designed in this module to fuse spatial features from various scales. In the traffic spatiotemporal feature extraction module, the relative position information of speed data and the relative weight information due to the periodicity unique to traffic are taken into account to select and utilize more valuable historical data when extracting spatiotemporal dependencies. The prediction module performs multi-step prediction on the data from which spatiotemporal features have been extracted to predict traffic speed at a specified future time step. Compared with other traffic data prediction models, this invention has significant advantages. It not only solves the problem that existing traffic speed prediction methods cannot accurately and quickly predict traffic speed due to the complex spatial structure of the traffic system and the strict order of spatiotemporal dependencies, but also has the following advantages:

[0107] The multi-scale spatial feature extraction module can extract spatial features comprehensively and in a targeted manner, improving prediction accuracy while reducing a lot of useless computation. In addition, the traffic spatiotemporal feature extraction module selects more valuable historical data based on traffic characteristics and the relative position information of the data to perform full spatiotemporal feature extraction, solving the problem of losing relative position information when extracting spatiotemporal dependencies. Attached Figure Description

[0108] Figure 1 This is a flowchart of the traffic speed prediction method based on multi-spatial-scale spatiotemporal Transformer of the present invention.

[0109] Figure 2 This is a structural diagram of the traffic speed prediction method based on multi-spatial-scale spatiotemporal Transformer of the present invention. Detailed Implementation

[0110] The embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0111] This invention discloses a traffic speed prediction method based on a multi-scale spatiotemporal Transformer. It utilizes urban traffic speed data to design a prediction model, achieving comprehensive and targeted extraction of spatial features, accurate modeling of spatiotemporal dependent features, and prediction of traffic speed over a future period. The prediction model includes a multi-scale spatial feature extraction (MCS) module, a traffic time-space feature extraction (TTW) module, and a prediction module. The method comprises the following steps:

[0112] Step 1: First, preprocess the obtained road segment sensor speed sequence data: including processing sensor node data and generating a sample set to obtain the preprocessed speed sample set and the weighted adjacency matrix of the road network.

[0113] In step 1, the sensor node data is processed and a sample set is generated;

[0114] The sensor node data refers to the average vehicle speed information over a period of time obtained from road sensors. The method for processing the sensor node data is as follows: the sensor data is aggregated every 5 minutes, missing values ​​are filled using linear interpolation, and finally the sensor data with missing values ​​filled is normalized using the z-score method to obtain the traffic dataset. The linear interpolation method is a method of using a straight line connecting two known quantities to determine the value of an unknown quantity between the two known quantities. The z-score method is a process of dividing the difference between a measured value and the mean by the standard deviation. The z-score method can transform data of different magnitudes into a uniform z-score score.

[0115] The method for generating the sample set is as follows:

[0116] Define a sliding window of length l with a step size of 1; make the sliding window move across the dataset [x1,...,x...]. T Slide the slider up to get the set of all data samples H = [X1,...,X...]. h ,...,X T-l+1 ],in

[0117] All data samples are subjected to feature extraction and prediction in sequence. The feature extraction and prediction processes are consistent, and both go through the multi-scale spatial feature extraction module, the traffic spatiotemporal feature extraction module, and the prediction module in sequence.

[0118] The traffic dataset includes speed data and a weighted adjacency matrix determined by the distances between sensor nodes. The speed data in the traffic dataset is time-series data, represented as follows: in, Let G represent the observations of N sensor nodes at time step t. These observations are represented as a traffic graph G = (V, E, W), where V represents the set of sensor nodes, |V| = N; and E represents the set of edges. Let W represent the weighted adjacency matrix of the traffic graph G; where the adjacency matrix and edge weights are determined by the distance between the locations of the sensors, and the edge weight matrix W is an adjacency matrix constructed based on connectivity. For sensor i and sensor j, w... ij =d ij Among them, w ij d represents the weight between sensor i and sensor j. ij It is the distance between sensor i and sensor j.

[0119] Step 2: Input the data processed in Step 1 into the multi-scale spatial feature extraction module to obtain speed data that integrates multi-scale dynamic spatial structure and static road network structure features;

[0120] In step 2, the multi-scale spatial feature extraction module includes a node feature extraction layer, a region feature extraction layer, a road feature extraction layer, a static road network feature extraction layer, and a fusion layer. The speed data samples obtained in step 1 are input into each feature extraction layer to extract dynamic spatial structure features and static road network structure features at three scales: node level, region level, and road level. Then, the dynamic features and static feature data at the three scales are input into the fusion layer for fusion to obtain speed data that integrates multi-scale dynamic spatial structure and static road network structure features. The multi-scale spatial feature extraction module can comprehensively and specifically extract spatial features, improving prediction accuracy while reducing the computation of a large amount of useless information.

[0121] The extraction process of the multi-scale spatial feature extraction module is as follows:

[0122] First, take the velocity sample X h For example, let's take the velocity sample X. h First, a 1×1 convolutional layer is used to expand the number of feature channels, resulting in data with expanded feature channels.

[0123]

[0124] Where: Conv represents the convolution operation;

[0125] Will Parallel input sensor node feature extraction layer, region feature extraction layer, road feature extraction layer, and static feature extraction layer are used to obtain node-level features S. node Regional-level characteristics S areaFeatures of the road layer S road and static road network structure characteristics S static The aforementioned features are then input into a fusion layer to fuse the dynamic and static features across the three scales, resulting in speed data X that integrates multi-scale dynamic spatial structure and static road network structure features. S :

[0126] X S =Fusion(S node ,S area ,S road ,S static )

[0127] Wherein: Fusion represents the fusion operation of the fusion layer.

[0128] The node feature extraction layer has its own unique traffic features for each sensor node, and does not need to aggregate the features of other sensor nodes. Therefore, the node feature extraction layer only extracts features from the original input features.

[0129] First, the velocity samples after expanding the number of feature channels. Perform layer normalization (LN) to ensure the stability of features in the data;

[0130] Secondly, a feedforward neural network is used to extract nonlinear features. The feedforward neural network consists of two linear layers and a nonlinear activation function.

[0131] Finally, to prevent gradient vanishing, a layer normalization operation with residual connections was added after feature extraction to obtain the feature values ​​S at the sensor node level. node The extraction process is as follows:

[0132]

[0133] Here, LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; and ReLU is a non-linear activation, used to learn the non-linear characteristics of the data.

[0134] The region feature extraction layer includes a region location embedding unit, a region multi-head self-attention unit, and a feedforward neural network unit;

[0135] First, a learnable spatial location embedding matrix is ​​used to learn the dynamic positional relationships between nodes and incorporate them into the original data;

[0136] Secondly, the data input region after location embedding is used to learn features from multi-head self-attention units;

[0137] Finally, it passes through a feedforward neural network unit to extract deeper features;

[0138] The embedding process of the region location embedding unit is as follows:

[0139] Using a learnable spatial location embedding matrix To learn the dynamic positional relationships between nodes, R area Initialize as a weighted adjacency matrix W to obtain the data after position embedding.

[0140]

[0141] Here, F is a 1×1 convolutional layer used to incorporate dynamic positional information into the input data;

[0142] The feature extraction process for the multi-head self-attention unit in the region is as follows:

[0143] Used in regional multi-head self-attention units Each attention head learns different features, and then the results of each attention head are aggregated; within each attention head, the input data is processed... Spatial feature extraction is performed, where, Parallel computing, the feature extraction process is as follows:

[0144] First, train three latent subspaces for a sequence of N sensor nodes, including the query subspace Q. area Key space K area Sum subspace V area : in, They are Q area ,K area V area The learnable weight matrix;

[0145] Secondly, the attention scores between nodes are calculated. When calculating the regional attention scores, nodes are filtered, and only the attention scores within their respective regions are calculated. The filtering process is as follows:

[0146]

[0147] in, express Query the corresponding value in the subspace. express The corresponding value in the key subspace d represents the transpose of a matrix; k K represents area Dimensions Used to prevent gradient vanishing and problems with excessively large input values; B ij This indicates the selection variable, where B is selected when node j is within the region of node i. ij The value is 0, otherwise it is set to negative infinity:

[0148]

[0149] Where R i This represents the set of all other nodes within the region centered at node i, determined by a given distance threshold K.

[0150] Next, the obtained attention scores are mapped to the range [0,1] using the softmax activation function to ensure that their sum is 1 throughout the entire sequence. Then, they are multiplied and added with the corresponding value subspace to obtain the data M from which the region features have been extracted. area :

[0151]

[0152] Finally, layer normalization with residual connections is used to stabilize the output of the unit, yielding data M′ with extracted regional spatial features. area :

[0153]

[0154] Among them: LN represents the layer standardization operation, which is used to ensure data stability;

[0155] The feature extraction process of the feedforward neural network unit is as follows:

[0156] The feedforward neural network consists of two linear layers and a non-activation function. To prevent gradient vanishing, a normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in data S with extracted regional spatial features. area :

[0157] S area =LN(Linear(ReLU(Linear(M′) area )))+M′ area )

[0158] Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

[0159] The road feature extraction layer first processes the input data. After passing through the linear mapping layer, three subspaces different from those of the region extraction layer are projected, including the query subspace Q. road Key space K road Sum subspace V road :

[0160]

[0161] in: They are Q road ,K road V road The learnable weight matrix;

[0162] Secondly, the correlation scores between nodes are calculated, and sparse self-attention is used to extract spatial features; by Q road With K road First, the attention score set S is obtained by dot product calculation. road :

[0163]

[0164] Among them: (K) road ) T Representation matrix K road The transpose of the result; in the resulting set of attention scores S road The highest attention scores are selected and multiplied by their corresponding value subspaces to obtain the data M from which road features have been extracted. road :

[0165] M road =F top-p (S road V

[0166] Among them, F top-p This represents the filtering function used to select from S road In the middle, select the top-p attention scores according to their numerical values ​​and keep them at their original values, while setting all other attention scores to 0;

[0167] Then, a layer normalization operation with residual connections is used to stabilize the output of the unit, resulting in data M′ with extracted road spatial features. road :

[0168]

[0169] Among them: LN represents the layer standardization operation, which is used to ensure data stability;

[0170] Finally, the data M′ from which road spatial features have been extracted will be used... roadA feedforward neural network unit is input to learn the nonlinear features of the data; the feedforward neural network unit consists of two linear layers and a nonlinear activation function; to prevent gradient vanishing, a layer normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in data S with extracted road spatial features. road :

[0171] S road =LN(Linear(ReLU(Linear(M′) road )))+M′ road )

[0172] Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

[0173] The static spatial feature extraction layer uses graph convolution operations to aggregate information from neighboring nodes to extract the static features S of the traffic network. static ;

[0174] The extraction process of the static spatial feature extraction layer is as follows:

[0175]

[0176] in, It is the input data. It is an adjacency matrix with additional self-connections; yes The degree matrix, W static σ is a trainable weight matrix, and σ is the activation function.

[0177] The fusion process of the fusion layer is as follows:

[0178] The fusion layer utilizes a gating mechanism to fuse dynamic and static spatial features across multiple spatial scales. First, a gate g is calculated based on the data to be fused. Then, the gate g is used to calculate a weighted approach to selectively process the input data. The data-calculated gate g is expressed as:

[0179] g = sigmoid(f node (S node )+f area (S area )+f road (S road )+f static (S static ))

[0180] Where: f node farea f road and f static They are S node S area S road and S static The linear function is converted into a one-dimensional vector; the gating mechanism uses a transformation gate and a carry gate to represent how much output is generated through the transformation input and carry output, respectively; the speed data X integrates the characteristics of multi-scale dynamic spatial structure and static road network structure. S Represented as:

[0181] X S =g(S node )+g(S area )+g(S road )+(1-g)(S static ).

[0182] Step 3: The speed data with extracted spatial structural features obtained in Step 2 is processed by the traffic spatiotemporal feature extraction module to construct spatiotemporal dependencies, thereby obtaining speed data with accurate spatiotemporal dependencies;

[0183] In step 3, the traffic spatiotemporal feature extraction module includes a traffic time-location embedding layer, a multi-head self-attention layer, and a feedforward neural network layer; the traffic spatiotemporal feature extraction module first extracts the speed data X with extracted spatial features. S The system inputs a traffic time and location embedding layer to learn the chronological relationship between corresponding times in the data. Then, it passes through a multi-head self-attention layer to learn the spatiotemporal features in the data. Finally, it inputs a feedforward neural network layer to learn the non-linear dependencies between the data.

[0184] The embedding process of the traffic time and location embedding layer is as follows:

[0185] Since absolute position embedding methods lose some positional information, relative positional information is injected into the sequence, and a trainable parameter representing the relative position is added when calculating the attention score later.

[0186] Velocity data with extracted spatial features The data is of length l, and there are 2l-1 relative positional relationships between them. The relative positional relationship table RPR is represented as follows:

[0187] RPR=[-l+1,…,-2,-1,0,1,2,…,l-1];

[0188] in, and The relative positional relationship between them is a ij =ji∈RPR; then generate the corresponding weight matrices for each of the 2l-1 relative positional relationships, where aij The corresponding weight vector is W i-j Since traffic data is highly time-fluid, future times will not affect the data flow of previous time steps. Therefore, the weight vector representing future values ​​in RPR is set to 0, i.e., W. i-j =0, ij<0;

[0189] The feature extraction process of the multi-head self-attention layer is as follows:

[0190] Used in multi-head self-attention layer Each attention head is used to learn different features, and then the results of each attention head are aggregated; within each attention head, the input data X is processed. S Spatiotemporal feature extraction is performed. The spatiotemporal feature extraction process is as follows:

[0191] First, for Training query subspace Q T Key space K T Sum subspace V T : in, They are Q T ,K T V T The learnable weight matrix;

[0192] Secondly, the dependencies between computing nodes are determined by... and The result obtained after dot product calculation is:

[0193]

[0194] in, Representation matrix transpose; d k K represents T Dimensions Used to prevent gradient vanishing and the problem of excessively large input values; W i-j It is a weight matrix representing the relative positions between sequences;

[0195] Next, multiply by the corresponding value subspace to obtain the data M from which time dependencies have been extracted. T :

[0196] M T =S T V T

[0197] Among them, S T Include This represents the attention score across all time steps in the sequence;

[0198] Finally, layer normalization with residual connections is used to stabilize the output of this unit, yielding the data M′ after the extraction of time-dependent features. T :

[0199] M′ T =LN(X S +M T )

[0200] Among them: LN represents the layer standardization operation, which is used to ensure data stability;

[0201] The feature extraction process of the feedforward neural network unit is as follows:

[0202] The feedforward neural network consists of two linear layers and a non-linear activation function. To prevent gradient vanishing, a normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in traffic speed data X with precise time dependence. ST :

[0203] X ST =LN(Linear(ReLU(Linear(M′) T )))+M′ T )

[0204] Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

[0205] Step 4: Take the velocity data X with precise spatiotemporal dependence obtained in Step 3. ST The input prediction module performs multi-step predictions to forecast traffic speeds over a future period; simultaneously, a loss function is used to train the traffic speed prediction model, progressively training and optimizing parameters to achieve accurate predictions of urban traffic speeds.

[0206] In step 4, the prediction module consists of two classic convolutional layers. The first convolutional layer reduces the dimensionality of the time step, and the second convolutional layer reduces the dimensionality of the feature dimension; ultimately, the future T is obtained. τ Traffic data at each time step, T τ This represents the number of future time steps to predict; the predicted data Y is obtained:

[0207]

[0208] Where: Conv represents the convolution operation;

[0209] After prediction, the Huber loss function is used for optimization:

[0210]

[0211] The Huber loss function is a parameterized loss function used for regression problems. δ represents a parameter that adjusts robustness. When the prediction bias is less than δ, the squared error is used; when the prediction bias is greater than δ, the linear error is used. Y represents the predicted value, and Y represents the actual value.

[0212] Application of a traffic speed prediction method based on multi-scale spatiotemporal Transformer: Validation was performed on a highway traffic dataset, PeMSD7(M), in a specific location. This dataset was collected in real-time every 30 seconds by the Caltrans Performance Measurement System (PeMS), covering weekdays in May and June of a given year. The traffic speed prediction steps are as follows:

[0213] Step 1: Preprocess the acquired road segment sensor speed sequence data: This includes processing sensor node data and generating a sample set to obtain the preprocessed speed sample set and the weighted adjacency matrix of the road network.

[0214] Step 2: Input the data processed in Step 1 into the multi-scale spatial feature extraction module to obtain speed data that integrates multi-scale dynamic spatial structure and static road network structure features;

[0215] Step 3: The speed data with extracted spatial structural features obtained in Step 2 is processed by the traffic spatiotemporal feature extraction module to construct spatiotemporal dependencies, thereby obtaining speed data with accurate spatiotemporal dependencies;

[0216] Step 4: Input the speed data with precise spatiotemporal dependence obtained in Step 3 into the prediction module for multi-step prediction to predict traffic speed in the future; at the same time, use the loss function to train the traffic speed prediction model, and gradually train and optimize the parameters to achieve accurate prediction of urban traffic speed.

[0217] Step 5: Experimental Environment and Hyperparameter Settings:

[0218] The deep learning framework used was PyTorch 1.10.0, and the programming language was Python 3.6. All experiments were conducted on a computer equipped with an NVIDIA Tesla P40 processor, with CUDA 10.2 and cuDNN 10.2 as the deep learning acceleration environment. The proposed model was trained for 50 epochs using the Adam optimizer with mean absolute error loss and a batch size of 32. The initial learning rate was 0.01, which was reduced at a rate of 0.7 every five epochs. The distance threshold K was set to 7000, and the prediction time step Tτ was set to 12.

Claims

1. A traffic speed prediction method based on a multi-spatial-scale spatiotemporal Transformer, characterized in that: A predictive model is designed using urban traffic speed data to achieve comprehensive and targeted extraction of spatial features, accurate modeling of spatiotemporal dependent features, and prediction of traffic speed over a future period. The predictive model includes a multi-scale spatial feature extraction module, a traffic spatiotemporal feature extraction module, and a prediction module; it includes the following steps: Step 1: Preprocess the acquired road segment sensor speed sequence data: This includes processing sensor node data and generating a sample set to obtain the preprocessed speed sample set and the weighted adjacency matrix of the road network. Step 2: Input the data processed in Step 1 into the multi-scale spatial feature extraction module to obtain speed data that integrates multi-scale dynamic spatial structure and static road network structure features; the multi-scale spatial feature extraction module includes a node feature extraction layer, a region feature extraction layer, a road feature extraction layer, a static road network feature extraction layer, and a fusion layer; input the speed data samples obtained in Step 1 into each feature extraction layer to extract dynamic spatial structure features and static road network structure features at three scales: node level, region level, and road level. Then, the dynamic features and static features at the three scales are input into the fusion layer for fusion to obtain speed data that integrates the features of multi-scale dynamic spatial structure and static road network structure. The multi-scale spatial feature extraction module can extract spatial features comprehensively and in a targeted manner, improving prediction accuracy while reducing the computation of a large amount of useless information. The extraction process of the multi-scale spatial feature extraction module is as follows: First, take the velocity sample X h For example, let's take the velocity sample X. h First, a 1×1 convolutional layer is used to expand the number of feature channels, resulting in data with expanded feature channels. Where: Conv represents the convolution operation; Will Parallel input sensor node feature extraction layer, region feature extraction layer, road feature extraction layer, and static feature extraction layer are used to obtain node-level features S. node Regional-level characteristics S area Features of the road layer S road and static road network structure characteristics S static The aforementioned features are then input into a fusion layer to fuse the dynamic and static features across the three scales, resulting in speed data X that integrates multi-scale dynamic spatial structure and static road network structure features. S : X S =Fusion(S node ,S area ,S road ,S static ) Where: Fusion represents the fusion operation of the fusion layer; The region feature extraction layer includes a region location embedding unit, a region multi-head self-attention unit, and a feedforward neural network unit; First, a learnable spatial location embedding matrix is ​​used to learn the dynamic positional relationships between nodes and incorporate them into the original data; Secondly, the data input region after location embedding is used to learn features from multi-head self-attention units; Finally, it passes through a feedforward neural network unit to extract deeper features; The embedding process of the region location embedding unit is as follows: Using a learnable spatial location embedding matrix To learn the dynamic positional relationships between nodes, R area Initialize as a weighted adjacency matrix W to obtain the data after position embedding. Here, F is a 1×1 convolutional layer used to incorporate dynamic positional information into the input data; The feature extraction process for the multi-head self-attention unit in the region is as follows: Used in regional multi-head self-attention units Each attention head learns different features, and then the results of each attention head are aggregated; within each attention head, the input data is processed... Spatial feature extraction is performed, where, Parallel computing, the feature extraction process is as follows: First, train three latent subspaces for a sequence of N sensor nodes, including the query subspace Q. area Key space K area Sum subspace V area : in, They are Q area ,K area V area The learnable weight matrix; Secondly, the attention scores between nodes are calculated. When calculating the regional attention scores, nodes are filtered, and only the attention scores within their respective regions are calculated. The filtering process is as follows: in, express Query the corresponding value in the subspace. express The corresponding value in the key subspace d represents the transpose of a matrix; k K represents area Dimensions Used to prevent gradient vanishing and problems with excessively large input values; B ij This indicates the selection variable, where B is selected when node j is within the region of node i. ij The value is 0, otherwise it is set to negative infinity: Where R i This represents the set of all other nodes within the region centered at node i, determined by a given distance threshold K. Next, the obtained attention scores are mapped to the range [0,1] using the softmax activation function to ensure that their sum is 1 throughout the entire sequence. Then, they are multiplied and added with the corresponding value subspace to obtain the data M from which the region features have been extracted. area : Finally, layer normalization with residual connections is used to stabilize the output of the unit, yielding data M′ with extracted regional spatial features. area : Among them: LN represents the layer standardization operation, which is used to ensure data stability; The feature extraction process of the feedforward neural network unit is as follows: The feedforward neural network consists of two linear layers and a non-activation function. To prevent gradient vanishing, a normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in data S with extracted regional spatial features. area : S area =LN(Linear(ReLU(Linear(M′ area )))+M′ area ) Where: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear characteristics of the data; Step 3: The speed data with extracted spatial structural features obtained in Step 2 is processed by the traffic spatiotemporal feature extraction module to construct spatiotemporal dependencies, resulting in speed data with precise spatiotemporal dependencies. The traffic spatiotemporal feature extraction module includes a traffic time-location embedding layer, a multi-head self-attention layer, and a feedforward neural network layer. The traffic spatiotemporal feature extraction module first processes the speed data X with extracted spatial features... S The system inputs a traffic time and location embedding layer to learn the chronological relationship between corresponding times in the data. Then, it passes through a multi-head self-attention layer to learn the spatiotemporal features in the data. Finally, it inputs a feedforward neural network layer to learn the non-linear dependencies between the data. The embedding process of the traffic time and location embedding layer is as follows: Since absolute position embedding methods lose some positional information, relative positional information is injected into the sequence, and a trainable parameter representing the relative position is added when calculating the attention score later. Velocity data with extracted spatial features The data is of length l, and there are 2l-1 relative positional relationships between them. The relative positional relationship table RPR is represented as follows: RPR=[-l+1,…,-2,-1,0,1,2,…,l-1]; in, and The relative positional relationship between them is a ij =ji∈RPR; then generate the corresponding weight matrices for each of the 2l-1 relative positional relationships, where a ij The corresponding weight vector is W i-j Since traffic data is highly time-fluid, future times will not affect the data trend of previous time steps. Therefore, the weight vector representing future values ​​in RPR is set to 0, i.e., W. i-j =0, ij<0; The feature extraction process of the multi-head self-attention layer is as follows: Used in multi-head self-attention layer Each attention head is used to learn different features, and then the results of each attention head are aggregated; within each attention head, the input data X is processed. S Spatiotemporal feature extraction is performed. The spatiotemporal feature extraction process is as follows: First, for Training query subspace Q T Key space K T Sum subspace V T : in, They are Q T ,K T V T The learnable weight matrix; Secondly, the dependencies between computing nodes are determined by... and The result obtained after dot product calculation is: in, Representation matrix transpose; d k K represents T Dimensions Used to prevent gradient vanishing and problems caused by excessively large input values; W i-j It is a weight matrix representing the relative positions between sequences; Next, multiply by the corresponding value subspace to obtain the data M from which time dependencies have been extracted. T : M T =S T V T Among them, S T Include This represents the attention score across all time steps in the sequence; Finally, layer normalization with residual connections is used to stabilize the output of this unit, yielding the data M′ after the extraction of time-dependent features. T : M′ T =LN(X S +M T ) Among them: LN represents the layer standardization operation, which is used to ensure data stability; The feature extraction process of the feedforward neural network unit is as follows: The feedforward neural network consists of two linear layers and a non-linear activation function. To prevent gradient vanishing, a normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in traffic speed data X with precise time dependence. ST : X ST =LN(Linear(ReLU(Linear(M′ T )))+M′ T ) Where: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear characteristics of the data; Step 4: Take the velocity data X with precise spatiotemporal dependence obtained in Step 3. ST The input prediction module performs multi-step predictions to forecast traffic speeds over a future period; simultaneously, a loss function is used to train the traffic speed prediction model, progressively training and optimizing parameters to achieve accurate predictions of urban traffic speeds.

2. The traffic speed prediction method based on multi-spatial-scale spatiotemporal Transformer according to claim 1, characterized in that: In step 1, the sensor node data is processed and a sample set is generated; The sensor node data refers to the average vehicle speed information over a period of time obtained from road sensors. The method for processing the sensor node data is as follows: the sensor data is aggregated every 5 minutes, missing values ​​are filled using linear interpolation, and finally the sensor data with missing values ​​filled is normalized using the z-score method to obtain the traffic dataset. The linear interpolation method is a method of using a straight line connecting two known quantities to determine the value of an unknown quantity between the two known quantities. The z-score method is a process of dividing the difference between a measured value and the mean by the standard deviation. The z-score method can transform data of different magnitudes into a uniform z-score score. The method for generating the sample set is as follows: Define a sliding window of length l with a step size of 1; make the sliding window move across the dataset [x1,…,x...]. T Slide the slider up to get the set H = [X1, ..., X...] of all data samples. h ,…,X T-l+1 ],in All data samples are subjected to feature extraction and prediction in sequence. The feature extraction and prediction processes are consistent, and both go through the multi-scale spatial feature extraction module, the traffic spatiotemporal feature extraction module, and the prediction module in sequence.

3. The traffic speed prediction method based on multi-spatial-scale spatiotemporal Transformer according to claim 2, characterized in that: The traffic dataset includes speed data and a weighted adjacency matrix determined by the distances between sensor nodes. The speed data in the traffic dataset is time-series data, represented as follows: in, Let G represent the observations of N sensor nodes at time step t. These observations are represented as a traffic graph G = (V, E, W), where V represents the set of sensor nodes, |V| = N; and E represents the set of edges. Let W represent the weighted adjacency matrix of the traffic graph G; where the adjacency matrix and edge weights are determined by the distance between the locations of the sensors, and the edge weight matrix W is an adjacency matrix constructed based on connectivity. For sensor i and sensor j, w... ij =d ij Among them, w ij d represents the weight between sensor i and sensor j. ij It is the distance between sensor i and sensor j.

4. The traffic speed prediction method based on multi-spatial scale spatiotemporal Transformer according to claim 1, characterized in that: the node feature extraction layer has its own unique traffic features for each sensor node, and does not need to aggregate the features of other sensor nodes, so the node feature extraction layer only extracts features from the original input features; First, the velocity samples after expanding the number of feature channels. Perform layer normalization (LN) to ensure the stability of features in the data; Secondly, a feedforward neural network is used to extract nonlinear features. The feedforward neural network consists of two linear layers and a nonlinear activation function. Finally, to prevent gradient vanishing, a layer normalization operation with residual connections was added after feature extraction to obtain the feature values ​​S at the sensor node level. node The extraction process is as follows: in, LN stands for Layer Normalization, used to ensure data stability; Linear stands for Linear Layer, used to expand and reduce the dimensionality of data. ReLU is a non-linear activation used to learn the non-linear characteristics of data.

5. A traffic speed prediction method based on a multi-spatial-scale spatiotemporal Transformer according to claim 1, characterized in that: the road feature extraction layer first processes the input data... After passing through the linear mapping layer, three subspaces different from those of the region extraction layer are projected, including the query subspace Q. road Key space K road Sum subspace V road : in: They are Q road ,K road V road The learnable weight matrix; Secondly, the correlation scores between nodes are calculated, and sparse self-attention is used to extract spatial features; by Q road With K road First, the attention score set S is obtained by dot product calculation. road : Among them: (K) road ) T Representation matrix K road The transpose of the result; in the resulting set of attention scores S road The highest attention scores are selected and multiplied by their corresponding value subspaces to obtain the data M from which road features have been extracted. road : M road =F top-p (S road )V Among them, F top-p This represents the filtering function used to select from S road In the middle, select the top-p attention scores according to their numerical values ​​and keep them at their original values, while setting all other attention scores to 0; Then, a layer normalization operation with residual connections is used to stabilize the output of the unit, resulting in data M′ with extracted road spatial features. road : Among them: LN represents the layer standardization operation, which is used to ensure data stability; Finally, the data M′ from which road spatial features have been extracted will be used... road A feedforward neural network unit is input to learn the nonlinear features of the data; the feedforward neural network unit consists of two linear layers and a nonlinear activation function; to prevent gradient vanishing, a layer normalization operation with residual connections is added after feature extraction to stabilize the output, resulting in data S with extracted road spatial features. road : S road =LN(Linear(ReLU(Linear(M′ road )))+M′ road ) Wherein: LN represents layer normalization, used to ensure data stability; Linear represents a linear layer, used to expand and reduce the dimensionality of the data; ReLU is a non-linear activation function, used to learn the non-linear features of the data.

6. The traffic speed prediction method based on a multi-spatial-scale spatiotemporal Transformer according to claim 1, characterized in that: the static spatial feature extraction layer uses graph convolution operations to aggregate information from neighboring nodes to extract the static features S of the traffic network. static ; The extraction process of the static spatial feature extraction layer is as follows: in, It is the input data. It is an adjacency matrix with additional self-connections; yes The degree matrix, W static σ is a trainable weight matrix, and σ is the activation function. The fusion process of the fusion layer is as follows: The fusion layer utilizes a gating mechanism to fuse dynamic and static spatial features across multiple spatial scales. First, a gate g is calculated based on the data to be fused. Then, the gate g is used to calculate a weighted approach to selectively process the input data. The data-calculated gate g is expressed as: g=sigmoid(f node (S node )+f area (S area )+f road (S road )+f static (S static )) Where: f node f area f road and f static They are S node S area S road and S static The linear function is converted into a one-dimensional vector; the gating mechanism uses a transformation gate and a carry gate to represent how much output is generated through the transformation input and carry output, respectively; the speed data X integrates the characteristics of multi-scale dynamic spatial structure and static road network structure. S Represented as: X S =g(S node )+g(S area )+g(S road )+(1-g)(S static )。 7. The traffic speed prediction method based on multi-spatial-scale spatiotemporal Transformer according to claim 1, characterized in that: In step 4, the prediction module consists of two classic convolutional layers. The first convolutional layer reduces the dimensionality of the time step, and the second convolutional layer reduces the dimensionality of the feature dimension; ultimately, the future T is obtained. τ Traffic data at each time step, T τ This represents the number of future time steps to predict; (The result is...) Predicted value: Where: Conv represents the convolution operation; After prediction, the Huber loss function is used for optimization: The Huber loss function is a parameterized loss function used for regression problems. δ represents a parameter that adjusts robustness. When the prediction bias is less than δ, the squared error is used; when the prediction bias is greater than δ, the linear error is used. Y represents the predicted value, and Y represents the actual value.

Citation Information

Patent Citations

  • Traffic prediction method and device based on dynamic space-time diagram convolution attention model

    CN113487088A

  • Chinese named entity recognition method fusing time sequence convolution and Transform encoder

    CN114169330A

  • Pollution source-water quality prediction model weight influence calculation method of two-stage space-time attention mechanism

    CN114358435A

  • Construction method of traffic flow prediction model based on space-time multi-scale graph convolutional network and traffic flow prediction method

    CN114611383A

  • Traffic flow prediction method based on improved space-time Transform

    CN115273464A