Traffic speed prediction method based on geometric algebra and hypergraph
Patent Information
- Application Number
- CN202211370158.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-11-03
Smart Images

Figure CN115762183B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of intelligent transportation, and in particular to a traffic speed prediction method based on geometric algebra and hypergraph. Background Art
[0002] The transportation system is one of the most important infrastructures in modern cities, supporting the daily travel of millions of people. With rapid urbanization and population growth, the transportation system is becoming more complex. Modern transportation systems include road vehicles, rail transportation, and various shared travel modes that have emerged in recent years. Expanding cities face many traffic-related problems, including air pollution and traffic congestion. Early intervention based on traffic prediction can be the key to improving the efficiency of the transportation system and alleviating related problems such as traffic congestion.
[0003] Traffic speed prediction methods are mainly data-driven, and most of them are based on historical speed data. Traffic prediction problems are more challenging than other time series prediction problems because they involve high-dimensional large data volumes and a variety of dynamic characteristics including emergency situations (such as traffic accidents), which may lead to non-stationarity of traffic time series, making it difficult to make long-term predictions.
[0004] The traffic status at a specific location has both temporal and spatial dependencies. Traditional linear time series models, such as autoregression and integrated moving average models, cannot handle this spatiotemporal prediction problem well. In recent years, machine learning and deep learning techniques have been introduced into this field to improve prediction accuracy. For example, by modeling the entire city as a grid and applying convolutional neural networks. However, methods based on convolutional neural networks are not the optimal solution for traffic network structures with graphical forms.
[0005] In recent years, graph neural networks have become the forefront of deep learning research and have shown excellent performance in various applications. Graph neural networks are well suited for traffic prediction problems because they can capture spatial dependencies and use non-Euclidean graphs. Sensors in road networks generally have complex relationships, and even two sensors that are very close in Euclidean space may exhibit very different behaviors. Therefore, the road network is naturally a non-Euclidean graph with road intersections as nodes and road connections as edges. With graphs as input, several models based on graph neural networks have shown superior performance to previous methods in problems such as road traffic flow and speed prediction.
[0006] Although the most advanced models can extract spatiotemporal features from traffic data by combining temporal feature extraction methods and graph convolutional networks, they still have some shortcomings. From the perspective of temporal information extraction, methods based on recurrent neural networks or convolutional neural networks are often used. The former is prone to gradient vanishing or gradient exploding problems. If variants of recurrent neural networks, such as long short-term memory models, are used, problems such as excessive resource consumption, difficulty in training, and inability to handle a large number of longer sequence predictions will occur; although the latter solves the above-mentioned problems based on recurrent neural networks, the usual convolutional neural network not only ignores the internal dependencies between different time segments, but also fails to model the external dependencies between convolution kernels and time segments, and has limitations in mining and analyzing massive high-dimensional related traffic data. From the perspective of spatial information extraction, previous work often captures spatial features through a fixed graph structure. However, the traffic prediction problem is a time series problem, and the relationship between sensor nodes may change over time, and a fixed graph structure cannot reflect such changes. At the same time, most of the current models are based on traditional non-Euclidean graph structures, in which the edges can only connect two vertices. However, the relationship between road network sensors in traffic prediction problems is not entirely such a pairwise relationship. To a large extent, it involves the joint action of multiple nodes, which is a high-order relationship that cannot be captured by traditional graph structures.
[0007] Geometric Algebra is a covariate algebra framework generated in a unified pattern and is an extension of vector algebra. Geometric algebra introduces the concepts of multiple vectors and geometric products, allowing it to use higher-dimensional subspaces for operations. At the same time, it has a unified and efficient expression for information in high-dimensional space and the interaction between information, and can be extended to any high-dimensional space. The characteristics of geometric algebra make it suitable for encoding different time segments in traffic data, modeling the internal dependencies between time segments, and constructing convolution kernels based on convolution operations.
[0008] A hypergraph is a generalization of a graph. Unlike a simple graph where two nodes are connected by an edge, each hyperedge can connect any number of nodes in the hypergraph. The degree of hyperedges in a hypergraph can be higher than that of edges in a simple graph. Compared with graph structures that can only use pairwise connections, hypergraphs have significant advantages in modeling the correlation of real data. Summary of the invention
[0009] In view of the above-mentioned deficiencies in the existing field of traffic speed prediction, the present invention proposes a traffic speed prediction method based on geometric algebra and hypergraph based on the encoding method of high-dimensional data in the geometric algebra framework and the expression of dependency by rotation in the time dimension, and based on the hypergraph that can represent high-order interactions in the spatial dimension. By encoding time segments into multiple vectors and constructing a rotation matrix under the geometric algebra framework as a convolution kernel, the internal dependency between different time segments of the time series and the external dependency between the time series and the convolution kernel are learned to extract time features. A set of hyperedges constructed by applying a multi-level clustering method to the data and a set of hyperedges constructed based on the road network structure are combined to form a multi-dimensional high-order hypergraph structure. The hypergraph convolution is combined with the diffusion graph convolution on the traditional traffic map to extract higher-order spatial information. By conducting multi-dimensional in-depth mining of traffic speed data, long-term prediction of traffic speed is achieved, thereby improving the accuracy of traffic speed prediction.
[0010] The technical solution to be protected in the present invention is:
[0011] A traffic speed prediction method based on geometric algebra and hypergraph, the steps of the method are as follows:
[0012] Step 1. Input traffic speed data into the model, and use a linear layer to increase the dimension of the speed data so that the speed value is converted from a scalar to a vector. Then, the clusters are obtained through the K-means unsupervised clustering method. The pre-trained clustering results of the entire training set and the traffic road network diagram are combined to construct the hypergraph in the spatial feature extraction module.
[0013] Traffic speed data is input into the model, and the speed data is dimensionally upgraded through a linear layer, converting the speed value from a scalar to a vector, so that the model can extract richer information. The K-means unsupervised clustering method is used to obtain the clustering results for the input data. Combined with the clustering results obtained by pre-training the multi-layer K-means clustering method for the entire training set, the clustering results are used as a set of hyperedges in the hypergraph, and each node in the traffic road network structure is combined with its neighboring nodes to construct a set of hyperedges. The multi-level hypergraph constructed by the two sets of hyperedges includes both hyperedges based on traffic data and hyperedges based on road network structures, which will be used in the hypergraph convolution of the spatiotemporal extraction module.
[0014] Step 2. Construct K-layer spatiotemporal feature extraction modules. In each module, a gated geometric algebraic temporal convolutional network based on a geometric algebra framework is constructed to extract temporal features from traffic data, and a multidimensional graph convolutional network that combines diffuse graph convolution and multi-level hypergraph convolution is used to extract spatial features.
[0015] In the temporal information extraction module, a sliding time window is first used to encode different time slices into multi-vectors in geometric algebra, so that the model can model the internal dependencies and convolutions between time slices. The convolution kernel of the temporal convolutional network is constructed based on the description of multi-vector rotation in geometric algebra, so that the model can model the external dependencies between temporal data and convolution kernels, and better mine temporal features. In the spatial information extraction module, two types of graph convolutions at different levels are used. One type is a diffusion graph convolution based on an adaptive adjacency matrix and a traditional road network structure, which can model the spatial relationship between paired nodes; the other type is a hypergraph convolution based on the multi-level hypergraph constructed in step 1, which can model the spatial relationship between multiple nodes.
[0016] Step 3. The periodic information of which day of the week the traffic data belongs to and which time of the day it belongs to is embedded through two linear layers, combined with the spatiotemporal features extracted by each layer of modules, and then the linear layer is used to predict the future traffic speed from the spatiotemporal features of the current input data.
[0017] Traffic data contains strong periodicity, and the changes in traffic speed are strongly correlated with a certain day of the week and a certain time of the day. This periodicity can be embedded by constructing a network consisting of two linear layers based on the specific time information of the traffic data, and used as auxiliary data for predicting future traffic speeds. Combining the spatiotemporal features extracted by the multi-layer module can ensure that the model will not encounter gradient vanishing or gradient explosion. Combining the periodic features with the extracted recent spatiotemporal features, the linear layer is used to predict future traffic speeds.
[0018] Step 4. Use an optimized loss function that combines two commonly used loss functions, continuously optimize the network parameters through back propagation and gradient descent, minimize the loss function, and finally obtain the optimal model.
[0019] The mean absolute error (MAE) and the root mean square error (RMSE) are combined so that the time segments with congestion or mutation in the real traffic speed data are calculated using RMSE, while the other normal speed segments are calculated using MAE. While ensuring that the normal speed time segments are not affected by outliers, the congested time segments are better fitted. Finally, back propagation and gradient descent are performed to obtain the optimal model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 System flow chart of the traffic prediction method based on geometric algebra and hypergraph.
[0021] Figure 2 Application scenario diagram of the present invention.
[0022] Figure 3 The overall model structure diagram based on geometric algebra and hypergraph in the present invention.
[0023] Figure 4 An example of a hypergraph.
[0024] Figure 5 Visual representation of bidirectional quantities in geometric algebra.
[0025] Figure 6 Visual representation of reflection in geometric algebra.
[0026] Figure 7 Fitting curve of the predicted value and true value of the example METR-LA at node 0.
[0027] Figure 8 Fitting curve of the predicted value and the true value of the METR-LA example at node 22. DETAILED DESCRIPTION
[0028] In view of the above-mentioned deficiencies in the existing field of traffic speed prediction, the present invention proposes a traffic speed prediction method based on geometric algebra and hypergraph, based on the encoding method of high-dimensional data in the geometric algebra framework and the expression of dependency relationships by rotation in the time dimension, and based on the hypergraph that can represent high-order interactions in the spatial dimension. The method encodes time segments into multi-vectors and constructs a rotation matrix in the geometric algebra framework as a convolution kernel to learn the internal dependency between different time segments of the time series and the external dependency between the time series and the convolution kernel, thereby extracting deeper time features. A set of hyperedges constructed by applying a multi-level clustering method to the data and a set of hyperedges constructed based on the road network structure are combined to form a multi-dimensional high-order hypergraph structure. The hypergraph convolution is combined with the diffusion graph convolution on the traditional traffic map to extract higher-order spatial information. By conducting in-depth multi-dimensional mining of traffic speed data, long-term prediction of traffic speed is achieved, thereby improving the accuracy of traffic speed prediction. The specific method process and logical relationship are as follows: Figure 1 shown.
[0029] The application of this invention in actual scenarios can help traffic management departments better strengthen traffic demand management, strengthen comprehensive management of urban traffic congestion, make urban traffic smoother, and make people's travel experience more comfortable. Specific application scenarios such as Figure 2 shown.
[0030] The present invention proposes a traffic speed prediction method based on geometric algebra and hypergraph. The overall architecture of the method is as follows: Figure 3 As shown, the specific steps are as follows:
[0031] Step 1. Input the traffic speed data sampled by the road speed sensor into the model, and use a linear layer to increase the dimension of the speed data so that the speed value is converted from a scalar to a vector. Then, clustering is obtained through the K-means unsupervised clustering method. The pre-trained clustering results of the entire training set and the traffic road network diagram are combined to construct the hypergraph in the spatial feature extraction module.
[0032] Step 1.1 Traffic speed data is obtained by sampling speed sensors on the road at certain time intervals. Each speed sensor can be used as a node in a non-Euclidean graph or hypergraph. In the present invention, the traditional non-Euclidean graph structure and hypergraph structure are used to model the spatial dependency in traffic speed data. First, the speed data and the traffic road network structure are used to construct a hypergraph.
[0033] A hypergraph is a generalization of a graph. Unlike a simple graph in which two nodes are connected by an edge, each hyperedge can connect any number of nodes in the hypergraph. The degree of hyperedges in a hypergraph can be higher than that of edges in a simple graph. Compared with graph structures that can only use pairwise connections, hypergraphs have significant advantages in modeling the correlation of real data. An example of a hypergraph is Figure 4 shown.
[0034] A hypergraph is constructed based on the results of pre-training the K-means clustering method on the entire training set and the results of applying the K-means clustering method to the current input traffic speed data, by treating each cluster as a hyperedge.
[0035] K-means clustering is an unsupervised clustering method that can classify nodes with similar characteristics in traffic data into the same cluster. The hyperedge constructed by the K-means clustering method is a spatial modeling of nodes with high data similarity. The number of clusters is K, and through continuous iteration, the algorithm divides the data into different K groups. When pre-training the K-means clustering method for the entire training set, different clustering results can be obtained by setting multiple different numbers of clusters, and hyperedges of different granularities can be constructed.
[0036] At the same time, for each traffic speed sensor node, it is combined with its neighbor nodes in the traffic network structure into a hyperedge, and a set of hyperedges based on the traffic network structure is constructed. The two types of hyperedges together constitute a hypergraph, which includes hyperedges built based on data and hyperedges built based on the road network structure, and can perform multi-dimensional high-level modeling of spatial dependencies. This hypergraph will be used in the hypergraph convolution of the spatiotemporal feature extraction layer.
[0037] Step 1.2 Initialize the model parameter settings, which includes randomly initializing the two node embeddings E used to construct the adaptive adjacency matrix1 , E 2 The input traffic data is then passed through a fully connected layer to increase its dimension, converting the speed value from a scalar to a vector so that the model can extract richer information from the data.
[0038] Step 2. Construct K-layer spatiotemporal feature extraction modules. In each module, a gated geometric algebraic temporal convolutional network based on a geometric algebra framework is constructed to extract temporal features from traffic data, and a multidimensional graph convolutional network that combines diffuse graph convolution and multi-level hypergraph convolution is used to extract spatial features.
[0039] Step 2.1 uses a gated geometric algebraic temporal convolutional network based on a geometric algebra framework to extract temporal features from traffic data.
[0040] Geometric Algebra is a covariate algebra framework generated in a unified pattern and is an extension of vector algebra. Geometric algebra introduces the concepts of multiple vectors and geometric products, allowing it to use higher-dimensional subspaces for operations. At the same time, it has a unified and efficient expression for information in high-dimensional space and the interaction between information, and can be extended to any high-dimensional space. The characteristics of geometric algebra make it suitable for encoding different time segments in traffic data, modeling the internal dependencies between time segments, and constructing convolution kernels based on convolution operations.
[0041] Geometric algebra first introduced an operator called the outer product, which is represented by the ∧ symbol. Given two vectors a and b, the outer product a∧b results in a directional two-dimensional subspace, which is called a bivector in geometric algebra, such as Figure 5 shown.
[0042] A vector can be decomposed into a linear combination of basis vectors. Consider the Euclidean plane R 2 The two vectors a=(α 1 ,α 2 )=α 1 e 1 +α 2 e 2 , b=(β 1 ,β 2 )=β 1 e 1 +β 2 e 2 , the outer product of the two is:
[0043] a∧b=(α 1 e 1 +α 2 e 2 )∧(β 1 e1 +β 2 e 2 )
[0044] According to the distributive law and anti-commutative law of the outer product, and the law that the outer product of a vector and itself is zero, the above formula can be simplified to:
[0045] a∧b=(α 1 β 2 -α 2 β 1 ) 1 ∧e 2
[0046] In the Euclidean plane, we can use I = e 12 =e 1 ∧e 2 . In this way, geometric algebra also defines a set of bases in two-dimensional space, namely {1,e 1 ,e 2 ,e 12}. A vector can be decomposed into a combination of several basis vectors. In geometric algebra, this definition is extended to the definition of multivectors: multivectors are linear combinations of different bases. 2 In space, a multivector can be decomposed into a scalar part, a vector part, and a double vector part:
[0047] A=α 1 +α 2 e 1 +α 3 e 2 +α 4 I
[0048] α i are all real numbers used to represent the components of a multivector and can be zero. Multivectors, as linear combinations of subspaces, can be used to express many different concepts in geometry, and the above definition can be extended to higher-dimensional spaces. The geometric product operation is defined in geometric algebra, combining the outer product and the dot product. For any multivector, the geometric product is calculated as:
[0049]
[0050] In geometric algebra, each subspace can be simplified by creating a so-called basis multiplication table based on the calculation laws of outer products and dot products. For example, R 2 The multiplication table of the basis in the space is as follows:
[0051] 1 <![CDATA[e 1 ]]> <![CDATA[e 2 ]]> I 1 1 <![CDATA[e 1 ]]> <![CDATA[e 2 ]]> I <![CDATA[e 1 ]]> <![CDATA[e 1 ]]> 1 I <![CDATA[e 2 ]]> <![CDATA[e 2 ]]> <![CDATA[e 2 ]]> -I 1 <![CDATA[-e 1 ]]> I I <![CDATA[-e 2 ]]> <![CDATA[e 1 ]]> -1
[0052] Applying geometric multiplication in R2 Multiple vectors in space, according to R 2 The multiplication table of the basis in the space can obtain the following result:
[0053]
[0054] In geometric algebra, it is defined that a rotation can be realized by two reflections, and such a rotation can be applied in a space of any dimension by defining a rotation plane. The present invention uses rotation in three-dimensional space, so reflection and rotation in three-dimensional space will be briefly introduced.
[0055] First, reflection is the reflection of a vector to the other side of a double vector, which can also be simply understood as the other side of the Euclidean plane. Suppose we have a double vector U and its dual U * is the normal vector u of the Euclidean plane. If a vector a is geometrically multiplied by vector u and vector -u, and a is decomposed and simplified, the following result can be obtained (see also Figure 6 ):
[0056]
[0057] Rotation is based on such reflection. 2 Take the rotation in space as an example. Assume that s and t are two unit normal vectors. A rotation in geometric algebra is to combine two reflections. ′ =-t(-svs -1 )t -1 =tsvs -1 t -1 , since s and t are unit normal vectors, v ′ =tsvst, the geometric product symbol is omitted for convenience. Definition is the conjugate of R, so From the definition of geometric multiplication, we know that R consists of a scalar and a double vector. If we want to 2 To rotate vector v by angle θ in space, we need to define:
[0058]
[0059] The vector is rotated relative to the Euclidean plane A, that is, A is a double vector, and R is called R 2 A spinor in space. The proof here is more complicated, so only one conclusion is provided.
[0060] In higher-dimensional geometric algebra rotations, if you want to rotate any vector or multi-vector by an angle θ, you also need to define such a spinor R consisting of a scalar and a double vector. In the present invention, the rotation in three-dimensional space is used, and the spinor in three-dimensional space is defined as:
[0061]
[0062] The spinor R in three-dimensional space consists of four parts, a scalar and three double vectors (rotation planes) in three-dimensional space. The definition of rotation is still make:
[0063]
[0064] Rotations in three dimensions can be composed as a matrix multiplied by a vector:
[0065]
[0066] After entering the spatiotemporal feature extraction module, the present invention constructs the convolution kernel in the temporal convolution network based on the matrix expression of rotation in the above-mentioned geometric algebra to extract the time information in the traffic data after dimensionality enhancement.
[0067] First, we need to initialize the two convolution kernels of the gated temporal convolutional network. We randomly initialize an angle θ in [-π,π], and randomly sample and normalize the three matrices v in the average distribution of [0,1]. b , v c , v d , in order to calculate the four parts of the convolution kernel that constitute the temporal convolutional network:
[0068]
[0069] in is a randomly sampled matrix. The four initialized matrices can be concatenated into the geometric algebra matrix in R 3 A rotation matrix in space:
[0070]
[0071] This forms a matrix form of a rotation in geometric algebra, and is also the convolution kernel of the temporal convolutional network. According to the geometric algebra rotation formula, the rotation is completed by a spinor left multiplication and a spinor conjugate right multiplication. After being simplified to a matrix left multiplication form, this matrix is used as the convolution kernel for convolution, which can achieve the effect of two convolution operations with one convolution operation, and the convolution kernel can also be continuously optimized during the model training process.
[0072] The traffic data after dimensionality enhancement is divided into different time segments using a sliding window. In this model, a time window of size 4 is used. The four time segments are encoded as a multi-vector in a three-dimensional space in geometric algebra, and the data is convolved with two randomly initialized geometric algebra rotation convolution kernels, which is the matrix expression of rotation in three-dimensional space in geometric algebra. Finally, the extracted information is filtered using a gating mechanism. Encoding different time slices as multi-vectors in geometric algebra enables the model to learn the internal dependencies between different time segments in the time series. By constructing the convolution kernel as a matrix description of rotation in geometric algebra, the model can learn the external dependencies between the time series and the convolution kernel. The specific process of time feature extraction is described as follows:
[0073] h=g(W 1 *X+b)⊙σ(W 2 *X+c)
[0074] Among them, X is the traffic data after dimension enhancement input to the model spatiotemporal feature extraction module, W 1 , W 2 are two different convolution kernels, * is the temporal convolution operation, and ⊙ is the multiplication of the corresponding elements of the matrix. g(·) is the activation function of the output data, which is set to the tanh activation function in this model, and σ(·) is the sigmoid function used to determine how much proportion of information can pass to the next layer. The final h will be used as the input of the spatial feature extraction module.
[0075] Step 2.2 Extract spatial features using a multi-dimensional graph convolutional network based on diffuse graph convolution and multi-level hypergraph
[0076] By E initialized in step 1 1 , E 2 The node embedding calculates the adaptive adjacency matrix and combines it with the static adjacency matrix to perform bidirectional diffusion graph convolution. The specific process is as follows:
[0077]
[0078]
[0079] Among them, A adp is an adaptive adjacency matrix that can continuously optimize itself during the back propagation of the model. K is the number of steps of random diffusion, W k1 , W k2 , W k3 are all learnable weight matrices in the model. Since the traffic graph is a directed graph, the diffusion graph convolution is also bidirectional, P f , P pThey are the forward transfer probability matrix and the backward transfer probability matrix respectively. Both transfer probability matrices are calculated through the static adjacency matrix A. The specific calculation formula is:
[0080]
[0081] Among them, rowsum(·) is a function that calculates the sum of each row of the matrix.
[0082] The high-order spatial dependency information of the data is extracted through the multi-level hypergraph constructed in step 1. It is divided into two steps. The first step is to form the hyperedge information by aggregating the node information within the hyperedge. The second step is to update the node information itself by aggregating the hyperedge information connected to the node. The specific process is as follows:
[0083]
[0084] Among them, h e is the hyperedge hidden feature obtained by aggregating the feature values of nodes within the hyperedge, W e is the weight matrix learned in the lesson, d i is the degree of node i, that is, the number of hyperedges connected to node i, d e is the average degree of hyperedge e, that is, the average degree of each node in the hyperedge.
[0085] The results after diffuse graph convolution and hypergraph convolution are added together, and the data input to this layer is used as a residual connection to finally obtain the output of this spatiotemporal layer.
[0086] The transition probability graph used in the diffusion graph convolution belongs to the traditional graph structure, which can model the pairwise relationship between nodes, including the static traffic graph structure and the dynamic traffic graph structure obtained by node embedding calculation. The hypergraph used in the hypergraph convolution can model the common relationship between multiple nodes, including a set of static hyperedges obtained by pre-training the entire training set with multiple K-means clustering methods with different clustering numbers, a set of static hyperedges built based on the traditional traffic graph structure, and dynamic hyperedges for K-means clustering of each traffic speed data input into the model. Combining the traditional graph structure with the high-order graph structure, and combining static and dynamic characteristics, it fully considers the characteristics of the traffic network structure and can conduct a deeper exploration of the spatial characteristics in the traffic data.
[0087] Step 3. The periodic information of which day of the week the traffic data belongs to and which time of the day it belongs to is embedded through two linear layers, combined with the spatiotemporal features extracted by each layer of modules, and then the linear layer is used to predict the future traffic speed from the spatiotemporal features of the current input data.
[0088] The spatiotemporal features of the input traffic data are extracted through stacked spatiotemporal feature extraction modules, and the output of each layer is superimposed to ensure that the gradient disappearance or gradient explosion problem does not occur.
[0089] Traffic data contains strong periodicity, and the changes in traffic speed are strongly correlated with a certain day of the week and a certain time of the day. This periodic feature can be used as auxiliary data for predicting future traffic speeds. It is embedded through two linear layers and connected with the extracted spatiotemporal features of traffic data. The prediction results of the current input traffic speed data are obtained through multiple linear layers.
[0090] Step 4. Use an optimized loss function that combines two commonly used loss functions, continuously optimize the network parameters through back propagation and gradient descent, minimize the loss function, and finally obtain the optimal model.
[0091] In the loss function part, this model combines the mean absolute error (MAE) and the root mean square error (RMSE). By using RMSE to calculate the loss for the time segments with sudden changes or congestion in the actual traffic speed data, the difference between the predicted result and the actual traffic speed is amplified, while the loss is calculated using MAE for the rest of the normal speed segments. This ensures that the normal speed time segments will not be affected by abnormal values, and better fits the time segments with congestion or sudden changes. The optimized loss function is used to optimize the prediction accuracy of traffic speed. The optimized loss function formula is as follows:
[0092]
[0093]
[0094] Among them, α is the change threshold of sudden changes in the real traffic speed data, β is the speed threshold for identifying congestion in the real traffic speed, and T is the length of the predicted traffic speed time series.
[0095] Perform gradient calculation and back propagation, repeat model training and make the model gradually reach the optimal point.
[0096] To prove the effectiveness of the present invention, the method of the present invention is applied to the METR-LA dataset, and the more outstanding models based on graph neural networks in recent years are selected as comparative experiments. The METR-LA dataset records the traffic speed of 207 road speed sensors on Los Angeles highways within four months. Each speed sensor can be used as a node in a graph neural network or a hypergraph neural network. All models use the same preprocessing method, with a five-minute time window. The model input is the traffic speed of the previous hour, and the output is the predicted traffic speed of the next hour. In order to enable the model to cover the entire time series, the present invention adopts a 5-layer spatiotemporal feature extraction module. The present invention selects the comparison between the actual value and the predicted value of the traffic speed on two road speed sensors for visualization, and the results are as follows Figure 7 and Figure 8 shown.
[0097] Compared with other models, the present invention achieved the best results in MAE and RMSE indicators at 30 minutes and one hour, and the MAE at 15 minutes was only 0.01 behind the most advanced model. In terms of MAPE indicator, the present invention achieved results close to those of the most advanced model. The experimental results are shown in Table 1:
[0098] Table 1 Comparison of prediction effects between the present invention and the existing model on the METR-LA dataset
[0099]
[0100] References are as follows:
[0101] [1] Y. Li, R. Yu, C. Shahabi, and Y. Liu, "Diffusion convolutional recurrent neural network: Data-driven traffic forecasting," in Proc. of ICLR, 2018.
[0102] [2] Z. Wu, S. Pan, G. Long, J. Jiang, and C. Zhang, "Graph wavenet for deepspatial-temporal graph modeling," in Proc.of IJCAI, 2019.
[0103] [3] C.Zheng,
[0104] [4] Z. Wu, S. Pan, G. Long, J. Jiang, X. Chang, and C. Zhang, "Connecting thedots: Multivariate time series forecasting with graph neural networks," 082020, pp.753–763.
[0105] [5] F. Li, J. Feng, H. Yan H, et al. "Dynamic graph convolutional recurrent network for traffic prediction: Benchmark and solution." ACM Transactions on Knowledge Discovery from Data (TKDD), 2021.
[0106] Innovation
[0107] In the field of traffic speed prediction, in view of the limitations of existing methods in the expression of high-dimensional data objects, the extraction of internal dependencies between time segments, and the lack of external dependencies between convolution kernels in temporal feature extraction, the present invention combines geometric algebra with temporal convolutional networks to achieve the encoding and structured operation of high-dimensional information. At the same time, in view of the problem that existing methods in spatial feature extraction can only model the pairwise relationship between nodes in the traditional graph structure, but cannot model the common relationship between multiple nodes, the present invention combines the multi-level hypergraph with the traditional graph structure to achieve the extraction of high-order spatial information in the traffic map.
[0108] The present invention first performs multi-vector encoding and unified expression on traffic speed data according to the multi-vector definition in geometric algebra, extracts the correlation between different time segments in the data through the vector rotation expression in geometric algebra, and further explores the time dimension information in the traffic data while reducing a large number of parameters. And by combining the traditional graph structure with the high-order graph structure, the static and dynamic are combined, and the spatial dimension information in the data is extracted in a high-order manner through multi-level bidirectional diffusion graph convolution and hypergraph convolution, so as to fully explore the temporal and spatial correlation contained in the traffic data. By combining the output of each layer, the present invention ensures that the gradient will not disappear or explode during back propagation, and at the same time, considering the special periodicity in the traffic speed data, the periodic time information is embedded in the output layer. The present invention also combines two loss functions to make the final prediction result better fit the true value, and finally achieve the accuracy and performance improvement of traffic speed data prediction.
Claims
1. Traffic speed prediction method based on geometric algebra and hypergraph, It is characterized in that The specific method includes the following steps: Step 1. Input the traffic speed data into the model, and use a linear layer to increase the dimension of the speed data so that the speed value is converted from a scalar to a vector. Then, the clusters are obtained by using the K-means unsupervised clustering method. The hypergraph in the spatial feature extraction module is constructed by combining the pre-trained K-means clustering results of the entire training set and the traffic road network graph. Step 2. Construct K-layer spatiotemporal feature extraction modules; in each module, construct a gated geometric algebraic temporal convolutional network based on a geometric algebra framework to extract temporal features from traffic data, and use a multidimensional graph convolutional network that combines diffuse graph convolution and multi-level hypergraph convolution to extract spatial features; Step 3. The periodic information of which day of the week the traffic data belongs to and which time of the day it belongs to is embedded through two linear layers, combined with the spatiotemporal features extracted by each layer of modules, and then the linear layer is used to predict the future traffic speed from the spatiotemporal features of the current input data; Step 4. Use an optimized loss function that combines two commonly used loss functions to continuously optimize network parameters through back propagation and gradient descent to minimize the loss function and finally obtain the optimal model; In step 2, in the time information extraction module, the traffic data after dimensionality enhancement is divided into different time segments by means of a sliding window. In this model, a time window of size 4 is used, and the four time segments are encoded as a multi-vector in a three-dimensional space in geometric algebra, so that the model can model the internal dependency and convolution between the time segments; and the convolution kernel of the time convolution network is constructed based on the description of multi-vector rotation in geometric algebra, so that the model can model the external dependency between the time data and the convolution kernel; the specific process of time feature extraction is described as follows: h=g(W 1 *X+b)☉σ(W 2 *X+c) Among them, X is the traffic data after dimension enhancement input to the model spatiotemporal feature extraction module, W 1 , W 2 are two different convolution kernels, * is the temporal convolution operation, ⊙ is the multiplication of the corresponding elements of the matrix; g(·) is the activation function of the output data, which is set to the tanh activation function in this model, and σ(·) is the sigmoid function used to determine how much proportion of information can pass through to the next layer. The final h will be used as the input of the spatial feature extraction module; in the spatial information extraction module, two different levels of graph convolution are used, one is the diffusion graph convolution based on the adaptive adjacency matrix and the traditional road network structure to model the spatial relationship between paired nodes; the other is the hypergraph convolution based on the multi-level hypergraph constructed in step 1 to model the spatial relationship between multiple nodes; By initializing E in the hypergraph in step 1 1 , E 2 The node embedding calculates the adaptive adjacency matrix and combines it with the static adjacency matrix to perform a bidirectional diffusion graph convolution operation. The specific process is as follows: Among them, A adp is an adaptive adjacency matrix that can continuously optimize itself during the back propagation of the model; K is the number of steps of random diffusion, W k1 , W k2 , W k3 are all learnable weight matrices in the model; since the traffic graph is a directed graph, the diffusion graph convolution is also bidirectional, P f , P b They are the forward transfer probability matrix and the backward transfer probability matrix respectively. Both transfer probability matrices are calculated through the static adjacency matrix A. The specific calculation formula is: Among them, rowsum(·) is a function that calculates the sum of each row of the matrix; The high-order spatial dependency information of the data is extracted through the multi-level hypergraph constructed in step 1. It is divided into two steps. The first step is to form the hyperedge information by aggregating the node information within the hyperedge. The second step is to update the node information itself by aggregating the hyperedge information connected to the node. The specific process is as follows: Among them, h e is the hyperedge hidden feature obtained by aggregating the feature values of nodes within the hyperedge, W e is the learnable weight matrix, d i is the degree of node i, that is, the number of hyperedges connected to node i, d e For super edge The average degree of is the average degree of each node in the hyperedge; The results after diffuse graph convolution and hypergraph convolution are added together, and the data input to this layer is used as a residual connection to finally obtain the output of this spatiotemporal layer.
2. The prediction method according to claim 1, It is characterized in that In step 1, traffic speed data is input into the model, and the speed data is dimensionally enhanced through a linear layer, converting the speed value from a scalar to a vector, so that the model can extract richer information; The K-means unsupervised clustering method is used to obtain the clustering results of the input data. Combined with the clustering results obtained by pre-training the multi-layer K-means clustering method on the entire training set, the clustering results are used as a set of hyperedges in the hypergraph. At the same time, each node in the traffic road network structure is combined with its neighboring nodes to construct a set of hyperedges. The multi-level hypergraph constructed by the two sets of hyperedges includes both hyperedges based on traffic data and hyperedges based on road network structure, which will be used in the hypergraph convolution of the spatiotemporal extraction module.
3. The prediction method according to claim 1, It is characterized in that In step 3, the traffic data contains strong periodicity, and the change in traffic speed is strongly correlated with a certain day of the week and a certain time of the day. This periodicity is embedded by constructing a network consisting of two linear layers based on the specific time information of the traffic data, and used as auxiliary data for predicting future traffic speeds. Combining the spatiotemporal features extracted by the multi-layer modules can ensure that the model will not encounter gradient vanishing or gradient exploding situations. Combining the periodic features with the extracted recent spatiotemporal features, future traffic speed prediction is performed through the linear layer.
4. The prediction method according to claim 1, It is characterized in that In step 4, the mean absolute error (MAE) and the root mean square error (RMSE) are combined, so that the loss of the time segment where congestion or mutation occurs in the real traffic speed data is calculated using RMSE, while the loss of the other parts with normal speed is calculated using MAE; the time segment where congestion occurs is fitted; and finally, back propagation and gradient descent are performed to obtain the optimal model.
Citation Information
Patent Citations
Time sequence prediction method based on time convolution and LSTM
CN108764460A
Traffic flow prediction method based on space-time hypergraph neural network
CN114944053A