Dynamic traffic flow direction migration method for traffic key road identification
By constructing Path2vec and GraphSAGE layer models in the traffic network, combining dynamic traffic flow direction walking method and k-means clustering, the accuracy and multi-dimensional feature fusion problems of critical road segment identification in the existing technology are solved, and a more efficient identification of critical road segments in the traffic network is achieved.
Patent Information
- Application Number
- CN202510409562.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing technology relies on artificial design indicators to lead to the deviation of the result in key road segment identification. The graph neural network method fails to accurately identify the dynamic transmittance characteristics of traffic flow, ignores the fusion of dynamic and static characteristics of multi-dimensional area, and the graph random sampling mechanism lacks targeted traffic road network.
Using a line graph mapping method based on traffic network, combined with Path2vec and GraphSAGE layer models, the dynamic traffic flow roaming model is constructed, and the dynamic and static characteristics of the path and nodes are obtained, and the k-means algorithm is used to cluster and identify key road sections.
It improves the accuracy of key road identification, enhances the ability to capture static and dynamic feature of the traffic network, improves the network efficiency, resistance to destruction, and the accuracy of evaluation of dynamic contagious indicators.
Smart Images

Figure CN120279705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of critical traffic road identification, and particularly to a dynamic traffic flow wandering method for traffic critical road identification. Background Art
[0002] In the field of critical section identification, researchers have been working on improving the accuracy and practicality of road network critical node identification by integrating traffic characteristics and advanced algorithm models, providing important technical support for urban traffic management. With the development of complex network theory and spatio-temporal data mining technology, the critical section identification framework based on graph neural network has been widely applied. With the proposal of the dual network structure of road entity topology and traffic flow topology, it breaks through the limitations of traditional single road network representation, provides the underlying data structure support for the identification method based on graph representation learning, so as to extract the low-dimensional embedding features of road section nodes by using graph neural network, and effectively capture the spatial correlation and traffic flow propagation characteristics of the road network. The evaluation system of critical sections has gradually shifted from single feature index evaluation to multi-dimensional feature index evaluation. By analyzing the research results in recent years, it can be seen that the selection and evaluation methods of critical sections in the road network are changing from traditional index selection methods to critical section selection and evaluation methods using dynamic feature learning.
[0003] Although these methods have achieved certain results in the field of critical section identification, there are still some problems:
[0004] (1) The traditional research on critical sections mainly adopts two methods: one is the index evaluation method, including single-index evaluation method and multi-index evaluation method, which establish an evaluation system by selecting one or more measurement indexes to identify network critical nodes; the other is to identify network critical nodes by means of node deletion method or node contraction method, and identify network critical nodes by comparing the changes in network performance before and after removing a certain node. The traditional critical section evaluation method depends on artificially designed indexes, which may lead to biased results. The method of iteratively deleting nodes is not suitable for large-scale traffic networks, and cannot fully obtain the spatial correlation between nodes from the perspective of the graph, and cannot capture complex network features, such as dynamic changes and high-order relationships between nodes.
[0005] (2) With the development of graph neural networks, the method of selecting critical nodes based on graph neural networks has gradually been applied to traffic networks. Graph neural networks automatically learn the features of nodes and edges through the node aggregation mechanism, can capture features that are difficult to obtain by traditional methods, and have advantages in dealing with complex large-scale graph classification tasks. In most existing works based on graph neural networks, the critical section acquisition method only obtains node embeddings from the perspective of random walk, fails to accurately identify the dynamic transfer characteristics of traffic flow, ignores the fusion of multi-dimensional regional dynamic and static features, and the graph random sampling mechanism lacks pertinence to traffic road networks. Summary of the Invention
[0006] In view of the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to more accurately identify key roads. To solve the above technical problem, the present invention adopts the following technical solutions:
[0007] A dynamic traffic flow wandering method for identifying key roads in traffic, comprising the following steps:
[0008] Select the traffic database X of a certain area:
[0009] S100: Use the networkx library and pandas library in Python to process X to obtain the traffic data information table corresponding to X, and convert the traffic data information table into a matrix to obtain the traffic network line graph G; all traffic data information is based on the records of all license plate numbers passing through the detector within each hour, calculate the hourly traffic volume between the detectors to which each group of related nodes belong, and according to the connection relationship of the detector road network;
[0010] The expression of the traffic network line graph G is as follows:
[0011] G=(V,E,S T ,F T ,Q T ,K T ,D T ,B T ,C T )
[0012] Wherein, V=(v1,v2,…,v n ) is the set of nodes, v n represents the nth node in G, n represents the total number of nodes in G, v n represents a node in G and also represents a section of a road in an actual road in a certain area; E=(v i ,v j ) represents the edge set in G, (v i ,v j ) represents an edge between v i and v j in G, i,j∈[1,n]; S T =(s1,s2,…,s n ) represents the set of average speeds of each node in G within a specific time T, s n represents the average speed of the vehicle passing through the nth node in G; F T =(f1,f2,…,f n ) represents the set of traffic volumes of each node in G within a specific time T, f n represents the traffic volume passing through the nth node in G; Q T=(q1,q2,…,q n ) represents the set of lengths of the actual road segments corresponding to the nodes, where q n represents the length of the actual road segment corresponding to the nth node in G; K T =(k1,k2,…,k n ) represents the set of vehicle flow densities of the nodes, where k n represents the vehicle flow density passing through the nth node in G; D T =(d1,d2,…,d n ) represents the set of node degrees, where d n represents the node degree of the nth node in G; B T =(b1,b2,…,b n ) represents the set of betweenness centralities of the nodes, where b n represents the betweenness centrality of the nth node in G; C T =(c1,c2,…,c n ) represents the set of closeness centralities of the nodes, where c n represents the closeness centrality of the nth node in G; T is generally a constant;
[0013] Compared with the traffic network of the original method, the traffic network emphasizes the important role of the links in the traffic system. This method enhances the clarity of the topological relationship between different links. In the traffic network model, the static and dynamic characteristics of the road links can be captured. The structure of the traffic network shows the connection relationship between different links and depicts the static characteristics of the road links. At the same time, the traffic flow flow, average speed speed, length length indicators of each road segment and other indicators reflect the dynamic demand characteristics of the traffic network.
[0014] S200: Construct a key road identification model M, where M includes a Path2vec layer model and a GraphSAGE layer model; the Path2vec layer model improves DeepWalk using traffic flow, and improves the completely random walk function of DeepWalk into a walk function that conforms to traffic flow; DeepWalk and the GraphSAGE layer model are existing technologies.
[0015] S300: Input G into M:
[0016] G outputs the path embedding feature vector Path through the Path2vec layer model v ; G outputs the node embedding feature vector through the GraphSAGE layer model N(v) represents the neighborhood of node v;
[0017] S400: Concatenate Path obtained in S300 v and Then, after linear transformation, we get the final embedded feature vector of v at layer l. The specific formula is as follows:
[0018]
[0019] Among them, CONCAT(·) represents the concatenation operation, represents any node u, represents the embedded feature vector of v at the l-1 layer, represents the embedded feature vector of u at layer l, represents the embedding vector obtained by v through the Path2vec layer model, σ is the nonlinear activation function, and W l is the learnable weight matrix of layer l;
[0020] S500: Set clustering algorithm parameter λ and use clustering algorithm to Clustering is performed to obtain the λ-class traffic sections;
[0021] S600: Calculate the average value of the traffic demand index of all nodes in each type of traffic section obtained in S500 The calculation formula is as follows:
[0022]
[0023]
[0024] in, h is The total number of nodes in the traffic segment, q h represents the length of the actual road section corresponding to the hth node, f h represents the traffic flow passing through the hth node, s h represents the average speed of the vehicle passing the hth node;
[0025] S700: Select the category with the highest average value of traffic demand index as the key category, sort the traffic demand indexes of all nodes in the key category in descending order, select the actual road sections corresponding to the traffic demand indexes of the first W nodes as the key road sections, output the W key road sections, and obtain the final key road list;
[0026] S800: Select the traffic database of the area Y to be tested, repeat S100-S700, and obtain a key road list of Y.
[0027] Preferably, in S300, G is output through the Path2vec layer model to obtain the path embedding feature vector Path v The specific steps are as follows:
[0028] S310: Calculate the adjacency relationship A between i and j ij , and the specific expression is as follows:
[0029]
[0030] The non - zero elements in the adjacency matrix A represent the connections between road segments. There can only be a transfer probability between adjacent road segments, otherwise the probability is 0;
[0031] S320: Obtain the traffic flow f of j that conforms to the adjacency relationship A with i from G ij ; j ;
[0032] S330: Calculate the transfer weight ω from i to j ij , and perform normalization processing. The specific expression is as follows:
[0033]
[0034]
[0035] Among them, represents the distance between nodes v i and v j in the graph G;
[0036] S340: Calculate the transfer matrix T for i to transfer to j in the next step ij , and the calculation formula is as follows:
[0037]
[0038] Among them, N(i) is the set of adjacent nodes of i, k ∈ N(i) means that k is the k - th adjacent node of i, and f k represents the traffic flow passing through k;
[0039] S350: Calculate the transfer probability P for i to transfer to j ij , and the calculation formula is as follows:
[0040] P ij = T ij ·ω ij
[0041] S360: Obtain the random walk sequence of node i , and the specific expression is as follows:
[0042]
[0043] Among them, represents the random walk sequence at the t - th node starting from v i , m represents the length of the random walk sequence, t = 1, 2,..., m;
[0044] S370: Calculate Path v , and the specific expression is as follows:
[0045]
[0046] Among them, Path v represents the embedding vector of node v, and Word2vec represents an existing word vector training model imported from a Python library.
[0047] Preferably, in the S300, G outputs the node embedding feature vector through the GraphSAGE layer model and the calculation formula is as follows:
[0048]
[0049] N(v) = {u: D(u, v) ≤ k, u ∈ G}
[0050] Among them, D(u, v) represents the minimum distance between node v and node u, k represents the number of sampling layers, and AGGREGATE l (·) represents the aggregation operation, represents the embedding feature vector of node u at the (l - 1)th layer.
[0051] Preferably, in the S500, the clustering algorithm is the kmeans algorithm, and λ = 3 is set. The importance of road connections is usually divided into three categories: critical, normal, and unimportant. The clustering algorithm is an existing model and can output the number of categories required by the model according to parameter settings.
[0052] Compared with the prior art, the present invention has at least the following advantages:
[0053] 1. Use the line graph mapping method based on the traffic network, where the road connection lines are represented as nodes, and the connection relationship between them is represented as edges. Incorporate the dynamic attributes such as traffic flow and speed, and the static attributes such as length and degree of the traffic section into the node attributes of the traffic network, expand the data collection range, and improve the accuracy of the final recognition.
[0054] 2. Aiming at the problem that in the graph embedding algorithm based on graph random walk such as Node2vec in the traffic network node representation, due to the physical conflict between the hyperparameter randomness and the spatio-temporal heterogeneity of traffic flow, the node embedding space cannot effectively encode the traffic flow dynamic transfer sequence features. The present invention innovatively proposes the dynamic traffic flow direction random walk model Path2Vec, constructs the traffic flow transfer matrix, and obtains the random walk path sequence features that match the historical traffic flow transfer probability through random walk path sampling.
[0055] 3. Aiming at the problems that traditional GraphSAGE has a single feature extraction in traffic network representation and fails to effectively integrate the dynamic and static attributes of road network nodes, the present invention innovatively combines Path2vec with GraphSAGE, constructs multi-dimensional regional dynamic and static features of nodes through a graph sampling aggregation model, and obtains key road segments.
[0056] 4. From the perspectives of the dynamic and static features of the traffic network, the accuracy of road segment selection is evaluated based on node network efficiency, network invulnerability, and dynamic contagion indicators. Experimental data from the Chongqing West Station network and PEMS08 show that the method proposed in the present invention has higher accuracy compared with traditional index evaluation methods. Through data analysis, it is found that compared with the selection methods based on traffic flow, degree, betweenness, and closeness in traditional indexes, the model proposed in the present invention is significantly superior to the key road segments obtained by traditional methods in terms of network efficiency indexes. It is superior in terms of the performance of the ratio index of the number of subgraphs and the number of nodes, and has a huge advantage in terms of network efficiency index compared with other methods, indicating that the model fully obtains the static regional features of the road network and the dynamic transfer sequence features of traffic flow, has a greater impact on the efficiency of information exchange in the network and the integrity of the road network, and improves the influence on network dynamic and static indexes. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is the architecture diagram of the hybrid model of the present invention;
[0058] Figure 2 It is the architecture diagram of the Path2vec model;
[0059] Figure 3 It is the architecture diagram of the GraphSAGE model;
[0060] Figure 4 It is the construction process of the traffic network line diagram;
[0061] Figure 5 It is the distribution diagram of the original detection equipment around the West Station;
[0062] Figure 6 It is the traffic network line diagram around the West Station;
[0063] Figure 7 It is the performance of the network efficiency during the morning peak period from 8:00 to 9:00 at the Chongqing West Station;
[0064] Figure 8 It is the performance of the number of infected nodes existing in the simulated infectious disease model network during the morning peak period from 8:00 to 9:00 at the Chongqing West Station;
[0065] Figure 9 It is the comparison of the performance of the subgraph number index during the morning peak period from 8:00 to 9:00 at the Chongqing West Station;
[0066] Figure 10 The comparison of the ratio index performance between the number of nodes in the largest connected component of the network after deleting nodes and the number of nodes in the original graph during the morning peak period from 8:00 to 9:00 at Chongqing West Station;
[0067] Figure 11 A comparison chart of the positions of the top 15 key road segments obtained by the model of the present invention and the traffic flow model only in the actual road during the morning peak period from 8:00 to 9:00 at Chongqing West Station. Detailed implementation manners
[0068] The present invention will be further described in detail below.
[0069] Identifying key road segments in the road network and taking them as the core for regional traffic control and optimization play a crucial role in traffic. The present invention proposes an improved traditional biased random walk model Path2vec to obtain path feature embeddings that conform to the traffic flow direction, aggregates the key road segment hybrid model of Graph Sampling Aggregation (GraphSAGE) to obtain graph embedding vectors, and then clusters the obtained embedding vectors to obtain key road segments in the network.
[0070] See Figures 1 - 11 , a dynamic traffic flow direction walking method for traffic key road recognition, including the following steps:
[0071] Select the traffic database X of a certain region:
[0072] S100: Use the networkx library and pandas library in Python to process X to obtain the traffic data information table corresponding to X, and convert the traffic data information table into a matrix to obtain the traffic network line graph G;
[0073] The expression of the traffic network line graph G is as follows:
[0074] G = (V, E, S T , F T , Q T , K T , D T , B T , C T )
[0075] Among them, V = (v1, v2,..., v n ) is the set of nodes, v n represents the nth node in G, n represents the total number of nodes in G, E = (v i , v j ) represents the edge set in G, (v i , v j ) represents v i and vj An edge between them, where i, j ∈ [1, n]; S T =(s1, s2, …, s n ) represents the set of average speeds of each node in G within a specific time T, and s n represents the average speed of the vehicle passing through the nth node in G; F T =(f1, f2, …, f n ) represents the set of traffic flows of each node in G within a specific time T, and f n represents the traffic flow passing through the nth node in G; Q T =(q1, q2, …, q n ) represents the set of lengths of the actual road segments corresponding to the nodes, and q n represents the length of the actual road segment corresponding to the nth node in G; K T =(k1, k2, …, k n ) represents the set of traffic flow densities of the nodes, and k n represents the traffic flow density passing through the nth node in G; D T =(d1, d2, …, d n ) represents the set of node degrees, and d n represents the node degree of the nth node in G·; B T =(b1, b2, …, b n ) represents the set of betweenness centralities of the nodes, and b n represents the betweenness centrality of the nth node in G; C T =(c1, c2, …, c n ) represents the set of closeness centralities of the nodes, and c n represents the closeness centrality of the nth node in G;
[0076] S200: Construct a key road identification model M, where M includes a Path2vec layer model and a GraphSAGE layer model; the Path2vec layer model improves DeepWalk using traffic flow, and improves the completely random walk function of DeepWalk to a walk function that conforms to traffic flow;
[0077] S300: Input G into M:
[0078] G outputs a path embedding feature vector Path through the Path2vec layer model v ; G outputs a node embedding feature vector through the GraphSAGE layer model N(v) represents the neighborhood of node v;
[0079] In the said S300, the specific steps for G to output a path embedding feature vector Path through the Path2vec layer model v are as follows:
[0080] S310: Calculate the adjacency relationship A between i and j ij , and the specific expression is as follows:
[0081]
[0082] S320: Obtain the traffic flow f of j that conforms to the adjacency relationship A with i from G ij ; j
[0083] S330: Calculate the transfer weight ω from i to j ij , and perform normalization processing. The specific expression is as follows:
[0084]
[0085] where θ ij represents the distance between nodes v i and v j in the graph G;
[0086] S340: Calculate the transfer matrix T for i to transfer to j in the next step ij , and the calculation formula is as follows:
[0087]
[0088] where N(i) is the set of adjacent nodes of i, and k ∈ N(i) means that k is the k-th adjacent node of i, and f k represents the traffic flow passing through k;
[0089] S350: Calculate the transfer probability P for i to transfer to j ij , and the calculation formula is as follows:
[0090] P ij = T ij · ω ij
[0091] S360: Obtain the random walk sequence of node i , and the specific expression is as follows:
[0092]
[0093] where represents the random walk sequence at the t-th node starting from v i , m represents the length of the random walk sequence, and t = 1, 2,..., m;
[0094] S370: Calculate Path v , and the specific expression is as follows:
[0095]
[0096] Among them, Path v represents the embedding vector of node v, and Word2vec represents an existing word vector training model imported from a Python library.
[0097] In S300, G outputs the node embedding feature vector through the GraphSAGE layer model The calculation formula is as follows:
[0098]
[0099] N(v) = {u: D(u, v) ≤ k, u ∈ G}
[0100] Among them, D(u, v) represents the minimum distance between node v and node u, k represents the number of sampling layers, and AGGREGATE l (·) represents the aggregation operation, represents the embedding feature vector of node u at the (l - 1)th layer.
[0101] S400: Concatenate Path obtained in S300 v and Then, through linear transformation, the final embedding feature vector of v at the lth layer is obtained The specific formula is as follows:
[0102]
[0103] Among them, CONCAT(·) represents the concatenation operation, represents any node u, represents the embedding feature vector of v at the (l - 1)th layer, represents the embedding feature vector of u at the lth layer, represents the embedding vector obtained by v through the Path2vec layer model, σ is a non - linear activation function, and W l is the learnable weight matrix at the lth layer;
[0104] S500: Set the clustering algorithm parameter λ, and use the clustering algorithm to perform clustering to obtain λ classes of traffic sections;
[0105] In S500, the clustering algorithm is the kmeans algorithm, and λ = 3 is set.
[0106] S600: Calculate the average value of the traffic demand indicators of all nodes in each class of traffic sections obtained in S500 The calculation formula is as follows:
[0107]
[0108] Among them, h is the total number of nodes in the traffic-like section, and q h represents the length of the actual section corresponding to the h-th node, and f h represents the traffic flow passing through the h-th node, and s h represents the average speed of the vehicle passing through the h-th node;
[0109] S700: Select the category with the highest average value of traffic demand indicators as the key category. Arrange the traffic demand indicators of all nodes in the key category in descending order. Select the actual sections corresponding to the traffic demand indicators of the first W nodes as the key sections, and output the W key sections to obtain the final list of key roads;
[0110] S800: Select the traffic database of the area Y to be measured, and repeat S100 - S700 to obtain the list of key roads of Y.
[0111] Experiments and Results
[0112] Data Description
[0113] This study uses the private dataset of Chongqing West Station: The private dataset of Chongqing West Station contains the original traffic data from July 1st, 2024 to July 31st, 2024 obtained by the external road detection equipment of the road network around Chongqing West Station. After data cleaning and processing, traffic flow data, road average speed data, and road length data of each external road detection equipment during the morning peak (8:00 - 9:00), noon peak (12:00 - 13:00), and evening peak (17:00 - 18:00) are obtained. Figure 5 and Figure 6 shows the distribution of the original detection equipment around the west station, and a traffic network line map is established based on the connection of the detection equipment. PEMS08: It is the PeMS traffic flow dataset of the California highway network. PEMS08 is data generated by 170 detectors collecting data every 5 minutes for a total of 62 days. The same as the private dataset of Chongqing West Station, traffic data during peak periods is extracted.
[0114] Baseline Description
[0115] Compare the key sections obtained by a key section fusion model based on pathtoVec and graph sampling aggregation Graphsage proposed in the present invention with traditional methods for selecting key sections based on traffic flow, methods for selecting key sections based on centrality indicators, and methods for selecting key sections based on graph neural networks. Evaluate the accuracy of the selected key sections from the perspectives of node network efficiency, network invulnerability, and dynamic contagion indicators.
[0116] The method for selecting key road sections based on centrality indicators includes the node degree selection method, the node betweenness centrality selection method, and the closeness centrality selection method. The degree of a node characterizes the number of edges connected to the node. The node degree formula is as follows:
[0117] DC i =k i
[0118] Among them, DC i represents the node degree, and k i represents the number of edges connected to node i;
[0119] The node betweenness centrality refers to the number of shortest paths passing through a node in a network. The node betweenness centrality formula is as follows:
[0120]
[0121] Among them, represents the number of paths passing through node i and being the shortest path, and g st represents the number of shortest paths between s and t;
[0122] The node closeness reflects the average shortest path length from a specific node to any point in the network. The node closeness centrality formula is as follows:
[0123]
[0124] Among them, d i represents the average distance from node i to other points, and CC i represents the node closeness centrality;
[0125] The following baseline methods are used to make comparisons through data:
[0126] Only flow sorting onlyflow: Flow is an important data in traffic flow. Traffic road sections with high flow often have a greater impact on the traffic network. This method selects the top K traffic road sections in terms of flow in the network.
[0127] Only degree sorting OnlyDegree: The node degree is one of the indicators of network node centrality. The degree reflects the number of connections of a node in the network. The higher the degree of a node, the more edges are connected to the node. When this node fails, the number of nodes that may be affected will be more. This method selects the top K traffic road sections in terms of degree in the network.
[0128] Only Betweenness: The betweenness of a node is one of the indicators of the centrality of network nodes. Nodes with high betweenness are usually bridges in the network, connecting different communities or groups. For example, they are the intermediaries for information dissemination in social networks. This method selects the top K traffic segments ranked by betweenness in the network.
[0129] Only Closeness: The closeness of a node is one of the indicators of the centrality of network nodes, representing the reciprocal of the average shortest distance from a node to all other nodes, reflecting the information dissemination efficiency of the node. Nodes with high closeness can reach other nodes quickly. This method selects the top K traffic segments ranked by closeness in the network.
[0130] Path2vec: To reduce the uncertainty brought by the fixed parameters p and q in the traditional node2vec model and make the random walk process more in line with the path tendency in the road traffic network, the Path2Vec Embedding model constructs a transition matrix reflecting traffic flow directivity based on the node2vec model through the traffic flow in the network, represents the paths of traffic vehicles as embedding vectors, and obtains important segments through clustering, and selects the top K traffic segments from them.
[0131] GraphSAGE: It is improved based on the graph convolutional network. The feature attributes of nodes are integrated into a feature matrix. By sampling a certain number of neighbor nodes, the embedding vectors of nodes are obtained through an aggregation function, important segments are obtained through clustering, and the top K traffic segments are selected from them.
[0132] Experimental Results
[0133] By training onlyflow, OnlyDegree, OnlyBetweenness, OnlyCloseness, Path2vec, GraphSAGE on the private dataset of Chongqing West Station and the PEMS08 dataset, and comparing them with the model proposed in the present invention. The experiment simulated the key road segments obtained by the seven models for more than 100 times, and calculated the network overall efficiency index, network invulnerability index of the key road segments in each model, including the number of subgraphs existing in the network after deleting nodes, the ratio of the number of nodes in the largest connected component of the network after deleting nodes to the number of nodes in the original graph, and the number of infected nodes existing in the network in the simulated infectious disease model. The network overall efficiency index selects the ratio of the network efficiency of the largest connected subgraph of the network after deleting nodes to the network efficiency of the original network. The lower the value of this index, the greater the impact on the communication efficiency of the overall network when the node list in the network fails. The number of subgraphs existing in the network after deleting nodes measures that when the nodes in the node list are removed, the network may be divided into multiple subgraphs, resulting in a significant decrease in communication efficiency. The more subgraphs existing in the network after deleting nodes, the stronger the node's ability to divide the network, and the more important the node. The ratio of the number of nodes in the largest connected component of the network after deleting nodes to the number of nodes in the original graph. A low ratio means that when the key node is deleted, the propagation of traffic flow will be significantly inhibited, and at the same time, the smaller the ratio, the greater the impact on the connectivity of the remaining network. The number of infected nodes existing in the network in the simulated infectious disease model reflects the diffusion impact on the surrounding roads when the key node is congested from the perspective of dynamic propagation. The selected key node in the simulation is used as the congested node to infect the surrounding for a certain number of time steps. The higher the number of congested nodes, the stronger the diffusion congestion ability of the initial node and the greater the impact on the network. In the present invention, the infection rate is set to 0.8, the recovery rate is set to 0.2, and the time step is set to 15 to simulate the propagation situation of high infection rate and low recovery rate. Figures 7 - 11It shows the network efficiency after the key sections found by onlyflow, OnlyDegree, OnlyBetweenness, OnlyCloseness, Path2vec, GraphSAGE and the aggregation model method proposed by the present invention during the morning rush hour, lunch rush hour and evening rush hour in the Chongqing West Station dataset. It also calculates the number of subgraphs existing in the network after deleting nodes, the ratio of the number of nodes in the largest connected component of the network after deleting nodes to the number of nodes in the original graph, and the number of infected nodes existing in the network in the simulated infectious disease model. The average values of these sections in each result are obtained, as well as the total amounts of these indicators after 100 experiments. There are a total of 134 section nodes in the traffic network line map based on the Chongqing West Station, and a total of 277 section nodes in the traffic network line map constructed based on the PEMS08 dataset. In actual traffic control and planning, due to cost considerations, not too many nodes will be selected as key nodes. Therefore, it is set that K is 15 in the Chongqing West Station dataset model and K is 30 in the PEMS08 dataset model as the appropriate number of key sections.
[0134] Table 1 shows the performance of the verification indicators of each model on the PEMS08 dataset. The integrated model of Path2vec and GraphSAGE proposed by the present invention has the highest values in the two indicators of the number of infected nodes existing in the network in the simulated infectious disease model and the number of subgraphs existing in the network after deleting nodes. Followed by the key sections obtained by the Path2vec model. It has a large lead compared to the key sections selected according to the flow, degree, and betweenness indicators. This shows that on the PEMS08 dataset, the method proposed by the present invention is more excellent than the traditional method from the perspectives of dynamic propagation and network segmentation. The integrated model of Path2vec and GraphSAGE proposed by the present invention has the lowest values in the two indicators of network efficiency and the ratio of the number of nodes in the largest connected component of the network after deleting nodes to the number of nodes in the original graph. This shows that the nodes found by the method proposed by the present invention simultaneously carry the functions of structural hubs and dynamic transmission cores, and have a greater impact on the communication and connectivity of the network.
[0135] Table 1. Index Table of Each Model on the PEMS08 Dataset
[0136]
[0137] Figure 7Shows the performance of network efficiency during the morning peak period from 8:00 to 9:00 at Chongqing West Station. The total value of the network efficiency index evaluated 100 times for the Path2vec and GraphSAGE fusion model proposed in the present invention is 58.9. Compared with 71.6 of OnlySAGE, 72.9 of Path2vec, and the traffic-based selection method and centrality-based selection method with an average total network efficiency exceeding 80, it has a huge impact on the network efficiency index and is significantly better than the other six methods. The key sections obtained by the method proposed in the present invention bear the energy bottleneck of cross-regional traffic transfer and have greater global influence.
[0138] Figure 8 Shows the performance of the number of infected nodes existing in the network in the simulated infectious disease model during the morning peak period from 8:00 to 9:00 at Chongqing West Station. The total value of the number of infected node indicators for 100 times of the Path2vec and GraphSAGE fusion model proposed in the present invention is 917. Compared with 840 of OnlySAGE, 820 of Path2vec, 829 based on OnlyCloseness, and other selection methods with a total value less than 800, it shows that the method of the present invention more stably predicts the key nodes with higher infection ability in multiple simulations, resulting in a wider spread and being more persistent and effective in predicting the spread dynamics.
[0139] Figure 9 Shows the performance of the subgraph number index during the morning peak period from 8:00 to 9:00 at Chongqing West Station. The total value of the subgraph number index for 100 times of the Path2vec and GraphSAGE fusion model proposed in the present invention is 359.4. The value of OnlySAGE is 355.7, which is close to the method proposed in the present invention. The rest of the methods do not exceed 300, indicating that the ability of OnlySAGE to aggregate neighbor information can capture the local and global characteristics of the road sections, obtain the topological attributes of the road sections in the graph, and already has good performance itself, having a greater advantage compared with traditional methods. The model integrating Path2vec proposed in the present invention has a certain similarity in mechanism with the OnlySAGE model, resulting in close results, but there is still a certain improvement.
[0140] Figure 10 Shows the performance of the ratio index of the number of nodes in the largest connected component of the network after deleting nodes to the number of nodes in the original graph during the morning peak period from 8:00 to 9:00 at Chongqing West Station. The node number ratio index of the Path2vec and GraphSAGE fusion model proposed in the present invention is better than the other six methods, indicating that the method proposed in the present invention takes into account both the local connection mode and the global path dependence, and more comprehensively evaluates the functions of nodes in complex traffic networks. After the key sections obtained by the present invention are deleted, the network cannot maintain a high integrity, and the nodes obtained by the fusion model have a greater impact on the network connectivity.
[0141] Figure 11 It shows the positions of the top 15 key road segments obtained by the Path2vec and GraphSAGE fusion model proposed by the present invention and the positions of the top 15 key road segments obtained only by traffic flow in the actual traffic roads during the morning peak period from 8:00 to 9:00 at Chongqing West Station. The present invention draws them in the figure with red line segments. The key road segments in the actual road network may be distributed on the main roads, bridges, tunnels connecting different regions, or the connecting road segments around transportation hubs. The results show that the key road segments based on traffic flow only include the traffic main roads with dense traffic flow. However, the key road segments obtained by the method proposed by the present invention not only include the main roads, but also include important weaving areas, which are more consistent with the frequently congested traffic road segments in reality and can better conform to the positions of traffic road segments in experience.
[0142] Table 2. Index table of each model for the dataset of Chongqing West Station
[0143]
[0144]
[0145] Table 2 shows the total average values of the four indicators mentioned above after 100 simulation experiments during the morning peak, afternoon peak, and evening peak periods at Chongqing West Station. Through data analysis, compared with the selection methods based on traffic flow, degree, betweenness, and closeness in traditional indicators, Path2vec, GraphSAGE, and the aggregation method proposed by the present invention in the graph neural network method are significantly superior to the key road segments obtained by traditional indicators in the four indicators mentioned above. This shows that the graph neural network-based method can fully obtain the connection situation and topological structure information between road segments in the traffic network, has stronger structure recognition ability, is stronger in the diffusion ability of traffic congestion and the ability to damage the network structure. Compared with the traditional method based on indicators for obtaining key road segments after integrating the GraphSAGE module, it has great advantages on the Chongqing West Station dataset and has a great impact on the propagation efficiency of traffic flow. The method for obtaining key road segments of the fusion model proposed by the present invention is superior to the single GraphSAGE method and Path2vec method in terms of network efficiency, SIR model quantity, subgraph quantity, and node quantity ratio indicators mentioned above. This shows that the fusion of GraphSAGE and Path2vec models has a greater impact on the efficiency of exchanging information in the network, stronger ability to damage the network, stronger ability to spread congestion, and the obtained nodes are more critical, and can better integrate the road segment attributes and the traffic flow direction in the actual traffic network.
[0146] From the perspectives of the dynamic and static characteristics of the traffic network, based on node network efficiency, network invulnerability, and the accuracy of selecting road segments by the dynamic contagion index evaluation. Experimental data from the Chongqing West Station network and PEMS08 show that the technical method proposed by the present invention has higher accuracy compared with the traditional index evaluation method, can better obtain traffic flow characteristic information, has stronger structure recognition ability, expands the impact on the communication efficiency of the overall network when the node list in the network fails, the node's ability to partition the network is also stronger than the traditional method, has a greater impact on the connectivity of the network, and has a stronger congestion diffusion ability.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A dynamic traffic flow wandering method for identifying key traffic roads, characterized in that: It includes the following steps: Select the traffic database X of a certain region: S100: Use the networkx library and pandas library in Python to process X to obtain the traffic data information table corresponding to X, and convert the traffic data information table into a matrix to obtain the traffic network line graph G; The expression of the traffic network line graph G is as follows: G = (V, E, S T , F T , Q T , K T , D T , B T , C T ) where \(V=(v_1, v_2, \ldots, v\) n ) is the set of nodes, \(v\) n represents the \(n\)th node in \(G\), \(n\) represents the total number of nodes in \(G\), \(E=(v\) i , v\) j ) represents the set of edges in \(G\), \((v\) i , v\) j ) represents an edge in \(G\) between \(v\) i and \(v\) j , where \(i, j \in [1, n]\); \(S\) T =(s_1, s_2, \ldots, s\) n ) represents the set of average speeds of each node in \(G\) within a specific time \(T\), \(s\) n represents the average speed of the vehicle passing through the \(n\)th node in \(G\); \(F\) T =(f_1, f_2, \ldots, f\) n ) represents the set of traffic flows of each node in \(G\) within a specific time \(T\), \(f\) n represents the traffic flow passing through the \(n\)th node in \(G\); \(Q\) T =(q_1, q_2, \ldots, q\) n ) represents the set of lengths of the actual road segments corresponding to the nodes, \(q\) n represents the length of the actual road segment corresponding to the \(n\)th node in \(G\); \(K\) T =(k_1, k_2, \ldots, k\) n ) represents the set of traffic densities of the nodes, \(k\) n represents the traffic density passing through the \(n\)th node in \(G\); \(D\) T =(d_1, d_2, \ldots, d\) n ) represents the set of node degrees, \(d\) n represents the node degree of the \(n\)th node in \(G\); \(B\) T =(b_1, b_2, \ldots, b\) n ) represents the set of betweenness centralities of the nodes, \(b\) n represents the betweenness centrality of the \(n\)th node in \(G\); \(C\) T =(c_1, c_2, \ldots, c\) n ) represents the set of closeness centralities of the nodes, \(c\) n represents the closeness centrality of the \(n\)th node in \(G\); S200: Construct a key road recognition model M, where M includes a Path2vec layer model and a GraphSAGE layer model; the Path2vec layer model improves DeepWalk using traffic flow, and improves the completely random walk function of DeepWalk to a walk function that conforms to traffic flow; S300: Input G into M: G outputs the path embedding feature vector Path through the Path2vec layer model v ; G outputs the node embedding feature vector through the GraphSAGE layer model N(v) represents the neighborhood of node v; S400: Concatenate the Path obtained in S300 v and Then, through linear transformation, obtain the final embedded feature vector of v at the l-th layer The specific formula is as follows: Among them, CONCAT(·) represents the concatenation operation, represents any node u, represents the embedding feature vector of v at the (l - 1)-th layer, represents the embedding feature vector of u at the l-th layer, represents the embedding vector obtained by v passing through the Path2vec layer model, σ is a non-linear activation function, and W l is the learnable weight matrix at the l-th layer; S500: Set the clustering algorithm parameter λ and use the clustering algorithm to perform clustering to obtain λ classes of traffic segments; S600: Calculate the average value of the traffic demand indicators of all nodes in each type of traffic section obtained by S500 respectively The calculation formula is as follows: Among them, h is the total number of nodes in the traffic-like section, q h represents the length of the actual section corresponding to the h-th node, f h represents the traffic flow passing through the h-th node, s h represents the average speed of the vehicle passing through the h-th node; S700: Select the category with the highest average traffic demand index as the key category, sort the traffic demand indexes of all nodes in the key category in descending order, select the actual sections corresponding to the traffic demand indexes of the first W nodes as the key sections, output the W key sections, and obtain the final list of key roads; S800: Select the traffic database of the area to be measured Y, and repeat S100 - S700 to obtain the list of key roads in Y.
2. The dynamic traffic flow wandering method for identifying key traffic roads according to claim 1, characterized in that: In the above S300, the path embedding feature vector Path is obtained by the G output through the Path2vec layer model v The specific steps are as follows: S310: Calculate the adjacency relationship A between i and j ij , and the specific expression is as follows: S320: Obtain the traffic flow f of j that conforms to the adjacency relationship A with i from G ij ; j ; S330: Calculate the transition weight ω from i to j ij , and perform normalization. The specific expression is as follows: Among them, θ ij represents the distance between nodes v i and v j in graph G; S340: Calculate the transition matrix T for the next transfer of i to j ij , and the calculation formula is as follows: Among them, N(i) is the set of adjacent nodes of i, k ∈ N(i) means that k is the k-th adjacent node of i, and f k represents the traffic flow passing through k; S350: Calculate the transition probability P from i to j ij , and the calculation formula is as follows: P ij = T ij · ω ij S360: Obtain the walk sequence of node i The specific expression is as follows: Among them, represents the walk sequence starting from v i as the starting point at the t-th node, m represents the length of the walk sequence, and t = 1, 2, …, m; S370: Calculate Path v , and the specific expression is as follows: Among them, Path v represents the embedding vector of node v, and Word2vec represents an existing word vector training model imported from a Python library.
3. The dynamic traffic flow wandering method for identifying key traffic roads according to claim 2, wherein: In the above S300, G outputs the node embedding feature vector through the GraphSAGE layer model The calculation formula is as follows: N(v) = {u: D(u, v) ≤ k, y ∈ G} Among them, D(u, v) represents the minimum distance between node v and node u, k represents the number of sampling layers, and AGGREGATE l (·) represents the aggregation operation, represents the embedded feature vector of node u at the (l - 1)-th layer.
4. A dynamic traffic flow wandering method for identifying key traffic roads as described in claim 3, characterized in that: In S500, the clustering algorithm is the kmeans algorithm, and λ = 3 is set.
Citation Information
Patent Citations
Macroscopic fundamental diagram-based road network key section identification method
CN105702031A
Road network key section identification method
CN119007146A
Key road section identification method based on two-stage feature learning
CN119229649A