Traffic bottleneck identification method based on reverse influence sampling and evolutionary algorithm

By using reverse influence sampling and evolutionary algorithms in traffic bottleneck identification, the problems of long-term impact assessment and excessive computing resource overhead in the existing technology are solved, and the effect of efficiently identifying traffic bottlenecks is achieved.

CN120148247APending Publication Date: 2025-06-13SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510414851.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When identifying traffic bottlenecks, the impact assessment takes a long time, the complex network topology leads to poor solution effects, excessive computing resource overhead, and difficult to apply to large-scale road network instances.

Method used

The proxy model and evolutionary algorithm based on reverse influence sampling are adopted to pre-calculate the influence propagation results through reverse influence sampling, avoid the high time overhead of Monte Carlo simulation, and improve the search efficiency through local search strategies and reduce the number of iterations.

Benefits of technology

While maintaining less computing time and memory overhead, it significantly improves the ability to find the best in the problem of maximizing influence on the transportation network, and obtains a higher-quality collection of key nodes in the road network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148247A_ABST
    Figure CN120148247A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic bottleneck identification method based on reverse influence sampling and an evolutionary algorithm. The method comprises the following steps: acquiring historical traffic data of a city; a city map is divided into grids to construct a road network, each grid comprises a central area and road sections communicated with the central area, the traffic flow of each area grid of the road network is constructed according to GPS information and traffic flow data of each grid, the traffic state of the area grids is further obtained, a vehicle density-congestion degree formula is obtained, and the traffic flow of each area grid is calculated according to the vehicle density-congestion degree formula. Establishing a relationship between the degree of congestion and the traffic flow; a reverse influence sampling method is adopted to establish a proxy model for the traffic bottleneck identification problem, the proxy model uses a two-dimensional adaptive value to describe the quality of a candidate solution, the maximum estimated value of the congestion degree of the region is selected, and the key traffic bottleneck region is identified by taking the minimum estimated experience error as the target. The problems that in the prior art, when the influence maximization problem on the road network is solved, the search efficiency is low, and the final solution quality is poor are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic planning, and particularly to a traffic bottleneck identification method based on reverse influence sampling and evolutionary algorithm. Background Art

[0002] With the acceleration of urbanization, traffic congestion has become a common problem faced by many large cities around the world. As one of the main causes of traffic congestion, traffic bottlenecks have a significant impact on the smooth operation of urban traffic flow. Considering the dynamic and complex nature of urban traffic networks, some researchers have extended the influence maximization problem to the traffic field to effectively identify traffic bottlenecks, that is, to simulate the process of traffic congestion spreading between road networks, and use the impact of the change in the blockage degree of these nodes on the traffic flow of the entire road network to measure the influence of the nodes, and then find a set of fixed-size road network node sets that maximize the impact of traffic congestion spread by this set. However, the influence maximization problem belongs to the NP-hard problem, and existing methods usually use heuristic greedy algorithms to solve it. This method is prone to falling into local optimal solutions during the solution process, resulting in poor solution quality. Moreover, the evaluation of influence requires time-consuming Monte Carlo simulations, while another type of method based on reverse influence sampling to evaluate the influence size may bring excessive memory overhead, which makes existing methods difficult to apply to large-scale road network instances.

[0003] Evolutionary algorithms are a class of gradient-free solvers that can be effectively applied to a variety of complex optimization problems. Multiple studies have shown that for the influence maximization problem with a sparse search space, evolutionary algorithms can achieve better results than traditional greedy algorithms. The effectiveness of evolutionary algorithms benefits from their genetic operation mechanisms, including selection, crossover, and mutation steps, which help the algorithm jump out of local optimal solutions and explore a wider potential solution space. However, a key step in the optimization process is to estimate the influence of candidate solutions, which usually needs to be completed through a large number of Monte Carlo simulations. Given the randomness of influence propagation and the complexity of network topologies, this evaluation process may involve significant computational overhead. In the context of large-scale networks, algorithms using the Monte Carlo simulation method may take hours or even days to obtain a solution. Especially in traffic bottleneck identification applications that require real-time response, long computational delays are unacceptable. And sampling methods represented by reverse influence sampling may bring memory overhead of hundreds of gigabytes or more, which limits the deployment of optimization algorithms in actual scenarios. In view of this, exploring how to effectively identify traffic bottleneck points using evolutionary algorithms within acceptable computational resource limitations has become an urgent research topic. Summary of the Invention

[0004] In order to overcome the problems in the prior art of poor solution effect and excessive computing resource overhead caused by time-consuming influence evaluation and complex network topology, the present invention provides a traffic bottleneck identification method based on reverse influence sampling and evolutionary algorithm.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A traffic bottleneck identification method based on reverse influence sampling and evolutionary algorithm, comprising:

[0007] Acquire the city's traffic history data, wherein the traffic history data includes a city map, traffic flow data and GPS information, and the number of traffic bottleneck areas that need to be selected;

[0008] The city map is divided into grids to construct a road network. Each grid includes a central area and a road section connected to the central area. The traffic flow of each regional grid of the road network is constructed based on the GPS information and traffic flow data of each grid. The traffic status of the regional grid is further obtained to obtain the vehicle density-congestion formula, and the relationship between congestion and traffic flow is established to describe the diffusion process of congestion on the road network.

[0009] The reverse influence sampling method is used to establish a proxy model for the traffic bottleneck identification problem. The proxy model uses a two-dimensional fitness value to characterize the quality of candidate solutions, selects the maximum estimated value of the congestion level of the area, and identifies the key traffic bottleneck areas with the goal of minimizing the estimated empirical error.

[0010] Furthermore, the traffic bottleneck identification problem is defined as: given a road network G, an integer k and GPS information of several vehicles, the spread of traffic congestion in the past is described by GPS records. The traffic bottleneck identification problem requires finding k areas in the road network that can cause the most congestion in the future.

[0011] Furthermore, the GPS information refers to the record provided by the GPS device installed in the vehicle about the vehicle's location at a certain time every day, which is used to correspond the vehicle's flow information to the regional grid described by the road network.

[0012] Furthermore, the reverse influence sampling method comprises the following steps:

[0013] Estimator initialization: use the reverse influence sampling algorithm to initialize the estimator R;

[0014] Population initialization: use the greedy algorithm to initialize the population;

[0015] Population selection: From N individuals, N / 2 pairs of individuals are selected through the binary tournament selection method to produce new offspring;

[0016] Population crossover: Apply the crossover operation to the selected N / 2 pairs of individuals. Each new offspring's chromosome includes the intersection of the two chromosomes of its parents. The remaining positions in the chromosome are randomly sampled from the symmetric difference of the two chromosomes of the parents to finally obtain an offspring population O of size N.

[0017] Local search: Used to perform local search operations on the offspring population O obtained by crossover to obtain the fitness values of the new offspring population O'.

[0018] The population evolves to generate a new population.

[0019] Update the estimator: Obtain a new evaluation of the surviving individuals to get the fitness values (Inf, error).

[0020] Algorithm termination determination: If the algorithm meets the termination conditions, that is, the number of iterations reaches the maximum number of iterations or the proportion of a certain genotype in the population P exceeds the given threshold δ, the algorithm terminates and outputs the genotype with the highest proportion in the final population as the identified result of the key traffic congestion nodes; if the algorithm does not reach the termination conditions, that is, the number of evolutionary iterations is less than the maximum number of iterations and no genotype in the population exceeds the given threshold δ, the algorithm jumps to the population selection step to continue evolving.

[0021] Furthermore, the initialization of the population using the greedy algorithm is specifically as follows:

[0022] First, use the greedy algorithm to select 2*k times from the set of influence propagation samples called the reverse reachable set saved in the estimator R. Each time, select the node that can bring the greatest influence gain to the set of nodes that have already been selected, that is, maximize the marginal benefit, to ensure that there are enough high-value nodes for initializing the population.

[0023] Then, perform N random samplings from the 2*k nodes to construct the initial population. Each random sampling selects k nodes as an initial individual.

[0024] Calculate the fitness values (Inf, error) of each initial individual using the influence estimator, where Inf represents the influence estimation value given by the influence estimator, and error represents the empirical error of the estimation value.

[0025] Furthermore, the local search is used to perform local search operations on the offspring population O obtained by crossover to obtain the fitness values of the new offspring population O', specifically as follows:

[0026] First, from the current influence estimator, k nodes are selected by applying the greedy algorithm to form a set S. Then, each individual in the offspring O is crossed with S, and the set of new individuals forms a new offspring population O'. The fitness values of the new offspring population O' are calculated using the current influence estimator.

[0027] Furthermore, the population evolution is to obtain the new population generated in this evolution from the parent population P and the offspring O' generated by local search using non-dominated sorting and crowding distance calculation.

[0028] Furthermore, in the algorithm termination determination step, the threshold is set to 0.6.

[0029] Furthermore, the initial value of error is 1*10 5 。

[0030] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0031] (1) The method for identifying road network traffic bottlenecks based on reverse influence sampling and evolutionary algorithm provided by the embodiment of the present invention establishes a congestion diffusion model on the road network from information such as road network structure, traffic flow data, and GPS, transforms the congestion area identification into an influence maximization problem on the road network, and selects the road network node set with the greatest influence. The proxy model based on reverse influence sampling avoids the high time cost brought by the Monte Carlo simulation method. At the same time, it alleviates the memory occupation brought by directly using reverse influence sampling, and solves the problem of large memory and time overhead in the prior art when solving the influence maximization problem on the road network.

[0032] (2) The local search strategy provided by the embodiment of the present invention uses the greedy algorithm to mine information from the influence estimator, selects valuable nodes to guide the optimization process, and can improve the search efficiency and reduce the number of iterations more than the general random mutation strategy, solving the problems of low search efficiency and poor final solution quality in the prior art when solving the influence maximization problem on the road network. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of the present invention;

[0034] Figure 2 is a flowchart of the present invention based on reverse influence sampling and evolutionary algorithm. DETAILED DESCRIPTION OF THE INVENTION

[0035] The following will further describe the present invention in detail with reference to the embodiments, but the embodiments of the present invention are not limited thereto.

[0036] Embodiment

[0037] The traffic bottleneck recognition method based on reverse influence sampling and evolutionary algorithm in the embodiments of the present invention uses reverse influence sampling to construct an influence estimator. Through pre-computation, a sufficient number of influence propagation results are obtained in advance to avoid time-consuming Monte Carlo simulations during the solution process. To alleviate the potentially high memory overhead brought by the reverse influence sampling method, the number of samples used to construct the influence estimator is limited to a small range during each iteration to save memory. To improve the accuracy of evaluation, a new influence estimator is constructed during each iteration. The evaluation of candidate solutions is based on, and a local search strategy is designed to select valuable nodes in the road network to be added to the population to accelerate the convergence speed and improve the quality of the final solution. This method directly approximates the traffic congestion diffusion model on the original road network and considers the error between the simulation results and the real expectations, ensuring the quality of the final optimization result.

[0038] See Figure 1 , a traffic bottleneck recognition method based on reverse influence sampling and evolutionary algorithm provided by the embodiments of the present invention includes the following steps:

[0039] S1 Obtain the traffic historical data of the city, where the traffic historical data includes the city map, traffic flow data, and GPS information, as well as the number of traffic bottleneck areas to be selected.

[0040] The GPS information is the record provided by the GPS device installed in the vehicle about the vehicle's position at a certain moment every day, which is used to map the vehicle's movement information to the topological structure described by the road network.

[0041] Traffic flow data: Describes the data of vehicle inflow and outflow in a certain area at a certain time, which is used to calculate the traffic flow in that area.

[0042] S2 Divide the city map into grids to construct a road network. Each grid includes a central area and the road segments connected to the central area. Construct the traffic flow of each area grid in the road network according to the GPS information and traffic flow data of each grid, further obtain the traffic state of the area grid, obtain the vehicle density-congestion formula, and establish the relationship between congestion and traffic flow.

[0043] Road network structure: It is a traffic flow framework established based on geographical space division and network topological relationships. It divides the urban space into network units and represents the passing relationships between regions through directed edges. The traffic network can be represented as G(V, E). The vertex set V represents the area grids divided from the urban map area in this example, and the edge set E represents that two grids are physically adjacent and are connected by a road, and its direction corresponds to the one-way and two-way of the road.

[0044] The node of the road network is an area in the city, and the road in the edge city of the road network is a road. The traffic bottleneck area is the key node in the road network, and the identification of blocked sections is converted into the problem of identifying key nodes in the network.

[0045] Further explanation: Based on the original data of the city map, the city is divided into a set of grids. Each grid contains a central area and the road segments connecting the areas. Then a road network is constructed. The nodes of the road network are regional grids. The existence of an edge between two nodes in the road network means that the two nodes are adjacent in the city map.

[0046] Then, GPS data is a record of the movement speed of multiple vehicles at a certain time and location, which is used to construct the traffic flow on the road network. Specifically, based on the latitude and longitude information of GPS, the movement of vehicles on the map will be mapped to the movement of vehicles on the road network. According to the speed of vehicles in a certain area, we can define the congestion level of a certain area. Furthermore, according to the density-congestion formula, we can establish the relationship between congestion and traffic flow to describe the diffusion process of congestion on the road network.

[0047] S3 uses the reverse influence sampling method to establish a proxy model for the traffic bottleneck identification problem. The proxy model uses a two-dimensional fitness value to characterize the quality of candidate solutions, selects the maximum estimated value of the congestion level in the area, and minimizes the estimated empirical error as the goal to identify key traffic bottleneck areas.

[0048] Further, the definition of the traffic bottleneck identification problem is: given a road network G, an integer k and several vehicle GPS records that describe the spread of traffic congestion in the past, the traffic bottleneck identification problem requires finding k areas in the road network that will cause the most congestion in the future.

[0049] The present invention adopts the reverse influence sampling method to establish a proxy model for the original problem, and designs an evolutionary algorithm combined with a local search strategy to efficiently find areas in the urban road network that have a significant impact on traffic flow.

[0050] Specifically: Initialize the influence estimator; Apply the greedy algorithm to the influence sample set in the influence estimator, select 2*k valuable nodes, where k is the number of congested areas to be selected, and then sample N times from them. Each time, sample k nodes to form an initial population of size N; Evaluate the initial population using the initial influence estimator; Apply the binary tournament selection operation and the crossover operator to obtain the offspring O; Apply the local search operation to the offspring O to generate the final offspring O'; According to the definition of the surrogate model, calculate the two-dimensional fitness value of the offspring O', that is, the influence estimation and the corresponding empirical error; Apply non-dominated sorting and crowding distance calculation to the population P and the offspring O' to generate a new population P'; Update the influence estimator and update the fitness values of all individuals in the new population P'; Count the frequency of occurrence of each genotype in the population. If the frequency of occurrence of a certain genotype in the current population exceeds the set threshold of 0.6, or the number of evolutionary iterations exceeds the maximum number of iterations, the algorithm reaches the termination condition; If the algorithm does not reach the termination condition, that is, no genotype has a frequency of occurrence in the current population exceeding the threshold or the set maximum number of evolutionary iterations has not been exhausted, repeat the evolutionary process of the population; If the algorithm reaches the termination condition, stop the algorithm and output the genotype with the highest frequency of occurrence in the current population as the set of traffic congestion nodes with the greatest global influence. After appropriately simplifying the problem using the surrogate model based on reverse influence sampling, an evolutionary algorithm incorporating local search operations is used to solve the surrogate model. This method can significantly improve the optimization ability of the influence maximization problem on the traffic network while maintaining less computational time and memory overhead, and obtain a set of key nodes of the road network with higher quality.

[0051] As Figure 2 shown, the specific steps of the reverse influence and evolutionary algorithm are as follows:

[0052] 1) Estimator initialization:

[0053] The present invention uses the reverse influence sampling algorithm to initialize the estimator R to avoid performing Monte Carlo simulations on each individual during the search process to calculate the corresponding influence estimation value, reducing the computational resource overhead required for the solution;

[0054] 2) Population initialization:

[0055] The present invention initializes the population using the greedy algorithm, improving the quality of the individuals in the initial population while ensuring the diversity of the population;

[0056] First, use the greedy algorithm to select from the influence propagation sample set called the reverse reachable set saved in the estimator R 2*k times. Each time, select the node that can bring the greatest influence gain to the set of nodes that have already been selected, that is, maximize the marginal benefit, to ensure that there are enough high-value nodes for initializing the population;

[0057] Then, N random samplings are performed from 2*k nodes to construct the initial population. Each random sampling selects k nodes as an initial individual.

[0058] Use the influence estimator to calculate the fitness value (Inf, error) of each initial individual. Here, Inf represents the influence estimation value given by the influence estimator, and error represents the empirical error of the estimation value, which is initialized to 1*10 5 ;

[0059] 3) Population selection:

[0060] In the present invention, N / 2 pairs of individuals are selected from N individuals through the binary tournament selection method to generate new offspring.

[0061] 4) Population crossover:

[0062] The present invention applies the crossover operation to the selected N / 2 pairs of individuals. Each new offspring's chromosome includes the intersection of the two chromosomes of its parents. The remaining positions in the chromosome are randomly sampled and filled from the symmetric difference of the two chromosomes of the parents, and finally, an offspring population O of size N is obtained.

[0063] 5) Local search:

[0064] The present invention performs a local search operation on the offspring population O obtained by crossover. First, k nodes are selected from the current influence estimator by applying the greedy algorithm to form a set S. Then, each individual in the offspring O is crossed with S, and the set of new individuals obtained forms a new offspring population O'. The fitness value of the new offspring population O' is calculated using the current influence estimator.

[0065] 6) Population evolution:

[0066] The present invention uses non-dominated sorting and crowding distance calculation to obtain the new population generated in this evolution from the parent population P and the offspring O' generated by local search.

[0067] 7) Update the estimator:

[0068] The present invention uses the updated estimator to obtain a new evaluation of the surviving individuals. The average of all the estimation values of an individual so far is used as its influence estimation value, that is, the first dimension Inf of the fitness value. The length of the 95% confidence interval under the normal distribution is calculated from the set of all the estimation values of an individual so far as the empirical error of the estimation value, that is, the second dimension error of the fitness value.

[0069] 8) Algorithm termination determination

[0070] If the algorithm meets the termination condition, that is, the number of iterations reaches the maximum number of iterations or the proportion of a certain genotype in the population P exceeds the given threshold δ, the algorithm terminates and outputs the genotype with the highest proportion in the final population as the identified result of the key traffic congestion nodes; if the algorithm does not meet the termination condition, that is, the number of evolutionary iterations is less than the maximum number of iterations and the proportion of no genotype in the population exceeds the given threshold δ, the algorithm jumps to step 3 and the population selection continues to evolve.

[0071] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A traffic bottleneck identification method based on reverse influence sampling and evolutionary algorithm, characterized in that: include: Acquire the city's traffic history data, wherein the traffic history data includes a city map, traffic flow data and GPS information, and the number of traffic bottleneck areas that need to be selected; The city map is divided into grids to construct a road network. Each grid includes a central area and a road section connected to the central area. The traffic flow of each regional grid of the road network is constructed based on the GPS information and traffic flow data of each grid. The traffic status of the regional grid is further obtained to obtain the vehicle density-congestion formula, and the relationship between congestion and traffic flow is established to describe the diffusion process of congestion on the road network. The reverse influence sampling method is used to establish a proxy model for the traffic bottleneck identification problem. The proxy model uses a two-dimensional fitness value to characterize the quality of candidate solutions, selects the maximum estimated value of the congestion level of the area, and identifies the key traffic bottleneck areas with the goal of minimizing the estimated empirical error.

2. The traffic bottleneck identification method according to claim 1, characterized in that: The traffic bottleneck identification problem is defined as follows: given a road network G, an integer k and the GPS information of several vehicles, the spread of traffic congestion in the past is described by GPS records. The traffic bottleneck identification problem requires finding k areas in the road network that can cause the most congestion in the future.

3. The traffic bottleneck identification method according to claim 1, characterized in that: The GPS information refers to the record of the vehicle's location at a certain time every day provided by the GPS device installed in the vehicle, which is used to correspond the vehicle's flow information to the regional grid described by the road network.

4. The traffic bottleneck identification method according to claim 1, characterized in that: The reverse influence sampling method comprises the following steps: Estimator initialization: use the reverse influence sampling algorithm to initialize the estimator R; Population initialization, using greedy algorithm to initialize the population; Population selection: From N individuals, N / 2 pairs of individuals are selected through the binary tournament selection method to produce new offspring; Population crossover: Apply the crossover operation to the selected N / 2 pairs of individuals. The chromosome of each new offspring includes the intersection of its two parent chromosomes. The remaining positions in the chromosome are randomly sampled from the symmetric difference of the two parent chromosomes to fill in, and finally a offspring population O of size N is obtained; Local search is used to perform local search operations on the offspring population O obtained by crossover to obtain the fitness value of the new offspring population O'; The new populations generated by population evolution; Update the estimator to obtain a new evaluation of the surviving individuals and obtain the fitness value (Inf, error); Algorithm termination judgment: if the algorithm meets the termination condition, that is, the number of iterations reaches the maximum number of iterations or the proportion of a certain genotype in the population P exceeds the given threshold δ, the algorithm terminates and outputs the genotype with the highest proportion in the final population as the selected traffic congestion key node identification result; If the algorithm does not reach the termination condition, that is, the number of evolutionary iterations is less than the maximum number of iterations and the proportion of no genotype in the population exceeds the given threshold δ, the algorithm jumps to the population selection step to continue evolution.

5. The traffic bottleneck identification method according to claim 4, characterized in that: The greedy algorithm is used to initialize the population, specifically: First, a greedy algorithm is used to select 2*k times from the influence propagation sample set called the reverse reachable set stored in the estimator R, and each time a node that can bring the maximum influence gain to the selected node set is selected, that is, the marginal benefit is maximized to ensure that there are enough high-value nodes for initializing the population; Then perform N random samplings from 2*k nodes to construct the initial population, and select k nodes as initial individuals each time; The influence estimator is used to calculate the fitness value (Inf, error) of each initial individual, where Inf represents the influence estimate given by the influence estimator and error represents the empirical error of the estimate.

6. The traffic bottleneck identification method according to claim 4, characterized in that: The local search is used to perform a local search operation on the offspring population O obtained by crossover to obtain the fitness value of the new offspring population O', which is specifically: First, from the current influence estimator, a greedy algorithm is applied to select k nodes to form a set S. Then, each individual in the offspring O is cross-operated with S. The resulting set of new individuals constitutes a new offspring population O' and the current influence estimator is used to calculate the fitness value of the new offspring population O'.

7. The traffic bottleneck identification method according to claim 4, characterized in that: The population evolution adopts non-dominated sorting and crowding distance calculation to obtain the new population generated by this evolution from the parent population P and the offspring O' generated by local search.

8. The traffic bottleneck identification method according to claim 4, characterized in that: In the algorithm termination determination step, the threshold is set to 0.

6.

9. The traffic bottleneck identification method according to claim 5, characterized in that: The initial value of error is 1*10 5 .

Citation Information

Patent Citations

  • Traffic bottleneck prediction method and system based on congestion diffusion and electronic equipment

    CN111325968A

  • Urban road section passage key bottleneck identification method based on congestion propagation

    CN116311892A

  • Method and device for minimizing negative influence of time perception in social network

    CN117876142A

  • Traffic bottleneck detection and classification on a transportation network graph

    US20150081196A1