Dynamic traffic guidance method based on traffic flow prediction under navigation information influence

By constructing a stochastic dynamic traffic network model and a dynamic traffic guidance method based on deep reinforcement learning, the problems of node delays and heterogeneity of driving behavior not being considered in existing technologies are solved, the guidance accuracy and system adaptability of navigation information are improved, and better route recommendations are achieved.

CN121122023BActive Publication Date: 2026-02-17CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511648219.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-17
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Existing dynamic traffic guidance methods ignore node delays, lack the integration and utilization of real-time and predictive information, and fail to consider the heterogeneity of driving behavior, resulting in low guidance accuracy and poor system adaptability.

Method used

A stochastic dynamic traffic network model is constructed, and weight coefficients are introduced to measure driver information dependence. The Logit model is used to calculate the path selection probability. Combined with a dynamic traffic guidance model based on deep reinforcement learning, the path selection strategy is optimized through a dual deep Q network (DDQN). A state-space function and a reward function that integrate node impedance, road segment congestion index and location information are constructed.

Benefits of technology

It significantly enhances the system's ability to model diverse travel behaviors, improves the response accuracy and system adaptability of the route guidance strategy, reduces total delay and average path impedance, and achieves better route recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122023B_ABST
    Figure CN121122023B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent transportation, and discloses a dynamic traffic guidance method based on traffic flow prediction under the influence of navigation information, specifically, a random dynamic traffic network model is constructed, and the time-varying characteristics of OD demand and road traffic flow are quantitatively described; a path travel time perception model under the influence of navigation information is established, and a Logit model is used to calculate the path selection probability; a hybrid traffic distribution model based on dynamic system optimization and dynamic user equilibrium is constructed, and the continuous average method with residual flow updating mechanism is used for iterative solution; a reinforcement learning environment is formed by constructing state space function, action space function and reward function; the model is trained by using DDQN algorithm, and the path selection strategy is optimized. The present application effectively solves the problems of low guidance accuracy and poor adaptability caused by ignoring node delay and lacking information fusion in the existing method, and significantly improves the effect of dynamic traffic guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation technology, and in particular relates to a dynamic traffic guidance method based on traffic flow prediction under the influence of navigation information. Background Technology

[0002] With the continuous growth of motor vehicle ownership, urban road networks have long been under high saturation, leading to increasingly prominent problems such as road overload, bottlenecks at key points, and traffic delays. Traditional traffic guidance strategies based on fixed route selection are no longer applicable in dynamic traffic environments, resulting in poor guidance effects and even triggering a "guided-reverse guided" effect, further exacerbating local congestion. In recent years, route guidance technology based on real-time information has become a core component of Intelligent Transportation Systems (ITS), playing a crucial role in guiding travelers to avoid congestion and improving overall traffic efficiency.

[0003] In related research, classic path selection models such as User Equilibrium (UE) and System Optimum (SO) models are generally based on the assumption of perfect rationality, assuming that travelers can accurately grasp the road network status and choose the optimal route. However, actual driver behavior is often affected by perception errors, navigation dependence, and information lag, resulting in significant deviations between path selection results and model predictions. To address this, scholars have proposed the Stochastic User Equilibrium (SUE) model, which introduces path selection probability and error distribution to simulate travel decision-making behavior under actual perception. However, the SUE model is mostly used for static traffic allocation and is difficult to meet the high-frequency, dynamically changing needs of urban traffic guidance. Dynamic Traffic Assignment (DTA) methods, combined with temporal evolution characteristics, can describe the distribution changes of traffic flow in the spatiotemporal dimensions and have become a key direction in traffic modeling and guidance research in recent years. Some studies have further proposed Dynamic System Optimum (DSO) and Dynamic User Equilibrium (DUE) models to balance system efficiency and individual choice. However, most current DTA models focus on impedance modeling at the road segment level, without fully considering the dynamic feedback effects of node delays, signal control strategies, and traffic fluctuations on route selection, resulting in biases in the application of the models.

[0004] In terms of guiding behavior, navigation systems, as the most direct source of information for travelers, have a decisive impact on route selection due to their accuracy, responsiveness, and feedback update mechanisms. Existing navigation systems typically employ shortest path calculation strategies based on current traffic conditions, lacking the ability to predict future traffic conditions and provide guidance feedback. Especially in urban core road networks with dense traffic signalized intersections and limited node capacity, node delays become a major factor affecting route costs; if left unaddressed, this will significantly reduce guidance accuracy and road utilization efficiency.

[0005] Furthermore, current path guidance methods generally lack modeling of behavioral heterogeneity and information dependence, failing to differentiate the response levels of different drivers to historical experience and real-time information. With the widespread deployment of vehicle-to-everything (V2X) and in-vehicle navigation terminals, vehicles possess the ability to perceive signal phase, road segment status, and target path information; however, there is still a lack of effective mechanisms for embedding this multi-source perception data into traffic guidance decisions. At the algorithmic implementation level, traditional heuristic or rule-based path selection algorithms struggle to adapt to complex and dynamic traffic environments, failing to effectively learn and update guidance strategies. In recent years, with the increasing validation of the effectiveness of reinforcement learning (RL) in traffic scheduling and signal optimization, some studies have attempted to introduce deep reinforcement learning into the field of path guidance; however, a complete system for state definition, delay modeling, and reward design is still lacking, and there are still shortcomings in node impedance and the design of guidance objective functions. Summary of the Invention

[0006] The purpose of this invention is to provide a dynamic traffic guidance method based on traffic flow prediction under the influence of navigation information, so as to solve the problems of low guidance accuracy and poor system adaptability of existing dynamic traffic guidance methods due to ignoring node delays, lack of fusion and utilization of real-time and predicted information, and failure to consider the heterogeneity of driving behavior.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is a dynamic traffic guidance method based on traffic flow prediction under the influence of navigation information, comprising the following steps:

[0008] S1. Construct a stochastic dynamic traffic network model to quantitatively describe the time-varying characteristics of origin-destination (OD) demand and road segment traffic flow, and build a traffic state representation system with spatiotemporal dimensions.

[0009] S2. Establish a route travel time perception model under the influence of navigation information. By introducing weight coefficients to measure the driver's dependence on real-time information and historical experience, update the traveler's perceived travel time, and use the Logit model to calculate the probability of route selection.

[0010] S3. Construct a hybrid traffic assignment model based on dynamic system optimization and dynamic user equilibrium, and use the continuous average method with residual flow update mechanism for iterative solution to achieve traffic flow assignment.

[0011] S4. Based on the traffic network topology defined in S1 and the allocation objective defined in the hybrid traffic assignment model in S3, a dynamic traffic guidance model based on deep reinforcement learning is established. By constructing a state space function that integrates node impedance, road segment congestion index and location information, an action space function that introduces dynamic congestion constraints, and a reward function that comprehensively considers travel time, remaining capacity and goal orientation, a reinforcement learning environment is formed.

[0012] S5. The dynamic traffic guidance model established in S4 is trained using the dual deep Q-network DDQN algorithm. Through experience playback, target network synchronization and exploration probability decay mechanism, the path selection strategy is optimized, and dynamic traffic guidance is finally realized.

[0013] Furthermore, in step S1, the construction of the stochastic dynamic traffic network model specifically includes:

[0014] Define the random travel demand between origin and destination points (OD) and destination (w) within time period k. Random segment traffic flow of segment a Random path traffic on path r and its corresponding mathematical expectation , , And satisfy the flow conservation relationship;

[0015] Assume that the random travel demand, random path flow and random segment flow all follow a log-normal distribution, and that the variance-mean ratio of the random travel demand, random path flow and random segment flow are the same, and that the flow of each path is independent of each other;

[0016] Determining segment impedance using the BPR function And derive the probability distribution of travel time for each road segment;

[0017] Among them, the road section impedance The determination method is as follows:

[0018]

[0019] In the formula: t = kΔt, represents the actual time t corresponding to the k-th time step in the discretized time system, where k = 0, 1, ..., K-1, and K is a set of discrete time intervals. A sufficiently small unit of time time representing discretization; This represents the impedance of road segment a at time t; It refers to the travel time under free-flow conditions on road segment a; It is the traffic flow of road segment a at time t; α represents traffic capacity; α is the congestion sensitivity coefficient; and β is the congestion impact index.

[0020] Furthermore, step S2 specifically includes:

[0021] The perceived travel time is defined as the sum of the actual travel time and the perceived error, wherein the perceived error follows a normal distribution:

[0022] Introducing weighting coefficients ∈[0,1] measures the degree to which a driver relies on real-time information and historical experience;

[0023] Update perceived error based on prediction bias using induced information;

[0024] The Logit model is used to determine the probability of path r being chosen at time t. :

[0025]

[0026] in, R represents the updated expected perceived travel time. w This represents the set of all feasible paths between the origin and destination pairs w.

[0027] Furthermore, the construction of the hybrid traffic assignment model in step S3 includes:

[0028] Construct an optimal traffic assignment model for a dynamic system with the objective of minimizing the total cost of the entire traffic network over time:

[0029]

[0030] Where Y is the total cost of all road segments within the time range [0,T]; T is the time range of the study; This represents the flow status on road segment a from the starting point r to the starting and ending points s at time t; Let A represent the impedance of road segment a at time t; let A represent the set of road segments; and let W represent the set of origin-end point (OD) pairs. The path / segment traffic flow distribution over time is represented; x(t) is the decision variable;

[0031] Constructing a dynamic user-balanced traffic assignment model :

[0032]

[0033] in, It is the traffic flow of road segment a at time t;

[0034] Then, the continuous averaging method iteratively updates the allocation results of the dynamic system optimal traffic assignment model and the dynamic user equilibrium hybrid traffic assignment model, and introduces a residual flow update mechanism to handle unfinished flow. The specific process is as follows: in each iteration, the remaining unfinished flow of the previous period is introduced as the input of the current period.

[0035] Furthermore, the state-space function S mentioned in step S4 t for:

[0036]

[0037] In the formula, Let n be the impedance at time t; Let be the congestion index of road segment a at time t. , The current position One-Hot indication for destination d; A represents the set of road segments; N represents the set of nodes in the transportation network; This represents the sum of the corresponding travel time variables for each road segment;

[0038] The impedance of node n at time t The method for determining this is as follows:

[0039]

[0040] In the formula: The average overflow queue is the total number of vehicles queuing in all lanes. Let be the saturation flow rate of road segment a at time t; The saturation level of the road segment, i.e. ; The intersection phase period; The duration of the green light phase; This represents the actual traffic flow on road segment a at time t.

[0041] Furthermore, the action space function mentioned in step S4 Specifically:

[0042]

[0043] Represents the set of nodes in a transportation network; Indicates the current vehicle location; Indicates the target node; This indicates that at time t, from node The congestion index of the edge to j. The congestion threshold;

[0044] The dynamic congestion constraint is specifically defined as follows: if the congestion index of a certain road segment exceeds the dynamically adjusted congestion threshold, then the path is removed from the action space; the congestion threshold is set to 0.5 during peak hours and 0.35 during off-peak hours.

[0045] Furthermore, the reward function R described in step S4 t Specifically:

[0046]

[0047] in, It is the reward for the total travel time on path p. The reward is based on the percentage of remaining capacity on the path. It is a goal-oriented reward, where ω1, ω2, and ω3 are the weight coefficients of the total travel time reward, the remaining capacity ratio reward, and the goal-oriented reward, respectively.

[0048] Furthermore, in step S5, training the model based on the DDQN algorithm specifically includes:

[0049] S501. Initialize the policy network and target network;

[0050] S502. For each training round, initialize the state, and for each time step t:

[0051] S503, Based on the current state s t And ε-greedy strategy for choosing actions ;

[0052] S504, Execution Action Observe the next state and rewards ;

[0053] S505, Calculate the target Q value ;

[0054] S506, Experience Store in the experience replay pool;

[0055] S507. Randomly sample a sample of size from the experience pool. Calculate the loss function using small batches of data. ;

[0056] Indicates the state s t ,action ,award and the next state The joint expectations This is the Q-value estimation function for the current policy network under parameter θ;

[0057] S508. Update the policy network parameters using gradient descent. ;

[0058] S509, Configure policy network parameters Copy to the target network to update the target network; decay the exploration probability. .

[0059] Compared with existing technologies, this invention has the following advantages: It proposes a dynamic traffic assignment model that integrates the heterogeneity of driving behavior. By introducing individualized features such as navigation dependence and path preference, it significantly enhances the system's modeling ability and response accuracy for diverse travel behaviors. In terms of path guidance strategy solving, a reinforcement learning framework based on Dual Deep Q-Network (DDQN) is designed, effectively alleviating the Q-value overestimation problem in the traditional DQN algorithm and improving the stability and convergence of strategy learning. Furthermore, a three-dimensional reward function mechanism integrating travel time, road resource utilization, and individual goal orientation is constructed, enabling a better balance between system efficiency and user needs in path recommendation. Simulation results show that this method outperforms existing benchmark models in reducing total system delay and average path impedance, and possesses good scalability and practical application potential. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a schematic diagram of the flow conservation of graph theory nodes in an embodiment of the present invention;

[0062] Figure 2 This is a schematic diagram of the vehicle network navigation method in an embodiment of the present invention;

[0063] Figure 3 This is a structural diagram of the dynamic traffic guidance method based on the DDQN algorithm in an embodiment of the present invention;

[0064] Figure 4 This is a schematic diagram of the simulated road network in an embodiment of the present invention; wherein, (a) is the real road network, (b) is the road network, and (c) is the network topology;

[0065] Figure 5 This is a distribution feature map under different individual attribute conditions in the embodiments of the present invention;

[0066] Figure 6 This refers to the system access cost under different penetration rates in the embodiments of the present invention;

[0067] Figure 7 These are the average hourly segment V / C maps for different time periods in the embodiments of the present invention; wherein, (a) is the hourly segment V / C map during off-peak hours, (b) is the hourly segment V / C map during the morning peak hours, (c) is the hourly segment V / C map during the afternoon peak hours, and (d) is the hourly segment V / C map during the evening peak hours.

[0068] Figure 8 These are the training rewards for different OD pairs in the embodiments of the present invention;

[0069] Figure 9 These are training rewards for different OD pairs for path travel time in the embodiments of the present invention;

[0070] Figure 10 The different OD pairs in the embodiments of the present invention are used for training rewards based on path utilization.

[0071] Figure 11 The diagram shows the road density variation under different induction methods based on SUMO simulation; (a) is without control, (b) is controlled by Dijkstra's algorithm, (c) is controlled by the improved Floyd algorithm, and (d) is controlled by the method of this embodiment. Detailed Implementation

[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0073] This embodiment provides a dynamic traffic guidance method based on traffic flow prediction under the influence of navigation information, specifically including the following steps:

[0074] S1. Construct a stochastic dynamic traffic network model; the purpose of this step is to establish a stochastic network model that can characterize the spatiotemporal distribution characteristics of traffic demand under dynamic traffic conditions, to quantitatively describe the time-varying characteristics of origin-destination (OD) demand and road segment traffic flow, thereby constructing a traffic state representation system with spatiotemporal dimensions, and providing a modeling foundation and support for dynamic path guidance strategies based on traffic state prediction.

[0075] Given a directed graph Let N represent a transportation network, where N is the set of nodes. A is the set of road segments. K is a set of discrete time intervals. ,in The unit time interval represents a sufficiently small discretization, while T is the length of the entire study time range. W represents the set of origin-destination (OD) pairs in G, where w is any OD pair. , It is the set of paths between origin and destination points (OD) and road segments w. . For random travel demand between origin and destination points OD and w within time period k; The random traffic flow of road segment a within time period k; This represents the random traffic on the path r between origin and destination points OD and w within time period k. , , Let be the expected origin-destination (OD) demand, the expected flow of a random road segment, and the expected flow of a random path, respectively. Based on the topological relationship between road segments and paths, and according to the law of flow conservation, we have:

[0076]

[0077]

[0078]

[0079] In the formula: This represents the expectation operation; This is a variable relating road segment and path (i.e., whether road segment a is on the path within time period k). If road segment a is located on the path r between the origin and destination points OD and w, the value is 1; otherwise, the value is 0.

[0080] To solve for the covariance between different path flows and the second moment of path flows involved in the path flow modeling process in step S1, this implementation adopts the following behavioral and statistical assumptions:

[0081] Assume a. The origin-destination (OD) demand follows a log-normal distribution, and the random path flow has the same probability distribution as the origin-destination (OD) demand.

[0082]

[0083]

[0084]

[0085] Assumption b: The variance-mean ratio of random OD demand, random path traffic, and random segment traffic. "same.

[0086] Assumption c: The traffic flows of each path within the road network are independent of each other.

[0087] Assume d: The probability that a traveler chooses path r within time period k is... For users traveling from r to destination s during time period k, the route selection... It is real-time information based on time period k. The decision of the path, that is ,in .

[0088] Based on the above assumptions, the standard deviation of OD demand during time period k is... Standard deviation of traffic on path r The standard deviation of the traffic flow on road segment a is... :

[0089]

[0090]

[0091]

[0092]

[0093] `var` represents the variance operator, used to calculate the variance of a random variable;

[0094] The section impedance is represented by the BPR function:

[0095]

[0096] In the formula: t=kΔt, which represents the actual time t corresponding to the k-th time step in the discretized time system, where k=0,1,…,K-1; This represents the impedance of road segment a at time t; It refers to the travel time under free-flow conditions on road segment a; This represents the flow vector of all road segments at time point t. It is the traffic flow of road segment a at time t; α represents traffic capacity; α and β are model parameters, where α is the congestion sensitivity coefficient and β is the congestion impact index, and β is a positive integer.

[0097] In this embodiment, to account for the time-varying nature of road traffic flow, the traffic volume and capacity of road segments are treated as random variables. The travel time of the path will then form a probability distribution:

[0098]

[0099]

[0100] This represents the travel time of a vehicle on road segment a within time period k.

[0101] Assumption e: Traffic volume on the road segment With traffic capacity They are independent of each other.

[0102] Based on the above assumptions, the travel time of road segment a under free-flow conditions within time period k will be... As a deterministic parameter, but considering actual travel conditions, If it should be a random variable, then:

[0103]

[0104]

[0105] Assume f: Road segment capacity Obeying the upper bound is a design capability The lower bound is the uniform distribution of the worst-degraded capacity of the road segment. , This represents a uniform distribution, meaning the probability of the variable taking any value within the interval is equal. This represents the traffic capacity degradation coefficient.

[0106] For uniformly distributed link capacity, The mean and variance are:

[0107]

[0108]

[0109]

[0110] In solving When expecting something, firstly, The random variable Z is standardized to a standard normal distribution, i.e.:

[0111]

[0112] Through standardized transformation, this implementation method will Expressed as:

[0113]

[0114] Will Substituting the expression And expand using the binomial theorem:

[0115]

[0116]

[0117] Since Z follows a standard normal distribution, if i is even, then If i is odd, then .

[0118] Therefore, this implementation simplifies the desired expression to include only the sum of even-numbered terms, with the summation performed only on even-numbered terms i=2j, where j is an index variable, and j is the integer part from 0 to β / 2:

[0119]

[0120] Similarly:

[0121]

[0122] Where ! represents the double factorial operation;

[0123] Based on the above formula, the mean travel time of road segment a within time period k is obtained. With variance for:

[0124]

[0125]

[0126] Therefore, the route travel time variable can be expressed as the sum of the corresponding segment travel time variables:

[0127]

[0128] in, Let w represent the path-segment association matrix, which represents the set of all feasible paths corresponding to OD pair w.

[0129] According to the central limit theorem, the travel time along a path follows a normal distribution. And the following relationship exists:

[0130]

[0131]

[0132] Combining the above formula, we can obtain:

[0133]

[0134]

[0135] Based on the aforementioned assumptions, this implementation method derives mathematical expectation and variance expressions for origin-destination (OD) traffic demand, route flow, and road segment travel time, thereby constructing a computable probability distribution model of route travel time, providing the necessary mathematical modeling foundation for solving subsequent route selection probabilities and optimizing route guidance strategies.

[0136] S2. Modeling of route travel time under the influence of navigation information.

[0137] By using real-time information such as navigation, travelers can obtain their travel time in advance, thus making route selection decisions. Travelers generally choose the route with the shortest path or the shortest travel time. Influenced by real-time information, travelers can accurately perceive the traffic conditions on various sections of the route, thereby finding the most ideal route. In reality, some people choose routes based on their personal travel experience. Therefore, these travelers cannot accurately judge the time-varying traffic flow on road sections and will rely on past travel experience to estimate their perceived travel time. Furthermore, it conforms to a normal distribution, and there is a perception error between the perceived travel time and the actual travel time, and the distribution of the perception error also conforms to a normal distribution.

[0138]

[0139]

[0140]

[0141] This is a variable that associates road segments with paths. If road segment a is located on path r between OD pairs w, the value is 1; otherwise, the value is 0.

[0142] Equations (32) to (34) describe the driver's perception mechanism of travel time during route selection. The perceived travel time consists of the actual travel time and the perception error, which follows a normal distribution. The modeling of this perception mechanism provides the expected value and variance inputs for subsequent route selection probability calculations (Logit model), helping to characterize the driver's response to information perception bias.

[0143] When considering real-time information guidance, the traveler's perceived bias is updated based on the predicted bias of the guidance information, as shown in the equation:

[0144]

[0145]

[0146] In the formula, ∈[0,1] represents the weighting coefficients that depend on real-time information and historical experience; Prediction bias in traffic guidance information; This represents the driver's perceived travel time error for route r in time period k; that is, the deviation between the actual travel time and the perceived travel time, which follows a normal distribution. This indicates the prediction bias of the traffic guidance information system regarding this error.

[0147] Therefore, the updated expected perceived travel time is:

[0148]

[0149] Based on expected perceived travel time, route selection probability The probability of path r being chosen at time t is calculated using a Logit model. Represented as:

[0150]

[0151] Represents the updated expected perceived travel time; R w This represents the set of all feasible paths between the origin and destination pairs w.

[0152] S3: Construct a hybrid traffic assignment model.

[0153] S301, Dynamic System Optimal Traffic Assignment Model

[0154] When providing traffic guidance information, this implementation achieves dynamic system optimization through traffic guidance. The objective of dynamic optimal system allocation is to minimize the total cost of the entire traffic network over time. In this implementation, the dynamic optimal state is the solution to the infinite-dimensional linear programming (ILP) problem, with the objective function being:

[0155]

[0156] st

[0157]

[0158]

[0159] ,

[0160]

[0161] Where st represents the constraint condition, Y is the total cost of all road segments within the time range [0,T], and T is the time range of the study; This represents the flow status on road segment a at time t, from the starting point r to the starting and ending point s. It is the cumulative value of the road segment at the previous time plus the value flowing into the road segment minus the value flowing out of the road segment. This represents the inflow from the starting point r to the starting and ending point s on road segment a at time t; This represents the outflow from the starting point r to the starting and ending point s on road segment a at time t; Let A represent the external demand flow at intersection n, and let B represent the flow that directly enters the node from the outside; A is the set of road segments. This represents a directed road segment with a starting point of n. Indicates the start and end points of a directed road segment. n Sections of road; such as Figure 1 As shown, if node n is the starting point, the total flow out of the source node n should equal the external demand flow generated by that node (i.e., external inflow flow). If node n is both a starting and ending point, the flow flowing into n from other nodes should equal the demand at the ending point. If node n is neither a starting point nor an intermediate intersection of a starting and ending point, it only acts as an intermediate transfer node, responsible only for transferring the incoming flow to the next segment. , represents the travel demand of OD pair (r,s) within time period k, that is, the number of vehicles or traffic flow from the origin r to the destination s, which is used as the input parameter in the dynamic traffic assignment model in this embodiment. This represents the upper limit of the traffic capacity of the OD pair (r,s) on road segment a at time t, i.e., the maximum traffic flow that can pass per unit time. This represents the distribution of traffic flow along a path / segment over time. Let Y be the decision variable, and optimize the objective function Y.

[0162] S302. Construct a dynamic user-balanced traffic allocation model.

[0163] In the Dynamic User Equilibrium (DUE) traffic assignment model, it is assumed that all travelers will choose an optimal path within a given traffic network, such that within a specific time period, no traveler can reduce their travel time by unilaterally changing their chosen path. The DUE model is based on Wardrop's second principle, which states that in equilibrium, all chosen paths have the same travel time, and the travel time of any unchosen path will not be lower than this value. The Dynamic User Equilibrium (DUE) for time period k is:

[0164]

[0165] in, Let represent the traffic impedance function of road segment a when the flow rate is xa(t).

[0166] The constraints of DUE are the same as those in formulas (40)-(43). For travelers on each path, it is assumed that they pass through the road segment according to the First-In-First-Out (FIFO) principle. In equilibrium, for each OD pair w, all paths selected at time t have the same travel time, and the travel time of any unselected path will not be less than the travel time of the selected path. This condition is based on Wardrop's second principle, that is, all selected paths have equal travel time in equilibrium, ensuring that no individual can reduce their travel time by changing their path.

[0167] S303. Construct a dynamic traffic guidance model.

[0168] To achieve dynamic traffic assignment, this implementation constructs an iteratively convergent traffic flow assignment framework based on the Method of Successive Averages (MSA), as shown in Table 1. Considering that traditional MSA algorithms often ignore uncompleted traffic flow that is stuck on road segments due to capacity constraints during peak hours, this invention introduces a "remaining traffic flow" update mechanism on top of the standard MSA. This mechanism records the traffic flow that failed to complete its passage within the current time step in each time period and uses it as input for the next time period, ensuring the continuous transmission of traffic demand over time, thereby improving the model's ability to characterize actual congestion propagation phenomena.

[0169] Table 1 Dynamic Traffic Assignment (MSA) Algorithm

[0170] Continued

[0171]

[0172] in, Let denot be the Dynamic System Optimal (DSO) flow value of segment a at time t in the ξ-th iteration. Let denot be the Dynamic User Balanced (DUE) traffic value of segment a at time t in the ξ-th iteration. This represents the traffic flow of the road segment under the DUE model in the previous iteration; Let N represent the residual unfinished traffic flow of road segment a at time t−1 under the DUE model in the previous iteration, and let N represent the set of nodes in the traffic network.

[0173] S4. Establish a dynamic traffic guidance model

[0174] In the vehicle-to-everything (V2X) environment of Intelligent Transportation Systems (ITS) (see...) Figure 2 The vehicle-mounted terminal can receive the intersection signal phase and remaining green / red light duration in real time, and simultaneously acquire vehicle occupancy or traffic flow information for each road segment. Leveraging this multi-source, high-frequency perception data, the cloud-based navigation platform can complete dynamic path evaluation and replanning within a rolling time window, pushing the optimal driving plan to the driver, thereby balancing network load, reducing local queuing, and improving overall road operating efficiency. To achieve efficient path guidance in complex dynamic traffic networks, this implementation establishes a dynamic traffic guidance model based on deep reinforcement learning. This method optimizes path selection strategies by perceiving traffic conditions in real time, aiming to minimize the number of stops, delays, and total travel time.

[0175] S401. Construct the state-space function.

[0176] In the dynamic traffic guidance method based on deep reinforcement learning, the state space S tIt is the core input for the agent's decision-making, used to describe the current state of the traffic network and the environmental information of the vehicles. Let the traffic network be represented as a directed graph. Where N represents the set of nodes and A represents the set of road segments. For time t, the state vector S t Defined as:

[0177]

[0178] In the formula, Let n be the impedance at time t; Let be the congestion index of road segment a at time t. , The current position One-Hot directions for destination d.

[0179] In actual traffic guidance, the impact of traffic light delays on road network route guidance needs to be considered. Formula (46) combines the load rate of road segments and the saturation flow rate of intersections, and can dynamically adjust the delay estimate of each intersection, thereby more accurately reflecting the delay situation under different traffic flow conditions. This delay estimation method has higher applicability to large-scale traffic guidance and optimization, and can provide more comprehensive decision support for navigation systems.

[0180]

[0181] In the formula: The average overflow queue in vehicles is the total number of vehicles queuing in all lanes. Let be the saturation flow rate of road segment a at time t; The saturation level of the road segment, i.e. ; The intersection phase period; The duration of the green light phase; This represents the actual traffic flow on road segment a at time t, i.e., the number of vehicles passing through the road segment per unit time, and is used to calculate road segment saturation and delay.

[0182] Let the road segment currently selected by the vehicle be... Then the average overflow queue Calculated using the following formula:

[0183]

[0184] In the formula: The flow period is a time interval measured in hours during which the average number of arriving vehicles remains constant. This represents the basic saturation level of the road segment. Below this level, the average overflow queue is approximately zero. .

[0185] Node delay This represents the additional delay incurred when passing node n due to traffic signals, traffic convergence, etc. The specific delay time is calculated using the node impedance function and is expressed in two parts: the first part represents the time increase caused by delays when the road segment is nearing saturation; the second part represents the delay caused by the node under fully saturated conditions. Linking node delays to road segment load rates and the saturation flow rate of the node Connect them. Thus, the delay at each node n... It will be dynamically adjusted according to changes in traffic flow.

[0186] The delay model proposed in formula (46) uses veh as the unit. To ensure that the unit of node impedance is the same as that of road segment impedance, the node impedance model is calculated using the node delay model. .

[0187] Use traffic congestion index Reflects actual travel time Travel time with free flow Increased percentage:

[0188]

[0189] S402, Constructing Action Space Functions

[0190] Action space function Defined as the agent in the current state S t Given a set of selectable paths, the agent optimizes its driving behavior by selecting paths, aiming to minimize the total path cost while satisfying congestion constraints and guidance conditions. At time t, the vehicle is currently located at position... The target node is Then the action space is defined as starting from the current node. Next hop set:

[0191]

[0192] Represents the set of nodes in a transportation network; Indicates the current vehicle location; Indicates the target node; This indicates that at time t, from node The congestion index of the edge leading to j;

[0193] The path set is constrained by the network topology and traffic rules, and the cost function and constraints of each path need to be dynamically evaluated. In dynamic traffic guidance, to avoid the agent selecting severely congested or unreachable paths, the following constraints need to be introduced into action selection: using a congestion index. It represents the real-time congestion situation of a road segment.

[0194]

[0195] If a certain segment of the route satisfy ( (for congestion threshold), so that If the path does not meet the congestion constraint, it should be removed from the action set. The threshold is adjusted based on the characteristics of different cities and road segments or real-time traffic data. A common threshold range is [0.3, 0.5], meaning that when the actual travel time for a road segment increases by more than 30%-50% compared to free-flow travel time, that road segment will be marked as severely congested. In this invention, the threshold is dynamically adjusted according to different time periods. During peak hours, even under free-flow conditions, it is difficult to completely avoid congestion. Setting a higher threshold can prevent the system from frequently misjudging and excluding too many paths. Furthermore, if the threshold is set too low during peak hours, the system may frequently change recommended routes, potentially causing secondary congestion or traffic fluctuations. Considering that users have a higher tolerance for congestion during peak hours, the threshold is adjusted accordingly. Set to 0.5. During off-peak hours, users expect smooth traffic; therefore, when congestion reaches [0.3, 0.4], the system should quickly adjust routes to ensure a high level of service. Furthermore, the low threshold during off-peak hours helps guide vehicles to avoid potential congestion, maintaining the efficient operation of the road network. Therefore, during off-peak hours... Set it to 0.35.

[0196] The agent, based on the current state S t and Q-value function Select the optimal action .use A greedy strategy is used to balance exploration and exploitation. Its mathematical expression is as follows:

[0197]

[0198] In the formula: This represents the action chosen by the agent at time step t. This indicates that the agent is in state S. t The action space below represents the set of all feasible paths from the current node to the target node. This represents the Q-value function, used to evaluate the state S. t Select action The expected reward that can be obtained. ϵExploration probability, representing the agent's tendency to choose random actions, used to avoid getting trapped in local optima.

[0199] During the exploration phase, the agent randomly selects actions with probability ε. This phase helps the agent discover potential optimal paths, especially in complex environments where path states are unknown. In the exploitation phase, the agent uses probability... Choose the action that maximizes the current Q-value function. ,Right now . The value typically decreases gradually during training, transitioning from a high exploration state in the early stages to a high utilization state in the later stages. A common strategy for this dynamic adjustment is:

[0200]

[0201] In the formula, It is the lowest probability of exploration. λ is the initial exploration probability, and λ is the decay rate.

[0202] S403. Constructing the reward space function:

[0203] reward function Used to measure the agent's performance in a specific state. Select action The advantages and disadvantages of different routes can be assessed. By designing a reasonable reward function, the agent can be guided to choose the optimal route, thereby optimizing the overall performance of the transportation network. In this implementation, the reward function considers delay time and route travel time to balance driving efficiency and travel experience.

[0204]

[0205] In the formula: It is the total travel time reward on path p, including the travel time of each segment on the travel path and the delay time of passing through nodes on the path, as shown in formula (54); The percentage reward for remaining capacity of the path is shown in formula (55); The goal-oriented reward is shown in formula (56); ω1, ω2, and ω3 are the weight coefficients of the total path travel time reward, the path remaining capacity ratio reward, and the goal-oriented reward, respectively, which are used to adjust the contribution of different sub-goals in the overall path induction strategy. Their values ​​can be set or optimized according to the training objectives or strategy preferences.

[0206] To avoid instability in the learning algorithm due to significant numerical differences in features across different dimensions within the path travel time estimation function, this invention introduces a normalization factor. The total delay cost function is scaled.

[0207]

[0208] In the formula, and Let be the segment impedance and node impedance on the path. Assume that at each time t, the traveler chooses an alternative route based on the total cost of the current path.

[0209] If the current agent selects a path in the OD pair (r, s), then the proportion of remaining capacity at the path level is defined as follows: Let the path... For a feasible path under OD pair (r, s), the set of all road segments contained in path p is given. Define the remaining capacity percentage of path p at time t as:

[0210]

[0211] Essentially, it measures the percentage of total remaining carrying capacity of path p at the current moment, reflecting the operational "margin" of that path. Based on real-time traffic flow calculations for road sections, it can dynamically respond to traffic pressure caused by sudden traffic surges, signal timing interference, and other factors. Positive incentives are provided through rewards. For higher-speed paths, the model can encourage agents to avoid near-saturation paths, thereby achieving network-level reallocation of traffic load.

[0212] To guide the agent toward the target node and avoid blind detours or deviations from the target during path selection, this invention introduces a guiding reward term into the reward function. Let the current node be... The first-hop target node of action path p is The target node is The shortest path length from a given node to a target node is defined as... The guiding reward item is:

[0213]

[0214] S5. Training the model based on the DDQN algorithm

[0215] This implementation algorithm is based on a Double Deep Q-Network (DDQN) and aims to solve the path optimization problem in dynamic traffic guidance. The DDQN algorithm structure is as follows: Figure 3 As shown, by introducing a dual-network structure (policy network and target network), this algorithm effectively avoids the Q-value overestimation problem in traditional DQN, thereby improving the stability and accuracy of the path selection strategy. In the algorithm, the agent interacts with the traffic environment and selects the optimal path based on real-time and predicted state information to maximize cumulative rewards. The reward function comprehensively considers delay time and total travel time, and optimizes path selection behavior through a negative feedback mechanism.

[0216] DDQN employs Experience Replay technology to store and randomly sample training samples, breaking the temporal correlation between samples and improving the model's training efficiency and generalization ability. Simultaneously, the algorithm introduces... The strategy achieves a dynamic balance between exploration and exploitation, gradually reducing the exploration probability as training progresses, thus transitioning from extensive exploration in the early stages to policy optimization in the later stages. Furthermore, the target network parameters are periodically updated synchronously from the policy network to ensure the stability of the target Q-value calculation. The overall framework combines the dynamic adaptability of reinforcement learning with the complexity of traffic guidance, providing an efficient and scalable solution for real-time traffic management. The DDQN algorithm is detailed in Table 2:

[0217] Table 2 Traffic Dynamic Guidance DDQN Algorithm

[0218] Continued

[0219]

[0220] Indicates the state s t ,action ,award and the next state s t+1 The joint expectation is achieved by randomly sampling multiple interaction samples from the experience replay pool, and is used to calculate the average loss of the model. Here, is the Q-value estimation function of the current policy network under parameter θ; is the Q-value estimation function of the current policy network under parameter θ, used to evaluate the Q-value in state s. t Take action The long-term rewards that can be obtained are derived from this. It approximates the optimal Q-function through a deep neural network, serving as the policy basis for the agent during training and decision-making. `mod` represents the modulo operation; `F` is the update frequency of the target Q-network, meaning the policy network parameters are copied and updated to the target network every `F` steps; `decay rate` is the exploration probability. The attenuation factor is used to control The diminishing rate of random exploration in a greedy strategy.

[0221] S6, Numerical and Case Analysis

[0222] S601, Numerical Case

[0223] This invention employs the Sioux Falls network ( Figure 4 Simulation analysis was conducted, involving 24 nodes and 76 links. All 24 nodes were designated as candidate transfer stations. To better describe the origin-destination (OD) demand of dynamic traffic assignment, this implementation method, based on previous dynamic traffic assignment OD demand, free-flow time, and road capacity, uses a Poisson distribution to describe traffic arrival during nighttime and low-traffic periods, a Weibull distribution to describe peak-hour traffic arrival, and a negative binomial distribution to describe traffic arrival during periods influenced by peak hours. The morning peak hours are 7:00-9:00, the afternoon peak hours are 11:00-13:00, the evening peak hours are 17:00-19:00, and the off-peak hours are 9:00-11:00 and 13:00-17:00.

[0224] S602, Traffic Assignment Results

[0225] Considering weight parameters Exhibiting significant nonlinear characteristics, traditional linear modeling methods struggle to fully reveal the coupling relationships between variables and their combined effect on behavioral weights. Therefore, this invention is based on age (… ), driving experience ( ) and gender ( Three individual attribute variables are used to construct a second-order polynomial regression model including interaction and nonlinear terms to improve the fitting ability to complex behavioral mechanisms. Given the behavioral decision weight parameters... To accommodate the probabilistic interpretation requirements, the implementation further introduces a Sigmoid mapping function to normalize the model regression results, ∈[0,1]. The Sigmoid function possesses excellent mathematical properties, including continuity, monotonicity, and differentiability, facilitating numerical solutions and model training, and also finding wide application in logistic regression and probability prediction. Through this mapping, the final result obtained... The value not only has a clear interpretation space, but also provides theoretical support and computational foundation for subsequent behavior modeling and path selection decisions.

[0226]

[0227] in:

[0228]

[0229] In the formula, Indicates age (unit: years). Indicates driving experience (unit: years). For gender dummy variables (male=1, female=2); coefficients of each term (i=0,1,...,9) are derived from the regression model fitting, and their meanings are as follows: The coefficients of the constant term; , , This represents the linear influence of individual attributes; , , Indicates nonlinear effects; , , It reflects the interactive effects between variables.

[0230] Model setting constraints This means that driving experience must not exceed the age minus the legal driving age.

[0231] Based on the above regression model, the weight parameters for all survey sample individuals are... Calculations are performed, and valid data is selected from samples that meet the logical constraints for analysis. Figure 5 The weight parameters for different genders are shown. The distribution of three-dimensional surfaces related to age and driving experience.

[0232] To simplify individual modeling and enhance parameter interpretability, this invention assumes an equal ratio of male to female travelers in the road network within the given scenario. Different weighting parameters for different genders are also employed. The mean is used as a representative parameter; the average navigation weight parameter for males. The average navigation weight parameter for females is 0.5818. The value is 0.4390. Based on the model, solution algorithm, and Sioux-Falls network proposed in this embodiment, the impact of different time periods and navigation information penetration rates on the total cost of the road network system is studied. Figure 6 and Figure 7Figures (a) to (d) illustrate the dynamic trends of system transportation costs under different navigation penetration rates and time periods. As can be observed from the figures, the total system cost increases significantly during periods of low navigation penetration or peak traffic hours, exhibiting a typical bimodal structure, corresponding to the intensified congestion during morning and evening rush hours, respectively. Conversely, as navigation penetration increases, the system cost decreases, indicating that the navigation system has a good stress-relief effect in guiding travel routes and dispersing traffic flow.

[0233] S603, Training Results

[0234] To study traffic guidance strategies, several typical travel paths were extracted based on simulated travel results during the morning rush hour. These paths correspond to different combinations of origin and destination, such as nodes 2 to 23, 7 to 14, and 4 to 19, reflecting the main travel distribution characteristics of vehicles in the traffic network during this period. This serves as the basis for subsequent path performance evaluation and traffic response optimization. Path 1 represents long-distance commuter routes spanning the east and west ends of the network, utilizing peripheral main roads with higher capacity and lower saturation. Path 2 represents commuter flows from suburban nodes into the city center, reflecting typical "tidal" travel characteristics. Path 3 represents short-distance trips within the city center grid, with most routes concentrated on sections with limited capacity and high saturation.

[0235] The node signal timing uses Webster's formula to calculate the phase period and green light time; no traffic lights are installed at nodes with only two adjacent nodes. The phase sequence adopts a single-segment release strategy, releasing traffic sequentially in a clockwise direction starting from the south of the intersection. The simulation experiment is divided into three scenarios: Scenario 1 is shortest path selection (Dijkstra's algorithm), Scenario 2 is path selection considering road flow (improved Floyd algorithm), and Scenario 3 is a dynamic traffic guidance method considering road traffic conditions (DDQN algorithm).

[0236] To comprehensively evaluate the effectiveness of the proposed dynamic path induction algorithm, this invention analyzes and compares the learning performance of three different OD (Original Path Induction) algorithms over 200 training rounds. Figure 8 It can be seen that the total reward of the three paths fluctuates significantly in the early stages of training, gradually converging as the number of training rounds increases. The shaded area represents the standard deviation of adjacent data. As can be seen from the shaded area, there is significant uncertainty in the choice of each path in the early stages of training because the policy has not yet converged. As learning progresses, the policy gradually stabilizes and the fluctuations decrease, verifying the learning convergence and stability of the DDQN algorithm in dynamic environments.

[0237] Figure 9The training process is shown when minimizing travel time is the objective function. It can be observed that the algorithm experienced significant fluctuations in the early stages, indicating that the agent was in an active exploration phase; subsequently, the amplitude rapidly decreased and entered a stable convergence range. Among the three candidate paths, Path 1 consistently maintained the highest average reward and eventually stabilized, reflecting its superior efficiency in long-distance commuting scenarios across networks; Path 3 was second; and Path 2 had the lowest, corresponding to its network characteristics of needing to traverse multiple moderately congested bottlenecks. The confidence band of the curve narrowed significantly with iteration, indicating that the model can effectively distinguish the performance differences of each path in the travel time dimension and form a stable selection preference, verifying the convergence and discriminative ability of the proposed induced framework under this objective.

[0238] Figure 10 This shows the training results when using a load-balanced reward function. Figure 9 In comparison, the average rewards of the three paths are more similar, and the confidence band remains relatively loose throughout the training process, indicating that the agent tends to distribute traffic evenly among the paths under utilization guidance to avoid saturation of a single path. Although there are also significant fluctuations in the early stages, the algorithm stabilizes after a short exploration phase, achieving the load balancing goal. As can be seen from both figures, the proposed deep reinforcement learning method can quickly identify advantageous paths in travel-time optimization scenarios and maintain fairness in load balancing scenarios, exhibiting good convergence characteristics and adaptability under different objective function settings.

[0239] Figure 11 This demonstrates the spatiotemporal density distribution of 75 observed road segments within a 0–120 min time window in the SUMO microscopic simulation framework, corresponding to uncontrolled ( Figure 11 (a) ), Scenario1 ( Figure 11 (b) ), Scenario2 ( Figure 11 (c) and Scenario3 ( Figure 11 (d) Four guidance strategies. The color gradation increases from blue to red to represent the instantaneous level of lane density per unit area on each road segment. Under uncontrolled conditions ( Figure 11 (a) Multiple instantaneous high-density peaks (orange-red stripes) occurred within the 60–110 min time period, accompanied by congestion waves propagating upstream, especially in lane index intervals 20–35 and 40–55, indicating that queuing was superimposed on multiple bottlenecks and generated a spillover effect. Scenario 1, which introduces static information redistribution, is then used. Figure 11 (b) Although it reduced some of the congestion strips, intermittent orange peaks still occurred in lanes 25–35 during the 70–100 min period; Scenario 2, which further employed a prediction-replanning mechanism, was implemented. Figure 11(c) Further compresses the frequency of occurrence in medium-to-high density areas, yet localized density congestion can still be detected during peak demand periods. In contrast, Scenario3 (which considers the impact of node delays) uses deep reinforcement learning-based adaptive scheduling. Figure 11 (d) The high-density stripes were almost completely eliminated, and the spatial-temporal plane showed a continuous blue-light green distribution, indicating that the congestion wave was effectively suppressed and the flow remained highly balanced among the road segments.

[0240] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0241] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A dynamic traffic guidance method based on traffic flow prediction under the influence of navigation information, characterized in that, The method comprises the following steps: S1, constructing a random dynamic traffic network model, quantitatively describing the time-varying characteristics of OD demand and road traffic flow, and constructing a traffic state representation system with space-time dimensions; S2, establishing a path travel time perception model under the influence of navigation information, introducing a weight coefficient to measure the dependence of the driver on real-time information and historical experience, updating the perceived travel time of the traveler, and calculating the path selection probability by using a Logit model; S3, constructing a hybrid traffic assignment model based on dynamic system optimization and dynamic user equilibrium, and iteratively solving the traffic flow assignment by using a continuous average method with a residual flow updating mechanism; The construction of the hybrid traffic assignment model comprises: Constructing a dynamic system optimization traffic assignment model to minimize the total cost of the entire traffic network within a time range: ; where Y is the total cost of all links in the time horizon [0, T]; T is the time horizon of the study; denotes the flow state on link a from origin r to destination s at time t; denotes the impedance of link a at time t; A denotes the set of links; W denotes the set of OD pairs; denotes the path / link traffic flow distribution over time; x(t) is the decision variable; Constructing dynamic user equilibrium traffic assignment model : ; wherein, is the flow of the link a at time t; Then, the continuous average method iteratively updates the assignment results of the dynamic system optimization traffic assignment model and the hybrid traffic assignment model of the dynamic user equilibrium, and introduces a residual flow updating mechanism to process the unfinished flow, the specific processing process being: introducing the residual unfinished flow of the last time period as the input quantity of the current time period in each iteration; S4, based on the traffic network model defined in S1 and the assignment target defined in the hybrid traffic assignment model in S3, establishing a dynamic traffic guidance model based on deep reinforcement learning, forming a reinforcement learning environment by constructing a state space function integrating node impedance, road congestion index and location information, introducing a dynamic congestion constraint action space function, and considering travel time, residual capacity and target-oriented reward function; The action space function is specifically as follows: Specifically, the action space function is as follows: ; represents a set of nodes in a transportation network; represents a current vehicle position; represents a target node; represents a congestion index of an edge from node to j at time t, is a congestion threshold; A represents a set of road segments; The dynamic congestion constraint is specifically: if the congestion index of a road exceeds the dynamically adjusted congestion threshold, the path is excluded from the action space; the congestion threshold is set to 0.5 during peak hours and 0.35 during non-peak hours; S5, training the dynamic traffic guidance model established in S4 by using a double deep Q network (DDQN) algorithm, optimizing the path selection strategy through experience replay, target network synchronization and exploration probability decay mechanism, and finally realizing dynamic traffic guidance.

2. The method of claim 1, wherein, In the step S1, the construction of the random dynamic traffic network model specifically comprises: Define the random travel demand between OD pairs w in the k-period between origin and destination points , the random link flow on link a , the random path flow on path r , and its corresponding mathematical expectation , , , and satisfy the flow conservation relationship; It is assumed that the random travel demand, random path flow and random road flow all obey a lognormal distribution, and the variance-mean ratio of the random travel demand, random path flow and random road flow is the same, and each path flow is independent; Determining link impedance by a BPR function and deriving a probability distribution of link travel times; where the link impedance The determination method is as follows: ; where t=kΔt represents the actual time t corresponding to the kth time step in the discretized time system, where k=0, 1, …, K−1, K is a set of discrete time periods, represents a sufficiently small unit time period that is discretized; represents the impedance of link a at time t; is the travel time of link a in free flow state; is the flow of link a at time t; is the capacity; α is the congestion sensitivity coefficient, and β is the congestion influence index.

3. The method of claim 1, wherein the traffic flow prediction is based on the navigation information. The step S2 specifically comprises: Defining the perceived travel time as the sum of the actual travel time and the perceived error, and the perceived error obeys a normal distribution: Introducing weight coefficients ∈ [0, 1] measures the degree of dependence of the driver on real-time information and historical experience; Updating the perceived error based on the prediction bias of the guidance information; A Logit model is used to determine the probability of route r being selected at time t : ; wherein, denotes the updated expected perceived travel time, R w denotes the set of all feasible paths between the origin-destination pair w.

4. The method of claim 1, wherein the traffic flow prediction is based on the navigation information. The state space function S in step S4 t is: ; wherein is the impedance of node n at time t; is the congestion index of link a at time t, , are One-Hot indications of the current location , the destination d, respectively; A denotes the set of links; N denotes the set of nodes in the traffic network; denotes the sum of the respective link travel time variables; impedance of the node n at the time t The determination method is: ; wherein: is the average spillback queue, i.e. the total number of vehicles queuing on all lanes; is the saturation flow rate of link a at time t; is the saturation degree of the link, i.e. ; is the phase cycle of the intersection; is the duration of the green phase; denotes the actual flow rate on link a at time t.

5. The method of claim 1, wherein, The reward function R in step S4 t Specifically: ; wherein, is the total travel time reward on path p, is the path remaining capacity proportion reward, is the target-oriented reward, and ω1, ω2, ω3 are the weight coefficients of the total travel time reward, the path remaining capacity proportion reward, and the target-oriented reward, respectively.

6. The method of claim 1, wherein, In the step S5, the model training based on the DDQN algorithm specifically comprises: S501, initializing the policy network and the target network; S502, for each training round, initializing the state, and for each time step t: S503、According to the current state s t and epsilon-greedy policy to select actions ; S504, performing an action , observing the next state and reward ; S505、Calculate target Q value ; S506, store the experience into the experience replay pool; S507, randomly sample a mini-batch data with size from the experience pool, compute the loss function ; represents the joint expectation of the state s t , action , reward , and next state , Q-value estimation function of the current policy network at parameters θ. S508, update the policy network parameters by gradient descent method ; S509、updating the target network with the policy network parameters copying to the target network to update the target network; decaying the exploration probability .

Citation Information

Patent Citations

  • Electric vehicle charging demand prediction method based on traffic balance

    CN119514798A

  • Cloud side-end integrated collaborative digital and intelligent traffic collaborative management and control method, system, equipment and medium

    CN119541202A