Expressway dynamic toll collection method, device, equipment and storage medium

By designing a simulation environment and a traveler route selection model in a real highway network, and using reinforcement learning algorithms to optimize the toll collection strategy, the problem of insufficient adaptability of existing dynamic toll collection models for highways is solved, and congestion is effectively alleviated and revenue is increased.

CN116227215BActive Publication Date: 2026-05-05YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNNAN UNIV
Filing Date
2023-03-15
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing studies on dynamic toll collection on expressways mainly focus on urban road networks, failing to effectively adapt to the characteristics of my country's expressway network. Travelers face difficulties in choosing routes, and the effectiveness of existing models in real traffic network environments has not been verified.

Method used

We designed a realistic simulation environment of China's highway network and a traveler route selection model, and used reinforcement learning algorithms for dynamic toll collection. By establishing a highway network simulation environment model, a traveler route selection model, and agent-based toll rate adjustments, we optimized the toll collection strategy to alleviate congestion.

Benefits of technology

It effectively alleviates congestion in a real highway network environment, increases benefits for all parties, adapts to travelers' route selection behavior, and optimizes toll collection strategies to improve revenue and traffic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116227215B_ABST
    Figure CN116227215B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, equipment, and storage medium for dynamic toll collection on highways. The method includes: establishing a highway network simulation environment model and determining the system state of a reinforcement learning model; establishing a traveler route selection model and determining the system strategy of the reinforcement learning model based on the system state and the traveler route selection model; determining the system action for adjusting the toll rate for each toll segment based on the system strategy; calculating the system reward after executing the current system action and continuously adjusting the system strategy of the reinforcement learning model based on the system reward; and determining the optimal toll rate adjustment scheme for each toll segment when the system reward reaches its maximum. This invention utilizes a reinforcement learning algorithm for dynamic toll collection in a designed real Chinese highway network simulation environment and a traveler route selection model to alleviate congestion and improve benefits for all parties involved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning technology, and in particular to a method, apparatus, equipment, and storage medium for dynamic toll collection on highways. Background Technology

[0002] The dynamic toll collection scheme implemented in my country's expressways is a differentiated toll collection method: based on payment method, vehicle type, time period, direction, and road segment. The "dynamic" nature of the toll collection is limited and cannot well adapt to the real-time dynamic changes of expressways.

[0003] Congestion pricing refers to charging travelers a certain fee on certain roads or areas during peak traffic hours on city roads that are normally toll-free, in order to alleviate congestion. It is essentially a traffic demand management measure. Congestion pricing is generally divided into two categories: static pricing and dynamic pricing. Static pricing only considers the spatial dimension and does not consider the time-varying nature of the system, ignoring the impact of the current pricing on the road network. Dynamic pricing, on the other hand, considers both time and space dimensions, and different road sections charge different fees at different times [8]. Dynamic tolling on highways is essentially an extension of dynamic congestion pricing on Chinese highways.

[0004] Dynamic tolling is a tolling scheme that adapts to traffic conditions and dynamically adjusts the toll rate based on congestion pricing. Works by Joksimovic et al. (2005) and Lu et al. (2008) have explored dynamic tolling. Zhang et al. (2013) proposed a traffic dynamics-based method, but this method assumes that traffic demand between different nodes is fixed, which is an overly idealistic assumption. Later, ∆-tolling (2017) implemented dynamic tolling based on real-time traffic flow, but because this method also does not actively consider traffic demand in the model, it may lead to the model failing to achieve optimal performance.

[0005] The first application of reinforcement learning to the dynamic tolling problem was by Chen et al. (2018), who formalized the traffic dynamics problem as a Markov Decision Process (MDP) and proposed a dynamic model, DyETC, based on the PG-β algorithm. This model significantly improved upon the problems of the ∆-tolling method and achieved good results in alleviating traffic congestion. Although DyETC performed well, it only worked on road networks with 11 zones and could not be scaled to large-scale networks. To enable the dynamic tolling model to robustly scale to larger problems, Qiu et al. first divided the entire road network into different zones based on geographical and economic characteristics, and then used a multi-agent reinforcement learning algorithm to train a tolling agent for each zone, proposing the DPG-β algorithm, which achieved better performance and scalability. Furthermore, Pandey et al. (2020) used existing reinforcement learning algorithms to study the dynamic tolling problem of managed lanes in the United States.

[0006] The aforementioned studies mainly focused on urban roads with dense road networks, and only implemented toll collection on certain sections during peak hours. These studies are not well adapted to the characteristics of my country's expressway network and the route selection behavior of travelers on it.

[0007] Existing domestic and international research on dynamic road tolling mainly falls into two categories: one is research on congestion pricing in urban roads both domestically and internationally, and the other is research on determining toll rates. While the research findings are abundant, the following issues still exist:

[0008] (1) Current research mainly focuses on urban road networks at home and abroad. These networks consist of toll roads during congestion periods and adjacent free roads. Urban road networks are dense, making it easier for travelers to change their routes due to tolls, and the distance difference between different route choices to reach the destination is limited. However, the density of my country's expressway network is much lower than that of urban roads, and all sections are tolled. Once travelers enter the expressway network, it is more difficult for them to change their routes. The two have significant differences in characteristics, making it difficult for existing research results on congestion tolls to adapt to the actual situation of dynamic tolling on my country's expressways.

[0009] (2) Based on the above-mentioned characteristics of urban road networks at home and abroad, existing dynamic toll collection schemes based on reinforcement learning assume that travelers at intersections only make route selections based on the current travel time and toll fees received, that is, they choose toll roads or free roads based on the incurred fees and travel time to weigh time. Due to the characteristics of China's expressway network described in (1), for travelers there, route selection often needs to consider the prior factor of the distance to the destination from different routes, so the assumptions in the above-mentioned existing studies are obviously not applicable.

[0010] (3) Existing dynamic toll collection schemes for Chinese expressways are mostly limited to model optimization design and theoretical analysis. The model verification is also mostly carried out on virtual road networks, and its effectiveness in real traffic network environments needs to be verified. Summary of the Invention

[0011] The main objective of this invention is to provide a method, apparatus, device, and storage medium for dynamic toll collection on highways. The aim is to solve the problem of using reinforcement learning algorithms for dynamic toll collection in a simulated environment of the real Chinese highway network and a traveler's route selection model, so as to alleviate congestion and improve the benefits for all parties.

[0012] To achieve the above objectives, the present invention provides a dynamic toll collection method for highways, the method comprising the following steps:

[0013] Establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model;

[0014] Establish a traveler route selection model, and determine the system strategy of the reinforcement learning model based on the system state and the traveler route selection model;

[0015] Based on the system strategy, determine the system actions for adjusting the toll rate for the corresponding intelligent agent of each toll road segment;

[0016] Calculate the system reward after executing the current system action, and continuously adjust the system policy of the reinforcement learning model based on the system reward;

[0017] When the system reward reaches its maximum, determine the optimal toll rate adjustment plan for each toll road segment.

[0018] Optionally, the steps for establishing a high-speed road network simulation environment model include:

[0019] Establishing a high-speed network of ternary groups , Let be a finite non-empty set, representing the set of nodes in the highway network. The number of nodes; It is the set of edges between adjacent nodes; Adjacency matrices represent the connectivity information between two nodes in a highway network;

[0020] Define each time step The set of all time steps T is Integer multiples of; defines the time interval for the system action of the agent to adjust the rate. , A fixed value determined by the fee management department; defining the time series set for fee rate adjustments. .

[0021] Optionally, the steps of establishing a high-speed road network simulation environment model also include:

[0022] Use the mean as Standard deviation is Modeling modified Gaussian random variables at time step From the source node to the destination node Traffic demand;

[0023] Using Daganzo's CTM model to model traffic flow, each road in the road network... Divide into independent sub-segments; use Indicates that it belongs to a road section The set of all sub-segments, For each Its length is used This indicates that the distance traveled by a vehicle within a time step according to its free-flow velocity is determined; at each time step, the traffic state of a road segment is composed of the number of vehicles located in each sub-segment; using... , , , Separate road sections The length of the road, the free-flowing speed of vehicles on the road segment, the capacity, and the density when congestion occurs.

[0024] Optionally, the steps in establishing a traveler route selection model include:

[0025] Establish a traveler route selection model based on the decision-making route model;

[0026] The traveler route selection model determines the distance and utility function value of different routes when the traveler selects a route at the toll station entrance and at the intersection while driving, and selects the route with the shorter distance or the larger utility function value.

[0027] The expression for the utility function value is:

[0028] ;

[0029] in, A path segment from the current node to the destination node. Let be the set of all road segments traversed from the current node to the destination node, and let represent a path from the current node to the destination node. It is the set of all distinct paths from the current node to the destination node. The total toll for a given route is given by This is approximated as the time required to traverse each road segment in the current state. For road section The current total number of vehicles;

[0030] The expression for the traveler's route selection model is:

[0031] ;

[0032] in, The function is used to calculate the first... The total physical distance of the path.

[0033] Optionally, the system state includes time steps. Traffic flow in the sub-segments under this update step The system state is described by the equation: System rewards include revenue, system uptime, and network throughput; among which:

[0034] The system strategy for continuously adjusting the reinforcement learning model based on system rewards is as follows:

[0035] ;

[0036] ;

[0037] ;

[0038] ;

[0039] in, Indicates from time step Time to step During the time period, from the road section Move to the section Total traffic volume Indicates at time step At that time, it refers to the total number of vehicles that exit the highway in the high-speed network.

[0040] Optionally, the system policy step of continuously adjusting the reinforcement learning model based on system rewards also includes:

[0041] Measure the maximum difference between the number of vehicles in different lanes in all time steps. Specifically:

[0042] ;

[0043] in, Indicates roads in different directions;

[0044] Define a network-independent statistic. Specifically:

[0045] ;

[0046] ;

[0047] use and Describe the operational status of the road network in multiple dimensions.

[0048] Optionally, the system policy step of continuously adjusting the reinforcement learning model based on system rewards also includes:

[0049] The definitions measure the proportion of speeding below the minimum speed limit on roads in different directions and the total speeding on both roads, respectively.

[0050] ;

[0051] ;

[0052] ;

[0053] in, and All are indicator variables, when time step Shizi Road Section The number of vehicles on the road is higher than When the required number of vehicles is Set to 1 if the time step is 1, otherwise set to 0; Shizi Road Section When the vehicle's speed exceeds the maximum speed limit allowed on the highway Select 1 if the value is 1, otherwise select 0.

[0054] The values ​​%-violationRight, %-violationLeft, and %-sumOverSpeedviolation are used to measure the number of times a vehicle violates speed limits while driving.

[0055] Furthermore, to achieve the above objectives, the present invention also provides a dynamic toll collection device for highways, the dynamic toll collection device comprising:

[0056] The first module is used to establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model.

[0057] The second module is used to build the traveler's route selection model and determine the system strategy of the reinforcement learning model based on the system state and the traveler's route selection model.

[0058] The first determining module is used to determine the system action for adjusting the rate of each toll road segment by the corresponding intelligent agent, based on the system strategy.

[0059] The computation module is used to calculate the system reward after executing the current system action, and to continuously adjust the system strategy of the reinforcement learning model based on the system reward.

[0060] The second determining module is used to determine the optimal toll rate adjustment plan for each toll road segment when the system reward reaches its maximum.

[0061] In addition, to achieve the above objectives, the present invention also provides a dynamic toll collection device for highways, the device comprising: a memory, a processor, and a dynamic toll collection program for highways stored in the memory and executable on the processor, wherein the dynamic toll collection program for highways, when executed by the processor, implements the steps of the dynamic toll collection method for highways described above.

[0062] In addition, to achieve the above objectives, the present invention also provides a storage medium storing a dynamic toll collection program for highways, which, when executed by a processor, implements the steps of the dynamic toll collection method for highways described above.

[0063] This invention proposes a dynamic toll collection method, apparatus, equipment, and storage medium for highways. The method includes: establishing a highway network simulation environment model and determining the system state of a reinforcement learning model; establishing a traveler route selection model and determining the system strategy of the reinforcement learning model based on the system state and the traveler route selection model; determining the system action for adjusting the toll rate for each toll segment based on the system strategy; calculating the system reward after executing the current system action and continuously adjusting the system strategy of the reinforcement learning model based on the system reward; and determining the optimal toll rate adjustment scheme for each toll segment when the system reward reaches its maximum. This invention utilizes a reinforcement learning algorithm for dynamic toll collection in a designed real Chinese highway network simulation environment and a traveler route selection model to alleviate congestion and improve benefits for all parties. Attached Figure Description

[0064] Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention;

[0065] Figure 2 This is a flowchart illustrating an embodiment of the dynamic toll collection method for expressways according to the present invention.

[0066] Figure 3 A diagram illustrating the reinforcement learning process;

[0067] Figure 4 This is a simplified diagram of a highway network structure.

[0068] Figure 5 A time-based relationship diagram;

[0069] Figure 6To design a simulated road network structure diagram for human beings;

[0070] Figure 7 This is a model architecture diagram of the dynamic toll collection method for highways according to the present invention;

[0071] Figure 8 This is a structural block diagram of a dynamic toll collection device for highways according to an embodiment of the present invention.

[0072] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0073] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0075] like Figure 1 As shown, Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention.

[0076] like Figure 1 As shown, the device may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0077] Those skilled in the art will understand that Figure 1 The structure of the device shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0078] like Figure 1As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a highway dynamic toll collection program.

[0079] exist Figure 1 In the terminal shown, network interface 1004 is mainly used to connect to the backend server and communicate with it; user interface 1003 is mainly used to connect to the client (user terminal) and communicate with it; while processor 1001 can be used to call the highway dynamic toll collection program stored in memory 1005 and perform the following operations:

[0080] Establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model;

[0081] Establish a traveler route selection model, and determine the system strategy of the reinforcement learning model based on the system state and the traveler route selection model;

[0082] Based on the system strategy, determine the system actions for adjusting the toll rate for the corresponding intelligent agent of each toll road segment;

[0083] Calculate the system reward after executing the current system action, and continuously adjust the system policy of the reinforcement learning model based on the system reward;

[0084] When the system reward reaches its maximum, determine the optimal toll rate adjustment plan for each toll road segment.

[0085] The specific embodiments of the present invention applied to the device are basically the same as the embodiments of the application of the dynamic toll collection method for highways described below, and will not be repeated here.

[0086] This invention provides a method for dynamic toll collection on highways, referring to... Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the dynamic toll collection method for highways according to the present invention.

[0087] In this embodiment, the dynamic toll collection method for highways includes the following steps:

[0088] Step S100: Establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model;

[0089] Step S200: Establish a traveler route selection model and determine the system strategy of the reinforcement learning model based on the system state and the traveler route selection model.

[0090] Step S300: Based on the system strategy, determine the system action for adjusting the rate for each toll road segment corresponding to the intelligent agent.

[0091] Step S400: Calculate the system reward after executing the current system action, and continuously adjust the system policy of the reinforcement learning model based on the system reward;

[0092] Step S500: When the system reward reaches its maximum, determine the optimal rate adjustment plan for each toll road segment.

[0093] This embodiment provides a dynamic toll collection method for highways. By using a reinforcement learning algorithm in a simulated environment of the real Chinese highway network and a traveler route selection model, dynamic toll collection is carried out to alleviate congestion and improve the benefits for all parties.

[0094] To explain this application more clearly, specific examples of its practical application are provided below.

[0095] (1) For the construction of the road network simulation environment:

[0096] The Cell Transmission Model (CTM) is used to describe the evolution of cellular velocity states and was first proposed by Daganzo. In the CTM model, each road segment is divided into many sub-segments, called "cells," and time is discretized into time steps. Then, traffic flow is transferred between sub-segments using traffic conservation and traffic wave principles, which can dynamically describe road traffic flow.

[0097] (2) For reinforcement learning algorithms:

[0098] Reinforcement learning (RL) is a Markov decision process (MDP) in which an agent interacts with its environment and learns the optimal policy through trial and error. The reinforcement learning process can be represented as follows: Figure 3 .

[0099] according to Figure 3 In reinforcement learning, RL can generally be represented as a quintuple:

[0100] ;

[0101] in For the system state set, For the action set of the intelligent agent, Let be the state transition probability. The reward set obtained after performing an action. This is a mapping from the state space to the action space. The algorithm's process is as follows: the agent is in the state... According to the strategy Select Action Rewards for interacting with the environment Then, the system uses The probability of transitioning to state .

[0102] Deep learning (DL) uses multi-layered network structures and non-linear transformations to combine low-level features into abstract, easily distinguishable high-level representations, thereby discovering distributed feature representations of data. The combination of DL and reinforcement learning (RL) is called deep reinforcement learning (DRL), which uses DL to learn abstract features of the input data and then uses RL to optimize the solution policy based on these features. DRL can be broadly divided into two categories: value function-based methods and policy gradient-based methods.

[0103] Policy gradient descent (PRD) is a method for policy optimization that updates policy parameters by calculating the gradient of the expected reward of the policy with respect to the policy parameters. In Dynamic Toll Collection (DRL), the policy is typically represented by a Discrete Neural Network (DNN) and parameterized with parameters before the policy gradient method is used to optimize the policy. Compared to value function-based DRL, policy gradient-based DRL effectively addresses the limitation of Discrete QN (DQN) outputs of discrete state-action values ​​that are not applicable to continuous action spaces. It is more direct, effective, and has a wider range of applications. Therefore, this paper selects this algorithm as the adjustment algorithm for the dynamic toll collection model of highways.

[0104] This embodiment is based on the comparison of reinforcement learning models, and adaptively rewrites the traditional feedback control heuristic algorithm, applying it to road network environments and lane selection models.

[0105] In light of the actual characteristics of my country's expressway network, this embodiment modifies the aforementioned heuristic algorithm of Pandey and Boyles' density heuristic: First, define... The set represents travelers entering the toll road. Then proceed to the next possible road segment. For example... Figure 4 As shown, , ...and specify the set in the definition. The space of all possible next-path segments is mutually exclusive and finite. That is:

[0106] ;

[0107] In the experiment, it was assumed that the heuristic feedback control algorithm was based on The traffic flow density values ​​of each future road segment in the dataset are used to update the toll rate or dynamic adjustment coefficient for that road segment. The time is based on the previous time. The value is updated by adjusting the difference between the required traffic flow for that road segment and the current traffic flow. Specifically, the update process is as follows.

[0108] ;

[0109] in, yes exist The total number of vehicles on this road segment is updated in real time before the dynamic adjustment coefficient is updated. It is the expected value of the number of vehicles. This is a parameter used to adjust the difference; the model dynamically adjusts the coefficient update by controlling the difference between the current number of vehicles and the desired number of vehicles. Generally, the desired value is... The number of vehicles corresponding to the critical density of a road segment in the set, i.e.:

[0110] ;

[0111] in, Indicates road segment Traffic density along the street, It is an adjustment parameter between 0 (not 0) and 1, meaning a certain proportion of the traffic flow expected to be at the critical density.

[0112] (3) For the construction of the high-speed road network simulation environment:

[0113] In this embodiment, the highway network is abstracted as a triple. finite nonempty set This represents the set of nodes in the highway network. The number of nodes is represented by the number of toll stations (including toll station exits and entrances, and interchanges) in the road network structure of this paper. It is the set of edges (path segments) between adjacent nodes. The adjacency matrix describes the connectivity information between two nodes in a highway network.

[0114] In a road network environment, each time step The set of all time steps is represented as T is Integer multiples of the rate. The rate applies every [number] times [the specified number]. Updated once, in which A fixed value determined by the fee management department. Define the time series set for rate updates. Its index increases in positive integers, then, .by and For example, its corresponding , The relationship represented by different times is as follows: Figure 5 As shown.

[0115] In a road network environment, to simulate the operation of a high-speed system, this embodiment uses the mean value as... Standard deviation is Modeling modified Gaussian random variables at time step From the source node to the destination node Traffic demand is calculated, ignoring the correlation between different source-destination node pairs (OD pairs) and time. The mean demand can be estimated by observing historical lane data.

[0116] The road network environment simulates the propagation process of traffic flow within a road network and continuously updates the road network status (i.e., traffic flow on the roads) over time. Daganzo's CTM model is used to model the traffic flow, representing each road in the road network. Divide into independent sub-road segments. (Use...) Indicates that it belongs to a road section The set of all sub-segments, For each Its length is used This is represented by the distance a vehicle travels within a time step at free-flow speed (the speed at which vehicles travel under low traffic volume and without traffic control). At each time step, the traffic state of a road segment is determined by the number of vehicles located in each sub-segment. Here, we use... , , , Separate road sections The length of the road, the free-flowing speed of vehicles on the road segment, the capacity, and the density when congestion occurs.

[0117] (4) For the construction of the route selection model:

[0118] Based on the actual travel conditions of travelers on Chinese highways, travelers need to choose routes when entering toll stations and when encountering intersections during their journey. For example:

[0119] for Figure 6The OD pair "17-20" indicates that the vehicle enters the highway from toll station 15 and exits from toll station 2 (nodes "17" and "20" are virtual nodes for model processing convenience, representing the entrance at toll station "15" and the exit at toll station "2" respectively). During the journey, the vehicle first needs to choose a route at interchange node 13, and at each exit along the way, it can choose to continue or exit the highway network from a fork depending on whether its destination node matches the exit node (for example, at node "8", all vehicles traveling to the left have two choices: 1. If the destination node is "19", follow the road...). Exit the highway. 2. Otherwise, continue along the current road to the next node "4" and then make another route selection.

[0120] At each highway entrance node or intersection node in the road network, unlike the route selection model in previous studies on dynamic toll collection for urban roads which "assume travelers choose routes solely based on received real-time information (current travel time and cost)," this embodiment argues that travelers' route selection behavior on highways is related to the prior factor of the distance to the destination via different routes, and the estimated cost and travel time accordingly. This is more consistent with the route selection habits of travelers on highways with sparse road networks and large differences in the distance between different routes. In fact, in modern travel modes, obtaining prior knowledge of the distance of different routes is not difficult (e.g., with the help of navigation software and other technical means). Research has shown that the decision route model has the smallest error; therefore, this paper chooses the decision route model as the basis for the route selection model. In the route selection model, when travelers choose routes at toll station entrances and intersections while traveling, they simultaneously judge the distance of different routes and the utility function value of different routes, choosing the route with the shorter distance or the larger utility function value. Utility function value calculation:

[0121] ;

[0122] in, A path segment from the current node to the destination node. Let be the set of all road segments traversed from the current node to the destination node, and let represent a path from the current node to the destination node. It is the set of all distinct paths from the current node to the destination node. The total toll for a given route is given by ( For road section The current total number of vehicles is approximated as the time required to traverse each road segment in the current state. In summary, the route selection model based on prior knowledge can be expressed as:

[0123] ;

[0124] in The function is used to calculate the first... The total physical distance of the path.

[0125] It should be noted that the dynamic toll collection problem on highways is modeled as a partially observable Markov Decision Process (POMDP) ​​to solve the problem using reinforcement learning algorithms.

[0126] (5) For reinforcement learning:

[0127] The following variables are defined:

[0128] Update timestep: every fixed time step Optimize and update the joint charging coefficient.

[0129] States: Definition Indicates at time step Shizi Road Section The number of vehicles on board. Similar to the "non-atomic flow assumption" commonly used in transportation literature, this assumes... It is a non-negative real number, not an integer. The update step uses the joint toll coefficient. Traffic flow in the sub-segments under this update step This indicates the current system state; here, it is allowed. Compare Big, that is This reduces the size of the state space and improves computational efficiency. The state space can be described by the following equation.

[0130] ;

[0131] Observation: The model uses loop detectors to measure the total number of vehicles moving from one sub-segment to the next. The observation vector representing the state includes each road segment. The total number of vehicles on the road. The observation can be expressed by the following equation:

[0132] ;

[0133] Here we assume that the actual observation has a standard deviation of . A Gaussian random variable with a specified mean. Values ​​less than zero are mapped to 0.

[0134] Transition function: given an action POMDPs from state Transition to state The process is controlled by the traffic flow equations of the CTM model, which are related to the route selection behavior of travelers in the highway network.

[0135] Reward: In the state Take action at the time The reward obtained by the system afterward is represented as The value of this reward depends on the choice of optimization objective. This embodiment first considers three single optimization objectives: 1. Revenue Maximization (RevMax); 2. Total System Travel Time Minimization (TSTTMin); 3. System Throughput Maximization (ThptMax).

[0136] It is important to emphasize that the revenue and total running time here are different from the estimated time and cost in the route selection model. Here, revenue and time represent the cost and travel time incurred by all the road segments that vehicles have already passed.

[0137] The rewards for different optimization objectives are defined as follows:

[0138] Revenue maximization (RevMax):

[0139] ;

[0140] in Indicates from time step Time to step During the time period, from the road section Move to the section Total traffic volume.

[0141] Total system travel time minimization (TSTTMin):

[0142] ;

[0143] The minus sign "-" is used to ensure the reward. Maximizing is equivalent to minimizing the total system travel time.

[0144] Throughput maximization (ThptMax):

[0145] ;

[0146] in Indicates at time step At that time, it refers to the total number of vehicles that exit the highway in the high-speed network.

[0147] In reality, for highway operators, the goal is often not just a single objective like revenue or highway network efficiency, but rather Pareto optimality for revenue, system uptime, and network throughput. This embodiment uses a single policy approach to achieve joint optimization of multiple objectives. This method transforms the multi-objective problem into a single objective by defining a weighted combination of multiple objectives, making it relatively simple and direct. Specifically, a single policy is defined using weighted summation to jointly optimize revenue, total system uptime, and throughput. A joint reward function is defined. It is a linear combination of three reward functions.

[0148] ;

[0149] in, It represents the relative weights of revenue and throughput with respect to the total system uptime. Physically, it represents the slope of the straight line on a graph with the sum of revenue and throughput and the total system uptime as the coordinate axes.

[0150] For dynamic tolling problems, under the objective of maximizing revenue, the jam and harvest (JAH) phenomenon often occurs. This means that in the simulation, toll-free or low-toll sections experience congestion earlier, while in the later stages of the simulation, to attract more vehicles to choose toll (or higher-toll) sections in order to generate more revenue. Therefore, to measure the performance of the model in adjusting the road network state, this embodiment also defines the following variables to display the operational status of the road network in multiple dimensions.

[0151] First, we define two statistics to measure the JAH phenomenon. : Measure the maximum difference between the number of vehicles in different lanes in all time steps, defined as follows:

[0152] ;

[0153] in Indicates roads in different directions.

[0154] In addition, a network-independent statistic was defined. First, define As shown in the formula:

[0155] ;

[0156] This formula represents the difference between the current number of vehicles on all sub-segments of a road in different directions and the maximum allowed number of vehicles on that sub-segment (corresponding to the vehicle density during congestion). It signifies the difference in occupancy rates on roads in different directions. It can then be defined as all time steps The maximum value. The value range is [-1, 1]. When it takes a positive value, the larger the value, the more likely it is to represent a positive value. Traffic congestion on the roads in that direction is relatively high The more serious the situation becomes, the more difficult it is to navigate the road. Represented as follows:

[0157]

[0158] Furthermore, to measure speed limit violations during vehicle operation, the model defines %-violationRight, %-violationLeft, and %-sumOverSpeedviolation to measure the proportion of speeding below the minimum speed limit on different roads and the total speeding rate across both roads, respectively. Their specific definitions are as follows:

[0159] ;

[0160] ;

[0161] ;

[0162] in, and Both are indicator variables. At time step... Shizi Road Section The number of vehicles on the road is higher than When the required number of vehicles is Set to 1, otherwise set to 0. (When time step...) Shizi Road Section When the vehicle's speed exceeds the maximum speed limit allowed on the highway Select 1 if the value is 1, otherwise select 0.

[0163] In this embodiment, each toll road segment is modeled as an agent in a reinforcement algorithm. The agent acts as the main body interacting with the road network environment, based on the real-time status of the road network. Figure 7 middle ) produces action ( Figure 7 middle The toll rates for different road segments are dynamically adjusted. The changing toll rates affect the route selection model and indirectly influence the traffic flow in the road network by influencing travelers' route selection behavior, thereby regulating the road congestion and throughput efficiency.

[0164] In a preferred embodiment, the reinforcement learning algorithm in the model can be a multi-agent reinforcement learning algorithm based on collaboration, to better capture the mutual influence of various road segments in the road network. Furthermore, the multi-agent reinforcement learning algorithm based on collaboration can be combined with a graph neural network (GNN) to extract the state features of the entire road network, achieving theoretically superior results.

[0165] The dynamic toll collection method for expressways provided in this embodiment designs and develops a reinforcement learning simulation environment adapted to the actual road network structure and corresponding traffic rules of my country's expressways. It then utilizes existing reinforcement learning algorithms and the developed expressway network environment to build a reinforcement learning-based dynamic toll collection model for the eastern expressway. Based on the characteristics of my country's expressways and travelers' travel habits, a model of travelers' route selection behavior on Chinese expressways is designed. The model incorporates the prior influencing factor of "distance" to the destination for different routes, allowing travelers to weigh route distance, estimated travel time, and cost when making route choices, thus achieving better results. Multiple objective values, including "revenue," "throughput," and "total system runtime," are individually and jointly optimized using multi-objective optimization to achieve optimal results that benefit both expressway operators and travelers.

[0166] In a specific example, this embodiment verifies the performance of the dynamic toll collection model on a real highway network in Yunnan.

[0167] The experimental results were compared with my country's fixed-rate toll scheme and heuristic algorithms. The revenue increase column represents the increase percentage compared with the fixed-rate toll scheme under different optimization objectives.

[0168] exist Figure 6 In a simple, man-made road network:

[0169]

[0170] In parts of Yunnan's expressway network:

[0171]

[0172] Reference Figure 8 , Figure 8 This is a structural block diagram of an embodiment of the dynamic toll collection device for highways of the present invention.

[0173] like Figure 8 As shown, the dynamic toll collection device for highways proposed in this embodiment of the invention includes:

[0174] The first module 301 is used to establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model.

[0175] The second module 302 is used to establish a traveler's route selection model and determine the system strategy of the reinforcement learning model based on the system state and the traveler's route selection model.

[0176] The first determining module 303 is used to determine the system action for adjusting the rate of each toll road segment by the intelligent agent corresponding to each toll road segment according to the system strategy.

[0177] The calculation module 304 is used to calculate the system reward after the current system action is executed, and to continuously adjust the system policy of the reinforcement learning model based on the system reward.

[0178] The second determining module 305 is used to determine the optimal toll rate adjustment plan for each toll road segment when the system reward reaches its maximum.

[0179] Other embodiments or specific implementations of the dynamic toll collection device for highways of the present invention can be referred to the above-described method embodiments, and will not be repeated here.

[0180] Furthermore, the present invention also proposes a dynamic toll collection device for highways, characterized in that the dynamic toll collection device for highways includes a memory, a processor, and a dynamic toll collection program for highways stored in the memory and executable on the processor, wherein: when the dynamic toll collection program for highways is executed by the processor, it implements the dynamic toll collection method for highways described in various embodiments of the present invention.

[0181] The specific implementation method of the expressway dynamic toll collection equipment in this application is basically the same as the above-mentioned embodiments of the expressway dynamic toll collection method, and will not be repeated here.

[0182] Furthermore, this invention also proposes a readable storage medium, which includes a computer-readable storage medium storing a dynamic toll collection program for a highway. The readable storage medium may be... Figure 1 The memory 1005 in the terminal can also be at least one of ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc. The readable storage medium includes several instructions to cause a highway dynamic toll collection device with a processor to execute the highway dynamic toll collection method described in various embodiments of the present invention.

[0183] The specific implementation of the highway dynamic toll collection program in the readable storage medium of this application is basically the same as the embodiments of the highway dynamic toll collection method described above, and will not be repeated here.

[0184] It is understood that in the description of this specification, references to terms such as "one embodiment," "another embodiment," "other embodiments," or "first embodiment to Nth embodiment," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0185] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0186] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0188] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A dynamic toll collection method for highways, characterized in that, The method includes the following steps: Establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model; The steps involved in establishing a high-speed road network simulation environment model include: Establishing a high-speed network of ternary groups , Let be a finite non-empty set, representing the set of nodes in the highway network. The number of nodes; It is the set of edges between adjacent nodes; Adjacency matrices represent the connectivity information between two nodes in a highway network; Define each time step The set of all time steps T is Integer multiples of; defines the time interval for the system action of the agent to adjust the rate. , A fixed value determined by the fee management department; defining the time series set for fee rate adjustments. ; The steps for establishing a high-speed road network simulation environment model also include: Use the mean as Standard deviation is Modeling modified Gaussian random variables at time step From the source node to the destination node Traffic demand; Using Daganzo's CTM model to model traffic flow, each road in the road network... Divide into independent sub-segments; use Indicates that it belongs to a road section The set of all sub-segments, For each Its length is used This indicates that the distance traveled by a vehicle within a time step according to its free-flow velocity is determined; at each time step, the traffic state of a road segment is composed of the number of vehicles located in each sub-segment; using... , , , Separate road sections The length of the road, the free-flowing speed of vehicles on the road section, the capacity, and the density when congestion occurs; Establish a traveler route selection model, and determine the system strategy of the reinforcement learning model based on the system state and the traveler route selection model; The steps involved in establishing a traveler route selection model include: Establish a traveler route selection model based on the decision-making route model; The traveler route selection model determines the distance and utility function value of different routes when the traveler selects a route at the toll station entrance and at the intersection while driving, and selects the route with the shorter distance or the larger utility function value. The utility function value is expressed as follows: ; in, A path segment from the current node to the destination node. Let be the set of all road segments traversed from the current node to the destination node, and let represent a path from the current node to the destination node. It is the set of all distinct paths from the current node to the destination node. The total toll for a given route. Indicates road segment Cost weighting coefficient; This is approximated as the time required to traverse each road segment in the current state. For road section The current total number of vehicles; The expression for the traveler's route selection model is as follows: ; in, The function is used to calculate the first... The total physical distance of the path; The system state includes time steps. Traffic flow in the sub-segments under this update step The system state is described by the equation: System rewards include revenue, system uptime, and network throughput; among which: The system strategy for continuously adjusting the reinforcement learning model based on system rewards is as follows: ; ; ; ; in, Indicates from time step Time to step During the time period, from the road section Move to the section Total traffic volume Indicates at time step At that time, the total number of vehicles exiting the highway in the high-speed network; Based on the system strategy, determine the system actions for adjusting the toll rate for the corresponding intelligent agent of each toll road segment; Calculate the system reward after executing the current system action, and continuously adjust the system policy of the reinforcement learning model based on the system reward; When the system reward reaches its maximum, determine the optimal toll rate adjustment plan for each toll road segment.

2. The dynamic toll collection method for expressways as described in claim 1, characterized in that, The system strategy steps for continuously adjusting the reinforcement learning model based on system rewards also include: Measure the maximum difference between the number of vehicles in different lanes in all time steps. Specifically: ; in, Indicates roads in different directions; Define a network-independent statistic. Specifically: ; ; use and Describe the operational status of the road network in multiple dimensions.

3. The dynamic toll collection method for expressways as described in claim 2, characterized in that, The system strategy steps for continuously adjusting the reinforcement learning model based on system rewards also include: The definitions measure the proportion of speeding below the minimum speed limit on roads in different directions and the total speeding rate on both roads, respectively. ; ; ; in, and All are indicator variables, when time step Shizi Road Section The number of vehicles on the road is higher than When the required number of vehicles is Set to 1 if the time step is 1, otherwise set to 0; Shizi Road Section When the vehicle's speed exceeds the maximum speed limit allowed on the highway Select 1 if the value is 1, otherwise select 0. The values ​​%-violationRight, %-violationLeft, and %-sumOverSpeedviolation are used to measure the number of times a vehicle violates speed limits while driving.

4. A dynamic toll collection device for highways, characterized in that, For use in the dynamic toll collection method for highways as described in any one of claims 1-3, the dynamic toll collection device for highways comprises: The first module is used to establish a high-speed road network simulation environment model and determine the system state of the reinforcement learning model. The second module is used to build the traveler's route selection model and determine the system strategy of the reinforcement learning model based on the system state and the traveler's route selection model. The first determining module is used to determine the system action for adjusting the rate of each toll road segment by the corresponding intelligent agent, based on the system strategy. The computation module is used to calculate the system reward after executing the current system action, and to continuously adjust the system strategy of the reinforcement learning model based on the system reward. The second determining module is used to determine the optimal toll rate adjustment plan for each toll road segment when the system reward reaches its maximum.

5. A dynamic toll collection device for highways, characterized in that, The highway dynamic toll collection device includes: a memory, a processor, and a highway dynamic toll collection program stored in the memory and executable on the processor. When the highway dynamic toll collection program is executed by the processor, it implements the steps of the highway dynamic toll collection method as described in any one of claims 1 to 3.

6. A storage medium, characterized in that, The storage medium stores a dynamic toll collection program for highways, which, when executed by a processor, implements the steps of the dynamic toll collection method for highways as described in any one of claims 1 to 3.