An event-triggered urban road network dynamic boundary control method
By dividing the urban road network into multiple homogeneous sub-regions and utilizing reinforcement learning and model predictive control algorithms, the boundaries of the urban road network are adjusted in real time. This solves the problem that existing boundary control methods cannot cope with spatiotemporal fluctuations in demand and dynamic evolution of congestion, thereby improving the flexibility and operational efficiency of the transportation network.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2025-05-30
- Publication Date
- 2026-08-04
AI Technical Summary
Existing urban road network boundary control methods cannot respond in real time to spatiotemporal fluctuations in demand and dynamic evolution of congestion. They suffer from problems such as static boundary delineation, fixed update mechanisms, and slow response, making it difficult to adapt to changes in complex urban traffic networks.
An event-triggered dynamic boundary control method for urban road networks is adopted. By dividing the urban road network into multiple homogeneous sub-regions, reinforcement learning algorithms are used to evaluate the state of the sub-regions in real time and trigger boundary adjustments. Combined with model predictive control algorithms, the boundary control rate is optimized to maximize the number of trips completed.
It enables real-time updates of boundary positions, improves the adaptability and response speed of control strategies, and enhances road network operation efficiency and traffic control effectiveness.
Smart Images

Figure CN120612829B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent traffic control, and in particular to an event-triggered dynamic boundary control method for urban road networks. Background Technology
[0002] With rapid socio-economic development, traffic congestion has become a key bottleneck for urban economic development and has reduced residents' quality of life. To alleviate these problems, various traffic control measures have been explored over the past few decades. Existing methods can be mainly divided into traffic control at isolated intersections and traffic control for small-scale road networks. However, these control strategies are typically computationally expensive and require detailed local traffic information.
[0003] Boundary control based on macroscopic fundamental graph models is one of the effective methods to avoid deadlock in large-scale road networks. However, most existing boundary control methods are based on static road network boundaries, which severely limits the ability of macroscopic fundamental graph models to analyze the evolution of road network congestion. A few studies have adopted dynamic boundary control methods, but these still suffer from problems such as fixed number of zones and fixed update intervals. They are unable to cope with the realities of urban road networks, such as spatiotemporal fluctuations in demand and the coupling between dynamic congestion evolution and control states.
[0004] For example, Chinese patent application CN202110725453.0 discloses a multimodal traffic network boundary control method. This method suffers from problems such as static boundary delineation, fixed update mechanism, and slow response. While it introduces multimodal traffic flow control and model predictive control techniques into the model, its boundary control relies on periodic updates at preset time intervals, lacking real-time perception of traffic state changes. This makes it difficult to cope with the evolution of traffic congestion caused by sudden events and spatiotemporal fluctuations in demand, thus limiting its widespread application in complex urban traffic networks.
[0005] Chinese patent application CN202210997796.7 discloses a traffic area boundary control method combining data-driven and reinforcement learning. However, this method suffers from shortcomings such as pre-defined zoning structures and a lack of dynamism in the boundary control strategy. While this method uses data-driven traffic modeling and incorporates deep reinforcement learning algorithms to optimize boundary control ratios, improving the system's adaptability to traffic conditions, its zoning structure is fixed from the initial modeling stage, failing to update boundary positions in real-time based on traffic conditions. This makes it difficult to adapt to the dynamic changes in regional congestion areas within urban road networks. Furthermore, this method primarily uses reinforcement learning for optimizing control ratios, neglecting to determine "when" boundary adjustment control should be triggered, lacking proactive identification mechanisms and event perception capabilities. In terms of evaluation, the method also fails to fully incorporate comprehensive indicators closely related to traffic efficiency, such as the number of vehicles completing attenuation, resulting in relatively limited objectivity and interpretability of the control effect. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a more flexible event-triggered dynamic boundary control method for urban road networks that adjusts controlled boundaries and implements control, taking into account the current situation of urban traffic flow characterized by spatiotemporal fluctuations in demand, dynamic evolution of congestion, and coupled correlation of control states.
[0007] The objective of this invention can be achieved through the following technical solutions:
[0008] A method for dynamic boundary control of urban road networks based on event triggering, comprising the following steps:
[0009] The urban road network is divided into multiple sub-regions that are homogeneous in terms of demand and density, and the macro-basic graph MFD equation is defined within each sub-region.
[0010] Based on the relative positional relationships of sub-regions, demand, and MFD equations of the macro-basic map, a multi-sub-region traffic system model based on the road network macro-basic map is constructed with the boundary control rate between sub-regions as the control input.
[0011] Based on a multi-sub-region traffic system model, the state of sub-regions is evaluated in real time. When a sub-region meets the event triggering conditions, the outermost boundary of the triggering sub-region is defined as the controlled boundary, where the event triggering conditions are learned by a reinforcement learning algorithm.
[0012] Based on the event-triggered control that determines the controlled boundary, the model predictive control algorithm is used to control the road network boundary. The boundary control rate of the controlled boundary is calculated with the goal of maximizing the number of trips completed within the sub-area, and the optimal boundary control strategy is obtained.
[0013] As a preferred technical solution, the multi-sub-region traffic system model is established based on the defined sub-region MFD equations and road network traffic flow conservation equations, as follows:
[0014]
[0015] Where, n ii (t) represents the cumulative number of vehicles within sub-region i; n ij (t) represents the cumulative number of vehicles from sub-region i to sub-region j; d ii (t) represents the traffic demand within sub-region i; d ij (t) represents the demand generated in sub-region i, ending in sub-region j; M ii (t) represents the flow completed within sub-region i at time t; sub-region h∈N i Let N be the neighboring unit of subregion i, where N is the neighboring unit of subregion i. i Represents the set of all neighboring cells i; control input u ih (t) and u hi (t) is defined as the boundary control law from sub-region i to sub-region h and from sub-region h to sub-region i in the multi-sub-region MFD system at time t; It is the effective flow from sub-region i through the adjacent sub-region h to sub-region j. It is the effective flow from adjacent sub-region h through sub-region i to sub-region j. It is the effective flow from the adjacent sub-region h through sub-region i to sub-region i.
[0016] As a preferred technical solution, the effective transfer flow is subject to supply constraints in downstream sub-regions and boundary capacity constraints between sub-regions. The effective flow transferred from sub-region i to sub-region j via sub-region h is... It is expressed as follows:
[0017]
[0018] Among them, M ihj (t) represents the transfer flow from sub-region i to sub-region j via sub-region h; θ ihj (t) represents the flow allocation ratio from sub-region i to sub-region j and passing through sub-region h; S ih Indicates supply restrictions; C ih This represents the boundary capacity between subregion i and subregion h.
[0019] As a preferred technical solution, the supply restriction is expressed as follows:
[0020]
[0021] Where, ω h It is the slope of the MFD congestion phase; The vehicle accumulation in sub-region h when it is completely congested represents the vehicle accumulation in sub-region i when it is completely congested; n h(t) represents the vehicle accumulation in subregion h; subregion m is the set of neighboring subregions N of subregion h. h A subregion; n ihk (t) represents the number of vehicles that travel from sub-region i through sub-region h to sub-region k at time t; n mhk (t) represents the number of vehicles that travel from sub-region m through sub-region h to sub-region k at time t; n hh (t) represents the cumulative number of vehicles within sub-region h.
[0022] As a preferred technical solution, the boundary capacity C ih The sum of the maximum flow rates that the roads connecting sub-region i to sub-region h can accommodate is expressed as:
[0023]
[0024] Among them, L ih It includes all boundary connecting roads on the boundary between sub-region i and sub-region h, where l is L. ih A road in the middle; a l C is the number of lanes in road l; l This refers to the capacity of the lane.
[0025] As a preferred technical solution, for the multi-sub-area traffic system model, based on event-triggered control, when a sub-area meets the event triggering condition, the outermost boundary of the triggering sub-area is taken as the controlled boundary according to the adjacent position relationship between sub-areas, and the control rate of the controlled boundary is determined by the model predictive control algorithm; the control rates of other boundaries are taken as set values.
[0026] As a preferred technical solution, the event triggering condition is: when the cumulative number of vehicles in the sub-region is greater than or equal to the trigger value at time t, the sub-region is determined to be a triggering sub-region.
[0027] As a preferred technical solution, the trigger value of each sub-region is obtained through a reinforcement learning algorithm based on deep Q-learning. The state space is defined as the vehicle accumulation and location information of each sub-region; the action space is whether each sub-region is triggered at the research time; the sub-region reward value is defined as the sum of the number of vehicles completed in the sub-region and its neighboring sub-regions under the consideration of decay within the control interval.
[0028] A deep neural network is constructed, and the state of each agent is input into the online network. According to the ε-greedy policy, actions are output and applied to the environment. The agents receive immediate rewards and move to the next state. Finally, the network outputs the Q-value of performing different actions in the corresponding state.
[0029] The Q value is continuously updated during the learning process. By summarizing the actions where the Q value is greater than the set threshold under different states, the trigger value of the sub-region is obtained.
[0030] As a preferred technical solution, the boundary control based on model predictive control establishes the objective function of the discrete optimization problem with the goal of maximizing the number of trips completed as follows:
[0031]
[0032] Among them, u d T represents the boundary control rate between regions. c To control the interval, N p Indicates the prediction time domain;
[0033] By combining initial conditions, vehicle accumulation, and boundary control rate constraints, the optimal control strategy at the boundary is determined through a model predictive control algorithm.
[0034] As a preferred technical solution, the controlled boundary position remains constant at different partition times until a new sub-region is triggered or an already triggered sub-region no longer meets the triggering conditions, at which point it is recalculated.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1) This invention provides an event-triggered dynamic boundary control method for urban road networks, which can dynamically adjust the controlled boundary according to real-time traffic conditions. Based on event triggering, dynamic zoning is performed. When the accumulated vehicle volume in a sub-zone exceeds the trigger threshold, the system automatically sets its outermost boundary as the controlled boundary, achieving real-time updating of the boundary position and improving the adaptability and response speed of the control strategy.
[0037] 2) This invention introduces a reinforcement learning algorithm into the dynamic partitioning. The trigger value of an event is determined using reinforcement learning. The state space is defined as the vehicle accumulation and location information of each sub-region. The reward value of a sub-region is defined as the sum of the number of vehicles completed in that sub-region and its neighboring sub-regions within the control interval, considering attenuation. The trigger threshold of each sub-region is adaptively determined through deep Q-learning, avoiding the subjectivity of manual setting and improving the accuracy of event recognition and the flexibility of boundary adjustment.
[0038] 3) In the boundary control stage, this invention combines model predictive control methods to maximize the number of vehicle trips completed as the optimization objective. It combines initial conditions with constraints such as vehicle accumulation and boundary control rate to solve the boundary control rate, thereby further improving the road network operation efficiency and traffic control effect. Attached Figure Description
[0039] Figure 1 This is a flowchart of an event-triggered dynamic boundary control method for urban road networks according to the present invention;
[0040] Figure 2 This is a schematic diagram of the multi-sub-area traffic system based on MFD constructed according to the present invention;
[0041] Figure 3 This is a schematic diagram of the event-triggered controlled boundary dynamic adjustment of the present invention. Detailed Implementation
[0042] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0043] Example 1
[0044] This invention addresses the real-time, dynamic adjustment of urban road network boundaries by proposing an event-triggered dynamic boundary control method for urban road networks. This method effectively responds to various emergencies and changes in traffic demand. Examples include... Figure 1 As shown, it mainly includes the following three stages, and the specific implementation steps are as follows:
[0045] Model building phase: Based on information such as the relative positional relationships of sub-regions, demand, and macroscopic fundamental diagram (MFD) equations, a multi-sub-region traffic system model based on the macroscopic fundamental diagram is constructed.
[0046] Dynamic partitioning stage: Based on the model building stage, the controlled boundary is determined by evaluating the sub-region state in real time and using event-triggered control, where the event triggering conditions are learned by reinforcement learning algorithm;
[0047] Boundary control phase: Based on the dynamic zoning phase, an optimal control model is constructed and a model predictive control algorithm is used to control the road network boundary in order to determine the optimal boundary control strategy.
[0048] Furthermore, the modeling methods in the model building phase are as follows:
[0049] (1.1) As Figure 2 As shown, a large heterogeneous city is selected as the object and divided into several sub-regions that are relatively homogeneous in terms of demand and density. A complete MFD equation is defined within each sub-region.
[0050] (1.2) The topological relationship of the sub-region is known. The initial vehicle accumulation of each sub-region and the demand OD between sub-regions can be obtained through observation.
[0051] (1.3) The subregion MFD equation is defined as follows:
[0052] n i (t) represents the accumulated vehicle volume in sub-region i. The external flow O of the sub-region... i (n i(t) should be limited only by the boundary capacity and the supply under oversaturation conditions, with everything else remaining unchanged, while the internal flow G i (n i (i) follows the MFD equation for subregion i. Therefore, the MFD relation for subregion i is expressed as:
[0053]
[0054] That is, for the traffic flow to the outer sub-region, its MFD equation can be written as:
[0055]
[0056] The relationship between the traffic flow in sub-region i and the accumulated vehicle volume in that sub-region is expressed as follows:
[0057]
[0058] Where, n i (t) must meet physical constraints, meaning its feasible quantity should fall between the conditions of no vehicles and complete congestion. v i and ω i It is the MFD parameter, v i It is the slope of the free-flow phase of MFD, ω i It is the slope of the MFD congestion phase; C i It is the maximum traffic volume in sub-region i; the optimal accumulation amount of sub-region i is between the lower bounds. and the Upper Realm between; This represents the number of vehicles accumulated in sub-region i when it is completely congested.
[0059] (1.4) The accumulated number of vehicles n in sub-region i i (t) satisfies:
[0060]
[0061] Where, n ij (t) is the accumulated number of vehicles in unit i whose destination is unit j.
[0062] (1.5) Based on the defined sub-region MFD equation and road network traffic flow conservation equation, the traffic flow model of the sub-region is established as follows:
[0063]
[0064] Where, n ii (t) represents the cumulative number of vehicles within sub-region i; n ij (t) represents the cumulative number of vehicles from sub-region i to sub-region j; d ii(t) represents the traffic demand within sub-region i, that is, the demand generated in sub-region i with sub-region i as the destination; d ij (t) represents the demand generated in sub-region i, ending in sub-region j, where i, j ∈ R. Sub-region h ∈ N. i (i∈R) represents the neighboring units of subregion i, where N i Represents the set of all neighboring cells i. Control input u ih (t) and u hi (t) is defined as the boundary control law in the multi-subregion MFD system at time t. Let u ih Taking (t) as an example, this variable represents the boundary control rate from sub-region i to sub-region h at time t. Wherein, It is the effective flow from sub-region i through the adjacent sub-region h to sub-region j. The definition is similar; M ii (t) represents the flow completed within sub-region i at time t, i.e., the number of vehicles that complete their journey within sub-region i.
[0065] (1.6) Taking the flow transferred from sub-region i to sub-region j via sub-region h as an example, since the downstream sub-region h needs to have sufficient capacity to provide space for flow from one or more upstream sub-regions, this transferred flow will be subject to the supply constraint of sub-region h. In addition to the supply constraint, It is also constrained by the capacity of the boundary between sub-region i and sub-region h. Boundary capacity refers to the sum of the maximum flows that the roads connecting sub-region i and sub-region h can accommodate, calculated as the sum of the road service capacities of all relevant boundary roads. Written as:
[0066]
[0067] Where, θ ihj (t) represents the flow distribution ratio from sub-region i to sub-region j and through sub-region h, which satisfies:
[0068]
[0069] (1.7)S ih It is a supply constraint, which can be expressed in the following form:
[0070]
[0071] Where subregion m is the set of neighboring subregions N of subregion h. h In a subregion, the number of vehicles that travel from subregion m through subregion h to subregion k at time t is n. mhk (t). Furthermore, the total road service capacity from sub-region i to sub-region h is described as the boundary capacity C. ih It can be expressed as:
[0072]
[0073] Among them, L ih It includes all boundary connecting roads on the boundary between sub-region i and sub-region h, where l is L. ih One of the roads, l contains a l There are 10 lanes. Assume the capacity C of each lane on road l is... l They are the same.
[0074] (1.8) Transfer flow M ihj (t) can be derived from O i (n i The function (t) is calculated as follows:
[0075]
[0076] (1.9) Internal completion flow rate M ii (t) can be written as:
[0077]
[0078] (1.10) Based on event-triggered control, the controlled boundary is defined as the outermost boundary of the "triggering sub-zone" according to the adjacent position relationship between sub-zones, and its control rate can be determined by an optimal algorithm. The control rates of other boundaries can be set to a certain value by the traffic manager.
[0079] The specific partitioning methods during the dynamic partitioning phase are as follows:
[0080] (2.1) The outermost boundary of the triggering sub-region that satisfies the event triggering conditions is defined as the controlled boundary, such as... Figure 2 As shown. Its boundary control rate is determined by the optimal algorithm. The event triggering condition is defined as: when the cumulative number of vehicles in sub-region i is greater than or equal to the trigger value at time t, that is:
[0081]
[0082] in, These are the trigger values for each sub-region.
[0083] (2.2) The trigger values for each sub-region are obtained through a reinforcement learning algorithm based on deep Q-learning. The state space is defined as the accumulated vehicle quantity and location information for each sub-region; the action space is whether each sub-region triggers at the current study time; the reward value for a sub-region is defined as the sum of the number of vehicles completed in that sub-region and its neighboring sub-regions within the control interval, considering decay. Assuming the current time is t and the control interval is T, the reward function for sub-region i can be written as:
[0084]
[0085] Where α is the attenuation factor.
[0086] (2.3) This implementation uses a deep neural network (DNN), whose network structure includes an input layer, multiple hidden layers, and an output layer. The input layer receives feature vectors from the input data, the hidden layers extract high-level features from the input features, and the output layer generates the final prediction result. The specific network architecture is as follows:
[0087] Input Layer: The dimension of the input layer is determined by the number of features in the input data. In traffic control scenarios, the dimension of the input layer is the number of sub-regions, with the state of each sub-region serving as an input unit.
[0088] Hidden layers: The network contains several hidden layers. Each hidden layer is connected to the previous layer through a linear transformation (fully connected layer) and uses a non-linear activation function (such as ReLU) to enhance the network's expressive power.
[0089] The first hidden layer receives the output of the input layer and maps it to a representation of 128 neurons.
[0090] The second hidden layer processes the output of the first hidden layer again, maintaining the structure of 128 neurons.
[0091] Output layer: The output layer generates the final prediction result. For each sub-region, two Q-values are output, corresponding to different action selections (action 0 and action 1).
[0092] (2.4) In the deep neural network of this embodiment, the output of each hidden layer is processed by a non-linear activation function. The present invention preferably uses the Rectified Linear Unit (ReLU) as the activation function, which is defined as follows:
[0093] f(x) = max(0,x)
[0094] The ReLU activation function truncates negative values to 0, ensuring that the gradient vanishing problem does not occur during backpropagation, making it suitable for deep neural networks.
[0095] (2.5) In the reinforcement learning framework, the goal is to minimize the error between the predicted Q-value and the target Q-value. This embodiment of the invention uses mean squared error (MSE) as the loss function, which is defined as:
[0096]
[0097] Where N is the number of mini-batch samples in each training iteration; Q(s) i ,a i ;θ) represents the current state s predicted by the online network. i and action ai The Q value, with parameter θ; r i For state s i Perform action a i The reward obtained later; γ is a discount factor, usually between 0 and 1, used to weigh the relative importance of current and future rewards; Q′(s i+1 ,a′;θ - ) represents the target network's response to the next state s i+1 Q-value estimation for action a′, with parameter θ - .
[0098] (2.6) Q-value update is a core process in deep Q-learning, representing the updating of network parameters by comparing the actual reward with the predicted Q-value. The Q-value update formula is as follows:
[0099]
[0100] Where Q(s) i ,a i ) is state s i and action a i The current Q value under r; i It is the immediate reward obtained from the environment; γ is the discount factor; max a ′Q(s i+1 ,a′) indicates the next state s i+1 The maximum Q-value among all possible actions a′ represents the estimated future optimal Q-value; α is the learning rate, used to control the step size for Q-value updates.
[0101] Using the above formula, the neural network gradually learns how to estimate the optimal Q-value for each state-action pair.
[0102] (2.7) In order to provide a stable target during training, this implementation uses a target network Q′(s,a′;θ). - The system synchronizes parameters between the target network and the online network using a soft update method. The update formula for the target network is:
[0103] θ - ←τ·θ+(1-τ)·θ -
[0104] Where: τ is the update rate, typically 0.001, used to control the step size when updating the target network each time; θ is a parameter of the online network; θ - These are the parameters of the target network.
[0105] (2.8) To accelerate network training and ensure the stability of gradient descent, this invention employs the Adam optimizer (Adaptive Moment Estimation). The Adam optimization algorithm combines the advantages of momentum and adaptive learning rate, and its update formula is as follows:
[0106]
[0107] Where η is the learning rate; m t It is the first moment estimate of the gradient, representing the average value of the gradient; v t is the second moment estimate of the gradient, representing the average of the squared gradient; ∈ is a small constant used to prevent the denominator from being zero.
[0108] The Adam algorithm effectively avoids oscillations during training by dynamically adjusting the learning rate, making it particularly suitable for training large neural networks.
[0109] (2.9) To improve sample utilization and avoid the impact of data correlation on training, this invention employs an experience replay mechanism. The specific process is as follows:
[0110] The agent's interaction with the environment, including its state, actions, rewards, and next state, is stored in an experience pool. During each training session, a small batch of data is randomly drawn from the experience pool for training to break the temporal correlation and improve the model's generalization ability.
[0111] (2.10) During training, the online network selects actions based on the current state and updates the network through the experience replay mechanism, while the target network is used to provide stable target values.
[0112] The specific control methods for the boundary control phase are as follows:
[0113] (3.1) In a multi-subregion MFD system, the boundary control process based on model predictive control (MPC) can be viewed as an optimization problem, the objective of which is to maximize the number of journeys completed within each subregion. In this scenario, the objective function for the discrete optimization problem is established as follows:
[0114]
[0115] Among them, u d T represents the boundary control rate of the controlled boundary. c To control the interval, N p This indicates the prediction time domain.
[0116] Discrete control time intervals can be described as control time domain N c The control process assumes that from the current time interval T... c Start, until time interval T c +Np -1 marks the end. Control inputs can only be switched within each control interval, up to T. c +N c And from T c +N c To T c +N p -1 remains constant during the period, where N p This is for the prediction time domain. Note that N needs to be considered in the formula. c ≤N p .
[0117] (3.2) Combining the initial conditions with constraints such as vehicle accumulation and boundary control rate, the optimal control rate at the boundary can be solved based on the model predictive control algorithm, thereby completing the boundary control.
[0118] (3.3) The vehicle accumulation in each sub-region of the system satisfies the following constraints:
[0119]
[0120] The inequality ensures that the number of vehicles in each sub-area is between 0 and the maximum number of vehicles the sub-area can accommodate. between.
[0121] (3.4) Boundary control rate u d It should be subject to the inequality constraints of minimum and maximum ratios:
[0122] u min ≤u d (t)≤u max
[0123] Among them, u min and u max These are the corresponding minimum and maximum boundary control rates, which are usually predefined by traffic managers.
[0124] (3.5) The last constraint is the initial state:
[0125]
[0126] Where n(0) represents the value at time t c The initial accumulation of time.
[0127] The controlled boundary position remains constant across different partition times until a new sub-region is triggered or an already triggered sub-region no longer meets the triggering conditions, at which point it is recalculated.
[0128] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for dynamic boundary control of urban road networks based on event triggering, characterized in that the steps include: include: The urban road network is divided into multiple sub-regions that are homogeneous in terms of demand and density, and the macro-basic graph MFD equation is defined within each sub-region. Based on the relative positions of sub-regions, demand, and MFD equations from the macroscopic basic map, a multi-sub-region traffic system model based on the road network macroscopic basic map is constructed, with the boundary control rate between sub-regions as the control input. This multi-sub-region traffic system model is established based on the defined sub-region MFD equations and the road network traffic flow conservation equations, as follows: in, Sub-region The cumulative number of vehicles inside; Indicates from sub-region To sub-region The cumulative number of vehicles; Sub-region Internal transportation needs; Sub-region Subregions generated in The demand at the endpoint; Is Time zone Internal traffic completion; sub-region Represented as a sub-region The neighboring units, where Represents all neighboring A collection of units; control input and Defined as being in Boundary control rates from subregion A to subregion B and from subregion B to subregion A in a time-multiple subregion MFD system; From sub-region Passing through adjacent sub-regions Arrival sub-region Effective traffic, From adjacent sub-regions Passing through sub-regions Arrival sub-region Effective traffic, From adjacent sub-regions Passing through sub-regions Arrival sub-region Effective traffic; Based on a multi-sub-region traffic system model, the state of sub-regions is evaluated in real time. When a sub-region meets the event triggering conditions, the outermost boundary of the triggering sub-region is defined as the controlled boundary, where the event triggering conditions are learned by a reinforcement learning algorithm. Based on the event-triggered control that determines the controlled boundary, the boundary control rate of the controlled boundary is calculated using an optimal control model that aims to maximize the number of trips completed within the sub-region. The model predictive control algorithm is then used to control the road network boundary to obtain the optimal boundary control strategy.
2. The event-triggered dynamic boundary control method for urban road networks according to claim 1, characterized in that, The effective transfer flow is constrained by the supply of downstream sub-regions and the boundary capacity constraints between sub-regions. Passing through sub-regions Xiangzi District Effective flow of transfer It is expressed as follows: in, For sub-region Passing through sub-regions Xiangzi District The transferred flow; From sub-region Departure to Sub-region And through sub-regions Traffic allocation ratio; Indicates supply restrictions; This represents the boundary capacity between subregions 𝑖 and ℎ.
3. The event-triggered dynamic boundary control method for urban road networks according to claim 2, characterized in that, The supply restriction is expressed as follows: in, It is the slope of the MFD congestion phase; Sub-region The sub-region representing the amount of vehicles accumulated during complete congestion. The amount of vehicles accumulated during a complete traffic jam; Sub-region Vehicle accumulation in the sub-region; It is a sub-region Adjacent sub-regions A sub-region; for From the sub-region Passing through sub-regions Arrival Sub-region The number of vehicles; for From the sub-region Passing through sub-regions Arrival Sub-region The number of vehicles; Sub-region The cumulative number of vehicles inside.
4. The event-triggered dynamic boundary control method for urban road networks according to claim 2, characterized in that, The boundary capacity The sum of the maximum flow rates that the roads connecting sub-regions A to B can accommodate is expressed as: in, Includes sub-regions sub-region All the boundary lines connecting the roads between them yes One of the roads in the middle; Let l be the number of lanes in road l; This refers to the capacity of the lane.
5. The event-triggered dynamic boundary control method for urban road networks according to claim 1, characterized in that, For the multi-sub-area traffic system model, based on event-triggered control, when a sub-area meets the event triggering condition, the outermost boundary of the triggering sub-area is taken as the controlled boundary according to the adjacent position relationship between sub-areas, and the control rate of the controlled boundary is determined by the model predictive control algorithm; the control rates of other boundaries are taken as set values.
6. The event-triggered dynamic boundary control method for urban road networks according to claim 1, characterized in that, The event is triggered when the cumulative number of vehicles in the sub-region reaches a certain threshold at time [time value missing]. If the value is greater than or equal to the trigger value, then the sub-region is determined to be a trigger sub-region.
7. The event-triggered dynamic boundary control method for urban road networks according to claim 6, characterized in that, The trigger value for each sub-region is obtained through a reinforcement learning algorithm based on deep Q-learning. The state space is defined as the vehicle accumulation and location information of each sub-region; the action space is whether each sub-region is triggered at the research time; the sub-region reward value is defined as the sum of the number of vehicles completed in that sub-region and its neighboring sub-regions under decay consideration within the control interval. Build a deep neural network, input the state of each agent into the online network, and then... ε The greedy strategy outputs actions and applies all agent actions to the environment, receiving immediate rewards and transitioning to the next state; ultimately, the network outputs the Q-values for performing different actions in the corresponding states. The Q value is continuously updated during the learning process. By summarizing the actions where the Q value is greater than the set threshold under different states, the trigger value of the sub-region is obtained.
8. The event-triggered dynamic boundary control method for urban road networks according to claim 1, characterized in that, The boundary control based on model predictive control establishes the objective function of the discrete optimization problem with the goal of maximizing the number of completed journeys within the sub-region as follows: in, Indicates the boundary control rate between regions. To control the interval, Indicates the prediction time domain; By combining initial conditions, vehicle accumulation, and boundary control rate constraints, the optimal control strategy at the boundary is determined through a model predictive control algorithm.
9. The event-triggered dynamic boundary control method for urban road networks according to claim 1, characterized in that, The controlled boundary position remains constant across different partition times until a new sub-region is triggered or an already triggered sub-region no longer meets the triggering conditions, at which point it is recalculated.