Channel area boundary control method for trunk green wave control
By establishing a trunk line correlation model and channel area division algorithm, and combining with the reinforcement learning model to optimize regional boundary control, the problem of the impact of the surrounding areas of the trunk line on traffic is solved, and stable and efficient management of trunk line traffic and coordinated optimization of channel areas are achieved.
Patent Information
- Application Number
- CN202311187069.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-09-14
AI Technical Summary
The existing green wave control scheme is difficult to effectively consider the traffic impact of the trunk surrounding areas, resulting in limited traffic efficiency of the trunk line and lack of targeted regional boundary control, making it difficult to achieve the expected traffic optimization effect.
Establish a trunk line correlation model, design a channel area division algorithm, optimize regional boundary control through reinforcement learning model, and combine macro and micro traffic parameters to achieve stable and efficient management of trunk line traffic.
The traffic conditions of the trunk line have been optimized, delays and parking times have been reduced, traffic operation in related passage areas has been ensured, and the overall efficiency of urban transportation has been improved.
Smart Images

Figure CN117095542B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a channel area boundary control method oriented to trunk green wave control. Background Art
[0002] Traffic congestion, environmental pollution, and accidents—among other issues—not only hinder economic and social development but also impact the daily lives of urban residents, causing severe pollution and damage to the city's natural environment. The road traffic environment is deteriorating, especially during peak hours, when congestion and delays are severe, and the probability of accidents is also soaring. All of this has made traffic a critical issue in urban life that demands urgent resolution. Increasing investment in road traffic construction, the simplest and most direct approach, can initially effectively alleviate traffic problems, but it quickly becomes constrained by urban space and funding, resulting in significant limitations. Therefore, within the current context of road traffic, improving traffic control capabilities has become a hot topic of research. In recent years, with the development of urban traffic signal control, signal control research theory has gradually evolved from point-to-point to line-to-surface, continuously moving towards macro-micro integration, intelligentization, and integration.
[0003] Urban arterial green wave coordinated control is a typical line control system. It focuses on the main arterial traffic of urban transportation, prioritizing the improvement of main arterial traffic conditions. Through analytical modeling, it systematically coordinates and manages the signal control schemes of successive intersections along the arterial. This ensures that each arterial intersection coordinates with an appropriate green-to-signal ratio and phase difference, reducing stops and delays caused by vehicles passing through arterial intersections and enabling convoys to pass through the series of intersections without interruption. This prioritizes the smooth flow of arterial traffic and alleviates various urban traffic issues. However, the operation of the main arterial green wave is affected by many factors, such as incoming traffic and the overall traffic situation in the corridor area, making it difficult to achieve the expected results.
[0004] A comprehensive analysis of the current research progress of domestic and foreign scholars in the fields of trunk control, regional control, reinforcement learning, etc. shows that extensive research results have been achieved in each field, but there are still the following limitations and shortcomings:
[0005] (1) Existing green wave control schemes have little research on the areas surrounding the trunk line, and mainly focus on the traffic conditions of the trunk line itself. However, the green wave strategy is actually affected by many factors and often fails to achieve the control effect studied. For example, congestion in the areas surrounding the trunk line causes a large amount of traffic flow to the trunk line, which will have an impact on the operation of the trunk line traffic. The traffic efficiency of the trunk line is closely related to the traffic conditions in its surrounding areas.
[0006] (2) Existing regional boundary control mostly focuses on the region as a whole, optimizing the control of regional traffic from a macro perspective, but lacks targeted consideration of specific trunk lines within the region.
[0007] (3) In the process of conducting zoning research on the traffic network, traditional community division lacks targeted consideration of the impact of implementing green wave traffic on specific trunk lines. Summary of the Invention
[0008] To overcome the aforementioned shortcomings of the existing technology, the present invention proposes a channel area boundary control method for arterial green wave control. Focusing on the overall regional direction, the present invention considers the mutual influence between regional network traffic around the arterial and the linear traffic of the arterial green wave. From the perspective of regional control, the present invention rationally obtains the target control area, delineates regional boundaries, and establishes a reasonable signal control strategy for boundary signal intersections. This provides a stable and efficient regional network traffic environment for the operation of the arterial green wave belt, thereby optimizing and improving the arterial green wave traffic and ensuring the traffic conditions of the relevant channel areas, thereby alleviating urban traffic pressure.
[0009] The technical solution adopted by the present invention to solve the technical problem is: a channel area boundary control method for trunk green wave control, comprising the following steps:
[0010] Step 1: Establish a green wave band model for the target trunk line;
[0011] Step 2: Establish a trunk line correlation model for describing the correlation between the relevant channel area sections and the trunk line;
[0012] Step 3: Based on the channel area division algorithm oriented to the target trunk line, obtain the channel area oriented to the trunk line;
[0013] Step 4: Establish a channel area boundary control model and method for trunk lines based on the reinforcement learning model.
[0014] Compared with the prior art, the present invention has the following positive effects:
[0015] Based on the correlation between road sections in the road network and specific trunk lines, the present invention designs a trunk line correlation model that describes the degree of correlation between road sections and trunk lines; then, by analyzing the traditional cell division algorithm, a channel area division algorithm for urban trunk lines is designed to obtain traffic control cells (i.e., channel areas) facing specific trunk lines; finally, the present invention takes the channel area as the control object, establishes a channel area boundary control model and method for trunk lines based on reinforcement learning technology, and dynamically controls the boundary intersections of the area by combining macro and micro perspectives: at the macro level, the overall number of vehicles in the control area is stabilized, and at the micro level, the traffic flow distribution within the control area is reasonable, thereby optimizing and improving the traffic conditions of the trunk lines, reducing negative indicators such as trunk line traffic delays and the number of stops, and ensuring that the traffic conditions in the relevant channel areas are at a reasonable level. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0017] Figure 1 It is a bidirectional green wave band space-time diagram;
[0018] Figure 2 is the action space list;
[0019] Figure 3 This is a flow chart of the boundary intersection environment interaction. DETAILED DESCRIPTION
[0020] The goal of the present invention is to improve and optimize the trunk traffic conditions and collaboratively guarantee the traffic operation of related channel areas from the perspective of channel area boundary control. First, a trunk correlation model that describes the degree of correlation between road sections and trunk lines is established, and a channel area partitioning algorithm for the target trunk line is designed to obtain the target controlled channel area; then, based on macro basic graphs, reinforcement learning and other theories and technologies, a boundary control model and method are established for the boundary intersections of the controlled area to achieve optimization and improvement of the trunk traffic performance indicators and collaboratively guarantee the traffic operation of related channel areas.
[0021] The present invention first introduces the green wave band model used for the trunk line and its related parameters. Then, based on the traditional correlation model, the present invention establishes a trunk line correlation model to describe the degree of correlation between road sections and trunk lines. Then, based on the traditional road network partitioning theory and the DBSCAN algorithm, a set of channel area division algorithms for specific trunk lines is designed to obtain the target controlled area. Finally, based on the macro basic graph and related theories and technologies such as reinforcement learning, a set of channel area boundary control models and methods are designed for channel area boundary intersections. Finally, the present invention uses SUMO simulation to build a simulation test environment, compares and analyzes the actual benefits of the channel area division algorithm designed by the present invention and the channel area boundary control strategy, and confirms the effectiveness and advantages of the model scheme of the present invention.
[0022] The present invention will be described in detail below with reference to the accompanying drawings.
[0023] 1. Trunk Green Wave Band Model and Related Parameters
[0024] 1.1 Trunk Green Wave Band Model
[0025] This paper establishes a green wave model for a target arterial route based on the maximum green wave band method. This method maximizes the bandwidth for continuous traffic flow along the arterial route. A larger bandwidth allows for more traffic to continuously pass through the green wave band, which in turn improves the effectiveness of green wave coordinated control. The MAXBAND model is the basic model for numerous green wave band model algorithms used for urban arterial route control. This paper builds on this MAXBAND model by taking into account coordinated speed to establish a related green wave band model.
[0026] To establish a green wave band model, we must first clarify the relationship between the various parameters of the green wave band, such as Figure 1As shown, the horizontal coordinate in the time-space diagram represents time, and the vertical coordinate represents the distance between two adjacent intersections. The vertical coordinate in the time-space diagram represents the distance between two adjacent intersections, and the horizontal coordinate represents time. i―1 and S i For two adjacent intersections on a trunk line, this time-space diagram describes the parameters of the trunk line's green wave band control and their relationships. The solid line represents the upward direction, and the dashed line represents the downward direction. All time variables are based on the cycle as the basic unit. The specific parameter meanings are as follows:
[0027] —Uplink (downlink) green wave bandwidth;
[0028] S i —The i-th signalized intersection;
[0029] —Signalized intersection S i The ratio of red light time at upbound (downbound) traffic;
[0030] —Signalized intersection S i The time difference between the end (start) of the red light and the edge of the uplink (downlink) green wave band;
[0031] —Intersection S i―1 To intersection S i (S i to S i―1 )’s travel time ratio;
[0032] —Signal S i―1 From the midpoint of the up (down) red light to signal S i Upward (downward) red light midpoint time difference ratio;
[0033] Δ i —Intersection S i r i and Red light midpoint time difference ratio;
[0034] — Mainline traffic arrives at intersection S i Before, the time it takes for the original queue of vehicles to be cleared;
[0035] C—the duration of the trunk green wave control cycle;
[0036] λ—λ∈Z.
[0037] According to the above time-distance diagram and the meaning of related parameters, the green wave band model used in the present invention is derived and established. The core goal of the model is to maximize the green wave bandwidth. The objective function of the model is Z=Max(b). Figure 1 The time interval relationship shown in the figure establishes a corresponding mathematical relationship. First, the green time ratio of the green wave band must be greater than or equal to the green wave bandwidth of the uplink and downlink, so that the green wave can pass through the green time interval. The green time ratio can be obtained from the period and red time difference according to the figure, and the relationship is:
[0038]
[0039] By analyzing the relationship between the uplink and downlink phase differences based on the time-distance diagram, it was found that the ratio of the midpoint time difference of the red lights at adjacent uplink and downlink intersections is related to the ratio of the midpoint time difference of the red lights at a single intersection. Specifically, the sum of the ratios of the midpoint time difference of the red lights at adjacent uplink and downlink intersections plus the difference between the ratios of the midpoint time difference of the red lights at each adjacent intersection should be an integer multiple of the signal period. The equation is as follows:
[0040]
[0041] Analyzing the relationship between the arrival of the green wave and related parameters between intersection i and intersection i-1, we find that there are two ways to express the time interval from the midpoint of the red light at the starting intersection to the arrival of the green wave at the next intersection. The equations for the related parameters are established as shown in Equation (1-3):
[0042]
[0043] Analyze equations (1-2) and (1-3) and observe their common variables. Add the two equations of equation (1-3) and subtract them from equation (1-2) to obtain the new relationship equation:
[0044]
[0045] The green wave belt operation shown in the time-distance diagram is related to the green wave belt slope, that is, the running speed of the green wave belt will affect the results of the green wave belt model. The speed is used as a parameter to limit the time it takes for the green wave belt to travel between adjacent intersections. The time it takes for the green wave belt to run from intersection i-1 to intersection i should be within the intersection spacing l. i―1,i The mathematical relationship between the ratio of the two extreme values of belt speed is:
[0046] l i―1,i / v max ≤t i―1,i ≤l i―1,i / v min (1-5)
[0047] Based on the above time-distance diagram and the relationship between relevant parameters, a trunk green wave belt model is established. Its linear programming model is shown in formula (1-6):
[0048] Z=Max(b)
[0049]
[0050] The present invention establishes the above-mentioned green wave band model for the target trunk line, establishes speed intervals according to the upper and lower limits of the actual trunk line operating speed, adds the green wave band speed to the model as a variable, and solves the linear programming model using the pulp model of the Python language by determining the green wave band system period and coordinating the phase red light time. With the maximum bandwidth as the solution target, different bandwidth conditions under different speeds are solved. Finally, the signal control scheme of the trunk line of the present invention is designed with the maximum green wave bandwidth, the phase difference corresponding to each intersection is calculated, the trunk line intersection signal scheme is imported into the simulation environment, and the speed v corresponding to the green wave band is obtained. b As one of the factors to be considered in subsequent regional border control.
[0051] 1.2 Mainline Green Band Performance Indicators
[0052] (1) Delay time:
[0053] Delay time refers to the time loss caused by various interference factors such as traffic and traffic control during the operation of a vehicle. This indicator can be used to measure the state of road traffic. Generally speaking, the shorter the delay time, the better the traffic efficiency of the trunk line and the better the effect of green wave coordinated control. Therefore, the average delay time can be used as one of the evaluation indicators of the trunk line green wave belt.
[0054] (2) Number of stops:
[0055] The number of stops is mainly used to describe the degree of obstruction caused by traffic control measures during vehicle operation, that is, the ratio of the number of times the vehicle comes to a complete stop during driving. The number of stops index is often less than 1. Vehicles that do not come to a complete stop but merely slow down are not included in the number of stops index. The present invention uses a speed below 0.1 m / s as the judgment standard for stopping.
[0056] 2. Channel Region Division Algorithm and Boundary Control Model and Method
[0057] The present invention establishes a trunk correlation model for specific target trunks in the road network and designs a channel area division algorithm to obtain the target controlled channel area. Then, taking the controlled area as the object, a channel area boundary control strategy is established for its boundary intersections, thereby achieving improvement and optimization of trunk traffic.
[0058] 2.1 Channel Region Division Algorithm
[0059] 2.1.1 Principles of channel area division
[0060] Obtaining the target controlled channel area facing the trunk line as the basis for regional control will affect the effect of boundary control to a certain extent. Before studying the partitioning algorithm, the algorithm goal should be clarified in combination with specific research. The fundamental starting point for obtaining the channel area in this invention is to study the optimization of trunk line traffic from the perspective of regional control. The area obtained by the channel area partitioning algorithm should comply with the following principles:
[0061] (1) The divided sub-areas are homogeneous:
[0062] The present invention combines the macro basic map to study the regional characteristic relationship and boundary control strategy, and should ensure that the acquired area presents good macro basic map characteristics and has a good MFD (macro basic map) curve. A large number of studies have confirmed that the MFD characteristics of the regional road network are closely related to the homogeneity of the regional road network. Homogeneous areas are more likely to present MFD characteristics. Therefore, the present invention incorporates the road section density index into the consideration of channel area division.
[0063] (2) The divided sub-areas have a strong correlation with the trunk line:
[0064] For sections that are not related to the main line, such as sections that have no traffic exchange with the main line or are too far away, even if they have good homogeneity, their traffic conditions will not affect the main line traffic. The present invention is aimed at the channel area division of the main line green wave belt. The demarcated area should have a considerable degree of correlation with the target main line, and the regional road section traffic should have a considerable influence on the main line.
[0065] (3) The sub-areas after division should have a certain scale:
[0066] The MFD of the study area is essentially an aggregate analysis of a large amount of data in the area. The channel area as the division result should not be too large or too small. A moderate channel area is easier to describe and analyze the regional MFD characteristics.
[0067] (4) The divided sub-areas have good geographical connectivity with the trunk line:
[0068] The channel area should present a channel shape surrounding the main line, which not only reflects the relevance to the main line, but also facilitates the implementation of actual control measures.
[0069] 2.1.2 Trunk Correlation Model
[0070] Most correlation models use factors such as traffic flow characteristics and travel time as the main indicators to design correlation models. The present invention designs a trunk line correlation model based on the consideration of traffic flow exchange and travel time. It is used to describe the degree of correlation between a specific road section and the trunk line. Therefore, the traffic flow passing through the trunk line among all traffic flows in the road section is regarded as the key traffic flow, and the average travel time taken by all vehicles in the key traffic flow to reach the trunk line is regarded as the travel time between the road section and the trunk line. The present invention comprehensively considers the travel time and traffic flow exchange between the road section and the trunk line to establish the trunk line correlation model as shown in formula (2-1):
[0071]
[0072]
[0073] Where q i,g —Exchange traffic between section i and the main road, i.e., the target vehicle's origin and destination include both section i and the main road;
[0074] ∑q i —Total traffic volume on the road section;
[0075] l i —average travel distance from section i to the main road;
[0076] l j,m —The actual travel distance of vehicle j from section i to the main road;
[0077] n i,g —The total number of vehicles passing through section i and the main line.
[0078] 2.1.3 Channel Region Division Algorithm
[0079] The present invention designs a channel region division algorithm based on the idea of DBSCAN algorithm. The idea of channel region division has a good fit with the DBSCAN algorithm. The DBSCAN algorithm uses the density of spatial data as the idea of clustering. According to the algorithm requirements, a sufficiently dense data set is obtained, that is, each data sample clustered into the same class must be closely related to each other. Each sample point is an indispensable part of the class and there are enough data samples around it that are closely related to each other. At the level of the channel region division algorithm, it is to obtain a unique data set that is closely related to the trunk line. Each road section in the set has enough road sections that are closely related to each other. This basic idea is used to obtain a unique channel area that is closely connected to the target trunk line.
[0080] Based on the DBSCAN algorithm, this paper designs a channel area partitioning algorithm for trunk lines. Different from the original algorithm, the channel area only obtains unique clustering results, and the result data set must take the trunk line as the primary core. In the road network structure, the trunk line correlation and road section density are used as the inherent attributes of each road section, and the extended channel area with the trunk line as the core node is delineated. The trunk line correlation of the road sections belonging to the trunk line is the maximum value Max{I G}, a partitioning algorithm is designed based on the logic of expanding channels around the trunk line. The algorithm is established with trunk line correlation, density, ∈, MinPts, and attenuation coefficient β as basic parameters, and the channel area is divided outward. The channel area obtained by this method is mutually correlated with the trunk line and can maintain homogeneity. This algorithm comprehensively considers the trunk line correlation and density indicators, and the road section density is an important indicator to describe the homogeneity of the area. The present invention sets traffic detectors at the corresponding locations of the road network sections, uses SUMO simulation to generate the detector input and output files, and analyzes the road section density information during the entire traffic simulation process, and calculates its mean as the actual density parameter of different sections.
[0081] The channel area division algorithm aims to obtain a single cluster. After determining the target trunk set, the algorithm is executed according to the following algorithm flow.
[0082] 1) For a certain road network object, its network topology is U = {O, L}, where the intersection set O = {O1, O2, ..., O m}, the road segment set L={L1,L2,…,L n};
[0083] 2) Based on the actual traffic data of the road network, the road section is taken as the object node, and the exchange traffic ratio between each node and the trunk line and the travel distance to the trunk line are obtained according to the travel path files of all vehicles in the network. For each node L i Assign the corresponding characteristic value (trunk correlation I G and density K);
[0084] 3) Integrate all sample data into the initial sample set D = {x1, x2, ..., x n}, the corresponding domain parameter is described as (∈, MinPts), the trunk level distance attenuation coefficient β is defined, and the core object set is initialized as all the road section sub-objects corresponding to the trunk to expand the class. The initialized core object is Ω = {x1, x2, ..., x g}, while also reducing the large number of queries generated by operations on the entire dataset and speeding up processing.
[0085] 4) Category number k = 1, initialize cluster partition C gk =Ω, update the unvisited sample set Γ = Γ-∑C gk .
[0086] 5) Partition C in cluster gk In the collection, traverse and extract any core object o, and make the current cluster queue the initial queue Initialize the set of unvisited states and continuously update Γ=Γ―∑C gk , update the distance factor to ∈.
[0087] 6) For the core object o, find its ∈-neighborhood subsample set N ∈ (o):
[0088] N ∈ (o) Sample x j Find its ∈ Ω -Domain subsample set If it satisfies Then add it to the core sample set Ω, update Ω=Ω∪{x j}, update cluster partition C a =C a ∪{x j Repeat this step until N ∈ (o) Medium sample.
[0089] 7) Go to step 6 until you have traversed C gk Element, update category number k=k+1, C gk =C a Update the unvisited sample set Γ=Γ―∑C gk , update the distance factor to ∈ = β·∈.
[0090] 8) Go to step 5 until
[0091] 9) Output the results, a single channel partition cluster C that meets the DBSCAN clustering conditions G =∑C gk .
[0092] At this point, the traffic control cell facing a specific trunk line has been successfully obtained, and this channel area serves as the target object for boundary control. The advantage of the channel area division algorithm of the present invention is that the cell shape is channel-shaped and has good connectivity with the trunk line. The channel area obtained by the algorithm diverges outward with the trunk line as the core area to form a channel-shaped area including the target trunk line, which fits the research scenario of the present invention and conforms to the actual control scenario; the traditional partitioning algorithm is based on the overall target road network, and the overall area is divided into several control cells according to research needs. Its shape is uncertain, and it may be difficult to achieve coordinated control with the trunk line in practice. The cell characteristics have both cell homogeneity and trunk line correlation. The area is closely related to the trunk line, which is conducive to the development of regional boundary control facing the trunk line. Homogeneous cells are more likely to present good macro basic map characteristics, and the traffic parameters in the area are more closely correlated, which is conducive to the development of regional boundary control.
[0093] 2.2 Boundary Control Model and Method
[0094] The present invention establishes a channel area boundary control strategy based on macro basic graph theory and reinforcement learning theory, analyzes the relationship between macro regional traffic parameters, controls the number of vehicles in the entire region at a macro level, provides a good regional environment for trunk traffic operation, and controls the intensity of the boundary intersection confluence with the trunk line at a micro level, so that the traffic flow within the region is reasonably distributed, the trunk line capacity is fully utilized, and the trunk line traffic conditions are improved and optimized.
[0095] 2.2.1 Calculation of regional MFD
[0096] The calculation of regional MFD related parameters is shown in formula (2-3):
[0097]
[0098] Where n is the cumulative number of vehicles in the area;
[0099] o w 、k w ,q w —Weighted time occupancy, weighted density, and weighted traffic flow;
[0100] i—section number;
[0101] o i 、k i ,q i —Time occupancy, density and flow of road segment i;
[0102] l i —the length of section i;
[0103] s—average vehicle length.
[0104] The third-order equation based on the completed traffic flow and the cumulative number of vehicles in the road network is shown in Equation (2-4), which can be used to describe the traffic characteristics of the region at a macro level.
[0105] G(n(t))=a·n 3 (t)+b·n 2 (t)+c·n(t)+d(2-4)
[0106] Where G(n(t))—regional completed flow (veh / h);
[0107] n(t)—the cumulative number of vehicles in the area (veh);
[0108] a, b, c, d—fitting parameters.
[0109] The network's completed traffic consists of both outbound traffic and inbound traffic. Using a single signal cycle as the sampling interval, this paper defines the network's completed traffic as the sum of inbound traffic and outbound traffic within the sampling period. Outbound traffic is determined by analyzing, at each step during the simulation, vehicles outside the region whose previous step was within the region, adding them to the outbound vehicle set. Inbound traffic is determined by analyzing, at each step, whether vehicles within the region completed all their trips within the region, adding vehicles that completed their trips within the region to the inbound traffic set.
[0110] 2.2.2 Regional Boundary Control Strategy
[0111] After determining the target trunk line and the target controlled area, the present invention establishes a boundary control strategy for the controlled area based on the reinforcement learning model, thereby achieving improvement and optimization of trunk line traffic. The core control objectives of boundary control include two parts: the stability of the overall cumulative number of vehicles in the control area at the macro level and the reasonable distribution of traffic within the control area at the micro level. On the one hand, based on the regional macro basic map, the relationship between the relevant traffic parameters of the controlled area is analyzed, and the cumulative number of vehicles in the control area is within the range suitable for the operation of the trunk line green wave band; on the other hand, the distribution of traffic in the control area is controlled, and when the trunk line has surplus capacity, the traffic is rewarded for merging into the trunk line, and vice versa, it is punished. That is, under different road network traffic environments, the reasonable distribution of traffic within the area is guided, thereby ensuring the efficient operation of the trunk line. Based on the above control objectives, the state space, action space, and reward function are scientifically established, and the three are coordinated with each other to achieve boundary control of the controlled area.
[0112] (1) State space:
[0113] The state space describes the environment in which the intelligent agent is located. A scientific and reasonable state space is the basis for the interaction of intelligent agents. According to research needs, the continuous traffic state is described as a suitable state space, and its composition will also affect the effectiveness of the entire control scheme.
[0114] Based on the consideration of the control objectives, the present invention establishes a state space for the boundary intersection from the perspective of macro-micro integration as shown in formula (2-5), where S1 and S2 represent the cumulative number of vehicles in the area and the waiting state of vehicles at each entrance of the intersection, respectively. The macro-regional state is combined with the micro-intersection state, which not only describes in detail the specific state information of the intersection under the scenario environment of the present invention, but also can be coordinated with the action space and reward function described later. The cumulative number of vehicles is coordinated to realize the control of the overall number of vehicles in the area, and the waiting state of vehicles at the entrance is coordinated to realize the control of the traffic distribution in the area.
[0115] State={S1,S2} (2-5)
[0116] The specific meanings of states S1 and S2 are shown in equations (2-6) and (2-7). S1 is based on the macro basic graph related characteristics of the controlled area and the regional traffic parameter curve relationship. The number of vehicles N in the corresponding area where the trunk vehicles run at the green wave model speed is b As the optimal operating range, S1 is divided into three states. The S2 state describes the waiting state of vehicles at the four entrances to the boundary intersection: the east, south, west, and north. Vehicles with speeds below 0.1 m / s are considered waiting vehicles. The maximum number of waiting vehicles is used as the upper limit and divided into three equal parts. The waiting state values are defined as 1 / 3, 2 / 3, and 1, respectively. The state of the entrance is measured by the ratio of the actual number of waiting vehicles to the maximum number. The four S2 entrances have a total of 81 state spaces. Combining S1 and S2, a total of 243 traffic states are used to describe the state space of the boundary intersection.
[0117] S1={(0,N b ―25),(N b ―25,N b +25),(N b +25,∞)} (2-6)
[0118]
[0119] (2) Action Space
[0120] The reinforcement learning-based action space for boundary intersections is specifically represented by the different actions of intersection signals, including signal phase sequence adjustment, signal phase change, signal timing adjustment, and other signal-related action measures. Considering that signal phase sequence changes may easily lead to traffic accidents, the present invention adopts a method of fixed phase sequence adjustment and timing to establish a related action space for intersection signals. For boundary intersections, the present invention adopts a signaling scheme with separate releases for each entrance lane and phase. For example, a four-phase signaling scheme is adopted for the intersection, corresponding to the vehicle right-of-way of the four entrance lanes (east, west, south, and north). On the one hand, this phase scheme is more consistent with the state space. The four entrance lanes can better distinguish the traffic flow entering and exiting the area. The separate release of the entrance lanes also better fits the state space of waiting vehicles on the entrance lanes, which is conducive to improving the efficiency of interactive learning of intelligent agents. On the other hand, this action space and state space can be used together with the reward function to control the intensity of traffic flow from the boundary intersection to the main road by controlling the ratio of the release time of different entrance lanes and combining relevant parameters such as the entrance lane waiting state, thereby controlling the reasonable distribution of traffic flow within the area. For details, see the reward function in the third point.
[0121] The present invention establishes 15 action strategies based on the different release time ratios of the four entrance channels. Figure 2 As shown in the figure, Action = {1, 2, 3…15}, the action space is distributed to each boundary intersection, and the signal timing is scaled proportionally for special intersections such as three-way intersections. The reinforcement learning training process uses one signal cycle as a training round. At the end of the cycle, the benefits of the actions and the changes in the state are calculated, and the action strategy is selected at the beginning of the next cycle.
[0122] (3) Reward Function
[0123] The above-mentioned state space and action space together constitute the basic environment for intelligent agent interaction, and the reward function is the key to influencing the intelligent agent's learning process. The reward function should be consistent with the control objectives of the research, be able to describe the benefits and tendencies of the intelligent agent's learning process, and guide the intelligent agent to make reasonable action choices.
[0124] The boundary control strategy of the present invention aims to improve and optimize trunk traffic, and to provide a stable and efficient traffic environment for trunk traffic from a regional perspective. Directly using trunk efficiency indicators to establish rewards can easily reduce the effectiveness of learning at each intersection. Trunk indicators are not determined by a single intersection agent, but are jointly affected by all boundary intersections. Measuring rewards with trunk indicators alone may cause the system to deviate from the evaluation of the state actions of specific intersection agents. The establishment of a reward function should fit the control goal. The present invention takes the stability of the overall number of vehicles in the control area and the reasonable distribution of traffic in the area as the core control goals, and establishes a reward function as shown in formula (2-8). The exchange of traffic within and outside the area is used as the main factor. When the overall capacity of the region is surplus, the action of the intersection flowing into the area is greater than the outflow area is rewarded. On the contrary, when there are too many vehicles in the area, the action of the intersection flowing out is greater than the inflow is rewarded. When the overall traffic conditions in the region are suitable, the unbalanced inflow and outflow behavior is punished, and the overall number of vehicles in the region is macro-controlled. The three situations are supplemented with a correction coefficient μ i ∈(0,1), so that the traffic entering the area is reasonably distributed within the area.
[0125]
[0126] Where r(t)—from state S at time t t Transfer to state S t+1 The reward value,
[0127] n i,in —Vehicles entering the intersection i from time t to time t+1,
[0128] n i,out —Vehicles leaving intersection i from time t to time t+1,
[0129] μ i —Correction coefficient, the specific meaning is shown in formula (2-9).
[0130] Correction coefficient μ i The specific meaning of is shown in formula (2-9). The correction coefficient is established by combining the waiting vehicles, release time and trunk line correlation of each entrance road of the intersection to measure the intensity of the boundary intersection merging into the trunk line. The ratio of waiting vehicles to release time is used to measure the intensity of traffic released at the intersection. The trunk line correlation includes the traffic exchange ratio between the entrance road and the trunk line and the travel distance factor. These parameters are combined to establish a correction coefficient to measure the intensity of the traffic released from the intersection into the trunk line. When the trunk line speed v g Below green wave speed v bIf the proportion of mainline traffic in the region is too high, the intersection agent's actions that merge into the mainline with excessive intensity are penalized, reducing the amount of traffic merging into the mainline at the boundary intersection; conversely, merging into the mainline is rewarded. Establishing a correction coefficient for the reward function in this way allows boundary intersections to maintain a reasonable distribution of traffic within the region while controlling the overall number of vehicles in the region, while fully utilizing both the state space and action space, thereby improving the effectiveness and accuracy of the learning and training process.
[0131]
[0132] Where j is the entrance road in the jth direction of intersection i,
[0133] W i,j —Vehicle waiting status at the j-th direction entrance of intersection i,
[0134] P i,j —Ratio of the time length of the entrance lane for the jth direction of intersection i signal scheme,
[0135] I i,j —The main line association degree of the entrance road in the jth direction of intersection i.
[0136] The main process of establishing a boundary control strategy based on the above state space, action space, and reward function is as follows:
[0137] 1. Establish simulation-related files and build the SUMO simulation environment;
[0138] 2. Analyze the road network information, determine the boundary intersection objects based on the controlled area range, and establish an initial Q table for each controlled intersection;
[0139] 3. During the traversal time, reinforcement learning training is performed on the intersection;
[0140] 4. Traverse the intersection at each moment;
[0141] 5. At the start of the moment, determine the current intersection's position in state space based on the macroscopic number of vehicles and the microscopic state of vehicles waiting at the intersection entrance, and record the state;
[0142] 6. Before the next moment begins, calculate the reward function value of the intersection in this training round based on the traffic exchange situation, state parameters, action parameters and other parameters of the intersection.
[0143] 7. At the end of the time, calculate the new state of the intersection agent and the reward value obtained under the action strategy, and update the intersection Q table. Return to step 4 until the time traversal is completed.
[0144] 8. Repeatedly train the model until the training is completed and record the training results.
[0145] 9. Verify the effect of reinforcement learning. The simulation process updates the state and decides the action based on the Q-table data obtained from training. Dynamic signal control is performed on all boundary intersections. The control effect is verified based on the monitoring data results.
[0146] The process of agent interaction in the reinforcement learning experiment is as follows Figure 3 As shown in the figure, the intelligent agent is in the state space that describes the traffic environment. It continuously takes actions by interacting with the environmental perception and establishing the regional boundary signal control model. Whenever a signal plan is input, the intersection executes the action and interacts with the traffic environment to update to a new state. In the process, it obtains feedback from the environment and evaluates the specific actions under the specific state. In the process of continuous trial and error interaction with the traffic environment, it gradually obtains the optimal strategy decision method.
[0147] 3. Build a SUMO simulation environment for simulation analysis
[0148] The present invention uses SUMO simulation software to build a simulation test environment and establish corresponding road network files, signal light files, traffic flow files, detector files, etc. The software system analyzes the road network data, establishes a trunk green wave control strategy, and obtains the corresponding bandwidth and speed. Then, based on the channel area division algorithm, the target controlled area is obtained. By comparing with other cell division algorithms, it is shown that the channel area division algorithm of the present invention can obtain a good controlled channel area. Finally, a reinforcement learning boundary control strategy is established for the controlled channel area by combining macro and micro perspectives. By comparing the optimization effects of different control strategies on trunk parking rates and trunk delay times, it is confirmed that the channel area boundary control strategy of the present invention can effectively improve and optimize trunk traffic conditions and collaboratively guarantee traffic operations in related channel areas.
Claims
1. A channel area boundary control method for trunk green wave control, characterized by: The steps include: Step 1: Establish a green wave band model for the target trunk line; Step 2: Establish a trunk line correlation model for describing the correlation between road sections and trunk lines; Step 3: Based on the channel area division algorithm oriented towards the target trunk line, obtain the channel area oriented towards the trunk line: 1) For a certain road network object, its network topology is U = {O, L}, where the intersection set O = {O1, O2, ..., O m }, the road segment set L={L1,L2,…,L n }; 2) Based on the actual traffic data of the road network, the road section is taken as the object node, and the exchange traffic ratio between each node and the trunk line and the travel distance to the trunk line are obtained according to the travel path files of all vehicles in the network. For each node L i Assign it the corresponding characteristic value: trunk correlation I G and density K; 3) Integrate all sample data into the initial sample set D = {x1, x2, ..., x n }, the corresponding domain parameter is described as (∈,MinPts), which defines the trunk level distance attenuation coefficient β, expanding outward around the trunk; 4) Category number k = 1, initialize cluster partition C gk =Ω, update the unvisited sample set Γ = Γ-∑C gk ; 5) Partition C in cluster gk In the collection, traverse and extract any core object o, and make the current cluster queue the initial queue Initialize the set of unvisited states and continuously update Γ=Γ―∑C gk , update the distance factor to ∈; 6) For the core object o, find its ∈-neighborhood subsample set N ∈ (o): N ∈ (o) Sample x j Find its ∈ Ω -Domain subsample set If it satisfies Then add it to the core sample set Ω, update Ω=Ω∪{x j }, update cluster partition C a =C a ∪{x j Repeat this step until N ∈ (o) medium sample; 7) Go to step 6) until you have traversed C gk Element, update category number k=k+1, C gk =C a Update the unvisited sample set Γ=Γ―∑C gk , update the distance factor to ∈ = β·∈; 8) Go to step 5) until 9) Output the results, a single channel partition cluster C that meets the DBSCAN clustering conditions G =∑C gk ; Step 4: Establish a channel area boundary control model and method for trunk lines based on reinforcement learning technology.
2. The channel area boundary control method for trunk green wave control according to claim 1 is characterized by: The green band model is: Z=Max(b) Where: b—uplink green wave bandwidth; —Downlink green wave bandwidth; S i —i-th signalized intersection; r i —Signalized intersection S i The ratio of red light time on the up lane; —Signalized intersection S i The ratio of red light time for downhill traffic; ω i —Signalized intersection S i The time difference between the end of the red light and the edge of the uplink green wave band; —Signalized intersection S i The time difference between the start of the red light and the edge of the downward green wave band; —Intersection S i―1 To intersection S i travel time ratio; —Intersection S i To intersection S i―1 travel time ratio; —Signal S i―1 From the midpoint of the upgoing red light to signal S i Upward red light midpoint time difference ratio; —Signal S i Midpoint of the down-going red light to signal S i―1 Downward red light midpoint time difference ratio; Δ i —Intersection S i r i and Red light midpoint time difference ratio; τ i — Main line traffic arrives at intersection S i Before, the time it takes for the original queue of vehicles to be cleared; — Main line traffic arrives at intersection S i Before, the time it takes for the original queued vehicles to be cleared; C—the cycle length of the trunk green wave control; λ—λ∈Z, where Z is the objective function.
3. The channel area boundary control method for trunk green wave control according to claim 1 is characterized in that: The trunk correlation model is: Where q i,g —Exchange traffic between section i and the main road; ∑q i —Total traffic volume of the road section; i —average travel distance from section i to the main road; l j,m —The actual travel distance of vehicle j from section i to the trunk line; n i,g —The total number of vehicles passing through section i and the main line.
4. The channel area boundary control method for trunk green wave control according to claim 1 is characterized in that: In step 3), when expanding outward around the trunk line, the core object set is initialized to all the road section sub-objects corresponding to the trunk line to expand the class. The initialized core object is Ω = {x1, x2, ..., x g }.
5. The channel area boundary control method for trunk green wave control according to claim 1 is characterized in that: The model and method for establishing channel area boundary control for trunk lines based on reinforcement learning technology include: (1) Establish state space for boundary intersections: State={S1,S2} S1={(0,N b ―25),(N b ―25,N b +25),(N b +25,∞)} Among them, S1 and S2 represent the cumulative number of vehicles in the area and the waiting status of vehicles at each entrance of the intersection respectively. b W represents the number of vehicles in the corresponding area of the green wave model running at high speed. e ,W s ,W w ,W n Respectively represent the waiting vehicle status of the four entrance lanes of the boundary intersection in the east, south, west and north; (2) A fixed phase sequence adjustment method is used to establish an action space for intersection signals, and a signal plan is established for boundary intersections using a one-entry lane-one-phase independent release method; (3) With the core control objectives of stabilizing the number of vehicles in the control area and ensuring a reasonable distribution of traffic flow within the area, a reward function is established: Where: r(t)—the time from state S to t Transfer to state S t+1 The reward value, n i,in —Vehicles entering the intersection i from time t to time t+1, n i,out —The number of vehicles leaving the intersection i from time t to time t+1, n—The total number of vehicles in the area, μ i —Correction coefficient μ i ∈(0,1), n—the cumulative number of vehicles in the area.
6. The channel area boundary control method for trunk green wave control according to claim 5 is characterized in that: Based on the different release time ratios for the four entrance lanes, 15 action spaces (Action = {1, 2, 3…15}) were established. The action spaces were allocated to each boundary intersection, and the signal timing was proportionally scaled for special intersections. The reinforcement learning training process used one signal cycle as a training round. At the end of the cycle, the benefits of the actions and the changes in state were calculated, and the action strategy was selected at the beginning of the next cycle.
7. The channel area boundary control method for trunk green wave control according to claim 5 is characterized in that: Correction coefficient μ i Determined by the following formula: Where: j—the j-th direction entrance road of intersection i, W i,j —The waiting state of vehicles on the j-th entrance lane of intersection i, P i,j —Ratio of the time length of the entrance lane in the jth direction of the intersection i signal scheme, I i,j —The main line correlation degree of the entrance road in the jth direction of intersection i, v g Indicates the main line speed, v b Indicates green wave speed.
Citation Information
Patent Citations
Regional boundary main intersection signal control method based on deep reinforcement learning
CN113392577A