A method for cooperative variable speed limit control of a merging area based on deep reinforcement learning
By applying deep reinforcement learning to the merging zone of ramps, traffic state information is obtained and a speed limit control model is trained, which solves the problem of insufficient coordinated control between the main line and ramps, realizes more efficient coordinated variable speed limit control in the merging zone, and improves the operational efficiency of traffic flow.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTH CHINA UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2025-09-17
- Publication Date
- 2026-04-21
AI Technical Summary
Existing collaborative variable speed limit control methods for merging zones mainly focus on one type of control means, with shortcomings in the collaborative control of the main line and ramps, an imperfect linkage control mechanism, and insufficient research on traffic operation characteristics under mixed traffic flow.
Based on deep reinforcement learning, this method acquires traffic state information by deploying loop detectors in the merging zone of ramps, sets demand scenarios and speed limit information, constructs a reward function, trains a speed limit control model, realizes the joint action of the main line and ramps, and optimizes the speed limit control strategy.
It improved the efficiency of coordinated control between the main line and ramps, perfected the linkage control mechanism, and enhanced the traffic efficiency and operational status of the merging area.
Smart Images

Figure CN121171027B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent transportation systems technology, and in particular relates to a collaborative variable speed limit control method for merging zones based on deep reinforcement learning. Background Technology
[0002] Currently, with the rapid development of my country's economy and the significant increase in car ownership, traffic congestion on highways during peak hours has become very common. To alleviate these problems, many scholars have proposed active traffic control methods, such as ramp control and variable speed limit control. Control strategies are mainly divided into mainline control and ramp control. Deep reinforcement learning has achieved significant results in variable speed limit and ramp signal control. However, research has mainly focused on one type of control method, and there are still shortcomings in the coordinated control of mainlines and ramps; the linkage control mechanism between mainlines and ramps is not yet perfect. Meanwhile, with the development of intelligent connected vehicle technology, significant breakthroughs have been made in vehicle automation, and mixed traffic flows composed of intelligent connected vehicles and manually driven vehicles will inevitably exist for a long time.
[0003] Therefore, studying the traffic operation characteristics under mixed traffic flow and formulating reasonable traffic management measures has become a necessary research direction. Summary of the Invention
[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a deep reinforcement learning-based collaborative variable speed limit control method for merging areas. This method solves the problems of existing collaborative variable speed limit control methods for merging areas, which mainly focus on a single type of control and have deficiencies in the collaborative control between the mainline and ramps, as well as the imperfect linkage control mechanism between the mainline and ramps.
[0005] To achieve the above objectives, the technical solution adopted by this invention is: a collaborative variable speed limiting control method for merging regions based on deep reinforcement learning, comprising the following steps:
[0006] S1. Based on the simulation environment of the ramp merging area, deploy loop detectors, use access interfaces to obtain traffic status information, set up demand scenarios, and set up road test units and speed limit panels to transmit speed limit information of the main line and ramps.
[0007] S2. Based on traffic state information, define the state space, combine the actions of the main line and ramps, obtain the action space and get the single-selection action, introduce the ramp dwell time, construct the reward function, and obtain the speed limit control model.
[0008] S3. Based on the demand scenario, road test unit, and speed limit panel, and combined with a single selection action, train the speed limit control model to obtain the trained speed limit control model.
[0009] S4. Using the trained speed limit control model, simulation is performed to obtain the control strategy for the merging zone of the ramp, and the coordinated variable speed limit control of the merging zone is completed.
[0010] The beneficial effects of this invention are as follows: Based on the environment of mixed traffic flow, this invention combines ramp control and variable speed limits on the main road to control the entire merging zone. The variable speed limit control method is applied to the ramps, and the action spaces of the ramps and the main road are discretized and combined. At the same time, the efficiency of the merging zone and the queue length on the ramps are incorporated into the reward function. This invention achieves control over multiple types of control methods, improves the coordinated control of the main line and ramps, and perfects the linkage control mechanism between the main line and ramps.
[0011] It also put forward reasonable suggestions for improving the traffic efficiency of merging zones and improving their operational status.
[0012] Further, S1 includes the following steps:
[0013] S101. Based on the merging area of highway ramps, obtain the simulation environment of the merging area through simulation;
[0014] S102. Deploy loop detectors in the simulation environment of the ramp merging area and use the access interface to obtain the traffic status information of the ramp merging area system in real time.
[0015] S103. By pre-setting the mainline traffic and ramp traffic, set up demand scenarios including low demand scenarios, moderate demand scenarios and high demand scenarios.
[0016] S104. Install road test units and speed limit panels for transmitting speed limit information for mainlines and ramps.
[0017] The beneficial effects of the above-mentioned further solutions are as follows: By arranging coil detectors in the simulation environment of the merging zone of the ramp, the present invention acquires data in real time, presets the main line flow and ramp flow, sets up demand scenarios to cope with the variability of real scenarios, and sets up road test units and speed limit panels to improve the efficiency and accuracy of data transmission, providing an accurate simulation environment for collaborative variable speed limit control in the merging zone.
[0018] Furthermore, S2 includes the following steps:
[0019] S201. Define the state space based on the speed limit information of the mainline and ramps in the traffic state information, as well as the occupancy and speed of the mainline, merging zone, and ramps.
[0020] S202. Based on the relationship between the mainline speed limit and the mainline actions, as well as the relationship between the ramp speed limit and the ramp actions, obtain the action space, set speed limit constraints, and obtain a single selected action.
[0021] S203. Calculate the merging zone travel time and total travel time, and incorporate ramp dwell time to construct a reward function;
[0022] S204. By integrating the state space, action space, and reward function, a speed limit control model is obtained.
[0023] Furthermore, the expression for the state space is as follows:
[0024] ;
[0025] in, and These represent the market share and speed of the upstream of the main line, respectively. and These represent the occupancy and velocity at the front of the merging zone, respectively. and These represent the occupancy and velocity in the middle of the merging zone, respectively. and These represent the occupancy rate and velocity at the rear of the merging zone, respectively. and These represent the occupancy rate and velocity measured by the detector in the bottleneck area downstream of the main line, respectively. and These represent the occupancy rate and speed on the ramp, respectively, and can be measured by the entrance ramp detector. Indicates period t Speed limit information for the main road. Indicates period t Speed limit information for ramps.
[0026] Furthermore, step S202 includes the following steps:
[0027] S2021. Based on the relationship between the mainline speed limit and the mainline action, and combined with the minimum value of the mainline speed limit, seven mainline actions are obtained.
[0028] S2022. Based on the relationship between ramp speed limits and ramp actions, and combined with the minimum value of the ramp speed limit, five types of ramp actions are obtained. The main line actions and ramp actions are combined to obtain a space of thirty-five actions.
[0029] S2023. Based on the speed limit changes of the main line and ramps in adjacent cycles, speed limit constraints are set to obtain the constrained action space of the main line and ramps, resulting in nine single-selection actions.
[0030] Furthermore, the expression for the speed limit constraint is as follows:
[0031] ;
[0032] ;
[0033] in, Indicates period t Speed limit information for the main road. Indicates period Speed limit information for the main road. Indicates period t Speed limit information for ramps, Indicates period Speed limit information for ramps, cycle time For period t The previous cycle.
[0034] Furthermore, step S203 includes the following steps:
[0035] S2031. The merging zone travel time is calculated based on the number of vehicles passing through the merging zone within a cycle.
[0036] S2032. The total travel time is calculated based on the number of vehicles leaving the merging zone in one cycle and the number of vehicles entering the merging zone in one cycle.
[0037] S2033. Introducing ramp dwell time: The ramp dwell time is calculated based on the number of vehicles passing through the merging zone within one cycle.
[0038] S2034. Construct a reward function based on minimizing the merging zone travel time, minimizing the total travel time, and minimizing the ramp dwell time.
[0039] Furthermore, the expression for the reward function is as follows:
[0040] ;
[0041] ;
[0042] in, Represents the reward function, This indicates the merging zone travel time for all vehicles within a cycle. This indicates the total travel time of the vehicle. This indicates the time a vehicle spends on a ramp within a given cycle. This indicates the number of vehicles on the ramp within one cycle. This indicates the vehicle's number on the ramp. Indicates the first The travel time for a vehicle to pass through the ramp. This represents the weighting coefficient of the sub-rewards related to traffic efficiency in the merging zone. This represents the weighting coefficient of the sub-rewards related to the total travel time in the merging zone. This represents the weight coefficients corresponding to the sub-reward function related to the ramp queue length.
[0043] The beneficial effects of the above-mentioned further solutions are as follows: This invention applies the variable speed limit control method to ramps, discretizes and combines the action spaces of ramps and main roads, and incorporates the efficiency of the merging zone and the queue length on the ramps in the reward function.
[0044] Regarding the state space, by integrating data from the main line, merging area, and ramps, a comprehensive state space is provided, clearly defining the collection of ramp environmental information and improving the agent's ability to perceive the environment;
[0045] For the action space, the agent interacts with the environment by combining the actions of the main line and the ramp, covering all actions of the main line and the ramp.
[0046] For the reward function, the influence of ramps is considered while improving efficiency, which improves the accuracy and effectiveness of the reward function and increases the speed at which the agent converges to the final policy.
[0047] Furthermore, step S3 includes the following steps:
[0048] S301. Based on the required scenario, set model parameters including the number of neural network layers, number of neurons, discount factor, learning rate, batch size, update frequency of the competing target network, initial exploration rate, decay rate of the initial exploration rate per round, and minimum exploration rate to obtain the set rate limiting control model.
[0049] S302. Based on the preset traffic simulation, the speed limit control model is iterated and the initial traffic environment state is obtained using a loop detector.
[0050] S303. Based on a greedy strategy, using a preset probability. For a single selection action, a random selection is made, or probability is used. Based on the status, the combined actions of the main line and the ramp are output to obtain the output action;
[0051] S304. Convert the output action into a speed limit value, use the road test unit and speed limit panel to publish the speed limit value to the intelligent connected vehicles traveling on the road segment, and then obtain the reward value and the next state space from the ramp merging area environment.
[0052] S305. Combine the current state space, output action, reward value and next state space into a quadruple, and store the quadruple as an experience sample in the experience replay pool.
[0053] S306. Based on the experience replay pool, experience samples are extracted using the priority experience replay mechanism. The online network in the speed limit control model is updated using the loss function, and the target network parameters are updated using soft updates to obtain the trained speed limit control model.
[0054] The beneficial effects of the above-mentioned further solutions are as follows: This invention improves the accuracy and efficiency of the speed limit control model by setting parameters, iterative simulation, selecting output actions based on a greedy strategy, and constructing a four-tuple to create an experience replay pool. Through soft updates, with the goal of maximizing cumulative rewards, the model is trained by the agent continuously exploring in the simulation environment. This also improves the effectiveness of subsequent control strategies for ramp merging zones, and achieves variable speed limit control by combining the speed limit actions of the mainline and ramps. Attached Figure Description
[0055] Figure 1 This is a flowchart of the method of the present invention.
[0056] Figure 2 This is a schematic diagram of the research scenario in this embodiment.
[0057] Figure 3 This is a graph showing the change in traffic flow over time in this embodiment.
[0058] Figure 4 This is a diagram of the algorithm structure for training the speed limit control model in this embodiment.
[0059] Figure 5 This is a schematic diagram of the training results of the speed limit control model in this embodiment.
[0060] Figure 6 This is a schematic diagram of the speed limit result in this embodiment. Detailed Implementation
[0061] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0062] Before describing this embodiment, the following terms will be explained:
[0063] SUMO: An open-source, microscopic, multimodal traffic simulation software;
[0064] Traci interface: An interface for accessing and controlling a running traffic simulation;
[0065] veh / h: Theoretical capacity;
[0066] RSU: Road Test Unit;
[0067] VMS: Speed Limit Panel;
[0068] HDV: Manually driven vehicle;
[0069] CAV: Intelligent Connected Vehicle.
[0070] Example
[0071] like Figure 1 As shown, this invention provides a collaborative variable speed limiting control method for merging regions based on deep reinforcement learning, the implementation of which is as follows:
[0072] S1. Based on the simulation environment of the ramp merging area, deploy loop detectors and use the access interface to obtain traffic status information, set up the required scenario, and set up road test units and speed limit panels to transmit speed limit information of the main line and ramps. The specific steps are as follows:
[0073] S101. Based on the merging area of highway ramps, obtain the simulation environment of the merging area through simulation;
[0074] S102. Deploy loop detectors in the simulation environment of the ramp merging area and use the access interface to obtain the traffic status information of the ramp merging area system in real time.
[0075] S103. By pre-setting the mainline traffic and ramp traffic, set up demand scenarios including low demand scenarios, moderate demand scenarios and high demand scenarios.
[0076] S104. Install road test units and speed limit panels for transmitting speed limit information for mainlines and ramps.
[0077] In this embodiment, as Figure 2 As shown, based on the merging area of highway ramps, a simulation environment for the merging area is established using SUMO, and loop detectors are deployed in the established simulation environment. Traffic status information of the merging area system is obtained in real time through the Traci interface.
[0078] To enable the collaborative speed limit control model to adapt to different traffic demand scenarios, the mainline flow and ramp flow are preset, and demand scenarios including low demand, moderate demand and high demand are set.
[0079] Specifically, in low-demand scenarios, the mainline traffic flow is 3000 veh / h, and the ramp traffic flow is 700 veh / h; in moderate-demand scenarios, the mainline traffic flow is 4000 veh / h, and the ramp traffic flow is 1000 veh / h; in high-demand scenarios, the mainline traffic flow is 5000 veh / h, and the ramp traffic flow is 1500 veh / h; and as... Figure 3 The traffic flow graph shown demonstrates how to adapt to changing demand scenarios and respond to real-world situations.
[0080] Configure RSU and VMS, and simultaneously send speed limit information for the mainline and ramps to the intelligent connected vehicles via RSU and VMS.
[0081] S2. Based on traffic state information, define the state space, combine the actions of the mainline and ramps to obtain the action space and obtain the single-selection action, introduce the ramp dwell time, construct the reward function, and obtain the speed limit control model. The specific steps are as follows:
[0082] S201. Define the state space based on the speed limit information of the mainline and ramps, the occupancy and speed of the mainline, merging zone and ramps in the traffic state information.
[0083] In this embodiment, the state space of the speed limit control model is defined using the speed limit information of the mainline and ramps in the traffic state information, as well as the occupancy and speed of the mainline, merging zone, and ramps. The expression of the state space is as follows:
[0084] ;
[0085] in, and These represent the market share and speed of the upstream of the main line, respectively. and These represent the occupancy and velocity at the front of the merging zone, respectively. and These represent the occupancy and velocity in the middle of the merging zone, respectively. and These represent the occupancy rate and velocity at the rear of the merging zone, respectively. and These represent the occupancy rate and velocity measured by the detector in the bottleneck area downstream of the main line, respectively. and These represent the occupancy rate and speed on the ramp, respectively, and can be measured by the entrance ramp detector. Indicates period t Speed limit information for the main road. Indicates period t Speed limit information for ramps;
[0086] Define a state space containing occupancy, speed, and speed limit information, and define a set of ramp environmental information for the agent to perceive changes in the environment.
[0087] S202. Based on the relationship between the mainline speed limit and mainline actions, and the relationship between the ramp speed limit and ramp actions, obtain the action space, set speed limit constraints, and obtain the single-selection action. The specific steps are as follows:
[0088] S2021. Based on the relationship between the mainline speed limit and the mainline action, and combined with the minimum value of the mainline speed limit, seven mainline actions are obtained.
[0089] S2022. Based on the relationship between ramp speed limits and ramp actions, and combined with the minimum value of the ramp speed limit, five types of ramp actions are obtained. The main line actions and ramp actions are combined to obtain a space of thirty-five actions.
[0090] S2023. Based on the speed limit changes of the main line and ramps in adjacent cycles, speed limit constraints are set to obtain the constrained action space of the main line and ramps, resulting in nine single-selection actions.
[0091] In this embodiment, the action space of the speed limit control model is defined. The action space is represented by the speed limit values of the main line and the ramp. Therefore, the action space is solved according to the relationship between the main line speed limit and the main line action, and the relationship between the ramp speed limit and the ramp action.
[0092] The relationship between the main story speed limit and the main story actions is as follows:
[0093] ;
[0094] in, This indicates the speed limit value for the main line. This represents the minimum speed limit for the main road, typically 60 km / h. Indicates the main action. This indicates the unit of length that increases the speed limit, typically 10 km / h.
[0095] Get seven main action lines , That is, the main line has a speed limit. km / h;
[0096] The relationship between ramp speed limits and ramp operation is as follows:
[0097] ;
[0098] in, Indicates the speed limit value of the ramp. This indicates the minimum speed limit for a ramp, typically 10 km / h. Indicates the main action;
[0099] Five types of ramp actions were obtained. , That is, the speed limit on the ramp has km / h, joint main line action and ramp operation The total action space of the intelligent agent is thirty-five types.
[0100] In this embodiment, to improve the safety of the mainline and ramps, avoid safety hazards caused by large changes in speed limits, and allow the intelligent agent sufficient exploration space, the speed limit changes of the mainline and ramps in adjacent cycles are set to be less than or equal to 10 km / h; the specific speed limit conditions are as follows:
[0101] ;
[0102] ;
[0103] in, Indicates period t Speed limit information for the main road. Indicates period Speed limit information for the main road. Indicates period t Speed limit information for ramps, Indicates period Speed limit information for ramps, cycle time For period t The previous cycle; the constrained mainline and ramp action spaces are respectively ;
[0104] in, This indicates that the variable speed limit on the main road or ramp has been reduced by 10 km / h. This indicates that the current speed limit remains unchanged. This indicates that the current speed limit has increased by 10 km / h;
[0105] In summary, there are nine selectable actions for each main line and ramp, resulting in nine single-action selection actions.
[0106] S203. Calculate the merging zone travel time and total travel time, and incorporate ramp dwell time to construct a reward function. The specific steps are as follows:
[0107] S2031. The merging zone travel time is calculated based on the number of vehicles passing through the merging zone within a cycle.
[0108] S2032. The total travel time is calculated based on the number of vehicles leaving the merging zone in one cycle and the number of vehicles entering the merging zone in one cycle.
[0109] S2033. Introducing ramp dwell time: The ramp dwell time is calculated based on the number of vehicles passing through the merging zone within one cycle.
[0110] S2034. Construct a reward function based on minimizing the merging zone travel time, minimizing the total travel time, and minimizing the ramp dwell time.
[0111] S204. By integrating the state space, action space, and reward function, a speed limit control model is obtained.
[0112] In this embodiment, to improve traffic efficiency in the merging area and minimize vehicle travel time within the merging area, the merging area travel time is calculated. The calculation of the merging zone travel time The expression is as follows:
[0113] ;
[0114] in, This represents the total travel time of all vehicles passing through the merging zone within one cycle. This indicates the number of vehicles passing through the merging zone within one cycle. This indicates the vehicle's number in the merging area. Indicates the first The travel time of a vehicle through the merging zone;
[0115] To reduce travel time in the merging zone, it's necessary to suppress traffic flow entering the merging zone, which could lead to congestion upstream of the ramp. Therefore, to encourage the system to release as many vehicles as possible, the total travel time should be factored in, minimizing the total travel time of each vehicle. The objective is to maximize the number of vehicles leaving the ramp system, and the calculation of the total travel time is... The expression is as follows:
[0116] ;
[0117] in, This indicates the number of vehicles leaving the merging zone within one cycle. This indicates the number of vehicles entering the merging zone within a single cycle.
[0118] To improve traffic flow efficiency in merging areas while reducing excessive queue lengths on ramps, thereby minimizing vehicle dwell time on ramps. The objective is to calculate the ramp dwell time. The expression is as follows:
[0119] ;
[0120] in, This indicates the time a vehicle spends on a ramp within a given cycle. This indicates the number of vehicles on the ramp within one cycle. This indicates the vehicle's number on the ramp. Indicates the first The travel time for a vehicle to pass through the ramp;
[0121] In summary, based on minimizing the merging zone travel time, minimizing the total travel time, and minimizing the ramp dwell time, a reward function is constructed, the expression of which is shown below:
[0122] ;
[0123] in, Represents the reward function, This indicates the merging zone travel time for all vehicles within a cycle. This indicates the total travel time of the vehicle. This indicates the time a vehicle spends on a ramp within a given cycle. This represents the weighting coefficient of the sub-rewards related to traffic efficiency in the merging zone. This represents the weighting coefficient of the sub-rewards related to the total travel time in the merging zone. This represents the weight coefficients corresponding to the sub-reward function related to the ramp queue length.
[0124] S3. Based on the required scenario, road test unit, and speed limit panel, and combined with a single selection action, train the speed limit control model to obtain the trained speed limit control model. The specific steps are as follows:
[0125] S301. Based on the required scenario, set model parameters including the number of neural network layers, number of neurons, discount factor, learning rate, batch size, update frequency of the competing target network, initial exploration rate, decay rate of the initial exploration rate per round, and minimum exploration rate to obtain the set rate limiting control model.
[0126] S302. Based on the preset traffic simulation, the speed limit control model is iterated and the initial traffic environment state is obtained using a loop detector.
[0127] S303. Based on a greedy strategy, using a preset probability. For a single selection action, a random selection is made, or probability is used. Based on the status, the combined actions of the main line and the ramp are output to obtain the output action;
[0128] S304. Convert the output action into a speed limit value, use the road test unit and speed limit panel to publish the speed limit value to the intelligent connected vehicles traveling on the road segment, and then obtain the reward value and the next state space from the ramp merging area environment.
[0129] S305. Combine the current state space, output action, reward value and next state space into a quadruple, and store the quadruple as an experience sample in the experience replay pool.
[0130] S306. Based on the experience replay pool, experience samples are extracted using the priority experience replay mechanism. The online network in the speed limit control model is updated using the loss function, and the target network parameters are updated using soft update to obtain the trained speed limit control model.
[0131] In this embodiment, as Figure 4 As shown, the algorithm structure for training the speed limit control model is illustrated. Based on the required scenario, the model parameters for the speed limit control model are set, including: the number of neural network layers. Number of neurons Discount Factor Learning rate Batch processing size Competition target network update frequency Initial exploration rate Initial exploration rate per round decay rate and minimum exploration rate ;
[0132] Based on a pre-set traffic simulation model (SUMO), the initial traffic environment state is obtained from loop detectors deployed on the road network through iterative processing using a speed limit control model. ;
[0133] based on Greedy strategy, using an agent to select output actions Specifically, this means: setting the initial exploration rate As a preset probability, using the preset probability For a single selection action, randomly select one action, or use probability. Output the combined actions of the main line and ramps based on the status;
[0134] Output action This is converted into a speed limit value, which is then published to connected vehicles traveling on the road segment via RSU and VMS, and finally, a reward value is obtained from the ramp merging area environment. and the next state space ;
[0135] Combine the current state space, output action, reward value, and next state space into a quadruple. The quadruple is stored as an experience sample in the experience replay pool. Experience samples are extracted from the experience replay pool using a priority experience replay mechanism, and the online network parameters are updated according to the loss function. and use soft update Update target network parameters .
[0136] In this embodiment, the reward value obtained by the agent in each training round reflects the quality of its training; a higher reward value indicates a better training effect. A simulation training of 200 rounds was conducted on the variable speed limiting cooperative control model in the merging zone. The trend of the cumulative reward value in each round is as follows: Figure 5 As shown, by Figure 5 It can be seen that under the variable speed control strategy, as the number of training rounds increases, the cumulative reward value of each round shows an upward trend, indicating that the agent gradually learns useful strategies during the training process. The training begins to converge around the 80th round, and the final reward value stabilizes at around -430.
[0137] S4. Using the trained speed limit control model, simulation is performed to obtain the control strategy for the merging zone of the ramp, and the coordinated variable speed limit control of the merging zone is completed.
[0138] In this embodiment, a trained speed-limiting control model is used for simulation to obtain the following results: Figure 6 The speed limit structure shown, in the variable speed limit collaborative control scenario of the ramp merging area, exhibits a trend of first decreasing and then increasing speed limits on the main road as traffic demand gradually changes, while the speed limits on the ramps also show a trend of first decreasing and then increasing. The overall variable speed limit trend for both the main road and the ramps shows a trend of first decreasing and then increasing, thus revealing the control strategy for the ramp merging area. The control strategy for the ramp merging area is a variable speed limit collaborative control strategy. Under this strategy, the speed limit on the main road decreases significantly, indicating that limiting the speed on the main road can better improve the traffic efficiency of the merging area. However, the speed limit on the ramp decreases less compared to the speed limit on the main road, indicating that the speed limit setting on the ramps not only ensures that the speed limit is reduced to increase the capacity of the merging area but also guarantees a certain minimum speed limit to avoid ramp congestion.
[0139] The effectiveness of the proposed control strategy for ramp merging areas was evaluated, and traffic flow characteristics were compared and analyzed under the control method of this application. Without variable speed limit control at ramp merging areas, the average travel time was 92.86 s and the average speed was 17.11 m / s. Under real-time variable speed limit control at ramp merging areas, the average travel time was 88.34 s and the average speed was 17.65 m / s. The average travel time decreased by 4.87% and the average speed increased by 3.15%. At the merging area, without variable speed limit control, the total travel time was 66.76 h, while with variable speed limit control, the total travel time was 63.00 h, a decrease of 5.63%. In conclusion, the control strategy generated in this application can alleviate congestion in merging areas to a certain extent and improve traffic efficiency at ramp merging areas.
[0140] In this embodiment, the present invention addresses the interaction between traffic flow on the main road and traffic flow on the ramps in the merging zone of highway ramps. It studies a collaborative variable speed limit control method for merging zones based on deep reinforcement learning. Existing highway variable speed limit control mainly focuses on single control of traffic flow on the main road. This invention, based on the mixed traffic flow environment, combines ramp control and variable speed limits on the main road to control the entire merging zone. The variable speed limit control method is applied to the ramps, discretizing and combining the action spaces of the ramps and the main road. Simultaneously, the reward function incorporates the efficiency of the merging zone and the queue length on the ramps. Finally, simulations are used to verify the model's effectiveness under a 30% connected vehicle penetration rate. This provides rational suggestions for improving merging zone traffic efficiency and operational status.
[0141] In this embodiment, the variable speed limit problem of the main road and ramps in the merging zone is transformed into a Markov decision process. At the same time, the speed limit values of the main road and ramps are combined with the design action space, and the reward function is considered to balance the traffic efficiency of the merging zone and the queue length of the ramps. The DDQN reinforcement learning algorithm based on the priority experience replay technique is used for continuous training and learning to obtain the control method of variable speed limit of the main road and ramps.
[0142] The scenario in this embodiment is a merging zone on a highway. The ramps in this area do not have signal control, but like the main road, they have variable speed limit control. This section includes a 3-lane main road and 1 entrance ramp, and the merging zone has 4 lanes. The maximum speed limit on the main road is 120 km / h, and the speed limit on the ramps is 50 km / h. The variable speed limit control system transmits the collected real-time traffic status information to the intelligent control system by setting up detection points and placing detectors in the road network. The intelligent control system uses a merging zone collaborative variable speed limit control method based on deep reinforcement learning to determine the speed limit strategies for the main road and ramps, and transmits the speed limit control information to the vehicles in the CAV and the speed limit panel within the control area.
Claims
1. A collaborative variable velocity limiting control method for merging regions based on deep reinforcement learning, characterized in that, Includes the following steps: S1. Based on the simulation environment of the ramp merging area, deploy loop detectors, use access interfaces to obtain traffic status information, set up demand scenarios, and set up road test units and speed limit panels to transmit speed limit information of the main line and ramps. S2. Based on traffic state information, define the state space, combine the actions of the mainline and ramps to obtain the action space and the single-selection action, introduce ramp dwell time, construct a reward function, and obtain the speed limit control model, specifically: S201. Define the state space based on the speed limit information of the mainline and ramps in the traffic state information, as well as the occupancy and speed of the mainline, merging zone, and ramps. S202. Based on the relationship between the mainline speed limit and the mainline actions, as well as the relationship between the ramp speed limit and the ramp actions, obtain the action space, set speed limit constraints, and obtain a single selected action. S203. Calculate the merging zone travel time and total travel time, and incorporate ramp dwell time to construct a reward function, specifically: S2031. The merging zone travel time is calculated based on the number of vehicles passing through the merging zone within a cycle. S2032. The total travel time is calculated based on the number of vehicles leaving the merging zone in one cycle and the number of vehicles entering the merging zone in one cycle. S2033. Introducing ramp dwell time: The ramp dwell time is calculated based on the number of vehicles passing through the merging zone within one cycle. S2034. Construct a reward function based on minimizing the merging zone travel time, minimizing the total travel time, and minimizing the ramp dwell time. S204. Integrate the state space, action space, and reward function to obtain the speed limit control model; S3. Based on the demand scenario, road test unit, and speed limit panel, and combined with a single selection action, train the speed limit control model to obtain the trained speed limit control model. S4. Using the trained speed limit control model, simulation is performed to obtain the control strategy for the merging zone of the ramp, and the coordinated variable speed limit control of the merging zone is completed.
2. The method for cooperative variable speed limiting control in merging regions based on deep reinforcement learning according to claim 1, characterized in that, S1 includes the following steps: S101. Based on the merging area of highway ramps, obtain the simulation environment of the merging area through simulation; S102. Deploy loop detectors in the simulation environment of the ramp merging area and use the access interface to obtain the traffic status information of the ramp merging area system in real time. S103. By pre-setting the mainline traffic and ramp traffic, set up demand scenarios including low demand scenarios, moderate demand scenarios and high demand scenarios. S104. Install road test units and speed limit panels for transmitting speed limit information for mainlines and ramps.
3. The method for cooperative variable speed limiting control in merging regions based on deep reinforcement learning according to claim 1, characterized in that, The expression for the state space is as follows: in, and These represent the market share and speed of the upstream of the main line, respectively. and These represent the occupancy and velocity at the front of the merging zone, respectively. and These represent the occupancy and velocity in the middle of the merging zone, respectively. and These represent the occupancy rate and velocity at the rear of the merging zone, respectively. and These represent the occupancy rate and velocity measured by the detector in the bottleneck area downstream of the main line, respectively. and These represent the occupancy rate and speed on the ramp, respectively, and can be measured by the entrance ramp detector. Indicates period t Speed limit information for the main road. Indicates period t Speed limit information for ramps.
4. The method for cooperative variable speed limiting control in merging regions based on deep reinforcement learning according to claim 1, characterized in that, S202 includes the following steps: S2021. Based on the relationship between the mainline speed limit and the mainline action, and combined with the minimum value of the mainline speed limit, seven mainline actions are obtained. S2022. Based on the relationship between ramp speed limits and ramp actions, and combined with the minimum value of the ramp speed limit, five types of ramp actions are obtained. The main line actions and ramp actions are combined to obtain a space of thirty-five actions. S2023. Based on the speed limit changes of the main line and ramps in adjacent cycles, speed limit constraints are set to obtain the constrained action space of the main line and ramps, resulting in nine single-selection actions.
5. The method for cooperative variable speed limiting control in merging regions based on deep reinforcement learning according to claim 1, characterized in that, The expression for the speed limit constraint is as follows: in, Indicates period t Speed limit information for the main road. Indicates period Speed limit information for the main road. Indicates period t Speed limit information for ramps, Indicates period Speed limit information for ramps, cycle time For period t The previous cycle.
6. The method for cooperative variable speed limiting control in merging regions based on deep reinforcement learning according to claim 1, characterized in that, The expression for the reward function is as follows: in, Represents the reward function, This indicates the merging zone travel time for all vehicles within a cycle. Indicates the total travel time of the vehicle. This indicates the time a vehicle spends on a ramp within a given cycle. This indicates the number of vehicles on the ramp within one cycle. This indicates the vehicle's number on the ramp. Indicates the first The travel time for a vehicle to pass through the ramp. This represents the weighting coefficient of the sub-rewards related to traffic efficiency in the merging zone. This represents the weighting coefficient of the sub-rewards related to the total travel time in the merging zone. This represents the weight coefficients corresponding to the sub-reward function related to the ramp queue length.
7. The method for cooperative variable speed limiting control in merging regions based on deep reinforcement learning according to claim 1, characterized in that, S3 includes the following steps: S301. Based on the required scenario, set model parameters including the number of neural network layers, number of neurons, discount factor, learning rate, batch size, update frequency of the competing target network, initial exploration rate, decay rate of the initial exploration rate per round, and minimum exploration rate to obtain the set rate limiting control model. S302. Based on the preset traffic simulation, the speed limit control model is iterated and the initial traffic environment state is obtained using a loop detector. S303. Based on a greedy strategy, using a preset probability. For a single selection action, a random selection is made, or probability is used. Based on the status, the combined actions of the main line and the ramp are output to obtain the output action; S304. Convert the output action into a speed limit value, use the road test unit and speed limit panel to publish the speed limit value to the intelligent connected vehicles traveling on the road segment, and then obtain the reward value and the next state space from the ramp merging area environment. S305. Combine the current state space, output action, reward value and next state space into a quadruple, and store the quadruple as an experience sample in the experience replay pool. S306. Based on the experience replay pool, experience samples are extracted using the priority experience replay mechanism. The online network in the speed limit control model is updated using the loss function, and the target network parameters are updated using soft updates to obtain the trained speed limit control model.
Citation Information
Patent Citations
Highway differential variable speed limit control method and device based on deep reinforcement learning and storage medium
CN117496721A
Secondary accident variable speed limit prevention and control method based on deep reinforcement learning
CN118430285A