Lane-level mixed traffic flow cooperative control method and system based on deep reinforcement learning

Through the lane-level mixed traffic collaborative control method based on deep reinforcement learning, the vehicle speed limit is dynamically adjusted, which solves the traffic optimization challenges of traditional methods in mixed traffic environments and improves the traffic efficiency and safety of the merging area.

CN120412291BActive Publication Date: 2025-09-19BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510912699.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-19
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional traffic control methods have difficulty accurately responding to the interaction between autonomous vehicles and manually driven vehicles in mixed traffic environments, resulting in challenges in optimizing traffic in merging areas, especially when there are large differences in traffic volume and speed between lanes.

Method used

A lane-level mixed traffic cooperative control method based on deep reinforcement learning is adopted. By constructing a multi-dimensional state vector and reward function, the vehicle speed limit is dynamically adjusted. Combined with the Markov decision process and two-layer network control strategy, the coordinated control of traffic flow between lanes is optimized.

Benefits of technology

It achieves precise control in dynamic traffic scenarios, improves traffic efficiency and safety in merging areas, adapts to the traffic characteristics of different lanes, and reduces traffic congestion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412291B_ABST
    Figure CN120412291B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of transportation information engineering, and discloses a lane-level mixed traffic collaborative control method and system based on deep reinforcement learning. The method takes into account the flow rate difference and speed difference between lanes, and is embedded in the reward and punishment function of deep reinforcement learning. According to the distribution of mixed traffic in different lanes, a lane-level double-layer network control strategy is constructed, which allows it to dynamically generate different vehicle speed limits for each lane based on the traffic flow, speed and relationship between the current lanes. The deep reinforcement learning model is used to perform feedback optimization at each moment, continuously adapt to traffic changes, and thus accurately respond to the operating needs of each lane, and is empirically studied. The deep reinforcement learning model can achieve a lane-level mixed traffic collaborative control effect that is better than the baseline model, can effectively alleviate the congestion situation in the merging area of ​​the highway, and has cross-scenario applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of transportation information engineering and discloses a lane-level mixed traffic collaborative control method and system based on deep reinforcement learning. Background Art

[0002] Freeway merging areas are critical for traffic flow management, and their dynamic and complex traffic characteristics often lead to severe traffic congestion. In mixed traffic environments where autonomous vehicles (CAVs) and human-driven vehicles (HVs) coexist, traditional traffic control methods struggle to achieve efficient coordination due to differences in driving behaviors and response mechanisms between the two types of vehicles, making traffic optimization in merging areas even more challenging.

[0003] The primary traffic conflict in merging areas stems from the merging conflict between ramp traffic and main road traffic. Variable speed limit (VSL) technology is one effective means of alleviating this problem. Traditional VSL control methods rely primarily on feedback control or model predictive control (MPC). The former dynamically adjusts speed limits by setting density thresholds, while the latter mitigates congestion based on multi-factor optimization. However, these methods lack adaptability in mixed traffic environments and struggle to accurately address the interactions between CAVs and HVs, limiting their control effectiveness.

[0004] In recent years, deep reinforcement learning (DRL) has demonstrated significant advantages in the field of VSL control due to its strong environmental adaptability and autonomous learning capabilities. Compared with traditional methods, DRL can autonomously optimize decision-making strategies based on real-time traffic conditions, thereby achieving more precise control in dynamically changing traffic scenarios. Existing studies have shown that DRL-based VSL strategies can effectively improve the traffic efficiency of merging areas, but existing studies mostly focus on single-lane scenarios and fail to fully consider the differences in traffic flow between lanes and their impact on overall efficiency. Especially in mixed traffic environments, the uneven speed distribution of inner and outer lanes, the response differences between CAVs and HVs, and changes in CAV penetration all place higher demands on the collaborative control methods of mixed traffic flows. Summary of the Invention

[0005] To overcome the limitations of existing technologies that ignore flow or speed differences between lanes and inaccurately publish vehicle speed limits, the present invention proposes a lane-level mixed traffic collaborative control method and system based on deep reinforcement learning. By constraining the flow and speed differences between lanes, embedding a deep reinforcement learning reward and penalty function, and constructing a variable speed limit strategy based on the traffic distribution of different lanes, this method achieves collaborative control of lane-level mixed traffic in the merging area of ​​highways.

[0006] The lane-level mixed traffic flow cooperative control method based on deep reinforcement learning provided by the present invention is characterized by comprising:

[0007] The roadside detectors sense the vehicle flow, density, and speed data upstream and downstream of the main line of the highway merging area, and construct a multi-dimensional state vector , which not only reflects the heterogeneity of traffic space, but also ensures the temporal continuity of speed adjustment. is the vehicle density upstream of the main line of the highway merging area, is the density of the upstream confluence area, is the vehicle density of the highway merging area ramp, is the vehicle density downstream of the main line of the expressway merging area, is the historical vehicle speed of the previous control cycle;

[0008] Defining the action space is the vehicle speed upper limit set, is the upper limit of the speed of a vehicle in A, and the value range of a is between The difference in vehicle speed limits between the same lane and adjacent lanes during adjacent time periods shall not exceed 20 km / h;

[0009] Select Markov decision process to model the state transition probability between vehicle speed limits in adjacent time periods , providing an environment interaction mechanism for deep reinforcement learning models;

[0010] A deep reinforcement learning reward function is defined, with the minimum total travel time as the objective function and the minimum difference in traffic volume and speed between the upstream and downstream main lines of the merging area as the reward conditions, to collaboratively control the operation of traffic flow.

[0011] Construct a deep reinforcement learning model to model the Markov decision process. Based on the traffic volume, speed, and relationship between lanes, a different vehicle speed limit is dynamically generated for each lane. The deep reinforcement learning model is used to perform feedback optimization at every moment, continuously adapting to traffic changes and dynamically updating the action space. , accurately responding to the operational needs of each lane;

[0012] Using the SUMO platform to build a simulation environment, the penetration rate of intelligent connected vehicles in mixed traffic was changed, and the performance of the model was evaluated and verified on two datasets.

[0013] According to the lane-level mixed traffic cooperative control method based on deep reinforcement learning provided by the present invention, the state transition probability between the vehicle speed upper limits in adjacent time periods is modeled by selecting the Markov decision process. ,include:

[0014] Use the SUMO platform to simulate the traffic behavior of mixed traffic flows and model the state transition probability between vehicle speed limits in adjacent time periods. , where the state transition process is state and actions To the next state process of change.

[0015] The lane-level mixed traffic cooperative control method based on deep reinforcement learning provided by the present invention further includes defining a deep reinforcement learning reward function, including:

[0016] Minimize the total travel time TTS of vehicles in the merging area within the time range K, where TTS is expressed by formula (1):

[0017] ; (1)

[0018] in, represents the simulation time step; K represents the overall time interval, k is a certain time interval, t represents a specific moment, represents the initial number of vehicles in the road section at the beginning of the simulation, represents the number of vehicles entering the road section at time t, represents the number of vehicles leaving the road section at time t; is the lower limit of the time corresponding to the time interval k; is the upper limit of the time corresponding to the time interval k;

[0019] Introduce the traffic flow difference between the upstream and downstream of the main line of the merging area at time t as a sub-item of the reward function , Expressed by formula (2):

[0020] ; (2)

[0021] Introducing the speed difference between the upstream and downstream vehicles in the merging area at time t as another component of the reward function , Expressed by formula (3):

[0022] ; (3)

[0023] in, represents the speed of vehicle i in the merging area at time t, represents the number of vehicles in the merging area at time t, H is the number of time steps in the control cycle, is the vehicle speed under free flow conditions;

[0024] The total reward function is obtained by taking the minimum difference in traffic volume and speed between the upstream and downstream of the main line in the merging area as the reward condition. , collaboratively control traffic flow operation status, among which, It is defined as the following formula:

[0025] ; (4)

[0026] in, 、 They represent the difference in traffic volume and speed between the upstream and downstream of the main line in the merging area in the total reward function. The corresponding weights in .

[0027] According to the lane-level mixed traffic cooperative control method based on deep reinforcement learning provided by the present invention, the construction of the deep reinforcement learning model includes:

[0028] The deep reinforcement learning model includes a first neural network model and a second neural network model. The first neural network model is used to evaluate the action value under the state, and the second neural network model is used to select the action. The deep reinforcement learning model models the Markov decision process and dynamically generates different vehicle speed limits for each lane based on the traffic volume, speed and relationship between the current lanes. The deep reinforcement learning model is used to perform feedback optimization at every moment, continuously adapt to traffic changes, and dynamically update the action space. , accurately responding to the operational needs of each lane;

[0029] During the training process of the deep reinforcement learning model, the interaction and feedback between the online network and the target network are utilized to improve the stability and robustness of the model in complex traffic flow environments, and stable training is achieved through the dual-network interaction system of the online network and the target network;

[0030] The online network adopts the Dueling architecture to achieve decoupled evaluation of state value, which conforms to the following formula Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))), where Q(s,a) is the action driving function representing the comprehensive expected benefit obtained by performing action a in state s, V(s) is the state value function, A(s,a) is the advantage function, and meanA(s,a) is the mean of the advantages of all optional actions in the current state. V(s) and A(s,a) respectively learn the macroscopic state characteristics and microscopic action benefits of traffic flow. The macroscopic state characteristics include lane density and vehicle speed distribution, and the microscopic action benefits include lane changing and acceleration and deceleration benefits. The target network suppresses Q value overestimation through periodic parameter synchronization.

[0031] Combining the Dueling structure's ability to generalize sparse rewards, it achieves a coordinated improvement in global traffic efficiency optimization and local action decision-making accuracy in lane merging areas in dynamic traffic scenarios. Ultimately, through delayed updates and a decoupled evaluation mechanism, it ensures stable policy gradient convergence in complex road conditions with frequent interactions and unbalanced rewards.

[0032] The lane-level mixed traffic flow cooperative control method based on deep reinforcement learning provided by the present invention further includes: constructing a lane-level double-layer network control strategy based on the dynamic change characteristics of mixed traffic flow in the merging area of ​​the highway;

[0033] The lane-level double-layer network control strategy includes:

[0034] Speed ​​smoothing is used for the inner high-traffic lane to stabilize traffic flow, while speed coordination is implemented for the outer lane with frequent lane changes to optimize lane change intervals.

[0035] By establishing a coordinated speed optimization model between lanes to minimize speed differences and setting up a dynamic response mechanism, the speed limit plan can be adjusted in real time when sudden changes in traffic conditions are detected, thereby improving the overall traffic efficiency of the merging area.

[0036] The lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the present invention uses the SUMO platform to build a simulation environment, changes the penetration rate of intelligent connected vehicles in the mixed traffic flow, and uses two data sets to perform performance evaluation and case verification of the model under different penetration environments, including:

[0037] The simulation environment was built using the SUMO platform. During the simulation process, a variety of mixed traffic flow scenarios were simulated, including traffic flows consisting entirely of manually driven vehicles, traffic flows consisting entirely of autonomous vehicles, and traffic flows consisting of a mixture of manually driven vehicles and autonomous vehicles.

[0038] Modeling different types of traffic flows. The simulation platform provides detailed data on lane flow, vehicle speed, and traffic density, generating simulation results that are used to evaluate the adaptability and effectiveness of the deep reinforcement learning model in different traffic environments.

[0039] By comparing with traditional control methods, the performance of the deep reinforcement learning model under mixed traffic conditions was evaluated, and the advantages of the deep reinforcement learning model in improving the traffic capacity of the merging area, reducing traffic congestion and optimizing vehicle speed control were verified.

[0040] The present invention also provides a lane-level mixed traffic cooperative control system based on deep reinforcement learning, comprising:

[0041] The data acquisition and state construction module is used to perceive the vehicle flow, vehicle density, and vehicle speed data upstream and downstream of the main line of the highway merging area in real time through roadside detectors; construct a multi-dimensional state vector based on the perceived data; obtain lane-level flow and speed information in real time, and model the state transition probability between vehicle speed limits in adjacent time periods based on the Markov decision process, providing an environmental interaction mechanism for the deep reinforcement learning model.

[0042] A model building module is used to build a lane-level mixed traffic collaborative control model based on a deep reinforcement learning model. This lane-level mixed traffic control model aims to achieve optimal traffic efficiency in highway merging areas and completes model training and testing. Through a dual-network architecture, this lane-level mixed traffic control model can perform more precise control under different mixed traffic penetration conditions, improving the overall efficiency of the merging area.

[0043] Stable training is achieved through a dual-network interaction system consisting of an online network and a target network. The online network adopts the Dueling architecture to achieve decoupled evaluation of state value, decomposing the Q value into a state-value function V(s) and an advantage function A(s,a), respectively learning the macroscopic state characteristics and microscopic action benefits of the traffic flow. The aggregation layer then generates Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))). The target network synchronously suppresses Q-value overestimation through periodic parameters. Combined with the Dueling structure's ability to generalize sparse rewards, the system achieves synergistic improvements in global traffic efficiency optimization and local action decision-making accuracy in dynamic traffic scenarios in merging areas. Ultimately, through delayed updates and a decoupled evaluation mechanism, the system ensures stable convergence of the vehicle's policy gradient in complex road conditions with frequent interactions and uneven rewards.

[0044] An example verification module is used to input the highway merging area dataset into a lane-level mixed traffic collaborative control model based on a deep reinforcement learning model to realize example verification of mixed traffic collaborative control.

[0045] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements any of the above-mentioned lane-level mixed traffic collaborative control methods based on deep reinforcement learning.

[0046] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned lane-level mixed traffic collaborative control methods based on deep reinforcement learning.

[0047] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned lane-level mixed traffic collaborative control methods based on deep reinforcement learning.

[0048] This paper proposes a lane-level mixed traffic collaborative control method based on deep reinforcement learning. This method innovatively introduces a traffic balance reward mechanism, with minimizing the difference in the number of vehicles entering and exiting the merging area as the optimization goal. At the same time, it dynamically adjusts the speed limit value based on factors such as real-time speed, flow difference, and CAV penetration rate. Through lane-level micro-control, the method can accurately identify the traffic characteristics of different lanes (such as high flow in the inner lane and frequent lane changes in the outer lane) and implement differentiated speed limit control accordingly, thereby effectively balancing the speed distribution between lanes and improving the overall traffic efficiency and safety of the merging area. Simulation experiments have verified the superior performance of the method in a mixed traffic environment, providing a new approach for the collaborative control of mixed traffic in merging areas of highways. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 This is one of the application flow diagrams of the lane-level mixed traffic flow cooperative control method based on deep reinforcement learning provided by an embodiment of the present invention;

[0051] Figure 2 This is the second schematic diagram of the application process of the lane-level mixed traffic cooperative control method based on deep reinforcement learning provided by an embodiment of the present invention;

[0052] Figure 3 Schematic diagram of the structure of a lane-level mixed traffic cooperative control system based on deep reinforcement learning provided by an embodiment of the present invention;

[0053] Figure 4 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0055] Figure 1 This is one of the application flow diagrams of the lane-level mixed traffic collaborative control method based on deep reinforcement learning provided by an embodiment of the present invention.

[0056] Figure 2 This is the second application flow diagram of the lane-level mixed traffic collaborative control method based on deep reinforcement learning provided by an embodiment of the present invention.

[0057] like Figure 1 and Figure 2 As shown, this embodiment provides a lane-level mixed traffic cooperative control method based on deep reinforcement learning, including:

[0058] Step 101: Use roadside detectors to sense the vehicle flow, vehicle density, and vehicle speed data upstream and downstream of the main line of the highway merging area, and construct a multi-dimensional state vector , which not only reflects the heterogeneity of traffic space, but also ensures the temporal continuity of speed adjustment. is the vehicle density upstream of the main line of the highway merging area, is the density of the upstream confluence area, is the vehicle density of the highway merging area ramp, is the vehicle density downstream of the main line of the expressway merging area, is the historical vehicle speed of the previous control cycle; the multidimensional state vector defined It can not only reflect the heterogeneity of traffic space, but also ensure the temporal continuity of speed adjustment;

[0059] Step 102: Define the action space is the vehicle speed upper limit set, is the upper limit of the speed of a vehicle in A, and the value range of a is between The difference in vehicle speed limits between the same lane and adjacent lanes during adjacent time periods shall not exceed 20 km / h;

[0060] Select Markov decision process to model the state transition probability between vehicle speed limits in adjacent time periods , providing an environment interaction mechanism for deep reinforcement learning models;

[0061] Step 103: Define a deep reinforcement learning reward function, with the minimum total travel time as the objective function and the minimum difference in traffic volume and speed between the upstream and downstream main lines of the merging area as the reward conditions, to collaboratively control the traffic flow operation status.

[0062] Construct a deep reinforcement learning model to model the Markov decision process. Based on the traffic volume, speed, and relationship between lanes, a different vehicle speed limit is dynamically generated for each lane. The deep reinforcement learning model is used to perform feedback optimization at every moment, continuously adapting to traffic changes and dynamically updating the action space. , accurately responding to the operational needs of each lane;

[0063] The deep reinforcement learning model includes a first neural network model and a second neural network model. The first neural network model is used to evaluate the value of actions under a state, and the second neural network model is used to select actions. The deep reinforcement learning model dynamically generates different vehicle speed limits for each lane based on the current lane's traffic volume, speed, and the relationship between lanes. The deep reinforcement learning model also performs feedback optimization at every moment, continuously adapting to changing traffic conditions and accurately responding to the operating needs of each lane.

[0064] During the training process of the deep reinforcement learning model, the interaction and feedback between the online network and the target network are utilized to improve the stability and robustness of the model in complex vehicle traffic environments, and stable training is achieved through the dual-network interaction system of the online network and the target network;

[0065] The online network adopts the Dueling architecture to achieve decoupled evaluation of state value, which conforms to the following formula Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))), where Q(s,a) is the action driving function representing the comprehensive expected benefit obtained by performing action a in state s, V(s) is the state value function, A(s,a) is the advantage function, and meanA(s,a) is the mean of the advantages of all optional actions in the current state. V(s) and A(s,a) respectively learn the macro-state characteristics and micro-action benefits of the mixed traffic flow. The macro-state characteristics include lane density and speed distribution, and the micro-action benefits include lane changing and acceleration and deceleration benefits. The target network suppresses Q value overestimation through periodic parameter synchronization.

[0066] Combining the Dueling structure's ability to generalize sparse rewards, it achieves a coordinated improvement in global traffic efficiency optimization and local action decision-making accuracy in lane merging areas in dynamic traffic scenarios. Ultimately, through delayed updates and a decoupled evaluation mechanism, it ensures stable policy gradient convergence in complex road conditions with frequent interactions and unbalanced rewards.

[0067] Step 104 : Use the SUMO platform to build a simulation environment, change the penetration rate of intelligent connected vehicles in the mixed traffic flow, and use two data sets to perform performance evaluation and case verification of the model under different penetration rate environments.

[0068] The simulation verification was carried out using the SUMO (Simulation of Urban Mobility) simulation platform. During the simulation process, a variety of mixed traffic scenarios were simulated, including traffic composed entirely of human-driven vehicles (HVs), traffic composed entirely of autonomous vehicles (CAVs), and a mixture of traffic composed of human-driven vehicles and traffic composed of autonomous vehicles. By modeling these different types of traffic, the simulation platform can provide detailed data on lane flow, vehicle speed, traffic density, etc., and obtain simulation results. The simulation results are used to evaluate the adaptability and effectiveness of the deep reinforcement learning model in different traffic environments. By comparing with traditional control methods, the performance of the deep reinforcement learning model in a mixed traffic environment is evaluated, and its advantages in improving the traffic capacity of the merging area, reducing traffic congestion and optimizing vehicle speed control are verified.

[0069] This embodiment provides a lane-level mixed traffic collaborative control method based on deep reinforcement learning. This method innovatively introduces a traffic balance reward mechanism, with the optimization goal of minimizing the difference in the number of vehicles upstream and downstream of the main line in the merging area. It also dynamically adjusts the speed limit based on factors such as real-time speed, traffic volume differences, and CAV penetration. Through lane-level micro-control, the strategy can accurately identify the traffic characteristics of different lanes (such as high traffic volume in the inner lane and frequent lane changes in the outer lane) and implement differentiated speed limits accordingly, effectively balancing the speed distribution between lanes and improving the overall traffic efficiency and safety of the merging area. Simulation experiments have verified the superior performance of this strategy in mixed traffic environments, providing a new approach for the collaborative control of mixed traffic in merging areas on highways.

[0070] In an exemplary embodiment, the method further includes defining a deep reinforcement learning reward function, including:

[0071] Minimize the total travel time TTS of vehicles in the merging area within the time range K, where TTS is expressed by formula (1):

[0072] ; (1)

[0073] in, represents the simulation time step; K represents the overall time interval, k is a certain time interval, t represents a specific moment, represents the initial number of vehicles in the road section at the beginning of the simulation, represents the number of vehicles entering the road section at time t, represents the number of vehicles leaving the road section at time t; is the lower limit of the time corresponding to the time interval k; is the upper limit of the time corresponding to the time interval k.

[0074] Introduce the traffic flow difference between the upstream and downstream of the main line of the merging area at time t as a sub-item of the reward function , Expressed by formula (2):

[0075] ; (2)

[0076] Introducing the speed difference between the upstream and downstream vehicles in the merging area at time t as another component of the reward function , Expressed by formula (3):

[0077] ; (3)

[0078] in, represents the speed of vehicle i in the merging area at time t, represents the number of vehicles in the merging area at time t, H is the number of time steps in the control cycle, is the vehicle speed under free flow conditions;

[0079] The total reward function is obtained by taking the minimum difference in traffic volume and speed between the upstream and downstream of the main line in the merging area as the reward condition. , collaboratively control traffic flow operation status, among which, It is defined as the following formula:

[0080] ; (4)

[0081] in, 、 They represent the difference in traffic volume and speed between the upstream and downstream of the main line in the merging area in the total reward function. The corresponding weights in .

[0082] Spatiotemporal distribution of vehicle arrivals Independent of the control measures, this means that only the number of vehicles leaving the road section can be directly controlled. By increasing the number of vehicles leaving the road section per unit time, the TTS can be reduced, the speed limit control target can be achieved, and the road section can be optimized. Vehicles entering the road section within a specific time period can be detected by deploying detectors at key points on the road section. Based on the above settings of the speed limit control target and detection method, the difference in traffic flow upstream and downstream of the main line of the merging area at time t is introduced as a sub-item of the reward function ;

[0083] The inflow and outflow volume alone is not enough to fully characterize the traffic conditions in the merging area. For example, when the downstream section of the merging area is congested, vehicles move slowly and the difference in inflow and outflow volume may be small. However, in this case, the low speed of vehicles leads to a decrease in overall traffic efficiency. To address this limitation, this embodiment introduces the speed difference between the upstream and downstream main lines of the merging area at time t as another component of the reward function. .

[0084] The total reward function is obtained by taking the minimum difference in traffic volume and speed between the upstream and downstream of the main line in the merging area as the reward condition. , collaboratively control traffic flow operation status.

[0085] The following is an example of a specific embodiment to illustrate the group classification method provided by the solution of this application.

[0086] (1) Working environment

[0087] The present invention uses the PyTorch framework to write the model code. All experimental codes are run on a Windows 10 workstation equipped with a "GenInter(R) Corei7-13700F @ 2.10GHz" CPU and an "NVIDIA GeForce RTX 4090 (128 GB RAM)" GPU.

[0088] (2) Introduction to public datasets

[0089] The experiment uses datasets A and B as two inputs respectively.

[0090] Dataset A is a simulated dataset, primarily used to simulate vehicle merging on the upstream section of the main road and on the on-ramp. Traffic flow parameters are set as follows: the mainline input flow rate is [2400, 2700, 3000, 2700, 2400] vehicles / hour, and the ramp input flow rate is [400, 600, 800, 600, 400] vehicles / hour. Vehicle paths are randomly generated using SUMO software, and their arrival times are randomized according to a binomial distribution.

[0091] Dataset B is the open-source dataset exiD, which primarily records the trajectories of manually driven vehicles in seven different highway merging areas. This paper specifically selected data from this dataset related to highway merging areas for experimental analysis.

[0092] (3) Experimental parameter setting

[0093] To simulate a realistic mixed-traffic environment, this study used the SUMO (Simulation of Urban Traffic) software platform, which has outstanding advantages in modeling the operating environment of autonomous vehicles (CAVs). In the simulation, the behavior of human-driven vehicles (HVs) was implemented using the Intelligent Driver Model (IDM) and the LC2013 lane change model, while the autonomous vehicles were simulated using the Adaptive Cruise Control (ACC) and Cooperative Adaptive Cruise Control (CACC) models.

[0094] The road network in SUMO is configured as follows: the basic section is 1600 meters long, with a 400-meter variable speed control zone, followed by a 200-meter acceleration zone and a 600-meter dissipation zone. The entrance ramp is located at 800 meters. The main line is designed as a three-lane road, which expands to a four-lane road (including a 200-meter acceleration lane) at the merge area. The specific layout is as follows: Figure 1 shown.

[0095] The simulation system interacts with the control model via the Python TraCI interface. Road detectors are deployed at key locations to collect traffic status parameters and reward value data in real time. The calculated speed limit is transmitted in real time to the intelligent connected vehicles, enabling dynamic and variable speed control.

[0096] Based on the SUMO simulation platform, this study established a complete mixed traffic flow simulation system. The simulation environment was initialized by inputting basic traffic parameters. To monitor traffic flow in real time, multiple road detectors were deployed at key road nodes to collect dynamic mixed traffic flow data.

[0097] (4) Selection of evaluation indicators

[0098] This study evaluated model performance using four core metrics: total trip time, merging area travel time, total average speed, and average merging area speed. Through a network of road detectors deployed on the SUMO simulation platform, the system collected real-time multi-dimensional data, including lane-by-lane traffic flow, regional speed distribution, and time-segment differentiated operation. This provided a comprehensive basis for evaluating the model's effectiveness in coordinating mixed traffic flows.

[0099] (5) Baseline model selection

[0100] In the comparison and ablation experiments, a baseline model is selected to compare the performance of the deep reinforcement learning model on a given dataset.

[0101] DQN (Deep Q-Network) model: A reinforcement learning algorithm that combines deep neural networks with Q-learning. Through two innovative mechanisms, the experience replay buffer and the target network, it effectively solves the instability problem of traditional reinforcement learning in high-dimensional state space and can learn optimal control strategies directly from raw input data.

[0102] The Deep Deterministic Policy Gradient (DDPG) model is a deep reinforcement learning algorithm based on the actor-critic framework, specifically designed to solve control problems in continuous action spaces. DDPG combines the core concepts of Deterministic Policy Gradient (DPG) and Deep Q-Network (DQN), improving training stability by employing experience replay and target network mechanisms.

[0103] Method-r1 variant: For the degree reinforcement learning model, only the speed difference constraint of vehicles upstream and downstream of the main line in the merging area is considered to perform mixed traffic flow control.

[0104] Method-R2 variant: For the deep reinforcement learning model, only the flow difference constraint of vehicles upstream and downstream of the main line in the merging area is considered to perform mixed traffic flow control.

[0105] (6) Experimental results and analysis

[0106] a) Comparison between overall speed limit control and lane-level speed limit control

[0107] Experimental results comparing the uncontrolled method, the overall speed limit control method, and the lane-level speed limit control method are shown in Table 1. The present invention verifies the effectiveness of different control methods on datasets A and B. The results highlighted in bold represent the optimal experimental results, while the underlined results represent the suboptimal experimental results. As can be seen, compared to other baseline models, the deep reinforcement learning model achieves the best control effect on mixed traffic flow across four evaluation metrics: total travel time, average travel time in the merging area, overall average vehicle speed, and average vehicle speed in the merging area. This is attributed to the interaction between outer lane vehicle speed and traffic dynamics captured in the simulation. When ramp vehicles enter the merging area, they must change lanes to merge with the mainline traffic flow, resulting in frequent interactions with vehicles in the outer lanes. The significant speed difference between ramp and mainline vehicles further amplifies this interaction, disrupting normal traffic flow and reducing speeds. The resulting congestion significantly reduces the efficiency of traffic in the entire merging area.

[0108] Table 1 Comparison of travel time and average speed indicators under different speed limit methods

[0109]

[0110] b) Comparison of speed limit control effects under different penetration rates

[0111] Table 2 shows the optimization effects of the deep reinforcement learning model under different penetration conditions. As can be seen from the table, the model's optimization effects on various metrics improve with increasing penetration. Furthermore, increased penetration significantly reduces the frequency of low-speed operating conditions and increases the duration of high-flow operating states. This is attributed to the fact that as the proportion of CAVs increases, more vehicles within the control area are coordinated and controlled by the autonomous driving system, further stabilizing the overall traffic flow. The average speed of HVs gradually approaches that of CAVs due to the guidance provided by CAVs, further demonstrating the robustness of the model across different simulation environments.

[0112] Table 2 Comparison of travel time and average speed indicators under different permeabilities

[0113]

[0114] c) Baseline model comparison

[0115] Table 3 shows the performance improvement of the deep reinforcement learning model. As can be seen from the table, the deep reinforcement learning model outperforms the other four models in terms of optimization, with better results for total travel time, average travel time in the merging area, overall average vehicle speed, and average vehicle speed in the merging area. This demonstrates that the dual flow and speed constraints considered in this invention help improve the performance of the speed limit control model.

[0116] Table 3 Comparison of travel time and average speed indicators of each model

[0117]

[0118] Based on the above examples, it can be determined that the method has the following beneficial effects:

[0119] (1) To address the congestion and safety issues in merging areas on urban highways, a lane-level mixed traffic flow cooperative control method based on deep reinforcement learning is proposed. The proposed deep reinforcement learning model models speed limit control as a Markov decision framework, comprehensively considering key traffic dynamic characteristics such as the flow difference between the upstream and downstream main lines of the merging area and the speed difference between adjacent lanes, and designs a reward function to guide the model to achieve optimal control.

[0120] (2) This research effectively serves the field of intelligent transportation and effectively alleviates traffic congestion. At the same time, the present invention has self-learning capabilities and can adapt to traffic conditions at different times and with different flow rates, significantly reducing the cost of manual parameter adjustment. It not only provides technical support for the coordinated control of mixed traffic flows, but also lays a theoretical foundation for the forward-looking layout of traffic management after the popularization of autonomous driving.

[0121] The lane-level mixed traffic flow collaborative control system based on deep reinforcement learning provided by the present invention is described below. The lane-level mixed traffic flow collaborative control system based on deep reinforcement learning described below and the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning described above can be referenced to each other.

[0122] Figure 3 It is a structural diagram of a lane-level mixed traffic cooperative control system based on deep reinforcement learning provided by an embodiment of the present invention.

[0123] like Figure 3 As shown, the lane-level mixed traffic cooperative control system based on deep reinforcement learning provided by the embodiment of the present invention includes:

[0124] The data acquisition and state construction module 301 is used to perceive the vehicle flow, vehicle density, and vehicle speed data upstream and downstream of the main line of the highway merging area in real time through roadside detectors; construct a multidimensional state vector based on the perceived data; and model the state transition probability between the upper limits of vehicle speeds in adjacent time periods based on the Markov decision process, providing an environmental interaction mechanism for the deep reinforcement learning model.

[0125] The model construction module 302 is used to construct a lane-level mixed traffic flow collaborative control model based on a deep reinforcement learning model. The lane-level mixed traffic flow control model aims to achieve optimal traffic efficiency in the merging area of ​​the highway and completes the model training and testing. The lane-level mixed traffic flow control model uses a dual-network architecture to perform more precise control under different mixed traffic flow penetration conditions, thereby improving the overall efficiency of the merging area.

[0126] Stable training is achieved through a dual-network interaction system consisting of an online network and a target network. The online network adopts the Dueling architecture to achieve decoupled evaluation of state value, decomposing the Q value into a state-value function V(s) and an advantage function A(s,a), respectively learning the macro-state characteristics and micro-action benefits of mixed traffic flows. The aggregation layer then generates Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))). The target network synchronously suppresses Q-value overestimation through periodic parameters. Combined with the Dueling structure's ability to generalize sparse rewards, the system achieves synergistic improvements in global traffic efficiency optimization and local action decision-making accuracy in dynamic traffic scenarios in merging areas. Ultimately, through delayed updates and a decoupled evaluation mechanism, the system ensures stable convergence of the policy gradient in complex road conditions with frequent interactions and uneven rewards.

[0127] The example verification module 303 is used to input the highway merging area dataset into the lane-level mixed traffic collaborative control model based on the deep reinforcement learning model to realize example verification of the mixed traffic collaborative control.

[0128] The specific implementation method of the lane-level mixed traffic cooperative control system provided in this embodiment can be implemented with reference to the above embodiment and will not be repeated here.

[0129] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lane-level mixed traffic flow cooperative control method based on deep reinforcement learning, characterized by: include: The roadside detectors sense the vehicle flow, density, and speed data upstream and downstream of the main line of the highway merging area, and construct a multi-dimensional state vector , which not only reflects the heterogeneity of traffic space, but also ensures the temporal continuity of speed adjustment. is the vehicle density upstream of the main line of the highway merging area, is the density of the upstream confluence area, is the vehicle density of the highway merging area ramp, is the vehicle density downstream of the main line of the expressway merging area, is the historical vehicle speed of the previous control cycle; Defining the action space is the vehicle speed upper limit set, for The upper limit of a vehicle speed in The value range is between The difference in vehicle speed limits between the same lane and adjacent lanes during adjacent time periods shall not exceed 20 km / h; Select Markov decision process to model the state transition probability between vehicle speed limits in adjacent time periods , providing an environment interaction mechanism for deep reinforcement learning models; A deep reinforcement learning reward function is defined, with the minimum total travel time as the objective function and the minimum difference in traffic volume and speed between the upstream and downstream main lines of the merging area as the reward conditions, to collaboratively control the operation of traffic flow. Construct a deep reinforcement learning model to model the Markov decision process. Based on the traffic volume, speed, and relationship between lanes, a different vehicle speed limit is dynamically generated for each lane. The deep reinforcement learning model is used to perform feedback optimization at every moment, continuously adapting to traffic changes and dynamically updating the action space. , accurately responding to the operational needs of each lane; Using the SUMO platform to build a simulation environment, we varied the penetration rate of connected vehicles in mixed traffic flows and used two datasets to evaluate the model's performance and conduct case studies under different penetration rates. Aiming at the dynamic characteristics of mixed traffic flow in highway merging areas, a lane-level double-layer network control strategy is constructed; The lane-level double-layer network control strategy includes: Speed ​​smoothing is used for the inner high-traffic lane to stabilize traffic flow, while speed coordination is implemented for the outer lane with frequent lane changes to optimize lane change intervals. By establishing a coordinated speed optimization model between lanes to minimize speed differences and setting up a dynamic response mechanism, the speed limit plan can be adjusted in real time when sudden changes in traffic conditions are detected, thereby improving the overall traffic efficiency of the merging area.

2. The lane-level mixed traffic flow cooperative control method based on deep reinforcement learning according to claim 1 is characterized in that: The Markov decision process is selected to model the state transition probability between the vehicle speed upper limits in adjacent time periods ,include: Use the SUMO platform to simulate the traffic behavior of mixed traffic flows and model the state transition probability between vehicle speed limits in adjacent time periods. , where the state transition process is state and actions To the next state process of change.

3. The lane-level mixed traffic flow cooperative control method based on deep reinforcement learning according to claim 1 is characterized in that: It also includes defining deep reinforcement learning reward functions, including: Minimize the total travel time TTS of vehicles in the merging area within the time range K, where TTS is expressed by formula (1): (1) in, represents the simulation time step; K represents the overall time interval, k is a certain time interval, t represents a specific moment, represents the initial number of vehicles in the road section at the beginning of the simulation, represents the number of vehicles entering the road section at time t, represents the number of vehicles leaving the road section at time t; is the lower limit of the time corresponding to the time interval k; is the upper limit of the time corresponding to the time interval k; Introduce the traffic flow difference between the upstream and downstream of the main line of the merging area at time t as a sub-item of the reward function , Expressed by formula (2): ;(2) Introducing the speed difference between the upstream and downstream vehicles in the merging area at time t as another component of the reward function , Expressed by formula (3): ;(3) in, represents the speed of vehicle i in the merging area at time t, represents the number of vehicles in the merging area at time t, H is the number of time steps in the control cycle, is the vehicle speed under free flow conditions; The total reward function is obtained by taking the minimum difference in traffic volume and speed between the upstream and downstream of the main line in the merging area as the reward condition. , collaboratively control traffic flow operation status, among which, It is defined as the following formula: ;(4) in, 、 They represent the difference in traffic volume and speed between the upstream and downstream of the main line in the merging area in the total reward function. The corresponding weights in .

4. The lane-level mixed traffic flow cooperative control method based on deep reinforcement learning according to claim 1 is characterized in that: The construction of the deep reinforcement learning model includes: The deep reinforcement learning model includes a first neural network model and a second neural network model. The first neural network model is used to evaluate the action value under the state, and the second neural network model is used to select the action. The deep reinforcement learning model models the Markov decision process and dynamically generates different vehicle speed limits for each lane based on the traffic volume, speed and relationship between the current lanes. The deep reinforcement learning model is used to perform feedback optimization at every moment, continuously adapt to traffic changes, and dynamically update the action space. , accurately responding to the operational needs of each lane; During the training process of the deep reinforcement learning model, the interaction and feedback between the online network and the target network are utilized to improve the stability and robustness of the model in complex mixed traffic environments, and stable training is achieved through the dual-network interaction system of the online network and the target network; The online network adopts the Dueling architecture to achieve decoupled evaluation of state value, which conforms to the following formula Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))), where Q(s,a) is the action driving function representing the comprehensive expected benefit obtained by performing action a in state s, V(s) is the state value function, A(s,a) is the advantage function, and meanA(s,a) is the mean of the advantages of all optional actions in the current state. V(s) and A(s,a) respectively learn the macro-state characteristics and micro-action benefits of the mixed traffic flow. The macro-state characteristics include lane density and speed distribution, and the micro-action benefits include lane changing and acceleration and deceleration benefits. The target network suppresses Q value overestimation through periodic parameter synchronization. Combining the Dueling structure's ability to generalize sparse rewards, it achieves a coordinated improvement in global traffic efficiency optimization and local action decision-making accuracy in lane merging areas in dynamic traffic scenarios. Ultimately, through delayed updates and a decoupled evaluation mechanism, it ensures stable policy gradient convergence in complex road conditions with frequent interactions and unbalanced rewards.

5. The lane-level mixed traffic flow cooperative control method based on deep reinforcement learning as claimed in claim 1, characterized in that: The simulation environment is built using the SUMO simulation model, and two data sets are used to evaluate the model's performance and conduct case verification under different permeability environments, including: The SUMO simulation platform was used for simulation verification. During the simulation process, a variety of mixed traffic flow scenarios were simulated, including traffic flows composed entirely of manually driven vehicles, traffic flows composed entirely of autonomous vehicles, and traffic flows composed of a mixture of manually driven vehicles and autonomous vehicles. Modeling different types of traffic flows. The simulation platform provides detailed data on lane flow, vehicle speed, and traffic density, generating simulation results that are used to evaluate the adaptability and effectiveness of the deep reinforcement learning model in different traffic environments. By comparing with traditional control methods, the performance of the deep reinforcement learning model under mixed traffic conditions was evaluated, and the advantages of the deep reinforcement learning model in improving the traffic capacity of merging areas, reducing traffic congestion and optimizing vehicle speed control were verified.

6. A lane-level mixed traffic cooperative control system based on deep reinforcement learning, applied to the lane-level mixed traffic cooperative control method based on deep reinforcement learning according to any one of claims 1-5, characterized in that: include: The data collection and status construction module is used to sense the vehicle flow, vehicle density, and vehicle speed data upstream and downstream of the main line of the highway merging area in real time through roadside detectors; Construct a multi-dimensional state vector based on the sensed data; The Markov decision process is used to model the state transition probability between vehicle speed limits in adjacent time periods, providing an environmental interaction mechanism for deep reinforcement learning models. A model building module is used to build a lane-level mixed traffic collaborative control model based on a deep reinforcement learning model. This lane-level mixed traffic control model aims to achieve optimal traffic efficiency in highway merging areas and completes model training and testing. Through a dual-network architecture, this lane-level mixed traffic control model can perform more precise control under different mixed traffic penetration conditions, improving the overall efficiency of the merging area. Stable training is achieved through a dual-network interaction system between an online network and a target network. The online network uses a Dueling architecture to achieve decoupled evaluation of state value, decomposing the Q value into a state-value function V(s) and an advantage function A(s,a), respectively learning the macro-state characteristics and micro-action benefits of mixed traffic flows. The aggregation layer then generates Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))). The target network synchronously suppresses Q-value overestimation through periodic parameters. Combined with the Dueling structure's ability to generalize sparse rewards, it achieves a synergistic improvement in global traffic efficiency optimization and local action decision-making accuracy in dynamic traffic scenarios. Ultimately, through delayed updates and a decoupled evaluation mechanism, it ensures stable convergence of the vehicle's policy gradient in complex road conditions with frequent interactions and uneven rewards. The example verification module is used to input the highway merging area dataset into the lane-level mixed traffic collaborative control model based on the deep reinforcement learning model to realize the example verification of the mixed traffic collaborative control.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the lane-level mixed traffic collaborative control method based on deep reinforcement learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Variable speed-limiting control method for divided lanes based on intelligent network connection special lane environment

    CN116229706A