Lane-level mixed traffic flow cooperative control method and system based on deep reinforcement learning

Through deep reinforcement learning methods, dynamically adjusting the vehicle speed upper limit and speed limit strategies, the shortcomings of traditional traffic control methods in hybrid traffic environments are solved, lane-level coordinated control is achieved, and the traffic efficiency and safety of the combined flow area are improved.

CN120412291AActive Publication Date: 2025-08-01BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510912699.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional traffic control methods are difficult to accurately deal with the interaction effects of autonomous vehicles (CAVs) and artificially driven vehicles (HVs) in a hybrid traffic environment, resulting in traffic congestion in the confluence area. Existing research has failed to fully consider the traffic flow differences between lanes and their impact on overall efficiency.

Method used

The lane-level hybrid flow collaborative control method based on deep reinforcement learning is adopted to perceive vehicle data through roadside detectors, build multi-dimensional state vectors, define action space and reward functions, use Markov decision-making process modeling, dynamically adjust vehicle speed upper limit, optimize speed limit strategies to reduce traffic and speed differences between lanes, and build a double-layer network control strategy to cope with traffic changes.

Benefits of technology

It realizes efficient coordinated control in a mixed traffic environment, improves the traffic efficiency and safety of the combined flow area, adapts to different traffic environments, reduces congestion, and improves the stability and robustness of the overall traffic flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412291A_ABST
    Figure CN120412291A_ABST
Patent Text Reader

Abstract

The invention relates to the field of traffic transportation information engineering, and discloses a lane-level mixed traffic flow cooperative control method and system based on deep reinforcement learning, and the method considers the flow difference and speed difference between lanes, is embedded into a reward and punishment function of deep reinforcement learning, and achieves the cooperative control of the mixed traffic flow of different lanes according to the distribution condition of the mixed traffic flow of different lanes. A lane-level double-layer network control strategy is constructed, different vehicle speed upper limits are dynamically generated for each lane according to the vehicle flow and the vehicle speed of the current lane and the correlation among the lanes, feedback optimization is carried out at each moment through the deep reinforcement learning model, the traffic change condition is continuously adapted, and the speed of the vehicle is improved. Therefore, the operation demand of each lane is accurately responded, and empirical research is carried out. The deep reinforcement learning model can obtain a lane-level mixed traffic flow cooperative control effect superior to that of a baseline model, can effectively relieve the current congestion situation of a highway confluence area, and has cross-scene applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of transportation information engineering, and discloses a lane-level mixed traffic flow collaborative control method and system based on deep reinforcement learning. Background Art

[0002] The highway merge area is a key merge area for traffic flow management, and its dynamic and complex traffic characteristics often lead to serious traffic congestion problems. In the mixed traffic environment where connected and automated vehicles (CAVs) and human-driven vehicles (HVs) coexist, due to the differences in driving behaviors and response mechanisms between the two types of vehicles, traditional traffic control methods are difficult to achieve efficient collaboration, making the traffic optimization in the merge area face greater challenges.

[0003] The main traffic conflict in the merge area stems from the merge conflict between the ramp traffic flow and the main-road traffic flow, and the variable speed limit (VSL) technology is one of the effective means to alleviate this problem. Traditional VSL control methods mainly rely on feedback control or model predictive control (MPC). The former dynamically adjusts the speed limit by setting density thresholds, and the latter suppresses congestion based on multi-factor optimization. However, the adaptability of these methods in the mixed traffic environment is insufficient, and it is difficult to accurately cope with the interactive effects of CAVs and HVs, resulting in limited control effects.

[0004] In recent years, deep reinforcement learning (DRL) has shown significant advantages in the field of VSL control due to its strong environmental adaptability and autonomous learning ability. Compared with traditional methods, DRL can autonomously optimize decision-making strategies based on real-time traffic states, so as to achieve more accurate control in dynamically changing traffic scenarios. Existing studies have shown that the VSL strategy based on DRL can effectively improve the traffic efficiency in the merge area, but existing studies mostly focus on single-lane scenarios and fail to fully consider the traffic flow differences between lanes and their impact on the overall efficiency. Especially in the mixed traffic environment, the uneven speed distribution between the inner and outer lanes, the response differences between CAVs and HVs, and the change of CAV penetration rate all pose higher requirements for the collaborative control method of mixed traffic flow. Summary of the Invention

[0005] To break through the limitations of the existing technology that ignores the flow difference or speed difference between lanes and issues inaccurate vehicle speed limits, the present invention proposes a lane-level mixed traffic flow collaborative control method and system based on deep reinforcement learning. By constraining the flow difference and speed difference between lanes, embedding the reward and punishment function of deep reinforcement learning, and constructing a variable speed limit strategy according to the traffic distribution of different lanes, the collaborative control of lane-level mixed traffic flow in the highway merge area is realized.

[0006] According to the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the present invention, it is characterized by including: The roadside detectors sense the vehicle flow, density, and speed data upstream and downstream of the main line of the highway merging area, and construct a multi-dimensional state vector , which not only reflects the heterogeneity of traffic space, but also ensures the temporal continuity of speed adjustment. is the vehicle density upstream of the main line of the highway merging area, is the density of the upstream confluence area, is the vehicle density of the highway merging area ramp, is the vehicle density downstream of the main line of the expressway merging area, is the historical vehicle speed of the previous control cycle; Defining the action space is the vehicle speed upper limit set, is the upper limit of the speed of a vehicle in A, and the value range of a is between The difference in vehicle speed limits between the same lane and adjacent lanes during adjacent time periods shall not exceed 20 km / h; Select Markov decision process to model the state transition probability between vehicle speed limits in adjacent time periods , providing an environment interaction mechanism for deep reinforcement learning models; A deep reinforcement learning reward function is defined, with the minimum total travel time as the objective function and the minimum difference in traffic volume and speed between the upstream and downstream main lines of the merging area as the reward conditions, to collaboratively control the operation of traffic flow. Construct a deep reinforcement learning model to model the Markov decision process. Based on the traffic volume, speed, and relationship between lanes, a different vehicle speed limit is dynamically generated for each lane. The deep reinforcement learning model is used to perform feedback optimization at every moment, continuously adapting to traffic changes and dynamically updating the action space. , accurately responding to the operational needs of each lane; Using the SUMO platform to build a simulation environment, the penetration rate of intelligent connected vehicles in mixed traffic was changed, and the performance of the model was evaluated and verified on two datasets. According to the lane-level mixed traffic cooperative control method based on deep reinforcement learning provided by the present invention, the state transition probability between the vehicle speed upper limits in adjacent time periods is modeled by selecting the Markov decision process. ,include: Use the SUMO platform to simulate the traffic behavior of mixed traffic flows and model the state transition probability between vehicle speed limits in adjacent time periods. , where the state transition process is state and actions To the next state process of change.

[0007] The lane - level mixed - traffic flow collaborative control method based on deep reinforcement learning provided by the present invention further includes defining a deep reinforcement learning reward function, including: Minimizing the total travel time TTS of vehicles in the merging area within the time range K, where TTS is expressed by formula (1): ; (1) Wherein, represents the simulation time step; K represents the overall time interval, k is a certain time interval, t represents a specific moment, represents the initial number of vehicles in the road section at the start of the simulation, represents the number of vehicles entering the road section at time t, represents the number of vehicles leaving the road section at time t; is the lower limit of the moment corresponding to the time interval k; is the upper limit of the moment corresponding to the time interval k; Introducing the traffic - flow difference between the upstream and downstream of the main line in the merging area at time t as a sub - item of the reward function , which is expressed by formula (2): ; (2) Introducing the vehicle - speed difference between the upstream and downstream of the main line in the merging area at time t as another sub - item of the reward function , which is expressed by formula (3): ; (3) Wherein, represents the speed of vehicle i in the merging area at time t, represents the number of vehicles in the merging area at time t, H is the number of time steps within the control period, is the vehicle speed under free - flow conditions; Taking the minimum traffic - flow difference and the minimum vehicle - speed difference between the upstream and downstream of the main line in the merging area as the reward conditions to obtain the total reward function , and collaboratively controlling the operation status of the traffic flow, wherein, is defined as the following formula: ; (4) Wherein, , respectively represent the corresponding weights of the traffic - flow difference and the vehicle - speed difference between the upstream and downstream of the main line in the merging area in the total reward function [[ID=6l]] .

[0008] The lane - level mixed - traffic flow collaborative control method based on deep reinforcement learning provided by the present invention, the constructing of the deep reinforcement learning model includes: The deep reinforcement learning model includes a first neural network model and a second neural network model. The first neural network model is used to evaluate the action value in a state, and the second neural network model is used to select actions, enabling the deep reinforcement learning model to model the Markov decision process. Based on the traffic flow, vehicle speed, and the mutual relationship between lanes in the current lane, different vehicle speed limits are dynamically generated for each lane, and feedback optimization is performed at each moment through the deep reinforcement learning model to continuously adapt to traffic changes and dynamically update the action space. , and accurately respond to the operation requirements of each lane; During the training process of the deep reinforcement learning model, the interaction and feedback between the online network and the target network are utilized to improve the stability and robustness of the model in a complex traffic flow environment, and stable training is achieved through the dual-network interaction system of the online network and the target network; Among them, the online network adopts the Dueling architecture to achieve decoupled evaluation of state values, which conforms to the formula Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))), where Q(s,a) is the action driving function representing the comprehensive expected return obtained by executing action a in state s, V(s) is the state value function, A(s,a) is the advantage function, meanA(s,a) is the mean of the advantages of all optional actions in the current state, V(s) and A(s,a) respectively learn the macroscopic state characteristics and microscopic action benefits of the traffic flow. The macroscopic state characteristics include lane density and vehicle speed distribution, and the microscopic action benefits include lane-changing and acceleration / deceleration benefits; the target network suppresses the overestimation of Q values through periodic parameter synchronization; Combining the generalization ability of the Dueling structure for sparse rewards, the global traffic efficiency optimization in the lane confluence area and the collaborative improvement of local action decision accuracy are realized in a dynamic traffic scenario. Finally, through the delayed update and decoupled evaluation mechanism, it is ensured that the policy gradient of the vehicle converges stably in a complex road condition with frequent interactions and unbalanced rewards.

[0009] According to the lane-level mixed traffic flow collaborative control method provided by the present invention, it further includes: constructing a lane-level double-layer network control strategy for the dynamic change characteristics of the mixed traffic flow in the highway confluence area; The lane-level double-layer network control strategy includes: For the inner high-flow lanes, speed smoothing control is adopted to stabilize the traffic flow, and for the outer lanes with frequent lane-changes, speed coordination control is implemented to optimize the lane-changing gap; By establishing a speed collaborative optimization model between lanes, the speed difference is minimized, and a dynamic response mechanism is set up to adjust the speed limit scheme in real time when a sudden traffic state change is detected, thereby improving the overall traffic efficiency of the confluence area.

[0010] According to the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the present invention, a simulation environment is built using the SUMO platform, the penetration ratio of connected and autonomous vehicles in the mixed traffic flow is changed, and two datasets are used to evaluate the performance of the model and verify the instances under different penetration rate environments, including: Build a simulation environment using the SUMO platform. During the simulation process, simulate a variety of different mixed traffic flow scenarios, where the mixed traffic flow scenarios include traffic flows composed entirely of manually driven vehicles, traffic flows composed entirely of autonomous vehicles, and situations where traffic flows composed of manually driven vehicles and traffic flows composed of autonomous vehicles are mixed; Model different types of traffic flows. The simulation platform provides detailed data on lane flow, vehicle speed, and traffic density, and obtains simulation results, which are used to evaluate the adaptability and effectiveness of the deep reinforcement learning model in different traffic environments; By comparing with traditional control methods, evaluate the performance of the deep reinforcement learning model under mixed traffic flow conditions, and verify the advantages of the deep reinforcement learning model in improving the passing capacity of the merging area, reducing traffic congestion, and optimizing vehicle speed control.

[0011] The present invention also provides a lane-level mixed traffic flow collaborative control system based on deep reinforcement learning, including: A data acquisition and state construction module, which is used to sense the vehicle flow, vehicle density, and vehicle speed data of the upstream and downstream of the main line in the highway merging area in real time through roadside detectors; construct a multi-dimensional state vector according to the sensed data; obtain lane-level flow and speed information in real time, and model the state transition probability between the vehicle speed upper limits in adjacent time periods based on the Markov decision process, so as to provide an environmental interaction mechanism for the deep reinforcement learning model.

[0012] A model construction module, which is used to build a lane-level mixed traffic flow collaborative control model based on the deep reinforcement learning model. The lane-level mixed traffic flow control model aims at the optimal passing efficiency of the highway merging area, and completes the training and testing of the model; the lane-level mixed traffic flow control model can perform more refined control under different mixed traffic flow penetration conditions through a dual-network architecture, improving the overall efficiency of the merging area; Stable training is achieved through a dual-network interaction system between an online network and a target network. The online network adopts a Dueling architecture to decouple and evaluate state values, decomposing the Q value into a state value function V(s) and an advantage function A(s,a), learning the macroscopic state characteristics of traffic flow and the microscopic action rewards respectively, and generating Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))) through an aggregation layer; the target network suppresses overestimation of the Q value through periodic parameter synchronization; meanwhile, combined with the generalization ability of the Dueling structure for sparse rewards, it realizes the coordinated improvement of the global traffic efficiency optimization and local action decision-making accuracy in the lane merging area under dynamic traffic scenarios. Finally, through a delayed update and decoupled evaluation mechanism, it ensures the stable convergence of the policy gradient of vehicles in complex road conditions with frequent interactions and unbalanced rewards.

[0013] An instance verification module is used to input the highway merging area data set into the lane-level mixed traffic flow collaborative control model based on a deep reinforcement learning model to realize the instance verification of the mixed traffic flow collaborative control.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements any one of the above-mentioned lane-level mixed traffic flow collaborative control methods based on deep reinforcement learning.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above-mentioned lane-level mixed traffic flow collaborative control methods based on deep reinforcement learning.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above-mentioned lane-level mixed traffic flow collaborative control methods based on deep reinforcement learning.

[0017] The present invention proposes a lane-level mixed traffic flow collaborative control method based on deep reinforcement learning. This method innovatively introduces a traffic flow balance reward mechanism, with the goal of minimizing the difference in the number of vehicles entering and leaving the merging area, and dynamically adjusts the speed limit value in combination with factors such as real-time speed, flow difference, and CAV penetration rate. Through lane-level microscopic control, the method can accurately identify the traffic characteristics of different lanes (such as high flow in the inner lane and frequent lane-changing behavior in the outer lane), and implement differential speed limit control accordingly, thereby effectively balancing the speed distribution between lanes and improving the overall traffic efficiency and safety of the merging area. Simulation experiments verify the excellent performance of the method in a mixed traffic flow environment, providing a new idea for the collaborative control of mixed traffic flow in highway merging areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is one of the schematic diagrams of the application process of the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the embodiment of the present invention; Figure 2 It is the second of the schematic diagrams of the application process of the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the embodiment of the present invention; Figure 3 It is the schematic diagram of the structure of the lane-level mixed traffic flow collaborative control system based on deep reinforcement learning provided by the embodiment of the present invention; Figure 4 It is the schematic diagram of the physical structure of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0021] Figure 1 It is one of the schematic diagrams of the application process of the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the embodiment of the present invention.

[0022] Figure 2 It is the second of the schematic diagrams of the application process of the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning provided by the embodiment of the present invention.

[0023] Such as Figure 1 and Figure 2 As shown, this embodiment provides a lane-level mixed traffic flow collaborative control method based on deep reinforcement learning, including: Step 101, sense the vehicle flow, vehicle density, and vehicle speed data of the upstream and downstream of the main line in the highway merging area through roadside detectors, and construct a multi-dimensional state vector , which not only reflects the heterogeneity of the traffic space but also ensures the time continuity of speed adjustment. Among them, is the vehicle density of the upstream of the main line in the highway merging area, is the density of the upstream merging area, is the vehicle density on the ramp in the highway merging area, is the vehicle density downstream of the main line in the highway merging area, is the historical vehicle speed of the previous control period; the defined multi-dimensional state vector can not only reflect the heterogeneity of the traffic space, but also ensure the time continuity of speed adjustment; Step 102, define the action space is the set of upper limits of vehicle speeds, is the upper limit value of a certain vehicle speed in A, and the value range of a is between ; the difference in the upper limits of vehicle speeds in the same lane and adjacent lanes in adjacent time periods does not exceed 20 km / h; Select the state transition probability between the upper limits of vehicle speeds in adjacent time periods by using the Markov decision process modeling , providing an environmental interaction mechanism for the deep reinforcement learning model; Step 103, define the deep reinforcement learning reward function, with the minimum total travel time as the objective function and the minimum differences in traffic flow and vehicle speed between the upstream and downstream of the main line in the merging area as the reward conditions to coordinately control the traffic flow operation status; Construct a deep reinforcement learning model to model the Markov decision process, and dynamically generate different upper limits of vehicle speeds for each lane based on the traffic flow, vehicle speed of the current lane, and the mutual relationship between lanes, and perform feedback optimization at each moment through the deep reinforcement learning model to continuously adapt to traffic changes and dynamically update the action space , accurately responding to the operation requirements of each lane; The deep reinforcement learning model includes a first neural network model and a second neural network model. The first neural network model is used to evaluate the action value in the state, and the second neural network model is used to select actions, enabling the deep reinforcement learning model to dynamically generate different upper limits of vehicle speeds for each lane based on the traffic flow, vehicle speed of the current lane, and the mutual relationship between lanes, and perform feedback optimization at each moment through the deep reinforcement learning model to continuously adapt to traffic changes and accurately respond to the operation requirements of each lane; During the training process of the deep reinforcement learning model, the interaction and feedback between the online network and the target network are used to improve the stability and robustness of the model in a complex vehicle traffic environment, and stable training is achieved through the dual-network interaction system of the online network and the target network; Among them, the online network uses the Dueling architecture to achieve decoupled evaluation of state values, which conforms to the following formula Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))), where Q(s,a) is the action driving function representing the comprehensive expected reward obtained by executing action a in state s, V(s) is the state value function, A(s,a) is the advantage function, meanA(s,a) is the mean of the advantages of all optional actions in the current state, V(s) and A(s,a) respectively learn the macroscopic state characteristics and microscopic action rewards of the mixed traffic flow. The macroscopic state characteristics include lane density and vehicle speed distribution, and the microscopic action rewards include lane-changing and acceleration / deceleration benefits; the target network suppresses overestimation of Q values through periodic parameter synchronization; Combined with the generalization ability of the Dueling structure for sparse rewards, it realizes the coordinated improvement of the global traffic efficiency optimization in the lane merging area and the local action decision-making accuracy in the dynamic traffic scenario. Finally, through the delayed update and decoupled evaluation mechanism, it ensures the stable convergence of the policy gradient of vehicles in complex road conditions with frequent interactions and unbalanced rewards.

[0024] Step 104: Use the SUMO platform to build a simulation environment, change the penetration ratio of intelligent connected vehicles in the mixed traffic flow, and use two datasets to evaluate the performance of the model and conduct example verification in different penetration rate environments.

[0025] The simulation verification is carried out using the SUMO (Simulation of Urban Mobility) simulation platform. During the simulation, various different mixed traffic flow scenarios are simulated, including traffic flows composed entirely of human-driven vehicles (HVs), traffic flows composed entirely of autonomous driving vehicles (CAVs), and the situation of a mixture of traffic flows composed of human-driven vehicles and traffic flows composed of autonomous driving vehicles. By modeling these different types of traffic flows, the simulation platform can provide detailed data on aspects such as lane flow, vehicle speed, and traffic density, and obtain simulation results. The simulation results are used to evaluate the adaptability and effectiveness of the deep reinforcement learning model in different traffic environments. By comparing with traditional control methods, the performance of the deep reinforcement learning model in the mixed traffic flow environment is evaluated, and its advantages in improving the traffic capacity of the merging area, reducing traffic congestion, and optimizing vehicle speed control are verified.

[0026] The lane-level cooperative control method for mixed traffic flow based on deep reinforcement learning provided in this embodiment innovatively introduces a traffic flow balance reward mechanism, with the goal of minimizing the difference in the number of vehicles upstream and downstream on the main line of the merging area. At the same time, it dynamically adjusts the speed limit value by combining factors such as real-time speed, flow difference, and CAV penetration rate. Through lane-level microscopic control, the strategy can accurately identify the traffic characteristics of different lanes (such as high flow in the inner lane and frequent lane-changing behavior in the outer lane), and implement differential speed limits accordingly, thereby effectively balancing the speed distribution between lanes and improving the overall traffic efficiency and safety of the merging area. Simulation experiments verify the superior performance of this strategy in the mixed traffic flow environment, providing a new idea for the cooperative control of mixed traffic flow in highway merging areas.

[0027] In an exemplary embodiment, it further includes defining a deep reinforcement learning reward function, including: Minimize the total travel time TTS of vehicles in the merging area within the time range K, where TTS is expressed by formula (1): ; (1) Where, represents the simulation time step; K represents the overall time interval, k is a certain time interval, t represents a specific moment, represents the initial number of vehicles in the section at the start of the simulation, represents the number of vehicles entering the section at time t, represents the number of vehicles leaving the section at time t; is the lower limit of the moment corresponding to the time interval k; is the upper limit of the moment corresponding to the time interval k.

[0028] Introduce the traffic flow difference between the upstream and downstream of the main line in the merging area at time t as a sub-item of the reward function , is expressed by formula (2): ; (2) Introduce the vehicle speed difference between the upstream and downstream of the main line in the merging area at time t as another sub-item of the reward function , is expressed by formula (3): ; (3) Where, represents the speed of vehicle i in the merging area at time t, represents the number of vehicles in the merging area at time t, H is the number of time steps within the control period, is the vehicle speed under free flow conditions; Take the minimum traffic flow difference and minimum vehicle speed difference between the upstream and downstream of the main line in the merging area as the reward conditions to obtain the total reward function , co - control the traffic flow operation status, where is defined as the following formula: ; (4) Among them, and respectively represent the corresponding weights of the traffic flow difference and the vehicle speed difference between the upstream and downstream of the main line in the confluence area in the total reward function .

[0029] The spatio - temporal distribution of vehicle arrivals has nothing to do with the control measures, which means that only the number of vehicles leaving the road section can be directly controlled. By increasing the number of vehicles leaving the road section per unit time, the TTS can be reduced, the speed limit control target can be achieved, and thus the road section can be optimized. The vehicles entering the road section within a specific time period can be detected by deploying detectors at key points of the section. Based on the above settings of the speed limit control target and the detection method, the traffic flow difference between the upstream and downstream of the main line in the confluence area at time t is introduced as a sub - item of the reward function ; The traffic conditions in the confluence area cannot be fully characterized by just the inflow and outflow. For example, when the downstream section of the confluence area is congested, the vehicle movement is slow, and the difference between the inflow and outflow may be small. However, in this case, the low driving speed of the vehicles leads to a reduction in the overall traffic efficiency. To solve this limitation, the vehicle speed difference between the upstream and downstream of the main line in the confluence area at time t is introduced as another sub - item of the reward function .

[0030] Taking the minimum traffic flow difference and the minimum vehicle speed difference between the upstream and downstream of the main line in the confluence area as the reward conditions, the total reward function is obtained to co - control the traffic flow operation status.

[0031] The following uses a specific embodiment to illustrate the group category division method provided by the solution of this application.

[0032] (1) Working environment The present invention uses the PyTorch framework to write the model code. All experimental codes are run on a Windows 10 workstation equipped with a "GenInter(R) Corei7 - 13700F @ 2.10GHz" CPU and an "NVIDIA GeForce RTX 4090 (128 GB RAM)" GPU.

[0033] (2) Introduction of public datasets Datasets A and B are respectively used as two kinds of inputs in the experiments.

[0034] Dataset A is a simulated dataset, mainly used to simulate the vehicle merging situation on the upstream section of the main road and the entrance ramp. The traffic flow parameters are set as follows: the main line input flow is [2400, 2700, 3000, 2700, 2400] vehicles per hour, and the ramp input flow is [400, 600, 800, 600, 400] vehicles per hour. The vehicle paths are randomly generated by the SUMO software, and their arrival times are randomly set according to the binomial distribution.

[0035] Dataset B is the open-source dataset exiD, which mainly records the driving trajectories of human-driven vehicles in 7 different highway merging areas. The present invention specifically selects the relevant data related to the highway merging area in this dataset for experimental analysis.

[0036] (3)Experimental parameter settings To simulate the real mixed traffic flow environment, this study uses the SUMO (Simulation of Urban Mobility) software platform, which has outstanding advantages in modeling the operating environment of connected and autonomous vehicles (CAVs). In the simulation modeling, the behaviors of human-driven vehicles (HVs) are realized through the Intelligent Driver Model (IDM) and the LC2013 lane-changing model, while the connected and autonomous vehicles are simulated using the Adaptive Cruise Control (ACC) and Cooperative Adaptive Cruise Control (CACC) models.

[0037] The specific configuration of the road network in SUMO is as follows: the length of the basic section is 1600 meters, among which a 400-meter variable speed limit control area is set, followed by a 200-meter acceleration area and a 600-meter dissipation area. The entrance ramp is located at 800 meters of the section. The main line is designed with three lanes and expands to four lanes (including a 200-meter acceleration lane) in the merging area. The specific layout is as Figure 1 shown.

[0038] The simulation system interacts with the control model through the TraCI interface of Python. Road detectors are arranged at key positions to collect traffic state parameters and reward value data in real time. The calculated speed limit values will be transmitted to the connected and autonomous vehicles in real time, so as to achieve dynamic variable speed limit control.

[0039] Based on the SUMO simulation platform, this study established a complete mixed traffic flow simulation system, and initialized the simulation environment by inputting basic traffic parameters. To monitor the traffic operation state in real time, multiple groups of road detectors are arranged at key section nodes to collect dynamic mixed traffic flow data.

[0040] (4)Selection of evaluation indicators Four core indicators are adopted in this study to evaluate the model performance: the whole journey time, the merging area journey time, the whole journey average speed, and the merging area average speed. Through the road detector network deployed on the SUMO simulation platform, the system can collect multi-dimensional data such as lane-by-lane traffic flow, regional speed distribution, and time-period differential operation in real time, providing detailed basis for comprehensively evaluating the collaborative control effect of the model on mixed traffic flow.

[0041] (5)Selection of baseline models In the comparison and ablation experiment sessions, a benchmark model is selected to compare its performance with the deep reinforcement learning model on a given data set.

[0042] DQN (Deep Q-Network) model: A reinforcement learning algorithm that combines a deep neural network with Q-learning. It effectively solves the instability problem of traditional reinforcement learning in high-dimensional state spaces through two innovative mechanisms, namely the experience replay buffer and the target network, and can directly learn the optimal control strategy from the original input data. DDPG (Deep Deterministic Policy Gradient) model: A deep reinforcement learning algorithm based on the Actor-Critic framework, specifically designed to solve the control problem of continuous action spaces. DDPG combines the core ideas of deterministic policy gradient (DPG) and deep Q-network (DQN), and improves the training stability by adopting the experience replay and target network mechanisms.

[0043] The method-r1 variant: For the deep reinforcement learning model, only consider the speed difference constraint between the upstream and downstream of the main line in the merging area to conduct mixed traffic flow control.

[0044] The method-r2 variant: For the deep reinforcement learning model, only consider the traffic flow difference constraint between the upstream and downstream of the main line in the merging area to conduct mixed traffic flow control.

[0045] (6)Experimental results and analysis a) Comparison between overall speed limit control and lane-level speed limit control Compare the experimental results of the no-control method, the overall speed limit control method, and the lane-level speed limit control method. As shown in Table 1, the present invention verifies the effects of different control methods on datasets A and B. The bold values in the table are the optimal experimental results, and the underlined values are the sub-optimal experimental results. It can be seen that, compared with other baseline models, the deep reinforcement learning model has the best control effect on mixed traffic flow in four evaluation indicators, namely the total travel time, the average travel time in the merging area, the overall average speed of vehicles, and the average speed of vehicles in the merging area. This is attributed to the interactive influence of the vehicle speed and flow dynamics captured in the simulation. When ramp vehicles enter the merging area, they must merge with the main-line traffic through lane-changing behavior, resulting in frequent interactions with the vehicles in the outer lane. The significant speed difference between ramp vehicles and main-line vehicles further amplifies this interactive effect, not only disturbing the normal traffic flow and reducing the vehicle speed, but ultimately the resulting congestion problem will significantly weaken the traffic operation efficiency of the entire merging area.

[0046] Table 1 Comparison of travel time and average speed indicators for different speed limit methods

[0047] b) Comparison of speed limit control effects under different lower permeability limits Table 2 reflects the optimization effect of the deep reinforcement learning model in different permeability environments. As can be seen from the table, as the permeability increases, the optimization effect of the model on different indicators also improves. At the same time, the increase in permeability can significantly reduce the frequency of low-speed working conditions and increase the duration of high-flow operation states. This is attributed to the fact that as the proportion of CAVs increases, more vehicles in the control area are coordinated and managed by the autonomous driving system, further stabilizing the overall traffic flow. The average speed of HVs will gradually approach the CAV speed level due to the guiding effect of CAVs, further proving the robustness of the model in different simulation environments.

[0048] Table 2 Comparison of travel time and average speed indicators under different permeabilities

[0049] c) Comparison with baseline models Table 3 reflects the performance improvement of the deep reinforcement learning model. As can be seen from the table, the optimization effect of the deep reinforcement learning model is better than that of the other 4 models, and the indicators of the total travel time, the average travel time in the merging area, the overall average speed of vehicles, and the average speed of vehicles in the merging area are better, fully demonstrating that the dual constraints of flow and speed considered in the present invention are helpful for improving the performance of the speed limit control model.

[0050] Table 3 Comparison of travel time and average speed indicators for each model

[0051] Based on the above examples, it can be determined that the method has the following beneficial effects: (1) Aiming at the congestion and safety problems in the merging area of urban expressways, a lane-level cooperative control method for mixed traffic flow based on deep reinforcement learning is proposed. The deep reinforcement learning model models speed limit control as a Markov decision framework, comprehensively considers key traffic dynamic characteristics such as the flow difference between the upstream and downstream of the main line in the merging area and the speed difference between adjacent lanes, and designs a reward function to guide the model to achieve optimal control.

[0052] (2) This research effectively serves the field of intelligent transportation, effectively alleviates traffic congestion problems. At the same time, the present invention has the ability of self-learning, can adapt to traffic conditions at different times and different flows, and greatly reduces the cost of manual parameter adjustment. It not only provides technical support for the cooperative management and control of mixed traffic flow, but also lays a theoretical foundation for the forward-looking layout of traffic management after the popularization of autonomous driving.

[0053] Next, the lane-level cooperative control system for mixed traffic flow based on deep reinforcement learning provided by the present invention will be described. The lane-level cooperative control system for mixed traffic flow based on deep reinforcement learning described below can be correspondingly referred to the lane-level cooperative control method for mixed traffic flow based on deep reinforcement learning described above.

[0054] Figure 3 It is a schematic structural diagram of the lane-level cooperative control system for mixed traffic flow based on deep reinforcement learning provided by an embodiment of the present invention.

[0055] As Figure 3 shown, the lane-level cooperative control system for mixed traffic flow based on deep reinforcement learning provided by an embodiment of the present invention includes: A data acquisition and state construction module 301, configured to sense the vehicle flow, vehicle density, and vehicle speed data of the upstream and downstream of the main line in the expressway merging area in real time through roadside detectors; construct a multi-dimensional state vector according to the sensed data; and model the state transition probability between the vehicle speed limits in adjacent time periods based on the Markov decision process, so as to provide an environment interaction mechanism for the deep reinforcement learning model.

[0056] A model construction module 302, configured to construct a lane-level cooperative control model for mixed traffic flow based on a deep reinforcement learning model. The lane-level mixed traffic flow control model aims at the optimal traffic efficiency in the expressway merging area, and completes the training and testing of the model; the lane-level mixed traffic flow control model can perform more refined control under different mixed traffic flow penetration conditions through a dual-network architecture, so as to improve the overall efficiency of the merging area.

[0057] Stable training is achieved through a dual-network interaction system between an online network and a target network. The online network uses a Dueling architecture to decouple and evaluate state values, decomposing the Q-value into a state value function V(s) and an advantage function A(s,a), learning the macroscopic state characteristics of mixed traffic flows and the microscopic action rewards respectively, and generating Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))) through an aggregation layer; the target network suppresses overestimation of the Q-value through periodic parameter synchronization; at the same time, combining the generalization ability of the Dueling structure for sparse rewards, it realizes the coordinated improvement of the global traffic efficiency and local action decision-making accuracy in the lane merging area under dynamic traffic scenarios. Finally, through the delayed update and decoupled evaluation mechanism, it ensures the stable convergence of the policy gradient of vehicles in complex road conditions with frequent interactions and unbalanced rewards.

[0058] The instance verification module 303 is used to input the highway merging area data set into the lane-level mixed traffic flow collaborative control model based on the deep reinforcement learning model to realize the instance verification of the mixed traffic flow collaborative control.

[0059] The specific implementation method of the lane-level mixed traffic flow collaborative control system provided in this embodiment can be implemented with reference to the above embodiments and will not be elaborated here.

[0060] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. Lane-level cooperative control method for mixed vehicle flows based on deep reinforcement learning, characterized in that Including: Perceive the vehicle flow, vehicle density, and vehicle speed data on the upstream and downstream of the main line in the highway merging area through roadside detectors, and construct a multi-dimensional state vector , which not only reflects the heterogeneity of the traffic space but also ensures the time continuity of speed adjustment. Among them, is the vehicle density on the upstream of the main line in the highway merging area, is the density of the upstream merging area, is the vehicle density on the ramp in the highway merging area, is the vehicle density on the downstream of the main line in the highway merging area, is the historical vehicle speed in the previous control cycle; Define the action space is the set of upper limits of vehicle speeds, is the upper limit value of the vehicle speed of a certain vehicle in A, and the value range of a is between ; the difference in the upper limits of vehicle speeds in the same lane and adjacent lanes in adjacent time periods does not exceed 20 km / h; Select the state transition probability between the upper limits of vehicle speeds in adjacent time periods by Markov decision process modeling to provide an environmental interaction mechanism for the deep reinforcement learning model; Define a deep reinforcement learning reward function, with the minimum total travel time as the objective function and the minimum differences in traffic flow and vehicle speed between the upstream and downstream of the main line in the merging area as the reward conditions to collaboratively control the traffic flow operation status; Build a deep reinforcement learning model to model the Markov decision process. Based on the traffic flow, vehicle speed of the current lane, and the mutual relationship between lanes, dynamically generate different upper limits of vehicle speed for each lane, and perform feedback optimization through the deep reinforcement learning model at each moment to continuously adapt to traffic changes and dynamically update the action space. , accurately respond to the operation requirements of each lane; Use the SUMO platform to build a simulation environment, change the penetration ratio of connected and autonomous vehicles in the mixed traffic flow, and use two datasets to evaluate the model performance and conduct case verification in environments with different penetration rates.

2. The lane-level mixed traffic flow collaborative control method based on deep reinforcement learning according to claim 1, wherein The state transition probability between the upper limits of vehicle speeds in adjacent time periods selected by Markov decision process modeling , including: Simulate the traffic behavior of mixed traffic flow using the SUMO platform and model the state transition probability between the upper limits of vehicle speeds in adjacent time periods , where the state transition process is the state and the action to the next state change process.

3. The lane-level mixed traffic flow collaborative control method based on deep reinforcement learning according to claim 1, characterized in that It also includes defining a deep reinforcement learning reward function, including: Minimize the total travel time TTS of vehicles in the merging area within the time range K, where TTS is expressed by formula (1): ; (1) Among them, represents the simulation time step; K represents the overall time interval, k represents a certain time interval, and t represents a specific moment. represents the initial number of vehicles in the road section at the start of the simulation. represents the number of vehicles entering the road section at time t. represents the number of vehicles leaving the road section at time t. is the lower limit of the moment corresponding to the time interval k. is the upper limit of the moment corresponding to the time interval k. Introduce the difference in traffic flow between the upstream and downstream of the main line in the merging area at the moment t as a sub-item of the reward function , Expressed by formula (2): ; (2) Introduce the vehicle speed difference between the upstream and downstream of the main line in the merging area at the moment t as another sub-item of the reward function , expressed by formula (3): ; (3) Among them, represents the speed of vehicle i in the merging area at time t, represents the number of vehicles in the merging area at time t, and H is the number of time steps within the control period, is the vehicle speed under free flow conditions; Taking the minimum difference in traffic flow and the minimum difference in vehicle speed between the upstream and downstream of the main line in the confluence area as the reward conditions, the total reward function is obtained , and the traffic flow operation status is collaboratively controlled, where is defined by the following formula: ; (4) Among them, , respectively represent the corresponding weights of the traffic flow difference and the vehicle speed difference between the upstream and downstream of the main line in the confluence area in the total reward function .

4. The lane-level mixed traffic flow collaborative control method based on deep reinforcement learning according to claim 1, characterized in that The construction of the deep reinforcement learning model includes: The deep reinforcement learning model includes a first neural network model and a second neural network model. The first neural network model is used to evaluate the action value in a state, and the second neural network model is used to select actions, enabling the deep reinforcement learning model to model the Markov decision process. Based on the traffic flow, vehicle speed, and the mutual relationship between lanes on the current lane, different vehicle speed limits are dynamically generated for each lane, and feedback optimization is performed at each moment through the deep reinforcement learning model to continuously adapt to traffic changes and dynamically update the action space. , accurately responding to the operation requirements of each lane; During the training process of the deep reinforcement learning model, utilize the interaction and feedback between the online network and the target network to improve the stability and robustness of the model in the complex mixed traffic flow environment, and achieve stable training through the dual-network interaction system of the online network and the target network; Among them, the online network adopts the Dueling architecture to achieve decoupled evaluation of the state value, which conforms to the formula Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))), where Q(s,a) is the action driving function representing the comprehensive expected benefit obtained by executing action a in state s, V(s) is the state value function, A(s,a) is the advantage function, meanA(s,a) is the mean of the advantages of all optional actions in the current state, V(s) and A(s,a) respectively learn the macroscopic state characteristics and microscopic action benefits of the mixed traffic flow. The macroscopic state characteristics include lane density and vehicle speed distribution, and the microscopic action benefits include lane-changing and acceleration / deceleration benefits; the target network suppresses the overestimation of Q values through periodic parameter synchronization; Combined with the generalization ability of the Dueling structure for sparse rewards, achieve the collaborative improvement of the global traffic efficiency optimization and local action decision-making accuracy in the lane merging area in the dynamic traffic scenario. Finally, through the delayed update and decoupled evaluation mechanism, ensure the stable convergence of the policy gradient of vehicles in the complex road conditions with frequent interactions and unbalanced rewards.

5. The lane-level mixed traffic flow collaborative control method based on deep reinforcement learning according to claim 1, characterized in that, It also includes: Construct a lane-level double-layer network control strategy for the dynamic characteristics of mixed traffic flow in the highway merging area; The lane-level double-layer network control strategy includes: Adopt speed smoothing control for the inner high-flow lane to stabilize the traffic flow, and implement speed coordination control for the outer lane with frequent lane-changes to optimize the lane-changing gap; By establishing a speed coordination optimization model between lanes, minimize the speed difference, and set up a dynamic response mechanism to adjust the speed limit scheme in real time when detecting sudden changes in traffic conditions, thereby improving the overall traffic efficiency of the merging area.

6. The lane-level mixed traffic flow collaborative control method based on deep reinforcement learning according to claim 1, characterized in that The use of the SUMO simulation model to build a simulation environment and the performance evaluation and case verification of the model using two datasets in environments with different penetration rates include: Use the SUMO simulation platform for simulation verification. During the simulation process, simulate a variety of different mixed traffic flow scenarios, where the mixed traffic flow scenarios include traffic flows composed entirely of human-driven vehicles, traffic flows composed entirely of autonomous vehicles, and situations where traffic flows composed of human-driven vehicles and traffic flows composed of autonomous vehicles are mixed; Model different types of traffic flows. The simulation platform provides detailed data on lane flow, vehicle speed, and traffic density, and obtains simulation results, which are used to evaluate the adaptability and effectiveness of the deep reinforcement learning model in different traffic environments; By comparing with traditional control methods, evaluate the performance of the deep reinforcement learning model under mixed traffic flow conditions, and verify the advantages of the deep reinforcement learning model in improving the traffic capacity of the merging area, reducing traffic congestion, and optimizing vehicle speed control.

7. Lane-level mixed traffic flow collaborative control system based on deep reinforcement learning, characterized in that, It includes: A data collection and state construction module, which is used to sense the vehicle flow, vehicle density, and vehicle speed data of the upstream and downstream of the main line in the highway merging area in real time through roadside detectors; Construct a multi-dimensional state vector according to the sensed data; And based on the Markov decision process, model the state transition probability between the upper limits of vehicle speeds in adjacent time periods, providing an environmental interaction mechanism for the deep reinforcement learning model; A model construction module, which is used to construct a lane-level mixed traffic flow collaborative control model based on the deep reinforcement learning model. The lane-level mixed traffic flow control model aims at the optimal traffic efficiency in the highway merging area to complete the training and testing of the model; the lane-level mixed traffic flow control model can perform more refined control under different mixed traffic flow penetration conditions through a dual-network architecture, improving the overall efficiency of the merging area; Achieve stable training through the dual-network interaction system of the online network and the target network. Among them, the online network uses the Dueling architecture to achieve the decoupled evaluation of the state value, decomposes the Q value into the state value function V(s) and the advantage function A(s,a), learns the macroscopic state characteristics and microscopic action benefits of the mixed traffic flow respectively, and generates Q(s,a)=V(s)+(A(s,a)−mean(A(s,a))) through the aggregation layer; the target network suppresses the overestimation of the Q value through periodic parameter synchronization; at the same time, combined with the generalization ability of the Dueling structure for sparse rewards, achieve the coordinated improvement of the global traffic efficiency and local action decision accuracy in the lane merging area under dynamic traffic scenarios. Finally, through the delayed update and decoupled evaluation mechanism, ensure the stable convergence of the policy gradient of vehicles in complex road conditions with frequent interactions and unbalanced rewards; An example verification module, which is used to input the highway merging area data set into the lane-level mixed traffic flow collaborative control model based on the deep reinforcement learning model to achieve the example verification of the mixed traffic flow collaborative control.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the lane-level mixed traffic flow collaborative control method based on deep reinforcement learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Variable speed-limiting control method for divided lanes based on intelligent network connection special lane environment

    CN116229706A

  • Speed coordination and cooperative confluence combined control method for intelligent network connection special lane scene

    CN116758739A

  • CAV speed guidance system and method based on deep reinforcement learning

    CN117612396A

  • Hard shoulder driving and variable speed limit cooperative control method and system

    CN119672981A

  • Method for intelligent traffic scheduling based on deep reinforcement learning

    US20230362095A1