Highway confluence area cooperative control method based on deep reinforcement learning

By applying a collaborative control method based on deep reinforcement learning in the highway confluence area, the collaborative control of variable speed limit and ramp metering is achieved, which solves the problem that existing traffic control methods are difficult to cope with varying traffic conditions, and improves traffic efficiency and safety.

CN120088976AActive Publication Date: 2025-06-03NANJING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510005116.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-03
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The existing traffic control methods are difficult to deal with the changing traffic conditions in highway confluence areas in real time, resulting in inefficient traffic efficiency and increased vehicle delays.

Method used

The coordinated control method of highway convergence zones based on deep reinforcement learning is adopted to achieve coordinated control of variable speed limit and ramp metering through a shared reward mechanism, and the method of sharing experience is applied to multiple convergence zones to achieve overall control optimization.

Benefits of technology

This method can more effectively reduce traffic congestion, improve road traffic efficiency, reduce traffic accident rates, and have higher flexibility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088976A_ABST
    Figure CN120088976A_ABST
Patent Text Reader

Abstract

The invention discloses a deep reinforcement learning-based cooperative control method for an expressway confluence area, and the method comprises the steps: building a LiikeSim-Python joint simulation environment, and setting a coil detector in the simulation environment to obtain the traffic flow data of the upstream and downstream of the expressway confluence area; an EM algorithm based on Gaussian mixture distribution is used as a traffic state classifier, traffic flow data of the highway confluence area are used as input, and traffic states of the highway confluence area are divided; designing a state space, an action space and a reward function; taking the state space of the highway confluence area as input, taking the actions of the variable speed-limiting intelligent agent and the ramp metering intelligent agent as output, and constructing a network model of multi-agent sharing experience under time sequence characteristics; an independent experience pool is set for each of the variable speed limit agent and the ramp metering agent, and interaction experience of the agents and the traffic simulation environment is collected with the control period as the frequency; training an agent model by using the sampled samples; and realizing cooperative control of the highway confluence area by using the trained intelligent agent model. According to the invention, the traffic travel delay of the highway confluence area can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to technical fields such as deep reinforcement learning and traffic control, and particularly relates to a cooperative control method for highway merge areas based on deep reinforcement learning. Background Art

[0002] Highway merge areas are regions in traffic flows where bottlenecks are easily formed, leading to frequent traffic congestion and accidents. Existing traffic control methods, including variable speed limit control and ramp control methods, can significantly regulate traffic flows in ramp bottleneck areas and are the most effective control methods for alleviating highway congestion. With the increase in traffic volume, traffic management and control in merge areas become increasingly complex and important. Traditional traffic control methods, such as fixed signal control and rule-based control methods, often struggle to respond to ever-changing traffic conditions in real time, resulting in low traffic efficiency and increased vehicle delays.

[0003] Deep Reinforcement Learning (DRL), as a method that combines the advantages of deep learning and reinforcement learning, has demonstrated excellent decision-making and adaptive capabilities in complex dynamic environments. Through continuous interaction with the environment, the agent can autonomously learn the optimal strategy and adapt to changes in the environment, making it particularly suitable for handling complex traffic problems with multiple variables and multiple objectives. The cooperative control method for highway merge areas based on deep reinforcement learning can, by perceiving information such as traffic volume, vehicle speed, and queue length in the environment, enable the agent to adjust merge strategies in real time, such as vehicle speed control and signal scheduling, to optimally coordinate traffic flow in the merge area. Compared with traditional methods, the deep reinforcement learning method is more flexible and adaptable, can effectively reduce traffic congestion, improve road traffic efficiency, and reduce the accident rate. By comprehensively utilizing advanced sensing technologies, communication technologies, and intelligent algorithms, the cooperative control system based on deep reinforcement learning represents an important direction for the future development of intelligent transportation. Summary of the Invention

[0004] The object of the present invention is to propose a cooperative control method for highway merge areas based on deep reinforcement learning. On the one hand, through a shared reward mechanism, the cooperative control of variable speed limit and ramp metering is achieved. On the other hand, through the method of sharing experience, the model can be effectively applied to highway corridors with multiple merge areas to achieve the optimization of overall control.

[0005] The technical solution for achieving the object of the present invention is: A cooperative control method for highway merge areas based on deep reinforcement learning, comprising the following steps:

[0006] Step 1: Establish a LiikeSim-Python joint simulation environment based on the real road network environment and traffic flow data, and set up loop detectors in the simulation environment to obtain traffic flow data upstream and downstream of the highway merge area;

[0007] Step 2: Use the EM algorithm based on Gaussian mixture distribution as a traffic state classifier, and take the traffic flow data of the highway merge area as input to divide the traffic state of the highway merge area;

[0008] Step 3: Design the state space, action space, and reward function according to the traffic flow data upstream and downstream of the highway merge area obtained by the loop detector in Step 1 and the traffic state of the merge area obtained by the traffic state classifier in Step 2;

[0009] Step 4: Take the state space of the highway merge area in Step 3 as input and the actions of the variable speed limit agent and ramp metering agent as output to construct a multi-agent shared experience network model under temporal characteristics, including an LSTM temporal feature fusion module for extracting the temporal characteristics of highway traffic flow and a D3QN decision-making module for outputting agent actions;

[0010] Step 5: Set up an independent experience pool for the variable speed limit and ramp metering agents respectively and collect the interaction experience between the agent and the traffic simulation environment at the frequency of the control period in Step 3;

[0011] Step 6: Randomly sample from the experience pool according to the set batch size, and use the sampled samples to train the agent models including the variable speed limit agent and ramp metering agent;

[0012] Step 7: Repeat Steps 5 and 6 until the reward reaches a convergent state, save the agent model parameters, and use the trained agent model to achieve the coordinated control of the highway merge area.

[0013] Further, in Step 1, to establish a LiikeSim-Python joint simulation environment based on the real road network environment and traffic flow data, and set up loop detectors in the simulation environment to obtain traffic flow data upstream and downstream of the highway merge area, the following specific steps are included:

[0014] Step 1-1: Use the visualization interface of LiikeSim to draw the road network of the highway merge area, and set the traffic flow information in the road network according to the real traffic flow data, including the maximum speed, maximum acceleration / deceleration, vehicle path, and departure time of the vehicle. Further, set up loop detectors at the end of the upstream section of the highway merge area, the end of the variable speed limit section, the ramp entrance, the merge area, and the downstream section of the merge area, and set the detection period to 60 seconds. The schematic diagram is as Figure 2as shown, and save it as the Scenario.xml road network file;

[0015] Step 1-2: Establish a LiikeSim-Python co-simulation environment based on the road network file and application programming interface in Step 1-1, and obtain the traffic flow data upstream and downstream of the highway merge area through the loop detectors set at fixed positions in Step 1-1, including the traffic flow q of the upstream section of the merge area in , the traffic flow q at the ramp entrance r , the queue length w of ramp vehicles r , the traffic flow q of the downstream section of the merge area out , the traffic flow q of the merge area m , the average vehicle density ρ of the merge area m and the average vehicle speed v of the merge area m , where the average vehicle density ρ of the merge area m The calculation formula is as follows, and the rest of the traffic flow data can be directly obtained by the detector.

[0016]

[0017] Among them, the average vehicle density ρ of the merge area m The unit is veh / km, N is the number of vehicles in the merge area, and L is the length of the merge area interval, with the unit of meter.

[0018] Further, in Step 2, use the EM algorithm based on Gaussian mixture distribution as the traffic state classifier, and take the traffic flow data of the highway merge area as the input to divide the traffic state of the highway merge area, which specifically includes the following steps:

[0019] Step 2-1: Use the EM algorithm based on Gaussian mixture distribution to construct a traffic state classifier, and take the traffic flow q of the merge area, the average vehicle density ρ of the merge area m , and the average vehicle speed v of the merge area m obtained by the loop detector at each time step as the input, and the output is the traffic state Q of the merge area corresponding to each time step, where the traffic state of the merge area is divided into 3 categories, including unobstructed, moderately congested, and congested; m

[0020] Step 2-2: Take the traffic flow q of the historical merge area, the average vehicle density ρ of the merge area m , and the average vehicle speed v of the merge area m collected in the simulation environment as the training set to train the traffic state classifier; for the trained traffic state classifier, input the traffic flow q of the merge area, the average vehicle density ρ of the merge area m collected in the simulation environment in real time m , the average vehicle density ρ of the merge area m ​The average vehicle speed v of the merging area m That is, the traffic state Q of the merging area at the current time step is obtained.

[0021] Furthermore, in step 3, according to the traffic flow data of the upstream and downstream of the highway merging area obtained by the loop detector in step 1 and the traffic state of the merging area obtained by the traffic state classifier in step 2, the state space, action space, and reward function are designed as follows:

[0022] (1) State space S

[0023] The state space is represented by a one-dimensional vector s = [q in , q r , w r , q out , q m , ρ m , v m , Q] to represent the traffic flow state of the highway merging area, where [q in , q r , w r , q out , q m , ρ m , v m are the traffic flow data of the upstream and downstream of the highway merging area obtained, and Q is the traffic state of the merging area obtained by the traffic state classifier;

[0024] (2) Action space

[0025] The actions of the variable speed limit and ramp metering agent are respectively set as the speed limit of the discrete variable speed limit section and the green light phase duration at the ramp entrance. The speed limit of the variable speed limit section is set as [60, 65, 70, 75, 80, 85, 90, 100, 110, 120] km / h, and the green light phase duration at the ramp entrance is set as [6, 12, 18, 24, 30, 36, 42, 48, 54, 60] seconds. Moreover, the control periods of the variable speed limit and ramp metering agent are both set as 60 seconds and kept synchronized;

[0026] (3) Reward function r

[0027] The variable speed limit and ramp metering agents in the same merging area share the reward, and the reward is jointly determined by the average vehicle speed in the merging area and the ramp vehicle queue length. The reward function is designed as where ω 1 , ω 2 are the weight parameters of the average vehicle speed and the ramp vehicle queue length, respectively represent the ramp vehicle queue length and the ideal ramp vehicle queue length, v m is the average vehicle speed in the merging area, v vslis the average vehicle speed on the variable speed limit section.

[0028] Further, in step 4, taking the state space of the highway merge area in step 3 as the input and the actions of the variable speed limit agent and the ramp metering agent as the output, a multi-agent shared experience network model under temporal features is constructed, including an LSTM temporal feature fusion module for extracting the temporal features of highway traffic flow and a D3QN decision-making module for outputting the actions of the agents, which specifically includes the following steps:

[0029] Step 4-1, construct a temporal feature fusion module based on the LSTM long short-term memory network, input the historical temporal features of the traffic flow in the highway merge area into the temporal feature fusion module, and effectively manage and capture the long-term and short-term dependencies of the historical temporal features of the traffic flow in the merge area through the input gate, forget gate, and output gate, and then output the traffic flow features that fuse the historical temporal features. At the same time, set the length of the input historical temporal features of the traffic flow to len;

[0030] Step 4-2, construct a decision-making module for the variable speed limit and ramp metering agents based on the D3QN deep reinforcement learning algorithm. The decision-making module takes the output of the temporal feature fusion module in step 4-1 as the input, and separates the state value and action advantage by introducing a dueling network. The dueling network has two branches: a state value network and an advantage network. Among them, the advantage network is composed of two fully connected layers in series, with sizes of 256×256 and 256×10 respectively. The state value network is composed of fully connected layers with sizes of 256×256 and 256×1 in series. The state-action value is obtained by taking the average difference of the output of the advantage network and summing it with the output of the state value network, that is, Q(s,a;θ) = V(s;θ) + A(s,a;θ) - mean a A(s,a;θ), where Q(s,a;θ) represents the state-action value function of the agent, V(s;θ) represents the output of the state value network of the agent, and A(s,a;θ) and mean a A(s,a;θ) respectively represent the output of the advantage network of the agent and the average value of the output of the advantage network; subsequently, select the section speed limit and the green light phase duration as the output strategy of the decision-making module according to the ε-greedy strategy;

[0031]

[0032] where r is a random number subject to a uniform distribution, that is a i represents the output strategy of agent i, that is, the section speed limit and the green light phase duration, and Q i (s,a i ;θ i ) is the state-action value function of agent i, is the action space of agent i, where i = 1, 2 represent the variable speed limit and ramp metering agents respectively.

[0033] Furthermore, in step 5, set up an independent experience pool for the variable speed limit and ramp metering agents respectively and collect the interaction experiences between the agents and the traffic simulation environment at the frequency of the control period in step 3. Specifically:

[0034] Create two python lists as the experience pools for the variable speed limit and ramp metering respectively. Each element in the list represents a state transition <s, a i , s next , r, T>, where s represents the traffic flow state in the highway merge area, a i represents the action selected by agent i based on the ε-greedy policy in step 4-2, s next represents the traffic flow state in the highway merge area in the next control period after the agent executes the action, r is the reward calculated from the next state, T is the end flag of the episode. Define the end flag of the episode as: when T is 1, it means the episode ends; when T is 0, it means the episode has not ended yet; the variable speed limit and ramp metering agents respectively interact with the simulation environment to generate experience sequences <s, a 1 , s next , r, T> and <s, a 2 , s next , r, T>, and cache the experience sequences in their respective experience pools respectively, i = 1, 2. If the number of experiences exceeds the capacity of the experience pool, the old experiences will be overwritten and discarded.

[0035] Furthermore, in step 6, randomly sample according to the set batch size from the experience pool, and use the sampled samples to train the agent models including the variable speed limit agent and the ramp metering agent. Specifically, it includes the following steps:

[0036] Step 6-1, respectively in the experience pools of the variable speed limit and ramp metering agents randomly draw a batch of experience samples as the input of each agent network according to the set fusion time series length len, and calculate the target value function and the loss function respectively, and update the parameters of the evaluation network of each agent according to the loss function L i , where i = 1, 2 represent the variable speed limit and ramp metering agents respectively, θ i is the parameter of the evaluation network, is the parameter of the target network, and γ is the discount factor set to 0.98; every certain number of steps Soft update the target network parameters of the variable speed limit and ramp metering agents, that is τ is the update ratio factor, set to 0.005;

[0037] Step 6-2, update the greedy coefficient ε of the agent in each episode, and the update formula of the greedy coefficient ε is where ε 0 is the initial greedy coefficient, ε end represents that the greedy coefficient ε converges to ε end , D is the decay coefficient, set to 90, and epsiode represents the number of episodes of the current simulation.

[0038] A cooperative control method for highway merge areas based on deep reinforcement learning. Implement the cooperative control method for highway merge areas based on deep reinforcement learning to achieve cooperative control of highway merge areas based on deep reinforcement learning.

[0039] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the cooperative control method for highway merge areas based on deep reinforcement learning is implemented to achieve cooperative control of highway merge areas based on deep reinforcement learning.

[0040] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the cooperative control method for highway merge areas based on deep reinforcement learning is implemented to achieve cooperative control of highway merge areas based on deep reinforcement learning.

[0041] Compared with the prior art, the significant advantage of the present invention is that: using the microscopic traffic simulation software LiikeSim and the application programming interface to build a LiikeSim-Python joint simulation environment, and designing a multi-agent reinforcement learning algorithm sharing experience, which can efficiently train the algorithm model. Description of the Drawings

[0042] Figure 1 is a schematic diagram of the cooperative control method for highways based on deep reinforcement learning of the present invention.

[0043] Figure 2 is a schematic diagram of the placement of loop detectors in the present invention.

[0044] Figure 3 is a relationship curve between the reward function and the vehicle queue length of the present invention.

[0045] Figure 4 is the time series feature fusion module of the present invention.

[0046] Figure 5It is the D3QN network structure of the present invention.

[0047] Figure 6 It is the pseudo-code of the control algorithm of the present invention.

[0048] Figure 7 It is the LiikeSim road network scenario of the embodiment of the present invention.

[0049] Figure 8 It is the relevant comparison result of the embodiment of the present invention. Detailed implementation manners

[0050] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than limiting the invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention fall within the scope defined by the appended claims of the present application.

[0051] The present invention provides a cooperative control method for highway merge areas based on deep reinforcement learning. The schematic diagram of the overall control method is as Figure 1 shown. This control scheme includes traffic state division and the cooperative work of variable speed limit agents and ramp metering agents. The variable speed limit agents and ramp metering agents obtain the traffic flow state data of the highway merge area through the simulation environment LiikeSim as the model input. Their outputs are respectively the speed limit of the variable speed limit section and the duration of the green light phase at the ramp entrance. Through the application programming interface, the two agents can effectively control the traffic flow of the simulation platform in real time.

[0052] Step 1: Establish a LiikeSim-Python joint simulation environment according to the real road network environment and traffic flow data, and set loop detectors in the simulation environment to obtain the traffic flow data of the upstream and downstream of the highway merge area. The specific steps are as follows:

[0053] Step 1-1: Use the visualization interface of LiikeSim to draw the road network of the highway merge area, and set the traffic flow information in the road network according to the real traffic flow data, including the maximum speed, maximum acceleration / deceleration, vehicle path and departure time of the vehicle. Further, loop detectors are set at the end of the upstream section of the highway merge area, the end of the variable speed limit section, the ramp entrance, the merge area and the downstream section of the merge area, and the detection period is set to 60 seconds. The schematic diagram is as Figure 2 shown, and it is saved as the Scenario.xml road network file.

[0054] Step 1-2: Establish a LiikeSim-Python co-simulation environment based on the road network file and application programming interface in Step 1-1, and obtain the traffic flow data upstream and downstream of the highway merge area through the loop detectors set at fixed positions in Step 1-1, including the traffic flow q of the upstream section of the merge area in , the traffic flow q at the ramp entrance r , the queue length w of ramp vehicles r , the traffic flow q of the downstream section of the merge area out , the traffic flow q of the merge area m , the average vehicle density ρ of the merge area m and the average vehicle speed v of the merge area m , where the average vehicle density ρ of the merge area m The calculation formula is as follows, and the rest of the traffic flow data can be directly obtained by the detector.

[0055]

[0056] Among them, the average vehicle density ρ of the merge area m The unit is veh / km, N is the number of vehicles in the merge area, and L is the length of the merge area interval, with the unit of meter.

[0057] Step 2: Use the EM algorithm based on Gaussian mixture distribution as a traffic state classifier, take the traffic flow data of the highway merge area as input, and divide the traffic state of the highway merge area as an additional feature of the state space of the reinforcement learning algorithm, specifically including the following steps:

[0058] Step 2-1: Use the EM algorithm based on Gaussian mixture distribution to construct a traffic state classifier, taking the traffic flow q of the merge area, the average vehicle density ρ of the merge area, and the average vehicle speed v of the merge area obtained by the loop detector at each time step in Step 1-2 m , the average vehicle density ρ of the merge area m and the average vehicle speed v of the merge area m as input, and the output is the traffic state Q of the merge area corresponding to each time step, where the traffic state of the merge area is divided into 3 categories, including unobstructed, relatively congested, and congested.

[0059] Step 2-2: Take the traffic flow q of the historical merge area, the average vehicle density ρ of the merge area, and the average vehicle speed v of the merge area collected in the simulation environment m , the average vehicle density ρ of the merge area m and the average vehicle speed v of the merge area m as the training set to train the traffic state classifier; for the trained traffic state classifier, input the traffic flow q of the merge area, the average vehicle density ρ of the merge area, and the average vehicle speed v of the merge area collected in the simulation environment in real time m , the average vehicle density ρ of the merge area m and the average vehicle speed v of the merge area mThe traffic state Q of the merging area at the current time step can be obtained.

[0060] Step 3: Design the state space, action space, and reward function based on the traffic flow data of the upstream and downstream of the highway merging area obtained by the loop detector in Step 1 and the traffic state of the merging area obtained by the traffic state classifier in Step 2:

[0061] (1) State space S

[0062] The state space is represented by a one-dimensional vector s = [q in , q r , w r , q out , q m , ρ m , v m , Q] to represent the traffic flow state of the highway merging area, where [q in , q r , w r , q out , q m , ρ m , v m is the traffic flow data of the upstream and downstream of the highway merging area obtained in Steps 1 - 2, and Q is the traffic state of the merging area obtained by the traffic state classifier in Step 2.

[0063] (2) Action space

[0064] The actions of the variable speed limit and ramp metering agents are respectively set as the speed limit (km / h) of the discrete variable speed limit section and the green light phase duration (s) of the ramp entrance. The speed limit of the variable speed limit section is set as [60, 65, 70, 75, 80, 85, 90, 100, 110, 120] km / h, and the green light phase duration of the ramp entrance is set as [6, 12, 18, 24, 30, 36, 42, 48, 54, 60] seconds. Moreover, the control periods of the variable speed limit and ramp metering agents are both set as 60 seconds and kept synchronized.

[0065] (3) Reward function r

[0066] The variable speed limit and ramp metering agents in the same merging area share the reward, which is jointly determined by the average speed of the merging area and the queue length of ramp vehicles. The reward function is designed as where ω 1 , ω 2 are the weight parameters of the average speed and the queue length of ramp vehicles, respectively represent the queue length of ramp vehicles and the ideal queue length of ramp vehicles, v m is the average speed of the merging area, v vslis the average vehicle speed of the variable speed limit section. The relationship between the reward function and the ramp vehicle queue length is as Figure 3 shown. This reward function takes into account the average vehicle speed in the merging area and the ramp vehicle queue length, and then ensures the smooth and efficient operation of the traffic flow in the ramp area through the coordinated control of variable speed limits and entrance ramps.

[0067] Step 4: Using the state space of the highway merging area in Step 3 as the input and the actions of the variable speed limit agent and the ramp metering agent as the output, construct a multi-agent shared experience network model under temporal characteristics, which includes an LSTM temporal feature fusion module for extracting the temporal characteristics of highway traffic flow and a D3QN decision-making module for outputting agent actions. Specifically, it includes the following steps:

[0068] Step 4-1: Based on the LSTM long short-term memory network, construct a temporal feature fusion module, as Figure 4 shown. Input the historical temporal characteristics of the traffic flow in the highway merging area into the temporal feature fusion module, and effectively manage and capture the long-term and short-term dependencies of the historical temporal characteristics of the traffic flow in the merging area through the input gate, forget gate, and output gate, and then output the traffic flow characteristics that fuse the historical temporal characteristics. At the same time, here it is set that the length of the input historical temporal characteristics of the traffic flow is len.

[0069] Step 4-2: Based on the D3QN deep reinforcement learning algorithm, construct a decision-making module for the variable speed limit and ramp metering agents. The decision-making module takes the output of the temporal feature fusion module in Step 4-1 as the input, and by introducing a dueling network to separate the state value and action advantage, the dueling network has two branches: a state value network and an advantage network, as Figure 5 shown. Among them, the advantage network is composed of two fully connected layers in series, with sizes of 256×256 and 256×10 respectively. The state value network is composed of fully connected layers with sizes of 256×256 and 256×1 in series. The state-action value is obtained by taking the average difference of the output of the advantage network and summing it with the output of the state value network, that is, Q(s,a;θ) = V(s;θ) + A(s,a;θ) - mean a A(s,a;θ), where Q(s,a;θ) represents the state-action value function of the agent, V(s;θ) represents the output of the state value network of the agent, A(s,a;θ) and mean a A(s,a;θ) represent the output of the advantage network of the agent and the average value of the output of the advantage network respectively. Subsequently, select the section speed limit and the green light phase duration as the output strategy of the decision-making module according to the ε-greedy strategy shown in the following formula.

[0070]

[0071] where r is a random number following a uniform distribution, that is a i represents the output strategy of agent i, that is, the road segment speed limit and the green light phase duration, and Q i (s, a i ; θ i ) is the state-action value function of agent i. is the action space of agent i, where i = 1, 2 represent the variable speed limit and ramp metering agents respectively.

[0072] Step 5, set up an independent experience pool for the variable speed limit and ramp metering agents respectively and collect the interaction experiences of the agents with the traffic simulation environment at the frequency of the control period in Step 3:

[0073] Create two python lists as the experience pools for the variable speed limit and ramp metering respectively. Each element in the list represents a state transition <s, a i , s next , r, T>, where s represents the traffic flow state in the highway merge area, a i represents the action selected by agent i based on the ε-greedy strategy in Step 4-2, s next represents the traffic flow state in the highway merge area in the next control period after the agent executes the action, r is the reward calculated from the next state, and T is the end flag of the episode. Define the end flag of the episode as: when T is 1, it means the episode ends; when T is 0, it means the episode has not ended yet. The variable speed limit and ramp metering agents respectively interact with the simulation environment to generate experience sequences <s, a 1 , s next , r, T> and <s, a 2 , s next , r, T>, and cache the experience sequences in their respective experience pools respectively. For i = 1, 2, if the number of experiences exceeds the capacity of the experience pool, the old experiences will be overwritten and discarded.

[0074] Step 6, randomly sample from the experience pool according to the set batch size, and use the sampled samples to train the agent model, which specifically includes the following steps:

[0075] Step 6-1, the pseudo-code for the training process of the variable speed limit and ramp metering agent networks is as Figure 6 shown. Randomly draw a batch of experience samples as the input of each agent network from the experience pools of the variable speed limit and ramp metering agents respectively according to the length len of the fusion time series sequence set in Step 4-1, and calculate the target value function and loss function of each agent respectively. And according to the loss function L i Update the parameters of the evaluation network of each agent, where i = 1, 2 represent the variable speed limit and ramp metering agents respectively, and θ i are the parameters of the evaluation network, are the parameters of the target network, and γ is the discount factor set to 0.98; the target network parameters of the variable speed limit and ramp metering agents are softly updated every certain number of steps c, that is τ is the update ratio factor, which is set to 0.005 here.

[0076] Step 6-2, update the greedy coefficient ε of the agent in each episode, and the update formula of the greedy coefficient ε is where ε 0 is the initial greedy coefficient, and ε end indicates that the greedy coefficient ε converges to ε end , D is the attenuation coefficient, set to 90, and epsiode represents the number of episodes of the current simulation.

[0077] Step 7, repeat Step 5 and Step 6 until the reward reaches the convergence state, and save the model parameters, and use the trained model to achieve the cooperative control of the highway merge area.

[0078] Embodiment

[0079] In order to verify the effectiveness of the proposed solution of the present invention, the following simulation experiments are carried out.

[0080] In this embodiment, the traffic simulation environment is built by the microscopic traffic simulation software LiikeSim. LiikeSim can be used in multiple fields such as studying traffic flow characteristics and traffic management strategy evaluation. At the same time, in the Python environment, the application programming interface can be used to achieve dynamic interaction with the LiikeSim simulation environment, such as obtaining vehicle data, lane data, signal phase control, road speed limit control, etc., so as to achieve the joint simulation of LiikeSim-Python.

[0081] In this embodiment, the simulation scenario is a highway merge area road network scenario built according to the real road network. The road network scenario is as Figure 7 shown, and based on the real traffic flow data as the vehicle input of the simulation environment. This embodiment contains 7200s of traffic flow data volume, and the traffic flow data is shown in Table 1. In order to effectively evaluate the control effect of the present invention, the average vehicle speed in the merge area is used as the evaluation index.

[0082] Table 1 Traffic demand flow

[0083]

[0084] In this embodiment, the traffic control unit consists of a variable speed limit and ramp metering agent. The traffic control unit obtains the traffic flow states s upstream and downstream of the highway merge area in real time, and selects the speed limit a 1 and the green light phase duration a of the traffic signal 2 , and can significantly change the traffic flow state s' within the current control cycle of the highway merge area by publishing speed limit information on the variable speed limit interval signs a 1 and changing the green light phase duration a 2 . At the same time, the variable speed limit and ramp metering agents receive the immediate reward r from the environmental state s', and cache the experience sequences <s, a 1 , r, s', T> and <s, a 2 , r, s', T>. The variable speed limit and ramp metering agent network continuously interacts with the environment to learn experience and training parameters. As the number of updates i of the network parameter θ i approaches infinity, the action value evaluation function Q(s, a; θ i ) of the network infinitely approaches the optimal action value evaluation function Q * (s, a), that is, Q(s, a; θ i ) → Q * (s, a). Finally, the trained variable speed limit and ramp metering agent network model can be loaded, and the optimal action policy can be obtained through the greedy strategy, that is, a * = argmax a Q * (s, a).

[0085] In this embodiment, the above method is trained 5 times through simulation. Each training lasts for 300 rounds, and the simulation time for each round is 7200 s. Finally, the results of the 5 trainings are averaged to be used as the final effect evaluation of this solution. To verify the effectiveness of this solution, this embodiment is compared with the no control strategy, fixed timing control, and PI_ALINEA control methods. The relevant comparison results are as Figure 8 shown. It can be concluded that this solution shows more significant advantages in improving the traffic efficiency of the highway merge area compared with the traditional control methods.

[0086] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as the scope recorded in this specification.

[0087] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A collaborative control method for highway merging areas based on deep reinforcement learning, characterized in that: The steps include: Step 1: Establish a LiikeSim-Python joint simulation environment based on the real road network environment and traffic flow data, and set up a coil detector in the simulation environment to obtain traffic flow data upstream and downstream of the highway merging area; Step 2, using the EM algorithm based on Gaussian mixture distribution as a traffic state classifier, taking the traffic flow data of the highway merging area as input, and dividing the traffic state of the highway merging area; Step 3, designing a state space, an action space, and a reward function based on the traffic flow data of the upstream and downstream of the highway merging area obtained by the coil detector in step 1 and the traffic state of the merging area obtained by the traffic state classifier in step 2; Step 4: Using the state space of the highway merging area in step 3 as input and the actions of the variable speed limit agent and the ramp metering agent as output, a network model for multi-agent shared experience under time series features is constructed, including an LSTM time series feature fusion module for extracting time series features of highway traffic flow and a D3QN decision module for outputting agent actions. Step 5: Set up a separate experience pool B for the variable speed limit and ramp metering agents respectively. i And collect the interaction experience between the agent and the traffic simulation environment at the control cycle in step 3; Step 6: randomly sample from the experience pool according to the set batch size, and use the sampled samples to train the agent model, including the variable speed limit agent and the ramp metering agent; Step 7: Repeat steps 5 and 6 until the reward reaches a convergence state, save the agent model parameters, and use the trained agent model to achieve collaborative control of the highway merging area.

2. The method for coordinated control of highway merging areas based on deep reinforcement learning according to claim 1 is characterized in that: Step 1: Establish a LiikeSim-Python joint simulation environment based on the real road network environment and traffic flow data, and set up a coil detector in the simulation environment to obtain traffic flow data upstream and downstream of the highway merging area. The specific steps include: Step 1-1, use the visualization interface of LiikeSim to draw the highway merging area road network, set the traffic information in the road network according to the real traffic flow data, including the maximum speed, maximum acceleration / deceleration, vehicle path and departure time of the vehicle, and further, set the coil detector at the end of the upstream section of the highway merging area, the end of the variable speed limit section, the ramp entrance, the merging area and the downstream section of the merging area, and set the detection cycle to 60 seconds, and save it as the Scenario.xml road network file; Step 1-2: Establish the LiikeSim-Python joint simulation environment based on the road network file and application program interface in step 1-1, and obtain the traffic flow data upstream and downstream of the merging area of ​​the expressway through the coil detector set at a fixed position in step 1-1, including the flow q of the upstream section of the merging area. in , the flow rate at the ramp entrance q r , ramp vehicle queue length w r , the flow rate q of the downstream section of the merging area out , the flow rate q in the confluence area m , the average vehicle density in the merging area ρ m and the average vehicle speed v in the merging area m , where the average vehicle density in the merging area is m The calculation formula is as follows. The rest of the traffic flow data can be directly obtained by the detector. Among them, the average vehicle density in the merging area is m The unit is veh / km, N is the number of vehicles in the merging area, and L is the length of the merging area in meters.

3. The method for coordinated control of highway merging areas based on deep reinforcement learning according to claim 1 is characterized in that: Step 2, using the EM algorithm based on Gaussian mixture distribution as a traffic state classifier, taking the traffic flow data of the highway merging area as input, and dividing the traffic state of the highway merging area, specifically including the following steps: Step 2-1: Use the EM algorithm based on Gaussian mixture distribution to build a traffic state classifier, using the flow rate q in the merging area obtained by the loop detector at each time step m , the average vehicle density in the merging area ρ m and the average speed v in the merging area m The input is the traffic state Q of the merging area corresponding to each time step, where the traffic state of the merging area is divided into three categories, including smooth, relatively congested, and congested; Step 2-2: The flow rate q of the historical confluence area collected in the simulation environment m , the average vehicle density in the merging area ρ m and the average speed v in the merging area m As a training set, the traffic state classifier is trained. For the trained traffic state classifier, the flow rate q in the merging area collected in the simulation environment is input in real time. m , the average vehicle density in the merging area ρ m and the average speed v in the merging area m That is, the traffic state Q of the merging area at the current time step is obtained.

4. The method for coordinated control of highway merging areas based on deep reinforcement learning according to claim 1 is characterized in that: Step 3: Design the state space, action space and reward function based on the traffic flow data of the upstream and downstream of the highway merging area obtained by the coil detector in step 1 and the traffic state of the merging area obtained by the traffic state classifier in step 2. Specifically, (1) State space S The state space is composed of a one-dimensional vector s = [q in ,q r ,w r ,q out ,q m ,ρ m ,v m ,Q] to represent the traffic flow state of the highway merging area, where [q in ,q r ,w r ,q out ,q m ,ρ m ,v m ] is the traffic flow data obtained upstream and downstream of the merging area of ​​the highway, Q is the traffic state of the merging area obtained by the traffic state classifier; (2) Action Space The variable speed limit and ramp metering agent actions are set as the speed limit of the discrete variable speed limit section and the green light phase duration of the ramp entrance, respectively. The speed limit of the variable speed limit section is set to [60, 65, 70, 75, 80, 85, 90, 100, 110, 120] km / h, and the green light phase duration of the ramp entrance is set to [6, 12, 18, 24, 30, 36, 42, 48, 54, 60] seconds. The control cycle of the variable speed limit and ramp metering agent is set to 60 seconds and kept synchronized. (3) Reward function r The variable speed limit and ramp metering agents in the same merging area share rewards, which are determined by the average speed of the merging area and the length of the ramp vehicle queue. The reward function is designed as Among them, ω1 and ω2 are weight parameters of average vehicle speed and ramp vehicle queue length, Represent the ramp vehicle queue length and the ideal ramp vehicle queue length, v m is the average vehicle speed in the merging area, v vsl is the average vehicle speed on the variable speed limit section.

5. The method for coordinated control of highway merging areas based on deep reinforcement learning according to claim 1 is characterized in that: Step 4: Taking the state space of the highway merging area in step 3 as input and the actions of the variable speed limit agent and the ramp metering agent as output, a network model for multi-agent shared experience under time series features is constructed, which includes an LSTM time series feature fusion module for extracting time series features of highway traffic flow and a D3QN decision module for outputting agent actions. Specifically, the following steps are included: Step 4-1, construct a time series feature fusion module based on the LSTM long short-term memory network, input the historical time series features of the traffic flow in the merging area of ​​the expressway into the time series feature fusion module, and effectively manage the historical time series features of the traffic flow in the merging area and capture the long-term and short-term dependencies through the input gate, forget gate and output gate, and then output the traffic flow features that integrate the historical time series features, and set the length of the input traffic flow historical time series features to len; Step 4-2, based on the D3QN deep reinforcement learning algorithm, a decision module for the variable speed limit and ramp metering agent is constructed. The decision module takes the output of the time series feature fusion module in step 4-1 as input, and separates the state value and action advantage by introducing the duel network. The duel network has two branches: the state value network and the advantage network. The advantage network is composed of two fully connected layers in series, with sizes of 256×256 and 256×10 respectively. The state value network is composed of fully connected layers in series with sizes of 256×256 and 256×1 respectively. The state action value is the average difference of the advantage network output and the sum of the state value network output, that is, Q(s,a;θ)=V(s;θ)+A(s,a;θ)-mean a A(s,a;θ), where Q(s,a;θ) represents the state-action value function of the agent, V(s;θ) represents the state-value network output of the agent, A(s,a;θ) and mean a A(s,a;θ) represents the dominant network output and the average value of the dominant network output of the agent, respectively; then, the road speed limit and green light phase duration are selected as the output strategy of the decision module according to the ε-greedy strategy; Among them, r is a random number that obeys uniform distribution, that is, a i represents the output strategy of agent i, namely the road speed limit and green light phase duration, Q i (s,a i θ i ) is the state-action value function of agent i, is the action space of agent i, where i=1, 2 represents the variable speed limit and ramp metering agents, respectively.

6. The method for coordinated control of highway merging areas based on deep reinforcement learning according to claim 1, characterized in that: Step 5: Set up a separate experience pool for the variable speed limit and ramp metering agents The interaction experience between the agent and the traffic simulation environment is collected at the control cycle in step 3, specifically: Create two python lists as experience pools for variable speed limits and ramp meters, respectively. Each element in the list represents a state transition. <s,a i ,s next ,r,T>, where s represents the traffic flow state of the highway merging area, a i represents the action selected by agent i based on the ε-greedy strategy in step 4-2, s next represents the traffic flow state of the highway merging area in the next control cycle after the agent performs the action, r is the reward calculated for the next state, T is the end mark of the round, and the end mark of the round is defined as: when T is 1, it means the round is over, and when T is 0, it means the round has not ended; the variable speed limit and ramp metering agents interact with the simulation environment to generate experience sequences s, a respectively. 1 ,s next ,r,T> and <s,a 2 ,s next ,r,T>, and cache the experience sequence in their respective experience pools In the example, i=1,2, if the amount of experience exceeds the capacity of the experience pool, the old experience will be overwritten and discarded.

7. The method for coordinated control of highway merging areas based on deep reinforcement learning according to claim 1 is characterized in that: Step 6: randomly sample from the experience pool according to the set batch size, and use the sampled samples to train the agent model including the variable speed limit agent and the ramp metering agent, which specifically includes the following steps: Step 6-1, respectively in the experience pool of the variable speed limit and ramp metering agent In the set fusion time series length len, a batch of experience samples are randomly selected as the input of each agent network, and the target value function of each agent is calculated respectively. And the loss function And according to the loss function L i Update the parameters of the evaluation network of each agent, where i=1,2 represents the variable speed limit and ramp metering agents, respectively, θ i To evaluate the parameters of the network, is the parameter of the target network, γ is the discount factor set to 0.98; every certain number of steps The target network parameters of the variable speed limit and ramp metering agent are soft updated, namely τ is the update scale factor, which is set to 0.005; Step 6-2: Update the greedy coefficient ε of the agent in each round, and the update formula of the greedy coefficient ε is Among them, ε0 is the initial greedy coefficient, ε end Indicates that the greedy coefficient ε converges to ε end , D is the attenuation coefficient, which is set to 90, and epsiode represents the number of rounds of the current simulation.

8. A collaborative control system for highway merging areas based on deep reinforcement learning, characterized in that: Implement the method for collaborative control of highway merging areas based on deep reinforcement learning as described in any one of claims 1-6 to achieve collaborative control of highway merging areas based on deep reinforcement learning.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for collaborative control of a highway merging area based on deep reinforcement learning according to any one of claims 1 to 6 is implemented to realize collaborative control of a highway merging area based on deep reinforcement learning.

10. A computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for collaborative control of a highway merging area based on deep reinforcement learning according to any one of claims 1 to 6 is implemented to realize collaborative control of a highway merging area based on deep reinforcement learning.

Citation Information

Patent Citations

  • MADDPG-based automatic driving vehicle ramp confluence cooperative control method and system

    CN115273501A

  • Reinforcement learning automatic driving fleet control method based on model predictive control guidance

    CN116088530A

  • Traffic control method, device and equipment for expressway confluence area and medium

    CN117315956A

Cited By

  • Confluence area cooperative variable speed limit control method based on deep reinforcement learning

    CN121171027A

  • Large-scale expressway and connection road network cooperative control method based on multi-agent reinforcement learning

    CN121811654A

  • Method for coordinated control of large-scale expressway and connecting road network based on multi-agent reinforcement learning

    CN121811654B