Cooperative control method for traffic signal and direction-variable lane based on multi-agent learning
By adopting a collaborative control method of multi-agent learning in the intelligent transportation system, combining traffic signal control and variable direction lane control, the traffic congestion problem caused by independent operation in the existing system is solved, and more efficient traffic flow management and lane utilization are achieved.
Patent Information
- Application Number
- CN202411349081.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-05-30
AI Technical Summary
In existing intelligent transportation systems, traffic signal control and variable direction lane control usually operate independently, lacking effective collaboration, making it difficult to quickly adapt to dynamically changing traffic flows, resulting in difficult to effectively alleviate traffic congestion problems.
The coordinated control method of traffic signals and variable direction lanes based on multi-agent learning is adopted. Through the traffic signal control module and the variable direction lane guide control module interact and iterate decision-making, the reinforcement learning framework is used to adapt to dynamic traffic conditions, and more efficient coordination between traffic lights and variable direction lanes is achieved.
Through collaborative control methods, the traffic volume at intersections is maximized, the delays and length of vehicles are reduced, the efficiency and flexibility of the transportation system are significantly improved, and urban traffic congestion is effectively alleviated.
Smart Images

Figure CN120071644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic management, and particularly relates to a cooperative control method for traffic signals and variable-direction lanes based on multi-agent learning. Background Art
[0002] With the continuous increase in urban population and vehicle numbers, the problem of traffic congestion has become increasingly serious, leading to an increase in vehicle idle emissions. The application of intelligent transportation systems has brought new opportunities for improving traffic conditions. Although there are already various control methods in existing intelligent transportation systems, such as traffic signal control methods and variable-direction lane technologies, they usually operate independently and lack effective cooperation.
[0003] The total traffic volume at intersections depends on two factors: traffic flow rate and traffic time. The traffic flow rate is related to the number of lanes at intersections; the traffic time is related to the traffic signal phases. The longer the green light phase duration, the longer the traffic time. To solve this problem, current research mainly focuses on two aspects: traffic signal control methods and variable-direction lane control methods. On the one hand, by controlling the time and operating efficiency of lane switching, the utilization rate of lanes and the operating efficiency of intersections are improved. However, these methods usually rely on historical data for phase control and cannot quickly update rules to adapt to dynamically changing traffic flows. On the other hand, traffic conditions are improved through traffic signal control methods: (1) Rule-based traffic signal control methods adjust the length and sequence of traffic signals according to dynamic traffic conditions, but it is difficult for them to effectively adapt to these changes. (2) More advanced traffic signal control based on reinforcement learning (RL) directly learns from observed data without relying on unrealistic assumptions about traffic models. Although it solves the problem of dynamic traffic conditions, it is not combined with variable lane control methods.
[0004] Traditional traffic signal control methods are divided into fixed-time methods and adaptive methods, which can adapt to regular changes to a certain extent. However, their key parameters are set based on human experience and cannot cope with rapidly changing road traffic conditions, resulting in slow responses. Existing methods using reinforcement learning algorithms to better solve traffic signal control problems improve the passing time by adjusting traffic lights, thereby enhancing the passing capacity of intersections. However, they cannot handle the situation where the vehicle flow in different directions shows an unbalanced distribution at different time periods at intersections. In terms of research on the timing of lane function switching, relevant studies include: establishing a threshold model for tidal lane control at intersections based on the left-turn and straight traffic flows, using the average delay of incoming vehicles as a criterion; proposing a threshold model for tidal lanes under different degrees of saturation considering the traffic flows and queue lengths in all directions; a variable lane control method based on video detection for adaptive switching of lane guidance. In terms of research on lane operation efficiency control, relevant studies include: improving the operation efficiency of left-turn exits at intersections through a mixed-integer nonlinear programming model; proposing an empirical rule considering the setting conditions of variable guide lanes and a control model for using variable guide lanes at intersections, and achieving an improvement in lane efficiency.
[0005] Some related studies have attempted to combine variable guide lanes, traffic signals, and vehicle path planning to improve traffic efficiency. For example, some studies have proposed a vehicle-road collaborative scheduling method, which includes two modules: real-time route planning and traffic signal control, and the two interact and cooperate to achieve joint decision-making; some studies have proposed a coupled control strategy to simultaneously optimize signal timing and entrance lane settings through a piecewise linear programming model. In addition, some studies have proposed a two-layer model and designed the interaction relationship between variable guide lanes and signal control. Some studies have optimized the traffic flow at intersections by combining variable guide lanes with signal timing schemes. Some studies have considered factors such as lane function division, main signal control, and pre-signal control based on a combined design optimization model to maximize the capacity of intersections. However, although these studies have been applied in different scenarios, they often ignore the joint control problem between conventional vehicles, variable guide lanes, and traffic lights. In addition, the potential of using reinforcement learning methods for joint control has not been fully explored.
[0006] Currently, the commonly used method for traffic signal control is Webster's calculation of signal timing control. Most signal timing controls use time-based control and adopt different timing schemes during peak and off-peak hours. However, this is mainly applicable to non-saturated or slightly saturated traffic conditions and does not consider the optimization of signal phases. Although there are studies on traffic signal control based on reinforcement learning methods, traffic flow is dynamically changing in both the time and space dimensions. Reinforcement learning-based traffic signal control can only increase the effective time of phases and cannot improve lane utilization.
[0007] To this end, the present invention proposes a collaborative control method for traffic signals and variable-direction lanes based on multi-agent learning. Through the interaction and iterative decision-making of the traffic signal control module and the variable-direction lane guidance control module, and by using reinforcement learning to adapt to dynamic traffic conditions, more efficient coordination between traffic lights and variable-direction lanes is achieved to maximize the traffic volume at intersections and alleviate urban traffic congestion. Summary of the Invention
[0008] In view of the problems existing in the above-mentioned background technology, the present invention provides a collaborative control method for traffic signals and variable-direction lanes based on multi-agent learning. Through the interaction and iterative decision-making of the traffic signal control module and the variable-direction lane guidance control module, and by using reinforcement learning to adapt to dynamic traffic conditions, more efficient coordination between traffic lights and variable-direction lanes is achieved to maximize the traffic volume at intersections and alleviate urban traffic congestion.
[0009] To achieve the above object, the present invention adopts the following technical solutions: A collaborative control method for traffic signals and variable-direction lanes based on multi-agent learning, called LCSL, makes decisions iteratively through a traffic signal control module and a lane guidance control module, and runs simultaneously to control traffic; The lane guidance control module adopts a double deep Q-network (DDQN) algorithm enhanced learning model. The agent selects actions based on traffic flow and traffic signal status to maximize lane utilization rate, and the reward is calculated based on the difference in lane queue lengths and the ratio of the maximum directional demand traffic flow. The lane guidance control module controls the state, action, and reward of the agent for the variable lanes at intersections as follows: State: The state observation is mainly the traffic flow volume in the direction selected by the variable guide lane and the traffic flow volume in other lanes; at time step t The state obtained by the lane The expression is as follows: Where is the traffic flow volume of the fixed left-turn lane on the road of the variable guide lane, is the traffic flow volume of the fixed straight lane on the road of the variable guide lane, is the traffic flow volume of the left turn on the road of the variable guide lane, is the traffic flow volume of the straight line on the road of the variable guide lane, is the action of the traffic signal at time step t ; Action: The action of the variable guide lane is to meet the dynamic changes between straight and left turns. The action space set is: Among them, indicates that the lane guidance is a left turn, indicates that the lane guidance is a straight line; Reward: The ratio of the difference in queue lengths on fixed lanes to the maximum directional demand traffic flow is used as a measurement standard, and the formula is as follows: ; The double deep Q-network algorithm stabilizes the training process by introducing a target network and an experience replay pool, uses a multi-layer perceptron (MLP) neural network to fit the expected reward for a given state-action pair, and uses an action network and a target network to reduce the overestimation problem, selects the phase action with the maximum reward, and updates the parameters of the action network by minimizing the loss function; The traffic signal control module agent selects actions based on the observed lane density and the current lane state, and obtains rewards according to the traffic flow satisfaction of the selected phase; The state, actions, and rewards of the agent of the traffic signal control module that controls the intersection signals are as follows: State: Starting from the incoming road on the left side of the north side, calculate the lane density clockwise and finally add the lane actions to obtain the state of the intersection; at the time step t The state obtained by the signal lamp The expression is as follows: Among them, is the lane density of the incoming lane at the time step , is the action of the lane control agent for the lane t at the time step Actions: The agent selects different available phase sets according to the structure of the road network and traffic demand; Reward: The reward mechanism is based on the utilization rate of each traffic signal phase. When the agent selects a signal lamp phase, it selects the phase that can maximize the satisfaction of the vehicle passing demand; by considering the phase required for the maximum traffic flow; The interaction method between the traffic signal control module and the lane guidance control module is as follows: First, the agent responsible for controlling the traffic signal is trained, and the training period is episodes, and each episode contains T steps; then, based on the output of the traffic signal control agent, the agent responsible for controlling the variable lane is trained, and the training period is episodes; the training alternates between the two; finally, the agent controlling the traffic signal is retrained rounds; after training two agents, the traffic signal control module and the lane guidance control module run simultaneously to control traffic; the LCSL algorithm is as follows: where, is the value adopted under using . is under the adoption of value.
[0010] Furthermore, the parameter update formula of the action network is as follows: where, T is the number of time steps, θ are the parameters of the neural network, is the Q value of the target network, the calculation formula of the is: where, r lane represents the reward obtained by the lane agent at time step t , γ represents the discount factor, Q is the value adopted under using . Parameters are copied from the action network every C time steps, that is, .
[0011] Furthermore, in the reward of the agent that the traffic signal control module controls the intersection signal, the utilization rate of the traffic signal phase is: where, represents the reward obtained by the signal agent at time step t , represents the signal light phase selected by the signal agent at time step t , represents the traffic flow on phase t at time step , represents the phase with the maximum traffic flow stage.
[0012] Furthermore, the experience replay pool is implemented through a data structure SumTree based on a binary tree structure.
[0013] Further, the implementation method of the experience replay pool is as follows: According to the priority of each sample given to the sample : wherein, is the absolute value of the TD error of each sample, ε is a constant, and the loss function optimization formula is as follows: wherein, is the Q value of the target network, is defined as: wherein, is the ratio of the priority of the current sample to the total priority, is the minimum priority ratio, is an adjustable parameter, and finally update the TD error of the sample: .
[0014] Compared with the disadvantages and deficiencies of the prior art, the present invention has the following beneficial effects: The present invention improves the lane utilization rate by controlling the lane guidance, and the traffic signal control module and the variable direction lane guidance control module interact with each other and make decisions iteratively. Using the reinforcement learning framework, these modules can control the dynamically changing traffic flow in the spatial and temporal dimensions. The present invention designs a state and reward function based on lane balance and traffic demand, and establishes an information sharing mechanism to promote effective interaction between agents. Through extensive experimental verification, the method of the present invention performs excellently on multiple real datasets, exceeding the current advanced benchmark methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a framework diagram of LCSL provided by an embodiment of the present invention; Figure 2 is a schematic diagram of an intersection, a phase stage, and variable direction lanes provided by an embodiment of the present invention; Figure 3 is the average delay and average queue length on the simulated dataset provided by an embodiment of the present invention; wherein, Figure 3 (a) and (b) are respectively the average delay and average queue length on the simulated dataset in the non-peak period, Figure 3 (c) and (d) are respectively the average delay and average queue length on the simulated dataset in the peak period; Figure 4 is the average delay and average queue length on the real dataset provided by an embodiment of the present invention; wherein,Figure 4 (a) and (b) are the average delay and average queue length on the real - world dataset during the off - peak period, Figure 4 (c) and (d) are the average delay and average queue length on the real - world dataset during the peak period; Figure 5 is the loss function during training provided by the embodiment of the present invention ( Figure 5 (a)), delay ( Figure 5 (b)) and queue length ( Figure 5 (c)) changes. Detailed implementation manners
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] The collaborative control method of traffic signals and variable - direction lanes based on multi - agent learning is called LCSL. The framework diagram of LCSL is as Figure 1 shown. LCSL makes decisions iteratively through a traffic signal control module and a lane guidance control module, and runs simultaneously to control traffic.
[0018] An intersection consists of two parts: the entrance road and the exit road . Vehicles on the entrance road drive towards the intersection, while vehicles on the exit road drive away from the intersection. The lanes on the entrance road are divided into four categories: straight - through lanes, left - turn lanes, right - turn lanes, and variable - direction lanes. The set of all entrance lanes is defined as , and the set of all exit lanes is defined as . Figure 2 is a schematic diagram of an intersection, phase stages, and variable - direction lanes. The intersection has four entrance roads and four exit roads. The north, south, and east entrance roads are divided into three lanes, and each lane corresponds to a different turning direction; the west entrance road has a variable - direction lane; the variable - direction lane in the middle can dynamically switch functions between straight - through and left - turn.
[0019] 1. Lane guidance control module: Adopt a double - deep Q - network (DDQN) algorithm - based reinforcement learning model. The agent selects actions based on traffic flow and traffic signal states to maximize lane utilization. The reward is calculated based on the difference in lane queue lengths and the ratio of the maximum directional demand traffic flow. The DDQN algorithm stabilizes the training process by introducing a target network and an experience replay pool, overcomes the limitations of traditional Q - learning in large state spaces, and makes the algorithm more effective in dealing with traffic system problems with strong spatio - temporal correlations. The lane guidance control agent calculates according to each time interval Observation Select its phase action.
[0020] The lane guidance control module controls the state, action, and reward of the agent of the variable lane at the intersection as follows: State: The state observation is mainly the traffic flow volume in the direction selected by the variable guidance lane and the traffic flow volume in other lanes; obtaining the action of the traffic signal agent is to better coordinate and cooperate with it. The lane guidance control module controls the agent of the variable lane at the intersection at time step t The state obtained by the lane The expression is as follows: Among them, is the traffic flow volume of the fixed left-turn lane on the road of the variable guidance lane, is the traffic flow volume of the fixed straight-through lane on the road of the variable guidance lane, is the traffic flow volume of left-turning on the road of the variable guidance lane, is the traffic flow volume of going straight on the road of the variable guidance lane, is the action of the traffic signal at time step t .
[0021] Action: The action of the variable guidance lane is to meet the dynamic changes between going straight and turning left. Obtaining the action of the traffic signal agent is to better coordinate and cooperate with it. The action space set is: Among them, indicates that the lane guidance is for left-turning, indicates that the lane guidance is for going straight; Reward: The main concern of the present invention is to maximize the lane utilization rate and ensure the balance of the number of vehicles on each lane. Therefore, the present invention uses the ratio of the difference in the queue length on the fixed lane to the maximum directional demand traffic flow volume as a measurement standard, and the formula is as follows: The smaller the difference, the closer the utilization rates of the two lanes are. In this way, the guidance of the variable lane can be adjusted more effectively, making the vehicle queue lengths on the lane more balanced, thereby improving the overall lane utilization rate.
[0022] The double deep Q-network algorithm stabilizes the training process by introducing a target network and an experience replay pool, and the long-term impact of the signal control action G is defined as follows: Among them, is based on the intersection Time action Reward; Use a multi-layer perceptron (MLP) neural network to fit the expected reward for a given state-action pair and use the actor network and the target network to reduce the overestimation problem: where is the value adopted under is adopted under value, and are training parameters, is the guiding number (variable lane action space), θ are the parameters of the neural network.
[0023] Select the phase action with the maximum reward, and update the parameters of the actor network by minimizing the loss function. The formula is as follows: where T is the number of time steps, is the Q value of the target network, and the is calculated as: where r lane represents the reward obtained by the lane agent at time step t , γ represents the discount factor, Q is the value adopted under , and copy the parameters from the actor network every C time steps, that is, .
[0024] 2. Traffic signal control module: The traffic signal control module agent selects actions based on the observed lane density and the current lane state, and obtains rewards according to the traffic flow satisfaction of the phase; the signal control agent selects its phase action according to the observation at each time interval .
[0025] The state, action, and reward of the agent that controls the intersection signal of the traffic signal control module are as follows: State: Starting from the incoming road on the left side of the north side, calculate the lane density clockwise, and finally add the lane action to obtain the state of the intersection; at time step tThe status obtained by the signal lamp The expression is: Among them, is the lane density of the lane t entered at time step , is the action of the control lane agent at time step t .
[0026] Action: The agent selects different available phase sets according to the structure of the road network and traffic demand.
[0027] Reward: The reward mechanism is based on the utilization rate of each traffic signal phase, that is, whether the phase selected by the agent maximally meets the traffic flow required by the phase. Since one of the main goals of traffic signals is to optimize the passing efficiency of traffic flow, selecting an appropriate signal lamp phase can more effectively allocate traffic flow, reduce congestion, and ensure that vehicles on the road can pass through the intersection smoothly. Therefore, when the agent selects the signal lamp phase, it should give priority to the current traffic flow situation and try to select the phase that can maximally meet the vehicle passing demand; by considering the phase required for the maximum traffic flow; the utilization rate of the traffic signal phase is: Among them, represents the reward obtained by the agent of the signal at time step t , represents the signal lamp phase selected by the signal agent at time step t , represents the traffic flow at phase t at time step , represents the phase of the maximum traffic flow of .
[0028] 3. Prioritized Experience Replay Algorithm: In traffic signal control, some traffic scenarios may occur more frequently or have a greater impact on traffic flow optimization. By replaying the prioritized experience pool, it can ensure that these key experiences are learned and utilized more frequently, thereby improving the learning efficiency, making the training samples more diverse, reducing the correlation between samples, and helping to improve the stability and convergence speed of the model.
[0029] The experience replay pool is implemented through a data structure SumTree based on a binary tree structure. In SumTree, each node contains a priority value, and the leaf nodes store the specific samples and their priorities. In this way, the priorities of the samples can be effectively updated and adjusted to reflect their importance for model training. The implementation method of the experience replay pool is as follows: Given the priority of each sample : where, is the absolute value of the TD error of each sample, ε is a constant to ensure that samples are also drawn when the TD error is 0. The optimization formula for the loss function is as follows: where, is the Q value of the target network, is the value adopted under of, is defined as: where, is the ratio of the priority of the current sample to the total priority, is the minimum priority ratio, is an adjustable parameter, and finally update the TD error of the sample: .
[0030] 4. LCSL algorithm: The interaction method between the traffic signal control module and the lane guidance control module is as follows: First, the agent responsible for controlling the traffic signal is trained for a training period of episodes, and each episode contains T steps; Then, based on the output of the traffic signal control agent, the agent responsible for controlling the variable lane is trained for a training period of episodes; The training alternates between the two; Finally, the agent controlling the traffic signal is trained for more episodes; After training the two agents, the traffic signal control module and the lane guidance control module run simultaneously to control the traffic.
[0031] The LCSL algorithm is as follows: .
[0032] 5. Simulation test and performance comparison: The present invention conducts experiments using an open-source traffic simulator SUMO. After inputting traffic data into the simulator, vehicles navigate to their destinations according to the environmental settings. The simulator provides the status to the traffic signal control method and performs traffic signal operations according to the control method. Traditionally, there are three seconds of yellow light and two seconds of all-red time after each green light signal. The simulator provides the status to the traffic signal control module and performs traffic signal actions from the control strategy. In addition, the present invention introduces a variable-direction lane guidance control module in the simulator to control the guidance of variable lanes.
[0033] Experimental setup: Simulated and real-world datasets are used to evaluate the effectiveness and efficiency of different methods. Both the simulated and real-world datasets include off-peak and peak periods. Two real traffic flow datasets are from intersections in Lanzhou City, Gansu Province, for evaluation under real and dynamic traffic conditions; Table 1 lists the statistical information of different datasets: Table 1 Statistical data of four datasets
[0034] A detailed description of how the present invention sets or preprocesses these datasets is as follows: Simulated dataset: From 0 to 3600 seconds, each time period is 300 seconds; the traffic flow is relatively stable (the range in the off-peak period is from 132 to 160 vehicles, and the range in the peak period is from 225 to 276 vehicles); the routes from west to east and west to north are the busiest.
[0035] Real dataset: From the intersection of Wannian Road and Jianning Road in Lanzhou City, Gansu Province. The main traffic flows are from west to east and west to north. The traffic flow data is collected from 8 am to 9 am and 10 am to 11 am on weekdays; from 0 to 3600 seconds, each time period is 300 seconds.
[0036] Performance test: The LCSL method of the present invention is compared with various baseline methods and variants of LCSL. The results are shown in Table 3 and are classified from two aspects. On the one hand, considering key technologies, such as whether variable lanes are used, whether reinforcement learning (RL) methods are used, etc., it can be classified into traditional traffic signal control methods, RL-based traffic signal control methods, and RL-based variable lane control methods. On the other hand, considering the cooperation method, it can be classified into learning-based traffic signal and variable lane co-control methods.
[0037] Baseline methods: (1) FixedTime: This is the most commonly used traffic signal control method in practice, which processes traffic flow according to the cycle length and phase time in a predetermined schedule.
[0038] (2)MaxPressure: This is the most popular network-level traffic signal control method in the transportation field, which greedily selects the phase with the maximum pressure using a pressure feedback mechanism.
[0039] (3)DDQN: This is the most commonly used RL-based traffic signal control method, which selects phase actions using the maximum Q value. In addition, a target network is introduced to enhance the stability of the agent.
[0040] Variants of LCSL: (1)LCSL-1: Remove the traffic signal control module from LCSL; this variant demonstrates the improvement brought by the variable lane control module designed in the present invention.
[0041] (2)LCSL-2: Remove the variable lane control module from LCSL; this variant demonstrates the importance of the variable lane control module in dealing with the spatial distribution of uneven traffic flow.
[0042] (3)LCSL-3: Remove the agent information sharing from LCSL and modify the reward function; this variant demonstrates the improvement in enhancing the coupling and coordination between two agents. Information sharing: Remove a mechanism where one agent cannot access the action information of another agent when taking actions. Reward function: When calculating the reward function, remove the terms related to variable lane traffic flow.
[0043] Table 2 Comparison of evaluation metrics of LCSL and other models on four datasets
[0044] Performance: Table 2 shows the average delay, average queue length, and maximum queue length of LCSL and various baselines or variants in four datasets.
[0045] (1)Advantages of LCSL over baseline methods: As can be seen from Table 2, in the case of the peak traffic flow dataset, compared with FixedTime, MaxPressure reduces the average queue delay by 18.7%, the average queue length by 57.3%, and the maximum queue length by 51.5%. This indicates that MaxPressure has a stronger ability to handle dynamic traffic flow. Although MaxPressure shows a slightly higher average delay (14.1%) in some off-peak situations, overall, MaxPressure can manage traffic flow more effectively and reduce congestion. However, neither FixedTime nor MaxPressure takes into account variable direction lane control and dynamic signal control, resulting in performance differences when traffic flow increases. LCSL integrates the traffic signal control and variable direction lane control modules while retaining the information sharing mechanism among agents. Compared with FixedTime, LCSL reduces the average queue delay, average queue length, and maximum queue length by up to approximately 34.8%, 64.9%, and 63.4% respectively during peak and off-peak periods. Compared with MaxPressure, these reductions are up to approximately 45.7%, 57.1%, and 54.5% respectively.
[0046] Although the DDQN method shows some improvements, it still does not reach the performance of LCSL. Compared with DDQN, LCSL reduces the average queue delay, average queue length, and maximum queue length by up to approximately 59.4%, 76.0%, and 77.3% respectively. Figure 3 are the average delay and average queue length on the simulation dataset, Figure 3 (a) and (b) are the average delay and average queue length on the simulation dataset during off-peak periods respectively. LCSL shows the smallest fluctuation range and high stability in terms of average queue length and average delay; the average queue length fluctuates between 1 and 3 vehicles, while the average delay fluctuates between 20 and 25 seconds. Figure 3 (c) and (d) are the average delay and average queue length on the simulation dataset during peak periods respectively. LCSL shows lower and more stable average delay and queue length, with the fluctuations reduced by up to 55% and 72% respectively. The average delay and average queue length on the real dataset are as Figure 4 shown, Figure 4 (a) and (b) are the average delay and average queue length on the real dataset during off-peak periods respectively, Figure 4 (c) and (d) are the average delay and average queue length on the real dataset during peak periods respectively; in the real dataset, the trends of average delay and average queue length are similar to those of the simulation dataset, and the queue length during off-peak periods is reduced by up to 81.6%.
[0047] Compared with LCSL-2, LCSL reduces the average delay by 13.8% and the queue length by 28.8% on average in four datasets. This is because the improved lane utilization shortens the green time required to clear the traffic in this direction, so that more green time can be allocated to other directions, thereby reducing the overall average delay at the intersection. These results show that the LCSL of the present invention has advantages in two key areas: 1) using a learning-based traffic signal control strategy instead of a rule-based strategy to handle dynamic traffic flow, 2) combining it with a variable direction lane control strategy to reduce delays and queue lengths at intersections with unbalanced traffic flows. Especially under high-density traffic conditions, LCSL can quickly reduce the queue length and delay, and effectively respond to sudden changes in traffic flow.
[0048] (2) Advantages of the variable direction lane guiding control module: To improve lane utilization and ensure vehicle balance on each lane, the LCSL-1 used introduces a variable direction lane control module based on FixedTime. In real data during off-peak hours, LCSL-1 is superior to MaxPressure and FixedTime in reducing the average queue length, reducing it by 62.5% and 57.1% respectively. In Figure 3 (a) and (b), LCSL-1 and FixedTime are almost the same. In Figure 4 (a) and (b), the trends are similar, but the average reductions are 15 seconds and 1.3 vehicles respectively. In Figure 3 (c), the average reduction is 13 seconds. In Figure 3 (d), the trends are similar, but the average reduction is 5 vehicles. In Figure 4 In (c) and (d), the queue length of LCSL-1 is slightly lower than that of FixedTime (0.7%). The variable direction lane control module in LCSL-1 allows for a more flexible vehicle distribution, especially under low to medium traffic conditions, where its effect is most obvious. However, under peak traffic conditions, due to the lack of support from the signal control module, LCSL-1 is slightly insufficient in handling high-density traffic, but it is still superior to FixedTime.
[0049] (3) Advantages of the traffic signal control module To improve vehicle passing time and agent stability, the LCSL-2 used modifies the state function and reward function and introduces a prioritized experience replay mechanism based on DDQN. LCSL-2 outperforms the FixedTime and Pressure methods in four datasets. Especially during peak hours, LCSL-2 reduces the average queue length and delay by approximately 20.3% and 57.4% respectively. Compared with FixedTime, DDQN reduces the average delay and queue length by up to 11.6% and 53.4% respectively during peak hours, but increases them by 60.4% and 56.25% respectively during off-peak hours ( Figure 4 in (a) and (b)). In contrast, compared with DDQN, LCSL-2 reduces these metrics by 9.8% and 11.3% respectively during peak hours and by 59.0% and 72% respectively during off-peak hours. In addition, Figure 5 is the change of the loss function during training (refer to Figure 5 (a)), delay (refer to Figure 5 (b)), and queue length (refer to Figure 5 (c)). It can be seen that the Pr version maintains lower loss values, average delays, and queue lengths throughout the training process and has less fluctuations. LCSL-2 demonstrates better adaptability and stability in dealing with complex traffic flows.
[0050] (4) Advantages of the information sharing mechanism and reward function LCSL-3 removes the information sharing between agents and modifies the reward function. It can be seen from Table 2 that the performance of LCSL-3 is inferior to FixedTime in many cases, especially during peak hours and in complex traffic scenarios. During the peak hours of the simulated data, the average delay of LCSL-3 is 60.0 seconds, which is 8.7% higher than 55.2 seconds of FixedTime. In the real peak data, the average delay of LCSL-3 is 73.0 seconds, which is 21.9% higher than 59.9 seconds of FixedTime. In Figure 3 (c) and Figure 4 (a), (c), and (d), LCSL-3 shows higher delays and queue lengths (86.8% and 66.7% higher respectively) and greater fluctuations. Without an information sharing mechanism and dynamic adjustment ability, LCSL-3 has difficulty coping with complex and dynamic traffic flows, while the fixed cycle and phase duration of FixedTime can provide more stable performance in some cases.
[0051] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A collaborative control method for traffic signals and variable-direction lanes based on multi-agent learning, characterized in that: The coordination method is to iteratively make decisions through the traffic signal control module and the lane guidance control module, and operate simultaneously to control traffic, which is called LCSL; The lane guidance control module adopts a dual deep Q network algorithm reinforcement learning model. The agent selects actions based on traffic flow and traffic signal status to maximize lane utilization, and rewards are calculated based on the difference in lane queue lengths and the maximum direction demand traffic flow ratio; The lane guidance control module controls the state, action and reward of the agent of the variable lane at the intersection as follows: State: The state observation mainly refers to the traffic volume in the direction of the variable-guide lane selection and the traffic volume in other lanes; the state obtained by the lane at time step t The expression is as follows: in, is the traffic volume of the fixed left-turn lane on a road with variable guide lanes, is the traffic volume of the fixed straight lane on a road with variable guide lanes, is the left-turn traffic volume on the road with variable guide lanes, is the straight traffic volume on the road with variable guide lanes, is the action of the traffic signal at time step t; Action: The action of the variable guide lane is to meet the dynamic changes between straight and left turn, action space set for: in, Indicates that the lane is directed to turn left. Indicates that the lane guidance is straight ahead; Reward: The ratio of the queue length difference on a fixed lane to the maximum direction demand traffic flow As a measure, the formula is as follows: ; The dual deep Q network algorithm stabilizes the training process by introducing a target network and an experience replay pool, uses a multilayer perceptron neural network to fit the expected reward of a given state-action pair, and uses the action network and the target network to reduce the overestimation problem, selects the phase action with the maximum reward, and updates the parameters of the action network by minimizing the loss function. The traffic signal control module agent selects an action based on the observed lane density and current lane status, and obtains a reward based on the traffic flow satisfaction of the selected phase; The traffic signal control module controls the states, actions and rewards of the agent of the intersection signal as follows: State: Starting from the north left entrance road, rotate clockwise to calculate the lane density, and finally add the lane action to obtain the state of the intersection; the state of the traffic light at time step t The expression is as follows: Among them, For the time step Enter the driveway The lane density, For the time step Control the actions of lane agents; Action: The agent selects different available phase sets according to the structure of the road network and traffic demand; Reward: The reward mechanism is based on the utilization rate of each traffic signal phase. When selecting the signal light phase, the agent chooses the phase that can maximize the satisfaction of vehicle passing needs; by considering the phase required for maximum traffic flow; The interaction between the traffic signal control module and the lane guidance control module is as follows: First, the agent responsible for controlling the traffic signal is trained, and the training cycle is rounds, each round contains T steps; then, based on the output of the traffic signal control agent, the agent responsible for controlling the variable lanes is trained, and the training cycle is rounds; training alternates between the two; finally, the agent controlling the traffic signal is retrained rounds; after training the two agents, the traffic signal control module and the lane guidance control module run simultaneously to control traffic; the LCSL algorithm is as follows: in, is Next adopt The value of yes Next adopt value.
2. The method for coordinated control of traffic signals and variable direction lanes based on multi-agent learning as claimed in claim 1, characterized in that: The parameter update formula of the action network is as follows: Where T is the number of time steps, is the Q value of the target network, The calculation formula is: Among them, r lane represents the reward obtained by the lane agent at time step t, γ represents the discount factor, and Q is Next The value of is copied from the action network every C time steps, that is .
3. The method for coordinated control of traffic signals and variable direction lanes based on multi-agent learning as claimed in claim 1, characterized in that: In the reward of the agent controlling the intersection signal by the traffic signal control module, the utilization rate of the traffic signal phase is: in, The agent representing the signal at time step Rewards received, The agent representing the signal at time step Select the signal light phase, Indicates that at time step Phase Traffic flow on Indicates the phase of maximum traffic flow of stage.
4. The method for coordinated control of traffic signals and variable direction lanes based on multi-agent learning as claimed in claim 1, characterized in that: The experience replay pool is implemented by a data structure SumTree based on a binary tree structure.
5. The method for coordinated control of traffic signals and variable direction lanes based on multi-agent learning as claimed in claim 4, characterized in that: The implementation method of the experience replay pool is as follows: Prioritize each sample based on the : in, is the absolute value of the TD error for each sample, is a constant, and the loss function optimization formula is as follows: in, is the Q value of the target network, , is defined as: in, is the ratio of the current sample’s priority to the total priority, is the minimum priority ratio, is an adjustable parameter, and finally updates the TD error of the sample: 。
Citation Information
Cited By
Traffic signal optimization method based on multi-agent deep reinforcement learning
CN120599840A
Dynamic traffic guidance method based on traffic flow prediction under influence of navigation information
CN121122023A
Dynamic reversible lane control method and system based on reinforcement learning
CN121212727A
A dynamic tidal lane control method and system based on reinforcement learning
CN121212727B
Traffic signal control method and system based on multi-source event and double-loop phase cooperation
CN121214705A