Urban traffic signal reinforcement learning control method considering phase switching and duration adjustment

By introducing two deep neural subnets into urban traffic signal control, it is used to estimate the value of different phases and phase durations and dynamically adjust the phase and duration of traffic signals, the existing methods are solved in the problem of insufficient flexibility in dealing with real-time traffic changes, and more efficient traffic flow scheduling and traffic signal control are achieved.

CN119992828AInactive Publication Date: 2025-05-13TAIZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115435.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing urban traffic signal control methods lack flexibility in dealing with real-time changing traffic, resulting in frequent air traffic in individual directions, and it is difficult to effectively improve the rapid traffic capacity of vehicles at controlled intersections and alleviate traffic congestion.

Method used

In the DQN reinforcement learning framework, two deep neural subnets are set up for the policy network, which are used to estimate the value of different phases and different phase durations respectively. Through shared experience and joint optimization, the traffic signal phase and phase duration are dynamically adjusted to optimize traffic scheduling.

Benefits of technology

By combining real-time collection of vehicle flow information and deep reinforcement learning, more flexible and efficient traffic signal control is achieved, the traffic flow efficiency at controlled intersections is optimized, and traffic congestion and vehicle waiting time is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992828A_ABST
    Figure CN119992828A_ABST
Patent Text Reader

Abstract

The invention provides an urban traffic signal reinforcement learning control method considering phase switching and duration adjustment, and the method comprises the steps: taking a traffic signal control system of an urban road intersection as a deep reinforcement learning agent, and collecting the traffic flow information of a controlled intersection in real time through a sensor; each time at the end moment of the current phase of the traffic signal, the intelligent agent calculates a state and an award based on the collected traffic flow information, trains a strategy network containing two deep neural sub-networks based on a DQN deep reinforcement learning framework, and then calculates a new action based on the state and the strategy network, so as to obtain a new action; adjusting the traffic signal to a new phase or prolonging the current phase duration; and after multiple times of training, an optimized traffic signal control strategy network is obtained. The traffic signal phase and the phase duration are dynamically adjusted based on the real-time traffic flow data and deep reinforcement learning, optimization of controlled intersection traffic flow scheduling is facilitated, and the method is suitable for different intersections and traffic flow change scenes and has wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep reinforcement learning and urban traffic signal control, and in particular to a method for reinforcing learning control of urban traffic signals taking into account phase switching and duration adjustment. Background Art

[0002] With the advancement of urbanization, traffic congestion has become an important issue facing digital city governance. Especially in the context of the increasing number of cars year by year, traffic congestion has become more serious. Improving traffic signal control strategies and enhancing their intelligence level are effective ways to alleviate congestion.

[0003] Adaptive traffic signal control can dynamically adjust the signal phase sequence or phase duration according to the real-time traffic flow, thereby optimizing the traffic flow at the controlled intersection. The phase sequence refers to the sequence of traffic lights in each direction of traffic at the controlled intersection, and the phase duration refers to the duration of traffic lights in each direction of traffic at the controlled intersection. Traditional traffic signal control mostly relies on preset phase sequences and phase durations, and lacks the ability to respond quickly to real-time changes in traffic flow. With the development of technology, optimized traffic signal control technology has begun to be gradually applied. These systems can dynamically adjust the phase or phase duration of traffic signals according to the collected road traffic flow. However, most of the current related landing application technologies remain at the shallow optimization level, and only preset different traffic signal phases and phase durations for different time periods based on the statistics of historical traffic flow data. Although it has improved the performance of traffic control to a certain extent, there is still much room for improvement in real-time and intelligence.

[0004] In recent years, deep reinforcement learning technology has begun to be introduced into the research of urban traffic signal control, and has shown certain advantages in improving the traffic capacity of controlled intersections. Most of the current common traditional and reinforcement learning urban traffic signal control methods use a fixed phase sequence, and only control the traffic flow at controlled intersections by adjusting the phase duration of the signal. This has brought certain limitations to the improvement of urban traffic signal control performance. For example, when the traffic flow in all directions of the controlled intersection is unevenly distributed, it may lead to frequent empty spaces in some directions. If a more flexible control strategy can be considered in the urban traffic signal control method through phase switching and duration adjustment, it can promote the further improvement of urban traffic signal control performance, which will have a positive effect on improving the rapid traffic capacity of vehicles at controlled intersections and alleviating traffic congestion. Summary of the invention

[0005] In the above background, the present invention proposes a reinforcement learning control method for urban traffic signals considering phase switching and duration adjustment. This method sets two deep neural sub-networks for the strategy network in the DQN reinforcement learning framework, which are used to estimate the value of deploying different phases and deploying different phase durations, respectively. The two sub-networks improve the learning efficiency and quality of the urban traffic signal reinforcement learning control strategy by sharing experience and joint optimization, and obtain a signal control strategy with good traffic control performance.

[0006] The specific content of the present invention is as follows:

[0007] The present invention uses the traffic signal control system of the road intersection as the intelligent agent of deep reinforcement learning. The intelligent agent collects the traffic information of each direction of the controlled intersection in real time through sensors, including the length of the fleet of each controlled lane, the number of stopped vehicles and the waiting time of vehicles; each time the predetermined duration of the current phase of the traffic signal expires, the intelligent agent calculates the state s of the current environment of the intelligent agent and the reward r of the currently completed execution action based on the collected traffic information, and trains the policy network used for action decision-making based on the DQN deep reinforcement learning framework, and then calculates a new action a based on the state s and the policy network, corresponding to adjusting the traffic signal to a new phase or extending the duration of the current phase; after multiple trainings, an optimized policy network is finally obtained, and the intelligent agent can make the best action decision based on the policy network and the real-time traffic state information of the controlled intersection, when the predetermined duration of the current phase of the traffic signal expires, and adjust the traffic signal to a suitable phase or extend the appropriate action time for the current phase.

[0008] The state s includes the information of the convoy length set QL, the vehicle number set VN, the vehicle waiting time set WT, the current phase p0 and the duration d0 of the current phase, etc. of each controlled lane at the road intersection, where QL = {QL l ,l∈L},VN={VN l ,l∈L},WT={WT l ,l∈L}, L is the set of controlled lanes at the road intersection, QL l is the length of the convoy on lane l at the latest moment, VN l is the number of vehicles on lane l that experienced a speed lower than 0.1 m / s during the current phase execution period, WT l It is the sum of the duration that all vehicles on lane l experienced a speed lower than 0.1 m / s during the current phase execution period.

[0009] The action a refers to the agent adjusting the traffic light signal to a given phase p1, and the new phase lasts for a given duration d1, in seconds, where p1 is the phase given by the agent's decision, and d1 is the duration of the phase given by the agent's decision; if p1 is different from the current phase p0 of the traffic signal, then when adjusting the phase from p0 to p1, an additional 3 seconds of yellow light time is executed, and then d1 seconds of p1 phase time is executed; if p1 is the same as p0, the p0 phase is extended by d1 seconds without inserting an additional yellow light time.

[0010] The reward r is used to evaluate the performance of the agent in executing action a, and is defined as r = -∑ l∈L (QL l +WT l ), where L is the set of controlled lanes at the road intersection, QL l is the length of the convoy on lane l at the latest moment, WT l It is the sum of the duration that all vehicles on lane l experienced a speed lower than 0.1 m / s during the current phase execution period.

[0011] The strategy network comprises two deep neural sub-networks Q1(s,p1) and Q2(s,d1), which are used to estimate the value of deploying different phases and the value of deploying different phase durations, respectively. Q1(s,p1) is the potential performance value of executing the traffic signal phase p1 for the controlled intersection traffic control estimated by the agent in state s, and Q2(s,d1) is the potential performance value of executing the traffic signal phase for d1 seconds for the controlled intersection traffic control estimated by the agent in state s.

[0012] The beneficial effects of the present invention are as follows: the present invention collects the traffic status information of the controlled intersection in real time through sensors, and by combining phase switching and phase duration adjustment, uses two deep neural sub-networks in the strategy network, which are respectively used to estimate the value of deploying different phases and the value of deploying different phase durations. The two sub-networks improve the learning efficiency and quality of the control strategy in the reinforcement learning control of urban traffic signals by sharing experience and joint optimization; the present invention dynamically adjusts the phase and phase duration of traffic signals based on real-time traffic data and deep reinforcement learning, which helps to optimize the traffic scheduling performance of controlled intersections, reduce traffic congestion and vehicle waiting time, and improve traffic efficiency. It is suitable for different intersection structures and traffic flow change scenarios, and has good compatibility and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flow chart of the present invention;

[0014] Figure 2 A schematic diagram of a traffic signal phase example for illustrating an embodiment;

[0015] Figure 3 Schematic diagram of the strategy network of the present invention. DETAILED DESCRIPTION

[0016] The following are specific embodiments of the present invention and combined with the accompanying drawings to further describe the technical solution of the present invention, but the present invention is not limited to the following embodiments.

[0017] like Figure 1 As shown, the present invention uses the traffic signal control system of the road intersection as the intelligent agent of deep reinforcement learning. The intelligent agent collects the traffic information of each direction of the controlled intersection in real time through sensors, including the length of the fleet of each controlled lane, the number of stopped vehicles and the waiting time of vehicles; each time the predetermined duration of the current phase of the traffic signal expires, the intelligent agent calculates the state s of the current environment of the intelligent agent and the reward r of the currently completed execution action based on the collected traffic information, and trains the policy network used for action decision-making based on the DQN deep reinforcement learning framework, and then calculates a new action a based on the state s and the policy network, corresponding to adjusting the traffic signal to a new phase or extending the duration of the current phase; after multiple trainings, an optimized policy network is finally obtained, and the intelligent agent can make the best action decision based on the policy network and the real-time traffic state information of the controlled intersection, when the predetermined duration of the current phase of the traffic signal expires, and adjust the traffic signal to a suitable phase or extend the suitable action time for the current phase.

[0018] The state s includes the information of the convoy length set QL, the vehicle number set VN, the vehicle waiting time set WT, the current phase p0 and the duration d0 of the current phase, etc. of each controlled lane at the road intersection, where QL = {QL l ,l∈L},VN={VN l ,l∈L},WT={WT l ,l∈L}, L is the set of controlled lanes at the road intersection, QL l is the length of the convoy on lane l at the latest moment, VN l is the number of vehicles on lane l that experienced a speed lower than 0.1 m / s during the current phase execution period, WT l It is the sum of the duration that all vehicles on lane l experienced a speed lower than 0.1 m / s during the current phase execution period.

[0019] The action a refers to the agent adjusting the traffic light signal to a given phase p1, and the new phase lasts for a given duration d1, in seconds, where p1 is the phase given by the agent's decision, and d1 is the duration of the phase given by the agent's decision. Figure 2Taking the crossroad scenario shown as an example, the options for p1 include four phases: east-west straight-up-right turn, east-west left turn, north-south straight-up-right turn, and north-south left turn. If the actual demand limits the traffic signal phase duration range to 5 to 60 seconds, the options for d1 include 5 seconds, 6 seconds, ..., 59 seconds, 60 seconds, a total of 56 durations; if p1 is different from the current phase p0 of the traffic signal, then when adjusting the phase from p0 to p1, an additional yellow light time of 3 seconds is executed, and then the p1 phase time of d1 seconds is executed; if p1 is the same as p0, the p0 phase is extended by d1 seconds without inserting an additional yellow light time.

[0020] The reward r is used to evaluate the performance of the agent in executing action a, and is defined as r = -∑ l∈L (QL l +WT l ), where L is the set of controlled lanes at the road intersection, QL l is the length of the convoy on lane l at the latest moment, WT l It is the sum of the duration that all vehicles on lane l experienced a speed lower than 0.1 m / s during the current phase execution period.

[0021] like Figure 3 As shown, the strategy network includes two deep neural sub-networks Q1(s,p1) and Q2(s,d1), which are respectively used to estimate the value of deploying different phases and the value of deploying different phase durations, wherein Q1(s,p1) is the potential performance value of executing the traffic signal phase p1 for the controlled intersection traffic control estimated by the agent in state s, and Q2(s,d1) is the potential performance value of executing the traffic signal phase for d1 seconds for the controlled intersection traffic control estimated by the agent in state s.

[0022] Compared with the existing urban traffic signal control method that relies on preset phase sequence and phase duration, the present invention can better adapt to different road intersection structures and traffic flow change scenarios, help optimize traffic scheduling at controlled intersections, reduce traffic congestion and vehicle waiting time, and improve traffic efficiency. It has good compatibility and broad application prospects.

[0023] The above is a specific implementation case of the present invention, which is used to make the present invention clearer. Any modification or equivalent replacement of the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A method for reinforcing learning control of urban traffic signals considering phase switching and duration adjustment, characterized in that: The traffic signal control system at a road intersection is used as an intelligent agent for deep reinforcement learning. The intelligent agent collects traffic information in each direction of the controlled intersection through sensors in real time, including the length of the fleet of each controlled lane, the number of stopped vehicles and the waiting time of vehicles. Each time the scheduled duration of the current phase of the traffic signal expires, the intelligent agent calculates the state s of the current environment of the intelligent agent and the reward r of the currently completed action based on the collected traffic information, and trains the policy network used for action decision-making based on the DQN deep reinforcement learning framework. Then, a new action a is calculated based on the state s and the policy network, corresponding to adjusting the traffic signal to a new phase or extending the duration of the current phase. After multiple trainings, an optimized policy network is finally obtained. The intelligent agent can make the best action decision based on the policy network and the real-time traffic status information of the controlled intersection when the scheduled duration of the current phase of the traffic signal expires, and adjust the traffic signal to a suitable phase or extend the appropriate action time for the current phase.

2. The urban traffic signal reinforcement learning control method considering phase switching and duration adjustment according to claim 1 is characterized in that: The state s includes the information of the convoy length set QL, the vehicle number set VN, the vehicle waiting time set WT, the current phase p0 and the duration d0 of the current phase, etc. of each controlled lane at the road intersection, where QL = {QL l ,l∈L},VN={VN l ,l∈L},WT={WT l ,l∈L}, L is the set of controlled lanes at the road intersection, QL l is the length of the convoy on lane l at the latest moment, VN l is the number of vehicles on lane l that experienced a speed lower than 0.1 m / s during the current phase execution period, WT l It is the sum of the duration that all vehicles on lane l experienced a speed lower than 0.1 m / s during the current phase execution period.

3. The urban traffic signal reinforcement learning control method considering phase switching and duration adjustment according to claim 1 is characterized in that: The action a refers to the agent adjusting the traffic light signal to a given phase p1, and the new phase lasts for a given duration d1, in seconds, where p1 is the phase given by the agent's decision, and d1 is the duration of the phase given by the agent's decision; if p1 is different from the current phase p0 of the traffic signal, then when adjusting the phase from p0 to p1, an additional 3 seconds of yellow light time is executed, and then d1 seconds of p1 phase time is executed; if p1 is the same as p0, the p0 phase is extended by d1 seconds without inserting an additional yellow light time.

4. The urban traffic signal reinforcement learning control method considering phase switching and duration adjustment according to claim 1 is characterized in that: The reward r is used to evaluate the performance of the agent in executing action a, and is defined as r = -∑ l∈L (QL l +WT l ), where L is the set of controlled lanes at the road intersection, QL l is the length of the convoy on lane l at the latest moment, WT l It is the sum of the duration that all vehicles on lane l experienced a speed lower than 0.1 m / s during the current phase execution period.

5. The urban traffic signal reinforcement learning control method considering phase switching and duration adjustment according to claim 1 is characterized in that: The strategy network comprises two deep neural sub-networks Q1(s,p1) and Q2(s,d1), which are used to estimate the value of deploying different phases and the value of deploying different phase durations, respectively. Q1(s,p1) is the potential performance value of executing the traffic signal phase p1 for the controlled intersection traffic control estimated by the agent in state s, and Q2(s,d1) is the potential performance value of executing the traffic signal phase for d1 seconds for the controlled intersection traffic control estimated by the agent in state s.