Decentralized adaptive signal control method considering the impact of downstream traffic pressure transmission

By deploying DTAIC agents in the road network and using deep reinforcement learning algorithms to optimize traffic signal control, the limitations of traditional systems in dynamic environments are overcome, coordination and efficiency improvements at different intersections are achieved, and it is suitable for a variety of traffic flow scenarios.

CN119207133BActive Publication Date: 2025-09-30TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411340275.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-09-30
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Traditional traffic signal control systems have limitations in dynamic traffic environments, especially the difficulty in effectively coordinating intersections with different intelligence levels. In addition, the optimization complexity increases exponentially with the increase of network nodes, making analysis difficult.

Method used

A decentralized adaptive signal control method is adopted, using deep reinforcement learning algorithms and proximal policy optimization algorithms. By deploying a small number of intelligent agents in the road network, a traffic-aware intersection controller (DTAIC) based on a "one-to-many" topology model is constructed to optimize traffic signal control in real time, taking into account the impact of downstream traffic pressure transmission.

Benefits of technology

It achieves effective coordination between low-level and high-level intersections, improves road network traffic efficiency, reduces queue lengths, is suitable for intersections with different traffic flows, and has efficient real-time data processing capabilities and wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207133B_ABST
    Figure CN119207133B_ABST
Patent Text Reader

Abstract

The present invention relates to a decentralized adaptive signal control method that considers the impact of downstream traffic pressure transmission. The method comprises the following steps: estimating the number of queued vehicles at the entrance to a downstream intersection to obtain real-time information on downstream traffic conditions; constructing a signal control agent that considers the traffic conditions of the agent and multiple downstream intersection entrances, incorporating the number of queued vehicles into both the agent's input state and reward function, and encompassing all possible signal phase combinations in the action set; introducing a correction coefficient to adjust the agent's control scheme selection strategy, and constructing a deep reinforcement learning framework based on a proximal policy optimization algorithm. Through a decentralized training method, each agent continuously optimizes its control strategy; and utilizing the trained agents to output the optimal control scheme for adaptively controlling intersection traffic lights. Compared to existing technologies, the present invention can achieve coordinated control optimization of traffic across an entire road network while deploying only a small number of agents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic signal control, and in particular to a decentralized adaptive signal control method that takes into account the influence of downstream traffic pressure transmission. Background Art

[0002] Signal control systems in traffic road networks play a vital role in coordinating vehicle flows within urban areas. Traditional signal control networks face several challenges. First, existing traffic models often focus on specific aspects and usually only simulate some typical network characteristics. For example, traffic flow theory, which ignores the driving behavior of individual vehicles, and queuing theory, fail to fully address the spatiotemporal dynamics in road networks. Second, as the number of network nodes increases, the optimization complexity based on traditional signal control models increases exponentially. Third, the increasing complexity of these problems poses significant challenges to most analytical optimization algorithms, including problems such as non-convexity, high dimensionality, and nonlinearity.

[0003] The continuous evolution of innovative approaches is enhancing the comprehensiveness and precision of traffic management and monitoring capabilities. Technologies such as holographic intersections and integrated perception systems leverage advanced sensor technology and real-time data processing to capture the immediate traffic conditions at intersections. The emergence of new intersection control technologies has opened up new prospects for traffic management, such as multimodal traffic management and regional traffic optimization. Furthermore, advances in perception technology and improvements in big data processing capabilities are facilitating the integration of artificial intelligence into traffic signal control systems.

[0004] Due to the varying capabilities of traffic perception and control equipment at different intersections, some will inevitably be upgraded to intelligent systems early, while others will continue to utilize traditional control methods. This means that a road network may contain intersections with varying levels of intelligence, such as a combination of traditional rule-based traffic lights and those driven by deep reinforcement learning algorithms. Therefore, achieving compatibility between intelligent and non-intelligent intersections within a limited road network is crucial. Summary of the Invention

[0005] The purpose of this invention is to provide a decentralized adaptive signal control method that takes into account the influence of downstream traffic pressure transmission. It is based on a deep reinforcement learning algorithm based on the Proximal Policy Optimization (PPO) algorithm. Under the decentralized control framework, based on the "one (entity) to many (downstream intersections)" topology model, by deploying a small number of intelligent agents in the road network, it successfully achieves coordinated control optimization of the entire road network traffic.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] The Autonomous Traffic System project is dedicated to optimizing urban traffic flow and traffic signal control through intelligent and adaptive means to address the limitations of traditional traffic signal control systems in dynamic traffic environments. The core goal of this project is to achieve autonomous optimization of traffic signal control, thereby improving the overall efficiency of the traffic system. Based on the Autonomous Traffic System project, the present invention relates to a decentralized adaptive signal control method that considers the influence of downstream traffic pressure transmission, referred to as the Downstream Traffic Aware Intersection Controller (DTAIC), which includes the following steps:

[0008] S1, traffic status acquisition: by calculating the number of vehicles flowing out of the main intersection and the downstream intersection, estimating the number of vehicles queuing at the entrance of the downstream intersection, and obtaining real-time information on the downstream traffic status;

[0009] S2, Agent Construction: Build a signal control agent that considers the traffic status of the agent and multiple downstream intersection entrances. The number of queued vehicles is incorporated into the agent's input state and reward function, and all possible signal phase combinations are included in the action set to achieve more accurate control decisions.

[0010] S3, Agent Training: Introducing correction coefficients to adjust the agent's control scheme selection strategy, dynamically adjusting the weight of the agent's influence and downstream traffic pressure, and constructing a deep reinforcement learning framework based on a proximal policy optimization algorithm. Through a distributed training method, each agent continuously optimizes its control strategy, achieving multi-agent convergence and optimization.

[0011] S4 uses the trained agent to output the optimal control plan and adaptively control the intersection traffic lights.

[0012] Each intersection consists of four entrances, each of which includes two roads: an entrance road and an exit road. Each road consists of one or more one-way lanes. The entrance lane of each intersection that forms an entrance road is represented by m. k , forming the leaving lane of the road is represented by n k , where m k and n k Maintain a one-to-one correspondence between them.

[0013] The agent takes lane occupancy, queue length measured by the number of queued vehicles, and the current phase state of the lane entering the intersection as input for the initial state.

[0014] The input volume for each lane is considered as a vector containing three values ​​as follows:

[0015]

[0016] Among them, at a certain time step t, occ t and q t They represent the occupancy and queue length of lane i, g t is the phase state, which is a binary value indicating whether lane i is in the released state at the current time step t.

[0017] The agent's field of view is expanded so that it can access the data of the R lanes entering the main body and the R lanes entering the four downstream roads. The state of the agent is represented as follows:

[0018]

[0019] in, represents the three-valued input vector of the lane entering the main intersection at time t, A three-valued input vector representing the incoming lanes of the downstream intersection at time t, where the number of lanes R is obtained by multiplying the number of lanes of each road by 4.

[0020] The agent's actions consist of all possible combinations of different directional traffic flows. Any two non-conflicting traffic flows constitute a phase. A set of feasible actions is constructed based on all phases. Each time a decision is made, the agent selects a phase from the set of feasible actions.

[0021] If the next phase selected by the agent is the same as the previous phase, the green light of the phase is directly executed for the first preset duration;

[0022] If the next phase selected by the agent is different from the previous phase, it means that the phase needs to be changed. First, the yellow light of the previous phase is executed for the second preset time. After the yellow light is executed, the green light of the next phase is executed for the third preset time, and the sum of the second preset time and the third preset time is equal to the first preset time.

[0023] The agent's reward includes the main body's intersection queue length and the downstream lane's queue length corrected by the correction coefficient.

[0024] The correction factor is:

[0025]

[0026] in, is the correction coefficient of the exit road j of the main intersection. The exit road j has multiple exit lanes, which are represented by {n1,...n k ,...} indicates that the corresponding entry lane is {m1,...,m k ,...}, K jis a set of exit lane indices belonging to exit road j, and represent lane n in the most recent simulation step k and m k The number of vehicles on.

[0027] The reward of agent i is expressed as:

[0028]

[0029] in, The time step t is the time when the main intersection enters lane m k The queue length, Enter lane n at the downstream intersection at time step t k The queue length.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] (1) Technological innovation: Currently, the cooperation mechanism of multi-agent collaboration is mostly based on the environment of the same type of agents. The one-to-many intelligent control mode proposed in this invention can effectively coordinate the cooperation between low-level intersections and high-level intersections in the road network. Simulation experiments have verified the effectiveness of the proposed algorithm.

[0032] (2) Efficient optimization: Adaptive signal control based on deep reinforcement learning can take advantage of its real-time big data processing capabilities and high accuracy. In a decentralized control framework, based on the "one (entity) to many (downstream intersections)" topology model, coordinated control optimization of the entire road network can be achieved with only a small number of agents deployed in the road network.

[0033] (3) Strong scalability: The adaptive signal control system based on deep reinforcement learning can be flexibly deployed in the road network. Intersections equipped with the corresponding hardware can perform signal control based on the control mode proposed by this invention. Each trained DTAIC intersection obtains data based on the integrated radar and visual environment, making a comprehensive assessment of the congestion status and downstream pressure, and quickly providing a signal control solution.

[0034] (4) Wide applicability: DTAIC intersections have no installation limitations and do not require specific traffic conditions at the intersection. DTAIC intersections select phases based on queue lengths transmitted by sensors, making it suitable for intersection signal control simulation scenarios with varying traffic flows. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the control process of the present invention;

[0036] Figure 2is a schematic diagram of elements of a typical intersection in one embodiment;

[0037] Figure 3 A schematic diagram of an action set of an intelligent agent in one embodiment;

[0038] Figure 4 A traffic network designed for a simulation experiment in one embodiment and the number of each intersection;

[0039] Figure 5 The changes in the queue lengths of all entrances to the road network when different numbers of DTAIC intersections are deployed in the road network;

[0040] Figure 6 A graph showing the change in queue length (unit: meter) over time at a fixed-timing intersection and a DTAIC intersection in a simulation verification in one embodiment;

[0041] Figure 7 Graph showing the change in queue length (unit: meter) over time at a fixed-timing intersection and an intersection with three DTAICs in a simulation verification in one embodiment. DETAILED DESCRIPTION

[0042] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0043] This embodiment provides a decentralized adaptive signal control method that considers the influence of downstream traffic pressure transmission, referred to as Downstream Traffic Aware Intersection Controller (DTAIC). Figure 1 As shown, the method includes the following steps:

[0044] S1, traffic status acquisition: by calculating the number of vehicles flowing out of the main intersection and the downstream intersection, estimating the number of vehicles queuing at the entrance of the downstream intersection, and obtaining real-time information on the downstream traffic status.

[0045] The typical intersection used in this embodiment is as follows Figure 2 As shown, each intersection consists of four entrances, each entrance includes two roads: an entrance road and an exit road, each road consists of three one-way lanes, and the entrance lane of each intersection forming the entrance road is represented by m k , forming the leaving lane of the road is represented by n k , where m k and n k Maintain a one-to-one correspondence between them.

[0046] S2, Agent Construction: Construct a signal control agent that considers the traffic status of the main body and multiple downstream intersection entrances, incorporates the number of queued vehicles into the agent's input state and reward function, and covers all possible signal phase combinations in the action set to achieve more accurate control decisions.

[0047] S21, the state of the agent

[0048] The agent takes the lane occupancy, queue length measured by the number of queued vehicles, and the current phase state of the lane entering the intersection as input for the initial state.

[0049] The input volume for each lane is considered as a vector containing three values ​​as follows:

[0050]

[0051] Among them, at a certain time step t, occ t and q t They represent the occupancy and queue length of lane i, g t is the phase state, which is a binary value indicating whether lane i is in the released state at the current time step t.

[0052] The agent's field of view is expanded so that it can access the data of the 12 entry lanes of the main body and the 12 entry lanes corresponding to the four downstream roads. The state of the agent is represented as follows:

[0053]

[0054] in, represents the three-valued input vector of the lane entering the main intersection at time t, A three-valued input vector representing the incoming lanes of the downstream intersection at time t.

[0055] S22, the agent’s actions

[0056] For each DTAIC agent, the main action consists of all possible combinations of different flow directions. Figure 3 As shown, any two non-conflicting traffic flows constitute a phase. There are eight possible action choices among the eight flows at an intersection, so a set of feasible actions can be constructed based on all phases. Specifically, all feasible action sets for each intersection are predefined, and each time the agent makes a decision, it selects a phase from the set.

[0057] If the next phase selected by the agent is the same as the previous phase, the green light for that phase is directly executed for Δt seconds.

[0058] If the next phase selected by the agent is different from the previous phase, it means that the phase needs to be changed. First, the yellow light of the previous phase is executed for Δt. y seconds, after the yellow light is executed, the green light of the next phase will last for Δt g seconds, and Δt y +Δt g =Δt.

[0059] In this embodiment, the attached Figure 3 The phase combination scheme presented in

[15] serves as an action set to further explore the potential of reinforcement learning agents to control and manage intersections.

[0060] S23, agent rewards

[0061] A well-designed reward can help the agent perceive surrounding traffic conditions and guide it to choose an appropriate downstream intersection to release traffic flow. This paper designs a method based on lane connectivity to modify the reward, allowing the agent to prioritize downstream pressure and compare congestion levels, thereby helping it choose the appropriate phase.

[0062] The original reward of agent i is:

[0063]

[0064] in, The time step t is the time when the main intersection enters lane m k The queue length (measured in number of vehicles).

[0065] To coordinate the cooperation between agents, the queue length of the downstream lane is included in the reward and adjusted by a coefficient. The formula for the correction coefficient is:

[0066]

[0067] in, is the correction coefficient of the exit road j of the main intersection. The exit road j has multiple exit lanes, which are represented by {n1,...n k ,...} indicates that the corresponding entry lane is {m1,...,m k ,...}, K j is a set of exit lane indices belonging to exit road j, and represent lane n in the most recent simulation step k and m k For example, in Figure 2In the figure, the exit road marked with a dotted box is exit road 2, and its corresponding exit lanes are n4, n5, and n6, so the index set is K2 = {4, 5, 6}. The entry lanes m4, m5, and m6 are associated with these three exit lanes, respectively. Therefore, the correction coefficient expression for exit road 2 is

[0068]

[0069] The modified reward for DTAIC agent i is obtained by multiplying the queue length of each downstream lane by its corresponding correction factor and then adding the largest resulting value to the reward:

[0070]

[0071] in, The time step t is the time when the main intersection enters lane m k The queue length, Enter lane n at the downstream intersection at time step t k The queue length.

[0072] S3, Agent Training: Introducing correction coefficients to adjust the agent's control scheme selection strategy, dynamically adjusting the influence weights of the agent and downstream traffic pressure, and constructing a deep reinforcement learning framework based on the proximal policy optimization algorithm. Through distributed training methods, each agent continuously optimizes the control strategy to achieve multi-agent convergence and optimization.

[0073] S4 uses the trained agent to output the optimal control plan and adaptively control the intersection traffic lights.

[0074] "One (entity) to many (downstream intersections)" topology model. The Downstream Traffic Aware Intersection Control (DTAIC) mechanism is a fully decentralized multi-agent control method based on the PPO algorithm framework. The DTAIC agent observes upstream characteristics and combines them with downstream characteristics, uses the connection relationship between the lanes of the intersection, and judges the difference in actual traffic flow in and out of the lanes, thereby adjusting the reward of the DTAIC intersection. This correction guides the DTAIC intersection to evaluate the downstream pressure, thereby selecting a downstream intersection with less congestion to release traffic flow, so as to balance the traffic flow in the upstream and downstream directions, aiming to reduce the queue length at the entity intersection and surrounding intersections. In a decentralized control framework, based on the "one to many" topology model, coordinated control optimization of the entire road network traffic is achieved under the condition of deploying only a small number of agents in the road network.

[0075] Figure 4 The 3×3 traffic network and intersection numbers designed for the simulation experiment of this embodiment are shown.

[0076] Figure 5The paper shows how the total queue length (number of vehicles) at all intersection entrances in the entire road network changes as the number of DTAIC agents increases. Compared to the baseline approach, the total queue length in networks with one, two, and three DTAIC agents decreases by 22.4%, 25.9%, and 30.8%, respectively. This demonstrates the effectiveness of our approach in improving overall network efficiency.

[0077] Figure 6 The figure shows the changes in the queue lengths (meters) at all intersections when a DTAIC agent is only located at intersection E. As can be seen from the figure, the DTAIC agent significantly reduces the queue length at intersection E and has an indirect impact on other intersections. Specifically, compared with the fixed timing (FT) method, the average queue length at intersection E is reduced by about 55.9%, while the queue lengths at the surrounding fixed timing intersections are also reduced by an average of 21.4%. This further highlights the effectiveness of the designed cooperation mechanism between the DTAIC agent and the FT agent. In addition, Figure 7 As shown, the average queue length at all intersections is reduced by approximately 33.8% when DTAIC is located at intersections B, D, and I. Notably, among all FT-controlled intersections, intersection E shows the most significant performance improvement, with a reduction of 33.2%.

[0078] This paper describes a decentralized adaptive signal control method, called the Downstream Traffic-Aware Intersection Controller (DTAIC) collaboration mechanism, that considers the impact of downstream traffic pressure transmission. This mechanism adjusts the rewards at DTAIC intersections by calculating traffic flow differences between inbound and outbound lanes based on lane-to-lane connectivity. This adjusted reward helps DTAIC intersections assess downstream pressure and promote strategic phase selection to balance traffic flow in both upstream and downstream directions. Simulation results demonstrate the effectiveness of the DTAIC approach.

[0079] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A decentralized adaptive signal control method considering the influence of downstream traffic pressure transmission, characterized in that: By deploying a small number of intelligent agents in the road network, the coordinated control and optimization of the entire road network traffic is achieved. The method includes the following steps: S1, traffic status acquisition: By calculating the number of vehicles flowing out of the main intersection and the downstream intersection, the number of vehicles queuing at the entrance of the downstream intersection is estimated, and the real-time information of the downstream traffic status is obtained; wherein, each intersection consists of four entrances, each entrance includes two roads: an entrance road and an exit road, each road consists of one or more one-way lanes, and the entrance lane of each intersection forming the entrance road is represented as , forming the leaving lane of the road is expressed as ,in, and Maintain a one-to-one correspondence between them; S2, Agent Construction: Build a signal control agent that considers the traffic status of the main body and multiple downstream intersection entrance lanes. The number of queued vehicles is incorporated into the agent's input state and reward function, and all possible signal phase combinations are included in the action set to achieve more accurate control decisions. The agent uses lane occupancy, queue length measured by the number of queued vehicles, and the current phase state of the main body's intersection entrance lane as initial state input. The input for each lane is considered a vector containing three values, as shown below: Among them, at a certain time step , and Entering the lane Occupancy and queue length, is the phase state, which is a binary value indicating the lane Is it at the current time step? Whether it is in the release status; Expand the agent's field of view so that it can access the ontology R The entry lanes and the four downstream roads correspond to R The state of the agent is represented as follows: in, Representing an agent exist t The three-value input vector of the lane entering the intersection at time t, Representing an agent exist t The three-valued input vector of the incoming lanes of the downstream intersection at time , where the number of lanes R It is obtained by multiplying the number of lanes of each road by 4; S3, Agent Training: Introducing correction coefficients to adjust the agent's control scheme selection strategy, dynamically adjusting the weight of the agent's influence and downstream traffic pressure, and constructing a deep reinforcement learning framework based on a proximal policy optimization algorithm. Through a distributed training method, each agent continuously optimizes its control strategy, achieving multi-agent convergence and optimization. S4 uses the trained agent to output the optimal control plan and adaptively control the intersection traffic lights.

2. The decentralized adaptive signal control method considering the influence of downstream traffic pressure transmission according to claim 1 is characterized in that: The agent's actions consist of all possible combinations of different directional traffic flows. Any two non-conflicting traffic flows constitute a phase. A set of feasible actions is constructed based on all phases. Each time a decision is made, the agent selects a phase from the set of feasible actions.

3. The decentralized adaptive signal control method considering the influence of downstream traffic pressure transmission according to claim 2 is characterized in that: If the next phase selected by the agent is the same as the previous phase, the green light of the phase is directly executed for the first preset duration; If the next phase selected by the agent is different from the previous phase, it means that the phase needs to be changed. First, the yellow light of the previous phase is executed for the second preset time. After the yellow light is executed, the green light of the next phase is executed for the third preset time, and the sum of the second preset time and the third preset time is equal to the first preset time.

4. The decentralized adaptive signal control method considering the influence of downstream traffic pressure transmission according to claim 1 is characterized in that: The agent's reward includes the main body's intersection queue length and the downstream lane's queue length corrected by the correction coefficient.

5. The decentralized adaptive signal control method considering the influence of downstream traffic pressure transmission according to claim 4 is characterized in that: The correction factor is: in, It is the exit road of the main intersection Correction factor for exit road There are multiple exit lanes, each with Indicates that the corresponding entry lane is , It indicates an exit road The collection of exit lane indices, and Represents the lanes in the most recent simulation step and The number of vehicles on.

6. The decentralized adaptive signal control method considering the influence of downstream traffic pressure transmission according to claim 5 is characterized in that: The intelligent agent The reward is expressed as: in, is the time step Entering the lane at the main intersection The queue length, is the time step Downstream intersection entry lane The queue length.

Citation Information

Patent Citations

  • Reinforcement learning area signal control method based on vehicle planning path

    CN113487902A