Method and system for adaptive cycle-level traffic signal control
By using a reinforcement learning model with a proximal policy optimization algorithm, the challenge of generating stage durations in a continuous action space for an adaptive traffic signal controller is solved, achieving more efficient and flexible traffic signal control and improving the decision-making optimization capability of the traffic signal cycle.
Patent Information
- Application Number
- CN202180063356.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-21
- Filing Date
- 2021-09-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2041-09-18
AI Technical Summary
Existing adaptive traffic signal controllers have problems with rough discretization or insufficient flexibility when generating the phase duration of traffic signal cycles in a continuous action space, which affects the performance and real-time response capability of the controller.
A reinforcement learning model using the Proximal Policy Optimization (PPO) algorithm generates the phase durations of traffic signal cycles in a continuous action space and uses an actor-critic model for training and adjustment to optimize the traffic signal control strategy.
This paper realizes forward-looking traffic signal control in a continuous action space, improves the predictability and flexibility of the controller, overcomes the limitations of discretization and insufficient response of existing methods, and optimizes the decision-making process of the traffic signal cycle.
Smart Images

Figure CN116235229B_ABST
Abstract
Description
[0001] Related application data
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 080,455 filed on September 18, 2020 and U.S. Non-Provisional Patent Application No. 17 / 327,523 filed on May 21, 2021, both entitled “Method and System for Adaptive Cycle-Level Traffic Signal Control.” Technical Field
[0003] The present application relates generally to methods and systems for traffic signal control, and more particularly to adaptive cycle-level traffic signal control. Background Art
[0004] Traffic congestion causes significant wasted time, fuel consumption, and pollution. Building new infrastructure to eliminate these problems is often impractical due to financial and spatial constraints, as well as environmental and sustainability concerns. Therefore, to increase the capacity of urban transportation networks, researchers have explored the use of technology to maximize the performance of existing infrastructure. Optimizing the operation of traffic signals holds promise for reducing delays for drivers in urban networks.
[0005] Traffic signals are used to communicate traffic regulations to drivers of vehicles operating in a traffic environment. A typical traffic signal controller controls a traffic signal that manages vehicular traffic within a traffic environment consisting of a single intersection in a traffic network. Thus, for example, a single traffic signal controller may control a traffic signal consisting of red / yellow / green traffic lights facing four directions (north, south, east, and west), but it should be understood that some traffic signals may control traffic in an environment consisting of more or fewer than four traffic directions and may include other signal types, such as different signals for different lanes facing the same direction, turn arrows, street-level public transportation signals, etc.
[0006] Traffic signals typically operate in cycles, with each cycle consisting of several phases. A single phase can correspond to a fixed state for the signal's lights, for example, green for north-south and red for east-west, or yellow for north-south and red for east-west. However, some phases may include additional, non-fixed states, such as a timer counting down for a crosswalk. Typically, a traffic signal cycle consists of each phase repeating once in the cycle, typically in a fixed order.
[0007] Figure 1 An exemplary traffic signal cycle 100 is shown consisting of eight phases, sequentially from a first phase 102 to an eighth phase 116. In this example, all other lights are red during a phase unless otherwise noted.
[0008] In a first stage 102, i.e., Stage 1, the traffic signal displays a green left-turn arrow, indicated as "NL," to northbound traffic (i.e., on a south-facing lamppost), and displays a green left-turn arrow, indicated as "SL," to southbound traffic (i.e., on a north-facing lamppost). In a second stage 104, i.e., Stage 2, the traffic signal displays a green left-turn arrow and a green "straight" light or arrow, indicated as "SL" and "ST," respectively, to southbound traffic. In a third stage 106, i.e., Stage 3, the traffic signal displays a green left-turn arrow and a green "straight" light or arrow, indicated as "NL" and "NT," respectively, to northbound traffic. In a fourth stage 108, i.e., Stage 4, the traffic signal displays a yellow left-turn arrow (shown as a dashed line) and a green "straight" light or arrow to both northbound and southbound traffic. In the fifth stage 110, i.e., Stage 5, the traffic signal displays a green left-turn arrow, indicated as "EL," to eastbound traffic (i.e., on west-facing lampposts), and displays a green left-turn arrow, indicated as "WL," to westbound traffic (i.e., on east-facing lampposts). In the sixth stage 112, i.e., Stage 6, the traffic signal displays a green left-turn arrow and a green "go straight" light or arrow, indicated as "WL" and "WT," respectively, to westbound traffic. In the seventh stage 114, i.e., Stage 7, the traffic signal displays a green left-turn arrow and a green "go straight" light or arrow, indicated as "EL" and "ET," respectively, to eastbound traffic. In the eighth stage 116, i.e., Stage 8, the traffic signal displays a yellow left-turn arrow (shown as a dashed line) and a green "go straight" light or arrow to both westbound and eastbound traffic.
[0009] After completing phase 8 116, the traffic signal returns to phase 1 102. Traffic signal controller optimization typically involves optimizing the duration of each phase of the traffic signal cycle to achieve traffic objectives.
[0010] The most common methods of traffic signal control are fixed-time and actuated. In a fixed-time traffic signal controller configuration, each phase of a traffic signal cycle has a fixed duration. The fixed-time controller uses historical traffic data to determine the optimal traffic signal pattern. This optimized fixed-time signal pattern (i.e., the set of phase durations for the cycle) is then deployed to control the actual traffic signal, and the pattern is fixed thereafter.
[0011] In contrast to fixed-time controllers, actuated signal controllers receive feedback from sensors to respond to traffic flow. However, actuated signal controllers do not explicitly optimize delays. Instead, they typically adjust signal patterns in response to immediate traffic conditions, without adapting to traffic flow over time. Therefore, the duration of a phase may be extended based on current traffic conditions based on sensor data, but there is no mechanism to use data from past phases or cycles to optimize traffic signal operation over time, or to make decisions based on optimizing performance metrics such as average or aggregate vehicle delay.
[0012] Adaptive traffic signal controllers (ATSCs) are more advanced and can outperform other controllers, such as fixed-time or actuated controllers. ATSCs continuously modify signal timing to optimize a predetermined goal or performance metric. Some ATSCs, including SCOOT, SCATS, PRODYN, OPAC, UTOPIA, and RHODES, optimize signals using an internal model of the traffic environment that is often simplistic and rarely keeps up with current conditions. The optimization algorithms used by these ATSCs are mostly heuristic and suboptimal. Designing accurate traffic models is difficult due to the stochastic nature of traffic and driver behavior. More realistic models are also more complex and difficult to control, sometimes resulting in computational delays that are too long to achieve real-time traffic control. Therefore, there is a trade-off between controller complexity and practicality.
[0013] However, this field has seen some progress with the emergence of reinforcement learning (RL), a model-free closed-loop control method for optimization. RL algorithms learn optimal control policies while interacting with the environment and evaluating their own performance. Recently, researchers have used deep reinforcement learning (DRL) using convolutional neural networks in ATSC. Examples of DRL traffic signal control systems are described in the following papers: W. Gessions and S. Razavi, “Using a Deep Reinforcement Learning Agent for Traffic Signal Control,” CoRR, vol. abs / 1611.0, 2016; J. Gao, Y. Shen, J. Liu, M. Ito, and N. Shiratori, “Adaptive Traffic Signal Control: Deep Reinforcement Learning Algorithm with Experience Replay and Target Network,” CoRR, vol. abs / 1705.0, 2017; S.M.A. Shabesary and B. Abdulhai, “Deep Learning vs. Discrete Reinforcement Learning for Adaptive Traffic Signal Control,” CoRR, vol. abs / 1706.0, 2017; Control), 2018 21st International Conference on Intelligent Transportation Systems (ITSC), 2018, pp. 286–293, all of which are incorporated herein by reference in their entirety.
[0014] Compared to other RL methods that use function approximation methods, deep reinforcement learning is able to handle large state space problems and achieve better performance. In some DRLATSCs, the street surface is discretized into small cells, which are grouped together to create a position and velocity matrix of vehicles approaching the intersection, and this matrix is used as input to a deep Q-network that performs the DRL task.
[0015] Existing DRL controllers are designed to take action every second, so-called second-based control. At every second, the DRL decides whether to extend the current green signal or switch to another phase. These controllers require a reliable high-frequency communication infrastructure and powerful computing units to effectively monitor the traffic environment and control traffic signals on a second-by-second time scale. Furthermore, since the controller's behavior cannot be known even one second in advance, some municipalities and transportation departments are dissatisfied with controllers that make decisions every second. Instead, they prefer to know in advance what each phase of the next cycle will look like, as is possible with fixed-time controllers. Furthermore, the possibility of a green signal terminating at any second may also conflict with crosswalk safety, as it may be difficult or impossible to configure a pedestrian countdown timer to allow pedestrians who have already entered the crosswalk to cross safely.
[0016] For these reasons, traffic signal controllers capable of making decisions less frequently than once per second may offer certain advantages. It is possible to implement traffic signal controllers that generate decision data for the entire cycle, which can be referred to as cycle-based control. A cycle-based controller can generate duration data for all phases of the next traffic signal cycle. However, by limiting the controller's interaction with the traffic signal, this approach can reduce the controller's flexibility in responding to changes in the traffic environment in real time. Literature on cycle-based RL-based traffic signal control is limited, at least in part due to the complexity and large action space. In a second-based control approach with a fixed order of phases within each cycle, the controller must decide whether to extend the current green phase or switch to the next phase, resulting in a discrete action space of size two (0 = extend, 1 = switch). At most, a second-based controller with a flexible ordering of phases within each cycle must decide not only whether to switch (extend or switch = 2 actions) but also which of the possible phases to switch to (n phases in a cycle = n actions). In this case, the action space size is a discrete set of n (n possible phases at each intersection, which in most cases is limited to a maximum of 8 phases, n = 8).
[0017] On the other hand, a cycle-based controller has to cope with a continuous action space. The traffic signal cycle and each of its phases can be of any time length. Compared to a second-based control problem, even with time discretization, the action space increases dramatically. In a first example, the traffic signal at the intersection has a cycle with 4 phases (e.g., north and south, north left turn and south left turn, east and west, east left turn and west left turn). Assuming a minimum green time of 10 seconds and a maximum green time of 30 seconds for all phases, the action space is the number of seconds duration values that can be selected for the current phase (i.e., 20) raised to the power of the number of phases to switch to (i.e., 4), i.e., 20 4 =160,000.
[0018] This problem is discussed in M. Aslani, M. S. Mesgari, and M. Wiering, “Adaptive traffic signal control with actor-critic methods in a real-world traffic network with different traffic disruption events,” Transp. Res. Part C Emerg. Technol., Vol. 85, pp. 732–752, 2017 (hereinafter referred to as “Aslani”), which is incorporated herein by reference in its entirety. Aslani solves this problem by discretizing the action space into intervals of 10 seconds. Therefore, the controller for each stage must choose the stage duration from the set [0 seconds, 10 seconds, 20 seconds, … 90 seconds], which is a very coarse discretization that may affect the performance of the controller.
[0019] Another approach, described in X. Liang, X. Du, G. Wang, and Z. Han, “A Deep Reinforcement Learning Network for Traffic Light Cycle Control” (IEEE Trans. Veh. Technol., Vol. 68, No. 2, pp. 1243–1253, 2019), which is incorporated herein by reference in its entirety, uses an incremental approach to setting signal timing. The controller does not directly define the stage duration, but it decides at each decision point to increase or decrease the timing of each stage by 5 seconds. This approach not only suffers from a coarse discretization of the action space, but also lacks the flexibility to react to sudden changes.
[0020] Therefore, there is a need for an adaptive traffic signal controller that can proactively generate one or more phase durations for a traffic signal cycle within a continuous range of duration values, thereby overcoming one or more limitations of the above-mentioned existing approaches. Summary of the Invention
[0021] The present disclosure describes methods, systems, and processor-readable media for adaptive cycle-level traffic signal control. An intelligent adaptive cycle-level traffic signal controller and control method that operates in a continuous action space are described. As described above, most existing adaptive traffic signal controllers operate second by second, which has the above-mentioned disadvantages in terms of safety, predictability, and communication and computing requirements. Existing adaptive cycle-level traffic signal controllers are model-based, offline, or preliminary. The embodiments described herein may include a continuous action adaptive cycle-level traffic signal controller or control method using a reinforcement learning algorithm called Proximal Policy Optimization (PPO), which is an action-critic model of reinforcement learning. In some embodiments, the controller does not treat the action space as discrete, but instead produces continuous values as output, making established RL methods such as deep Q networks (DQNs) unusable. In some embodiments, for intersections with four phases in the traffic signal cycle, the controller generates four consecutive numbers, each indicating the duration of a phase of the cycle.
[0022] In some aspects, the present disclosure describes a method for training a reinforcement learning model to generate traffic signal cycle data. A training data sample indicating an initial state of a traffic environment affected by a traffic signal is processed by performing several operations. The reinforcement learning model is used to generate traffic signal cycle data by applying a policy to the training data sample and one or more past training data samples. The traffic signal cycle data includes one or more phase durations for one or more corresponding phases of the traffic signal cycle. Each phase duration is a value selected from a continuous range of values. After applying the generated traffic signal cycle data to the traffic signal, an updated state of the traffic environment is determined. A reward is generated by applying a reward function to the initial state of the traffic environment and the updated state of the traffic environment. The policy is adjusted based on the reward. The step of processing the training data sample is repeated one or more times. The training data sample indicates the updated state of the traffic environment.
[0023] In some aspects, the present disclosure describes a method and system for training a reinforcement learning model to generate traffic signal cycle data. The system includes a processor device and a memory. The memory stores a reinforcement learning model and machine-executable instructions. When executed by the processing device, the machine-executable instructions cause the system to process a training data sample indicating an initial state of a traffic environment affected by a traffic signal by performing several operations. The training data sample indicating the initial state of the traffic environment affected by the traffic signal is processed by performing several operations. The reinforcement learning model is used to generate traffic signal cycle data by applying a policy to the training data sample and one or more past training data samples. The traffic signal cycle data includes one or more phase durations for one or more corresponding phases of the traffic signal cycle. Each phase duration is a value selected from a continuous range of values. After applying the generated traffic signal cycle data to the traffic signal, an updated state of the traffic environment is determined. A reward is generated by applying a reward function to the initial state of the traffic environment and the updated state of the traffic environment. The policy is adjusted based on the reward. The step of processing the training data sample is repeated one or more times. The training data sample indicates the updated state of the traffic environment.
[0024] In some examples, the traffic environment is a simulated traffic environment, and the traffic signal is a simulated traffic signal.
[0025] In some examples, the one or more phase durations include a phase duration for each phase of at least one cycle of the traffic signal.
[0026] In some examples, the one or more phase durations consist of a phase duration of one phase of a cycle of the traffic signal.
[0027] In some examples, the reinforcement learning model is an actor-critic model, the policy is an actor-policy, and the reward function is a critic-reward function.
[0028] In some examples, the actor-critic model is a proximal policy optimization (PPO) model.
[0029] In some examples, each training data sample includes traffic data, including position data and speed data for each of a plurality of vehicles in a traffic environment.
[0030] In some examples, each training data sample includes traffic data, including traffic density data and traffic speed data for each of a plurality of areas in the traffic environment.
[0031] In some examples, determining an updated state of a traffic environment includes determining a length of each of one or more queues of stationary vehicles in the traffic environment. The length indicates the number of stationary vehicles in the queue. The one or more past training data samples include one or more past training data samples corresponding to one or more queue peak times (each queue peak time is a time when the length of a queue is at a local maximum), and one or more past training data samples corresponding to one or more queue valley times (each queue valley time is a time when the length of a queue is at a local minimum).
[0032] In some examples, the one or more past training data samples correspond to one or more phase transition times. Each phase transition time is the time at which a traffic signal transitions between two phases of a traffic signal cycle.
[0033] In some examples, a reward function is applied to an initial state of a traffic environment and an updated state of the traffic environment to calculate a reward based on an estimated number of stationary vehicles in the traffic environment during a previous traffic signal period.
[0034] In some examples, the one or more past training data samples correspond to one or more phase transition times. Each phase transition time is the time at which a traffic signal transitions between two phases of a traffic signal cycle.
[0035] In some examples, each training data sample includes traffic signal phase data indicating a current phase of a traffic signal cycle and a time elapsed during the current phase.
[0036] In some examples, these one or more phase durations include the phase duration of each phase of at least one cycle of the traffic signal. The reinforcement learning model is a proximal policy optimization (PPO) behavior-critic model. The policy is a behavioral policy. The reward function is a judgment reward function. Each training data sample includes: traffic signal phase data and traffic data. The traffic signal phase data indicates the current phase of the traffic signal cycle and the time elapsed during the current phase. The traffic data includes traffic density data and traffic speed data for each of multiple areas of the traffic environment. The reward function is applied to the initial state of the traffic environment and the updated state of the traffic environment to calculate the reward based on the estimated number of stationary vehicles in the traffic environment during the previous traffic signal cycle. One or more past training data samples correspond to one or more phase transition times. Each phase transition time is the time when the traffic signal transitions between two phases of the traffic signal cycle.
[0037] In some aspects, the present disclosure describes a system for generating traffic signal cycle data. The system includes a processor device and a memory. The memory stores a trained reinforcement learning model trained according to the above-described method steps, and machine-executable instructions that, when executed by the processor device, cause the system to perform several operations. Traffic environment state data indicating a state of a real traffic environment is received from a traffic monitoring system. The traffic environment used to train the reinforcement learning model is a real traffic environment or a simulated version thereof. The reinforcement learning model is used to generate traffic signal cycle data by applying a policy to at least the traffic environment state data. The traffic signal cycle data is transmitted to a traffic control system.
[0038] In some aspects, the present disclosure describes a non-transitory processor-readable medium having stored thereon a trained reinforcement learning model trained according to the above method steps.
[0039] In some aspects, the present disclosure describes a non-transitory processor-readable medium having stored thereon machine-executable instructions that, when executed by a processor device, cause the processor device to perform the above-described method steps. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Reference will now be made, by way of example, to the accompanying drawings which show exemplary embodiments of the present application, in which:
[0041] Figure 1 An exemplary operating environment for the exemplary embodiments described herein is illustrated in a table of eight phases of an exemplary traffic signal cycle.
[0042] Figure 2 A block diagram of an exemplary traffic environment at an intersection including traffic signals and sensors in communication with traffic signal controllers for embodiments described herein.
[0043] Figure 3 is a block diagram of an exemplary traffic signal controller according to embodiments described herein.
[0044] Figure 4 A flowchart of steps of an exemplary method for training a reinforcement learning model to generate traffic signal cycle data according to embodiments described herein.
[0045] Figure 5 A top view of the traffic environment at an intersection where traffic lanes are divided into cells according to the embodiments described herein.
[0046] Figure 6 The traffic environment state data converted into training data samples for the embodiments described herein Figure 5 Schematic diagram of vehicle speed and vehicle density data for the medium cell.
[0047] Figure 7A For the embodiments described herein Figure 1 At the end of Phase 4 of the traffic signal cycle Figure 5 A top view of the traffic environment, showing different vehicle queue lengths in different traffic directions.
[0048] Figure 7B For the embodiments described herein Figure 1 At the end of phase 8 of the traffic signal cycle Figure 5 A top view of the traffic environment, showing different vehicle queue lengths in different traffic directions.
[0049] Figure 8 A block diagram of an exemplary behavior module of a traffic signal controller according to an embodiment described herein, showing traffic environment state data at a point in time as input and generated traffic signal cycle data as output.
[0050] Figure 9 A block diagram of an exemplary behavior module of a traffic signal controller according to an embodiment described herein, showing traffic environment state data at multiple time points as input and generated traffic signal cycle data as output.
[0051] Figure 10A For the embodiments described herein Figure 5 In the exemplary traffic environment Figure 1 Plot of the vehicle queue length versus time for southbound traffic during the first four phases of the traffic signal cycle.
[0052] Figure 10B yes Figure 10A Graph showing queue lengths approximated as linear interpolations between queue lengths at phase transition times according to embodiments described herein.
[0053] Figure 10C yes Figure 10B Graph showing queue lengths for southbound and northbound traffic at phase transition times according to embodiments described herein.
[0054] Figure 10D yes Figure 10C , illustrating training data samples generated at phase transition times according to embodiments described herein.
[0055] Figure 11A A graph of stationary vehicle distance from an intersection over time for embodiments described herein shows total delay calculated as the sum of the lengths of the stationary vehicle queues.
[0056] Figure 11B A graph of stationary vehicle distance from an intersection over time for the embodiments described herein shows a total delay calculated as the sum of the duration of the stationary period for each vehicle or unit.
[0057] Figure 11C A graph of stationary vehicle distance from an intersection over time for embodiments described herein shows the total delay calculated as the area of the triangle defined by the queue of stationary vehicles over time.
[0058] Like reference numerals may be used in different drawings to identify like components. DETAILED DESCRIPTION
[0059] In various examples, the present disclosure describes methods, systems, and processor-readable media for adaptive cycle-level traffic signal control in a continuous action space. Various embodiments are described below with reference to the accompanying drawings. The description of the exemplary embodiments is divided into multiple sections. The exemplary controller device section describes an exemplary device or computing system suitable for implementing the exemplary traffic signal controller and method. The exemplary reinforcement learning model section describes how the controller learns and updates the parameters of the RL model. The traffic signal cycle data example section describes the action space and outputs of the controller. The traffic environment state data example section describes the state space and inputs of the controller. The exemplary reward function section describes the reward function of the controller. The exemplary system for controlling traffic signals section describes the operation of the trained controller when used to control traffic signals in an actual traffic environment.
[0060] Exemplary Controller Device
[0061] Figure 2 2 is a block diagram illustrating an exemplary traffic environment 200 at an intersection 201 including traffic signals and sensors in communication with an exemplary traffic signal controller 220. The traffic signals are shown as four traffic lights: a south-facing light 202, a north-facing light 204, an east-facing light 206, and a west-facing light 208. Each traffic light 202, 204, 206, 208 includes a sensor, shown as a long-range camera 212 facing the same direction as the light. (In all figures showing a top-down view of the traffic environment, north corresponds to the top of the page.) The controller device 220 sends control signals to the four traffic lights 202, 204, 206, 208 and receives sensor data from the four cameras 212. The controller device 220 is also in communication with a network 210, through which the controller device can communicate with one or more servers or other devices, as described in more detail below.
[0062] It should be understood that although embodiments are described herein with reference to a traffic environment consisting of a single intersection managed by a single signal (e.g., a single set of traffic lights), in some embodiments, the traffic environment may include multiple nodes or intersections within a transportation grid and may control multiple traffic signals.
[0063] Figure 32 is a block diagram illustrating a simplified example of a controller device 220, such as a computer or cloud computing platform, suitable for executing the examples described herein. Other examples suitable for implementing the embodiments described in this disclosure may be used and may include components different from those described below. Although Figure 3 A single instance of each component is shown, but multiple instances of each component in the controller device 220 may exist.
[0064] The controller device 220 may include one or more processor devices 225, such as a processor, a microprocessor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a dedicated logic circuit, a dedicated artificial intelligence processing unit, or a combination thereof. The controller device 220 may also include one or more optional input / output (I / O) interfaces 232 that can be connected to one or more optional input devices 234 and / or optional output devices 236.
[0065] In the illustrated example, input devices 234 (e.g., a maintenance console, keyboard, mouse, microphone, touch screen, and / or keypad) and output devices 236 (e.g., a maintenance console, display, speakers, and / or printer) are shown as optional and external to controller device 220. In other examples, no input devices 234 and output devices 236 may be present, in which case I / O interface 232 may not be required.
[0066] The controller device 220 may include one or more network interfaces 222 for wired or wireless communication with one or more devices or systems in a network, such as the network 210. The network interfaces 222 may include wired links (e.g., Ethernet cables) and / or wireless links (e.g., one or more antennas) for intra-network and / or inter-network communication. The one or more network interfaces 222 may be used to send control signals to the traffic signals 202, 204, 206, 208 and / or to receive sensor data from sensors (e.g., camera 212). In some embodiments, the traffic signals and / or sensors may communicate with the controller device directly or indirectly via other means (e.g., I / O interface 232).
[0067] The controller device 220 may also include one or more storage units 224, which may include mass storage units such as solid-state drives, hard disk drives, magnetic disk drives, and / or optical disk drives. The storage unit 224 may be used for long-term storage of some or all of the data stored in the memory 228 described below.
[0068] The controller device 220 may include one or more memories 228, which may include volatile or non-volatile memory (e.g., flash memory, random access memory (RAM), and / or read-only memory (ROM)). The non-transitory memory 228 may store instructions executed by the processor device 225, for example, to perform the examples described in the present disclosure. The memory 228 may include software instructions 238, for example, to implement an operating system and other applications / functions. In some examples, the memory 228 may include software instructions 238 for execution by the processor device 225 to implement the reinforcement learning model 240, as further described below. In some examples, the memory 228 may include software instructions 238 for execution by the processor device 225 to implement the simulator module 248, as further described below. By executing the instructions 238 using the processor device 225, the reinforcement learning model 240 and the simulator module 248 may be loaded into the memory 228.
[0069] In some embodiments, the simulator module 248 can be a traffic micro-simulation software, such as the Simulation of Urban Mobility (SUMO) software. SUMO is an open source microscopic traffic simulator software that provides users and developers with the option to customize simulation model parameters and functionality through a functional interface or application programming interface (API). The API can be used to train cycle-level traffic signal controllers in a simulation environment that is very close to reality, as described in more detail below. It should be understood that the simulator module 248 is only required during training and not during reasoning (e.g., deployment in a traffic environment). Therefore, the simulator module 248 can be present on the training device, but not on the controller device 220 that controls the actual traffic signal using the trained RL model 240.
[0070] In some embodiments, the RL model 240 can be coded in the Python programming language using the tensorflow machine learning library and other widely used libraries including NumPy. To create a link between the RL model 240 and the simulator module 248, a wrapper can be written in Python to apply the actions of the behavior module 244 of the RL model 240 to the SUMO network (i.e., the simulator module 248), extract state and reward information, and pass this information back to the RL model 240 (specifically, to the judge module 246). It should be understood that other embodiments may use different simulator software, different software libraries, and / or different programming languages.
[0071] The memory 228 may also include one or more samples of traffic environment state data 250, which may be used as training data samples for training the reinforcement learning model 240 and / or as input to the reinforcement learning model 240 for generating traffic signal cycle data after training the reinforcement learning model 240 and deploying the controller device 220 to control traffic signals in an actual traffic environment, as described in detail below.
[0072] In some examples, the controller device 220 may additionally or alternatively execute instructions from an external memory (e.g., an external drive in wired or wireless communication with the controller device 220), or the executable instructions may be provided by a transient or non-transitory computer-readable medium. Examples of non-transitory computer-readable media include RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, CD-ROM, or other portable memory.
[0073] The controller device 220 may also include a bus 242 that provides communication between components of the controller device 220, including those discussed above. The bus 242 may be any suitable bus architecture including, for example, a memory bus, a peripheral bus, or a video bus.
[0074] It should be understood that in some embodiments, various components and operations described herein may be implemented on multiple separate devices or systems.
[0075] Example reinforcement learning model
[0076] In some embodiments, a self-learning traffic signal controller interacts with a real or simulated traffic environment and gradually finds the optimal strategy for traffic signal control. A controller (e.g., controller device 220) generates traffic signal cycle data by applying a function to traffic environment state data and using a learned strategy to determine an action (i.e., a traffic signal control action in the form of traffic signal cycle data) based on the output of the function. A model is used to approximate the function, and the model is trained using reinforcement learning, which is sometimes referred to here as a "reinforcement learning model." In some embodiments, the reinforcement learning model (e.g., reinforcement learning model 240) can be an artificial neural network, such as a convolutional neural network. In some embodiments, the traffic environment state data (e.g., traffic environment state data 250) can be formatted as one or more two-dimensional matrices, so that a convolutional neural network or other RL model applies known image processing techniques to generate traffic signal cycle data.
[0077] Reinforcement learning (RL) is a technique well-suited for optimal control problems with highly complex dynamics. These problems can be difficult to model, difficult to control, or both. In RL, the controller can be functionally represented as an agent that has no knowledge of the environment in which it operates. In the early stages of training, the agent begins taking random actions, known as exploration. For each action, the agent observes changes in the environment (for example, by monitoring the real traffic environment through sensors or by receiving simulated traffic from a simulator) and receives a numerical value called a reward that indicates the desirability of its action. The agent's goal is to optimize the cumulative reward over time, rather than the immediate reward after any given action. This optimization of cumulative reward is essential in fields such as traffic signal control, where the agent's actions affect the future state of the system, requiring the agent to consider the future impact of its actions rather than their immediate effects. As training progresses, the agent begins to understand the environment and takes less random actions; in fact, it takes actions that, based on its experience, lead to better system performance.
[0078] In some embodiments, the controller uses an actor-critic reinforcement learning model. In particular, a Proximal Policy Optimization (PPO) model, which is a variant of a deep actor-critic RL model, may be used in some embodiments. An actor-critic RL model may generate continuous action values (e.g., the duration of a traffic signal cycle phase) as output. An actor-critic RL model has two parts: an actor, which defines the agent's policy; and a critic, which helps the actor optimize its policy during training. The output of the actor can be represented as a policy:
[0079] π(a t |s t ;θ)(Equation 1)
[0080] Represents the state s under given model parameters θ t Next select action a t The probability π. The output of the judgment can be expressed as:
[0081] V π (s t |θ v )(Equation 2)
[0082] Represents a given policy π and model parameters θ v Status t The estimated expected value V of .
[0083] As mentioned above, the goal of RL is to optimize the agent’s expected cumulative reward, also known as return:
[0084]
[0085] where r t is the reward signal at time step t, and γγ∈(0,1] is a discount factor for stability reasons and taking into account the future characteristics of the agent. The lower the value of γ, the more important the immediate reward is to the agent, and the less important the future rewards are.
[0086] The agent attempts to use a v The function approximator estimates the current policy in state s t The expected return (i.e., value function) under:
[0087]
[0088] Such that the value function is the probabilistic expectation of the cumulative time-discounted reward.
[0089] In the action-critic model, the value function error (also called loss function) calculated by the critic is defined as:
[0090] L(θ v,t )=E[(r t +γV(s t+1 θ v,t-1 )-V(s t θ v,t )) 2 ]. (Equation 5)
[0091] Therefore, the parameter θ of the value function V is v According to the error function L(θ v,t ) with respect to the gradient update of the neural network weights, as follows:
[0092] R=r t +γV(s t+1 θ v,t-1 )(Equation 6)
[0093]
[0094] The behavior generates the agent's policy. The policy is generated in state s t The probability of choosing action a is:
[0095] π(a|s t θ a ). (Equation 8)
[0096] If the agent takes an action that yields a better reward than expected, the policy should be adjusted to increase the probability of choosing that action; similarly, actions with lower-than-expected rewards should cause the policy to be adjusted to decrease the probability of taking those actions. Therefore, the behavior updates the parameters θ that define the policy as follows:
[0097]
[0098] Where V(s t θ v,t ) Estimated state s t The expected return, and RV(s t θ v,t ) indicates the advantage of action a compared to the expected reward. In this example, the term RV(s t θ v,t ) is called the "advantage function". However, since other methods can be used to calculate the advantage, the advantage function can be generally referred to as A t An exemplary advantage function used in the current example is A t =RV(s t θ v,t ); however, other methods of computing advantages may also be used.
[0099] This policy update is only available from the policy π(a|s t ;θ) or at least from something very similar to π(a|s t ; θ) is valid only when samples are extracted from the policy (i.e., state, action, reward, and next state) of the previous policy. Therefore, if the behavior updates the policy according to Equation 5-9 above, the old samples cannot be used to update the policy. To use older samples, PPO can be used to adjust the values used in Equation 5-9 above. In the PPO algorithm, as long as the policy from which the sample was extracted does not differ from the current policy by more than a certain amount, the older samples can be used to update the current policy.
[0100] PPO makes two modifications to the actor-critic algorithm described above. The first change is to state that the update is based on older samples from different policies. The second change is to ensure that the policy where the older sample is located (π(a|s t θ old )) and the current strategy (π(a|s t ; θ)) has no significant difference. The loss function is calculated based on these two changes. Therefore, r t (θ) represents the probability ratio So that r(θ old )=1. The loss function L(θ) can be defined as:
[0101]
[0102] However, without constraints, maximizing L(θ) will lead to excessively large policy updates. Therefore, the objective should be modified to penalize making r t (θ) deviates from the policy change of 1. Therefore, the behavior loss function L in the PPO model CLIP The final form of (θ) is:
[0103] L CLIP (θ)=E[min(r t (θ)A t ,clip(r t (θ),1-∈,1+∈)A t )](Equation 11)
[0104] where ∈ is a hyperparameter that limits the policy change due to updates. The critic loss function is the same as in traditional actor-critic RL models, as defined in Equation 5 above.
[0105] Exemplary Training Methods
[0106] The RL model 240 used by the controller device 220 must be trained before it can be deployed to control traffic signals in a real traffic environment. Training is performed by providing traffic environment data to the RL model 240, using traffic signal cycle data generated by the RL model 240 to control traffic signals in the traffic environment, and then providing traffic environment data representing an updated state of the traffic environment data to the RL model for use in adjusting the RL model strategy and generating data for future traffic signal cycles. Figure 5-11C Describe the traffic environment data in more detail.
[0107] Training can be performed using data from a simulated traffic environment, for example, using the simulator module 248. The simulator module 248 can generate simulated traffic environment data and provide the simulated traffic environment data to the RL model 240. The RL model 240 generates traffic signal cycle data, which is provided to the simulator module 248 and used to model the response of the traffic environment to the traffic signal cycles applied by the simulated traffic signal. In some embodiments, the RL model 240 can be trained using the simulated traffic environment during early training stages, but later fine-tuning of the RL model 240 can be performed using the actual traffic environment.
[0108] Figure 4 An exemplary method 400 for training a reinforcement learning model to generate traffic signal cycle data is shown.
[0109] At 402, training data samples are generated based on an initial state of a traffic environment. If the traffic environment is a real traffic environment, the state of the traffic environment can be determined by a traffic monitoring system based on inputs from sensors (e.g., camera 212) that monitor the traffic environment. The traffic monitoring system can be separate from the controller device 220 and can communicate with the controller device 220 via the network 210. In some embodiments, the controller device 220 can directly receive sensor data from the sensors, relay the sensor data to the traffic monitoring system, and then receive the processed traffic environment data from the traffic monitoring system. In other embodiments, the traffic monitoring system can be implemented as part of the controller device. The traffic monitoring system can include, for example, a computer vision system for determining vehicle speed, position, and / or density data within the traffic environment based on the sensor data, as described in more detail below with reference to Figure 5 When the controller device 220 receives the traffic environment data, the traffic environment data can be formatted as a training data sample, or the controller device 220 can reformat the traffic environment data as a training data sample and then provide it to the RL model 240.
[0110] At 404, when a training data sample is received, the behavior module 244 of the RL model applies its policy to the training data sample (corresponding to s t ) and one or more past training data samples (corresponding to s j , where j < t) to generate traffic signal cycle data, as described in more detail below with reference to Figure 5-11C , particularly Figure 9 In some embodiments, as described with reference to Figure 10D , one or more past training data samples can correspond to time points when the traffic signal transitions between two phases. In other embodiments, one or more past training data samples can correspond to time points when the length of a queue of stationary vehicles in the traffic environment is at a maximum or minimum value.
[0111] The traffic signal cycle data generated in step 404 can be one or more phase durations of one or more corresponding phases of a traffic signal cycle. In some embodiments, each phase duration is a value selected from a continuous range of values. In some examples, such selection of phase durations from a continuous range of values can be achieved by using an actor-critic RL model, as described in detail above.
[0112] In some embodiments, the traffic signal cycle data generated in step 404 includes the phase durations of each phase of at least one cycle of the traffic signal. In other embodiments, the traffic signal cycle data generated in step 404 includes the phase duration of only one phase of the traffic signal cycle. There may be a trade-off between granularity and predictability in cycle-level control and phase-level control.
[0113] At 406, the traffic signal cycle data is applied to a real or simulated traffic signal. In the case of a real traffic environment using real traffic signals, the controller device 220 may send control signals to the traffic signals (e.g., lights 202, 204, 206, 208) to implement the phase durations specified by the traffic signal cycle data. In the case of a simulated traffic environment, the RL model provides the traffic signal cycle data to the simulator module 248, which simulates the response of the traffic environment to the phase durations of the traffic signal cycle data implemented by the simulated traffic signals.
[0114] At 408 , an updated state of the traffic environment is determined. As at step 402 , the state of the traffic environment may be determined by the traffic monitoring system based on sensor data from the traffic environment (if a real traffic environment is used) or by the simulator module 248 (if a simulated traffic environment is used).
[0115] Optionally, step 408 may include sub-step 409, in which the traffic monitoring system or simulator module 248 determines the length of one or more queues of stationary vehicles in the traffic environment. This data may be used to calculate rewards, as described below with reference to Figures 10A-11C Describe in more detail.
[0116] At 410, new training data samples are generated based on the updated state of the traffic environment determined at step 408. In some embodiments, the frequency with which step 410 is performed may differ from the frequency with which step 408 (and optionally step 409) is performed: for example, training data samples may be generated by step 410 only at time points corresponding to transitions between phases of a traffic signal cycle (i.e., a new training data sample is generated when phase 1 102 ends and phase 2 104 begins, and another training data sample is generated when phase 2 104 ends and phase 3 106 begins), while the updated state of the traffic environment may be determined by step 408 every second or even more frequently.
[0117] At 412, the reward function is applied to the initial state of the traffic environment and the updated state of the traffic environment to generate a reward value, as described above. For the purpose of calculating the reward using the judgment module 246, the initial state can be considered as s t , and the updated state can be regarded as s t+1 , as shown in Equation 1-11 above.
[0118] At 414, the behavior module 244 adjusts its policy based on the reward generated at step 412. The weights or parameters of the RL model can be adjusted using RL techniques including the PPO actor-critic technique described in detail above.
[0119] Method 400 then returns to step 404 to repeat the step 404 of processing the training data sample, which (generated at step 410) now indicates the updated state of the traffic environment (determined at step 408). This loop may be repeated one or more times (typically at least hundreds or thousands of times) to continue training the RL model.
[0120] Therefore, method 400 can be used to train an RL model and update the parameters of its policy based on the above-mentioned actor-critic RL technique or other RL techniques.
[0121] Traffic signal cycle data example
[0122] In some embodiments, the action space used by the behavior module 244 of the RL model 240 can be a continuous action space, such as a natural number space. Compared to existing second-level methods, embodiments operating under cycle-level or stage-level control of traffic signals have relatively low-frequency interactions with traffic signals: a cycle-level controller can send a control signal to the traffic signal once per cycle, for example, at the beginning of a cycle, while a stage-level controller can send a control signal to the traffic signal once per stage, for example, at the beginning of a stage.
[0123] Therefore, for a traffic signal with P phases per cycle (e.g., Figure 1 In the example of , P = 8), the output of the reinforcement learning model 240 using cycle-level control is P natural numbers, each natural number indicating the length of the traffic signal phase. Figure 9 An example of traffic signal cycle data including multiple phase durations is described. The reinforcement learning model 240 using phase-level control may generate only one natural number indicating the length of a traffic signal phase. Other embodiments may generate a different number of phase durations.
[0124] In some embodiments, the phase durations generated by the reinforcement learning model 240 are selected from a continuous range of positive real numbers. Using an actor-critic RL model allows for the generation of phase durations that are selected from a continuous range of values rather than a finite number of discrete values (e.g., 5-second or 10-second intervals as in prior art methods).
[0125] Traffic environment status data example
[0126] As described above, the controller device 220 provides traffic environment data to the RL model 240, which generates traffic signal cycle data and adjusts its strategy based on the traffic environment data. An example of how traffic environment data is collected, represented, and formatted as training data samples or traffic environment state data for use by the RL model will now be described.
[0127] In different embodiments, different state spaces may be used to represent the state of the traffic environment as traffic environment data. Typically, the traffic environment data will include traffic data indicating some aspect of the behavior and / or existence of vehicle traffic in the environment. In some embodiments, the traffic environment data includes a queue length for each traffic lane in the traffic environment, each queue length indicating the length of a queue of stationary vehicles in that lane. Thus, if the traffic environment includes a 50-meter radius around the center of a four-way intersection of two four-lane roads (i.e., two lanes northbound, two lanes southbound, two lanes eastbound, and two lanes westbound), the traffic environment data may include eight queue lengths indicating the number of stationary vehicles within 50 meters of the intersection in each lane at time t. Reference is made below to Figures 10A-11C An example describing queue length.
[0128] In some embodiments, the traffic environment data includes position data of each of the plurality of vehicles in the traffic environment. In some embodiments, the traffic environment data includes speed data of each of the plurality of vehicles in the traffic environment. For example, the traffic data included in the traffic environment data may include the position and speed of each vehicle within 50 meters of an intersection.
[0129] In some embodiments, the traffic environment data includes traffic density data for each of a plurality of areas of the traffic environment. In some embodiments, the traffic environment data includes traffic speed data for each of a plurality of areas of the traffic environment. For example, the traffic data included in the traffic environment data may include vehicle density (e.g., number of vehicles) and vehicle speed (e.g., average speed of each vehicle) for each of a plurality of area units within 50 meters of an intersection. Figure 5 Examples of such embodiments are described.
[0130] Figure 5 A top view of a traffic environment 500 at an intersection is shown, illustrating the segmentation of traffic lanes into cells. The northbound lane is segmented into northbound cells 502, the southbound lane is segmented into southbound cells 504, the eastbound lane is segmented into eastbound cells 506, and the westbound lane is segmented into westbound cells 508. Each cell 510 corresponds to a square area of road surface of the traffic environment. In this example, the traffic environment data provided to the RL model 240 may include traffic data indicating vehicle density data and vehicle speed data for each cell 510. The vehicle density data may be, for example, a count of the number of vehicles present within the cell at time t. The vehicle speed data may be, for example, an average or other aggregate measure of the speeds of vehicles present within the cell at time t. In some embodiments, each cell 510 may be the same width as a single lane. In some embodiments, the shape of the cells may be irregular: for example, each cell may be aligned with the directionality of the lane and may be the same width as the lane, but the length of the cell may be longer or shorter than its width.
[0131] It should be understood that Figure 5 The number of cells shown in the figures is not intended to be to scale: although each direction of traffic is divided into a matrix of cells seven cells wide, the figure is not necessarily intended to indicate that intersection 500 includes seven lanes of traffic in each direction. Rather, in each figure, the number and size of cells are arbitrary and do not necessarily correspond to the description of the corresponding traffic environment.
[0132] Figure 6 Show Figure 5 6 and 604 as traffic environment state data. The data from the cells in each lane direction 502, 504, 506, and 508 are concatenated into a two-dimensional matrix of vehicle speed data 602 and a two-dimensional matrix of vehicle density data 604, wherein the data from each set of lanes is represented as a horizontal band of the matrix. Thus, for example, the leftmost cell of the vehicle speed data matrix 602, starting from the top, may indicate the average speed of vehicles in each northbound lane immediately adjacent to the intersection (e.g., stopped at a red light adjacent to a pedestrian crossing), followed by the average speed of vehicles in each southbound lane immediately adjacent to the intersection, followed by each eastbound lane, and then each westbound lane.
[0133] The vehicle speed data matrix 602 and the vehicle density data matrix 604 may be converted into training data samples 606 before being provided to the RL model 240. In some embodiments, the training data samples 606 may be represented as two two-dimensional matrices, such as Figure 6 , and can be processed by RL model 240 using techniques for processing images with two channels. In other embodiments, training data samples 606 can be represented as a single two-dimensional matrix by concatenating or appending vehicle velocity data matrix 602 alongside vehicle density data matrix 604 to form a single larger two-dimensional matrix, which can be processed by RL model 240 using techniques for processing images with one channel.
[0134] In some embodiments, the traffic environment data may take other forms. For example, the traffic environment data may simply include the number of vehicles in each lane or each group of lanes in a given direction, and / or the aggregate speed measurement for each lane or each group of lanes in a given direction. It should be understood that other forms of traffic environment data may be used in different embodiments.
[0135] As described above, the traffic environment data may be generated by a traffic monitoring system based on sensor data collected from the traffic environment, and in some embodiments, the traffic monitoring system may be part of the controller device 220 .
[0136] Therefore, the traffic environment data provided as input to the reinforcement learning model 240 may represent the state of the traffic environment as of time t, thereby causing the reinforcement learning model 240 to generate traffic signal cycle data in response to the current traffic conditions.
[0137] Figure 8 An exemplary behavior module 244 of a traffic signal controller is shown receiving traffic environment state data (e.g., training data sample 606) at a single time point t. In response, the behavior module 244 generates traffic signal cycle data 804 as output using the strategy 802. The traffic signal cycle data 804 includes eight natural numbers p1 to p8, each of which represents Figure 1 The phase durations of the phases in the eight-phase traffic signal cycle 100 (P=8).
[0138] However, while information representing the state of the traffic environment at a single time t may be sufficient for a second-level controller or for some cycle-level or phase-level controllers, some embodiments may use traffic environment data representing the state of the traffic environment at more than one point in time. Data corresponding to each of the multiple points in time may be provided to the RL model 240 (e.g., to the behavior module 244), and the RL model 240 may generate traffic signal cycle data in response to receiving the data.
[0139] Figure 7A and 7B The potential importance of using historical traffic environment data (ie, data from one or more past times before time t) to generate traffic signal cycle data is shown. Figure 7A The traffic environment 700 is shown at the end of phase 4 108 of the traffic signal cycle 100, showing different stationary vehicle queue lengths for the blacked-out cells in each lane for different traffic directions. Specifically, at the end of phase 4, northbound cell 702 and southbound cell 704 may exhibit short stationary vehicle queues because vehicles in these lanes were able to move freely through the intersection during phase 4. However, eastbound cell 706 and westbound cell 708 may exhibit longer stationary vehicle queues because these vehicles were unable to proceed through the intersection during phase 4. Similarly, Figure 7B The traffic environment 750 at the end of phase 8 116 of the traffic signal cycle 100 is shown, showing different vehicle queue lengths for different traffic directions. Specifically, at the end of phase 8, northbound unit 752 and southbound unit 754 may exhibit long stationary vehicle queues because vehicles in these lanes were unable to proceed through the intersection during phase 8. However, eastbound unit 756 and westbound unit 758 may exhibit shorter stationary vehicle queues because these vehicles were able to move freely through the intersection during phase 8.
[0140] If the RL model 240 makes a decision at the end of stage 4 108 and thus receives a response corresponding to Figure 7A If the traffic environment data of the traffic environment 700 is obtained, the traffic signal cycle generated by the RL model may try to prioritize phases that will relieve the long queues of the eastbound unit 706 and the westbound unit 708, such as phase 5 110 to phase 8 116 (i.e., phases 5 110 to 8 will all have relatively long phase durations), while phases that relieve the northbound and southbound queues will be deprioritized (i.e., phases 1102 to 4 108 will all have relatively short phase durations). In contrast, if the RL model 240 makes a decision at the end of phase 8 116 and thus receives a signal corresponding to Figure 7B Based on the traffic environment data of the traffic environment 750, the traffic signal cycle generated by the RL model may attempt to prioritize phases that will alleviate the long queues of the northbound unit 702 and the southbound unit 704, while deprioritizing phases in phases that will alleviate the eastbound and westbound queues.
[0141] Thus, some embodiments may provide traffic environment data corresponding to one or more past times corresponding to one or more past phases of the cycle as input to the RL model 240. In some embodiments, traffic environment data is provided for points within each phase of the entire cycle.
[0142] Figure 9 An exemplary behavior module 244 of a traffic signal controller is shown, showing traffic environment state data (e.g., training data samples) at multiple time points as input. A first training data sample 902 corresponds to the traffic environment state at time t (e.g., the current time), a second training data sample 904 corresponds to the traffic environment state at time t-1 (e.g., the time during the previous phase), and so on, until the Tth training data sample 906 corresponds to the traffic environment state at time tT (e.g., the time during the phase T phases before, where T may be equal to P or P-1 in some embodiments).
[0143] In this example, the behavior module 244 also receives traffic signal phase data input indicating the current phase of the traffic signal cycle 910 and the time elapsed during the current phase 912. These additional inputs 910, 912 are used to place the current time t and its corresponding traffic environment state within the traffic signal cycle.
[0144] According to the techniques described above in the exemplary reinforcement learning model section, the behavior module 244 uses the policy 902 to generate traffic cycle data 804 based on the training data samples 902 , 904 , 906 .
[0145] It should be understood that in various embodiments, the historical traffic environment data (e.g., training data samples 904 to 906) may correspond to multiple time points within a given phase, time points spanning more than one cycle, or other distributions of historical state data. For example, each time interval of 1 unit (i.e., the time between t-1 and t) may correspond to the duration of a particular phase of a traffic signal cycle, or it may correspond to a fixed duration, such as one second.
[0146] Different embodiments may use different methods to select historical traffic state data for use as RL model input. In some embodiments, each past time point (i.e., t-1 to tT) may correspond to the time when the queue length of the traffic environment reached a local maximum or local minimum. Other methods may approximate these local minima and local maxima based on the time when the traffic signal cycle transitions from one phase to the next.
[0147] Figures 10A to 10D Various methods of selecting past time points (ie, t-1 to tT) to select traffic environment data as inputs to the RL model 240 are shown.
[0148] Figure 10A Shown in Figure 5 Example traffic environment 500 southbound traffic at Figure 1 FIG2 is a graph of vehicle queue length 1002 as a function of time 1004 during the first four phases 102, 104, 106, 108 of a traffic signal cycle 100. It can be observed that the southbound queue length 1014 (e.g., the longest queue length in each southbound lane) increases during phase 1 102 until reaching a local maximum 1016, and then decreases during phase 2 104 until reaching a local minimum 1018 of zero. During phase 3 106, the southbound queue length 1014 begins to increase again.
[0149] Figure 10B Show Figure 10A The queue length 1014 in FIG. 1 can be approximated as a linear interpolation between local maxima (e.g., maximum 1016) and local minima (e.g., minimum 1018). Furthermore, these local maxima and local minima are likely located at or near times corresponding to phase transitions within the traffic signal cycle. For example, the southbound queue length 1014 begins to decrease shortly after the transition 1022 between phase 102 and phase 2 104, and begins to increase again shortly after the transition 1024 between phase 2 104 and phase 3 106. Therefore, a linear interpolation 1022 can be drawn between the queue lengths at these phase transition times (e.g., 1022 and 1024) to approximate the graph 1014 of the actual queue length.
[0150] Figure 10CA linear interpolation of southbound queue length 1022 and a linear interpolation of northbound queue length 1040 are shown. Queue lengths at phase transition times are shown as circles: an estimated maximum southbound queue length 1032 at the transition from phase 1 102 to phase 2 104; an estimated minimum southbound queue length 1034 and an estimated maximum northbound queue length 1036 at the transition from phase 2 104 to phase 3 106; and an estimated minimum northbound queue length 1038 at the transition from phase 3 106 to phase 4 108.
[0151] Figure 10D Shown in Figure 10C The first training data sample 1052 represents the traffic environment state data from the transition 1022 between stage 1 102 and stage 2 104; the second training data sample 1054 represents the traffic environment state data from the transition 1024 between stage 2 104 and stage 3 106; and the third training data sample 1056 represents the traffic environment state data from the transition between stage 3 106 and stage 4 108. Each training data sample 1052, 1054, 1056 can be referred to as above. Figure 6 Generated as described.
[0152] Therefore, in some embodiments, past training data samples are selected from time points corresponding to one or more queue peak times and one or more queue valley times. Each queue peak time is a time when the length of a queue is at a local maximum, and each queue valley time is a time when the length of a queue is at a local minimum. In other embodiments, past training data samples correspond to one or more phase transition times. Each phase transition time is the time when a traffic signal transitions between two phases of a traffic signal cycle.
[0153] According to the above reference Figures 10A-10D The described phase transition times are based on estimating queue lengths based on the dynamics of traffic accumulation and release following a semi-linear pattern. Using this assumption, some embodiments can significantly reduce the input size of the controller relative to embodiments using historical data every second or every s seconds (s<<[p1...p8]), while providing the controller with sufficient information to make decisions based on traffic conditions within one or more traffic signal cycles.
[0154] In some reference Figures 10A-10D In the embodiment of the described historical traffic environment data, the state spaces are stacked together (eg Figure 9Traffic environment data at a critical point of at least one cycle of the traffic signal cycle (as shown in the training data samples 902 to 906 in FIG), and traffic signal phase data input indicating the current phase 910 of the traffic signal cycle and the time 912 elapsed during the current phase.
[0155] Example Reward Function
[0156] Different embodiments may use different reward functions.The reward function may use a traffic flow metric or performance indicator that aims to achieve some optimal result.
[0157] In some embodiments, the reward is based on the negative average number of vehicles that stopped (i.e., stationary) in the traffic environment during the previous cycle. If the speed of a vehicle (e.g., the magnitude of its velocity vector, or the scalar projection of its velocity vector on the traffic directionality of its lane or area) is below a speed threshold, the vehicle may be considered stationary. In some examples, a speed threshold of 2 meters per second may be used. The total delay consumed at the intersection during the cycle may be calculated by summing the delay of each vehicle present in the traffic environment during the cycle (i.e., the time spent stationary). In an embodiment using area-based speed and density data, an aggregate measure of total delay may be calculated by treating any cell with a vehicle speed below the speed threshold as a stationary cell and treating the stationary cell as representing a number of stationary vehicles equal to the number of vehicles present in the cell. Other methods may be used to calculate the total delay for a traffic signal cycle.
[0158] Figures 11A-11C Three methods of calculating the total delay of vehicles present in traffic lanes within a traffic environment are shown.
[0159] Figure 11AA diagram 1100 showing the distance from an intersection 1102 over time 1104 illustrates the position of stationary vehicles in a single traffic lane. During the first phase of a traffic signal cycle 1118 corresponding to traffic in the red light facing lane, the length of the queue increases. At a first time near the start of the first cycle, the queue has a first length 1112, indicating several vehicles that have stopped at the intersection boundary (e.g., at a crosswalk). At a second time later, the queue has a second length 1114, indicating several more vehicles that have stopped behind the previous vehicles. This continues in the first phase 1118 until the traffic signal cycle transitions to the second phase 1119. Once the traffic signal cycle enters the second phase 1119 corresponding to traffic in the green light facing lane, the queue begins to shorten as vehicles can enter the intersection. The vehicles closest to the intersection begin to move first, while vehicles farther from the intersection in the stationary queue have no chance to begin moving until the vehicle ahead of them begins to move; therefore, the position of the front of the queue shifts away from the intersection. Finally, at the last time during the second phase 1119, the queue has a final length 1116, after which the last vehicle in the queue begins to move and the queue length becomes zero.
[0160] It should be appreciated that the various queue lengths 1112 through 1116 may be defined by a generally triangular shape (shown in dashed outline).
[0161] Depend on Figure 11A The total delay represented by the stationary vehicles shown in can be calculated by summing the various queue lengths 1112 to 1116 at each point in time, such as every second. This calculation yields the area between the two lines indicating the front and end of the queue.
[0162] Figure 11B A second diagram 1120 is shown, which uses Figure 11A Graph 1100 shows the same traffic environment data but illustrates the platoon over time by showing the duration that each vehicle is stationary. Thus, a first vehicle is stationary for a first duration 1122, a second vehicle behind the first vehicle is stationary for a second duration 1124, and so on until a final vehicle is stationary for a final duration 1126.
[0163] exist Figure 11B The total delay represented by the stationary vehicles can be calculated by summing the various durations 1122 to 1126 for each vehicle. Thus, the cumulative delay for a cycle consisting of the first phase 1118 and the second phase 1119 is is the sum of the delays of all vehicles in the queue during the period Where d1 is the first duration 1122, d2 is the second duration 1124, d k The final duration is 1126. Figure 11AThis sum is calculated as the area between the two lines indicating the front and end of the queue. Another method.
[0164] Figure 11C Show Figures 11A-11B Same illustration of queue length and stationary vehicle positions for , but here the calculation of the total delay is performed directly by calculating the area of triangle 1132, i.e. the area between the two lines indicating the front and end of the queue
[0165] It should be understood that in some embodiments, Figures 11A-11C The calculations described can be performed using regional data rather than individual vehicle data, such as from Figure 5 Vehicle density and speed data for each unit.
[0166] Once the area of triangle 1132 has been determined by the above reference Figures 11A-11C If calculated using one of the methods described above, this area can be used to determine the cumulative delay of a cycle. Thus, the average number of vehicles stopped during the cycle length (i.e., the duration of the first phase 1118 and the second phase 1119) indicates the average delay faced by each vehicle during that cycle. The average delay over the cycle length can be calculated and used as a reward to discourage the controller from choosing short cycles to avoid larger penalties.
[0167] Thus, in some embodiments, a controller determines a state of a traffic environment at least in part by determining the length of each of one or more queues of stationary vehicles in the traffic environment, wherein the length indicates the number of stationary vehicles in the queue. This state data is used to generate training data samples. In some embodiments, a reward function is applied to the initial state of the traffic environment and the updated state of the traffic environment to calculate a reward based on an estimated number of stationary vehicles in the traffic environment during a previous traffic signal cycle.
[0168] The algorithm for calculating rewards based on this technique can be expressed in pseudocode as:
[0169] Status = null
[0170] Reward = 0
[0171] For each t in period_length:
[0172] Reward + = Number of vehicles stopped
[0173] If the signal has just turned yellow:
[0174] status = append(current_traffic_status)
[0175] Reward / =cycle_length
[0176] In this example, the number of stopped vehicles recorded at each time point t during the cycle is summed together and then divided by the cycle duration (i.e., cycle_length) to produce a final value for the reward (where a high reward indicates poor performance and a low reward indicates high performance). Each time the light turns yellow, the state of the traffic environment is sampled (e.g., to generate additional training data samples).
[0177] It will be appreciated that some embodiments may determine rewards using different performance metrics such as total throughput (number of vehicles passing through the intersection per cycle), the longest single delay of a single vehicle in one or more cycles, or any other suitable metric.
[0178] Exemplary system for controlling traffic signals
[0179] Once the RL model 240 is trained as described above, the controller device 220 can be deployed for use in controlling a real traffic signal in a real traffic environment. When deployed for controlling an actual traffic signal, the operation of the RL model 240 and other components described above is very similar to that described with reference to the training method 400. However, references to "training data samples" should be understood to refer to traffic environment state data, as they are not primarily used for training purposes. When deployed to control a real traffic signal, the controller device 220 constitutes a system for generating traffic signal cycle data. The controller device 220 includes a reference Figure 3 The components described include a processor device 225 and a memory 228. The RL model 240 stored in the memory 228 is now a trained RL model 240 that has been trained according to one or more of the techniques described above. The traffic environment used to train the reinforcement learning model is the same real traffic environment that is now being controlled, or a simulated version thereof, as described above. Instructions 238, when executed by the processor device 225, cause the system to perform steps very similar to those of method 400. As described above, traffic environment state data indicating the state of the real traffic environment is received from the traffic monitoring system (in some embodiments, it may be identical in format and content to the training data samples 606). The reinforcement learning model 240 is used to generate traffic signal cycle data by applying a policy (e.g., policy 802 or policy 902) to at least the traffic environment state data. The controller device 220 then sends the traffic signal cycle data to the traffic control system. The traffic controller system may be part of the controller device 220, or may be separate and communicate with the controller device 220, for example, via a network. The traffic controller system controls traffic signals (eg, lights 202 , 204 , 206 , 208 ) to execute traffic signal cycles according to traffic signal cycle data.
[0180] Thus, using the embodiments described herein, reinforcement learning can be used to implement a cycle-level traffic signal controller with an accuracy of one second or better. Some embodiments can use proximal policy optimization to achieve second-level or better accuracy in their outputs. The embodiments described herein can use a state-space definition that is concise yet captures all the necessary information to control traffic signals on a cycle-level basis. A reward function can be used to minimize the average vehicle delay of a cycle-level traffic signal controller at a signalized intersection.
[0181] Overview
[0182] Although the present disclosure describes methods and processes by steps in a certain order, one or more steps in the methods and processes may be omitted or changed as appropriate. Where appropriate, one or more steps may be performed in an order other than the order described.
[0183] Although the present disclosure is described at least in part in terms of methods, it will be understood by those skilled in the art that the present disclosure also relates to various components for performing at least some aspects and features of the methods, whether through hardware components, software, or any combination thereof. Accordingly, the technical solutions of the present disclosure may be embodied in the form of software products. Suitable software products may be stored in pre-recorded storage devices or other similar non-volatile or non-transitory computer-readable media, including DVDs, CD-ROMs, USB flash drives, removable hard drives, or other storage media. The software product includes instructions tangibly stored thereon, which enable a processing device (e.g., a personal computer, a server, or a network device) to perform examples of the methods disclosed herein.
[0184] The present disclosure may be embodied in other specific forms without departing from the subject matter of the claims. The exemplary embodiments described are intended to be illustrative in all respects and not restrictive. Selected features from one or more of the above-described embodiments may be combined to create alternative embodiments not explicitly described, and features suitable for such combinations are understood to be within the scope of the present disclosure.
[0185] All values and subranges within the disclosed ranges are also disclosed. Furthermore, although the systems, devices, and processes disclosed and illustrated herein may include a specific number of elements / components, these systems, devices, and components may be modified to include more or fewer of such elements / components. For example, although any disclosed element / component may be referred to in the singular, the embodiments disclosed herein may be modified to include a plurality of such elements / components. The subject matter described herein is intended to cover and encompass all suitable technical variations.
Claims
1. A method for training a reinforcement learning model to generate traffic signal cycle data, characterized in that The method comprises: The training data samples indicating the initial state of the traffic environment affected by the traffic signal are processed in the following manner: generating traffic signal cycle data using the reinforcement learning model by applying a policy to the training data sample and one or more past training data samples, the traffic signal cycle data comprising one or more phase durations for one or more corresponding phases of a traffic signal cycle, each phase duration being a value selected from a continuous range of values; determining an updated state of the traffic environment after applying the generated traffic signal cycle data to the traffic signal; generating a reward by applying a reward function to the initial state of the traffic environment and the updated state of the traffic environment; adjusting the strategy based on the reward; Repeating the step of processing training data samples one or more times, said training data samples indicating said updated state of said traffic environment; The reinforcement learning model is a proximal policy optimization (PPO) model; the strategy is a behavioral strategy; and the reward function is a judgment reward function.
2. The method according to claim 1, wherein: The traffic environment is a simulated traffic environment; and The traffic signal is a simulated traffic signal.
3. The method according to claim 1 or 2, characterized in that The one or more phase durations include a phase duration for each phase of at least one cycle of the traffic signal.
4. The method according to claim 1, wherein The one or more phase durations consist of a phase duration of one phase of a cycle of the traffic signal.
5. The method according to claim 1, wherein Each training data sample includes traffic data, including position data and speed data of each vehicle among a plurality of vehicles in the traffic environment.
6. The method according to claim 1, characterized in that Each training data sample includes traffic data, including traffic density data and traffic speed data for each of a plurality of areas of the traffic environment.
7. The method according to claim 1, wherein: Determining an updated state of the traffic environment includes determining a length of each of one or more queues of stationary vehicles in the traffic environment, the length indicating a number of stationary vehicles in the queue; and The one or more past training data samples include: one or more past training data samples corresponding to one or more queue peak times, each queue peak time being a time when a length of one of the queues is at a local maximum; One or more past training data samples corresponding to one or more queue valley times, each queue valley time being a time when a length of one of the queues is at a local minimum.
8. The method according to claim 1, characterized in that The one or more past training data samples correspond to one or more phase transition times, each phase transition time being a time when the traffic signal transitions between two phases of the traffic signal cycle.
9. The method according to claim 1, characterized in that The reward function is applied to the initial state of the traffic environment and the updated state of the traffic environment to calculate the reward based on an estimated number of stationary vehicles in the traffic environment during a previous traffic signal period.
10. The method according to claim 9, characterized in that The one or more past training data samples correspond to one or more phase transition times, each phase transition time being a time when the traffic signal transitions between two phases of the traffic signal cycle.
11. The method according to claim 1, wherein Each training data sample includes traffic signal phase data, wherein the traffic signal phase data indicates: the current phase of the traffic signal cycle; and The time that has elapsed during the current phase.
12. The method according to claim 1, wherein: The one or more phase durations include a phase duration for each phase of at least one cycle of the traffic signal; The reinforcement learning model is a proximal policy optimization (PPO) behavior-critic model; The strategy described is a behavioral strategy; The reward function is a judgment reward function; Each training data sample consists of: Traffic signal phase data, which indicates: the current phase of the traffic signal cycle; and The time that has elapsed during the current phase; traffic data including traffic density data and traffic speed data for each of a plurality of areas of the traffic environment; The reward function is applied to the initial state of the traffic environment and the updated state of the traffic environment to calculate the reward based on an estimated number of stationary vehicles in the traffic environment during a previous traffic signal cycle; and The one or more past training data samples correspond to one or more phase transition times, each phase transition time being a time when the traffic signal transitions between two phases of the traffic signal cycle.
13. A system for training a reinforcement learning model to generate traffic signal cycle data, characterized in that include: Processor device; and Memory, which stores: the reinforcement learning model; machine-executable instructions thereon, which, when executed by the processor device, cause the system to: The training data samples indicating the initial state of the traffic environment affected by the traffic signal are processed in the following manner: generating traffic signal cycle data using the reinforcement learning model by applying a policy to the training data sample and one or more past training data samples, the traffic signal cycle data comprising one or more phase durations for one or more corresponding phases of a traffic signal cycle, each phase duration being a value selected from a continuous range of values; determining an updated state of the traffic environment after applying the generated traffic signal cycle data to the traffic signal; generating a reward by applying a reward function to the initial state of the traffic environment and the updated state of the traffic environment; adjusting the strategy based on the reward; Repeating the step of processing training data samples one or more times, wherein the training data samples indicate the updated state of the traffic environment; wherein the reinforcement learning model is a proximal policy optimization (PPO) behavior-critic model; and wherein the strategy is a behavior strategy; And the reward function is a judgment reward function.
14. The system according to claim 13, wherein: The reward function is applied to the initial state of the traffic environment and the updated state of the traffic environment to calculate the reward based on an estimated number of stationary vehicles in the traffic environment during a previous traffic signal cycle; and The one or more past training data samples correspond to one or more phase transition times, each phase transition time being a time when the traffic signal transitions between two phases of the traffic signal cycle.
15. A system for generating traffic signal cycle data, characterized in that include: Processor device; and Memory, which stores: A trained reinforcement learning model trained according to the method of claim 1; and machine-executable instructions that, when executed by the processor device, cause the system to: receiving traffic environment state data indicating a state of a real traffic environment from a traffic monitoring system, wherein the traffic environment used to train the reinforcement learning model is the real traffic environment or a simulated version thereof; generating traffic signal cycle data using the reinforcement learning model by applying the policy to at least the traffic environment state data; The traffic signal cycle data is sent to a traffic control system.
16. A non-transitory processor-readable medium, characterized in that The non-transitory processor-readable medium stores a trained reinforcement learning model trained according to the method of claim 1.
17. A non-transitory processor-readable medium, characterized in that The non-transitory processor-readable medium has machine-executable instructions stored thereon, which, when executed by a processor device, cause the processor device to perform the method of claim 1 .
Citation Information
Patent Citations
Picture self-leaning traffic signal control method based on convolutional neural network
CN109410608A
Signal lamp control method, and related equipment and system
CN110114806A