Hybrid reinforcement learning traffic signal hierarchical control method
Through hybrid reinforcement learning of traffic signal layered control methods, combining the underlying maximum pressure control and global proximity strategy optimization, the intersection signal lights are dynamically adjusted, which solves the problem of efficient traffic flow control at multiple intersections in the urban traffic network, and achieves the improvement of traffic efficiency and safety.
Patent Information
- Application Number
- CN202510633514.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to achieve efficient global traffic flow control in urban traffic networks, especially when facing dynamic and unpredictable traffic patterns, fixed time control and maximum pressure control have limitations, and it is impossible to effectively coordinate traffic signals at multiple intersections.
The hybrid reinforcement learning traffic signal layered control method is adopted, and the queue length and phase pressure values of the intersection are collected, and the global traffic data is encoded in combination with the graph neural network. The underlying maximum pressure controller and the global proximity strategy optimization controller work together, and the signal timing strategy is dynamically adjusted to optimize traffic flow.
It realizes efficient traffic flow under dynamic traffic conditions, reduces queue length and vehicle waiting time, and improves traffic efficiency and safety of the entire network.
Smart Images

Figure CN120356351A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic control, and relates to, but is not limited to, a hierarchical control method for traffic signals based on hybrid reinforcement learning. Background Art
[0002] Urban traffic congestion is an increasingly serious global challenge, leading to increased travel time, energy consumption, and environmental pollution. Due to the increasing complexity of urban traffic networks, urban traffic requires an adaptive signal control mechanism that can cope with dynamic and unpredictable traffic patterns. Although a large amount of research and progress has been made in adaptive traffic signal control, there are still some challenges in developing efficient traffic signal control methods. These challenges include the accurate representation of traffic conditions, the adaptation to fluctuating traffic demands, and the consideration of different traffic densities and lane occupancies at intersections.
[0003] In related technologies, fixed-time control (FTC) relies on a preset timing plan and is difficult to cope with the dynamic changes of traffic flow; maximum pressure (MP) control achieves distributed decision-making through local queue pressure, but lacks global coordination ability across intersections and is prone to regional suboptimal solutions.
[0004] Therefore, how to improve the efficient traffic flow of the entire network has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a hierarchical control method for traffic signals based on hybrid reinforcement learning, which at least solves the problem that the related technology cannot ensure the efficient traffic flow of the entire network.
[0006] According to the first aspect of the embodiment of the present invention, a hierarchical control method for traffic signals based on hybrid reinforcement learning is provided, including:
[0007] Collect the first queue length corresponding to the upstream lane and the second queue length corresponding to the downstream lane of each intersection, and obtain the phase pressure value through the weighted sum of the first queue length and the second queue length;
[0008] Combine the first queue length, the weighted sum, and the phase pressure value to obtain local traffic data;
[0009] When no congestion or emergency vehicle is detected, use the underlying maximum pressure controller to obtain the optimal phase corresponding to each intersection through the maximum pressure algorithm and the local traffic data; and control the traffic signals of each intersection based on the optimal phase, and each intersection corresponds to a underlying maximum pressure controller;
[0010] Integrate the regional queue matrix, traffic speed vector, historical congestion pattern matrix, emergency vehicle position vector and road network topology adjacency matrix to obtain global traffic data;
[0011] The global traffic data is encoded by a graph neural network to obtain encoded features; and a global proximal strategy optimization controller is used to learn a global optimal strategy through the features;
[0012] When a congested or emergency vehicle is detected, the global proximal strategy optimization controller is used to perform queue weight correction on the current phase through the global optimal strategy, generate a green wave coordination offset and / or emergency phase coverage processing, and obtain a target optimal phase, and each intersection corresponds to a global proximal strategy optimization controller;
[0013] The target optimal phase corresponding to each of the intersections is issued, and the traffic signals of each intersection are collaboratively controlled through the target optimal phase.
[0014] According to a second aspect of an embodiment of the present invention, there is provided an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.
[0015] According to a third aspect of an embodiment of the present invention, there is provided a computer storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.
[0016] According to the solution provided by the embodiments of the present invention, the first queue length corresponding to the upstream lane of each intersection and the second queue length corresponding to the downstream lane are collected, and the phase pressure value is obtained through the weighted sum of the first queue length and the second queue length; the first queue length, the weighted sum, and the phase pressure value are combined to obtain local traffic data; when no congestion or emergency vehicle is detected, the underlying maximum pressure controller uses the maximum pressure algorithm and the local traffic data to obtain the optimal phase corresponding to each intersection; and based on the optimal phase, the traffic signals of each intersection are controlled, and each intersection corresponds to an underlying maximum pressure controller; the regional queue matrix, the traffic flow speed vector, the historical congestion pattern matrix, the emergency vehicle position vector, and the road network topology adjacency matrix are integrated to obtain global traffic data; the global traffic data is encoded through a graph neural network to obtain the encoded features; and the global proximal policy optimization controller is used to learn the global optimal policy through the features; when congestion or an emergency vehicle is detected, the global proximal policy optimization controller uses the global optimal policy to perform queue weight correction, generate a green wave coordination offset, and / or perform emergency phase coverage processing on the current phase to obtain the target optimal phase, and all intersections correspond to a global proximal policy optimization controller; the target optimal phase corresponding to each intersection is sent to each intersection, and the traffic signals of each intersection are cooperatively controlled through the target optimal phase. During the process, the underlying maximum pressure controller uses the maximum pressure algorithm to make local real-time decisions at each intersection, responsible for local real-time phase selection; the global proximal policy optimization controller uses the proximal policy optimization algorithm to learn the global policy to dynamically adjust the maximum pressure parameter (MP parameter), thereby coordinating the signal timing strategies of each intersection. This combination utilizes the local efficiency of the maximum pressure algorithm and the global adaptability of the proximal policy optimization algorithm, and can effectively handle dynamic traffic conditions. The reward function aims to minimize the queue length and vehicle waiting time and ensure efficient traffic flow throughout the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0018] Figure 1 It is a schematic flowchart of a hybrid reinforcement learning traffic signal hierarchical control method provided by the embodiments of the present invention;
[0019] Figure 2 It is a schematic diagram of the effect of road network pressure provided by the embodiments of the present invention;
[0020] Figure 3 Schematic diagram of the effect of hierarchical control of traffic signals by hybrid reinforcement learning provided by an embodiment of the present invention;
[0021] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0024] It should be noted that the terms "first \ second \ third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first \ second \ third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.
[0025] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the technical field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0026] Figure 1 Schematic flow diagram of a method for hierarchical control of traffic signals by hybrid reinforcement learning provided by an embodiment of the present invention. The method for hierarchical control of traffic signals by hybrid reinforcement learning provided by the embodiments of the present invention can be executed by an electronic device, which can be, for example, a computer, a server, etc.
[0027] As Figure 1 shown, the method for hierarchical control of traffic signals by hybrid reinforcement learning includes:
[0028] S101. Collect the first queue length corresponding to the upstream lane and the second queue length corresponding to the downstream lane at each intersection, and obtain the phase pressure value through the weighted sum of the first queue length and the second queue length.
[0029] In the embodiments of the present invention, the upstream lane refers to all road segments entering a certain intersection or section, which is the road part passed by vehicles before reaching the intersection or section. The downstream lane refers to all road segments leaving a certain intersection or section and heading to the next destination, which is the road part where vehicles continue to drive after passing through the intersection. The queue length is the number of vehicles waiting to pass at the intersection or section, used to measure the degree of traffic congestion. The traffic data of the intersection can be collected in real time through traffic environment sensors. Among them, the traffic data includes the queue length x(l,m) from the upstream lane to the downstream lane, and the weighted sum corresponding to the queue length of the downstream lane is calculated. Further calculate the phase pressure value w(l,m)=x(l,m)-∑ p r(m,p)x(m,p), where r(m,p) is the fixed turning ratio from the upstream lane m to the downstream lane p, p is an element of the downstream lane set, and x(m,p) is the queue length from the upstream lane m to the downstream lane, that is, the second queue length.
[0030] S102. Combine the first queue length, the weighted sum, and the phase pressure value to obtain local traffic data.
[0031] In the embodiments of the present invention, the local traffic data is data reflecting the current local traffic conditions. Combine the first queue length, the weighted sum corresponding to the second queue length, and the phase pressure value to obtain local traffic data.
[0032] S103. When no congestion or emergency vehicle is detected, use the underlying maximum pressure controller to obtain the optimal phase corresponding to each intersection through the maximum pressure algorithm and the local traffic data; and control the traffic signals of each intersection based on the optimal phase. Each intersection corresponds to an underlying maximum pressure controller.
[0033] In an embodiment of the present invention, the optimal phase refers to the signal light phase configuration calculated according to the current traffic conditions, which can optimize traffic flow to the greatest extent. This configuration not only considers the immediate traffic demand but may also incorporate a prediction model to estimate the best strategy for a period of time in the future. Implementing these phases can effectively reduce vehicle queuing time, alleviate congestion, and improve road capacity. Each intersection corresponds to a bottom-layer maximum pressure controller (bottom-layer MP controller), and the bottom-layer maximum pressure controller is the local maximum pressure controller (each local maximum pressure controller controls one intersection). Each bottom-layer maximum pressure controller is responsible for managing and optimizing the switching of traffic lights in all directions and phases (such as straight, left turn, right turn, etc.) at its corresponding intersection to improve traffic efficiency, reduce congestion, and enhance traffic safety. Therefore, when no congestion or emergency vehicle is detected, the bottom-layer maximum pressure controller uses the maximum pressure algorithm and local traffic data to obtain the optimal phase corresponding to each intersection, and finally adjusts and manages the behavior of traffic lights at each intersection based on the optimal phase. Among them, controlling traffic signals includes signal light time allocation, phase switching, and priority for emergency vehicles to coordinate upstream and downstream intersections. Among them:
[0034]
[0035] In the above formula, c(l,m) is the saturation flow rate of lane (l,m), S n is the set of permitted phases of intersection n. Each phase in the phase set corresponds to a phase pressure value representing the current demand intensity of this phase, is the optimal phase.
[0036] S104. Integrate the regional queue matrix, vehicle flow speed vector, historical congestion pattern matrix, emergency vehicle position vector, and road network topological adjacency matrix to obtain global traffic data.
[0037] In an embodiment of the present invention, the queue matrix is a combination of queues at each intersection; the traffic flow speed vector represents the driving speeds of vehicles in different directions, which can help analyze traffic flow and identify potential bottleneck locations. The historical congestion pattern is based on past data to identify traffic congestion patterns that occur during specific time periods, dates, or conditions, which is very useful for predicting future traffic flow and planning traffic management strategies. The emergency vehicle location refers to the location information of special vehicles such as ambulances and fire trucks, which is crucial for timely adjusting signal settings to ensure that emergency vehicles can pass quickly. The road network topology adjacency matrix is a mathematical representation used to describe the connection relationships between various intersections (nodes) in the road network. This matrix can be used to calculate the shortest path, evaluate traffic mobility, etc. Integrating the queue matrix, traffic flow speed vector, historical congestion pattern, emergency vehicle location, and road network topology adjacency matrix means combining all these different types of data to obtain global traffic data for more effective management and optimization of the entire traffic system.
[0038] S105. Encode the global traffic data through a graph neural network to obtain encoded features; and use the global proximal policy optimization controller to learn the global optimal policy through the features.
[0039] In an embodiment of the present invention, the global traffic data is input into a graph neural network for encoding to obtain high-dimensional (the dimension is greater than the preset dimension) features, and the global proximal policy optimization controller (global PPO controller) derives the best behavior policy applicable to the entire intersection, that is, the global optimal policy, based on the features corresponding to this global traffic data.
[0040] S106. When congestion or an emergency vehicle is detected, use the global proximal policy optimization controller to perform queue weight correction, generate a green wave coordination offset, and / or emergency phase coverage processing on the current phase through the global optimal policy to obtain the target optimal phase. Each intersection corresponds to one global proximal policy optimization controller.
[0041] In an embodiment of the present invention, the queue weight is corrected to adjust the importance weights of queues in each direction, the green wave coordination offset is to optimize the green wave bandwidth and reduce the number of vehicle stops, and the emergency phase coverage is to prioritize special requirements (such as ambulances, fire trucks, etc.). The target optimal phase can be considered from a global perspective to achieve regional-level coordinated optimization and avoid a decrease in global performance caused by single-point optimization. Each intersection corresponds to a global proximal policy optimization controller, and the global controller works in cooperation with each underlying maximum pressure controller to optimize the traffic flow of the entire region or network. The main responsibility of the global controller is to coordinate the signal control strategies between multiple intersections, thereby achieving smoother traffic flow and higher efficiency within a larger range. Therefore, when congestion or an emergency vehicle is detected, that is, when an element in the regional queue matrix exceeds the threshold or an emergency vehicle is detected, that is, when an emergency or severe congestion situation at the intersection is detected based on the queue matrix of each intersection, the global proximal policy optimization controller is used to perform queue weight correction, generate a green wave coordination offset, and / or emergency phase coverage processing on the current phase through the global optimal policy to obtain the target optimal phase. Finally, the traffic signals at each intersection are cooperatively controlled through the target optimal phase to solve situations such as lane congestion. Among them, the current phase is the phase corresponding to each intersection when congestion or an emergency occurs.
[0042] S107. Send the corresponding target optimal phase to each intersection and cooperatively control the traffic signals at each intersection through the target optimal phase.
[0043] In an embodiment of the present invention, after obtaining the target optimal phase corresponding to each intersection, the target optimal phase is sent to each intersection, so that each intersection controls its corresponding traffic signal through the target optimal phase, achieving the purpose of cooperatively controlling the traffic signals of the entire region.
[0044] It can be understood that, in the embodiments of the present invention, the first queue length corresponding to the upstream lane of each intersection and the second queue length corresponding to the downstream lane are collected, and the phase pressure value is obtained through the weighted sum of the first queue length and the second queue length; the first queue length, the weighted sum, and the phase pressure value are combined to obtain local traffic data; when no congestion or emergency vehicle is detected, the underlying maximum pressure controller uses the maximum pressure algorithm and the local traffic data to obtain the optimal phase corresponding to each intersection; and based on the optimal phase, the traffic signals of each intersection are controlled, and each intersection corresponds to an underlying maximum pressure controller; the regional queue matrix, the traffic flow speed vector, the historical congestion pattern matrix, the emergency vehicle position vector, and the road network topology adjacency matrix are integrated to obtain global traffic data; the global traffic data is encoded through a graph neural network to obtain the encoded features; and the global proximal policy optimization controller uses the features to learn the global optimal policy; when congestion or an emergency vehicle is detected, the global proximal policy optimization controller uses the global optimal policy to correct the queue weights of the current phase, generate a green wave coordination offset, and / or perform an emergency phase coverage process to obtain the target optimal phase, and all intersections correspond to a global proximal policy optimization controller; the target optimal phase corresponding to each intersection is sent down, and the traffic signals of each intersection are cooperatively controlled through the target optimal phase. In this process, a hierarchical hybrid control architecture is proposed, which combines the real-time performance of the maximum pressure control (MP) and the global optimization ability of the proximal policy optimization (PPO) algorithm of deep reinforcement learning to achieve efficient and stable multi-intersection cooperative control to solve the limitations of existing methods. The proposed framework operates in a hierarchical manner: the underlying maximum pressure controller makes local real-time decisions at each intersection and is responsible for local real-time phase selection; the global proximal policy optimization controller learns the global policy to dynamically adjust the maximum pressure parameters, thereby coordinating the signal timing strategies of multiple intersections. This combination utilizes the local efficiency of the underlying maximum pressure controller and the global adaptability of the global proximal policy optimization controller, enabling the system to effectively handle dynamic traffic conditions. The reward function aims to minimize the queue length and vehicle waiting time to ensure efficient traffic flow throughout the network.
[0045] In some embodiments of the present invention, after S103, it can also be implemented through S201, which is described through the following steps.
[0046] S201. Obtain the local reward of the underlying maximum pressure controller after controlling the traffic signals of each intersection with the optimal phase based on the first reward formula.
[0047] In an embodiment of the present invention, the local reward refers to an evaluation index designed for the performance of a single intersection such as each intersection. It is used to measure the behavior effect of the traffic control system within each intersection and serves as the basis for optimizing and adjusting the control strategy. Based on the first reward formula, the local reward of the underlying maximum pressure controller after using the optimal phase to control the traffic signals of each intersection can be obtained. The first reward formula is as follows:
[0048] R local =-∑ (l,m) x(l,m)+λ1∑ (l,m) (c(l,m)S(l,m)-d(l,m))
[0049] In the above formula, x(l,m) represents the queue length from the upstream lane l to the downstream lane m of each intersection. The queue length includes the first queue length and the second queue length. λ1 is a weight coefficient used to balance the queue penalty and the service efficiency reward, 0 < λ1 < 1. S(l,m) is a binary control variable indicating whether vehicles are allowed to turn from the upstream lane l to the downstream lane m within the decision period. c(l,m) represents the saturation flow rate of the lane (l,m), and d(l,m) represents the vehicle demand arriving at the lane (l,m) from outside, where
[0050] Among them, -∑ (l,m) x(l,m) is to minimize the overall queue length to avoid local or global congestion. The longer the queue, the greater the penalty. λ1∑ (l,m) (c(l,m)S(l,m)-d(l,m)) is to reward the service efficiency, that is, the difference between the actual number of serviced vehicles c(l,m)S(l,m) and the number of demanded vehicles d(l,m). When the number of serviced vehicles c(l,m)S(l,m) exceeds the demand d(l,m), it indicates that the resources are sufficient and a positive reward is given. If the service capacity is insufficient, c(l,m)S(l,m) < d(l,m), then a penalty is imposed to promote policy optimization.
[0051] In some embodiments of the present invention, the traffic signals of each intersection are coordinated and controlled through the target optimal phase in S106, which can be implemented through S301 and is described in the following steps.
[0052] S301. Obtain the global reward of the global proximal policy optimization controller after using the target phase to control the traffic signals of each intersection based on the second reward formula, and obtain the final reward based on the local reward and the global reward.
[0053] In some embodiments of the present invention, the global reward refers to a mechanism for evaluating the overall performance of multiple intersections. Different from focusing on the local rewards of individual intersections, the global reward aims to optimize the performance of the entire traffic network and ensure the maximization of the overall traffic efficiency. The global reward takes into account the state of the entire traffic network and focuses on achieving a dynamic balance among reducing travel time, increasing throughput, and alleviating congestion in regional traffic. Based on the second reward formula, the global rewards of each intersection after controlling the traffic signals of each intersection based on the target phase are obtained. Finally, the global reward and the local reward are summed to obtain the final reward. The second reward formula is as follows:
[0054] R global =-AvgTravelTime + λ2·Throughput - λ3·CongestionIndex
[0055] In the above formula, AvgTravelTime is the average travel time of all vehicles from the starting point to the end point at each intersection, and CongestionIndex is a comprehensive index reflecting the overall congestion degree of each intersection, usually calculated by weighting the congestion duration and the congestion ratio of the road section. If 30% of the road section is in a congested state and lasts for 20 minutes, then CongestionIndex = 0.3×20 = 6. λ2 and λ3 are used to adjust the importance of throughput and congestion index in the reward, and Throughput is the total number of vehicles passing through all intersections per unit time.
[0056] Furthermore, the final reward is obtained based on the local reward and the global reward. The calculation method is as follows:
[0057] R = αR local +(1 - α)R global
[0058] In the above formula, α is the adaptive weight, and R is the final reward.
[0059] In some embodiments of the present invention, the utilization of the global proximal policy optimization controller in S106 to correct the queue weights, generate the green wave coordination offset, and / or perform the emergency phase coverage processing on the current phase through the global optimal policy to obtain the target optimal phase can be achieved through S1061 to S1062, and the specific description is as follows.
[0060] S1061. Use the global proximal policy optimization controller to apply the proximal policy optimization algorithm to process the global optimal policy to obtain the continuous parameter adjustment amount and the phase selection probability.
[0061] In some embodiments of the present invention, the Proximal Policy Optimization (PPO) algorithm is responsible for learning how to act according to the current environmental state. The continuous parameter adjustment amount is a numerical value used to finely tune certain continuous variables or parameters. The phase selection probability is the likelihood or probability of selecting a specific phase (i.e., a specific signal light configuration at an intersection). The global proximal policy optimization controller applies the PPO algorithm to process the global optimal policy, obtaining the continuous parameter adjustment amount and the phase selection probability.
[0062] Among them, the policy network of the PPO algorithm uses two fully connected layers, with 512 nodes in each layer. The input is the global state encoded by the graph neural network, and the output is the continuous parameter adjustment amount and the offline phase selection probability. The value network (Critic) of the PPO algorithm: uses two fully connected layers, with 512 nodes in each layer. The input is the global state encoded by the graph neural network, and the output is the evaluation result of the state value. The loss function is The model is trained in stages. First, offline pre-training is carried out. Based on the SUMO simulation environment, maximum pressure baseline data is generated, and the parameters of the proximal policy optimization network are initialized. Secondly, online fine-tuning is carried out. After deployment, traffic data is collected in real time, and the policy is dynamically updated in combination with safety constraints (such as action space limitations) to ensure control stability.
[0063] S1062. Optimize the current phase based on the continuous parameter adjustment amount and the phase selection probability to obtain the target optimal phase. The optimization process includes queue weight correction, green wave coordination offset, and / or emergency phase coverage processing.
[0064] In some embodiments of the present invention, the current phase is optimized, such as queue weight correction, green wave coordination offset, and / or emergency phase coverage, through the calculated continuous parameter adjustment amount and phase selection probability to obtain the target optimal phase.
[0065] In some embodiments of the present invention, after the traffic signals of each intersection are cooperatively controlled by the target optimal phase in S106, it is also implemented through S401 to S404, which will be described in the following steps.
[0066] S401. Adjust the policy of the global proximal policy optimization controller through the final reward, and collect the new queue lengths corresponding to the upstream lanes to the downstream lanes of each intersection again until new optimal phase local traffic data and new global traffic data are obtained.
[0067] S402. When no congestion or emergency vehicle is detected, use the underlying maximum pressure controller to obtain the new optimal phase corresponding to each intersection through the maximum pressure algorithm and the new local traffic data; and control the traffic signals of the intersections based on the new optimal phase. Each intersection corresponds to an underlying maximum pressure controller;
[0068] S403, encoding the new global traffic data through a graph neural network to obtain encoded new features, and using a global proximal strategy optimization controller to learn a new global optimal strategy through the new features.
[0069] S404. When congestion or an emergency vehicle is detected, the global proximal strategy optimization controller is used to correct the queue weight of the new current phase through a new global optimal strategy, generate a green wave coordination offset and / or emergency phase coverage processing, and obtain a new target optimal phase. Each intersection corresponds to one global proximal strategy optimization controller.
[0070] S405: Send the corresponding new target optimal phase to each intersection, and coordinately control the traffic signals of each intersection through the new target optimal phase.
[0071] In an embodiment of the present invention, after obtaining the final reward, the global proximal policy optimization controller is adjusted by the final reward, that is, the policy network parameters in the global proximal policy optimization controller are updated with the feedback obtained from the environment, so as to make better decisions in the future, and obtain the adjusted global proximal policy optimization controller. The process from S101 to S107 is executed again, and the queue lengths of the upstream lanes and downstream lanes of each intersection are collected in real time until new local traffic data and new global traffic data are obtained. When no congestion or emergency vehicles are detected, the new local traffic data is processed by the bottom layer maximum pressure controller to obtain the new optimal phase corresponding to each intersection, and the traffic signals of each intersection are controlled based on the new optimal phase. The new global traffic data is encoded through the graph neural network to obtain the encoded new features, and the global proximal strategy optimization controller is used to learn the new global optimal strategy through the new features. When congestion or emergency vehicles are detected, the global proximal strategy optimization controller is used to correct the queue weight of the new current phase through the new global optimal strategy, generate green wave coordination offset and / or emergency phase coverage, and obtain the new target optimal phase. Each intersection corresponds to a global proximal strategy optimization controller; the corresponding new target optimal phase is issued to each intersection, and the traffic signals of each intersection are coordinated and controlled through the new target optimal phase. The adjusted global proximal strategy optimization controller is adjusted again with the new final reward, and then data collection is performed again, and the cycle continues.
[0072] In an embodiment of the present invention, a novel hybrid control framework for multi-intersection traffic signal control is proposed, which integrates the maximum pressure algorithm and the proximal policy optimization algorithm, called HybridMax-PPO, which demonstrates a scalable framework suitable for dynamic urban traffic environments and a hierarchical approach combining local stability with global adaptability for adaptive optimization of signal timing parameters.
[0073] In addition, the road network topological features encoded by the graph convolutional neural network effectively capture the spatial dependence relationships between intersections. The hybrid reward mechanism takes both local efficiency and global balance into account, and the safety constraints ensure the stability of the control process. The hierarchical structure of the present invention can be extended to large-scale urban road networks, providing an efficient and adaptive signal control solution for intelligent transportation systems.
[0074] An embodiment of the present invention provides a hierarchical control system for traffic signals based on hybrid reinforcement learning. The system includes a traffic environment perception module, a state encoding module, a hierarchical control module, a global proximal policy optimization and coordination module, a signal execution module, a reward calculation and feedback module, and a training and update module. Among them, the traffic environment perception module collects traffic data in real time through traffic environment sensors. The state encoding module calculates the phase pressure value of the downstream queue based on the queue length, further obtains local traffic data, integrates the queue matrix, vehicle flow speed vector, historical congestion pattern, emergency vehicle position, and road network topological adjacency matrix of each intersection to obtain global traffic data. The hierarchical control module is used to obtain the optimal phase corresponding to each intersection based on the maximum pressure algorithm and local traffic data, and encode the global traffic data through a graph neural network to obtain the encoded features. The global proximal policy optimization and coordination module learns the global optimal policy based on the encoded features and the proximal policy algorithm, and performs queue weight correction, green wave coordination offset, and / or emergency phase coverage processing on the current phase through the global optimal policy to obtain the target optimal phase, and controls the traffic signals of each intersection based on the target optimal phase. The reward calculation and feedback module is used to calculate the local reward and the global reward, and the training and update module is used to train the policy network and value network in the proximal policy optimization algorithm.
[0075] In summary, as Figure 3 shown, Figure 3 is a schematic diagram of the effect of the hierarchical control of traffic signals based on hybrid reinforcement learning provided by the embodiment of the present invention. Figure 3 shows how an intelligent transportation system optimizes road traffic management through data exchange between roadside units and cloud servers. In Figure 3It includes a global proximal policy optimization controller, network communication, an operation server, a traffic control center, a base station, a bottom-layer maximum pressure controller, a roadside unit, and an emergency vehicle. There are multiple roads in the figure, and each road represents a section of road or an intersection. There is a roadside unit beside each road, which includes a bottom-layer maximum pressure controller for collecting and transmitting road information. The yellow dotted line represents the data transmission path. Data is transmitted from the roadside unit to the global proximal policy optimization controller and then from the global proximal policy optimization controller to other roadside units. The global proximal policy optimization controller is responsible for coordinating and managing the operation of the intelligent transportation system in the entire area, such as adjusting the time of traffic lights, providing real-time traffic conditions, etc. It exchanges data with the operation server through network communication. The operation server is used to process and analyze the data collected from each roadside unit. The arrows of the network communication represent the data flow direction between different components, indicating that information flows bidirectionally, not only with data upload but also with instruction issuance. The traffic control center is the workstation of human operators, and they can monitor and control the operating status of the entire traffic system through this interface. There is two-way communication between the traffic control center and the operation server, allowing real-time data transmission and instruction issuance. The red line is a partial magnification of the global traffic, and the blue line can represent a specific type of communication or data flow. When no vehicle congestion or emergency is detected, the bottom-layer maximum pressure controller controls the traffic signals at each intersection based on local traffic data. When vehicle congestion or an emergency is detected, the global proximal policy optimization controller controls the traffic signals at each intersection based on the global traffic data integrated from local traffic data.
[0076] Referring to Figure 4 , a schematic structural diagram of an electronic device according to an embodiment of the present invention is shown. The specific implementation of the present invention does not limit the specific implementation of the electronic device.
[0077] As Figure 4 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.
[0078] Wherein:
[0079] The processor 502, the communication interface 504, and the memory 506 communicate with each other through the communication bus 508.
[0080] The communication interface 504 is used to communicate with other electronic devices or servers.
[0081] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above method embodiments.
[0082] Specifically, the program 510 may include program code that includes computer operation instructions.
[0083] The processor 502 may be a central processing unit (CPU), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0084] The memory 506 is used to store the program 510. The memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0085] The program 510 may specifically be used to cause the processor 502 to perform the operations corresponding to the methods described in the above method embodiments.
[0086] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and descriptions in the corresponding units in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.
[0087] It should be noted that according to the needs of implementation, the various components / steps described in the embodiments of the present invention may be split into more components / steps, or two or more components / steps or partial operations of components / steps may be combined into new components / steps to achieve the objectives of the embodiments of the present invention.
[0088] The method according to an embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as software or computer code stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or can be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored as such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0089] Those of ordinary skill in the art can realize that the units and method steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.
[0090] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention shall be defined by the claims.
Claims
1. A hierarchical control method for traffic signals based on hybrid reinforcement learning, characterized in that, Including: Collecting the first queue length corresponding to the upstream lane of each intersection and the second queue length corresponding to the downstream lane, and obtaining a phase pressure value through the weighted sum of the first queue length and the second queue length; Combining the first queue length, the weighted sum, and the phase pressure value to obtain local traffic data; When no congestion or emergency vehicle is detected, using the underlying maximum pressure controller to obtain the optimal phase corresponding to each intersection through the maximum pressure algorithm and the local traffic data; and controlling the traffic signals of the intersections based on the optimal phase, with each intersection corresponding to an underlying maximum pressure controller; Integrating the regional queue matrix, the traffic flow speed vector, the historical congestion pattern matrix, the emergency vehicle position vector, and the road network topology adjacency matrix to obtain global traffic data; Encoding the global traffic data through a graph neural network to obtain encoded features; And using the global proximal policy optimization controller to learn the global optimal policy through the features; When congestion or an emergency vehicle is detected, using the global proximal policy optimization controller to perform queue weight correction, generate a green wave coordination offset, and / or perform emergency phase coverage processing on the current phase through the global optimal policy to obtain the target optimal phase, with all intersections corresponding to a global proximal policy optimization controller; Sending the target optimal phase corresponding to each intersection to the intersections, and coordinately controlling the traffic signals of the intersections through the target optimal phase.
2. The method according to claim 1, wherein After controlling the traffic signals of the intersections based on the optimal phase, the method further includes: Obtaining the local reward of the underlying maximum pressure controller after controlling the traffic signals of the intersections based on the optimal phase according to the first reward formula, and the first reward formula is as follows: R local = -∑ (l,m) x(l, m) + λ1∑ (l,m) (c(l, m)S(l, m) - d(l, m)) In the above formula, x(l,m) represents the queue length from the upstream lane l to the downstream lane m of each intersection, the length includes the first queue length and the second queue length, λ1 is a weight coefficient for balancing queue penalty and service efficiency reward, S(l,m) is a binary control variable indicating whether vehicles are allowed to turn from the upstream lane l to the downstream lane m within the decision period, c(l,m) represents the saturation flow rate of the lane (l,m), and d(l,m) represents the vehicle demand arriving at the lane (l,m) from outside.
3. The method according to claim 2, wherein After coordinately controlling the traffic signals of the intersections through the target optimal phase, the method further includes: Obtaining the global reward of the global proximal policy optimization controller after controlling the traffic signals of the intersections based on the target phase according to the second reward formula, and obtaining the final reward based on the local reward and the global reward, and the second reward formula is as follows: R global = -AvgTravelTime + λ2·Throughput - λ3·CongestionIndex In the above formula, AvgTravelTime is the average travel time of all vehicles from the starting point to the ending point at each of the intersections, CongestionIndex is a comprehensive index reflecting the overall congestion level at each of the intersections, λ2 and λ3 are used to adjust the importance of throughput and congestion index in the reward, and Throughput is the total number of vehicles passing through all intersections in the area per unit time.
4. The method according to claim 1, characterized in that, The utilization of the global proximal policy optimization controller to perform queue weight correction, generate a green wave coordination offset, and / or emergency phase coverage processing on the current phase through the global optimal policy to obtain the target optimal phase includes: The global proximal policy optimization controller applies the proximal policy optimization algorithm to process the global optimal policy to obtain a continuous parameter adjustment amount and a phase selection probability; Based on the continuous parameter adjustment amount and the phase selection probability, the current phase is optimized to obtain the target optimal phase, and the optimization processing includes the queue weight correction, the green wave coordination offset, and / or the emergency phase coverage processing.
5. The method according to claim 3, wherein After the traffic signals at each intersection are collaboratively controlled through the target optimal phase, the method further includes: Adjusting the policy of the global proximal policy optimization controller through the final reward, and collecting the new queue lengths corresponding to the upstream lanes to the downstream lanes at each of the intersections again until new local traffic data and new global traffic data are obtained; When no congestion or emergency vehicle is detected, the underlying maximum pressure controller uses the maximum pressure algorithm and the new local traffic data to obtain the new optimal phase corresponding to each of the intersections; and controls the traffic signals at each intersection based on the new optimal phase; Encoding the new global traffic data through the graph neural network to obtain encoded new features, and using the global proximal policy optimization controller to learn a new global optimal policy through the new features; When congestion or an emergency vehicle is detected, the global proximal policy optimization controller performs queue weight correction, generates a green wave coordination offset, and / or emergency phase coverage processing on the new current phase through the new global optimal policy to obtain a new target optimal phase, and there is one global proximal policy optimization controller corresponding to all the intersections; Sending the new target optimal phase corresponding to each intersection to each intersection, and collaboratively controlling the traffic signals at each intersection through the new target optimal phase.
Citation Information
Cited By
Traffic signal cooperative control method and system based on multi-scale space attention
CN121545371A