A traffic signal optimization method based on multi-agent deep reinforcement learning
By constructing a context-enhanced state space and dynamically adjusting rewards, combined with multi-agent deep learning to optimize traffic signal phase switching, the problems of insufficient flexibility and efficiency in traffic signal control in existing technologies are solved, and an efficient response to complex traffic environments is achieved.
Patent Information
- Application Number
- CN202511108992.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing deep reinforcement learning in traffic signal control has the problem that the state space is limited to local features, the reward function is statically configured, and it cannot adapt to dynamic changes in traffic flow in real time, resulting in insufficient flexibility in highly dynamic and complex traffic environments.
A context-enhanced state space is constructed, regional traffic flow and directional flow imbalance are introduced, reward weights are dynamically adjusted, a multi-agent dual deep Q network is combined to train the reinforcement learning agent, traffic signal phase switching is optimized, and prior knowledge of traffic engineering is incorporated to improve the robustness of the control strategy.
It achieves comprehensive perception of macro and micro traffic dynamics, can identify sudden congestion in advance and make accurate decisions, improve the flexibility and efficiency of traffic signal control, and ensure optimal response in different traffic stages.
Smart Images

Figure CN120599840B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic signal control, and in particular relates to a traffic signal optimization method based on multi-agent deep reinforcement learning. Background Art
[0002] In recent years, urban traffic congestion has become one of the important factors restricting social and economic development and affecting the quality of life of residents. Traffic signal control is an important technical means to alleviate urban traffic congestion. At present, the widely used traffic signal control methods mainly include fixed timing control and rule-based adaptive control. The fixed timing control method usually adopts a pre-set signal timing scheme, which is suitable for environments with stable traffic flow, but has slow response and poor adaptability under complex and dynamic traffic conditions; although the rule-based adaptive control method can respond to changes in traffic flow to a certain extent, the rules are usually static and preset, and lack flexibility, especially when dealing with highly dynamic, multi-intersection coordinated control problems.
[0003] Deep reinforcement learning can adapt to dynamically changing traffic conditions through environmental interaction and autonomous decision-making. It has good development prospects and has gradually been applied to the field of traffic signal control. Some existing studies have shown that deep reinforcement learning has obvious advantages in traffic signal control, but it also exposes some shortcomings, mainly in the following aspects: (1) the state space used only contains local feature information such as queue length, lane density and current signal phase, and does not consider macroscopic traffic flow information at the regional scale, which limits the agent's ability to perceive the dynamic changes of the broader traffic environment; (2) the reward function is statically configured and cannot adapt to the dynamic changes of traffic flow in real time. In particular, the control effect is limited when the real-time traffic state changes drastically; (3) in highly congested and complex traffic environments, it lacks flexibility due to the limited dimension of the state space. Summary of the Invention
[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides a traffic signal optimization method based on multi-agent deep reinforcement learning, which solves the problems of insufficient flexibility and efficiency of traffic signal control in complex scenarios.
[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0006] The present invention provides a traffic signal optimization method based on multi-agent deep reinforcement learning, comprising the following steps:
[0007] S1. Track and simulate vehicles based on vehicle operation data, and build an urban traffic simulation model and an enhanced intelligent learning agent corresponding to traffic lights at each intersection;
[0008] S2. Constructing a context-enhanced state space, normalizing the feature parameters in the context state space, and combining them to obtain a real-time traffic environment state vector;
[0009] S3. Based on the basic reward indicators of each intersection in the urban traffic simulation model, dynamically adjust the weight of the basic reward indicators and calculate the congestion index adaptive reward;
[0010] S4. Based on the heuristic reward shaping method, define the flow matching index and the signal cycle position reward, and combine them with the congestion index adaptive reward to obtain the traffic signal optimization reward;
[0011] S5. Based on the traffic signal optimization reward, a multi-agent dual-depth Q network training is used to strengthen the intelligent learning agent to control the traffic signal phase switching.
[0012] The beneficial effects of the present invention are as follows: the present invention provides a traffic signal optimization method based on multi-agent deep reinforcement learning. By introducing regional traffic flow, directional flow imbalance and signal cycle position into state representation, a high-dimensional context-aware state vector is constructed, which enables the reinforced intelligent learning agent to fully perceive macro and micro traffic dynamics, realize early identification of sudden congestion or traffic fluctuations, and thus make more accurate phase switching decisions; based on the real-time calculated traffic congestion index, the present invention dynamically adjusts the weights of basic rewards such as waiting time, queue length, average speed and intersection pressure: in high congestion conditions, the intelligent agent focuses on optimizing the relief of queuing and waiting, and in low congestion conditions, it turns to increasing the passing speed and reducing the intersection pressure, ensuring that the strategy maintains the optimal response in different traffic phases, significantly improving the robustness of the control strategy; the present invention introduces two heuristic rewards, flow matching shaping and cycle position shaping, and integrates prior knowledge in the field of traffic engineering, such as dominant flow priority and reasonable phase switching timing, into reinforcement learning, thereby effectively guiding the intelligent agent to prioritize serving the dominant flow direction and avoid unnecessary phase extension, reducing ineffective actions in the exploration phase, accelerating model convergence and improving signal utilization efficiency.
[0013] Furthermore, the S1 includes the following steps:
[0014] S11. Use drones to collect real-time traffic video data from the air and extract vehicle location, speed, queue length, and traffic volume through target detection as vehicle operation data;
[0015] S12. Using a multi-target tracking algorithm, obtain trajectory data of each vehicle based on the vehicle operation data, and map the image coordinates in the trajectory data to an actual geographic coordinate system through perspective transformation;
[0016] S13, according to the path inference algorithm, converting the trajectory data into a corresponding routing file as simulation input data;
[0017] S14. Based on the simulation input data, an urban traffic simulation model is constructed, and an independent reinforcement learning agent is established for each traffic light at each intersection in the urban traffic simulation model.
[0018] Furthermore, the characteristic parameters in the context-enhanced state space in S2 include: current signal phase, minimum green light duration, queue length, lane density, regional traffic flow, directional flow imbalance and signal cycle position.
[0019] Furthermore, the S3 includes the following steps:
[0020] S31. Calculate a real-time traffic congestion index for the intersection based on basic reward indicators for each intersection in the urban traffic simulation model, where the basic reward indicators include real-time waiting time, real-time queue length, real-time average speed, and real-time intersection pressure;
[0021] The calculation expression of the real-time traffic congestion index of the intersection is as follows:
[0022] ,
[0023] in, represents the real-time traffic congestion index of the intersection, Indicates the real-time waiting time at the intersection, represents the real-time queue length at the intersection, represents the real-time average speed at the intersection, Indicates the real-time intersection pressure at the intersection;
[0024] S32. Dynamically adjust the weight of the basic reward indicator based on the real-time congestion index of the intersection;
[0025] The calculation expression of the weight of the basic reward indicator is as follows:
[0026] , , , ,
[0027] ,
[0028] in, represents the weight of real-time waiting time, represents the real-time congestion index adjustment factor, represents the weight of the real-time queue length, Indicates the weight of the real-time average speed, The weight representing the real-time intersection pressure;
[0029] S33. Calculate the congestion index adaptive reward based on the weight of the basic reward indicator;
[0030] The calculation expression of the congestion index adaptive reward is as follows:
[0031] ,
[0032] in, represents the congestion index adaptive reward, represents the waiting time reward, represents the queue length reward, represents the average speed reward, Represents the intersection pressure reward sub.
[0033] Furthermore, the S4 includes the following steps:
[0034] S41. Define the traffic matching index based on the heuristic reward shaping method;
[0035] The calculation expression of the traffic matching index is as follows:
[0036] ,
[0037] in, Indicates the traffic matching index. Indicates the real-time traffic flow in the north-south direction of the intersection. Indicates the real-time traffic flow in the east-west direction of the intersection. represents a constant that prevents the denominator from being zero, Indicates green light in the north-south direction. Indicates a green light in the east-west direction;
[0038] S42. Set traffic matching rewards based on traffic matching indicators;
[0039] The calculation expression of the traffic matching reward is as follows:
[0040] ,
[0041] ,
[0042]
[0043] in, Indicates traffic matching rewards, Indicates the dominant direction control strength, represents the number of dominant phase directions, represents the dominant direction control strength coefficient, Indicates the normalization factor of the current phase direction, represents the period indicator variable;
[0044] S43, defining an indication signal period position reward;
[0045] The calculation expression of the indication signal period position reward is as follows:
[0046] ,
[0047] in, Indicates the indicator signal period position reward, represents the positive coefficient of intensity adjustment, Indicates the normalized ratio of the current green light duration to the minimum green light duration;
[0048] S44. Calculate a traffic signal optimization reward based on the congestion index adaptive reward function, the traffic matching reward, and the indicator signal cycle position reward;
[0049] The calculation expression of the traffic signal optimization reward is as follows:
[0050]
[0051] in, represents the traffic signal optimization reward.
[0052] Furthermore, the S5 includes the following steps:
[0053] S51. Using a multi-agent dual-depth Q network, obtain the real-time traffic environment state vector at the current moment;
[0054] S52. According to the greedy strategy, based on the real-time traffic environment state vector, using the reinforcement learning agent to select and execute an action, wherein the action refers to switching the phase of the traffic signal;
[0055] S53. Obtaining a traffic signal optimization reward and a traffic environment state vector at the next moment in the urban traffic simulation model based on the action selected and executed by the reinforcement intelligent learning agent;
[0056] S54, repeat S51 to S53 several times, and store the historical interaction data in the experience replay pool;
[0057] S55. According to a preset period, batches of historical interaction data are extracted from the experience replay pool to update the network weights of the multi-agent dual-depth Q network to obtain an updated multi-agent dual-depth Q network;
[0058] S56. Utilize the updated multi-agent dual-depth Q network to enhance the intelligent learning agent's selection and execution of actions based on the newly acquired real-time traffic environment state vector to control the traffic signal phase switching.
[0059] Furthermore, the historical interaction data in S53 refers to the real-time traffic environment state vector at each moment, the action selected and executed by the reinforcement intelligent learning body under the real-time traffic environment state vector, the corresponding traffic signal optimization reward after selecting and executing the action, and the traffic environment state vector at the next moment.
[0060] Other advantages of the present invention will be analyzed in more detail in subsequent embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0062] Figure 1 This is a flowchart of the steps of a traffic signal optimization method based on multi-agent deep reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0064] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides a traffic signal optimization method based on multi-agent deep reinforcement learning, comprising the following steps:
[0065] S1. Track and simulate vehicles based on vehicle operation data, and build an urban traffic simulation model and an enhanced intelligent learning agent corresponding to traffic lights at each intersection;
[0066] The S1 comprises the following steps:
[0067] S11. Use drones to collect real-time traffic video data from the air and extract vehicle location, speed, queue length, and traffic volume through target detection as vehicle operation data;
[0068] In this embodiment, traffic video data is collected from the air by a drone, and target detection is performed on the traffic video data to obtain vehicle motion data, which can obtain more realistic traffic conditions. In this embodiment, target detection is implemented based on the YOLOv12 neural network, which can effectively extract data such as vehicle location, speed, queue length, and traffic flow.
[0069] S12. Using a multi-target tracking algorithm, obtain trajectory data of each vehicle based on the vehicle operation data, and map the image coordinates in the trajectory data to an actual geographic coordinate system through perspective transformation;
[0070] In this embodiment, the multi-target tracking algorithm used is BoT-SORT-ReID, which is an advanced tracking framework that combines a Transformer-based multi-target tracking algorithm and pedestrian re-identification technology. It is mainly used to improve the accuracy and robustness of multi-target tracking, especially when dealing with complex scenes such as occlusion and target similarity. The homography transformation used in the perspective transformation is Homography, which is a core concept in computer vision and geometry. It describes the projective transformation relationship between two planes and can establish a correspondence between three-dimensional space points and two-dimensional image points, thereby converting the vehicle trajectory in the video image into actual geographic coordinates for representation, providing a basis for accurate urban traffic simulation.
[0071] S13, according to the path inference algorithm, converting the trajectory data into a corresponding routing file as simulation input data;
[0072] In this embodiment, the acquired trajectory data of each vehicle is converted into a rou xml routing file required by the SUMO simulation platform through a path inference algorithm, which can be effectively used to accurately simulate actual traffic conditions.
[0073] S14. Based on the simulation input data, an urban traffic simulation model is constructed, and an independent reinforcement learning agent is established for each traffic light at each intersection in the urban traffic simulation model.
[0074] In this embodiment, the SUMO simulation platform is used to construct an urban traffic simulation model, which realizes detailed modeling of multi-lane roads, intersections, and special turning lanes. With micro-simulation as the core, it accurately simulates the acceleration, deceleration, lane changing, and following behaviors of a single vehicle, which can reflect the dynamic characteristics of real traffic flow. The SUMO simulation platform provides an API interface to support dynamic adjustment of traffic signal timing, lane functions, speed limit strategies, etc. It also supports the dynamic insertion of sudden traffic events to simulate their impact on traffic flow. It is also compatible with the collaborative simulation of motor vehicles, non-motor vehicles, pedestrians and public transportation.
[0075] In this solution, each intelligent agent autonomously optimizes the signal control strategy based on the real-time traffic data collected in the simulation platform, which can improve the realism of the simulation and enhance the adaptability of the reinforcement learning model in the actual traffic environment.
[0076] S2. Constructing a context-enhanced state space, normalizing the feature parameters in the context state space, and combining them to obtain a real-time traffic environment state vector;
[0077] The characteristic parameters in the context-enhanced state space in S2 include: current signal phase, minimum green light duration, queue length, lane density, regional traffic flow, directional flow imbalance and signal cycle position.
[0078] In this embodiment, the current signal phase uses a one-hot encoding format to identify the current green light control phase of the intersection. For example, for an intersection with four signal phases, if the current phase is 2, it is represented as [0, 1, 0, 0]. The minimum green light duration indicates whether the current signal has met the minimum green light duration, which is used to constrain signal phase switching. The queue length refers to the ratio of stationary vehicles to the number of vehicles that can be accommodated on each entrance lane, measuring the degree of vehicle backlog. The lane density refers to the density of vehicles per unit length on each entrance lane, reflecting the level of lane congestion. For example, a lane is 300 meters long, with an average vehicle length of 5 meters and a minimum vehicle spacing of 2 meters. It can accommodate a maximum of approximately 42 vehicles. If there are currently 30 vehicles, the density is approximately 0.71. Regional traffic flow refers to dividing all lanes into four directions: north, south, east, and west. The number of vehicles in each direction is counted and normalized to obtain the traffic flow [north proportion, south proportion, east proportion, west proportion], which reflects the overall traffic load distribution in the intersection area. Directional flow imbalance refers to summing the total north-south flow and the total east-west flow, respectively, and normalizing them to obtain the flow proportion [NS proportion, EW proportion] to quantify the degree of difference in traffic flow in each direction. If there are 20 vehicles in the north-south direction and 30 vehicles in the east-west direction, the value is [0.4, 0.6]. Signal cycle position indicates the relative position of the current time within the signal cycle. For example, if the entire signal cycle is 60 seconds and the current time is at the 15th second, it can be normalized to 15 / 60 = 0.25. The above features are normalized before being input into the reinforcement learning model, uniformly mapped to the interval [0, 1], and combined to form a high-dimensional state vector. This vector covers signal control status, lane operation status, regional traffic structure and timing characteristics, providing relatively complete environmental perception support for intelligent agent strategy learning.
[0079] This scheme obtains a high-dimensional real-time traffic environment state vector by normalizing and combining the feature parameters in the context-enhanced state space to fully describe the real-time traffic environment.
[0080] S3. Based on the basic reward indicators of each intersection in the urban traffic simulation model, dynamically adjust the weight of the basic reward indicators and calculate the congestion index adaptive reward;
[0081] In this scheme, in order to achieve the adaptive response of traffic signal control strategy to traffic congestion conditions, this scheme constructs the dynamic reward mechanism in S3;
[0082] The S3 comprises the following steps:
[0083] S31. Calculate a real-time traffic congestion index for the intersection based on basic reward indicators for each intersection in the urban traffic simulation model, where the basic reward indicators include real-time waiting time, real-time queue length, real-time average speed, and real-time intersection pressure;
[0084] The calculation expression of the real-time traffic congestion index of the intersection is as follows:
[0085] ,
[0086] in, represents the real-time traffic congestion index of the intersection, Indicates the real-time waiting time at the intersection, represents the real-time queue length at the intersection, represents the real-time average speed at the intersection, Indicates the real-time intersection pressure at the intersection;
[0087] S32. Dynamically adjust the weight of the basic reward indicator based on the real-time congestion index of the intersection;
[0088] The calculation expression of the weight of the basic reward indicator is as follows:
[0089] , , , ,
[0090] ,
[0091] in, represents the weight of real-time waiting time, represents the real-time congestion index adjustment factor, represents the weight of the real-time queue length, Indicates the weight of the real-time average speed, The weight representing the real-time intersection pressure;
[0092] In this scheme, when the intersection is highly congested, the weight of the basic reward indicators tends to reduce queue length and waiting time. When the congestion level is low, the weight tends to increase the average speed and reduce intersection pressure. By calculating the traffic congestion index in real time and dynamically adjusting the weights of basic reward items such as waiting time, queue length, average speed and intersection pressure, adaptive optimization for different congestion levels can be achieved.
[0093] S33. Calculate the congestion index adaptive reward based on the weight of the basic reward indicator;
[0094] The calculation expression of the congestion index adaptive reward is as follows:
[0095] ,
[0096] in, represents the congestion index adaptive reward, represents the waiting time reward, represents the queue length reward, represents the average speed reward, Represents the intersection pressure reward sub.
[0097] S4. Based on the heuristic reward shaping method, define the flow matching index and the signal cycle position reward, and combine them with the congestion index adaptive reward to obtain the traffic signal optimization reward;
[0098] In order to effectively guide the intelligent agent to learn the control strategy that meets the actual needs of traffic engineering, this scheme introduces a heuristic reward shaping method;
[0099] The S4 comprises the following steps:
[0100] S41. Define the traffic matching index based on the heuristic reward shaping method;
[0101] The calculation expression of the traffic matching index is as follows:
[0102] ,
[0103] in, Indicates the traffic matching index. Indicates the real-time traffic flow in the north-south direction of the intersection. Indicates the real-time traffic flow in the east-west direction of the intersection. represents a constant that prevents the denominator from being zero, Indicates green light in the north-south direction. Indicates a green light in the east-west direction;
[0104] In this scheme, the flow matching index reflects the consistency between the current green light phase and the dominant flow direction.
[0105] S42. Set traffic matching rewards based on traffic matching indicators;
[0106] The calculation expression of the traffic matching reward is as follows:
[0107] ,
[0108] ,
[0109]
[0110] in, Indicates traffic matching rewards, Indicates the dominant direction control strength, represents the number of dominant phase directions, represents the dominant direction control strength coefficient, Indicates the normalization factor of the current phase direction, represents the period indicator variable;
[0111] In this embodiment, during peak traffic hours, During off-peak hours By setting the time period indicator variable, this scheme can effectively encourage the agent to give priority to serving the dominant flow direction, i.e., the flow direction with greater traffic volume, during peak hours.
[0112] S43, defining an indication signal period position reward;
[0113] In this scheme, by defining the reward of the signal cycle position, unnecessary phase extension is suppressed and the agent is encouraged to switch the signal at a reasonable time point;
[0114] The calculation expression of the indication signal period position reward is as follows:
[0115] ,
[0116] in, Indicates the indicator signal period position reward, represents the positive coefficient of intensity adjustment, Indicates the normalized ratio of the current green light duration to the minimum green light duration;
[0117] In this scheme, the reward for the indicator signal period position is based on a positive reward at the beginning of the phase. As time progresses, the reward gradually decreases or even turns into a penalty, eventually inducing phase switching, thereby improving timing utilization efficiency.
[0118] S44. Calculate a traffic signal optimization reward based on the congestion index adaptive reward function, the traffic matching reward, and the indicator signal cycle position reward;
[0119] The calculation expression of the traffic signal optimization reward is as follows:
[0120]
[0121] in, represents the traffic signal optimization reward.
[0122] In this scheme, in addition to the basic rewards, traffic matching rewards and indicator signal cycle position rewards are superimposed. Through traffic signal optimization rewards, it is possible to simultaneously achieve a balance between macro-level traffic congestion response and micro-level traffic signal optimization decision-making, guiding the intelligent agent to learn more efficient and robust traffic signal control strategies.
[0123] S5. Based on the traffic signal optimization reward, a multi-agent dual-depth Q network training is used to strengthen the intelligent learning agent to control the traffic signal phase switching.
[0124] The S5 comprises the following steps:
[0125] S51. Using a multi-agent dual-depth Q network, obtain the real-time traffic environment state vector at the current moment;
[0126] S52. According to the greedy strategy, based on the real-time traffic environment state vector, using the reinforcement learning agent to select and execute an action, wherein the action refers to switching the phase of the traffic signal;
[0127] S53. Obtaining a traffic signal optimization reward and a traffic environment state vector at the next moment in the urban traffic simulation model based on the action selected and executed by the reinforcement intelligent learning agent;
[0128] The historical interaction data in S53 refers to the real-time traffic environment state vector at each moment, the action selected and executed by the reinforcement intelligent learning agent under the real-time traffic environment state vector, the corresponding traffic signal optimization reward after selecting and executing the action, and the traffic environment state vector at the next moment.
[0129] S54, repeat S51 to S53 several times, and store the historical interaction data in the experience replay pool;
[0130] S55. According to a preset period, batches of historical interaction data are extracted from the experience replay pool to update the network weights of the multi-agent dual-depth Q network to obtain an updated multi-agent dual-depth Q network;
[0131] S56. Utilize the updated multi-agent dual-depth Q network to enhance the intelligent learning agent's selection and execution of actions based on the newly acquired real-time traffic environment state vector to control the traffic signal phase switching.
[0132] In this paper, under the framework of multi-agent dual deep Q network (DDQN), the context state space, dynamic adaptive reward mechanism and heuristic reward shaping are jointly applied to form an end-to-end optimized control of traffic signals.
[0133] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A traffic signal optimization method based on multi-agent deep reinforcement learning, characterized in that: The steps include: S1. Track and simulate vehicles based on vehicle operation data, and build an urban traffic simulation model and an enhanced intelligent learning agent corresponding to traffic lights at each intersection; S2. Constructing a context-enhanced state space, normalizing the feature parameters in the context-enhanced state space, and combining them to obtain a real-time traffic environment state vector; S3. Based on the basic reward indicators of each intersection in the urban traffic simulation model, dynamically adjust the weight of the basic reward indicators and calculate the congestion index adaptive reward; The S3 comprises the following steps: S31. Calculate a real-time traffic congestion index for the intersection based on basic reward indicators for each intersection in the urban traffic simulation model, where the basic reward indicators include real-time waiting time, real-time queue length, real-time average speed, and real-time intersection pressure; The calculation expression of the real-time traffic congestion index of the intersection is as follows: , in, represents the real-time traffic congestion index of the intersection, Indicates the real-time waiting time at the intersection, represents the real-time queue length at the intersection, represents the real-time average speed at the intersection, Indicates the real-time intersection pressure at the intersection; S32. Dynamically adjust the weight of the basic reward indicator based on the real-time congestion index of the intersection; The calculation expression of the weight of the basic reward indicator is as follows: , , , , , in, represents the weight of real-time waiting time, represents the real-time congestion index adjustment factor, represents the weight of the real-time queue length, Indicates the weight of the real-time average speed, The weight representing the real-time intersection pressure; S33. Calculate the congestion index adaptive reward based on the weight of the basic reward indicator; The calculation expression of the congestion index adaptive reward is as follows: , in, represents the congestion index adaptive reward, represents the waiting time reward, represents the queue length reward, represents the average speed reward, Indicates intersection pressure reward; S4. Based on the heuristic reward shaping method, define the flow matching index and the signal cycle position reward, and combine them with the congestion index adaptive reward to obtain the traffic signal optimization reward; The S4 comprises the following steps: S41. Define the traffic matching index based on the heuristic reward shaping method; The calculation expression of the traffic matching index is as follows: , in, Indicates the traffic matching index. Indicates the real-time traffic flow in the north-south direction of the intersection. Indicates the real-time traffic flow in the east-west direction of the intersection. represents a constant that prevents the denominator from being zero, Indicates green light in the north-south direction. Indicates a green light in the east-west direction; S42. Set traffic matching rewards based on traffic matching indicators; The calculation expression of the traffic matching reward is as follows: , , , in, Indicates traffic matching rewards, Indicates the dominant direction control strength, represents the number of dominant phase directions, represents the dominant direction control strength coefficient, Indicates the normalization factor of the current phase direction, represents the period indicator variable; S43, defining an indication signal period position reward; The calculation expression of the indication signal period position reward is as follows: , in, Indicates the indicator signal period position reward, represents the positive coefficient of intensity adjustment, Indicates the normalized ratio of the current green light duration to the minimum green light duration; S44. Calculate a traffic signal optimization reward based on the congestion index adaptive reward function, the traffic matching reward, and the indicator signal cycle position reward; The calculation expression of the traffic signal optimization reward is as follows: , in, represents the traffic signal optimization reward; S5. Based on the traffic signal optimization reward, a multi-agent dual-depth Q network training is used to strengthen the intelligent learning agent to control the traffic signal phase switching.
2. The traffic signal optimization method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The S1 comprises the following steps: S11. Use drones to collect real-time traffic video data from the air and extract vehicle location, speed, queue length, and traffic volume through target detection as vehicle operation data; S12. Using a multi-target tracking algorithm, obtain trajectory data of each vehicle based on the vehicle operation data, and map the image coordinates in the trajectory data to an actual geographic coordinate system through perspective transformation; S13, according to the path inference algorithm, converting the trajectory data into a corresponding routing file as simulation input data; S14. Based on the simulation input data, an urban traffic simulation model is constructed, and an independent reinforcement learning agent is established for each traffic light at each intersection in the urban traffic simulation model.
3. The traffic signal optimization method based on multi-agent deep reinforcement learning according to claim 2 is characterized in that: The feature parameters in the context-enhanced state space in S2 include: current signal phase, minimum green light duration, queue length, lane density, regional traffic flow, directional flow imbalance and signal cycle position; Regional traffic flow refers to dividing all lanes into four directions: north, south, east, and west, counting the number of vehicles in each direction, and normalizing the traffic flow; directional traffic imbalance refers to summing the total north-south flow and the total east-west flow, respectively, and normalizing the resulting traffic flow ratio; the signal cycle position indicates the relative position of the current moment within the signal cycle.
4. The traffic signal optimization method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The S5 comprises the following steps: S51. Using a multi-agent dual-depth Q network, obtain the real-time traffic environment state vector at the current moment; S52. According to the greedy strategy, based on the real-time traffic environment state vector, using the reinforcement learning agent to select and execute an action, wherein the action refers to switching the phase of the traffic signal; S53. Obtaining a traffic signal optimization reward and a traffic environment state vector at the next moment in the urban traffic simulation model based on the action selected and executed by the reinforcement intelligent learning agent; The historical interaction data in S53 refers to the real-time traffic environment state vector at each moment, the action selected and executed by the reinforcement intelligent learning agent under the real-time traffic environment state vector, the corresponding traffic signal optimization reward after selecting and executing the action, and the traffic environment state vector at the next moment; S54, repeat S51 to S53 several times, and store the historical interaction data in the experience replay pool; S55. According to a preset period, batches of historical interaction data are extracted from the experience replay pool to update the network weights of the multi-agent dual-depth Q network to obtain an updated multi-agent dual-depth Q network; S56. Utilize the updated multi-agent dual-depth Q network to enhance the intelligent learning agent's selection and execution of actions based on the newly acquired real-time traffic environment state vector to control the traffic signal phase switching.
Citation Information
Patent Citations
Traffic signal optimization control method
CN115171408A
Adaptive traffic signal control method based on reinforcement learning and self-attention mechanism
CN118942261A