Intelligent collaborative optimization method for regional traffic signal control
Through the collaborative acquisition of traffic video by multiple drones and combined with deep reinforcement learning, dynamic optimization of regional traffic signals is achieved, solving the problems of poor adaptability and insufficient coordination of signal control in traditional methods, and improving traffic efficiency and stability.
Patent Information
- Application Number
- CN202510416537.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional traffic signal timing methods cannot effectively deal with dynamic changes in different time periods and traffic flow, resulting in poor adaptability of signal control, lagging response, and insufficient coordination, making it difficult to alleviate traffic congestion during peak periods.
Through multiple drones collaboratively collecting traffic videos within the region, combined with deep reinforcement learning algorithms, coordinated optimization of traffic signal timing within the region can be achieved. The specific steps include: optimizing the aerial photography path of the drone based on the road network structure, using YOLOv10s and DeepSORT algorithms for target detection and tracking, obtaining traffic flow data and signal timing parameters, and using multi-agent deep reinforcement learning for signal coordination and control.
The dynamic, accurate and efficient optimization of regional traffic signals has been achieved, the intelligence and flexibility of urban traffic management has been improved, traffic delays have been reduced, and overall traffic efficiency and stability have been improved.
Smart Images

Figure CN120236402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep reinforcement learning and intelligent transportation, and more specifically to a regional traffic signal optimization method that combines drone sensing data and multi-agent collaboration. The method aims to achieve regionalized, dynamic, and intelligent traffic signal collaboration optimization. Background Art
[0002] With the acceleration of the urbanization process, the global motor vehicle ownership has been increasing continuously, and the pressure faced by urban traffic networks has been growing. Problems such as traffic congestion and low traffic efficiency have become increasingly prominent. In addition, the uneven distribution of traffic flow and the unreasonable configuration of traffic signal control have made it difficult to improve the traffic efficiency of the road network during peak hours. Even in some areas with dense traffic flows, the phenomenon of "secondary congestion" caused by inappropriate signal timing has become more obvious. Frequent starts, stops, accelerations, and brakings not only increase the risk of traffic accidents but also pose a severe challenge to urban traffic safety.
[0003] Traditional traffic signal timing methods mostly rely on fixed-time control and fixed timing strategies, and are unable to effectively cope with the dynamic changes in different time periods and traffic flows. Most existing optimization methods still focus on the independent optimization of a single intersection, lacking the collaborative control of traffic flows within the region, resulting in poor adaptability, lagging response, and insufficient coordination of signal control. This makes it difficult to alleviate the traffic congestion problem during peak hours.
[0004] To address this problem, this study proposes an intelligent collaborative optimization method for regional traffic signal control, which combines drone aerial photography technology to achieve intelligent signal collaborative optimization, aiming to improve the accuracy and efficiency of signal scheduling and achieve more intelligent and flexible urban traffic management. Summary of the Invention
[0005] The present invention uses multi-drone collaboration to collect traffic videos within the region, calculates traffic parameters based on the videos, and reverse-infers traffic signal timing. Subsequently, it collaboratively optimizes the traffic signal timing within the region. The optimization process is based on a deep reinforcement learning algorithm. Through the interaction process of agents, continuous learning and optimization are carried out, enabling the signal timing to cope with different traffic conditions, reducing delays, and improving the overall efficiency and stability of urban traffic.
[0006] To achieve the above objectives, the specific technical solutions adopted by the present invention are as follows:
[0007] S1: Based on the road network structure of the target area, a multi-drone collaborative scheduling method is used for path optimization to ensure task balance among different drones, reduce the overlap of aerial photography ranges, and improve the overall data collection efficiency.
[0008] S2: Apply the YOLOv10s object detection algorithm to the collected aerial videos for object detection, and combine the DeepSORT algorithm to achieve multi-object tracking, obtaining the continuous trajectories of traffic participants. Through in-depth analysis of their movement patterns and calculation of relevant data, traffic flow data and signal timing parameters of each intersection in the area are obtained.
[0009] S3: In the regional traffic signal control model based on deep reinforcement learning, multiple agents are used for regional coordinated control. By refining and compressing state information, the dimension of state information is reduced. On the premise of ensuring the optimization effect, the learning efficiency and calculation speed of the agents are improved. The reward function is redesigned, comprehensively considering the comprehensiveness of the evaluation system and the overall efficiency of the algorithm. By weighted fusion of multiple optimization indicators, efficient, accurate and dynamic adjustment of signal timing optimization within the region is achieved.
[0010] Further, in step S1, it includes:
[0011] S101: Based on the road network structure of the target area, use an improved genetic algorithm for initial aerial photography path planning. First, determine the target area of the UAV aerial photography and the positions of key intersections. Then, according to the distribution and characteristics of the intersections, initialize the UAV flight route population. On this basis, use the fitness function to evaluate factors such as the coverage rate, flight cost, and task requirements of the path. The fitness function comprehensively evaluates multiple indicators such as the coverage effect of the aerial photography area, flight distance, and power consumption. Through selection, crossover, and mutation operations in the genetic algorithm, iteratively optimize the path planning to ensure that the UAV can cover all key intersections in the shortest time and optimize the flight path to maximize flight efficiency.
[0012] S102: To improve the flexibility and efficiency of task allocation, use the auction algorithm to allocate the optimal monitoring area for each UAV. Under this mechanism, the system allocates the monitoring area for the UAVs through real-time bidding, and at the same time ensures the balanced allocation of tasks through the auction process, avoiding task duplication and overlap among UAVs and improving the overall monitoring efficiency.
[0013] Further, in step S2, it includes:
[0014] S201: Based on the collected aerial videos, use the YOLOv10s object detection algorithm to accurately identify traffic participants (including motor vehicles, non-motor vehicles, and pedestrians) within the scope of each intersection. On the basis of object detection, combine the DeepSORT multi-object tracking algorithm to track the trajectories of the detected traffic participants to ensure the continuity of target identity and the integrity of the trajectory.
[0015] S202: Additionally, calculate key traffic parameters using trajectory data, including vehicle speed, flow, density, etc., and perform inverse inference of signal timing based on the motion state of vehicles to obtain a complete signal timing plan for the intersection, providing accurate data support for regional traffic signal collaborative optimization.
[0016] Further, in step S3, it includes:
[0017] S301: Adopt multi-agent deep reinforcement learning for regional coordinated control. Each intersection signal control agent makes autonomous decisions based on the traffic flow state it perceives, and shares key information through the communication mechanism between adjacent agents to solve the local observation problem and improve the collaborative efficiency of overall signal timing optimization. In addition, in order to reduce the dimension of the state space, this study conducts feature extraction and compression on traffic flow state information to ensure the optimization effect while improving the convergence speed and computational efficiency of the reinforcement learning training process.
[0018] S302: In response to the requirements of regional traffic signal optimization, redesign the reward function, comprehensively consider multiple indicators such as average vehicle speed, queue length, delay time, signal cycle utilization rate, etc., construct it using a weighted fusion strategy, and reduce the negative reward threshold to better meet the requirements of arterial coordinated signals. Through such a reward function design, the agent can learn strategies that are more in line with actual traffic management goals during the training process. The objective function of the agent is to maximize the cumulative reward, ultimately realizing the dynamic optimization of regional signal control, thereby improving the global traffic efficiency.
[0019] S303: Perform phase control and timing control on the signal control process through a neural network. Among them, the phase control module selects the optimal signal phase through the Markov decision process (MDP). Each intersection agent adopts a mini-batch data update strategy during the training process and optimizes the control decision through interactive learning. The timing control module then dynamically adjusts the signal phase duration according to the current phase decision to ensure that the optimized timing plan can adapt to different traffic states, realizing precise and efficient scheduling of regional traffic signal control.
[0020] Benefits of the present invention:
[0021] 1. By optimizing the task allocation strategy, ensure that each drone can cover different traffic flow areas, reduce the overlap of the aerial photography range, improve the overall monitoring efficiency, and thus achieve more comprehensive and efficient traffic flow data collection.
[0022] 2. Adopt the advanced object detection algorithm YOLOv10s and the tracking algorithm DeepSORT to accurately identify and track different types of traffic participants, thereby obtaining detailed traffic flow data and inferring the traffic signal timing plan, improving the adaptability and intelligent level of signal control.
[0023] 3. A traffic signal collaborative control method based on deep reinforcement learning is proposed. By embedding the multi-head self-attention (MHA) mechanism into a multi-layer perceptron, the spatio-temporal feature expression ability of traffic flow data is improved, and the signal timing optimization is combined with the Markov decision process (MDP) to ensure the coordination and global optimality of signal control at different intersections, thereby effectively improving the traffic efficiency and safety of the regional traffic. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0025] Figure 1 is a flowchart of the present invention;
[0026] Figure 2 is a detailed flowchart of the present invention;
[0027] Figure 3 is a schematic aerial view of the area;
[0028] Figure 4 is an overall architecture diagram of the multi-agent traffic signal optimization control environment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the technical solutions, objectives, and advantages of the present invention clearer and more understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings. It should be noted that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments that can be obtained by those of ordinary skill in the art without creative work according to the described embodiments of the present invention belong to the scope of protection of the present invention.
[0030] As Figure 2 shown, the present invention example proposes an intelligent collaborative optimization method for regional traffic signal control, and the specific implementation is as follows:
[0031] S1 UAV Aerial Photography Path Optimization and Task Collaborative Allocation
[0032] (1) In the implementation process of the present invention, for the road network structure of the target area, an improved genetic algorithm is used for initial aerial photography path planning to ensure that the UAV can efficiently cover all key intersections in the shortest time.
[0033] Based on the positions of the target area and key intersections of UAV aerial photography, an initial aerial photography path population is constructed. Each path consists of a series of aerial photography points, and these aerial photography points correspond to the central positions and surrounding areas of the intersections.
[0034] The aerial photography path is evaluated using a fitness function. The construction of the fitness function is as follows:
[0035] F = w1C cover + w2C energy + w3C time
[0036] where C cover represents the coverage rate of the aerial photography area, C energy represents the energy consumption of the UAV, and C time represents the completion time of the aerial photography task. w1, w2, and w3 are weight parameters.
[0037] (2) During the optimization process, the aerial photography path is iteratively updated through selection, crossover, and mutation operations in the genetic algorithm. The specific steps are as follows:
[0038] Selection: According to the fitness function value, select the paths with higher fitness from the current population to enter the next generation.
[0039] Crossover: Randomly select two path individuals and exchange part of the paths with a certain probability to generate a new path scheme.
[0040] Mutation: Randomly adjust the order of some aerial photography points in the path to improve the diversity of solutions and prevent falling into local optimal solutions.
[0041] After multiple rounds of iteration, the optimal aerial photography path is finally converged, as Figure 3 shown, ensuring that the UAV can cover all key intersections with the optimal flight efficiency.
[0042] (3) To improve the monitoring efficiency of the UAV swarm, an auction algorithm based on the market bidding mechanism is used for task allocation to ensure the balance of tasks and avoid duplicate coverage. The system collects the initial states of all UAVs, including information such as the current position, power status, and task load, and initializes the aerial photography area to be allocated.
[0043] Each UAV evaluates the area to be monitored according to its state and submits a bid to the system. The bidding strategy is calculated based on the following formula:
[0044] B i = αE i + βD i + γC i
[0045] where E i represents the remaining power of the UAV, D i represents the distance between the UAV and the target area, C i represents the current task load, and α, β, and γ are weight coefficients.
[0046] (4) Based on the bidding results, the system selects the optimal bidder to assign an executor for the monitoring tasks in a certain area. This process continues until all monitoring areas are successfully assigned. The drones execute the aerial monitoring tasks according to the assigned tasks and upload the traffic flow data in real time.
[0047] Through the above task allocation method based on the auction mechanism, each drone can execute the optimal task within its capacity, maximizing the monitoring efficiency, avoiding task overlap among drones, and improving the accuracy and integrity of overall traffic data collection.
[0048] S2 Target Detection, Tracking and Traffic Flow Data Extraction
[0049] (1) When the drone takes videos, it uses an encoding and decoding standard to transmit the high-resolution video stream to the computer. The video is processed through a corresponding program, and a trained neural network model is used to detect and track traffic participants in the aerial videos of each intersection to obtain basic data.
[0050] (2) Calculate traffic parameters based on the data obtained in S2. The reasoning methods for vehicle speed, flow, and density are as follows:
[0051]
[0052] Among them, represents the two-dimensional
[0053] plane displacement distance between adjacent timestamps, and t′ i -t i is the difference in timestamps, and Δt represents the reciprocal of the video acquisition frequency.
[0054]
[0055] Among them, I cross (t) represents the instantaneous flow rate passing through the observation point.
[0056]
[0057] Among them, K i represents the instantaneous density at the i-th observation point, represents the value of normalizing the density to a unit length.
[0058] Analyze according to the movement laws of traffic participants and the changes in vehicle queues, speculate on the signal timing parameters of the current intersection, and provide support for optimization decisions. Detect whether a new vehicle starts. If a new vehicle starts, record the starting position and time of the vehicle. If the current signal phase changes, start updating the phase list, record the effective vehicle information and convert it into phase information. If a new vehicle start is detected and the phase information changes, update the phase information, traverse the entire phase list, and synthesize the green and red light times of each phase to form a complete signal timing plan for this intersection.
[0059] S3 Multi-Agent Deep Reinforcement Learning for Optimizing Signal Control
[0060] (1) Each intersection is a signal control agent. The agent learns based on the historical traffic flow state information collected from video data and decides the corresponding signal timing strategy. At the same time, extract relevant traffic features to form the decision basis for optimizing signal timing. Through the communication mechanism between agents, agents at adjacent intersections are allowed to share traffic data to achieve cross-intersection signal coordination. The communication between agents can not only improve the understanding of the overall traffic state but also optimize the global signal timing strategy. Test the signal control optimization plan by simulating different traffic scenarios.
[0061] (2) Refine the traffic flow state information of each intersection, including key indicators such as average vehicle speed, traffic flow density, and queue length. Process the original data through feature engineering methods to reduce the dimension of the state information and extract the features most valuable for signal timing control, improving the learning efficiency and calculation speed of the agent while ensuring the optimization effect.
[0062] (3) Re-optimize and improve the state space and reward function of the function for the regional coordinated control scenario. In addition to introducing common road network performance parameters such as average delay and number of stops, propose indicators such as branch balance, signal cycle utilization rate, and number of conflict points. Maximize the traffic capacity of the arterial road and minimize the average vehicle delay on the arterial road in the regional coordinated control scenario, construct a multi-objective optimization model, and at the same time reduce the negative reward threshold to better meet the needs of regional signal coordination.
[0063] The objective function f and the reward function r are composed of a weighted combination of multiple sub-objectives:
[0064] f = α1f1 + α2f2 + α3f3 + α4f4 + α5f5
[0065] where α1, α2, α3, α4, α5 are weight parameters.
[0066]
[0067] f1 is the arterial road delay, Nm is the total number of intersections, d i is the average delay time of the i-th intersection.
[0068]
[0069] f2 is the number of stops, P i is the average number of stops at the i-th intersection.
[0070]
[0071] f3 is the balance of branch roads, N s is the total number of branch roads, d j is the average delay time of the branch road.
[0072]
[0073] f4 is the utilization rate of signal cycle, C i is the total signal cycle, is the actual passing time.
[0074]
[0075] f5 is the number of potential conflict points, X i is the actual number of conflict points at the i-th intersection.
[0076] r = w1R1 + w2R2 + w3R3 + w4R4 + w5R5
[0077] where w1, w2, w3, w4, w5 are weight parameters, R1, R2 are the trunk line delay reward and the stop number reward respectively, which are negative values; R3, R4 are the branch road balance reward and the signal cycle utilization rate reward respectively, which are positive values; R5 is the conflict point number penalty.
[0078] (4) Perform reinforcement learning optimization for phase control and timing control, such as Figure 4 shown, use the Markov decision process (MDP) to select the optimal signal phase for each intersection agent, and at the same time embed the multi-head self-attention (MHA). The MDP selects the most suitable signal phase according to the current traffic flow state to reduce traffic delay. During the training process, use <S,a d ,a p ,r,S′> mini-batch update strategy. Where the current agent state is S, and the executed action is <a d ,a p >,a d and a pThey respectively represent the phase action and the timing action, and then generate a new state S′ and a reward r. By extracting a part of the data from the experience pool each time for training and continuously updating the Q-value function, the agent can learn and optimize the signal control strategy.
Claims
1. An intelligent collaborative optimization method for regional traffic signal control, characterized in that: The following steps are involved: S1: Based on the road network structure of the target area, a multi-UAV collaborative scheduling method is used to optimize the path, ensure the task balance between different UAVs, reduce the overlap of aerial photography ranges, and improve the overall data collection efficiency; S2: Apply the YOLOv10s target detection algorithm to the collected aerial videos for target detection, and combine it with the DeepSORT algorithm to achieve multi-target tracking, obtain the continuous trajectory of traffic participants, and obtain the traffic flow data and signal timing parameters of each intersection in the area through in-depth analysis of their movement patterns and calculation of related data; S3: In the regional traffic signal control model based on deep reinforcement learning, multiple intelligent agents are used for regional coordinated control. By refining and compressing state information and reducing the dimension of state information, the learning efficiency and calculation speed of the intelligent agents are improved while ensuring the optimization effect. The reward function is redesigned, and the comprehensiveness of the evaluation system and the overall efficiency of the algorithm are comprehensively considered. By weighted fusion of multiple optimization indicators, efficient, accurate and dynamic adjustment of signal timing optimization within the region is achieved.
2. The intelligent collaborative optimization method for regional traffic signal control according to claim 1 is characterized in that: In step S1, a fitness function is introduced to comprehensively evaluate factors such as aerial photography area coverage, flight cost and power consumption, and the flight route is optimized through the selection, crossover and mutation operations of the genetic algorithm to improve the aerial photography efficiency and coverage accuracy of the drone. The auction algorithm is used to dynamically allocate the monitoring area of the drone to ensure balanced task distribution, avoid repeated coverage of aerial photography tasks, and improve data collection efficiency.
3. The intelligent collaborative optimization method for regional traffic signal control according to claim 1 is characterized in that: In step S2, the latest YOLOv10s target detection algorithm is combined to perform high-precision target recognition, and the DeepSORT algorithm is used to achieve multi-target trajectory tracking to ensure the continuity of traffic participants' movements and trajectory integrity. The extracted trajectory data is used to calculate key traffic parameters and reverse the signal timing plan of the intersection, providing accurate input data for regional traffic signal optimization.
4. The intelligent collaborative optimization method for regional traffic signal control according to claim 1 is characterized in that: In step S3, a multi-agent deep reinforcement learning framework is adopted to enable each intersection agent to make autonomous decisions and share traffic status information through a communication mechanism between adjacent agents, thereby improving the coordination of global signal control. By feature extraction and state information dimensionality reduction, the complexity of agent input data is reduced, the convergence speed and computational efficiency of the reinforcement learning training process are improved, and the signal control strategy is made more efficient.
5. The intelligent collaborative optimization method for regional traffic signal control according to claim 1, characterized in that: In step S3, the reward function is redesigned to comprehensively consider multiple traffic efficiency indicators such as average vehicle speed, queue length, delay time, and signal cycle utilization, and the signal timing scheme is optimized through a weighted fusion strategy to make the optimization process more in line with actual traffic management needs. The Markov decision process (MDP) is used for signal phase decision-making, combined with a small batch data update strategy, so that the intelligent agent can gradually learn the optimal signal phase and timing control scheme during the training process, ensuring that the signal control system can dynamically adapt to different traffic flow conditions and achieve accurate scheduling and optimization.
Citation Information
Cited By
Traffic signal control method, equipment and medium
CN121528007A