A Traffic Signal Timing Optimization Method Based on UAV Edge Computing
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]固定摄像机受安装角度、监测范围和建筑遮挡影响较大,难以同时覆盖机动车道、非机动车道和人行横道区域,导致机动车目标、非机动车目标和行人目标存在漏检、误检和轨迹中断问题,复杂交叉口中的异构交通流状态难以准确获取;无人机采集的视频流存在俯视视角变化、边缘区域透视压缩和目标尺度变化明显的问题,现有目标检测网络在多尺度交通目标识别过程中容易出现小目标特征丢失、目标边缘信息衰减和跨帧检测不稳定现象,导致机动车目标、非机动车目标和行人目标检测精度下降;传统多目标跟踪算法主要依据边界框交并比和运动信息进行轨迹关联,在非机动车密集通行、目标遮挡频繁和多相位交替放行场景下容易发生轨迹漂移、轨迹切换和身份混淆,导致交通轨迹连续性不足以及交通流统计结果偏差较大;现有交通信号配时方法大多侧重机动车通行效率优化,对非机动车和行人通行需求考虑不足,难以对机动车、非机动车和行人之间的通行冲突关系进行协同平衡,导致复杂交叉口区域中等待时间增加、排队长度扩大以及整体通行效率下降
[0064](1)通过改进YOLOv8目标检测网络与视漂弥散机制的结合,针对无人机俯视场景中交通目标尺度变化明显、边缘区域透视压缩严重和小目标特征易丢失问题,采用目标尺度变化梯度、边缘区域透视压缩量和跨帧尺寸偏移量建立尺度漂移场,实现不同尺度特征图之间的特征补偿,提高机动车目标、非机动车目标和行人目标的检测精度与检测稳定性;
Smart Images

Figure CN122575153A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent traffic control technology, and in particular to a traffic signal timing optimization method based on UAV edge computing. Background Technology
[0002] With the continuous growth of urban traffic flow and the increasing demand for intelligent transportation systems, technologies for traffic flow perception, traffic target tracking, and traffic signal timing optimization in complex intersection areas have received widespread attention. Existing traffic signal control systems mainly rely on geomagnetic detectors, loop detectors, or fixed cameras for traffic flow detection, and combine these with preset cycles or inductive control methods to execute traffic signal timing. However, in practical applications, the following problems commonly exist:
[0003] Fixed cameras are significantly affected by installation angle, monitoring range, and building obstruction, making it difficult to simultaneously cover motor vehicle lanes, non-motor vehicle lanes, and pedestrian crossings. This leads to issues such as missed detections, false detections, and trajectory interruptions for motor vehicle targets, non-motor vehicle targets, and pedestrian targets, and makes it difficult to accurately acquire the heterogeneous traffic flow status in complex intersections. Video streams acquired by drones suffer from problems such as changes in top-down perspective, perspective compression of edge regions, and significant changes in target scale. Existing target detection networks are prone to small target feature loss, target edge information attenuation, and cross-frame detection instability during multi-scale traffic target recognition, resulting in issues with the detection of motor vehicle targets, non-motor vehicle targets, and pedestrian targets. The detection accuracy is reduced; traditional multi-target tracking algorithms mainly rely on the intersection-union ratio of bounding boxes and motion information to associate trajectories. In scenarios with dense non-motorized vehicle traffic, frequent target occlusion, and multi-phase alternating release, trajectory drift, trajectory switching, and identity confusion are prone to occur, resulting in insufficient continuity of traffic trajectories and large deviations in traffic flow statistics. Existing traffic signal timing methods mostly focus on optimizing the efficiency of motor vehicle traffic, and do not adequately consider the traffic needs of non-motorized vehicles and pedestrians. It is difficult to coordinate and balance the traffic conflict relationship between motor vehicles, non-motorized vehicles, and pedestrians, resulting in increased waiting time, increased queue length, and decreased overall traffic efficiency in complex intersection areas.
[0004] Therefore, how to provide a traffic signal timing optimization method based on UAV edge computing is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a traffic signal timing optimization method based on UAV edge computing. This invention fully utilizes an improved YOLOv8 target detection network, an improved Bytetrack multi-target tracking algorithm, a perspective transformation algorithm, and an NSGA-II multi-target genetic algorithm. It details the processing steps of heterogeneous traffic target identification, heterogeneous traffic trajectory generation, traffic flow parameter extraction, comprehensive traffic efficiency function construction, and target signal timing scheme optimization in the target intersection area. It has the advantages of high heterogeneous traffic target detection accuracy, high trajectory correlation stability in complex traffic scenarios, strong non-motorized vehicle occlusion compensation capability, high accuracy of traffic flow parameter statistics, and good traffic signal timing optimization effect.
[0006] A traffic signal timing optimization method based on UAV edge computing according to an embodiment of the present invention includes the following steps:
[0007] S1. Collect real-time video streams of the target intersection area using drones, perform time-stamp alignment and coordinate calibration processing, and generate air-ground traffic data;
[0008] S2. Input the air and ground traffic data into the improved YOLOv8 target detection network, introduce a visual drift diffusion mechanism in the scale correlation module, identify motor vehicle targets, non-motor vehicle targets and pedestrian targets, and generate heterogeneous traffic targets;
[0009] S3. An improved Bytetrack multi-target tracking algorithm is used to perform trajectory association on the heterogeneous traffic targets. A phase trajectory constraint mechanism is introduced to generate motor vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories. If a non-motor vehicle target is occluded, Kalman filtering is used to perform trajectory prediction on the non-motor vehicle trajectory to generate a compensation trajectory. A perspective transformation algorithm is used to map the motor vehicle trajectory, compensation trajectory, and pedestrian trajectory to the world coordinate system to generate a heterogeneous traffic trajectory set.
[0010] S4. Extract traffic flow feature parameters based on the heterogeneous traffic trajectory set to generate a traffic flow parameter set;
[0011] S5. Calculate the weights of motor vehicles, non-motor vehicles, and pedestrians based on the traffic flow parameter set, and construct a comprehensive traffic efficiency function in combination with the traffic flow parameter set;
[0012] S6. Use the NSGA-II multi-objective genetic algorithm to perform signal timing optimization on the comprehensive traffic efficiency function to generate a target signal timing scheme;
[0013] S7. Send the target signal timing scheme to the roadside signal controller to execute signal control, and perform closed-loop feedback update based on the air and ground traffic data collected in the next control cycle, and output the traffic signal timing optimization result.
[0014] Optionally, S1 specifically includes:
[0015] Control the drone to hover at a preset height above the target intersection, and control the high-definition gimbal camera on the drone to collect real-time video streams of the target intersection area at a preset pitch angle;
[0016] The continuous image frames in the real-time video stream are time-stamp aligned according to the acquisition timestamp to generate an aligned video frame sequence.
[0017] Based on the location of the stop line, the boundary of the approach lane, the boundary of the non-motorized vehicle lane, and the boundary of the pedestrian crossing in the target intersection area, coordinate calibration processing is performed on the aligned video frame sequence to generate air and ground traffic data.
[0018] Optionally, the improved YOLOv8 target detection network specifically includes a backbone network, a scale correlation module, and a detection head;
[0019] The backbone network divides continuous image frames in air-ground traffic data into image blocks, and extracts texture features, edge features and semantic features from the image blocks. These features are then arranged according to the image resolution to form a feature map with a first spatial size and a feature map with a second spatial size, where the first size is larger than the second size.
[0020] The scale association module introduces a visual drift diffusion mechanism to perform scale association on the same target in feature maps of different scales. This mechanism includes: performing scale association on the same target in feature maps with a first spatial size and a second spatial size, where the targets include motor vehicle targets, non-motor vehicle targets, and pedestrian targets; extracting the bounding box width and height changes corresponding to the same target in adjacent image frames, and generating a target scale change gradient based on these changes; and extracting the distance from the target bounding box center point to the image center point and the distance from the target bounding box center point to the image edge. The distance to the edge and the change in the aspect ratio of the target bounding box are calculated. Based on the distance from the center point of the target bounding box to the center point of the image, the distance from the center point of the target bounding box to the edge of the image, and the change in the aspect ratio of the target bounding box, the perspective compression of the edge region is generated. The area change of the bounding box corresponding to the same target in adjacent image frames is extracted, and the cross-frame size offset is generated based on the area change of the bounding box. Based on the target scale change gradient, the perspective compression of the edge region, and the cross-frame size offset, the scale drift correspondence between the same target position in the feature map with a spatial size of the first size and the feature map with a spatial size of the second size is established, and the scale drift field is generated.
[0021] Based on the scale drift field, the feature compensation positions and feature compensation weights between the feature map with a first spatial size and the feature map with a second spatial size are determined; the target edge features in the feature map with a first spatial size are compensated to the corresponding target positions in the feature map with a second spatial size, and the target semantic features in the feature map with a second spatial size are compensated to the corresponding target positions in the feature map with a first spatial size, thereby generating stable target response features;
[0022] The detection head determines the target center point position, target bounding box width, target bounding box height, and target category confidence based on the stable target response characteristics, and outputs the bounding box coordinates and category confidence of motor vehicle targets, non-motor vehicle targets, and pedestrian targets, generating heterogeneous traffic targets.
[0023] Optionally, the improved Bytetrack multi-target tracking algorithm is used to perform trajectory association on the heterogeneous traffic targets, introducing a phase trajectory constraint mechanism to generate vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories, specifically as follows:
[0024] Based on the relationship between the category confidence scores of motor vehicle targets, non-motor vehicle targets, and pedestrian targets and the preset confidence thresholds, heterogeneous traffic targets with category confidence scores greater than or equal to the preset confidence thresholds are classified into a high-confidence target set, and heterogeneous traffic targets with category confidence scores less than the preset confidence thresholds are classified into a low-confidence target set.
[0025] The motor vehicle passage area, non-motor vehicle passage area, and pedestrian passage area are determined based on the location of the parking line, the boundary of the entrance lane, the boundary of the non-motor vehicle lane, and the boundary of the pedestrian crossing.
[0026] Based on the change in the position of the center point of the bounding box of the heterogeneous traffic target in consecutive image frames, the motion direction of the heterogeneous traffic target is determined.
[0027] Establish phase trajectory constraint relationships based on the traffic area, direction of movement, stop line location, and phase release status of heterogeneous traffic targets;
[0028] The high-confidence target set is associated with the existing trajectory according to the intersection-union ratio of the bounding box, the consistency of the motion direction, and the phase trajectory constraint relationship to generate the first trajectory association result;
[0029] The low-confidence target set is associated with the unassociated trajectories in the first trajectory association result according to the bounding box intersection-union ratio, motion direction consistency and phase trajectory constraint relationship to generate the second trajectory association result;
[0030] Based on the first trajectory association result and the second trajectory association result, update the trajectory identifier, trajectory category and trajectory location to generate motor vehicle trajectory, non-motor vehicle trajectory and pedestrian trajectory.
[0031] Optionally, if the non-motorized vehicle target is obstructed, Kalman filtering is used to predict the trajectory of the non-motorized vehicle and generate a compensated trajectory, specifically as follows:
[0032] Calculate the speed and direction of movement of the non-motorized vehicle based on the world coordinates and timestamps of adjacent trajectory points in the non-motorized vehicle trajectory.
[0033] If the detection box for the non-motorized vehicle trajectory is missing in consecutive image frames, and the predicted trajectory position corresponding to the non-motorized vehicle trajectory is located within the non-motorized vehicle passage area, then it is determined that the non-motorized vehicle target is occluded.
[0034] The position of the non-motorized vehicle trajectory point, the speed of the non-motorized vehicle, and the direction of the non-motorized vehicle are used as the state variables of the Kalman filter.
[0035] Predict the location of non-motorized vehicle trajectory points in missing image frames based on Kalman filter state variables; generate a compensation trajectory based on the associated trajectory points in the non-motorized vehicle trajectory and the location of non-motorized vehicle trajectory points in the missing image frames.
[0036] Optionally, the step of using a perspective transformation algorithm to map the vehicle trajectory, compensated trajectory, and pedestrian trajectory to the world coordinate system to generate a heterogeneous traffic trajectory set is as follows:
[0037] A calibration point in the image coordinate system is selected based on the location of the parking line, the boundary of the approach lane, the boundary of the non-motorized vehicle lane, and the boundary of the pedestrian crossing; the world coordinate point corresponding to the calibration point is determined based on the actual road plane coordinates corresponding to the target intersection area; and the perspective transformation matrix is calculated based on the calibration point and the world coordinate point.
[0038] Based on the perspective transformation matrix, the trajectory points in the vehicle trajectory, the compensated trajectory, and the pedestrian trajectory are mapped from the image coordinate system to the world coordinate system;
[0039] The mapped trajectory points are merged according to trajectory identifier, trajectory category, world coordinate position, and timestamp to generate a heterogeneous traffic trajectory set.
[0040] Optionally, S4 specifically includes:
[0041] Based on the trajectory categories, world coordinates, and timestamps of the heterogeneous traffic trajectory set, the vehicle trajectory, compensation trajectory, and pedestrian trajectory are matched to the corresponding approach lane, turning direction, and signal phase, respectively.
[0042] Based on the vehicle trajectories, compensation trajectories, and pedestrian trajectories matched to the same entrance lane, the same turning direction, and the same signal phase, the number of motor vehicles, non-motor vehicles, and pedestrians passing through the stop line per unit time is counted to generate traffic flow parameters.
[0043] Based on the world coordinates of the vehicle trajectories and compensation trajectories in the direction from the stop line upstream, the queue lengths of motor vehicles and non-motor vehicles are calculated, and traffic flow queuing parameters are generated.
[0044] Based on the timestamps of vehicle trajectories, compensation trajectories, and pedestrian trajectories when entering the detection area and when passing the parking line, calculate the waiting time for motor vehicles, non-motor vehicles, and pedestrians, and generate traffic flow waiting parameters.
[0045] A traffic flow parameter set is generated based on traffic flow traffic parameters, traffic flow queuing parameters, and traffic flow waiting parameters.
[0046] Optionally, S5 specifically includes:
[0047] Based on the traffic flow parameters, traffic flow queuing parameters, and traffic flow waiting parameters in the traffic flow parameter set, the vehicle density index, non-motor vehicle density index, and pedestrian density index are determined respectively.
[0048] Motor vehicle weights are calculated based on motor vehicle density indicators and preset motor vehicle basic priority coefficients; non-motor vehicle weights are calculated based on non-motor vehicle density indicators and preset non-motor vehicle basic priority coefficients; pedestrian weights are calculated based on pedestrian density indicators and preset pedestrian basic priority coefficients.
[0049] The weights of motor vehicles, non-motor vehicles, and pedestrians are normalized to generate normalized motor vehicle weights, normalized non-motor vehicle weights, and normalized pedestrian weights.
[0050] A comprehensive traffic efficiency function is constructed based on traffic flow parameters, normalized motor vehicle weights, normalized non-motor vehicle weights, and normalized pedestrian weights.
[0051] Optionally, S6 specifically includes:
[0052] Encode the green light duration of the signal phase at the target intersection into a real number vector to generate the initial timing individual;
[0053] The initial timing individuals are constrained and screened based on green light duration constraints, cycle duration constraints, and phase sequence constraints to generate a timing population.
[0054] The combined traffic efficiency function is used to calculate the traffic efficiency of motor vehicles, non-motor vehicles, and pedestrians for the timed individuals in the timed population, and a multi-objective fitness value is generated.
[0055] Based on the multi-objective fitness values, the timing individuals in the timing population are non-dominated and sorted to generate non-dominated levels; crowding distance is calculated using the fitness distance between timing individuals within the same non-dominated level.
[0056] Based on the non-dominance level and the crowding distance, parent timing individuals are selected from the timing population; the green light duration of the signal phase in the parent timing individuals is cross-combined and mutated to generate child timing individuals;
[0057] The parent and child mating individuals are merged, and target mating individuals are selected according to the non-dominance level and the crowding distance.
[0058] A target signal timing scheme is generated based on the green light duration of the signal phase in the target timing individual.
[0059] Optionally, S7 specifically includes:
[0060] The green light duration, signal cycle duration, and phase sequence of the target signal timing scheme are sent to the roadside signal controller; the roadside signal controller controls the signal light status of the target intersection according to the green light duration, signal cycle duration, and phase sequence.
[0061] In the next control cycle, real-time video streams of the target intersection area are acquired to generate air and ground traffic data for the next control cycle; based on the air and ground traffic data of the next control cycle, the heterogeneous traffic targets, heterogeneous traffic trajectory sets, traffic flow parameter sets, comprehensive traffic efficiency functions, and target signal timing schemes are updated.
[0062] The updated target signal timing scheme is sent to the roadside signal controller, which outputs the traffic signal timing optimization results.
[0063] The beneficial effects of this invention are:
[0064] (1) By improving the combination of YOLOv8 target detection network and visual drift diffusion mechanism, in order to address the problems of significant changes in traffic target scale, severe perspective compression in edge regions and easy loss of small target features in UAV top-down scenes, a scale drift field is established by using target scale change gradient, perspective compression in edge regions and cross-frame size offset to achieve feature compensation between feature maps of different scales, thereby improving the detection accuracy and detection stability of motor vehicle targets, non-motor vehicle targets and pedestrian targets;
[0065] (2) In the improved Bytetrack multi-target tracking algorithm, a phase trajectory constraint mechanism is introduced. In response to the problems of frequent traffic target occlusion, serious trajectory drift and high trajectory switching rate in complex intersection areas, a phase trajectory constraint relationship is established by using the passage area, movement direction, stop line position and phase release status to achieve stable trajectory association between high confidence target set and low confidence target set, thereby improving the continuity and association accuracy of motor vehicle trajectory, non-motor vehicle trajectory and pedestrian trajectory.
[0066] (3) By combining Kalman filtering and perspective transformation algorithm, in order to address the problems of missing detection boxes, trajectory interruption and spatial position deviation of non-motorized vehicle targets in dense traffic scenarios, compensation trajectory is generated by using non-motorized vehicle speed, non-motorized vehicle direction and trajectory prediction position, and the trajectory mapping between image coordinate system and world coordinate system is completed by combining perspective transformation matrix, thereby improving the position accuracy of heterogeneous traffic trajectory set and the accuracy of traffic flow parameter statistics.
[0067] (4) Based on the comprehensive traffic efficiency function and NSGA-II multi-objective genetic algorithm, in order to address the problem that existing traffic signal timing methods cannot take into account the traffic needs of motor vehicles, non-motor vehicles and pedestrians, a comprehensive traffic efficiency function is constructed by using motor vehicle weight, non-motor vehicle weight and pedestrian weight, and the target timing individuals are selected by combining non-dominated sorting and congestion distance, so as to realize the dynamic optimization of the target signal timing scheme, improve the overall traffic efficiency of the target intersection area and reduce traffic waiting time and queue length. Attached Figure Description
[0068] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0069] Figure 1 This is a flowchart of a traffic signal timing optimization method based on UAV edge computing proposed in this invention;
[0070] Figure 2 This is a schematic diagram of the improved YOLOv8 target detection network proposed in this invention;
[0071] Figure 3 This is a data flow diagram of a traffic signal timing optimization method based on UAV edge computing proposed in this invention. Detailed Implementation
[0072] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0073] refer to Figures 1-3 A traffic signal timing optimization method based on UAV edge computing includes the following steps:
[0074] S1. Collect real-time video streams of the target intersection area using drones, perform time-stamp alignment and coordinate calibration processing, and generate air-ground traffic data;
[0075] S2. Input air and ground traffic data into the improved YOLOv8 target detection network, introduce a visual drift diffusion mechanism in the scale association module, identify motor vehicle targets, non-motor vehicle targets and pedestrian targets, and generate heterogeneous traffic targets;
[0076] S3. An improved Bytetrack multi-target tracking algorithm is used to perform trajectory association on heterogeneous traffic targets. A phase trajectory constraint mechanism is introduced to generate motor vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories. If a non-motor vehicle target is occluded, Kalman filtering is used to perform trajectory prediction on the non-motor vehicle trajectory to generate a compensation trajectory. A perspective transformation algorithm is used to map the motor vehicle trajectory, compensation trajectory, and pedestrian trajectory to the world coordinate system to generate a heterogeneous traffic trajectory set.
[0077] S4. Extract traffic flow feature parameters based on heterogeneous traffic trajectory sets to generate traffic flow parameter sets;
[0078] S5. Calculate the weights of motor vehicles, non-motor vehicles, and pedestrians based on the traffic flow parameter set, and construct a comprehensive traffic efficiency function in combination with the traffic flow parameter set;
[0079] S6. Use the NSGA-II multi-objective genetic algorithm to optimize the signal timing of the comprehensive traffic efficiency function and generate the target signal timing scheme.
[0080] S7. Send the target signal timing scheme to the roadside signal controller to execute signal control, and perform closed-loop feedback update based on the air and ground traffic data collected in the next control cycle, and output the traffic signal timing optimization result.
[0081] In this embodiment, S1 specifically refers to:
[0082] Control the drone to hover above the target intersection. The hovering height is determined based on the detection range from the parking line to the upstream, the coverage range of the pedestrian crossing, and the field of view of the high-definition gimbal camera. Preferably, the hovering height is between 50m and 100m.
[0083] The high-definition PTZ camera is controlled to face the center area of the target intersection. The pitch angle is determined based on the coverage of the boundary of the approach lane, the boundary of the non-motorized vehicle lane, the boundary of the pedestrian crossing, and the parking line in the picture. Preferably, the pitch angle is between -30° and -60°.
[0084] The system acquires real-time video streams of the target intersection area, covering the stop line, approach lane boundary, non-motorized vehicle lane boundary, pedestrian crossing boundary, and the area from the stop line to 100m upstream. It extracts continuous image frames from the real-time video stream and records the acquisition timestamps corresponding to the continuous image frames. The continuous image frames are arranged in ascending order according to the acquisition timestamps, and the time intervals between adjacent image frames are adjusted to the same sampling period to generate an aligned video frame sequence.
[0085] Mark the endpoints of the parking line, the boundary endpoints of the approach lane, the boundary endpoints of the non-motorized vehicle lane, and the corner points of the pedestrian crossing in the aligned video frame sequence to generate image coordinate calibration points;
[0086] Based on the road plan of the target intersection or on-site measurement results, determine the road plan coordinate points corresponding to the image coordinate calibration points; establish the coordinate correspondence between the image coordinate calibration points and the road plan coordinate points, and generate the coordinate calibration relationship;
[0087] The aligned video frame sequence, acquisition timestamps, and coordinate calibration relationships are merged into air and ground traffic data for the target intersection area.
[0088] In this embodiment, the improved YOLOv8 target detection network specifically includes a backbone network, a scale correlation module, and a detection head;
[0089] The backbone network receives continuous image frames from air-to-ground traffic data. Each continuous image frame contains image coordinates and a collection timestamp. The continuous image frames are input into the network with the same pixel size and divided into image blocks with a fixed pixel step size. Each image block retains the image coordinate range and collection timestamp from the continuous image frames. Gray-level and color variation distributions are extracted within each image block to generate texture features. Pixel gray-level abrupt change locations and target contour locations are extracted within each image block to generate edge features. Spatial relationships between target areas, road areas, parking line areas, non-motorized vehicle lane areas, and pedestrian crossing areas are extracted within each image block to generate semantic features. Texture features, edge features, and semantic features are mapped to corresponding image frames according to image coordinates and arranged according to image resolution to form a feature map with a first spatial size and a feature map with a second spatial size, where the first size is larger than the second size.
[0090] The scale association module introduces a view drift diffusion mechanism to perform scale association on the same target in feature maps of different scales. This mechanism includes: receiving a feature map with a first spatial size and a feature map with a second spatial size; locating the position of candidate targets corresponding to the same traffic target in both feature maps; traffic targets including motor vehicle targets, non-motor vehicle targets, and pedestrian targets; extracting the bounding box width and height corresponding to the same traffic target in adjacent image frames; and calculating the change in bounding box width as the difference between the width of the bounding box in the subsequent image frame and the width of the bounding box in the previous image frame. The bounding box width in the image frame; the change in bounding box height is the difference between the bounding box height in the previous image frame and the bounding box height in the subsequent image frame; the target scale gradient consists of the changes in bounding box width and height, representing the direction and magnitude of scale change for the same traffic target in adjacent image frames; the center point of the traffic target bounding box, the image center point, and the image edge points are extracted; the distance from the center point of the traffic target bounding box to the image center point is the pixel distance between the center point of the traffic target bounding box and the image center point; the distance from the center point of the traffic target bounding box to the image edge is the distance between the traffic target bounding box center point and the image center point. The minimum pixel distance from the center point of the bounding box to the left, right, top, and bottom edges of the image is used; the aspect ratio change of the traffic target bounding box is the ratio of the width to the height of the bounding box in the subsequent image frame minus the ratio of the width to the height of the bounding box in the previous image frame; the perspective compression of the edge region is composed of the distance from the center point of the traffic target bounding box to the center point of the image, the distance from the center point of the traffic target bounding box to the edge of the image, and the aspect ratio change of the traffic target bounding box. The perspective compression of the edge region indicates the degree of size compression of the traffic target in the image edge region due to the top-down perspective effect; The bounding box area corresponding to the same traffic target in adjacent image frames is taken; the bounding box area is the product of the bounding box width and the bounding box height; the cross-frame size offset is the bounding box area in the later image frame minus the bounding box area in the previous image frame; the target scale change gradient, edge region perspective compression, and cross-frame size offset are written into the candidate target positions corresponding to the same traffic target in the feature map with a spatial size of the first size and the feature map with a spatial size of the second size; the scale drift field consists of the candidate target positions corresponding to the same traffic target, the target scale change gradient, the edge region perspective compression, and the cross-frame size offset;
[0091] The feature compensation position between a feature map with a first spatial size and a feature map with a second spatial size is determined based on the scale drift field. The feature compensation position is the corresponding position of the same traffic target in the feature map with the first spatial size and the feature map with the second spatial size after correction by the candidate target position. The feature compensation weight between the feature map with the first spatial size and the feature map with the second spatial size is determined based on the scale drift field. The feature compensation weight is jointly determined by the target scale change gradient, the perspective compression of the edge region, and the cross-frame size offset. When the target scale change gradient, the perspective compression of the edge region, and the cross-frame size offset increase, the feature compensation weight of the corresponding traffic target position is increased. The target edge features located at the feature compensation position in the feature map with the first spatial size are written into the corresponding target position in the feature map with the second spatial size. The target semantic features located at the feature compensation position in the feature map with the second spatial size are written into the corresponding target position in the feature map with the first spatial size. The stable target response feature is composed of the feature map with the second spatial size after writing the target edge features and the feature map with the first spatial size after writing the target semantic features.
[0092] The detection head receives stable target response features and locates the center point of traffic targets within these features. It then determines the target bounding box width and height based on the target edge features corresponding to the center point location. Finally, it determines the target category confidence scores for motor vehicle, non-motor vehicle, and pedestrian targets based on the target semantic features corresponding to the center point location. The center point location, bounding box width, height, and category confidence scores are then converted into bounding box coordinates and category confidence scores for motor vehicle, non-motor vehicle, and pedestrian targets, respectively. The heterogeneous traffic targets consist of bounding box coordinates, category confidence scores, image coordinates, and acquisition timestamps for motor vehicle, non-motor vehicle, and pedestrian targets.
[0093] In this embodiment, both the improved YOLOv8 target detection network and the YOLOv8 target detection network use a backbone network to extract image features and a detection head to output target bounding box coordinates and target category confidence. Both can perform target localization and target classification in consecutive image frames. The improved YOLOv8 target detection network introduces a visual drift diffusion mechanism in the scale association module to establish a scale drift correspondence between the same traffic target in a feature map with a first spatial size and a feature map with a second spatial size. The visual drift diffusion mechanism combines the target scale change gradient, edge region perspective compression, and cross-frame size offset to characterize the scale drift state of traffic targets in the UAV top-down scene. The target scale change gradient reflects the direction and magnitude of the size change of the traffic target in adjacent image frames, the edge region perspective compression reflects the degree of size compression caused by the top-down perspective when the traffic target is located in the image edge region, and the cross-frame size offset reflects the area change state of the traffic target in consecutive image frames. The scale association module is based on the scale drift... The field performs bidirectional compensation on the target edge features and target semantic features between the feature map with a first spatial size and the feature map with a second spatial size. The target edge features in the feature map with a first spatial size are written into the corresponding target positions in the feature map with a second spatial size, and the target semantic features in the feature map with a second spatial size are written into the corresponding target positions in the feature map with a first spatial size. The improved YOLOv8 target detection network can enhance the cross-scale feature association ability between motor vehicle targets, non-motor vehicle targets, and pedestrian targets in UAV-viewed traffic scenes, reduce the impact of target size distortion caused by perspective compression when traffic targets are located in the image edge region, reduce the impact of scale jump of traffic targets in consecutive image frames, enhance the target position correspondence ability between the feature map with a first spatial size and the feature map with a second spatial size, improve the target detection stability of non-motor vehicle targets and pedestrian targets in complex traffic scenes, and improve the accuracy of bounding box coordinates and category confidence of heterogeneous traffic targets.
[0094] In this embodiment, an improved Bytetrack multi-target tracking algorithm is used to perform trajectory association on heterogeneous traffic targets. A phase trajectory constraint mechanism is introduced to generate motor vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories, specifically as follows:
[0095] Obtain the bounding box coordinates, category confidence, image coordinates, and acquisition timestamp from heterogeneous traffic targets;
[0096] The preset confidence threshold is determined by the category confidence distribution of motor vehicle targets, non-motor vehicle targets, and pedestrian targets in the historical video stream of the target intersection. The confidence value that can distinguish between valid targets and noisy targets in the category confidence distribution is selected as the preset confidence threshold.
[0097] The category confidence scores corresponding to motor vehicle targets, non-motor vehicle targets, and pedestrian targets are compared with preset confidence thresholds respectively; heterogeneous traffic targets with category confidence scores greater than or equal to the preset confidence thresholds are classified into a high-confidence target set; heterogeneous traffic targets with category confidence scores less than the preset confidence thresholds are classified into a low-confidence target set.
[0098] The motor vehicle passage area is obtained by defining the driving range of motor vehicles based on the location of the parking line and the boundary of the entrance lane; the non-motor vehicle passage area is obtained by defining the driving range of non-motor vehicles based on the boundary of the non-motor vehicle lane and the location of the parking line; and the pedestrian passage area is obtained by defining the pedestrian crossing range based on the boundary of the pedestrian crossing and the location of the parking line.
[0099] Extract the center point position of the bounding box of the heterogeneous traffic target in adjacent consecutive image frames; subtract the center point position of the bounding box in the previous image frame from the center point position of the bounding box in the subsequent image frame to obtain the center point displacement vector; use the direction of the center point displacement vector as the motion direction of the heterogeneous traffic target.
[0100] Based on the positional relationship between the center point of the bounding box of the heterogeneous traffic target and the motor vehicle traffic area, non-motor vehicle traffic area and pedestrian traffic area, the traffic area where the heterogeneous traffic target is located is determined.
[0101] Phase trajectory constraints are established based on the traffic area, direction of movement, stop line location, and phase release status of the heterogeneous traffic target. The phase trajectory constraints include traffic area constraints, direction of movement constraints, stop line constraints, and phase status constraints.
[0102] The traffic area constraints are: motor vehicle targets are located in the motor vehicle traffic area, non-motor vehicle targets are located in the non-motor vehicle traffic area, and pedestrian targets are located in the pedestrian traffic area;
[0103] The motion direction constraint is that the motion direction of the heterogeneous traffic target is consistent with the direction of traffic at the entrance lane or the direction of traffic at the pedestrian crossing corresponding to the traffic area.
[0104] The stop line constraint allows the trajectory points of motor vehicle targets and non-motor vehicle targets to approach or pass the stop line in the direction of travel of the approach lane, and the trajectory points of pedestrian targets to cross the pedestrian crossing boundary in the direction of travel of the pedestrian crossing; the phase state constraint allows heterogeneous traffic targets in the corresponding approach lane, non-motor vehicle lane or pedestrian crossing to pass the stop line or pedestrian crossing boundary in the phase release state.
[0105] Calculate the intersection-union ratio of bounding boxes between heterogeneous traffic targets in the high-confidence target set and existing trajectories; calculate the directional angle between the motion direction of the heterogeneous traffic targets in the high-confidence target set and the motion direction of the end of existing trajectories;
[0106] If the bounding box intersection-union ratio satisfies the intersection-union ratio threshold, the direction angle satisfies the direction consistency threshold, and the phase trajectory constraint relationship satisfies the traffic area constraint, movement direction constraint, stop line constraint, and phase state constraint, then the heterogeneous traffic targets in the high-confidence target set will be associated with the corresponding existing trajectories to generate the first trajectory association result.
[0107] Existing trajectories that have not entered the first trajectory association result are considered as unassociated trajectories; the intersection-union ratio of bounding boxes between heterogeneous traffic targets in the low-confidence target set and unassociated trajectories is calculated; the directional angle between the motion direction of heterogeneous traffic targets in the low-confidence target set and the motion direction of the end of the unassociated trajectory is calculated.
[0108] If the bounding box intersection-union ratio meets the intersection-union ratio threshold, the direction angle meets the direction consistency threshold, and the phase trajectory constraint relationship meets the traffic area constraint, movement direction constraint, stop line constraint, and phase state constraint, then the heterogeneous traffic targets in the low confidence target set will be associated with the corresponding unassociated trajectories to generate the second trajectory association result.
[0109] Based on the first trajectory association result and the second trajectory association result, retain the same trajectory identifier for the same heterogeneous traffic target in consecutive image frames; update the trajectory category based on the category label of the heterogeneous traffic target; update the trajectory position based on the bounding box center point position and acquisition timestamp of the heterogeneous traffic target.
[0110] Trajectories categorized as motor vehicle targets are classified as motor vehicle trajectories; trajectories categorized as non-motor vehicle targets are classified as non-motor vehicle trajectories; and trajectories categorized as pedestrian targets are classified as pedestrian trajectories.
[0111] In this embodiment, both the improved Bytetrack multi-target tracking algorithm and the Bytetrack multi-target tracking algorithm adopt a two-stage trajectory association method using a high-confidence target set and a low-confidence target set. Both are associated based on the spatial positional correspondence between heterogeneous traffic targets and existing trajectories by comparing and intersecting bounding boxes, and generate continuous trajectories based on trajectory identifiers, trajectory categories, and trajectory positions. The improved Bytetrack multi-target tracking algorithm introduces a phase trajectory constraint mechanism in the two-stage trajectory association process, incorporating traffic area constraints, motion direction constraints, stop line constraints, and phase state constraints into the trajectory association process. The traffic area constraint limits the trajectory distribution range of motor vehicle targets, non-motor vehicle targets, and pedestrian targets based on the stop line position, approach lane boundary, non-motorized vehicle lane boundary, and pedestrian crossing boundary. The motion direction constraint limits the trajectory extension direction based on the displacement direction of the bounding box center point of the heterogeneous traffic targets in continuous image frames. The stop line constraint limits the trajectory based on the positional relationship between the heterogeneous traffic targets and the stop line. The improved Bytetrack multi-target tracking algorithm, in the process of associating high-confidence target sets, low-confidence target sets, and existing trajectories, simultaneously employs bounding box intersection-union ratio, motion direction consistency, and phase trajectory constraint relationships for joint screening, reducing the probability of erroneous trajectory association between motor vehicle targets, non-motor vehicle targets, and pedestrian targets in UAV-viewed traffic scenarios. This improved algorithm also reduces trajectory drift of traffic targets in intersection areas, trajectory switching near stop lines, pedestrian crossings, and intersections of approach lanes, reduces interference from noisy targets in low-confidence target sets on trajectory association results, improves the category consistency and temporal continuity among motor vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories, and enhances the trajectory position accuracy and trajectory stability in heterogeneous traffic trajectory sets.
[0112] In this embodiment, if the non-motorized vehicle target is obstructed, Kalman filtering is used to predict the trajectory of the non-motorized vehicle and generate a compensated trajectory, specifically as follows:
[0113] Extract the world coordinates and timestamps of adjacent trajectory points in the non-motorized vehicle trajectory. The world coordinates include horizontal and vertical coordinates.
[0114] The horizontal velocity is obtained by subtracting the horizontal coordinate of the previous trajectory point from the horizontal coordinate of the next trajectory point and dividing by the time difference between the timestamps of the next and previous trajectory points. The vertical velocity is obtained by subtracting the vertical coordinate of the previous trajectory point from the vertical coordinate of the next trajectory point and dividing by the time difference between the timestamps of the next and previous trajectory points.
[0115] The speed of a non-motorized vehicle is composed of its lateral and longitudinal speeds; the direction of movement of the non-motorized vehicle is composed of the direction from the previous trajectory point to the next trajectory point.
[0116] Extract the bounding box association state and trajectory prediction position of non-motorized vehicle trajectories in consecutive image frames;
[0117] If the detection box association status of the non-motorized vehicle trajectory in consecutive image frames is missing, and the trajectory prediction position corresponding to the non-motorized vehicle trajectory is located in the non-motorized vehicle passage area, then it is determined that the non-motorized vehicle target is occluded.
[0118] The position of the non-motorized vehicle trajectory point, the lateral speed, the longitudinal speed, and the direction of the non-motorized vehicle movement are used to form the state variables of the Kalman filter. The position of the non-motorized vehicle trajectory point in the Kalman filter state variables is used as the state position term, the lateral speed and the longitudinal speed are used as the state speed term, and the direction of the non-motorized vehicle movement is used as the state direction term.
[0119] Based on the state position item, state velocity item, state direction item and the timestamp corresponding to the missing image frame, the predicted position of the non-motor vehicle trajectory point in the missing image frame is obtained;
[0120] If the predicted location of the non-motorized vehicle trajectory point is located within the non-motorized vehicle passage area, and the predicted location of the non-motorized vehicle trajectory point extends along the direction of movement of the non-motorized vehicle, then the predicted location of the non-motorized vehicle trajectory point is retained.
[0121] Arrange the associated trajectory points in the non-motorized vehicle trajectory in ascending order of timestamp; insert the predicted positions of the non-motorized vehicle trajectory points in the missing image frames between the associated trajectory points according to their timestamps; merge the associated trajectory points and the inserted predicted positions of the non-motorized vehicle trajectory points to generate a compensated trajectory.
[0122] In this embodiment, a perspective transformation algorithm is used to map the vehicle trajectory, compensated trajectory, and pedestrian trajectory to the world coordinate system, generating a heterogeneous traffic trajectory set, specifically:
[0123] Read the parking line position, approach lane boundary, non-motorized vehicle lane boundary, and pedestrian crossing boundary from the aligned video frame sequence of the target intersection area; select the endpoints of the parking line, approach lane boundary, non-motorized vehicle lane boundary, and pedestrian crossing corner as calibration points in the image coordinate system; record the pixel x-coordinate and pixel y-coordinate of the calibration points in the image coordinate system;
[0124] Based on the actual road plane coordinates corresponding to the target intersection area, determine the world coordinate points corresponding to the endpoints of the stop line, the boundary endpoints of the approach lane, the boundary endpoints of the non-motorized vehicle lane, and the corner points of the pedestrian crossing; record the lateral and longitudinal coordinates of the world coordinate points in the road plane coordinate system; and form a coordinate correspondence group between each calibration point and its corresponding world coordinate point.
[0125] The perspective transformation matrix between the image coordinate system and the world coordinate system is determined based on the coordinate correspondence set; the perspective transformation matrix is determined by the homography relationship between the pixel horizontal and vertical coordinates in the image coordinate system and the horizontal and vertical coordinates in the world coordinate system.
[0126] Extract trajectory points from vehicle trajectories, compensated trajectories, and pedestrian trajectories; trajectory points include trajectory identifier, trajectory category, image coordinate position, and timestamp;
[0127] Multiply the image coordinates of the trajectory point by the perspective transformation matrix to obtain the homogeneous coordinates of the trajectory point in the world coordinate system; divide the horizontal component of the homogeneous coordinates by the scale component to obtain the horizontal coordinates of the trajectory point in the world coordinate system; divide the vertical component of the homogeneous coordinates by the scale component to obtain the vertical coordinates of the trajectory point in the world coordinate system; combine the horizontal and vertical coordinates of the trajectory point in the world coordinate system to form the world coordinate position.
[0128] Track points belonging to the same track are grouped into the same track sequence according to track identifiers; track points in the same track sequence are arranged in ascending order according to timestamps; and track sequences are divided into motor vehicle track sequences, non-motor vehicle track sequences, and pedestrian track sequences according to track category.
[0129] The trajectory sequences of motor vehicles, non-motor vehicles, and pedestrians are merged to generate a heterogeneous traffic trajectory set.
[0130] In this embodiment, S4 specifically refers to:
[0131] Read the trajectory identifier, trajectory category, world coordinate position and timestamp from the heterogeneous traffic trajectory set; divide the heterogeneous traffic trajectory set into motor vehicle trajectory, compensated trajectory and pedestrian trajectory according to the trajectory category;
[0132] The road plane range corresponding to the entrance lane is determined based on the boundary of the entrance lane; the vehicle trajectory and the compensation trajectory are matched to the corresponding entrance lane based on the positional relationship between the world coordinate position of the trajectory points in the vehicle trajectory and the road plane range corresponding to the entrance lane; the pedestrian trajectory is matched to the corresponding pedestrian crossing based on the positional relationship between the world coordinate position of the trajectory points in the pedestrian trajectory and the boundary of the pedestrian crossing.
[0133] Based on the direction of change of the world coordinate position of adjacent trajectory points in the vehicle trajectory and the compensation trajectory, the turning direction corresponding to the vehicle trajectory and the compensation trajectory is determined. The turning direction includes the straight direction, the left turn direction and the right turn direction. Based on the direction of change of the world coordinate position of adjacent trajectory points in the pedestrian trajectory, the crossing direction corresponding to the pedestrian trajectory is determined. Based on the approach lane, turning direction, crossing direction and phase release status, the vehicle trajectory, compensation trajectory and pedestrian trajectory are matched to the corresponding signal phase.
[0134] Using a fixed statistical period as the unit of time, the system counts the number of vehicle trajectories crossing the stop line within the same approach lane, turning direction, and signal phase to generate vehicle traffic volume; using a fixed statistical period as the unit of time, the system counts the number of compensated trajectories crossing the stop line within the same approach lane, turning direction, and signal phase to generate non-motorized vehicle traffic volume; using a fixed statistical period as the unit of time, the system counts the number of pedestrian trajectories crossing the pedestrian crossing boundary within the same pedestrian crossing, crossing direction, and signal phase to generate pedestrian traffic volume.
[0135] Traffic flow parameters are composed of motor vehicle traffic volume, non-motor vehicle traffic volume, and pedestrian traffic volume;
[0136] The queuing statistics range from the parking line to the upstream direction is determined based on the parking line location and the direction of traffic in the entrance lane; the number of vehicle trajectory points with speeds less than the parking judgment speed within the queuing statistics range and their corresponding world coordinate positions are counted to generate a vehicle queuing point set;
[0137] The number of compensation trajectory points with speeds less than the stopping determination speed within the queuing statistics range and their corresponding world coordinate positions are counted to generate a non-motorized vehicle queuing point set; the stopping determination speed is determined by the speed distribution of historical stationary target trajectories and low-speed target trajectories at the target intersection, and a speed value that can distinguish between a stopped state and a slow-moving state is selected.
[0138] The motor vehicle queue length is generated based on the road plane distance between the motor vehicle trajectory point furthest from the stop line and the stop line; the non-motor vehicle queue length is generated based on the road plane distance between the non-motor vehicle queue point compensation trajectory point furthest from the stop line and the stop line; the motor vehicle queue length and the non-motor vehicle queue length are combined to form traffic flow queuing parameters.
[0139] The detection zone is determined based on the boundary of the detection zone and the location of the parking line. The boundary of the detection zone is determined by the fixed road surface area from the parking line to the upstream direction.
[0140] Extract the timestamps of the vehicle's trajectory entering the detection area and passing the stop line; the vehicle waiting time is the timetamp of the vehicle's trajectory passing the stop line minus the timestamp of the vehicle's trajectory entering the detection area.
[0141] Extract the timestamps of the compensated trajectory entering the detection area and passing the stop line; the waiting time for non-motorized vehicles is the timestamp of the compensated trajectory passing the stop line minus the timestamp of the compensated trajectory entering the detection area.
[0142] Extract the timestamps of pedestrian trajectory entering and leaving the pedestrian crossing boundary; pedestrian waiting time is the timestamp of pedestrian trajectory leaving the pedestrian crossing boundary minus the timestamp of pedestrian trajectory entering the pedestrian crossing boundary; combine the waiting time of motor vehicles, the waiting time of non-motor vehicles, and the waiting time of pedestrians to form traffic flow waiting parameters;
[0143] Traffic flow parameters, traffic flow queuing parameters, and traffic flow waiting parameters are categorized according to approach lane, turning direction, crossing direction, and signal phase to generate a traffic flow parameter set.
[0144] In this embodiment, S5 specifically refers to:
[0145] Read the traffic flow parameters, traffic flow queuing parameters, and traffic flow waiting parameters from the traffic flow parameter set; the traffic flow parameters include the volume of motor vehicles, the volume of non-motor vehicles, and the volume of pedestrians; the traffic flow queuing parameters include the queue length of motor vehicles and the queue length of non-motor vehicles; the traffic flow waiting parameters include the waiting time of motor vehicles, the waiting time of non-motor vehicles, and the waiting time of pedestrians.
[0146] The arrival rate of motor vehicles is obtained based on the traffic volume of motor vehicles and the statistical period; the arrival rate of non-motor vehicles is obtained based on the traffic volume of non-motor vehicles and the statistical period; and the arrival rate of pedestrians is obtained based on the traffic volume of pedestrians and the statistical period.
[0147] The vehicle density index consists of vehicle arrival rate, vehicle waiting time, and vehicle unit throughput capacity. The vehicle unit throughput capacity is determined based on the number of vehicle lanes at the target intersection and the design saturation flow rate of the vehicle lanes.
[0148] The non-motorized vehicle density index consists of non-motorized vehicle arrival rate, non-motorized vehicle waiting time, and non-motorized vehicle unit capacity. The non-motorized vehicle unit capacity is determined based on the width of the non-motorized vehicle lane and the designed capacity of the non-motorized vehicle lane at the target intersection.
[0149] The pedestrian density index consists of pedestrian arrival rate, pedestrian waiting time, and pedestrian unit capacity. The pedestrian unit capacity is determined based on the width of the pedestrian crossing at the target intersection and the designed capacity of the pedestrian crossing.
[0150] The preset basic priority coefficients for motor vehicles, non-motor vehicles, and pedestrians are determined based on the traffic management strategy of the target intersection. The traffic management strategy includes a motor vehicle priority strategy, a non-motor vehicle priority strategy, a pedestrian priority strategy, and a balanced traffic strategy.
[0151] The weight of motor vehicles is obtained by multiplying the motor vehicle density index by the preset basic priority coefficient of motor vehicles; the weight of non-motor vehicles is obtained by multiplying the non-motor vehicle density index by the preset basic priority coefficient of non-motor vehicles; the weight of pedestrians is obtained by multiplying the pedestrian density index by the preset basic priority coefficient of pedestrians; the weights of motor vehicles, non-motor vehicles, and pedestrians are added together to obtain the total weight.
[0152] The normalized weight of motor vehicles is calculated by dividing the total weight of motor vehicles by the total weight; the normalized weight of non-motor vehicles is calculated by dividing the total weight of non-motor vehicles by the total weight; the normalized weight of pedestrians is calculated by dividing the total weight of pedestrians by the total weight; the motor vehicle efficiency term is obtained by multiplying the motor vehicle traffic volume by the normalized motor vehicle weight; the non-motor vehicle efficiency term is obtained by multiplying the non-motor vehicle traffic volume by the normalized non-motor vehicle weight; the pedestrian efficiency term is obtained by multiplying the pedestrian traffic volume by the normalized pedestrian weight.
[0153] Add the efficiency terms for motor vehicles, non-motor vehicles, and pedestrians to generate a comprehensive traffic efficiency function.
[0154] In this embodiment, S6 specifically refers to:
[0155] The system reads the number and phase sequence of signal phases at the target intersection; arranges the green light durations of the signal phases according to the phase sequence, and uses the green light duration of each signal phase as an element in a real number vector to generate initial timing individuals; the green light duration constraints include the minimum and maximum green light durations for each signal phase, with the minimum green light duration determined based on pedestrian crossing distance and vehicle start-up loss time, and the maximum green light duration determined based on the upper limit of signal cycle duration and the upper limit of waiting time for traffic participants; the cycle duration constraints include the lower and upper limits of signal cycle duration, which are obtained by adding the green and yellow light durations of the signal phases; the phase sequence constraint ensures that the existing phase sequence of the target intersection remains unchanged.
[0156] Initial timing individuals that do not meet the green light duration constraint, cycle duration constraint, and phase sequence constraint are removed, while initial timing individuals that meet the green light duration constraint, cycle duration constraint, and phase sequence constraint are retained to generate a timing population. The timing individuals in the timing population are then substituted into traffic flow parameters, normalized motor vehicle weights, normalized non-motor vehicle weights, and normalized pedestrian weights.
[0157] Motor vehicle traffic efficiency is obtained from motor vehicle traffic volume and normalized motor vehicle weight; non-motor vehicle traffic efficiency is obtained from non-motor vehicle traffic volume and normalized non-motor vehicle weight; pedestrian traffic efficiency is obtained from pedestrian traffic volume and normalized pedestrian weight; motor vehicle traffic efficiency, non-motor vehicle traffic efficiency, and pedestrian traffic efficiency are used as the multi-objective fitness values corresponding to the timing individual;
[0158] Compare the multi-objective fitness values of any two timing individuals in the timing population; if a timing individual has a higher fitness value than another timing individual in at least one of the objectives of motor vehicle traffic efficiency, non-motor vehicle traffic efficiency, and pedestrian traffic efficiency, and no lower fitness value than another timing individual in the remaining objectives, then the former timing individual dominates the latter timing individual.
[0159] The timing population is stratified according to the dominance relationship between timing individuals to generate non-dominance levels; within the same non-dominance level, timing individuals are arranged in ascending order according to motor vehicle traffic efficiency, non-motor vehicle traffic efficiency, and pedestrian traffic efficiency; the fitness difference between adjacent timing individuals under the same objective is accumulated to generate congestion distance.
[0160] Parent timing individuals are selected from the timing population in order of non-dominance level from low to high and crowding distance from large to small. Two timing individuals are selected from the parent timing individuals, and the green light durations corresponding to the same signal phase in the two timing individuals are mixed in proportion to generate cross-timing individuals.
[0161] For a randomly selected signal phase in the cross-timing individual, the green light duration is increased or decreased by the green light adjustment amount to generate a variant timing individual; the green light adjustment amount is determined according to the green light duration constraint range of the corresponding signal phase, and the green light duration of the adjusted signal phase is kept between the minimum green light duration and the maximum green light duration;
[0162] Cross-timed individuals and variant timed individuals are used as offspring timed individuals; parent timed individuals and offspring timed individuals are merged to generate a merged timed population; target timed individuals are selected from the merged timed population in order of non-dominance level from low to high and crowding distance from large to small; the signal phase green light duration, signal period duration and phase sequence of the target timed individuals are combined to generate the target signal timed scheme.
[0163] In this embodiment, S7 specifically refers to:
[0164] The system reads the green light duration, signal cycle duration, and phase sequence of the target signal timing scheme; converts the green light duration, signal cycle duration, and phase sequence into timing instructions that can be received by the roadside signal controller; the timing instructions include the signal phase number, the green light duration, the signal cycle duration, the phase sequence, and the effective timestamp; and sends the timing instructions to the roadside signal controller via the 5G / 4G communication module carried by the drone.
[0165] The roadside signal controller reads the signal phase number, green light duration, signal cycle duration, phase sequence, and effective timestamp from the timing instruction; activates the timing instruction according to the effective timestamp; sequentially switches the signal phases of the target intersection according to the phase sequence; controls the green light duration of the corresponding signal phase according to the green light duration; and controls the signal cycle period of the target intersection according to the signal cycle duration.
[0166] In the next control cycle, real-time video streams of the target intersection area are collected via UAVs. The collected video streams are then time-stamped according to their timestamps and calibrated using coordinates based on the stop line position, approach lane boundary, non-motorized vehicle lane boundary, and pedestrian crossing boundary to generate air-ground traffic data for the next control cycle. This data is then input into an improved YOLOv8 target detection network to generate heterogeneous traffic targets for the next control cycle. An improved Bytetrack multi-target tracking algorithm is used to correlate the heterogeneous traffic targets for the next control cycle, generating vehicular, non-motorized vehicle, and pedestrian trajectories. If non-motorized vehicle targets are occluded in the next control cycle, Kalman filtering is used to predict their trajectories, generating compensated trajectories. A perspective transformation algorithm is then used to... The vehicle trajectories, compensated trajectories, and pedestrian trajectories of each control cycle are mapped to the world coordinate system to generate a heterogeneous traffic trajectory set for the next control cycle. Traffic flow feature parameters are extracted from this set to generate a traffic flow parameter set for the next control cycle. The weights of vehicles, non-motorized vehicles, and pedestrians for the next control cycle are calculated based on this parameter set, and a comprehensive traffic efficiency function for the next control cycle is constructed using this parameter set. The NSGA-II multi-objective genetic algorithm is used to optimize the signal timing of the comprehensive traffic efficiency function for the next control cycle, generating an updated target signal timing scheme. The updated target signal timing scheme is sent to the roadside signal controller. The target signal timing scheme, the updated target signal timing scheme, the traffic flow parameter set for the next control cycle, and the execution records from the roadside signal controller constitute the traffic signal timing optimization result.
[0167] Example 1: To verify the feasibility of this invention in practice, it was applied to a traffic signal control scenario at a main road intersection in a city. The target intersection includes four vehicular lanes, four non-motorized vehicle lanes, and four sets of pedestrian crossings. The area surrounding the intersection includes commercial complexes, bus stops, and schools. During morning and evening rush hours, there are problems such as high vehicular traffic density, severe non-motorized vehicle mixing, and frequent pedestrian crossings. Traditional fixed-cycle signal control methods cannot dynamically adjust the green light duration according to real-time traffic flow changes, easily leading to continuously increasing vehicular queue lengths, excessively long waiting times for non-motorized vehicles, and frequent pedestrian crossing conflicts. Especially under the overhead view of an unmanned aerial vehicle (UAV), small targets in the edge areas exhibit significant scale changes and frequent local occlusion. Traditional target detection and trajectory association methods are prone to missed detections, trajectory drift, and frequent target ID switching, thus affecting traffic flow statistics and subsequent signal timing optimization.
[0168] In actual deployment, a quadcopter drone was deployed approximately 55m above the center of the target intersection. The drone was equipped with a 4K high-definition gimbal camera, with a video capture frame rate of 30fps and a view angle of 65°. The drone continuously captured real-time video streams of the target intersection area, performed time-stamp alignment based on the capture timestamps, and then performed coordinate calibration based on the stop line position, approach lane boundary, non-motorized vehicle lane boundary, and pedestrian crossing boundary to generate air-ground traffic data. This air-ground traffic data was then input into an improved YOLOv8 target detection network. A visual drift diffusion mechanism was used to correlate motorized, non-motorized, and pedestrian targets in feature maps of different scales, reducing target detection offset issues caused by drone viewpoint changes and perspective compression in edge areas. An improved Bytetrack multi-target tracking algorithm was then used to complete trajectory correlation, combined with a phase trajectory constraint mechanism to limit abnormal trajectory jumps. Kalman filtering was used for trajectory compensation for non-motorized targets in occluded areas. Finally, a perspective transformation algorithm was used to map the image coordinates to the world coordinate system, generating a heterogeneous traffic trajectory set.
[0169] The system ran continuously for 14 days, and the recognition results of motor vehicle targets, non-motor vehicle targets, pedestrian targets, occluded non-motor vehicle targets, and small targets in the edge area were statistically analyzed and compared with traditional detection and tracking methods. The data is shown in Table 1.
[0170] Table 1 Comparison of Target Recognition and Trajectory Generation
[0171]
[0172] As shown in Table 1, the target recognition capability of this invention in complex traffic scenarios is significantly superior to traditional methods. Specifically, the accuracy rate for non-motorized vehicle targets increased from 82.4% to 94.3%, and the accuracy rate for occluded non-motorized vehicle targets increased from 76.5% to 91.2%, indicating that the visual drift diffusion mechanism can effectively reduce the scale drift problem under UAV top-down conditions. The accuracy rate for small targets in edge regions increased to 92.5%, indicating that feature compensation between feature maps of different scales can enhance the ability to preserve edge details of small targets. The number of non-motorized vehicle target ID switching times decreased from 46 times / hour to 13 times / hour, and the number of occluded non-motorized vehicle target ID switching times decreased from 58 times / hour to 16 times / hour, indicating that the phase trajectory constraint mechanism can effectively improve the trajectory association stability. At the same time, the overall missed detection rate decreased significantly, verifying that Kalman filter trajectory compensation can maintain trajectory continuity under occlusion conditions, providing a stable data foundation for traffic flow statistics and signal timing optimization.
[0173] After generating the heterogeneous traffic trajectory set, the system further statistically analyzes traffic flow parameters such as the number of motor vehicles, non-motor vehicles, pedestrians, motor vehicle queue length, non-motor vehicle queue length, and waiting time. It then constructs a comprehensive traffic efficiency function and dynamically optimizes the green light duration of the signal phase using the NSGA-II multi-objective genetic algorithm. Data from the weekday morning rush hour from 07:00 to 09:00 is continuously selected and compared with traditional fixed timing methods and traditional induction control methods, yielding the results shown in Table 2.
[0174] Table 2 Comparison of Signal Timing Optimization Performance Data
[0175]
[0176] As shown in Table 2, this invention has significant advantages in dynamic traffic signal optimization. The average waiting time for motor vehicles is reduced to 42.3 seconds, for non-motorized vehicles to 48.9 seconds, and for pedestrians to 55.6 seconds, indicating that the comprehensive traffic efficiency function can balance the traffic needs of motor vehicles, non-motorized vehicles, and pedestrians. The maximum queue length for motor vehicles decreases from 138.5 meters to 91.4 meters, demonstrating that dynamic signal timing effectively alleviates congestion at approach lanes. The single-cycle comprehensive traffic volume increases to 613 vehicles and passengers, indicating that the NSGA-II multi-objective genetic algorithm can achieve a better balance among multiple traffic objectives. The signal cycle update interval is shortened to 3 minutes, demonstrating that this invention can quickly respond to real-time traffic flow changes and improve the flexibility of traffic signal control. Overall operational results show that this invention can improve the accuracy and trajectory stability of heterogeneous traffic target recognition in complex urban traffic scenarios, and effectively reduce intersection waiting time and queue length, possessing high engineering application value.
[0177] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A traffic signal timing optimization method based on UAV edge computing, characterized in that, Includes the following steps: S1. Collect real-time video streams of the target intersection area using drones, perform time-stamp alignment and coordinate calibration processing, and generate air-ground traffic data; S2. Input the air and ground traffic data into the improved YOLOv8 target detection network, introduce a visual drift diffusion mechanism in the scale correlation module, identify motor vehicle targets, non-motor vehicle targets and pedestrian targets, and generate heterogeneous traffic targets; S3. An improved Bytetrack multi-target tracking algorithm is used to perform trajectory association on the heterogeneous traffic targets. A phase trajectory constraint mechanism is introduced to generate motor vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories. If a non-motor vehicle target is occluded, Kalman filtering is used to perform trajectory prediction on the non-motor vehicle trajectory to generate a compensation trajectory. A perspective transformation algorithm is used to map the motor vehicle trajectory, compensation trajectory, and pedestrian trajectory to the world coordinate system to generate a heterogeneous traffic trajectory set. S4. Extract traffic flow feature parameters based on the heterogeneous traffic trajectory set to generate a traffic flow parameter set; S5. Calculate the weights of motor vehicles, non-motor vehicles, and pedestrians based on the traffic flow parameter set, and construct a comprehensive traffic efficiency function in combination with the traffic flow parameter set; S6. Use the NSGA-II multi-objective genetic algorithm to perform signal timing optimization on the comprehensive traffic efficiency function to generate a target signal timing scheme; S7. Send the target signal timing scheme to the roadside signal controller to execute signal control, and perform closed-loop feedback update based on the air and ground traffic data collected in the next control cycle, and output the traffic signal timing optimization result.
2. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, Specifically, S1 is: Control the drone to hover at a preset height above the target intersection, and control the high-definition gimbal camera on the drone to collect real-time video streams of the target intersection area at a preset pitch angle; The continuous image frames in the real-time video stream are time-stamp aligned according to the acquisition timestamp to generate an aligned video frame sequence. Based on the location of the stop line, the boundary of the approach lane, the boundary of the non-motorized vehicle lane, and the boundary of the pedestrian crossing in the target intersection area, coordinate calibration processing is performed on the aligned video frame sequence to generate air and ground traffic data.
3. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, The improved YOLOv8 target detection network specifically includes a backbone network, a scale correlation module, and a detection head; The backbone network divides continuous image frames in air-ground traffic data into image blocks, and extracts texture features, edge features and semantic features from the image blocks. These features are then arranged according to the image resolution to form a feature map with a first spatial size and a feature map with a second spatial size, where the first size is larger than the second size. The scale association module introduces a visual drift diffusion mechanism to perform scale association on the same target in feature maps of different scales. This mechanism includes: performing scale association on the same target in feature maps with a first spatial size and a second spatial size, where the targets include motor vehicle targets, non-motor vehicle targets, and pedestrian targets; extracting the bounding box width and height changes corresponding to the same target in adjacent image frames, and generating a target scale change gradient based on these changes; and extracting the distance from the target bounding box center point to the image center point and the distance from the target bounding box center point to the image edge. The distance to the edge and the change in the aspect ratio of the target bounding box are calculated. Based on the distance from the center point of the target bounding box to the center point of the image, the distance from the center point of the target bounding box to the edge of the image, and the change in the aspect ratio of the target bounding box, the perspective compression of the edge region is generated. The area change of the bounding box corresponding to the same target in adjacent image frames is extracted, and the cross-frame size offset is generated based on the area change of the bounding box. Based on the target scale change gradient, the perspective compression of the edge region, and the cross-frame size offset, the scale drift correspondence between the same target position in the feature map with a spatial size of the first size and the feature map with a spatial size of the second size is established, and the scale drift field is generated. Based on the scale drift field, the feature compensation positions and feature compensation weights between the feature map with a first spatial size and the feature map with a second spatial size are determined; the target edge features in the feature map with a first spatial size are compensated to the corresponding target positions in the feature map with a second spatial size, and the target semantic features in the feature map with a second spatial size are compensated to the corresponding target positions in the feature map with a first spatial size, thereby generating stable target response features; The detection head determines the target center point position, target bounding box width, target bounding box height, and target category confidence based on the stable target response characteristics, and outputs the bounding box coordinates and category confidence of motor vehicle targets, non-motor vehicle targets, and pedestrian targets, generating heterogeneous traffic targets.
4. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, The improved Bytetrack multi-target tracking algorithm is used to perform trajectory association on the heterogeneous traffic targets, and a phase trajectory constraint mechanism is introduced to generate vehicle trajectories, non-motor vehicle trajectories, and pedestrian trajectories. Specifically: Based on the relationship between the category confidence scores of motor vehicle targets, non-motor vehicle targets, and pedestrian targets and the preset confidence thresholds, heterogeneous traffic targets with category confidence scores greater than or equal to the preset confidence thresholds are classified into a high-confidence target set, and heterogeneous traffic targets with category confidence scores less than the preset confidence thresholds are classified into a low-confidence target set. The motor vehicle passage area, non-motor vehicle passage area, and pedestrian passage area are determined based on the location of the parking line, the boundary of the entrance lane, the boundary of the non-motor vehicle lane, and the boundary of the pedestrian crossing. Based on the change in the position of the center point of the bounding box of the heterogeneous traffic target in consecutive image frames, the motion direction of the heterogeneous traffic target is determined. Establish phase trajectory constraint relationships based on the traffic area, direction of movement, stop line location, and phase release status of heterogeneous traffic targets; The high-confidence target set is associated with the existing trajectory according to the intersection-union ratio of the bounding box, the consistency of the motion direction, and the phase trajectory constraint relationship to generate the first trajectory association result; The low-confidence target set is associated with the unassociated trajectories in the first trajectory association result according to the bounding box intersection-union ratio, motion direction consistency and phase trajectory constraint relationship to generate the second trajectory association result; Based on the first trajectory association result and the second trajectory association result, update the trajectory identifier, trajectory category and trajectory location to generate motor vehicle trajectory, non-motor vehicle trajectory and pedestrian trajectory.
5. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, If the non-motorized vehicle target is obstructed, Kalman filtering is used to predict the trajectory of the non-motorized vehicle and generate a compensated trajectory. Specifically: Calculate the speed and direction of movement of the non-motorized vehicle based on the world coordinates and timestamps of adjacent trajectory points in the non-motorized vehicle trajectory. If the detection box for the non-motorized vehicle trajectory is missing in consecutive image frames, and the predicted trajectory position corresponding to the non-motorized vehicle trajectory is located within the non-motorized vehicle passage area, then it is determined that the non-motorized vehicle target is occluded. The position of the non-motorized vehicle trajectory point, the speed of the non-motorized vehicle, and the direction of the non-motorized vehicle are used as the state variables of the Kalman filter. Predict the location of non-motorized vehicle trajectory points in missing image frames based on Kalman filter state variables; generate a compensation trajectory based on the associated trajectory points in the non-motorized vehicle trajectory and the location of non-motorized vehicle trajectory points in the missing image frames.
6. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, The perspective transformation algorithm is used to map vehicle trajectories, compensated trajectories, and pedestrian trajectories to the world coordinate system, generating a heterogeneous traffic trajectory set. Specifically: The calibration points in the image coordinate system are selected based on the location of the parking line, the boundary of the approach lane, the boundary of the non-motorized vehicle lane, and the boundary of the pedestrian crossing; the world coordinate points corresponding to the calibration points are determined based on the actual road plane coordinates corresponding to the target intersection area. Calculate the perspective transformation matrix based on the calibration point and world coordinate point; Based on the perspective transformation matrix, the trajectory points in the vehicle trajectory, the compensated trajectory, and the pedestrian trajectory are mapped from the image coordinate system to the world coordinate system; The mapped trajectory points are merged according to trajectory identifier, trajectory category, world coordinate position, and timestamp to generate a heterogeneous traffic trajectory set.
7. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, Specifically, S4 is: Based on the trajectory categories, world coordinates, and timestamps of the heterogeneous traffic trajectory set, the vehicle trajectory, compensation trajectory, and pedestrian trajectory are matched to the corresponding approach lane, turning direction, and signal phase, respectively. Based on the vehicle trajectories, compensation trajectories, and pedestrian trajectories matched to the same entrance lane, the same turning direction, and the same signal phase, the number of motor vehicles, non-motor vehicles, and pedestrians passing through the stop line per unit time is counted to generate traffic flow parameters. Based on the world coordinates of the vehicle trajectories and compensation trajectories in the direction from the stop line upstream, the queue lengths of motor vehicles and non-motor vehicles are calculated, and traffic flow queuing parameters are generated. Based on the timestamps of vehicle trajectories, compensation trajectories, and pedestrian trajectories when entering the detection area and when passing the parking line, calculate the waiting time for motor vehicles, non-motor vehicles, and pedestrians, and generate traffic flow waiting parameters. A traffic flow parameter set is generated based on traffic flow traffic parameters, traffic flow queuing parameters, and traffic flow waiting parameters.
8. The traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, Specifically, S5 is: Based on the traffic flow parameters, traffic flow queuing parameters, and traffic flow waiting parameters in the traffic flow parameter set, the vehicle density index, non-motor vehicle density index, and pedestrian density index are determined respectively. The weights of motor vehicles are calculated based on the motor vehicle density index and the preset basic priority coefficients for motor vehicles; the weights of non-motor vehicles are calculated based on the non-motor vehicle density index and the preset basic priority coefficients for non-motor vehicles. Pedestrian weights are calculated based on pedestrian density indicators and preset pedestrian basic priority coefficients. The weights of motor vehicles, non-motor vehicles, and pedestrians are normalized to generate normalized motor vehicle weights, normalized non-motor vehicle weights, and normalized pedestrian weights. A comprehensive traffic efficiency function is constructed based on traffic flow parameters, normalized motor vehicle weights, normalized non-motor vehicle weights, and normalized pedestrian weights.
9. A traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, Specifically, S6 is: Encode the green light duration of the signal phase at the target intersection into a real number vector to generate the initial timing individual; The initial timing individuals are constrained and screened based on green light duration constraints, cycle duration constraints, and phase sequence constraints to generate a timing population. The combined traffic efficiency function is used to calculate the traffic efficiency of motor vehicles, non-motor vehicles, and pedestrians for the timed individuals in the timed population, and a multi-objective fitness value is generated. Based on the multi-objective fitness values, the timing individuals in the timing population are non-dominated and sorted to generate non-dominated levels. Crowding distance was calculated using fitness distances between individuals within the same non-dominant rank. Based on the non-dominance level and the crowding distance, parent timing individuals are selected from the timing population; the green light duration of the signal phase in the parent timing individuals is cross-combined and mutated to generate child timing individuals; The parent and child mating individuals are merged, and target mating individuals are selected according to the non-dominance level and the crowding distance. A target signal timing scheme is generated based on the green light duration of the signal phase in the target timing individual.
10. A traffic signal timing optimization method based on UAV edge computing according to claim 1, characterized in that, Specifically, S7 is: The green light duration, signal cycle duration, and phase sequence of the target signal timing scheme are sent to the roadside signal controller; the roadside signal controller controls the signal light status of the target intersection according to the green light duration, signal cycle duration, and phase sequence. Real-time video streams of the target intersection area are collected during the next control cycle to generate air and ground traffic data for the next control cycle. Based on the air-ground traffic data of the next control cycle, update the heterogeneous traffic targets, heterogeneous traffic trajectory sets, traffic flow parameter sets, comprehensive traffic efficiency functions, and target signal timing schemes. The updated target signal timing scheme is sent to the roadside signal controller, which outputs the traffic signal timing optimization results.