Intersection traffic state completion method and system of multi-view unmanned aerial vehicle image
By constructing a multi-view UAV imagery method based on graph neural networks, the problem of spatiotemporal blind spots in UAV traffic condition monitoring is solved, achieving full-domain, continuous, and high-precision traffic condition reconstruction, improving data integrity and accuracy, and making it suitable for intelligent traffic management.
Patent Information
- Application Number
- CN202511671303.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-10
AI Technical Summary
Existing UAV traffic condition monitoring technologies suffer from problems such as time blind spots caused by perspective switching, spatial blind spots caused by limited field of view, and low accuracy and weak semantic expression of traditional data completion methods. These issues result in low accuracy and poor continuity of traffic condition monitoring, making them unsuitable for intelligent traffic management.
A multi-view UAV imagery method based on graph neural networks is adopted to construct a structured graph model. Road semantic information is fused through a message passing mechanism to identify and complete traffic state parameters in temporal and spatial blind spots. Graph attention networks and temporal modeling techniques are used for data completion.
It enables full-area, continuous, and high-precision reconstruction of traffic conditions at complex intersections, improving data integrity and accuracy, supporting dynamic multi-view data fusion, and is suitable for emergency monitoring and temporary control. It also has good scalability and practicality.
Smart Images

Figure CN121505886A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traffic condition perception and data completion technology, specifically to a method and system for completing intersection traffic conditions based on graph neural networks and multi-view UAV imagery for intelligent transportation systems. Background Technology
[0002] With the acceleration of urbanization and the continuous growth of motor vehicle ownership, problems such as traffic congestion, unreasonable signal timing, and delayed accident response are becoming increasingly prominent, placing higher demands on the real-time performance and refined management of Intelligent Transportation Systems (ITS). Accurate and comprehensive acquisition of traffic status information at road intersections is fundamental to realizing core functions such as traffic flow prediction, optimized signal control, congestion warning, and emergency management.
[0003] Currently, traffic condition perception mainly relies on fixed monitoring equipment, such as geomagnetic sensors, induction coils, and video surveillance cameras. However, these facilities are costly to deploy, difficult to maintain, and have limited coverage, making it difficult to flexibly respond to sudden traffic events or temporary traffic organization adjustments. In addition, their spatial perspective is limited, especially in complex intersection areas, making it difficult to achieve simultaneous, large-scale dynamic monitoring of traffic flow in multiple directions.
[0004] In recent years, unmanned aerial vehicles (UAVs) have gradually become an important supplementary means of traffic monitoring due to their advantages such as high mobility, wide field of view, and flexible deployment. Equipped with high-definition cameras, UAVs can acquire large-scale, high-resolution road traffic images from an aerial perspective, supporting real-time extraction of key traffic parameters such as vehicle distribution, driving trajectories, and queue lengths. Especially in complex areas with converging traffic flows, such as traffic light intersections, UAVs can provide a global observation capability that traditional ground equipment cannot match.
[0005] However, existing UAV-based traffic monitoring technologies still face significant technical bottlenecks, particularly the data loss caused by spatiotemporal blind spots during dynamic data acquisition. On one hand, limited by the camera's field of view, a single shot cannot simultaneously cover multiple road segments extending in various directions at an intersection. Periodic adjustments to the gimbal angle or flight attitude are necessary to switch observation perspectives, resulting in certain directions being unable to be continuously observed within specific time periods, creating temporal blind spots. On the other hand, during peak traffic hours, vehicle queues may exceed the effective coverage of the UAV's tilted field of view. If the UAV's position is fixed, it cannot fully capture information at the end of long queues. Even with small-scale movement to expand the field of view, reduced resolution or occlusion at the edge of the viewpoint can lead to the loss of some vehicle information, creating spatial blind spots. These blind spots result in discontinuous and partially missing traffic state data, severely impacting the completeness and accuracy of traffic state modeling.
[0006] To address these issues, existing research has attempted to use interpolation, Kalman filtering, or traditional machine learning models to complete missing data. However, these methods typically assume that the data exhibits linear variation or stationary statistical properties, making it difficult to effectively model the nonlinear and strongly spatiotemporally dependent relationships of traffic flow under complex road network structures. Furthermore, existing methods often treat lanes or road segments as independent units, lacking effective integration of semantic information such as road topology, vehicle interaction behavior, and traffic light control logic, resulting in a lack of physical plausibility at the microscopic level in the completed data.
[0007] Graph Neural Networks (GNNs), as deep learning models capable of effectively processing non-Euclidean structured data, have demonstrated excellent performance in traffic prediction and flow completion in recent years. By constructing a graph structure composed of nodes and edges, they can naturally represent the topological connections of road networks and utilize message passing mechanisms to capture the spatial propagation characteristics and temporal evolution patterns of traffic states. However, how to effectively map heterogeneous traffic observation data collected dynamically from multiple perspectives by UAVs onto graph structures, and how to design semantically aware completion mechanisms to address their unique spatiotemporal blind spots, remains a gap in current research.
[0008] The following are the key problems existing in the current UAV traffic condition monitoring technology: 1) Time blind zone problem caused by perspective switching: A single UAV is limited by the camera's field of view. When taking tilted pictures of road intersections extending in multiple directions, it needs to periodically switch the observation direction, resulting in no effective data collection for road sections in directions other than the current perspective within a certain time period, forming a time observation gap; 2) Spatial blind zone problem caused by limited field of view: During peak traffic hours, the length of vehicle queues may exceed the observable range of the UAV at a fixed position. Even if the field of view is expanded by moving in a small range, there is still a loss of local information due to insufficient resolution or occlusion in the edge areas; 3) Problems of low accuracy and weak semantic expression in missing data completion: Traditional data completion methods are difficult to model the complex spatiotemporal dependencies of traffic flow and do not fully integrate prior semantic information such as road topology, lane connection relationships and traffic light control, resulting in a lack of physical rationality and spatial consistency in the completion results.
[0009] In summary, existing UAV traffic condition monitoring technologies cannot effectively overcome the blind spots in UAV observation, resulting in low accuracy and poor continuity in traffic monitoring, making them unsuitable for intelligent traffic monitoring and management scenarios. Therefore, there is an urgent need for a traffic condition modeling method that can integrate multi-view UAV image data with road semantic information. This method would overcome the problem of missing spatiotemporal data caused by perspective switching and field of view limitations, enabling comprehensive, continuous, and high-precision reconstruction and visualization of traffic conditions at road intersections, providing reliable technical support for urban traffic operation monitoring and decision support. Summary of the Invention
[0010] The purpose of this invention is to address the shortcomings of existing technologies by providing a traffic state completion system based on a multi-view UAV imagery method for intersection traffic state completion. This system employs a road intersection state modeling method based on graph neural networks and multi-view UAV imagery to effectively complete and globally reconstruct traffic state parameters in spatiotemporal blind spots generated during dynamic UAV acquisition, thereby improving the completeness, continuity, and accuracy of traffic perception. This invention utilizes a structured graph model and the message passing mechanism of graph neural networks to fully leverage the spatial continuity and temporal correlation of traffic flow, achieving high-precision reconstruction of missing traffic states in temporal and spatial blind spots, thus improving data integrity. By deeply integrating road semantic information, lane connections, upstream and downstream structures, and traffic light control logic are explicitly modeled as edges and special nodes in the graph, enhancing the model's understanding of traffic operation mechanisms and making the completion process more physically plausible and interpretable. Furthermore, it can integrate heterogeneous imagery data collected by a single UAV at different times and in different postures, breaking through the limitations of a single perspective and achieving continuous modeling of the entire traffic state of complex intersections, realizing dynamic multi-view data fusion. This method does not rely on densely deployed ground sensors and is applicable to various scenarios such as emergency monitoring, temporary control, and assessment of newly built roads. It can also be adapted to different types of drones and urban road networks through software upgrades, exhibiting good scalability and practicality. It enables refined visualization and provides intuitive and accurate decision support for traffic command and dispatch.
[0011] The technical solution to achieve the objective of this invention is: a method for completing the traffic status of intersections using multi-view UAV imagery, characterized by the following specific steps: Step S1: Multi-view traffic condition data collection and extraction Step S1-1: Control the drone at different times above the target intersection, periodically switch shooting directions and slightly adjust the drone's flight position to collect video images from multiple tilted perspectives; Step S1-2: Perform vehicle target detection and trajectory tracking on the video images. Combine the real-time positioning information of the UAV, camera internal and external parameters and attitude data, convert the detected vehicle pixel coordinates into geographic coordinates, and match them with preset digital road vector data to extract the micro traffic state parameters of each lane at different times. The micro traffic state parameters include at least vehicle speed, vehicle spacing, traffic flow, lane change rate, queue length and vehicle type distribution.
[0012] Step S2: Traffic State Map Structure Construction Step S2-1: Divide the road within the coverage area of the target intersection into multiple discrete lane segment units; Step S2-2: Using each lane segment unit as a graph node and the connection relationship reflecting the traffic flow propagation relationship as an edge, construct a traffic state graph; wherein, the establishment of the edge is based on the upstream and downstream connection relationship of the lane, the vehicle driving direction and the traffic light control logic. Step S2-3: Map the traffic state parameters extracted in step S1 to the corresponding graph nodes to form an initial traffic state graph with spatiotemporal characteristics.
[0013] Step S3: Identification and Completion of Missing Data Step S3-1: Identify the time blind zone caused by the rotation of the drone gimbal to switch perspectives, and the spatial blind zone caused by the change of the drone's flight position, which causes some road areas to be outside the effective observation range; Step S3-2: Mark the graph nodes in the initial traffic state map that are in the time blind zone or spatial blind zone as missing nodes; Step S3-3: Input the traffic state map containing known node features and complete edge adjacency relationships into the pre-trained graph neural network model. Utilize the message passing mechanism of the graph neural network to aggregate the spatiotemporal feature information of known nodes, predict and complete the traffic state parameters of the missing nodes, and obtain a complete traffic state map.
[0014] Step S4: Result Visualization and Output Step S4-1: Remap the traffic state parameters in the complete traffic state map obtained in step S3 onto the digital map to generate and display a visualized traffic state map containing dynamic traffic flow and spatiotemporal evolution information.
[0015] The initial traffic state diagram in step S2 also includes traffic light control nodes; the traffic light control nodes are connected to the approach lane segment nodes controlled by them; the attributes of the connecting edges include the current traffic light phase and whether passage is permitted at the current time node; during the red light phase, the capacity weight of the lane segment node adjacent to the stop line is set to zero; during the green light phase, for lane segment nodes that simultaneously allow multiple driving directions, the historical turning flow ratio of each driving direction is calculated, and the capacity weight of the lane segment node is assigned to different downstream lane segment nodes.
[0016] In step S3, the weights of the edges are dynamically calculated and adjusted during the data prediction and completion process based on the correlation of the turning vehicle proportions detected in real time by the upstream and downstream lane segment units.
[0017] The graph neural network model described in step S3 adopts a graph attention network (GAT) architecture, which enables the model to assign different attention weights to different neighbor nodes during message passing, so as to reflect the differences in their impact on the target node.
[0018] After step S3, the method further includes: calculating the prediction variance of the graph neural network model as an uncertainty index of the completed traffic state parameters, and rendering the visualized traffic state map in step S4 according to the uncertainty index at a flashing frequency.
[0019] A traffic state completion system constructed using a multi-view UAV imagery method for intersection traffic state completion is characterized by comprising: a UAV acquisition module, a graph structure construction module, a data completion module, and a visualization output module. The UAV acquisition module performs the function of step S1 in claim 1, and is required to include at least a UAV flight platform, a high-definition camera, a gimbal controller, and a GNSS / IMU positioning unit to acquire multi-view images and extract initial traffic state data. The graph structure construction module performs the function of step S2 in claim 1, constructing a traffic state graph containing node, edge, and traffic light information. The data completion module performs the function of step S3 in claim 1, using a graph neural network model to complete missing data caused by temporal and spatial blind spots. The visualization output module is used to perform the function of step S4 in claim 1, generating and displaying a complete traffic status visualization map; the data completion module includes a pre-trained spatiotemporal graph neural network model, which is obtained by training on a dataset containing blind spots of simulated or real UAV observations.
[0020] Compared with the prior art, the present invention has the following beneficial pricing effects and significant technological advancements: 1) Effectively mitigate the impact of spatiotemporal blind spots: By constructing a structured graph model and introducing the message passing mechanism of graph neural networks, the spatial continuity and temporal correlation of traffic flow are fully utilized to achieve high-precision reconstruction of missing traffic conditions in time and space blind spots, thereby improving data integrity.
[0021] 2) Deep integration of road semantic information: Lane connection relationships, upstream and downstream structures and traffic light control logic are explicitly modeled as edges and special nodes in the graph, which enhances the model's ability to understand traffic operation mechanisms and makes the completion process more physically reasonable and interpretable.
[0022] 3) Supports dynamic multi-view data fusion: It can integrate heterogeneous image data collected by a single UAV at different times and in different attitudes, break through the limitations of a single viewpoint, and realize continuous modeling of the traffic status of the entire complex intersection.
[0023] 4) It has good scalability and practicality: The proposed method does not rely on densely deployed ground sensors, and is suitable for various scenarios such as emergency monitoring, temporary control, and new road construction assessment. It can also be adapted to different types of drones and urban road networks through software upgrades.
[0024] 5) Achieve refined and visualized presentation: The output results can be directly connected to the smart city traffic management platform, providing intuitive and accurate decision support for traffic command and dispatch. Attached Figure Description
[0025] Figure 1 This is a flowchart of the traffic status completion method of the present invention; Figure 2 This is a schematic diagram of the traffic status completion system of the present invention; Figure 3 This is a schematic diagram of the traffic state diagram in Example 1; Figure 4 This is a schematic diagram of the time blind zone in the dynamic data acquisition by the UAV in Example 1; Figure 5 This is a schematic diagram of the spatial blind zone in the dynamic data acquisition by the UAV in Example 1; Figure 6 is a schematic diagram of the graph structure for traffic state modeling in Example 1; Figure 7 This is a schematic diagram of the dynamic visualization interface for traffic status parameters in Example 1. Detailed Implementation
[0026] See Figure 1 The traffic status completion method of the present invention specifically includes: Step S1: Extraction of traffic state parameters based on UAV multi-view imagery Step S1-1: Obtain multiple video sequences of the target intersection and its extended road section taken by a single drone at different times and with different gimbal attitudes. These sequences include geographic coordinate information, camera pose, image images, and timestamp information at the time of shooting. Step S1-2: Perform the following processing sequentially on each group of video frames. 1) The deep learning model for vehicle detection and recognition in UAV images is pre-trained using the open-source YOLOv8 model and trained using the VisDrone public dataset. The model is then used to infer and recognize vehicle targets in the image, and the bounding boxes of each vehicle in the image coordinate system are obtained. 2) Apply well-known multi-target tracking algorithms such as DeepSORT to establish vehicle trajectories across frames and obtain the motion path and velocity information of each vehicle in the image coordinate system; 3) Using the UAV's own positioning data, camera intrinsic and extrinsic parameters, a projection transformation model from image pixel coordinates to ground geographic coordinates is applied to map the vehicle's position to the real-world coordinate system. The coordinate transformation is based on well-known imaging models in photogrammetry, including but not limited to pinhole camera models and direct linear transformation (DLT). 4) Match the mapped vehicle position with the lane vector data in the pre-stored high-precision map to determine its lane; 5) Based on the matching results, statistically analyze and calculate micro-traffic state parameters by lane unit, including: average vehicle speed, vehicle spacing, time occupancy, queue length, flow rate, etc.
[0027] Step S2: Constructing a graph structure and data mapping for traffic state modeling Step S2-1: Divide the road network within the coverage area of the UAV monitoring into basic units with semantic meaning, and construct a graph structure G=(V,E) to represent traffic status and their interrelationships.
[0028] Where V is a set of nodes, representing discretized lane segment nodes, each node corresponding to a lane unit of fixed length (e.g., 20 meters), used to carry the traffic state feature vector of the area; E is a set of edges including three types of connection relationships: front and rear straight edges, control-related edges, and spatially adjacent edges. The front and rear straight edges connect adjacent lane segment nodes within the same lane, as well as adjacent lane nodes with merging / diversion relationships, reflecting the vehicle's travel path. If it is the former, the passage weight of this edge is 1; if it is the latter, the weight is based on the historical statistical probability of merging and diversion. The control association edge abstracts the traffic light phase into a special type of traffic light control node and establishes a directed connection with all related lane segment nodes under its control, reflecting the impact of signal control on traffic flow. If it is a red light phase, the passage weight of this edge is 0; if it is a green light phase and there is only one direction of passage, the passage weight of this edge is 1; otherwise, the weight is based on the historical statistical probability of passage. The spatial adjacent edge constructs two directed connections between adjacent lanes, reflecting the impact of vehicle lane changes on traffic flow. If it is a lane-change dashed line, the passage weight of this edge is based on the lane-change direction and the historical statistical probability of passage; otherwise, the passage weight is 0.
[0029] Step S2-2: The traffic state feature vector of each node contains real-time traffic parameters extracted from step S1, and optional historical statistical data. If sufficient historical statistical passage probabilities are not collected, a weight preset of 0.5 is used as the default value. If the passage weight is 0, it is considered that the edge does not exist in the current state.
[0030] Steps S2-3: Map the traffic state parameters extracted from different perspectives and times in Step S1 to the corresponding lane segment nodes in the graph structure based on their corresponding geographical locations and timestamps. For ease of statistical analysis of traffic state parameters, it is recommended to set the time granularity to 10 seconds or more. Step S2-4: Create a mask corresponding to the graph structure. For lane segment nodes where data could not be collected due to time or space blind spots, mark the feature vectors corresponding to the geographical location and timestamp as "missing" in the mask graph to form a sparse graph input of partial observations.
[0031] Step S3: Traffic State Completion Based on Graph Neural Network Design and train a graph neural network model that integrates spatiotemporal modeling capabilities to predict traffic state features of missing nodes. The graph neural network model includes a graph convolutional encoder module, a temporal modeling module, and a message passing and feature reconstruction module. The output layer generates complete traffic state feature vectors for all nodes (including the original missing nodes) and outputs of uncertainty estimation.
[0032] The graph convolutional encoder utilizes a graph attention mechanism (GAT) or a graph convolutional layer (GCN) to aggregate known features of neighboring nodes and generate local context embeddings of nodes. The temporal modeling module employs a gated recurrent unit (GRU) or a Transformer structure to capture the trend of traffic state evolution of each node over time. The message passing and feature reconstruction module uses a multi-layer spatiotemporal graph neural network (such as ST-GNN or DCRNN architecture variants) to perform message propagation on the complete graph structure and infer the state of missing nodes by combining information from observed nodes.
[0033] Step S4: Integrated Traffic Status Analysis and Dynamic Visualization Step S4-1: Remap the complete traffic status map output in step S3 back to the Geographic Information System (GIS) platform or digital road network base map; Step S4-2: Dynamically display parameters such as speed, density, and queue length of each lane segment on a two-dimensional planar map using heat maps, flow arrows, color coding, etc. Step S4-3: Support playing the historical state change process by timeline, or generating a global traffic situation snapshot at a specified time. Step S4-4: Provide an API interface for traffic management systems to call for upper-level applications such as signal optimization and congestion warning.
[0034] The following is a detailed description of the invention, using a four-phase signal-controlled intersection where a main road and a secondary road intersect in a city as the monitoring object. The invention deploys a traffic status completion system based on graph neural networks and multi-view UAV imagery to achieve full-area, continuous, and high-precision traffic status perception and data completion for the intersection.
[0035] See Figure 2A multi-view UAV imagery-based intersection traffic status completion system includes: a UAV acquisition module, a graph structure construction module, a data completion module, and a visualization output module. The UAV acquisition module requires at least a UAV flight platform, a high-definition camera, a gimbal controller, and a GNSS / IMU positioning unit to acquire multi-view images and extract initial traffic status data. The graph structure construction module constructs a traffic status graph containing node, edge, and traffic light information. The data completion module uses a graph neural network model to complete missing data caused by temporal and spatial blind spots. The visualization output module generates and displays a complete traffic status visualization. The data completion module includes a pre-trained spatiotemporal graph neural network model, which is trained on a dataset containing simulated or real UAV observation blind spots.
[0036] The following example, using the intersection traffic state completion system constructed using this invention to achieve full-area, continuous, and high-precision traffic state perception and data completion for the intersection, will be used to further illustrate this invention in detail:
[0037] Example 1 The specific steps to achieve full-area, continuous, and high-precision traffic condition perception and data completion for this intersection are as follows: Step S1: Multi-view traffic condition data collection and extraction A DJI Matrice 3TD drone equipped with a high-definition camera, a three-axis gimbal, and a positioning module was used as the data acquisition platform. The drone hovered at an altitude of approximately 120 meters directly above the target intersection, maintaining a stable flight attitude. By periodically switching the gimbal's pitch and yaw angles, it sequentially captured video from the four entrance directions (east, south, west, and north) at 15-second intervals over a continuous 10-minute period, obtaining a total of 40 video sequences from different perspectives. Each frame of the image was accompanied by a precise timestamp, the drone's position coordinates (latitude, longitude, and altitude), camera intrinsic parameters (focal length, principal point coordinates), and extrinsic parameters (attitude angle).
[0038] Perform the following processing on each video segment: 1) Using a YOLOv8 object detection model pre-trained and fine-tuned on the VisDrone dataset, vehicle detection is performed on each frame of the image, and vehicle bounding boxes are output. 2) The DeepSORT multi-target tracking algorithm is used to establish cross-frame vehicle trajectories and obtain the motion trajectory and instantaneous image velocity of each vehicle; 3) Combining the UAV's pose (yaw, pitch, roll) and camera intrinsic parameters (focal length, principal point, pixel size), a formula is applied to transform pixel coordinates (u, v) to a coordinate system centered on the principal point. This is then divided by the focal length (in pixels) to obtain normalized planar coordinates. A rotation matrix is constructed based on the UAV's pose. The direction vector pointing from the UAV's position to the ground target point is then transformed from the camera coordinate system to the world coordinate system. Finally, the ray equation is constructed, and the intersection with the ground is solved to calculate the WGS-84 geographic coordinates corresponding to the vehicle's pixel coordinates. 4) Match the geographic coordinates with the pre-stored high-precision digital road vector map (including lane center lines, stop lines, turn arrows, etc.) to determine the lane segment to which each vehicle belongs; 5) Statistically analyze the micro traffic state parameters of each lane segment unit at a 10-second time granularity, including: average vehicle speed, average vehicle spacing, traffic flow per unit time, queue length (measured as the distance from the last stationary vehicle to the stop line), lane change frequency, and vehicle type distribution ratio.
[0039] Step S2: Traffic State Map Structure Construction See Figure 3 a) Divide the road within 200 meters upstream and downstream of the intersection into discrete lane segment units of 20 meters in length, generating a total of 178 lane segment nodes. Construct a traffic state diagram G=(V,E), where the node set V contains the aforementioned 178 lane segment nodes and 8 traffic light control nodes.
[0040] Referring to Figure b, the edge set E includes the following three types of edges: 1) Front and rear straight edges: connect adjacent lane segments in the same lane, with a traffic weight of 1; at merging and diverging points (such as the left-turn lane merging into the straight lane), based on historical turning traffic statistics, the merging weight is set to 0.3 and the diverging weight is set to 0.7. 2) Control Associated Edges: Each traffic light control node is connected to the lane segment node of the approach road it controls. When the current signal cycle is "East-West Straight Green Light", the control edge weight of the corresponding lane segment is set to 1, and the weight of the other directions is set to 0; 3) Spatial Adjacent Edges: Two-way edges are established between adjacent lanes. If the lane change area is separated by a dashed line, the weight of the lane change impact is set to 0.4 based on the historical lane change rate; the weight is 0 for solid line areas.
[0041] The traffic state parameters extracted in step S1 are mapped to corresponding nodes according to timestamps and geographical locations to form an initial traffic state map. Due to the switching of PTZ cameras, the far lanes of the south and north entrances have no monitoring information at multiple times. The system automatically identifies and marks these nodes as "missing" and generates corresponding mask maps.
[0042] See Figures 4-5 For lane segment nodes where data could not be collected due to time or space blind spots, the feature vectors corresponding to the geographical location and timestamp are marked as "missing" in the mask image, forming a sparse graph input of partial observations.
[0043] Step S3: Identification and Completion of Missing Data See Figure 6 An initial traffic state map, including a mask, is input into a pre-trained graph neural network model. This model employs a graph attention network (GAT) as the spatial encoder and a gated recurrent unit (GRU) as the temporal modeling module, forming a spatiotemporal graph neural network (ST-GNN) architecture. The model has been trained on a synthetic dataset containing simulated drone viewpoint switching and spatial occlusion. The training data comes from the publicly available pNEUMA dataset from Greece, and is superimposed with distortion and localization noise from real drone imagery.
[0044] During inference, the GAT layer dynamically calculates the influence weights of neighboring nodes on the target node through an attention mechanism. For example, when completing the speed of a missing lane segment at the south entrance, the model automatically assigns higher attention weights to downstream straight-ahead nodes and nodes in the same phase green light direction, while simultaneously suppressing the influence of nodes in the red light direction by combining traffic light control edge information. After three layers of message passing and feature aggregation, the model outputs a complete traffic state feature vector for all nodes. Furthermore, the model estimates the prediction variance using the Monte Carlo Dropout method, generating an uncertainty index for each completion parameter. For example, when the uncertainty of a queue length prediction is high, its variance value exceeds the threshold of 0.15.
[0045] Step S4: Result Visualization and Output See Figure 7 The completed traffic status map is mapped back to the GIS digital map platform to generate a dynamic and visualized traffic situation map. The system displays the vehicle speed of each lane segment in the form of a heat map (green indicates smooth traffic, red indicates congestion), and the length and color intensity of the arrows indicate the flow rate. The queue length is displayed as a red bar chart extending along the lanes.
[0046] For areas with high uncertainty (e.g., variance > 0.15), the system displays a flashing indicator (1Hz) on the visual interface to alert the user that the data in that area has low reliability and recommends cross-validation with other information sources. Simultaneously, the system supports replaying the traffic state evolution over the past 10 minutes along a timeline and provides a RESTful API interface to push the completed traffic flow data to the city's traffic signal optimization platform in real time for dynamic adjustment of signal timing schemes.
[0047] The above are merely specific implementations of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, are included within the scope of protection of the claims of the present invention.
Claims
1. A method for completing intersection traffic conditions using multi-view UAV imagery, characterized in that, This method employs a graph attention network architecture, dynamically assigns attention weights to neighboring nodes, and incorporates traffic light phase information to adjust capacity weights, thereby completing the traffic status update for intersections. Specifically, it includes the following steps: Step S1: Multi-view traffic condition data collection and extraction Step S1-1: Control the drone at different times above the target intersection, periodically switch shooting directions and slightly adjust the drone's flight position to collect video images from multiple tilted perspectives; Step S1-2: Perform vehicle target detection and trajectory tracking on the video images. Combine the real-time positioning information of the UAV, camera internal and external parameters and attitude data, convert the detected vehicle pixel coordinates into geographic coordinates, and match them with the preset digital road vector data to extract the micro traffic state parameters of each lane at different times. The micro traffic state parameters include at least vehicle speed, vehicle spacing, traffic flow, lane change rate, queue length and vehicle type distribution. Step S2: Construction of the traffic state map Step S2-1: Divide the road within the coverage area of the target intersection into multiple discrete lane segment units; Step S2-2: Take each lane segment unit as a graph node and use the connection relationship that reflects the traffic flow propagation relationship as an edge to construct a traffic state graph. The establishment of the edge is based on the upstream and downstream connection relationship of the lane, the vehicle driving direction and the traffic light control logic. Step S2-3: Map the traffic state parameters extracted in step S1 to the corresponding graph nodes to form an initial traffic state graph with spatiotemporal characteristics; Step S3: Identification and Completion of Missing Data Step S3-1: Identify the time blind zone caused by the rotation of the drone gimbal to switch perspectives, and the spatial blind zone caused by the change of the drone's flight position, which causes some road areas to be outside the effective observation range; Step S3-2: Mark the graph nodes in the initial traffic state map that are in the time blind zone or spatial blind zone as missing nodes; Step S3-3: Input the traffic state map containing known node features and complete edge adjacency relationships into the pre-trained graph neural network model. Utilize the message passing mechanism of the graph neural network to aggregate the spatiotemporal feature information of known nodes, predict and complete the traffic state parameters of the missing nodes, and obtain a complete traffic state map. The graph neural network model adopts a graph attention network architecture, which enables the model to assign different attention weights to different neighbor nodes during message passing to reflect the differences in their impact on the target node. S4: Results Visualization and Output The traffic state parameters in the complete traffic state map obtained in step S3 are remapped onto the digital map to generate and display a visualized traffic state map containing dynamic traffic flow and spatiotemporal evolution information.
2. The method for completing intersection traffic conditions from multi-view UAV imagery according to claim 1, characterized in that, The initial traffic state diagram in step S2 includes traffic light control nodes, which are connected to the lane segment nodes of the approach lanes controlled by them. The attributes of the connecting edges include the current traffic light phase and whether passage is permitted at the current time. During the red light phase, the capacity weight of the lane segment node adjacent to the stop line is set to zero. During the green light phase, for lane segment nodes that simultaneously allow multiple directions of travel, the historical turning flow ratio of each direction of travel is calculated, and the capacity weight of the lane segment node to different downstream lane segment nodes is assigned.
3. The method for completing intersection traffic conditions from multi-view UAV imagery according to claim 1, characterized in that, The missing data identification and completion in step S3 is dynamically calculated and adjusted based on the correlation of the turning vehicle ratio detected in real time by the upstream and downstream lane segment units, i.e., the weight of the edge.
4. The method for completing intersection traffic conditions from multi-view UAV imagery according to claim 1 or claim 3, characterized in that, In step S3, the graph neural network model calculates the prediction variance of the graph neural network model, which serves as an uncertainty index for the completed traffic state parameters. In step S4, the visualized traffic state map is rendered with a flashing frequency based on the uncertainty index.
5. A traffic state completion system constructed from a multi-view UAV imagery method for intersection traffic state completion, characterized in that, The traffic status completion system includes: a UAV acquisition module, a graph structure construction module, a data completion module, and a visualization output module. The UAV acquisition module includes: a UAV flight platform, a high-definition camera, a gimbal controller, and a GNSS / IMU positioning unit to acquire multi-view images and extract initial traffic status data. The graph structure construction module includes constructing a traffic status graph with node, edge, and traffic light information. The data completion module uses a graph neural network model to complete missing data caused by temporal and spatial blind spots. The visualization output module generates and displays a complete traffic status visualization graph.
6. The traffic state completion system constructed using the multi-view UAV imagery intersection traffic state completion method according to claim 5, characterized in that, The data completion module includes a pre-trained spatiotemporal graph neural network model, which is trained on a dataset containing blind spots from simulated or real UAV observations.