Road scene multi-target tracking system based on visual perception of unmanned aerial vehicle

Through the drone platform and ground control station combined with one-shot detector model and optical flow method, video data is collected and processed in real time and flight paths are dynamically adjusted, which solves the problem of slow video data processing speed and fuzzy abnormal behavior recognition in traffic management of drones, achieving efficient and intelligent traffic management.

CN120298460AActive Publication Date: 2025-07-11合肥众安睿博智能科技有限公司

Patent Information

Application Number
CN202510795576.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-11
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing drones are difficult to quickly process massive video frames in traffic management, and abnormal behavior recognition in complex highway scenarios is fuzzy, and video data lacks effective integration and analysis, resulting in waste of resources.

Method used

The drone platform, ground control station, object detection and tracking module and motion trajectory analysis and visualization module are adopted, combined with one-shot detector model, optical flow method and deep learning algorithm, video data is collected in real time, object detection and tracking, dynamically adjust the flight path, analyze the motion trajectory and visualize it.

Benefits of technology

It improves the timeliness and integrity of monitoring data, enhances the robustness and monitoring accuracy of target detection, can promptly detect potential risks and perform hierarchical processing, and improves the intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298460A_ABST
    Figure CN120298460A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traffic management, in particular to a road scene multi-target tracking system based on visual perception of an unmanned aerial vehicle, which realizes efficient acquisition and low-delay return of road scene videos through a high-definition camera and a 5G communication technology, and provides a reliable data basis for subsequent processing. Through data preprocessing and flight control, the video quality and the flight efficiency are improved, and the path of the unmanned aerial vehicle is dynamically adjusted according to the motion trail analysis result of the key dynamic target; a key dynamic target is obtained through a one-shot detector model and an optical flow method, and effective tracking is realized in combination with a deep learning algorithm and a multi-target tracking technology, so that the accuracy of target detection and the stability of tracking are improved; the motion trail of the key dynamic target is visually displayed by analyzing the motion trail of the key dynamic target and visualizing the analysis result of the motion trail of the key dynamic target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic management, and particularly to a multi-object tracking system for highway scenarios based on unmanned aerial vehicle (UAV) vision perception. Background Art

[0002] With the acceleration of urbanization and the increase in traffic flow, traditional traffic management methods are difficult to meet the needs of the new situation. UAV technology provides an efficient and intelligent solution for traffic management by collecting data in real time and processing and analyzing it. In the current application of UAV technology in the field of traffic management, there are indeed many challenges: since UAVs need to collect video data in real time, and the processing speed of conventional object detection algorithms is limited, it is difficult to efficiently process a large amount of video frames in a short time, which greatly limits its application efficiency; at the same time, in complex and changeable highway scenarios, the definition and recognition criteria of abnormal behaviors are still vague and cannot facilitate traffic management; in addition, there is a lack of effective integration and analysis means for the large amount of collected video data and motion trajectory information, making it difficult to achieve in-depth data mining and application, resulting in a waste of resources. Therefore, the present invention proposes a multi-object tracking system for highway scenarios based on UAV vision perception. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems in the background art, and to propose a multi-object tracking system for highway scenarios based on UAV vision perception.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A multi-object tracking system for highway scenarios based on UAV vision perception includes: a UAV platform, a ground control station, an object detection and tracking module, and a motion trajectory analysis and visualization module; UAV platform: Collects highway scenario videos in real time and transmits the collected video data back. The UAV platform includes an image acquisition unit and a data transmission unit; Ground control station: Processes the video data transmitted by the UAV and generates flight control instructions to dynamically adjust the path. The ground control station includes a data preprocessing unit and a flight control unit; Target Detection and Tracking Module: Extract the key dynamic target features of the highway scene through the one-shot detector model and perform cross-frame matching and statistics. At the same time, combine the optical flow method to analyze the motion vectors of static objects to offset the UAV interference and obtain the real motion information of the key dynamic targets; Use deep learning algorithms to detect targets in the video frames captured by the UAV; Adopt multi-object tracking algorithms to track the detected key dynamic targets, so as to obtain the motion trajectories of the key dynamic targets; Among them, the target detection and tracking module includes a key dynamic target acquisition unit, a key dynamic target association unit, a motion compensation unit, a key dynamic target detection unit, and a key dynamic target tracking unit; Motion Trajectory Analysis and Visualization Module: Analyze the motion trajectories of the key dynamic targets and visualize the analysis results of the motion trajectories of the key dynamic targets; Among them, the motion trajectory analysis and visualization module includes a motion trajectory analysis unit and a visualization unit.

[0005] Furthermore, the image acquisition unit is equipped with a 4K high-definition camera and a pan-tilt stabilization system for collecting real-time video data of the highway scene; The data transmission unit uses 5G communication technology for low-latency transmission of video data.

[0006] Furthermore, the data preprocessing unit is used to denoise, perform illumination equalization, and geometric correction on the received video data; The flight control unit is used to control the flight of the UAV and dynamically adjust the flight path according to the analysis results of the motion trajectories of the key dynamic targets.

[0007] Furthermore, the key dynamic target acquisition unit is used to obtain the preprocessed video data, extract key features through the one-shot detector model training, and identify and count the key dynamic targets in the highway scene.

[0008] Furthermore, the key dynamic target association unit is used to use the structural similarity index to judge whether the key dynamic targets in consecutive frames are the same target. For the key dynamic targets in the current frame and the next frame, calculate the similarity of the key dynamic targets; If the value of the structural similarity index of the key dynamic targets in consecutive frames exceeds the preset cross-frame structural similarity matching threshold, it is determined to be the same target.

[0009] Furthermore, the motion compensation unit is used to calculate the motion vectors of pixel points in the image between consecutive frames for static objects in the highway scene using the optical flow method, and fuse the weighted average of the average and mode of the moving direction and speed of the key pixel points to obtain the real pixel motion vector of the key dynamic target.

[0010] Further, the key dynamic target detection unit is used to extract features from the input video frames by using the optimized YOLO-X network structure, and combine the temporal feature fusion and dynamic threshold adjustment strategies to perform fine-grained detection on the determined key dynamic targets. The process of identifying the key dynamic target types includes: After the initial feature extraction in the YOLO backbone network, based on the analysis of target scale differences, a dynamic feature selection mechanism is constructed. This mechanism is used to adaptively allocate different levels of feature weights according to the key dynamic target scale to generate a preliminary target feature map. Based on the generated preliminary target feature map, it is combined with the temporal characteristics of the video sequence. Specifically: each frame's target feature map is regarded as a temporal node in the temporal feature map; at the same time, the target feature relationship between adjacent frames is defined as the temporal edge in the temporal feature map; the feature vector of the key dynamic target is used as the attribute of the temporal node in the temporal feature map, and the relative position relationship between the key dynamic targets is used as the weight of the temporal edge in the temporal feature map to capture the feature information and spatial relationship of the key feature targets in the temporal sequence. After the construction of the temporal feature map is completed, it is input into the graph neural network for processing, where the graph neural network operates on the temporal nodes and temporal edges in the temporal feature map; the graph neural network learns the temporal feature associations of the key dynamic targets between different frames and encodes the dynamic change information of the key dynamic targets in the temporal sequence into the feature representation; through the processing of multiple layers of graph neural networks, a feature representation integrating temporal information is obtained, where this feature representation contains the dynamic change information of the key dynamic targets in the entire temporal sequence; a splicing operation is used to fuse the feature representation integrating temporal information output by the graph neural network with the current frame's target feature map to generate a target feature map after integrating temporal information; based on the target feature map after integrating temporal information, the key dynamic target types are identified.

[0011] Further, the key dynamic target tracking unit is used to receive the output of the key dynamic target detection unit, and generate the unique ID of the key dynamic target and the motion trajectory by constructing an incremental dynamic feature memory, combining the dynamic adjustment threshold of the scene density, realizing trajectory association and initialization through the memory bank, and managing short-term and long-term motion trajectories. The process includes: For each identified key dynamic target, calculate the acceleration change rate and detection confidence of consecutive A frames, and combine the acceleration change rate and detection confidence into a dynamic attribute vector; where A frames is the length of the dynamic attribute calculation window, which is the number of consecutive frames used to calculate the dynamic attributes. Extract the appearance feature descriptors of key dynamic targets and store them in association with the dynamic attribute vectors; establish a memory pool to store only the dynamic attributes and appearance feature descriptors of the most recent G motion trajectory segments. When a new key dynamic target enters the memory pool, if the memory pool is full, remove the earliest dynamic attributes according to the FIFO rule; where G is the maximum number of trajectory segments that the memory pool can store. Count the number of key dynamic targets in the current highway scene image and estimate the density of the distribution of key dynamic targets, that is, calculate the highway scene density factor based on the number of key dynamic targets in the current frame and the area of the highway scene area where they are located. Based on the highway scene density factor, adjust the basic association matching threshold and the basic detection confidence threshold. Perform motion trajectory association and initialization based on the memory bank: For each key dynamic target in the current frame, compare it with the motion trajectory segments in the memory pool and calculate the association cost, that is, fuse the Mahalanobis distance and the appearance distance as the association cost. If a key dynamic target continuously satisfies that the association cost is less than the basic association matching threshold and the detection confidence is greater than the basic detection confidence threshold for N consecutive frames, it is considered that this key dynamic target belongs to this motion trajectory, assign it a unique ID, and officially initialize the motion trajectory to generate an initial motion trajectory with a unique ID; where N frames is the length of the motion trajectory initialization window, which is the number of consecutive matching frames for motion trajectory initialization. Perform motion trajectory life cycle management for the initial motion trajectory, including short-term trajectory management and long-term trajectory management. Among them, for the detection results that are not associated with any motion trajectory segments in the memory pool, manage them as short-term motion trajectories; for the key dynamic targets that are successfully associated with any motion trajectory segments in the memory pool, manage them as long-term motion trajectories. After completing the motion trajectory life cycle management, output the unique ID of the key dynamic target and the motion trajectory route.

[0012] Furthermore, the motion trajectory analysis unit is used to perform spatial and temporal enhancement processing on the motion trajectory data output by multi-target tracking, fuse appearance features and dynamic attributes to generate multi-modal motion trajectory features, and extract initial behavior pattern labels based on the motion trajectory point sequence; initialize the behavior map through traffic rules, combine historical abnormal motion trajectories to cluster and mine new abnormal behavior nodes, and dynamically update the map structure to form a dynamic behavior map containing normal and abnormal patterns; use spatio-temporal attention Transformer to encode the motion trajectory features, calculate the abnormal score and path deviation degree in combination with the dynamic behavior map, comprehensively judge the trajectory abnormality and map it to a natural language description, and finally classify according to the degree of abnormality; where the specific process is as follows: Obtain the motion trajectory data output by the multi-target tracking unit, including the unique ID, the sequence of motion trajectory points, the appearance feature identifier, and the dynamic attribute vector; perform enhancement processing on the motion trajectory points, appearance feature identifiers, and dynamic attribute vectors of key dynamic targets, and fuse them into multi-modal motion trajectory features; Among them, perform spatial interpolation on the motion trajectory points to generate a smooth motion trajectory curve, and obtain a spatially enhanced sequence of motion trajectory points; introduce a timestamp sequence, calculate the instantaneous velocity, acceleration, direction angle, and acceleration change rate of the motion trajectory points for time enhancement; perform temporal weighted averaging on the appearance feature vectors of the same target in the motion trajectory to generate an enhanced appearance feature representation; extract behavior pattern labels based on the sequence of motion trajectory points as initial semantic annotations; Construct an initial graph based on traffic rules. In the constructed initial graph, the behavior state nodes are the core components of the initial graph, representing various traffic behavior states, and the state transition edges are used to represent the transition relationships between traffic behavior states; Obtain the historical motion trajectory database, cluster the abnormal motion trajectories in the historical motion trajectory database, and extract abnormal behavior patterns; add the mined abnormal patterns as new behavior state nodes to the initial graph and initialize the transition probabilities; receive new motion trajectory data in real time and dynamically update the graph structure, thereby constructing a dynamic behavior graph; Introduce a spatio-temporal attention Transformer to encode the multi-modal motion trajectory features; compare the encoded multi-modal motion trajectory features with the feature distribution of normal behavior patterns in the dynamic behavior graph to calculate the anomaly score; calculate the path deviation degree of the multi-modal motion trajectory on the dynamic behavior graph; Integrate the anomaly score and the graph deviation degree to determine whether the motion trajectory of the key dynamic target is abnormal; map the path of the abnormal motion trajectory on the dynamic behavior graph to a natural language description; perform weighted calculation according to the anomaly score and the graph deviation degree to obtain the anomaly level, including minor anomaly, serious anomaly, and dangerous anomaly.

[0013] Furthermore, the visualization unit is used to overlay and display the motion trajectory analysis results and the anomaly degree classification of the key dynamic target on the electronic map in real time.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: By collecting highway scene videos in real time and transmitting them back with low latency, the timeliness and integrity of monitoring data are ensured, providing a reliable basis for subsequent analysis; By preprocessing the video data, the data quality is effectively improved, and at the same time, the UAV path is dynamically adjusted according to the target movement trajectory, significantly improving the accuracy and efficiency of monitoring; By combining the one-shot detector model and the optical flow method, the key dynamic target features are quickly and accurately extracted and the UAV interference is cancelled, obtaining the real movement information of the target, enhancing the robustness of target detection; By combining deep learning algorithms with multi-target tracking technology, the movement trajectories of key dynamic targets are obtained, providing reliable data for subsequent trajectory analysis; By analyzing the movement trajectories of key dynamic targets, the abnormality of the movement trajectories is judged and classified, which helps to timely discover and respond to potential risks. At the same time, the analysis results of the movement trajectories are visualized, enabling monitoring personnel to intuitively understand the target movement pattern and improving the overall intelligence level of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 FIG. is a module diagram of a multi-target tracking system for highway scenes based on UAV vision perception proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0017] Refer to Figure 1 , a multi-target tracking system for highway scenes based on UAV vision perception, the system includes a UAV platform, a ground control station, a target detection and tracking module, and a movement trajectory analysis and visualization module; UAV platform: Collect highway scene videos in real time and transmit the collected video data back; The UAV platform includes an image acquisition unit and a data transmission unit; The image acquisition unit is equipped with a 4K high-definition camera and a gimbal stabilization system for collecting real-time video data of the highway scene; The data transmission unit uses 5G communication technology for low-latency transmission of video data; Ground control station: Process the video data transmitted by the UAV and generate flight control instructions to dynamically adjust the path; Among them, the ground control station includes a data preprocessing unit and a flight control unit; The data preprocessing unit is used to denoise, perform illumination equalization and geometric correction on the received video data; The flight control unit is used to control the flight of the UAV and dynamically adjust the flight path according to the analysis result of the movement trajectory of the key dynamic target; Target detection and tracking module: Extract the key dynamic target features of the road scene through the one-shot detector model and perform cross-frame matching and statistics. At the same time, combine the optical flow method to analyze the motion vectors of static objects to offset the UAV interference and obtain the real motion information of the key dynamic targets; Use deep learning algorithms to detect targets in the video frames captured by the UAV; Adopt multi-object tracking algorithms to track the detected key dynamic targets, so as to obtain the motion trajectories of the key dynamic targets; Among them, the target detection and tracking module includes a key dynamic target acquisition unit, a key dynamic target association unit, a motion compensation unit, a key dynamic target detection unit, and a key dynamic target tracking unit; Motion trajectory analysis and visualization module: Analyze the motion trajectories of the key dynamic targets and visualize the analysis results of the motion trajectories of the key dynamic targets; Among them, the motion trajectory analysis and visualization module includes a motion trajectory analysis unit and a visualization unit.

[0018] It should be further noted that in the specific implementation process, the key dynamic target acquisition unit is used to obtain the preprocessed video data, extract key features through the training of the one-shot detector model, and identify and count the key dynamic targets in the road scene , where represents the key dynamic target, represents the number of key dynamic targets; The specific process is as follows: Input the preprocessed video frame by frame into the one-shot detector model, and use the model to quickly adapt to the target detection task in the road scene; Directly generate candidate target regions through the model; Combine a small number of target annotation samples to quickly identify the potential key dynamic targets in the candidate target regions; It can be understood that the One-shot detector model is a deep learning model that can quickly learn and identify target features through only a small number of samples (even a single sample), and is used to efficiently extract the features of key dynamic targets in video data.

[0019] It should be further noted that in the specific implementation process, the key dynamic target association unit is used to determine whether the key dynamic targets in consecutive frames are the same target by using the structural similarity index. For the key dynamic target in the current frame and the key dynamic target in the next frame ( both are key dynamic target indexes), calculate the similarity of the key dynamic targets , and its calculation formula is: , in the formula, respectively represent the similarity of brightness, contrast, and structure; are preset different weight coefficients, corresponding to brightness, contrast, and structure respectively; among them, the structural similarity index is used to measure the similarity of images, and measures similarity from three aspects: brightness, contrast, and structure; If the value of the structural similarity index of the key dynamic target in consecutive frames exceeds the preset cross-frame structural similarity matching threshold, it is determined to be the same target; among them, the cross-frame structural similarity matching threshold is used as the criterion for determining whether the key dynamic targets in consecutive frames are the same entity.

[0020] It should be further noted that in the specific implementation process, the motion compensation unit is used to calculate the motion vector of pixel points in the image between consecutive frames for static objects in the road scene by using the optical flow method, and fuse the weighted average of the average and mode of the moving direction and speed of key pixel points to offset the influence of the UAV's own motion on target tracking, and obtain the true pixel motion vector of the key dynamic target, that is, the offset value of the moving direction and speed of the key dynamic target in the road scene. The combination of the two obtains the true moving direction and speed of the key dynamic target in the scene: , where, is the true pixel motion vector of the key dynamic target, is the pixel motion vector calculated by the optical flow method (including the influence of the UAV's own motion), is the set of motion vectors of key pixel points; is the average of the motion vectors of key pixel points (used to offset global drift), is the mode of the motion vectors of key pixel points (used to suppress the interference of outliers), is the weight coefficient of UAV motion compensation, is the weighted coefficient of the average and the mode.

[0021] It should be further noted that in the specific implementation process, the key dynamic target detection unit is used to extract features from the input video frame by using the optimized YOLO-X network structure, and combine the temporal feature fusion and dynamic threshold adjustment strategies to perform fine-grained detection on the already determined key dynamic targets. The process of identifying the key dynamic target type includes: After the initial feature extraction by the YOLO backbone network, based on the analysis of target scale differences, a dynamic feature selection mechanism is constructed. This mechanism is used to adaptively allocate feature weights of different levels according to the key dynamic target scales to generate a preliminary target feature map. Specifically, the analysis of target scale differences refers to the analysis of feature differences of targets at different scales (such as small-scale pedestrians in the distance and large-scale vehicles nearby), so as to adaptively allocate feature weights of different levels according to the target scale during feature extraction. That is, for small-scale targets (such as pedestrians in the distance), more weights are assigned to deep semantic features to obtain richer semantic information; for large-scale targets (such as vehicles nearby), the weights of shallow spatial features are appropriately increased to retain more detailed information, so as to obtain a more target-specific feature representation. Based on the generated preliminary target feature map, it is combined with the temporal characteristics of the video sequence. Specifically: each frame of the target feature map is regarded as a temporal node in the temporal feature map, and these temporal nodes are arranged in sequence in the time dimension, representing the states of target features at different times; at the same time, the relationship between target features of adjacent frames is defined as the temporal edge in the temporal feature map, and this relationship is quantified by calculating the difference degree index between target features of adjacent frames. Among them, the difference degree index can be quantified by calculating the Euclidean distance or cosine similarity between target feature vectors of adjacent frames to measure the amplitude of feature change; in the video sequence, there are continuous motions and state changes between different frames. By constructing the temporal feature map, the temporal correlation relationship of key dynamic targets in the time dimension is represented, that is, the temporal correlation information between different frames; using the feature vectors of key dynamic targets as the attributes of temporal nodes in the temporal feature map and the relative position relationship between key dynamic targets as the weights of temporal edges in the temporal feature map to capture the feature information and spatial relationship of key feature targets in the temporal sequence. After the construction of the temporal feature map is completed, it is input into the graph neural network for processing. The graph neural network operates on the temporal nodes and temporal edges in the temporal feature map: during the information transmission process, the temporal nodes will receive information from adjacent temporal nodes and at the same time transmit their own information to adjacent temporal nodes; in the information aggregation stage, the temporal nodes integrate the received information to update their own feature representations; the graph neural network learns the temporal feature correlation of key dynamic targets between different frames and encodes the dynamic change information of key dynamic targets in the temporal sequence into the feature representation; through the processing of multiple layers of graph neural networks, a feature representation integrating temporal information is obtained, where this feature representation contains the dynamic change information of key dynamic targets in the entire temporal sequence; a concatenation operation is used to fuse the feature representation integrating temporal information output by the graph neural network with the current frame target feature map to generate a target feature map after fusing temporal information; according to the target feature map after fusing temporal information, key dynamic target types are identified, including vehicles and pedestrians.

[0022] It should be further noted that in the specific implementation process, the key dynamic target tracking unit is used to receive the output of the key dynamic target detection unit. By constructing an incremental dynamic feature memory, combining with the dynamic adjustment of the threshold based on the scene density, realizing trajectory association and initialization through the memory bank, and managing short-term and long-term motion trajectories, the process of generating the unique ID of the key dynamic target and the motion trajectory includes: For each identified key dynamic target, calculate the acceleration change rate and detection confidence of consecutive A frames, and combine the acceleration change rate and detection confidence into a dynamic attribute vector; Among them, the acceleration change rate of consecutive A frames is: , where is the instantaneous acceleration of the t-th frame; is the target speed of the t-th frame (calculated by the position difference of adjacent frames); is the frame time interval (default 1 frame); A frames is the length of the dynamic attribute calculation window, which is used to calculate the consecutive number of frames of the dynamic attribute. This window length determines the sensitivity of the dynamic attribute to the short-term motion changes of the key dynamic target (usually shorter, such as 3 - 5 frames); The dynamic attribute vector is: , where is the detection confidence; Extract the appearance feature descriptor of the key dynamic target (i.e., the ReID feature vector), and store it in association with the dynamic attribute vector; establish a memory pool, and only store the dynamic attributes and appearance feature descriptors of the most recent G motion trajectory segments. When a new key dynamic target enters the memory pool, if the memory pool is full, remove the earliest dynamic attribute according to the FIFO rule (first in, first out); among them, G is the maximum number of stored trajectory segments in the memory pool; Count the number of key dynamic targets in the current highway scene picture, and estimate the density of the distribution of key dynamic targets: According to the number of key dynamic targets in the current frame and the area of the highway scene area where they are located, calculate the highway scene density factor; Among them, the highway scene density factor The calculation formula is: , where is the number of key dynamic targets in the current frame, is the highway scene area; Based on the highway scene density factor, adjust the basic association matching threshold and adjust the basic detection confidence threshold: , where is the adjusted basic association matching threshold, is the basic association matching threshold, which is used to judge the association degree between the key dynamic target and the motion trajectory segment in the memory pool. When the association cost is less than this threshold, the association is considered successful; is the adjusted basic detection confidence threshold, is the basic detection confidence threshold, which is used to filter out the detection results with low confidence. Only the key dynamic targets with detection confidence greater than this threshold will be further processed; are respectively the positive adjustment coefficient of the road scene density factor to the basic association matching threshold and the negative adjustment coefficient to the basic detection confidence threshold; For each key dynamic target in the current frame, compare it with the motion trajectory segments in the memory pool and calculate the association cost: the Mahalanobis distance and the appearance distance are weighted and fused as the association cost; Among them, the association cost The calculation formula is: , where is the fusion weight of the Mahalanobis distance and the appearance distance; is the Mahalanobis distance, , is the motion trajectory predicted by the Kalman filter position, is the key dynamic target detection position, is the matrix transpose symbol, is the Kalman filter covariance matrix; is the appearance distance, , is the motion trajectory and the key dynamic target appearance feature descriptors; among them, the Mahalanobis distance is based on the difference between the motion trajectory position predicted by the Kalman filter and the key dynamic target detection position, and the appearance distance is based on the Euclidean distance between the key dynamic target and the appearance feature descriptor; If the key dynamic target continuously satisfies the condition that the association cost is less than the basic association matching threshold and the detection confidence is greater than the basic detection confidence threshold for N consecutive frames, then it is considered that the key dynamic target belongs to this motion trajectory, assign a unique ID to it, and formally initialize the motion trajectory to generate an initial motion trajectory with a unique ID: , where is the starting frame number when the current key dynamic target first satisfies the association condition (i.e., the first frame to start counting), which is used to mark the starting point of the successful match to ensure the stability of N consecutive frames; for example, assume N = 5 (5 consecutive frames need to be successfully matched), if the key dynamic target first satisfies the condition at frame number first, then it is necessary to check whether frames 10 - 14 all satisfy the conditions; if frames 10 - 14 all pass, then at Initialize the trajectory at the beginning; if a certain frame in the middle fails (such as frame 12 not being matched), reset the count; N frames are the window length for initializing the motion trajectory, which is the number of consecutive matching frames for motion trajectory initialization. This window length determines the reliability requirement for motion trajectory initialization (usually longer, such as 5 - 10 frames). Perform motion trajectory lifecycle management for the initial motion trajectory, including short - term trajectory management and long - term trajectory management. For the detection results that are not associated with any motion trajectory segments in the memory pool (i.e., the case marked as a new target), manage them as short - term motion trajectories; after the short - term motion trajectory lasts for Z frames, if it is still not associated with any motion trajectory segments in the memory pool, it is determined that there is a misdetection or a transient key dynamic target in this motion trajectory, and it is destroyed; where Z frames are the window length for maintaining the short - term trajectory, which is used to control the maximum survival frames of the motion trajectory not associated with the memory pool. This window length determines the tolerance time of the lifecycle management for transient key dynamic targets (usually set to 10 - 20 frames, and the specific value needs to be adjusted according to the scene dynamics); for the key dynamic targets that are successfully associated with any motion trajectory segments in the memory pool, manage them as long - term motion trajectories, and continuously update the dynamic attribute vector (acceleration change rate, detection confidence) and appearance feature identifier of the long - term motion trajectory to reflect the latest state of the key dynamic target during the motion process; the formula for updating the dynamic attributes and appearance features by moving average is: , In the formula, is the forgetting factor, that is, the retention ratio of historical information when updating the long - term trajectory, is the updated dynamic attribute vector, is the appearance feature identifier, is the updated appearance feature identifier; After completing the motion trajectory lifecycle management, output the unique ID of the key dynamic target and the motion trajectory route (where the motion trajectory route consists of a series of motion trajectory points, which is used to reflect the motion trajectory of the key dynamic target in the video sequence).

[0023] It should be further noted that in the specific implementation process, the motion trajectory analysis unit is used to perform spatial and temporal enhancement processing on the motion trajectory data output by multi-object tracking, fuse appearance features and dynamic attributes to generate multi-modal motion trajectory features, and extract initial behavior pattern labels (such as going straight, lane changing) based on the motion trajectory point sequence; initialize the behavior graph through traffic rules, combine historical abnormal motion trajectories to cluster and mine new abnormal behavior nodes, and dynamically update the graph structure (such as transition probability) to form a dynamic behavior graph containing normal and abnormal patterns; use spatio-temporal attention Transformer to encode the motion trajectory features, calculate the anomaly score and path deviation degree in combination with the dynamic behavior graph, comprehensively judge the trajectory anomaly and map it to a natural language description, and finally classify according to the anomaly degree (slight, severe, dangerous); the specific process is as follows: Obtain the motion trajectory data output by the multi-object tracking unit , including the unique ID, the motion trajectory point sequence, the appearance feature identifier, and the dynamic attribute vector: , where is the key dynamic target index after the output of the multi-object tracking unit; is the number of key dynamic targets after the output of the multi-object tracking unit; is the unique identifier of the key dynamic target; is the motion trajectory point sequence, is the motion trajectory point, is the motion trajectory length, m is the motion trajectory index; is the set of appearance feature identifiers; is the dynamic attribute vector sequence, is the dynamic attribute vector value, including the instantaneous speed, acceleration, direction angle, and acceleration change rate, that is ; Perform enhancement processing on the motion trajectory points, appearance feature identifiers, and dynamic attribute vectors of the key dynamic targets, and fuse them into multi-modal motion trajectory features : , where is the enhanced dynamic attribute vector, is the vector concatenation operation; Among them, perform spatial interpolation (such as cubic spline interpolation) on the motion trajectory points to generate a smooth motion trajectory curve, and obtain the spatially enhanced motion trajectory point sequence and the set of appearance feature identifiers ; Introduce a timestamp sequence, calculate the instantaneous speed, acceleration, direction angle, and acceleration change rate of the motion trajectory points, and perform time enhancement: ; Perform temporal weighted averaging on the appearance feature vectors of the same target in the motion trajectory to generate an enhanced appearance feature representation : , where is the detection confidence is the appearance feature identifier Extract behavior pattern labels (such as going straight, changing lanes, staying) based on the motion trajectory point sequence as the initial semantic annotation Construct an initial graph based on traffic rules. In the constructed initial graph , the behavior state nodes are the core components of the initial graph , representing various traffic behavior states (such as normal driving, speeding, reverse driving), and the state transition edges are used to represent the transition relationships between traffic behavior states. Initially, it is assumed that the probabilities of all state transitions are uniformly distributed: , where are the set of initial behavior state nodes, the initial state transition edges respectively is the behavior state node is the behavior state 's initial transition probability are all indices of behavior state nodes is the number of behavior state nodes Obtain the historical motion trajectory database (including normal behavior patterns and known abnormal patterns), cluster the abnormal motion trajectories in the historical motion trajectory database (such as DBSCAN), and extract abnormal behavior patterns (such as frequent lane changes, abnormal stays); add the mined abnormal patterns as new behavior state nodes to the initial graph and initialize the transition probabilities; receive new motion trajectory data in real time and dynamically update the graph structure (including adding new abnormal behavior state nodes and adjusting transition probabilities), thereby constructing a dynamic behavior graph (including normal behavior patterns and known abnormal patterns): , where is the transition probability of the behavior state after the initial graph is updated is the smoothing factor is the actual number of transitions of the behavior state is a temporary index variable for summation, used to traverse all behavior state nodes is the actual number of transitions of all behavior states ​Introduce a spatio-temporal attention Transformer to encode multi-modal motion trajectory features, where the spatio-temporal attention includes temporal attention, spatial attention, and graph attention. Temporal attention is used to capture the temporal dependencies between motion trajectory points, spatial attention is used to strengthen the spatial correlation of motion trajectory points (such as the mutual influence of adjacent motion trajectory points), and graph attention fuses the behavior state node features of the dynamic behavior graph with the motion trajectory features to enhance behavior semantic understanding; Compare the encoded multi-modal motion trajectory features with the feature distribution of normal behavior patterns in the dynamic behavior graph to calculate the anomaly score : , where, is the multi-modal motion trajectory feature, are the mean and covariance of the normal motion trajectory features respectively, is the matrix transpose symbol; Calculate the path deviation degree of the multi-modal motion trajectory on the dynamic behavior graph (i.e., the difference between the current trajectory path and the reference shortest path): , where, is the current trajectory path, is the reference shortest path; Integrate the anomaly score and the graph deviation degree to judge whether the motion trajectory of the key dynamic target is abnormal; Map the path of the abnormal motion trajectory on the dynamic behavior graph to a natural language description (such as the key dynamic target reverses from lane A to lane B and lasts for 5 seconds); For example, use a pre-trained language model (such as BERT) to perform grammar and logic verification on the generated semantics to ensure consistency with traffic rules; Perform weighted calculation based on the anomaly score and the graph deviation degree to obtain anomaly levels, including minor anomaly, serious anomaly, and dangerous anomaly. For example, minor anomaly - short-term speeding, lasting < 3 seconds, speed > 10% of the speed limit; serious anomaly - reverse, lasting > 3 seconds, direction opposite to the mainstream direction; dangerous anomaly - collision risk, distance from adjacent motion trajectory < preset safety threshold.

[0024] It should be further noted that in the specific implementation process, the visualization unit is used to overlay and display the motion trajectory analysis results of the key dynamic target and the anomaly level classification situation on the electronic map in real time, facilitating users to intuitively understand the highway traffic conditions.

[0025] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0026] It should be understood that determining B based on A does not mean determining B solely based on A, and B can also be determined based on A and / or other information.

[0027] As described above, it is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0028] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-target tracking system for highway scenes based on UAV visual perception, characterized in that: It includes a drone platform, a ground control station, a target detection and tracking module, and a motion trajectory analysis and visualization module; Drone platform: It collects road scene videos in real time and transmits the collected video data back. The drone platform includes an image acquisition unit and a data transmission unit; Ground control station: It processes the video data transmitted by the drone and generates flight control instructions to dynamically adjust the path. The ground control station includes a data preprocessing unit and a flight control unit; Target detection and tracking module: It extracts the key dynamic target features of the road scene through the one-shot detector model and performs cross-frame matching and statistics. At the same time, it combines the optical flow method to analyze the motion vectors of static objects to offset the drone interference and obtains the real motion information of the key dynamic targets; It uses deep learning algorithms to detect targets in the video frames captured by the drone; it uses multi-target tracking algorithms to track the detected key dynamic targets, thereby obtaining the motion trajectories of the key dynamic targets. Among them, the target detection and tracking module includes a key dynamic target acquisition unit, a key dynamic target association unit, a motion compensation unit, a key dynamic target detection unit, and a key dynamic target tracking unit; Motion trajectory analysis and visualization module: It analyzes the motion trajectories of the key dynamic targets and visualizes the analysis results of the motion trajectories of the key dynamic targets. Among them, the motion trajectory analysis and visualization module includes a motion trajectory analysis unit and a visualization unit.

2. The multi-object tracking system for highway scenarios based on UAV vision perception according to claim 1, characterized in that: The image acquisition unit is equipped with a 4K high-definition camera and a gimbal stabilization system for collecting real-time video data of the road scene; the data transmission unit uses 5G communication technology for low-latency transmission of video data.

3. A multi-object tracking system for highway scenarios based on UAV visual perception according to claim 1, characterized in that: The data preprocessing unit is used to denoise, perform illumination equalization, and geometric correction on the received video data; the flight control unit is used to control the flight of the drone and dynamically adjust the flight path according to the analysis results of the motion trajectories of the key dynamic targets.

4. The multi-object tracking system for highway scenarios based on UAV vision perception according to claim 1, wherein: The key dynamic target acquisition unit is used to obtain the preprocessed video data, perform key feature extraction using the one-shot detector model training, and identify and count the key dynamic targets in the road scene.

5. The multi-object tracking system for highway scenarios based on UAV vision perception according to claim 1, wherein: The key dynamic target association unit is used to use the structural similarity index to judge whether the key dynamic targets in consecutive frames are the same target. For the key dynamic target in the current frame and the key dynamic target in the next frame, calculate the similarity of the key dynamic targets; if the value of the structural similarity index of the key dynamic targets in consecutive frames exceeds the preset cross-frame structural similarity matching threshold, it is determined to be the same target.

6. The multi-object tracking system for highway scenes based on UAV visual perception according to claim 1, characterized in that: The motion compensation unit is used to calculate the motion vectors of pixel points in the image between consecutive frames for static objects in the road scene using the optical flow method, and fuse the weighted average of the average and mode of the moving direction and speed of the key pixel points to obtain the real pixel motion vector of the key dynamic target.

7. A multi-target tracking system for highway scenarios based on UAV vision perception according to any one of claims 1-6, characterized in that: The key dynamic target detection unit is used to extract features from the input video frames using the optimized YOLO-X network structure, combine the temporal feature fusion and dynamic threshold adjustment strategy, and perform fine-grained detection on the already determined key dynamic targets. The process of identifying the types of key dynamic targets includes: After the initial feature extraction by the YOLO backbone network, a dynamic feature selection mechanism is constructed based on the analysis of target scale differences. This mechanism is used to adaptively allocate feature weights at different levels according to the key dynamic target scales, generating a preliminary target feature map. Based on the generated preliminary target feature map, it is combined with the temporal characteristics of the video sequence. Specifically: each frame's target feature map is regarded as a temporal node in the temporal feature map; at the same time, the target feature relationship between adjacent frames is defined as the temporal edge in the temporal feature map; the feature vector of the key dynamic target is used as the attribute of the temporal node in the temporal feature map, and the relative position relationship between the key dynamic targets is used as the weight of the temporal edge in the temporal feature map to capture the feature information and spatial relationship of the key feature targets in the temporal sequence. After the construction of the temporal feature map is completed, it is input into the graph neural network for processing, where the graph neural network operates on the temporal nodes and temporal edges in the temporal feature map; the graph neural network learns the temporal feature associations of the key dynamic targets between different frames, encoding the dynamic change information of the key dynamic targets in the temporal sequence into the feature representation; through the processing of multiple layers of graph neural networks, a feature representation integrating temporal information is obtained, where this feature representation contains the dynamic change information of the key dynamic targets in the entire temporal sequence; a concatenation operation is used to fuse the feature representation integrating temporal information output by the graph neural network with the current frame's target feature map, generating a target feature map after fusing temporal information; based on the target feature map after fusing temporal information, the types of key dynamic targets are identified.

8. A multi-object tracking system for highway scenarios based on UAV vision perception according to claim 7, characterized in that: The key dynamic target tracking unit is used to receive the output of the key dynamic target detection unit. The process of generating the unique ID of the key dynamic target and the motion trajectory by constructing an incremental dynamic feature memory, dynamically adjusting the threshold in combination with the scene density, and realizing trajectory association and initialization through the memory bank and managing short-term and long-term motion trajectories includes: For each identified key dynamic target, calculate the acceleration change rate and detection confidence for consecutive A frames, and combine the acceleration change rate and detection confidence into a dynamic attribute vector; where A frames is the length of the dynamic attribute calculation window, which is the number of consecutive frames used to calculate the dynamic attributes. Extract the appearance feature descriptor of the key dynamic target and store it in association with the dynamic attribute vector; establish a memory pool that only stores the dynamic attributes and appearance feature descriptors of the most recent G motion trajectory segments. When a new key dynamic target enters the memory pool, if the memory pool is full, the earliest dynamic attribute is removed according to the FIFO rule; where G is the maximum number of stored trajectory segments in the memory pool. Count the number of key dynamic targets in the current highway scene picture and estimate the density of the distribution of key dynamic targets, that is, calculate the highway scene density factor according to the number of key dynamic targets in the current frame and the area of the highway scene area where they are located. Based on the highway scene density factor, adjust the basic association matching threshold and adjust the basic detection confidence threshold. Motion trajectory association and initialization based on the memory pool: For each key dynamic target in the current frame, compare it with the motion trajectory segments in the memory pool and calculate the association cost, that is, weighted fusion of the Mahalanobis distance and the appearance distance as the association cost; If the key dynamic target continuously satisfies that the association cost is less than the basic association matching threshold and the detection confidence is greater than the basic detection confidence threshold for N consecutive frames, it is considered that the key dynamic target belongs to this motion trajectory, assign a unique ID to it, and formally initialize the motion trajectory to generate an initial motion trajectory with a unique ID; where N frames is the motion trajectory initialization window length, which is the number of consecutive matching frames for motion trajectory initialization; Perform motion trajectory life cycle management for the initial motion trajectory, including short-term trajectory management and long-term trajectory management. Among them, for the detection results that are not associated with any motion trajectory segments in the memory pool, manage them as short-term motion trajectories; for the key dynamic targets that are successfully associated with any motion trajectory segments in the memory pool, manage them as long-term motion trajectories; After completing the motion trajectory life cycle management, output the unique ID and motion trajectory route of the key dynamic target.

9. The multi-object tracking system for highway scenes based on UAV visual perception according to claim 8, wherein: The motion trajectory analysis unit is used to perform spatial and temporal enhancement processing on the motion trajectory data output by multi-object tracking, fuse appearance features and dynamic attributes to generate multi-modal motion trajectory features, and extract initial behavior pattern labels based on the motion trajectory point sequence; Initialize the behavior graph through traffic rules, combine historical abnormal motion trajectory clustering to mine new abnormal behavior nodes, and dynamically update the graph structure to form a dynamic behavior graph containing normal and abnormal patterns; use spatio-temporal attention Transformer to encode motion trajectory features, combine the dynamic behavior graph to calculate the abnormal score and path deviation degree, comprehensively judge the trajectory abnormality and map it to a natural language description, and finally classify according to the degree of abnormality; where the specific process is: Obtain the motion trajectory data output by the multi-object tracking unit, including the unique ID, motion trajectory point sequence, appearance feature identifier, and dynamic attribute vector; perform enhancement processing on the motion trajectory points, appearance feature identifiers, and dynamic attribute vectors of the key dynamic target, and fuse them into multi-modal motion trajectory features; Among them, perform spatial interpolation on the motion trajectory points to generate a smooth motion trajectory curve and obtain a spatially enhanced motion trajectory point sequence; introduce a timestamp sequence, calculate the instantaneous speed, acceleration, direction angle, and acceleration change rate of the motion trajectory points for temporal enhancement; perform temporal weighted averaging on the appearance feature vectors of the same target in the motion trajectory to generate an enhanced appearance feature representation; extract behavior pattern labels based on the motion trajectory point sequence as the initial semantic annotation; Construct an initial graph based on traffic rules. In the constructed initial graph, the behavior state nodes are the core components of the initial graph, representing various traffic behavior states, and the state transition edges are used to represent the transition relationships between traffic behavior states; Obtain the historical movement trajectory database, cluster the abnormal movement trajectories in the historical movement trajectory database, and extract abnormal behavior patterns; Add the mined abnormal patterns as new behavior state nodes to the initial graph and initialize the transition probabilities; Receive new movement trajectory data in real time and dynamically update the graph structure, thereby constructing a dynamic behavior graph; Introduce a spatio-temporal attention Transformer to encode multi-modal movement trajectory features; Compare the encoded multi-modal movement trajectory features with the feature distribution of normal behavior patterns in the dynamic behavior graph and calculate the anomaly score; Calculate the path deviation degree of the multi-modal movement trajectory on the dynamic behavior graph; Based on the comprehensive anomaly score and graph deviation degree, determine whether the movement trajectory of the key dynamic target is abnormal; Map the path of the abnormal movement trajectory on the dynamic behavior graph to a natural language description; Perform weighted calculation according to the anomaly score and graph deviation degree to obtain the anomaly level, including minor anomaly, severe anomaly, and dangerous anomaly.

10. A multi-object tracking system for highway scenarios based on UAV vision perception according to claim 9, characterized in that: The visualization unit is used to overlay and display the analysis results of the movement trajectory of the key dynamic target and the anomaly level classification in real time on the electronic map.

Citation Information

Patent Citations

  • Target detection-based tracking method and device, and terminal equipment

    CN110047095A

  • Target tracking method based on graph convolution and trajectory convolution network learning

    CN110660082A

  • Method and system for tracking vehicle by unmanned aerial vehicle based on computer vision

    CN118534934A

  • Traffic tracking detection system for view angle of unmanned aerial vehicle

    CN118918148A

  • Road traffic anomaly detection and related equipment based on holographic perception

    CN119540832A

Cited By

  • Dynamic bright spot target statistics and analysis method based on visual tracking

    CN120726540A

  • Anti-unmanned aerial vehicle intelligent identification and tracking system based on multi-source data fusion

    CN120850103A