Expressway abnormal event detection and tracking method under view angle of unmanned aerial vehicle
Through the improved YOLO v11 and DeepSORT network structure, combined with attention mechanism and wavelet analysis method, the small object detection accuracy and tracking stability problems in highway traffic event detection from the perspective of the drone are solved, and the accurate identification and processing of abnormal events on the highway is achieved.
Patent Information
- Application Number
- CN202510497686.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
The existing highway traffic event detection technology from the perspective of drones has problems such as insufficient detection accuracy of small targets, poor target tracking stability, poor noise processing of trajectory data, and low accuracy in determining abnormal events.
The improved YOLO v11 network structure and DeepSORT network structure are adopted, combined with the CBAM attention mechanism, Focal-WIOU Loss function and wavelet analysis method, object detection and tracking are performed, and abnormal events are identified through coordinate conversion and trajectory processing.
It significantly improves the detection accuracy and tracking stability of small targets, improves the accuracy of identification of abnormal events, and realizes the timely detection and processing of various types of abnormal events on the highway.
Smart Images

Figure CN120339885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection and tracking, and particularly to a method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle. Background Art
[0002] In recent years, the technology of Intelligent Transport System (ITS) has continuously become a hot topic in the field of traffic research. Among them, Automatic Incident Detection (AID), as a key component of the ITS system, realizes the instant detection of traffic events by collecting various traffic data in real time and using automatic detection algorithms. The automatic detection of traffic events is of great significance for improving traffic safety. It can quickly and accurately capture traffic events and immediately notify relevant departments to intervene and handle them, effectively reducing the risk of casualties caused by accidents and preventing the occurrence of secondary accidents. In addition, while automatically detecting traffic events, it also provides valuable traffic data for traffic management departments, helping them develop the digital economy, optimize management decisions, and improve work efficiency.
[0003] With the rapid development of artificial intelligence and intelligent detection technology, in recent years, the research on automatic detection algorithms has been continuously deepened. Currently, in the research on highway traffic event detection algorithms, intelligent algorithms widely used in target detection and target tracking. Target detection is to extract the target information in highway pictures and then classify and locate the target. Target tracking is to capture the movement trajectory of vehicles in consecutive frames, obtain feature information such as the position, speed, and direction of vehicles, and then use this feature information to determine traffic events. However, the detection and tracking of most vehicles are based on traditional fixed road surveillance cameras, and fixed cameras have disadvantages such as poor flexibility and limited field of view. Therefore, many researchers have begun to look forward to using unmanned aerial vehicles (UAVs) with small volume, wide field of view, flexible movement, and strong adaptability to collect traffic information. When UAVs are deployed for road traffic monitoring, they can serve as important auxiliary tools to fill the monitoring gaps in areas not covered by fixed cameras during specific periods. With the excellent mobility and flexibility of UAVs, they can break free from the confinement of traditional fixed cameras and effectively identify, detect, and track target vehicles, thus greatly expanding the breadth and depth of monitoring. Therefore, the detection and tracking of vehicles in traffic videos from the perspective of UAVs is a research topic worthy of study and has high practical significance and future hotspots for realizing the detection of traffic abnormal events on highways.
[0004] However, there are still a series of significant deficiencies in the existing traffic event detection technology from the perspective of drones. First, the existing object detection models perform poorly in processing vehicle images from the top-down perspective of drones. In particular, the detection accuracy of small-sized targets is insufficient, and there are prone to missed detections and false detections under complex backgrounds and lighting conditions. Although the traditional YOLO series models have fast detection speeds, in processing highway scenes captured by drones, due to factors such as small target sizes, high densities, and special perspectives, the detection effects far from meet the actual requirements. Second, the existing object tracking algorithms have poor tracking stability in high-speed movement, occlusion, and dense scenarios. The target IDs frequently switch, and it is impossible to accurately maintain the long-term vehicle identity association, resulting in broken trajectories. Traditional tracking algorithms based on Kalman filtering are prone to losing targets in complex highway scenes and are difficult to handle complex behaviors such as vehicle lane changes and temporary parking. In addition, the jitter, altitude changes, and perspective conversions of the drones themselves also pose additional challenges to detection and tracking, and the existing methods have insufficient adaptability to these factors. Finally, for the detected trajectory data, the existing abnormal event recognition algorithms lack effective processing of data noise, and the judgment criteria for different types of traffic abnormal events are not precise enough, making it difficult to meet the requirements of rapid and accurate identification of abnormal events in actual traffic management. Summary of the Invention
[0005] In view of this, the present invention proposes a method for detecting and tracking abnormal events on highways from the perspective of drones to solve the technical problems of insufficient detection accuracy of small targets, poor target tracking stability, poor processing of trajectory data noise, and low accuracy of abnormal event determination in the existing technology from the perspective of drones.
[0006] The technical solution of the present invention is realized as follows: The present invention provides a method for detecting and tracking abnormal events on highways from the perspective of drones, including the following steps:
[0007] S1. Establish an aerial photography historical data set from the perspective of drones;
[0008] S2. Construct an object detection model and an object tracking model. Among them, the object detection model adopts an improved YOLO v11 network structure, and the object tracking model adopts an improved DeepSORT network structure, and train the object detection model and the object tracking model based on the historical data set;
[0009] S3. Obtain real-time aerial photography images of the highway through the drone, input the real-time aerial photography images into the trained object detection model, detect the target vehicles in the images, and obtain the position information and environmental information of the target vehicles;
[0010] S4. Input the detected target vehicle information into the trained object tracking model to obtain the motion trajectory information of the target vehicle;
[0011] S5. Analyze and determine whether the target vehicle has abnormal operating status, abnormal spatial position, or abnormal group behavior based on the position information, environmental information, and motion trajectory information of the target vehicle;
[0012] S6. Output the detection results of the abnormal events.
[0013] Based on the above technical solutions, preferably, the improved YOLO v11 network structure includes:
[0014] Backbone network, adopting the CSPDarknet network structure, where some conventional convolutional layers in the deep C3 module of the backbone network are replaced by deformable convolutional DCNv4, and the deep C3 module corresponds to the generation of feature maps of P3, P4, and P5 layers;
[0015] Neck network, adopting a feature pyramid structure, including P5, P4, P3 feature layers and a P2 layer feature processing unit. The P2 layer feature processing unit is formed by concatenating the P3 layer feature after 2x upsampling with the P2 layer feature generated by the second downsampling of the backbone network in the channel dimension, enabling the network to obtain a local receptive field corresponding to the original Figure 4 ×4 pixels; the CBAM mechanism is embedded in the feature fusion process of the neck network. The CBAM attention mechanism includes channel attention and spatial attention, which are used to enhance the extraction ability of small target-related features;
[0016] Detection head network, including P5, P4, P3 detection heads and a P2 detection head, which are respectively used to detect targets of different sizes. Among them, the CBAM attention mechanism is embedded in front of the P2 detection head to improve the positioning accuracy of small targets;
[0017] Output layer, which is used to output target position coordinates, confidence, and category information.
[0018] Based on the above technical solutions, preferably, the target detection model is trained using the Focal-WIOU Loss function. The Focal-WIOU loss function is composed of the WIoU loss function and the Focal loss function. The calculation formula of the Focal-WIOU loss function is as follows:
[0019] Focal―WIOU Loss = box × L WIoUv3 + cls × L focal
[0020] where,
[0021]
[0022] L focal = ―α t × (1 ― pt ) γ × log(p t )
[0023] Wherein, L WIoUv3 is the bounding box regression loss, x and y respectively represent the horizontal and vertical coordinates of the center of the predicted box, x gt , y gt respectively represent the horizontal and vertical coordinates of the center of the ground truth box, r represents the dynamic non-monotonic focusing coefficient, L IoU represents the basic intersection over union loss, L focal is the classification loss, W g and H g are the width and height of the smallest bounding box, the superscript * represents separation from the computational graph, p t is the predicted probability of the model for the true class, α t is the sample balance factor, γ represents the focusing parameter; box is the weight coefficient of the bounding box regression loss L WIoUv3 , used to adjust the positioning error, with a value of 0.05; cls is the weight coefficient of the classification loss L focal , used to adjust the classification loss, with a value of 0.5.
[0024] Based on the above technical solutions, preferably, the object detection model uses the soft non-maximum suppression algorithm for object box filtering. The soft non-maximum suppression algorithm realizes the dynamic screening of overlapping object boxes by applying a confidence decay mechanism based on the Gaussian function to the candidate boxes with higher IoU values. The formula of the soft non-maximum suppression algorithm is as follows:
[0025]
[0026] Wherein, score i represents the initial confidence of object box i, IOU(i, j) represents the intersection over union between object box i and object box j, σ represents the parameter controlling the decay rate, and Score′ i is the corrected score after IOU penalty.
[0027] Based on the above technical solutions, preferably, the improved DeepSORT network structure is based on the ResNet-18 architecture, including an input layer, a convolutional layer, a pooling layer, a residual layer, and a fully connected layer. The residual layer includes a residual module and an SE attention module. The SE attention module is embedded after each residual module. The SE attention module extracts channel information through global average pooling, generates weights through linear transformation, and adjusts the feature map by weighting.
[0028] Based on the above technical solutions, preferably, coordinate transformation and trajectory processing are further included between step S3 and step S4.
[0029] Coordinate transformation includes the transformation from the image coordinate system to the camera coordinate system and the transformation from the camera coordinate system to the ground coordinate system; trajectory processing includes identifying and correcting outliers in the trajectory data obtained from detection and tracking, processing the trajectory data using the wavelet analysis method, decomposing the trajectory data through wavelet transform to identify abnormal fluctuations; using a wavelet filter for noise reduction processing, and implementing noise reduction using soft thresholding and hard thresholding methods.
[0030] Based on the above technical solutions, preferably, the position information includes the coordinate position, width, and height of the target vehicle in the image; the environmental information includes lane line position information and lane type information; the motion trajectory information includes the real-time speed, driving direction, and vehicle spacing of the target vehicle, where the pixel position in the image coordinate system is converted into the actual position in the ground coordinate system through coordinate transformation.
[0031] Based on the above technical solutions, preferably, the determination of the abnormal operation state events includes:
[0032] Parking event: When the real-time speed of the vehicle is less than the preset value and the duration exceeds the set threshold, it is determined that a parking event has occurred;
[0033] Speeding event: When the real-time speed of the vehicle exceeds the specified maximum speed limit, it is determined that a speeding event has occurred;
[0034] Low-speed driving event: When the real-time speed of the vehicle is lower than the specified minimum speed limit and the duration exceeds the set threshold, it is determined that a low-speed driving event has occurred.
[0035] Based on the above technical solutions, preferably, the determination of the abnormal spatial position events includes:
[0036] Illegal use of the emergency lane event: When the center coordinate of the vehicle crosses the road edge line and enters the emergency lane area, it is determined that an illegal use of the emergency lane event has occurred;
[0037] Pedestrian getting off event: When a pedestrian is detected on the highway, it is determined that a pedestrian getting off event has occurred;
[0038] Reverse driving event: When the direction of the vehicle is different from the direction of the lane where the vehicle appears or when the angle between the driving direction of the vehicle and the specified driving direction of its lane is greater than 90 degrees, it is determined that a reverse driving event has occurred;
[0039] Lane crossing event: When the vehicle crosses or runs over the solid line, or continuously deviates from the lane center line by more than the preset threshold, it is determined that a lane crossing event has occurred.
[0040] Based on the above technical solutions, preferably, the determination of the abnormal group behavior events includes:
[0041] Following risk event: By calculating the relationship between the headway distance between adjacent vehicles in the same lane and the speed of the following vehicle, when the headway distance is less than the safe following distance, it is determined that a following risk event has occurred;
[0042] Traffic congestion event: by calculating the traffic density and average vehicle speed on the lane, when the vehicle density is greater than the preset threshold of vehicle density and the average vehicle speed is lower than the preset threshold of vehicle speed, it is determined that a traffic congestion event has occurred;
[0043] Traffic flow overload event: By setting a virtual detection line on the lane cross section, the vehicle flow crossing the detection line per unit time is calculated. When the vehicle flow exceeds the lane capacity, it is determined that a traffic flow overload event has occurred.
[0044] The method for detecting and tracking abnormal events on highways from the perspective of a drone of the present invention has the following beneficial effects compared with the prior art:
[0045] (1) An improved YOLO v11 algorithm is proposed for small target detection in drone aerial photography. The ability to capture small target features is improved by adding a shallow detection head, adding an attention mechanism, and optimizing the loss function. Based on VisDrone2019 and self-built datasets for training, the improved model has a mAP@0.5 of 72.1%, an increase of 7.6% over the baseline model, and the detection frame rate remains at 48FPS;
[0046] (2) Design a lightweight DeepSORT improvement solution for dynamic tracking scenarios. Reconstruct the feature extraction network based on ResNet-18 and integrate the channel attention mechanism to enhance target identification capabilities. In the VisDrone2019-MOT dataset test, the comprehensive accuracy index of multi-target tracking was improved to 60.3%, and the ground coordinate system mapping of the target trajectory was achieved simultaneously.
[0047] (3) Establish a highway abnormal event detection system in combination with traffic regulations. By analyzing vehicle kinematic parameters, abnormal events are divided into three categories: operating status, spatial position, and group behavior. Trajectory feature lines and lane-based reasoning algorithms are designed to achieve 10 types of event recognition, including speeding, wrong-way driving, illegal parking, and crossing the line. The practicality and reliability of the system are verified by using drone aerial video. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 Flowchart of the detection and tracking method for highway abnormal events from the perspective of drones;
[0050] Figure 2 Annotation map of the VisDrone2019 dataset;
[0051] Figure 3 Graph of data augmentation processing for VisDrone2019;
[0052] Figure 4 Structural diagram of the improved YOLO v11 algorithm model;
[0053] Figure 5 Comparison graph of the YOLO v11 algorithm model before and after improvement;
[0054] Figure 6 Structural diagram of the attention mechanism module;
[0055] Figure 7 Detection effect diagrams of ordinary convolution and deformable convolution;
[0056] Figure 8 Curve graph of the change of the loss function during the training process of the object detection model;
[0057] Figure 9 PR curve graph of the object detection model;
[0058] Figure 10 Visual comparison graph of the detection effects of the object detection model before and after improvement;
[0059] Figure 11 Schematic diagram of the SE Block attention mechanism structure;
[0060] Figure 12 Curve graph of the change of the loss function during the training process of the object tracking model;
[0061] Figure 13 Flowchart of the detection of abnormal events in the running state;
[0062] Figure 14 Schematic diagram of lane lines marked on the pixel coordinate axis;
[0063] Figure 15 Schematic diagram of the lane lines and lane positions;
[0064] Figure 16 Schematic diagram of the detection effect of marked lane lines;
[0065] Figure 17 Schematic diagram of the vehicle speed detection effect;
[0066] Figure 18Schematic diagram of the detection effect of speeding events;
[0067] Figure 19 Schematic diagram of the detection effect of illegal use of the emergency lane event;
[0068] Figure 20 Schematic diagram of the detection effect of the event of a person getting out of the vehicle;
[0069] Figure 21 Schematic diagram of the detection effect of the reverse driving event;
[0070] Figure 22 Schematic diagram of the detection effect of the event of crossing the line while driving;
[0071] Figure 23 Schematic diagram of the detection effect of the following - vehicle risk event;
[0072] Figure 24 Schematic diagram of the detection effect of traffic congestion events;
[0073] Figure 25 Schematic diagram of the detection effect of the flow overload event. Specific implementation manners
[0074] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0075] As Figure 1 shown, the present invention provides a method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle, including the following steps:
[0076] S1. Establish an aerial photography historical data set from the perspective of an unmanned aerial vehicle;
[0077] S2. Construct a target detection model and a target tracking model. Among them, the target detection model adopts an improved YOLO (You Only Look Once) v11 network structure, and the target tracking model adopts an improved DeepSORT (Deep Simple Online and Realtime Tracking) network structure, and train the target detection model and the target tracking model based on the historical data set;
[0078] S3. Obtain real-time aerial images of the highway through a drone, input the real-time aerial images into the trained object detection model, detect the target vehicles in the images, and obtain the position information and environmental information of the target vehicles;
[0079] S4. Input the detected target vehicle information into the trained target tracking model to obtain the motion trajectory information of the target vehicle;
[0080] S5. According to the position information, environmental information and motion trajectory information of the target vehicle, analyze and judge whether the target vehicle has abnormal operating status, abnormal spatial position or abnormal group behavior;
[0081] S6. Output the detection results of the abnormal events.
[0082] Through the organic combination of deep learning and computer vision technology, the present invention constructs a complete method for detecting and tracking abnormal events on the highway from the perspective of a drone, greatly improving the accuracy, real-time performance and adaptability of traffic monitoring, and providing a new technical path for the development of intelligent transportation systems. First, through the improved YOLO v11 network structure, the present invention introduces DCN (Deformable Convolutional Network) v4 in the network backbone to replace some conventional convolutional layers in the deep C3 module, enhancing the feature extraction ability for deformed targets; at the same time, adding a P2 layer feature processing unit and embedding a CBAM (Convolutional Block Attention Module) attention mechanism in the neck network, enabling the network to obtain a local receptive field corresponding to the original Figure 4 ×4 pixels, significantly improving the detection accuracy of small-sized target vehicles from the aerial perspective of the drone, and reducing the missed detection rate and false detection rate. Secondly, adopting the improved DeepSORT network structure, by embedding an SE (Squeeze-and-Excitation) attention module after the residual module of the ResNet-18 (Residual Network Type 18) architecture, the feature extraction and association ability are enhanced, improving the target tracking stability and ID switching accuracy in complex highway scenarios, and effectively solving the tracking problems under vehicle occlusion, dense scenarios and fast motion states. Finally, the present invention establishes a comprehensive abnormal event judgment mechanism. By analyzing the position information, environmental information and motion trajectory information of the target vehicle, it can accurately identify various types of highway abnormal events, including abnormal operating status, abnormal spatial position and abnormal group behavior, providing timely and effective warning information for traffic management departments.
[0083] In step S1, an aerial photography historical dataset from the perspective of a drone is established. The historical dataset is a drone perspective aerial photography dataset constructed by combining the VisDrone2019 dataset and the self-acquired dataset. After constructing the historical dataset, it also includes preprocessing of the historical data. The preprocessing includes converting the aerial photography data into video format, cropping the images, and performing frame extraction and annotation.
[0084] In one embodiment, the method of using a public dataset + self-collection is adopted. The aim is to obtain diverse data from the public dataset and at the same time self-obtain a sufficient number of pictures related to the research background. A total of 10,548 pictures are obtained, and 1,452 pictures are obtained through data augmentation. Finally, a dataset of 12,000 pictures and their annotation files is constructed, which is divided into 7,600 training sets, 650 validation sets, and 3,750 test sets.
[0085] S11. The public dataset uses VisDrone2019, with a total of 10,209 pictures. VisDrone2019 is a large-scale benchmark dataset specifically constructed for drone vision tasks. This dataset is captured by various drone cameras, has a wide coverage range, and covers 10 labels such as pedestrians, cars, and bicycles. It is especially suitable for researching object detection from the perspective of a drone. Although the performance metrics may be low due to problems such as small targets, target overlap, and complex backgrounds, it still has high research value. Since the present invention is to detect abnormal events of vehicles on highways and has little connection with some labels in VisDrone2019, the VisDrone2019 dataset is sorted out, and label data with less connection such as people in non-walking states (person), bicycles (Bicycle), motorcycles (Motor), etc. are deleted. The pedestrian (pedestrian) category is retained, the labels of other models of motor vehicles are unified and named as car, and pictures with less than 5 targets are deleted. Finally, 8,629 pictures are sorted out, as Figure 2 shown, Figure 2 shows the effect diagram after annotation of some images in the VisDrone2019 dataset.
[0086] S12. The self-collected data is obtained by aerial photography video using a DJI Air 2S drone. The DJI Air 2S is a high-performance drone equipped with a 1-inch CMOS sensor, supports 5.4K video recording, and has excellent image capture capabilities. It is also equipped with six visual obstacle avoidance sensors to provide four-way environmental perception and enhance flight safety. In addition, the image transmission system of the DJI Air 2S supports a maximum transmission distance of 12 km to ensure stability during remote flight. Among them, the basic parameters of the drone are shown in Table 1:
[0087] Table 1 Basic parameter table of DJI Air 2S drone
[0088] Device Parameter Image sensor 1-inch CMOS, 20 million pixels Video resolution 5472×3648, 5.4K@30fps Field of view (FOV) 88° (equivalent focal length 22mm) Lens focal length 22mm (equivalent in 35mm format) Aperture f / 2.8 - f / 11
[0089] In this embodiment, the shooting site of the UAV aerial photography image dataset obtained independently is selected as the highway and urban expressway in Hongshan District, Wuhan City, Hubei Province. The shooting time is mainly from 8:00 to 9:00 in the morning and from 6:00 to 7:00 in the afternoon. To obtain a better shooting view and avoid the influence of buildings, the flight altitude of the UAV is maintained within the range of 80 meters to 150 meters. The UAV is operated through a remote controller. After reaching the designated location and altitude, the UAV hovers and continuously records the road. To ensure the diversity of data, multiple angles and postures are adopted for shooting. Finally, a total of 60 minutes of original data with a size of about 64G and a frame rate of 30fps are obtained.
[0090] S13. After obtaining the original video data, preprocess the data.
[0091] S131. Convert the video format and crop the image. Since the original resolution is large and the image contains a large number of targetless areas outside the road, the image is uniformly cropped into an image with a resolution of 1360×765 (the same as the VisDrone2019 image size).
[0092] S132. Perform frame extraction on the image. The extracted pictures are screened to remove pictures with fewer targets, severe jitter, and duplicates. Finally, 1919 high-quality pictures are obtained.
[0093] S133. Label the above pictures. The pictures are labeled using the Label Img software and manually labeled according to the same format and labels as the processed VisDrone2019 to obtain 1919 corresponding txt format annotation files. The annotation principles are as follows: Try not to miss the details of small targets; for targets blocked by trees, buildings, or light, if the blocked area is less than half, continue to label, otherwise give up labeling; for blurred targets, if the category can be judged by the human eye, label it; if it cannot be recognized, give up labeling.
[0094] S134. Combine the self - collected pictures and annotation files with the processed VisDrone2019 dataset. A total of 10,548 pictures and annotation files are obtained, and data augmentation is performed on the aggregated data. The operations for data augmentation mainly include rotation, flipping, brightness adjustment, and motion blur processing, etc. Randomly select some pictures for the above operations, and finally a total of 1,452 new pictures are obtained. In addition, YOLO v11 also performs Mosaic (puzzle - type) data augmentation and Dropblock operations on all input pictures at the input end. Mosaic data augmentation randomly stitches 4 pictures together in a way of random cropping, scaling, and combination to form a new picture, which not only enriches the dataset but also generates many small targets, improving the model's detection ability for small objects. The role of Dropblock is to discard an entire random area of the picture (covered by a gray area), which simplifies the network and alleviates overfitting. The pictures after data augmentation are as shown in Figure 3 as follows.
[0095] After completing all the above steps, a total of 12,000 drone image pictures and their annotation files are obtained. After dividing all the data, a total of 7,600 training sets, 650 validation sets, and 3,750 test sets are obtained. Thus, the construction of the drone aerial image dataset is completed. Among them, the various categories and the number of their ground - truth boxes in the training set are shown in Table 2.
[0096] Table 2 Categories and the number of their ground - truth boxes in the training set
[0097] Label Quantity pedestrian 13969 car 17040
[0098] S2. Construct an object detection model and an object tracking model. Among them, the object detection model uses an improved YOLO v11 network structure, and the object tracking model uses an improved DeepSORT network structure, and train the object detection model and the object tracking model based on the historical dataset; the improved DeepSORT network structure uses a newly added shallow detection head and an attention mechanism to improve the detection effect of small targets, uses data augmentation and deformable convolution to deal with perspective changes and motion blur, and improves the loss function and the NMS (Non - maximum Suppression) algorithm to improve the detection effect of dense targets; the improved DeepSORT network structure is based on the ResNet - 18 network and is improved on it.
[0099] Furthermore, the improved YOLO v11 network structure includes:
[0100] The backbone network adopts the CSPDarknet network structure. Some of the conventional convolutional layers in the deep C3 modules of the backbone network are replaced by deformable convolutional DCNv4. The deep C3 modules correspond to the generation of feature maps of P3, P4, and P5 layers.
[0101] The neck network adopts a feature pyramid structure, including P5, P4, P3 feature layers and a P2 layer feature processing unit. The P2 layer feature processing unit is formed by concatenating the P3 layer feature after 2x upsampling with the P2 layer feature generated by the second downsampling of the backbone network in the channel dimension, enabling the network to obtain a local receptive field corresponding to the original Figure 4 ×4 pixels. The CBAM attention mechanism is embedded in the feature fusion process of the neck network. The CBAM attention mechanism includes channel attention and spatial attention, which are used to enhance the extraction ability of small target-related features.
[0102] The detection head network includes P5, P4, P3 detection heads and a P2 detection head, which are respectively used to detect targets of different sizes. Among them, the CBAM attention mechanism is embedded before the P2 detection head to improve the positioning accuracy of small targets.
[0103] The output layer is used to output the target position coordinates, confidence, and class information.
[0104] Figure 4 The improved YOLO v11 network structure diagram is shown. In the figure, new structures such as C3k2 and C2PSA are mainly introduced, which improve the feature extraction and processing capabilities. At the same time, the Head idea of YOLO 10 is applied to the Head of YOLO v11, and the depthwise separable method is used to reduce redundant calculations and improve efficiency.
[0105] Specifically, in object detection, taking the COCO dataset as the standard, small targets usually refer to objects with a size smaller than 32×32 pixels. From the perspective of an unmanned aerial vehicle (UAV), due to the influence of flight altitude and viewing angle, most vehicles will appear as small targets and dense targets, with fewer pixels in the image and are prone to being missed. And due to the shaking during the hovering of the UAV, problems such as viewing angle change and motion blur are likely to occur. The improvement of the YOLO v11 algorithm aims at the above problems and mainly starts from three aspects:
[0106] (1) Adding a shallow detection head and an attention mechanism to improve the small target detection effect.
[0107] To better achieve the detection of small targets, the present invention introduces a shallow detection head strengthening mechanism in the YOLO v11 architecture and adds a detection head P2. The purpose is to enable the backbone network to additionally generate a feature map with 4x (160×160) downsampling, corresponding to the original Figure 4 ×4 pixels of local receptive field. The structure of adding a shallow detection head is asFigure 5 As shown, where Figure 5 The improved comparison of the YOLO v11 network structure is shown, mainly focusing on the optimization of the network architecture to enhance the small object detection ability. The original network structure is on the left, which includes a basic Input input layer, multiple Conv convolutional layers, C1 / C2 / C3 feature extraction modules, and a conventional feature pyramid structure, and finally connects to the detection output layer. This structure has problems with insufficient detection accuracy when dealing with small target vehicles from the top-down perspective of drones. The improved network structure is on the right. The most significant change is the addition of a P2 layer feature processing unit in the part marked by the red box. Its improvements include: retaining the basic structure of CSPDarknet in the backbone network, including the input layer and the initial convolutional layer sequence; introducing deformable convolution DCNv4 in the C3 module (marked in pink) to replace some conventional convolutional layers, enhancing the feature extraction ability for deformed targets; the newly added P2 layer feature processing unit (marked by the red box) is constructed as follows: (1) Design the P2 layer after the second downsampling of the backbone network. After upsampling the P3 layer features by a factor of 2 once and concatenating them with the P2 layer features in the channel dimension (Concat), the final features are used as the output of the P2 detection head; (2) After downsampling the concatenated P2 layer once again, continue to use the Concat operation with the P3 layer to fuse the features again. Here, a four-level detection system is successfully constructed. The newly added P2 detection head in step (1) can obtain a local receptive field of Figure 4 ×4 pixels, significantly improving the model's detection ability for small targets and greatly enhancing the model's performance and accuracy at the cost of increasing the model size slightly. In step (2), the P2 detection head is combined with the original three detection heads to ensure the symmetry and feasibility of the network structure. A special Small Detect detection head is connected after the P2 layer for small target detection.
[0108] In the detection of small targets from the drone perspective, one of the key problems is the missed detection caused by complex background interference. Therefore, the present invention applies the CBAM attention mechanism to the YOLO 11 network to improve its detection effect on small targets. CBAM is an efficient and lightweight attention mechanism that can precisely regulate the focus of attention in the spatial and channel dimensions. As Figure 6 shown, Figure 6 The structure diagram of the attention mechanism module is shown. It can be seen from the figure that the channel attention and spatial attention enable the network to adaptively adjust the weights of the key channels and regions in the feature map, thereby enhancing the model's attention to important features.
[0109] For the objectives of the present invention, consider adding CBAM at the following positions: (1) After the P2 layer. The shallow feature map contains rich detailed information but is vulnerable to complex backgrounds. Adding CBAM at this position enhances the weights of the channels related to small targets through channel attention and focuses on the target area using spatial attention, significantly improving the feature discrimination of targets from 4×4 to 16×16 pixels; (2) After the Concat operation. Feature fusion (such as concatenating P2 and upsampled P3) integrates semantic information of different scales but may have feature conflicts. After introducing CBAM, it can adaptively screen important channels and spatial regions, reduce redundant information, and make the fused features more suitable for small target detection; (3) Before the P2 detection head. The detection head is responsible for bounding box regression and classification, and the quality of the input features directly affects the detection accuracy. Adding CBAM before prediction can dynamically adjust the feature response and particularly improve the localization accuracy of small targets. Generally speaking, after adding the CBAM attention mechanism, the YOLO v11 network will perform better in small target detection. By focusing on the key features of the target and reducing the influence of redundant information, it significantly improves the localization accuracy and recognition effect of the target.
[0110] (2) Use data augmentation and deformable convolutional network (DCN) to address perspective changes and motion blur.
[0111] In the present invention, the two main objectives of data augmentation are to address perspective transformation and motion blur. Perspective transformation uses methods such as rotation, flipping, and brightness adjustment to simulate perspectives at different heights, angles, and illuminations, enabling the YOLO v11 network to effectively perform target detection under diverse UAV perspectives. At the same time, motion blur simulates the image blur caused by the shaking during the UAV detection process, enhancing the model's ability to recognize blurred images. By applying these transformations to the model training data, YOLO v11 can maintain better detection accuracy when facing different perspectives and blurred images, especially showing more prominent performance in UAV perspective detection.
[0112] DCN is an improvement over traditional convolutional operations. It adjusts the position of the convolutional kernel by introducing learnable offsets, enabling the convolutional kernel to adaptively accommodate geometric changes in the image, such as Figure 7 shown, Figure 7The left figure in the middle is the effect diagram of ordinary convolution, and the right figure is the effect diagram of DCN. It can be seen that by dynamically adjusting the convolution kernel, the shape and size of deformed objects can be extracted more accurately, improving the detection effect of deformed targets. In the present invention, DCNv4 is introduced into the convolution layer in the C3k2 structure of the YOLO v11 backbone network, replacing the standard convolution operation with C3k2_DCNv4, which is applied to the deep C3 module in the backbone network corresponding to the feature generation of the P3, P4, and P5 layers, enhancing the ability of deep features to model deformed targets. By adding the learned offset during convolution, the convolution kernel can be flexibly adjusted to adapt to the local changes of the target, improving the accuracy of feature extraction, especially when there is vehicle occlusion or perspective tilt change. In addition, DCN helps to alleviate the image blurring problem caused by motion blur. Combining with the data augmentation strategy, it can more effectively improve the localization and classification accuracy of small targets.
[0113] (3) Improve the detection effect of dense targets by improving the loss function and NMS algorithm.
[0114] The classification loss calculation of YOLO v11 still uses Cross-Entropy Loss (Cross-Entropy Loss, cross-entropy loss function), and the regression branch uses DFL_loss (Distribution Focal Loss, distribution loss) and CIOU Loss (Complete-IoU Loss, complete intersection over union loss). However, when the class is unbalanced, the background samples dominate in the cross-entropy loss, resulting in the model ignoring small targets or minority classes; when calculating CIOU Loss, due to the vague definition of the aspect ratio, regression samples with poor quality will have a greater impact on the loss, making it difficult to optimize samples with good quality. Both face the problem of imbalance between positive and negative samples, affecting the detection accuracy.
[0115] Therefore, aiming at the problems of dense targets and imbalance between positive and negative samples in the drone perspective, the present invention adopts Focal-WIOU Loss (focal intersection over union loss) combining Focal Loss + WIoU Loss (focal weighted intersection over union loss) to help optimize localization and classification. Among them, WIoU (WeightedIoU Loss, weighted intersection over union loss) constructs a bounding box loss based on attention, adopts scale normalization, introduces the target size weight, reduces the penalty coefficient of small target IoU, and adopts the center point priority strategy to increase the penalty term of the center point distance, improving the localization stability.
[0116] The calculation method of LWIoUv3 is as follows:
[0117]
[0118] Among them, x and y respectively represent the horizontal and vertical coordinates of the center of the predicted bounding box, and x gt , y gt respectively represent the horizontal and vertical coordinates of the center of the ground truth bounding box, L WIoUv3 is the bounding box regression loss, and L WIoUv1 represents the weighted base intersection over union loss, R WIoU represents the weight of the distance weight force, L IoU represents the base intersection over union loss, r is the dynamic non-monotonic focusing coefficient, Wg and Hg are the width and height of the smallest bounding box, and the superscript * represents separation from the computational graph.
[0119] In a dense scene, the number of negative samples is much larger than that of positive samples. Focal Loss is a loss function specifically designed to handle the class imbalance problem. By modulating the factor, it reduces the loss contribution of easy-to-classify samples and focuses on difficult samples, thus solving the problem of severe imbalance between positive and negative samples. The calculation method of Focal Loss is as follows:
[0120] L focal = -α t ×(1 - p t ) γ ×log(p t ) (2)
[0121] Among them, p t is the predicted probability of the model for the true class, and α t is the sample balance factor, and γ = 2.
[0122] Combining Equation (1) and Equation (2), the Focal-WIOU Loss is finally obtained, and the calculation formula is as follows:
[0123] Focal - WIOU Loss = box × L WIoUv3 + cls × L focal (3)
[0124] Among them, box = 0.05 is the weight of L WIoUv3 for adjusting the positioning error; cls = 0.5 is the weight of L focal for adjusting the classification loss. Through this combination, Focal-WIOU Loss can effectively improve the detection accuracy of the YOLO v11 model in complex environments, especially for small targets and dense scenes.
[0125] NMS is used in YOLO v11 to retain the optimal detection results. Since the targets in UAV images are usually small and densely distributed, the prediction boxes of multiple targets often have a high overlap (e.g., IoU > 0.7). To solve the problem of missed detections of traditional NMS in dense scenes, this study uses Soft-NMS to replace NMS in YOLO v11. In Soft-NMS, when the IoU of two boxes is high, instead of completely eliminating one of the boxes, the confidence of the box is attenuated according to the IoU value, so as to retain more effective target boxes. This process is calculated by the following formula:
[0126]
[0127] where score i represents the initial confidence of target box i, IOU(i, j) represents the intersection over union between target box i and target box j, σ represents the parameter controlling the attenuation rate, and Score′ i is the corrected score after IOU penalty. Through this formula, Soft-NMS can effectively reduce the interference of highly overlapping boxes, avoid misdeletion caused by a fixed threshold, and improve the retention rate of target boxes.
[0128] This experiment is based on the improved YOLO v11 network. On the NVIDIA RTX 3070 GPU platform, the improved network structure is trained for 100 epochs (training rounds) using the UAV aerial photography dataset constructed in this chapter, with a total time consumption of 5 hours and 11 minutes. Finally, the model weights of the optimal round and the last round are obtained. Subsequently, the weight file of the optimal round will be used for inference and system construction.
[0129] The training process records the change curves of box loss (regression loss) and cls loss (classification loss) of the training set and the validation set with the number of training rounds, as Figure 8As shown. The loss curve is an important means to evaluate the training effect of the model and can effectively reflect the convergence of the model during training. From the training loss curve in the figure, it can be seen that in the first 20 epochs, the train loss of the model decreased rapidly, while the val loss also showed a relatively obvious downward trend in the early stage. This indicates that in the initial stage of training, the model can quickly learn effective features from the data and start to converge. As the training continues, in the epochs 20 - 60, the slopes of the double loss curves slow down, and the validation set loss also remains at a low level, indicating that the model enters the fine-grained feature optimization stage. During the convergence stage (epochs 60 - 100), the training loss is relatively stable, and no obvious overfitting phenomenon occurs. Generally speaking, this training process conforms to the normal rules, and both the training and validation losses show a good downward trend, verifying the rationality of the data augmentation strategy and optimizer parameters, providing a strong guarantee for subsequent model evaluation.
[0130] To comprehensively evaluate the improved YOLO v11 model, Figure 9 The PR (precision-recall) curves of the experimental results at different IoU thresholds are shown. The PR curve is a comprehensive representation of precision and recall and can evaluate the overall performance of the model at different thresholds.
[0131] In this embodiment, a quantitative comparison is made between the improved YOLO v11 model, the original YOLO v11 model, and the YOLOv8 model. The key performance indicators such as the number of parameters (M), single-threshold evaluation (mAP@0.50, that is, when the intersection over union (IoU) between the predicted box and the ground truth box is ≥ 50%, it is regarded as a correct detection), multi-threshold comprehensive evaluation (mAP@0.50:0.95, that is, the average precision when IoU ranges from 50% to 95% (step size 5%)), and real-time performance indicator (FPS, that is, the number of frames processed per second) are focused on. The comparison results are shown in Table 3.
[0132] Table 3 Comparison of Algorithm Performance
[0133]
[0134] As can be seen from Table 3, the improved YOLO v11 model performs better than the original YOLO v11 model and the YOLOv8 model in multiple key metrics. Specifically, the improved YOLO v11 has increased by 7.6% and 3.1% respectively in mAP@0.50 and mAP@0.50:0.95, and it can also meet the requirements of real-time detection in terms of FPS (frames per second). This indicates that although the parameter quantity of the improved model has increased, its comprehensive performance is still better than that of YOLO v8 and YOLO v11, especially in the detection of small targets and the target recognition in dynamic scenarios. In addition, the improvement in FPS of the improved YOLO v11 also enables it to have strong real-time processing capabilities in actual deployment, making it suitable for efficient and real-time target detection tasks, such as vehicle detection on highways from the perspective of drones.
[0135] To further verify the effectiveness of the improvement of YOLO v11, in this embodiment, the ablation experiment method is used to compare the index changes before and after the improvement item by item, and quantitatively analyze the specific contribution of each improvement to the model performance. The results of the ablation experiment are shown in Table 4.
[0136] Table 4 Results of Ablation Experiment
[0137]
[0138] In Table 4, Group A represents the original YOLO v11 algorithm. After adding a new shallow detection head (P2 layer) in Group B, mAP0.5 has increased from 67.0% to 69.3%, which indicates that the shallow detection head can significantly improve the detection ability of small targets, and both the detection accuracy and recall rate have increased. The experimental results of Group C show that after adopting CBAM, the model has increased by 1.9% and 0.6% respectively in mAP@0.50 and mAP@0.50:0.95. This improvement enhances the model's focus ability on small targets. The experimental results of Group D show that after replacing with deformable convolution DCNv4, the model has increased by 1.1% in mAP@0.50, which indicates that deformable convolution improves the performance of the model in the detection of targets with large pose changes, and the model has stronger robustness in the drone shooting scenario, but the computational overhead increases by 4%. The experimental results of Group E show that Focal-WIOU Loss and Soft-NMS effectively solve the problems of sample imbalance and target overlap. After the improvement, the mAP@0.50 of the model has increased by 0.98%. The improvement makes the model perform better in dense target scenarios and effectively reduces the detection conflict when targets overlap. The analysis of the ablation experiment results shows that the improved YOLO v11 algorithm of the present invention maintains a good balance between accuracy and efficiency, and has more advantages in the actual drone aerial photography scenario compared with before the improvement.
[0139] Figure 10Shows the detection effects of the model before and after improvement in typical complex scenarios: By comparing the detection effects of YOLOv11 and the improved YOLO v11 in the same scenario, it can be clearly seen that the improved YOLO v11 can effectively detect some targets that were missed by the original YOLO v11. The improved model shows stronger robustness and accuracy. For example, Figure 10 , the original YOLO v11 model did not detect the red bus on the far right of the picture, but it was captured by the improved model (car, conf = 0.44). According to statistics, the original model missed 10 vehicles and 11 pedestrians, and the improved model missed 5 vehicles and 9 pedestrians, indicating that the detection effect of this algorithm has been greatly improved in high-difficulty scenarios.
[0140] Through the above analysis, it can be concluded that the improved YOLO v11 algorithm has better detection effects and higher detection accuracy than the original version in multiple scenarios such as small targets, dense distribution, and occlusion in UAV aerial photography, laying a foundation for the subsequent construction of a highway vehicle abnormal event detection system.
[0141] Furthermore, the improved DeepSORT network structure is improved based on the ResNet-18 architecture, specifically including an input layer, a convolutional layer, a pooling layer, a residual layer, and a fully connected layer. The residual layer includes a residual module and an SE attention module. The SE attention module is embedded after each residual module. The SE attention module extracts channel information through global average pooling, generates weights through linear transformation, and weights and adjusts the feature map.
[0142] Specifically, the biggest improvement of DeepSORT compared to SORT is the introduction of the appearance features of the target. The appearance information is extracted through a convolutional neural network, which improves the multi-tracking effect in the case of target occlusion and fast movement. The convolutional neural network structure adopted by DeepSORT. This network structure contains 10 layers, including 2 convolutional layers Conv, 1 pooling layer MaxPool, and 6 residual layers Residual. Finally, a 128-dimensional feature vector is generated through 1 fully connected layer Dense.
[0143] The improved DeepSORT structure is based on the original ResNet-18 and makes the following improvements: (1) The input size is adjusted from 224×224 (for pedestrian detection) of the original ResNet-18 to 80×160, which is more suitable for the aspect ratio of vehicles under the drone's perspective, and the total number of pixels is only 25.5% of 224×224 (50,176), making the model more lightweight; (2) The original 7×7 convolution Conv1 (stride = 2) is changed to 3×3 convolution (stride = 1), which helps to reduce the loss of target details during the early downsampling process. Using a larger convolution kernel and stride will result in stronger spatial downsampling, which is likely to compress the local features of the target at the early stage of the network, having a greater impact on small targets. (3) An SE Block is added after each residual block. The SE Block is an attention mechanism that enhances the representation ability of deep learning models by learning the feature weights between channels. As Figure 11 shown, Figure 11 the schematic diagram of the SE Block attention mechanism is shown in Figure 11 . It can be seen that the SE Block first extracts channel information from the feature map through global average pooling (Squeenze), generates weights through linear transformation (Excitation), and adjusts the feature map by weighting (Scale), thereby enhancing the effective features and improving the model performance; (4) Replace the original classification head. After using the global average pooling (Global Average Pooling, GAP) operation to output a 512×1×1 vector, the 512-dimensional feature map is mapped to a 128-dimensional space through the fully connected layer Dense, replacing the 1000-dimensional classification head of the original network to meet the requirements of the re-identification (ReID) task.
[0144] In this embodiment, the publicly available dataset Visdrone2019-MOT is used to train and test the improved DeepSORT algorithm. Visdrone2019-MOT is a dataset specifically for multi-object detection in the Tianjin University VisDrone dataset, which contains various continuous video frames captured by drones, including a training set (56 video clips), a validation set (7 video clips), and a test set (17 video clips). These files are crucial for data processing, algorithm training, and performance evaluation. The internal structure of the dataset mainly consists of continuous video frame image data and original annotation files, and these files perform different functions and work together to support the research and development of multi-object tracking tasks.
[0145] For the detector: The improved detector only detects pedestrian and car labels. Therefore, the Visdrone2019-MOT dataset is synchronized, the label data with less relevance is deleted, the naming of vehicle labels of other models is unified as car, and the pedestrian label is retained to ensure the intercommunication of detection and tracking data.
[0146] For the tracker: Most multi-object tracking algorithms (such as DeepSORT, FairMOT (Fair Multi-Object Tracking)) use annotation files in the MOTChallenge (Multi-Object Tracking International Evaluation Benchmark) format. Therefore, it is necessary to convert the format of the original annotation file, retain the objects with a confidence level of 1 in the original annotation file, and ignore the low-confidence or invalid annotations; integerize the box coordinates, and at the same time ensure that the converted positions do not exceed the image boundaries. If they exceed, they are truncated to the image boundaries; calculate the visibility = 1 - occlusion degree.
[0147] Result analysis: The loss function curve in the training process of the improved DeepSORT algorithm is as Figure 12 shown, and the curve gradually converges, indicating that the training is completed.
[0148] The tracking performance metrics of the improved DeepSORT algorithm are calculated using the Visdrone2019-MOT test set, and then the performance of the DeepSORT algorithm with the improved feature extraction network is verified through horizontal comparison. Table 5 shows the results of the tracking experiments of the DeepSORT algorithm before and after improvement and the SORT algorithm.
[0149] Table 5 Algorithm Performance Comparison
[0150] Algorithm MT ML IDs MOTA / % MOTP / % Times / ms SORT 142.3 89.6 137 45.9 73.5 18.7 DeepSORT 175.4 68.2 96 56.7 77.2 29.4 Improved DeepSORT 207.6 54.2 71.8 60.3 78.6 26.9
[0151] In the table: MT (mostly tracked, most successfully tracked trajectories), ML (mostly lost, most lost trajectories), IDs (Identity switches, number of identity switches), MOTA (multiple object tracking accuracy, multi-object tracking comprehensive accuracy), MOTP (multiple object tracking precision, multi-object localization accuracy).
[0152] The performance advantages of the improved DeepSORT algorithm proposed in this invention are demonstrated on the VisDrone2019-MOT test set. Compared with the original DeepSORT algorithm, the MOTA of the improved model has increased by 8.61%, and the MOTP has increased by 11.26%, indicating that it has made comprehensive progress in the tracking task. In terms of the identity consistency index, the number of IDs has decreased by 5 times, verifying its more stable tracking ability in complex scenarios. Although the average processing time per frame of the improved algorithm has increased to 32.6 ms (including object detection) due to the introduction of a more complex feature extraction network, it can still meet the real-time detection and tracking tasks from the perspective of drones. Compared with the BoT-SORT (Bidirectional Optimization Tracker) and ByteTrack (Byte Trajectory Tracker) algorithms, the improved algorithm achieves a better balance between tracking accuracy and processing time. Especially in terms of the IDs index, it shows outstanding performance, with a 4.2% improvement compared to BoT-SORT and a 2.8% improvement compared to ByteTrack, attributed to the stronger feature extraction ability of the improved DeepSORT algorithm, which can more effectively distinguish and maintain the identity of targets, thereby reducing ID switches and improving stability and reliability in complex environments.
[0153] As can be seen from Table 5, the improved DeepSORT algorithm proposed in this invention demonstrates good performance advantages on the VisDrone2019-MOT test set. Compared with the original DeepSORT algorithm, the indicators of the improved model have been comprehensively improved: MOTA has increased by 3.6% from 56.7%, and MOTP has increased by +1.4%, verifying its dual optimization in target tracking accuracy and positioning accuracy. In terms of the identity consistency index, the number of IDs has decreased from 96 times to 71.8 times, proving that the dynamic matching strategy and the feature enhancement module have effectively improved the tracking stability in complex scenarios. The average processing time per frame of the improved algorithm has decreased from 32.4 ms of the original version to 26.9 ms, achieving efficiency improvement while maintaining accuracy. The experimental data shows that this algorithm achieves a better balance in the three dimensions of tracking accuracy, identity consistency, and real-time performance, providing an effective solution for the target tracking task from the perspective of drones.
[0154] The above analysis shows that the improved DeepSORT algorithm performs well in tracking from the perspective of drones and has practical functions in the real environment. In complex scenarios such as small targets, dense distribution, and occlusion, the algorithm shows better detection effects and higher accuracy, laying a solid foundation for the construction of subsequent highway vehicle abnormal event detection systems.
[0155] Furthermore, between step S3 and step S4, coordinate transformation and trajectory processing are also included. Coordinate transformation includes the transformation from the image coordinate system to the camera coordinate system and the transformation from the camera coordinate system to the ground coordinate system. Trajectory processing includes identifying and correcting outliers in the trajectory data obtained from detection and tracking, processing the trajectory data using the wavelet analysis method, decomposing the trajectory data through wavelet transform to identify abnormal fluctuations, and using a wavelet filter for noise reduction processing, and implementing noise reduction using soft thresholding and hard thresholding methods. The specific implementation is as follows:
[0156] In this study, the DJI Air 2S drone will be used to hover at a fixed height directly above the road section to be detected, and the camera will be directed vertically downward to take pictures centered on the road. The movement of targets within a large area in the center of the picture will be focused on, and the influence caused by lens distortion at the boundary can be ignored. Converting the pixel position in the image coordinate system to the actual position in the ground coordinate system requires the following key steps: the transformation from the image coordinate system to the camera coordinate system and the transformation from the camera coordinate system to the ground coordinate system.
[0157] To convert to the ground coordinate system, let the position of the drone in the ground coordinate system be (X drone , Y drone , h), the target pixel coordinates be (u, v), and the center point coordinates be (c x , c y ). Then the conversion formula for the ground coordinates (X g , Y g ) is:
[0158]
[0159] In Equation (5), the negative sign of Y g is caused by the reverse direction of the Y c axis of the camera coordinate system and the ground coordinate system. GSD x and GSD y respectively represent the horizontal (x-axis) and vertical (y-axis) components of the ground sampling distance in the image coordinate system. The relevant parameters of the DJI Air 2S drone are calculated as shown in Table 6.
[0160] Table 6 Calculation parameter table
[0161] Device Parameter Effective sensor size 13.2mm×8.8mm Video resolution 5472×3648, 5.4K@30fps Pixel size 2.41μm×2.41μm Equivalent focal length (35mm format) 22mm Actual physical focal length 8.06mm
[0162] As can be seen from Table 6, the horizontal focal length f x = 8.06 mm / 2.41 μm = 3344 pixels. Similarly, f y = 3344 pixels. For example, when the drone hovering height h = 100, assuming the drone is located at (0, 0) in the ground coordinate system and the target pixel coordinates (u, v) = (3000, 2000), and the camera center point (c x,c y ) = (2736, 1824), the ground coordinate calculation formula is:
[0163]
[0164] This coordinate indicates that the target ground point is 7.9 meters due east and 5.26 meters due south of the vertical projection of the UAV. However, since the built-in GPS of the UAV can only display the absolute height, in the actual calculation process, the height h needs to subtract the actual road surface height, taking the shooting site as the standard. Therefore, when the UAV conducts aerial photography to detect abnormal events on the highway, the choice of flight height directly affects the detection effect. If the height is too low, although high-resolution images can be obtained, it is difficult to cover the road; while when the height is too high, the vehicle pixels are too small, which is prone to missed detection. Table 7 shows the relationship between the aerial photography flight height of DJI Air 2S, the GSD value, and the shooting range.
[0165] Table 7 Flight Height, GSD, and Shooting Range
[0166]
[0167] In the subsequent aerial photography detection process, according to the different shooting site conditions, select an appropriate flight height, and then the corresponding GSD value can be used to convert the image coordinate system to the ground coordinate system to obtain the position information of the target on the ground, further study the relationship between them, and provide basic data for constructing an abnormal event detection system.
[0168] After using the improved YOLO v11 algorithm and the improved DeepSORT algorithm proposed in this study to detect and track the UAV video, the results are saved in the 6-dimensional vector format of frame, id, x, y, w, h. frame represents the frame number, id represents the target number, x and y respectively represent the horizontal and vertical coordinates of the upper left corner of the target detection frame, and w and h respectively represent the width and height of the target detection frame. Run a video file. Taking the target with id = 1 as an example, Table 8 shows the position information of this target within frames 1 to 10. Connecting the center point coordinates of its detection frames into a line can visualize its driving trajectory.
[0169] Table 8 Target Trajectory Information
[0170] frame Id x y w h 1 1 1884.5548 425.2365 70.6926 67.8073 2 1 1878.9812 424.0641 67.5428 64.4713 3 1 1871.6970 423.6837 64.5352 60.5753 4 1 1865.6394 423.3439 63.6936 58.1304 5 1 1857.8315 423.2059 66.6500 58.3569 6 1 1850.4008 423.6301 70.2061 58.1272 7 1 1842.5354 423.6396 73.2352 56.4603 8 1 1831.0759 423.9340 78.2699 56.2538 9 1 1818.3120 423.6325 83.4531 56.3243 10 1 1804.5640 423.2957 88.3752 56.0794
[0171] However, due to the possible irregular jitter during the hovering of the UAV at high altitude, especially affected by factors such as wind speed and environmental changes, the sensor often produces slight displacement and vibration, resulting in short-term deviation of the target position, thus affecting the smoothness and continuity of the trajectory. In addition, the positioning error of the target detection algorithm will also have a negative impact on the trajectory data. These errors will lead to fluctuations and outliers in the trajectory, affecting the accuracy of subsequent traffic analysis and event detection. Therefore, before subsequent event judgment, it is necessary to clean the abnormal trajectory data and smooth the trajectory to improve the quality and reliability of the data.
[0172] Outlier identification and correction is a crucial step in trajectory data cleaning. Outliers in trajectory data are usually caused by sensor errors, frame loss, or other unforeseen factors, manifested as sudden changes in data points, resulting in abnormal fluctuations in the trajectory. These outliers not only affect the smoothness of the trajectory but may also mislead subsequent analysis. In this study, the wavelet analysis method is used to identify and process outliers in trajectory data. Wavelet analysis is a powerful signal processing tool with good time-frequency localization characteristics, capable of decomposing signals at multiple scales, and then identifying abnormal fluctuations. By performing wavelet transform on trajectory data, normal trajectories and abnormal fluctuations can be effectively separated. The basic formula of wavelet transform is as follows:
[0173]
[0174] where ψ(t) is the mother wavelet, s is the scale factor, u is the translation factor, and t′ represents the continuous independent variable in the time domain. By appropriately selecting the scale and position, abnormal fluctuations in the trajectory can be extracted and outliers can be removed through reconstruction, thereby improving the quality of trajectory data.
[0175] After outlier correction, the trajectory still contains Gaussian noise and quantization errors, and further smoothing is required. Affected by the video frame rate and environmental factors (such as brightness changes, motion blur, etc.), trajectory data often contains random noise, resulting in jitter and instability of the trajectory. To reduce the impact caused by these noises, a wavelet filter is used for noise reduction. The wavelet filter suppresses high-frequency noise in the signal and retains low-frequency information in the trajectory by selecting an appropriate threshold, thereby achieving effective noise reduction. The noise reduction process is usually implemented using the soft threshold method and the hard threshold method. The soft threshold method is shown in Equation (8), and the hard threshold method is shown in Equation (9).
[0176]
[0177] where λ is the threshold, and X(t) is the original signal. is the filtered signal, and sgn(X(t)) is the sign function, representing the direction of the signal. The hard threshold method preserves the invariance of the signal, while the soft threshold method is smoother. Through the wavelet filter, the noise in the trajectory is effectively suppressed while important motion information is retained.
[0178] After cleaning the outliers through the above wavelet analysis and noise reduction through the wavelet filter, the trajectory data becomes smoother and more stable after cleaning, providing higher-quality data support for subsequent target recognition, traffic parameter calculation, and abnormal event detection.
[0179] Further, combining the key feature parameters of highway vehicle abnormal events, vehicle abnormal events can be classified into the following three categories: abnormal operating state, abnormal spatial position, and abnormal group behavior.
[0180] Obtain real-time aerial images of the highway through a drone, input the real-time aerial images into the trained object detection model to detect the target vehicles in the images, obtain the position information and environmental information of the target vehicles, input the detected target vehicle information into the trained target tracking model to obtain the motion trajectory information of the target vehicles, analyze and judge whether the target vehicles have abnormal operating states, abnormal spatial positions, or abnormal group behaviors based on the position information, environmental information, and motion trajectory information of the target vehicles, and finally output the detection results of abnormal events according to the analysis results. Among them, the position information includes the coordinate position, width, and height of the target vehicle in the image; the environmental information includes lane line position information and lane type information; the motion trajectory information includes the real-time speed, driving direction, and vehicle spacing of the target vehicle, and the pixel position in the image coordinate system is converted into the actual position in the ground coordinate system through coordinate transformation.
[0181] The judgment of abnormal operating state events includes:
[0182] Stopping event: When the real-time speed of the vehicle is less than the preset value and the duration exceeds the set threshold, it is determined that a stopping event has occurred; speeding event: When the real-time speed of the vehicle exceeds the specified maximum speed limit, it is determined that a speeding event has occurred; low-speed driving event: When the real-time speed of the vehicle is lower than the specified minimum speed limit and the duration exceeds the set threshold, it is determined that a low-speed driving event has occurred. As Figure 13 shown, the specific implementation method is as follows:
[0183] (1) Obtain the detection and tracking results of the aerial images from the perspective of the UAV. For each frame of the UAV aerial images, the improved YOLO v11 and improved DeepSORT algorithms constructed above are used for detection and tracking. For each frame of the image, the information of all the targets in the image is extracted to obtain the 6D vector of the targets, including the frame number, the ID of each target, and the position information (x, y, w, h) of the detection box in each frame. Further calculate that the central pixel coordinates of vehicle i at time t are
[0184] (2) Calculate the vehicle speed every 30 frames. According to the instantaneous speed of vehicle with ID i between the τ-th frame and the τ - Q-th frame It can be estimated by the average speed between two frames, as shown in the formula (where the numerator represents the displacement between the τ-th frame and the τ - Q-th frame, and the denominator represents the length of the time interval). Calculate the instantaneous speed of all target vehicles every 30 frames Since the actual flight altitude of the UAV is known, according to the above coordinate conversion method, the pixel value is converted into the real ground distance, so as to obtain the real speed of vehicle i at time t
[0185] (3) Judge the events of parking, speeding, and low-speed driving on the highway according to the characteristic parameters and laws and regulations. When the above abnormal events are detected, the system will issue corresponding warnings. The judgment methods are as follows:
[0186] Parking event: If the real speed of the vehicle is less than a certain set value (such as 5 km / h), start timing for it. If the duration times exceeds the set value (such as 3 seconds), it is considered that a parking event has occurred. When the number of simultaneous parking events is greater than 1, a more serious vehicle collision accident may have occurred. The judgment method for the parking event is:
[0187]
[0188] Speeding and low-speed driving events: If the real speed of the vehicle exceeds the specified maximum speed (such as 120 km / h), it is considered that a speeding event has occurred; when it is lower than the specified minimum speed (such as 60 km / h) and the duration times exceeds the set value (such as 3 seconds), it is considered that a low-speed driving event has occurred. The judgment methods for speeding and low-speed driving events are:
[0189]
[0190] The judgment of spatial position abnormal events includes:
[0191] Emergency lane violation event: When the vehicle center coordinate crosses the road edge line and enters the emergency lane area, it is determined that an emergency lane violation event has occurred; Pedestrian getting off event: When a pedestrian is detected on the highway, it is determined that a pedestrian getting off event has occurred; Wrong-way driving event: When the vehicle direction is different from the lane direction where the vehicle appears or when the included angle between the vehicle driving direction and the specified driving direction of its lane is greater than 90 degrees, it is determined that a wrong-way driving event has occurred; Lane violation event: When the vehicle crosses or runs over the solid line, or continuously deviates from the lane center line by more than a preset threshold, it is determined that a lane violation event has occurred. The specific implementation method is as follows:
[0192] (1) Calibrate the lane lines: At the start of the UAV shooting, take the first frame of the incoming image and manually calibrate the pixel start and end points of the lane lines, as Figure 14 shown.
[0193] First step, for the straight road section to be detected Figure 14 (a) As shown, sequentially mark the start point 1 and end point 2 of the carriageway edge line (the start point 3 and end point 4 for the oncoming lane). At this time, connect the two points into a straight line and use the least squares method to obtain the expression of the edge line r 外 r 外 : y r外 = k r外 x + b r外 , (y r外 represents the y coordinate of a certain point on the straight line, x represents the x coordinate of a certain point on the straight line, k r外 is the slope of the straight line, b r外 is the intercept of the straight line), and the lane direction θ r , that is, the included angle between the straight line and the x-axis is θ r = atan2(k 外 ), calculate the angle corresponding to the slope k 外 ; If it is a circular curve section Figure 14 (b) As shown, then sequentially mark its start point 1, any point 2 on the line, and end point 3 (the oncoming lane is 4→5→6). At this time, connect the three points to obtain the expression of the edge line r 外 r 外 : (x, y) represents the coordinates of any point on the edge line, (x rj , y rj ) represents the coordinates of the center of the circle, R ij represents the radius of the arc, and its lane driving direction changes with a certain point (x, y) on the curve. The angle (x r外 , y 外 ) represents the coordinates of the reference point. For the section where the straight line is connected to the circular curve, the piecewise expression can be obtained by the method of straight line + circular curve;
[0194] In the second step, mark the center dividing line of the road section in the same way, and select its line type (solid line / dashed line) to obtain its expression r 中 , and the center dividing line does not calculate the direction feature; in the last step, mark the white solid line or dashed line (there may be more than one pair) on the road section in the same way, select its line type, and obtain its expression r 内j , where j is the lane number. Thus, the type and position of the lane lines to be detected are obtained.
[0195] From Figure 14 , it can be seen that since vehicles drive on the right side, in the perspective of the drone's downward shooting, the driving direction θ of the traffic flow on the upper side in the figure always points to the negative x-axis direction ( the angle between the pixel coordinate system and the Cartesian coordinate system is equivalent), so it is defined that Figure 14 the driving direction of the vehicles on the upper side of the center dividing belt and the lane orientation are in the negative direction, and the driving direction of the vehicles on the lower side and the lane orientation are in the positive direction.
[0196] (2) Perform detection and tracking. Obtain the detection and tracking results of the aerial photography image in the perspective of the drone, obtain the trajectory information of vehicle i, and calculate the pixel coordinates of the vehicle center point
[0197] (3) Judge the vehicle direction. Calculate the vehicle speed and angle every 30 frames, and calculate the true speed of vehicle i at time t every 30 frames and direction When vehicle i performs the first calculation, according to its moving direction, if it is classified as a vehicle driving in the positive direction. Otherwise, it is classified as the negative direction.
[0198] (4) Judge the spatial position abnormal events on the highway according to the characteristic parameters and laws and regulations. When the above abnormal events are detected, the system will issue corresponding warnings. The judgment method is as follows:
[0199] Event of illegal use of the emergency lane: When a vehicle drives into the emergency lane area, it can be regarded as an abnormal event, that is, the center coordinates of vehicle i cross any one of the edge lines. If is substituted into r 外 , then is greater than or less than the y of any one direction r外 , it is considered that the vehicle has driven into the emergency lane. The judgment method for the event of illegal use of the emergency lane is:
[0200]
[0201] Pedestrian getting off: If a pedestrian is detected on the highway, it is considered that an event of a person getting off the vehicle (breaking into the highway) has occurred. The method for judging a person getting off the vehicle is as follows:
[0202] classes == pedestrian → Person getting off (13)
[0203] Reversing: First, determine the direction in which the vehicle appears, that is, substitute the abscissa of the center point of vehicle i into r 中 After that, if is greater than y r中 , it means that the vehicle is in the positive lane at this time, otherwise it is in the negative lane.
[0204] According to the vehicle direction category (positive / negative) obtained in step (3), if it is different from the lane direction (positive / negative) when the vehicle appears, it is considered that the vehicle i is reversing when it appears. The reverse driving judgment formula 1 is as follows:
[0205] Vehicle direction category (+ / ―) ≠ Lane direction (+ / ―) → Reversing (14)
[0206] Subsequently, continue to check for U-turn behavior during driving. When on a straight section, if the vehicle driving direction forms an angle greater than 90° with the direction θ 外 of the boundary line r r外 , it is considered that a reversing event has occurred; when on a circular curve section, connect the center coordinates of vehicle i with the center (x r外 , y r外 ) of the circular curve of r 外 , and intersect at a point (x, y) on r 外 , substitute it into equations (5)-(6) to find θ r外 . If the vehicle i driving direction forms an angle greater than 90° with the tangent θ r外 , it is considered that a reversing event has occurred. The reverse driving judgment formula 2 is:
[0207]
[0208] Crossing the line: It means that the vehicle illegally runs over or crosses the solid line, or continuously deviates from the center line of the lane by more than the allowable threshold. Calculate the distance d from the center coordinates of vehicle i to all white solid lines and double yellow solid lines in the same direction (i,rj) . After coordinate conversion, if there exists GSD(d (i,rj) ) < 0.5m, it is considered that a crossing the line event has occurred. The straight section crossing the line judgment formula is (16), and the circular curve section judgment formula is (17):
[0209]
[0210] where the GSD ground uses distance, k rj , b rj represent the slope and intercept of the lane edge line (obtained by least squares fitting), is the coordinate of the vehicle's current position at time t, d(i, r j ) is the vertical distance from the vehicle to the lane boundary, (x rj , y rj ) is the coordinate of the center of the curve (the center of the lane fitted by an arc), and R rj is the radius of the curve.
[0211] The determination of abnormal group behavior events includes:
[0212] Following risk event: By calculating the relationship between the headway of adjacent vehicles in the same lane and the speed of the following vehicle, when the headway is less than the safe following distance, it is determined that a following risk event has occurred; Traffic congestion event: By calculating the traffic density and average speed on the lane, when the vehicle density is greater than the preset threshold of the vehicle density and the average speed is lower than the preset threshold of the vehicle speed, it is determined that a traffic congestion event has occurred; Traffic flow overload event: By setting a virtual detection line on the lane cross-section and calculating the traffic flow passing through the detection line per unit time, when the traffic flow exceeds the lane capacity, it is determined that a traffic flow overload event has occurred. The specific implementation is as follows:
[0213] (1) Obtain the information of the lane lines. Similar to the previous section, calibrate the type and position of the lane line rj to be detected. After dividing it into the positive-direction lane and the negative-direction lane according to the start and end point directions, then arrange r j in ascending order. The value of brj in the lane line rj: yrj = krjx + brj (for the curve segment, use the ordinate yrj of the center of the circle), and obtain the dictionary information of all lanes. Taking a two-way four-lane as an example, there are a total of 5 lane lines including the central separation belt and the boundary lines and lane lines in the positive and negative directions after calibration. After sorting, record them as Define the area between two lane lines as lane R, that is and are the two lanes in the negative direction, and are the two lanes in the positive direction, as shown in Figure 15 .
[0214] (2) Calibrate the virtual line. Draw a straight line on the lane cross-section where you want to detect the traffic flow. Prepare for the subsequent detection of traffic flow overload events using the dynamic flow statistics method based on the detection line.
[0215] (3) Obtain the detection and tracking results of the aerial photography images from the perspective of the UAV. Calculate the pixel coordinates of the center point of vehicle i at time t.
[0216] (4) Judge the vehicle direction and calculate the vehicle speed every 30 frames. Calculate the true speed and direction of vehicle i in the image at time t every 30 frames. and direction When vehicle i performs the above calculations for the first time, according to its moving direction classify it as a vehicle traveling in the positive / negative direction.
[0217] (5) Synchronously conduct a statistical analysis of the lane-level data and record the vehicles on the lane every 30 frames. At step 3, synchronously calculate the positional relationship between vehicle i and the lane lines in the same direction to obtain the lane information. The calculation method is as follows: Starting from the r in the dictionary calculate the central coordinate of vehicle i at time t 中 and the vertical coordinate relationship with the lane line r First, confirm the relationship between vehicle i and the center line. If in r j then vehicle i is below the central median strip and belongs to the positive-direction lane; then, calculate the positional relationship between vehicle i and the next lane line in this direction. If then vehicle i is located on the positive-direction lane at this time, record the lane information at this moment Otherwise, continue to judge the positional relationship between vehicle i and the next lane line in the dictionary. Finally, obtain the lane information
[0218] (6) Judge the following-distance risk anomalies, traffic congestion, and traffic flow overload events on the highway according to the characteristic parameters and laws and regulations. When the above abnormal events are detected, the system will issue corresponding warnings. The judgment method is as follows:
[0219] Following-distance risk event: Statistically analyze the lane information obtain all the vehicle information on lane R at time t (j,k) and then sort it. If it is a positive-direction lane, sort it in ascending order according to the abscissa of vehicle i otherwise, sort it in descending order according to the value to obtain the serial number dictionary [i + n, i + n - 1,..., i + 1, i] from the leading vehicle to the trailing vehicle on lane R at time t. According to the dictionary order, calculate the headway between i + n and i + n - 1 (j,k) the headway between i + n - 1 and i + n - 2 ...... and the headway between i + 1 and i ...... Judge the event according to the speed of the following vehicle. The method for judging the following - vehicle risk event is as follows:
[0220]
[0221] Traffic congestion event: The traffic - flow statistical methods include the global statistical method based on detection boxes and the dynamic traffic - flow statistical method based on virtual detection lines. The global statistical method can count the lane information at time t to obtain the number of vehicles on lane R at this time Calculate the length L of lane R R , and obtain the traffic density For all vehicles on lane R at time t, calculate their average speed The method for judging the traffic congestion event is as follows:
[0222]
[0223] Flow - overload event: For the virtual detection line demarcated on the lane cross - section in step (2), using the least - squares method according to the starting and ending points (x1, y1), (x2, y2) to obtain its expression l: y = k sign x + b sign x ∈ (x1, x2), count the number of vehicles N passing through the detection line within a period of time (such as 15 minutes) R . When calculating for vehicle i, calculate its center point to the straight - line distance d from the line segment l l :
[0224]
[0225] where k sign is the slope of the lane boundary line, b sign is the intercept of the lane boundary line, is the coordinate position of the vehicle at time t
[0226] Only pay attention to the sign, and simplify the calculation to obtain the signed distance d sign :
[0227]
[0228] During the traffic - flow statistical process, calculate dsign of vehicle i every 30 frames. Let the signed distances of the vehicle at times t1 and t2 be and Then the traffic - flow statistical method is as follows:
[0229]
[0230] Calculate the traffic flow V passing through the detection line R(PCU / h), and compared with the lane capacity C (PCU / h), where C is generally taken as 2000. The method for judging traffic flow overload events is as follows:
[0231]
[0232] Detection effect display
[0233] The following are the detection effects of various abnormal events designed in this chapter by the highway abnormal event detection system constructed using the present invention. Since the occurrence probability of abnormal events on highways is low, the number of abnormal event shots from the drone's perspective that can be captured is very limited. In some test scenarios, due to the lack of real event shots on highways, the urban expressway scenario was used as a substitute simulation. The two have certain commonalities in traffic flow and lane types, which can effectively evaluate the detection performance of the system and provide feedback for system optimization.
[0234] Preparation before detection: After starting to shoot, the lane lines of the first frame of the picture are manually marked, and the detection line is set (skip the setting if traffic flow statistics are not required). The white solid line represents the boundary line, the white dotted line represents the lane separation line, and the yellow solid line represents the central separation line. The marked lane lines are as Figure 16 shown. It can be seen that the lane lines marked on the straight road can approximately reflect the situation of the real lane lines.
[0235] Detection of abnormal events in the running state: Figure 17 Shows the detection results of the id (left) and vehicle speed (right) of some vehicles by the detection system.
[0236] To verify the accuracy of the detection results, the vehicle speed in the unobstructed state of the urban expressway was detected using a drone, and at the same time, the manual method was used for speed measurement. The test results are compared in Table 9. The experiment shows that the algorithm can complete vehicle speed detection in real time. The average absolute error value reaches 3.7 km / h, and the relative error reaches 5.04%, and the error is within a reasonable range.
[0237] Table 9 Comparison test results of vehicle speeds
[0238]
[0239] Combined with the "Regulations for the Implementation of the Road Traffic Safety Law of the People's Republic of China", the effects of abnormal events such as parking, speeding, and low-speed driving detected by the highway abnormal event detection system constructed using the present invention are as Figure 18 shown, and a warning will be automatically issued when an abnormal vehicle is detected.
[0240] Detection of abnormal events in spatial location: The detection effects of the illegal use of the emergency lane, people getting off the vehicle, reverse driving, and lane crossing by the highway abnormal event detection system constructed using the present invention are asFigures 19 - 22 as shown, where Figure 19 shows a schematic diagram of the detection effect of the event of illegally using the emergency lane, Figure 20 shows a schematic diagram of the detection effect of the event of a person getting out of the vehicle, Figure 21 shows a schematic diagram of the detection effect of the event of driving in the reverse direction, Figure 22 shows a schematic diagram of the detection effect of the event of driving over the line. When an abnormal vehicle is detected, a warning will be automatically issued. It can be seen that the system can detect these types of abnormal events of spatial position in a timely and accurate manner and is universal for general roads.
[0241] Detection of abnormal group behavior events: The detection effects of abnormal events such as following risk abnormality, traffic congestion, and traffic flow overload using the highway abnormal event detection system constructed by the present invention are as Figures 23 - 25 shown, and a warning will be automatically issued when an abnormal situation is detected. Figure 22 Among them, for 3 lanes, 3 detection lines are respectively set for vehicle counting. When the statistical time ends (15 minutes is taken in the experiment), according to Equation (23), it is judged whether a traffic flow overload abnormal event occurs in the lane.
[0242] As can be seen from the above, from the perspective of the vertical downward aerial view of the drone, the system can accurately identify almost all vehicles in the picture, and through the designed event algorithm, further detect several types of abnormal group behavior events, which can cover most scenarios and help ensure highway traffic safety.
[0243] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle, characterized in that, Including the following steps: S1. Establish an aerial photography historical dataset from the perspective of the drone; S2. Construct a target detection model and a target tracking model. Among them, the target detection model adopts an improved YOLO v11 network structure, and the target tracking model adopts an improved DeepSORT network structure, and train the target detection model and the target tracking model based on the historical dataset; S3. Obtain real-time aerial photography images of the highway through the drone, input the real-time aerial photography images into the trained target detection model, detect the target vehicles in the images, and obtain the position information and environmental information of the target vehicles; S4. Input the detected target vehicle information into the trained target tracking model to obtain the motion trajectory information of the target vehicle; S5. According to the position information, environmental information and motion trajectory information of the target vehicle, analyze and judge whether the target vehicle has abnormal operating status, abnormal spatial position or abnormal group behavior; S6. Output the detection results of the abnormal events.
2. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 1, characterized in that The improved YOLO v11 network structure includes: The backbone network, which adopts the CSPDarknet network structure. Some of the conventional convolutional layers in the deep C3 module of the backbone network are replaced by deformable convolutional DCN v4. The deep C3 module corresponds to the generation of feature layers of P3, P4, and P5; The neck network, which adopts a feature pyramid structure, includes P5, P4, P3 feature layers and a P2 layer feature processing unit. The P2 layer feature processing unit is formed by concatenating the P3 layer features after 2-fold upsampling with the P2 layer features generated by the second downsampling of the backbone network in the channel dimension, so that the network obtains a local receptive field corresponding to 4×4 pixels of the original image. The CBAM attention mechanism is embedded in the feature fusion process of the neck network. The CBAM attention mechanism includes channel attention and spatial attention, which are used to enhance the extraction ability of small target-related features; The detection head network, which includes P5, P4, P3 detection heads and a P2 detection head, is used to detect targets of different sizes respectively. Among them, the CBAM attention mechanism is embedded in front of the P2 detection head to improve the positioning accuracy of small targets; The output layer, which is used to output the target position coordinates, confidence and category information.
3. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 2, wherein The Focal-WIOU Loss function is used for the training of the target detection model. The Focal-WIOU loss function is composed of the WIoU loss function and the Focal loss function. The calculation formula of the Focal-WIOU loss function is as follows: Focal - WIOU Loss = box × L WIoUv3 + cls × L focal Where, L focal = -α t ×(1 - p t ) γ × log(p t ) Where, L WIoUv3 is the bounding box regression loss, x and y represent the horizontal and vertical coordinates of the center of the predicted box respectively, x gt , y gt represent the horizontal and vertical coordinates of the center of the ground truth box respectively, r represents the dynamic non-monotonic focusing coefficient, L IoU represents the base intersection over union loss, L focal is the classification loss, W g and H g are the width and height of the smallest bounding box, the superscript * represents separation from the computational graph, p t is the predicted probability of the model for the true class, α t is the sample balance factor, γ represents the focusing parameter; box is the weight coefficient of the bounding box regression loss L WIoUv3 for adjusting the localization error, with a value of 0.05; cls is the weight coefficient of the classification loss L focal for adjusting the classification loss, with a value of 0.
5.
4. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 1, wherein The target detection model uses the soft non-maximum suppression algorithm to filter the target boxes. The soft non-maximum suppression algorithm realizes the dynamic screening of overlapping target boxes by applying a confidence decay mechanism based on the Gaussian function to the candidate boxes with higher IoU values. The formula of the soft non-maximum suppression algorithm is as follows: Among them, score i represents the initial confidence of the target box i, IOU(i, j) represents the intersection over union between the target box i and the target box j, σ represents the parameter controlling the attenuation rate, Score′ i is the corrected score after IOU penalty.
5. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 1, characterized in that The improved DeepSORT network structure is based on the ResNet-18 architecture, including an input layer, a convolutional layer, a pooling layer, a residual layer, and a fully connected layer. The residual layer includes a residual module and an SE attention module. The SE attention module is embedded after each residual module. The SE attention module extracts channel information through global average pooling, generates weights through linear transformation, and adjusts the feature map by weighting.
6. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 1, wherein, There is also coordinate transformation and trajectory processing between step S3 and step S4. Coordinate transformation includes the transformation from the image coordinate system to the camera coordinate system and the transformation from the camera coordinate system to the ground coordinate system; Trajectory processing includes identifying and correcting outliers in the trajectory data obtained by detection and tracking, and processing the trajectory data using the wavelet analysis method. The trajectory data is decomposed through wavelet transform to identify abnormal fluctuations. Noise reduction processing is carried out using a wavelet filter, and soft thresholding and hard thresholding methods are used to achieve noise reduction.
7. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 1, characterized in that, The position information includes the coordinate position, width, and height of the target vehicle in the image; The environmental information includes lane line position information and lane type information; The motion trajectory information includes the real-time speed, driving direction, and vehicle spacing of the target vehicle. Among them, the pixel position in the image coordinate system is converted into the actual position in the ground coordinate system through coordinate transformation.
8. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 7, wherein The judgment of the abnormal operation state event includes: Parking event: When the real-time speed of the vehicle is less than the preset value and the duration exceeds the set threshold, it is determined that a parking event has occurred; Speeding event: When the real-time speed of the vehicle exceeds the specified maximum speed limit, it is determined that a speeding event has occurred; Low-speed driving event: When the real-time speed of the vehicle is lower than the specified minimum speed limit and the duration exceeds the set threshold, it is determined that a low-speed driving event has occurred.
9. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 7, wherein: The judgment of the abnormal spatial position event includes: Event of illegal use of the emergency lane: When the center coordinate of the vehicle crosses the road edge line and enters the emergency lane area, it is determined that an event of illegal use of the emergency lane has occurred; Pedestrian getting off event: When a pedestrian is detected on the highway, it is determined that a pedestrian getting off event has occurred; Reverse driving event: When the direction of the vehicle is different from the direction of the lane where the vehicle appears or when the included angle between the driving direction of the vehicle and the specified driving direction of its lane is greater than 90 degrees, it is determined that a reverse driving event has occurred; Lane crossing event: When the vehicle crosses or crushes the solid line, or continuously deviates from the lane center line by more than the preset threshold, it is determined that a lane crossing event has occurred.
10. The method for detecting and tracking highway abnormal events from the perspective of an unmanned aerial vehicle according to claim 7, characterized in that, The judgment of the abnormal group behavior event includes: Following risk event: By calculating the relationship between the headway of adjacent vehicles in the same lane and the speed of the following vehicle, when the headway is less than the safe following distance, it is determined that a following risk event has occurred; Traffic congestion event: By calculating the traffic density and average vehicle speed on the lane, when the vehicle density is greater than the preset threshold of the vehicle density and the average vehicle speed is lower than the preset threshold of the vehicle speed, it is determined that a traffic congestion event has occurred; Traffic flow overload event: By setting a virtual detection line on the cross-section of the lane and calculating the traffic flow passing through the detection line per unit time, when the traffic flow exceeds the lane capacity, it is determined that a traffic flow overload event has occurred.
Citation Information
Cited By
Target motion track prediction and servo control optimization method based on AI
CN120762446A
Visual intelligent identification technology applied to unmanned aerial vehicle
CN120808221A
A visual intelligent identification technology used on a drone
CN120808221B
Method, system and equipment for detecting road vehicle speed and vehicle distance based on unmanned aerial vehicle vision
CN120877528A
Vehicle-road cooperation roadside signal processing device and method
CN120932460A