Traffic asset identification and tracking method based on YOLO and ByteTrack algorithms

By improving the combination of the YOLOv8 model and the ByteTrack algorithm, efficient and accurate traffic asset identification and tracking were achieved, solving the problems of low efficiency in traditional manual inspections and repetitive reports in YOLO detections, and improving the accuracy and efficiency of the inspection system.

CN120877240APending Publication Date: 2025-10-31NANJING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510966888.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional manual inspections are inefficient and struggle to comprehensively and accurately identify road infrastructure problems. YOLO detection results in repetitive reports, and existing technologies are insufficient for efficient and accurate identification and tracking of traffic assets.

Method used

Combining the YOLO and ByteTrack algorithms, the YOLOv8 model is improved through a small target detection enhancement module for detection, and the ByteTrack tracking algorithm with a motion compensation mechanism is introduced for target tracking. Finally, a logical judgment mechanism is used for event reporting.

Benefits of technology

It has improved the automatic detection and recognition capabilities of traffic signs and road facilities, reduced the rate of repeated incident reporting, enhanced recognition speed and accuracy, and improved the accuracy and image clarity of the inspection system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877240A_ABST
    Figure CN120877240A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic asset identification and tracking method based on YOLO and ByteTrack algorithms, and aims to improve the intelligence and working efficiency of road inspection, when an inspection vehicle enters a designated inspection area, a visible light camera is started to collect video images in real time; a YOLOv8 deep learning algorithm is used to carry out target detection on a video frame, a ByteTrack tracking algorithm fused with a motion compensation mechanism is used to carry out stable tracking on a detected target, and at the same time, a unique ID is allocated to each target. And cutting corresponding video frames based on a target detection frame generated by the tracker, naming the video frames according to the target ID and the event category, and uploading the video frames to an inspection system. Target tracking stability is improved through motion compensation, ID change caused by camera jitter or detection errors is effectively reduced, repeated reporting of inspection events is reduced through tracking information, and accuracy and efficiency of an inspection system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation system technology, specifically a method for traffic asset identification and tracking based on YOLO and ByteTrack algorithms. Background Technology

[0002] Driven by intelligent transportation systems and artificial intelligence technologies, traffic management and road inspection technologies are undergoing rapid innovation. Road inspection is a crucial link in ensuring road safety and maintaining the good condition of road facilities. Traditional methods rely on manual inspections to identify problems through regular patrols. However, manual inspections are inefficient, susceptible to subjective judgment and environmental conditions, and struggle to comprehensively and accurately identify various issues. Therefore, intelligent inspection technologies have emerged to improve inspection efficiency, reduce labor costs, and enhance the quality of road maintenance.

[0003] YOLO is an advanced real-time object detection algorithm that achieves efficient object detection by dividing an image into multiple grids and predicting bounding boxes and class probabilities in each grid. Due to its end-to-end training method and fast detection speed, YOLO is particularly suitable for traffic asset recognition. In road inspection, YOLO can detect and classify traffic assets in real time and record the location and type of signs. However, because the objects detected by YOLO are updated in real time, it can lead to a large number of duplicate reports during event reporting. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the present invention aims to provide a traffic asset identification and tracking method based on YOLO and ByteTrack algorithms.

[0005] The technical solution for achieving the purpose of this invention is: a traffic asset identification and tracking method based on YOLO and ByteTrack algorithms, comprising:

[0006] S1. Collect traffic sign and road information data;

[0007] S2. At the edge computing end, the improved YOLOv8 model using the small target detection enhancement module is used to detect real-time video captured by the vehicle camera and obtain target detection boxes.

[0008] S3. The ByteTrack tracking algorithm, which uses a fusion motion compensation mechanism, tracks the detected target at the edge computing end;

[0009] S4. Use the ByteTrack tracker to detect targets above the threshold in the current frame trajectory set and obtain a result array containing tracking ID and category;

[0010] S5. Cropping video frames based on the target detection boxes obtained by the tracker.

[0011] Compared with the prior art, the significant advantages of this invention are:

[0012] (1) It effectively improves the automatic detection and recognition capabilities of inspection vehicles for traffic signs and road facilities during operation;

[0013] (2) By combining YOLO and ByteTrack algorithms, rapid detection and accurate identification of traffic signs and road facilities were achieved, improving the identification speed and accuracy.

[0014] (3) When training with the YOLO algorithm, a small target detection layer and CBAM and SE attention modules were added to take into account the characteristics of traffic signs.

[0015] (3) An extended Kalman filter with a motion compensation mechanism was introduced into the ByteTrack algorithm for target tracking, which improved the stability of target tracking;

[0016] (4) The method of assigning a unique ID to the detection target was adopted, which greatly reduced the rate of repeated event reporting and increased the accuracy of the inspection system.

[0017] (5) By introducing a logical judgment mechanism, events can be reported according to different inspection vehicle speeds and different detection categories, which enhances the clarity and accuracy of the images. Attached Figure Description

[0018] Figure 1 This is a flowchart of the present invention.

[0019] Figure 2 This is a structural diagram of the YOLO network model of the present invention. Detailed Implementation

[0020] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and specific examples:

[0021] The present invention is conceived as a method for traffic asset identification and tracking based on YOLO and ByteTrack algorithms, the specific steps of which are as follows:

[0022] S1. Use a visible light camera to collect traffic signs and road information data.

[0023] A visible light camera was installed on the vehicle. A 1080P resolution RGB camera was selected to ensure high dynamic range (HDR) and good night vision capabilities to adapt to different lighting conditions. Videos of traffic signs and road guardrails under different environmental conditions were acquired. The videos were split into images frame by frame and saved in a folder. Several images of traffic signs and road information under different environmental conditions were obtained as a training set for object detection.

[0024] S2. Use the YOLOv8 model improved by the small object detection enhancement module to detect objects in real-time video captured by the vehicle camera and obtain object detection boxes.

[0025] Specifically, the network structure of the YOLOv8 model is shown in the appendix. Figure 2 The YOLOv8 model mainly consists of three parts: a backbone network, a neck network, and a detection head.

[0026] The backbone network consists of a set of basic convolutional modules, four sets of sliding window modules (Swin Transformers), four sets of C2f modules, a spatial pyramid fast pooling module (SPPF), and a channel attention module (SE). The processing flow is as follows: the image first passes through basic convolutional layers to extract low-level edge features, then sequentially passes through four sets of alternately stacked sliding window modules and C2f modules to progressively enhance the modeling of contextual semantic information. Finally, an SE attention module and a spatial pyramid fast pooling module are embedded after the last set of C2f modules, ultimately outputting highly discriminative features, significantly improving detection accuracy and robustness in complex scenes.

[0027] The neck network employs an improved bidirectional cross-layer feature fusion structure, enhancing feature interaction capabilities between different layers through collaborative modeling of multi-scale semantic and spatial information. First, the deep semantic features output from the end of the backbone network are upsampled by a factor of 2 and then channel-aligned and concatenated with the medium-resolution features output from the third C2f module of the backbone network to integrate deep semantic and mid-level structural information. Subsequently, the fused features are upsampled by a factor of 2 again and concatenated with the high-resolution features output from the second C2f module of the backbone network to form a fine-grained feature map for small and medium-sized target perception. Based on this, the fused features are upsampled by a factor of 2 again and concatenated with the shallow high-resolution features output from the first C2f module of the backbone network to generate a new 160×160 resolution feature path, serving as a high-resolution branch specifically for small target detection. This branch retains detailed target edge, texture, and positional features; to reduce computational complexity and match subsequent fusion paths, it is compressed using a 1×1 convolution. The compressed high-resolution features are then input into the CBAM attention module to enhance effective features in both spatial and channel dimensions while suppressing redundant background information.

[0028] Simultaneously, the compressed high-resolution features are introduced into the fusion process of the neck network, and spatially concatenated with the feature maps output by the second and third C2f modules after deep upsampling of the backbone network. Finally, they are concatenated with the deep features output by the SPPF module of the backbone network. During each fusion process, the spatially effective features after fusion are enhanced using the CBAM attention module. Ultimately, the neck network outputs a four-level multi-scale feature map enhanced by the CBAM module, corresponding to resolutions of 160×160 (fine-grained localization), 80×80 (medium-scale structure), 40×40 (semantic global perception), and 20×20 (deep global understanding), providing the detection head with sufficient information and clearly layered feature support.

[0029] The detection head, as the final prediction module of the YOLO network, is primarily responsible for generating the confidence scores for bounding box coordinates and target categories. While maintaining the efficiency of the original YOLOv8 architecture, the detection head performs parallel predictions on fused feature maps at four scales to adapt to the detection needs of targets of different sizes. Finally, the model is exported as an ONNX format model, adapted to edge device deployment environments, enabling model compression and deployment.

[0030] The real-time video stream captured by the vehicle-mounted camera is decoded using a high-efficiency software decoding library (such as OpenCV), and the decoded frame data is input into an optimized YOLO model detector for object detection. The location, category, and confidence score of all targets in the detection results are stored as bounding boxes, and the array of target bounding boxes is as follows:

[0031] results=[x1,y1,x2,y2,score,cl]

[0032] (x1,y1) and (x2,y2) are the coordinates of the top left and bottom right corners of the target bounding box, score is the confidence value of the target, and cl is the target category.

[0033] S3. Use the ByteTrack algorithm at the edge computing end to track the detected target.

[0034] For scenarios with high real-time requirements for in-vehicle video, the ByteTrack multi-target tracking algorithm is used at the edge computing end to track targets detected in the video stream. The core process of the ByteTrack algorithm in processing video frame sequences to achieve multi-target tracking is as follows:

[0035] S3.1: Set ByteTrack algorithm parameters:

[0036] track_thresh float = 0.25 track_buffer int = 30 match_thresh float = 0.8 aspect_ratio_thresh float = 3.0 min_box_area float = 1.0 mot20 bool = False

[0037] Among them, track_thresh represents the tracking confidence threshold, track_buffer is the number of frames used to retain lost tracks, match_thresh represents the tracking matching threshold, aspect_ratio_thresh represents the threshold for the aspect ratio of the target bounding box, min_box_area represents the threshold for the area of ​​the target bounding box, and mot20 indicates whether to use the mot20 dataset for testing.

[0038] S3.2: Initialization: Set an empty trajectory set T N It is used to store the tracking trajectory information of each target in the video stream.

[0039] S3.3: Target result classification: Classify the tracking trajectory and bounding box. All tracking trajectories are divided into two categories: active and inactive. All current frame bounding boxes are divided into two categories: high confidence (above the tracking confidence threshold) and low confidence (below the tracking confidence threshold).

[0040] An active trajectory is a detection box in the current frame that has a bounding box value higher than the tracking confidence threshold, and the intersection-over-union (IoU) value between this detection box and the corresponding detection box in the previous frame is higher than the set tracking matching threshold. Such trajectories are considered to be continuously active and participate in the matching and association of detection boxes in the current frame.

[0041] Inactive tracks refer to tracks that fail to match any detection boxes in the current frame, and whose inactivation threshold has been reached or exceeded since the last successful match, but have not yet met the conditions for complete removal. These tracks are considered temporarily lost and are reserved for subsequent compensatory matching with low-confidence detection boxes.

[0042] S3.4: Motion Compensation: The angular velocity ω = (ω_o) of the visible light camera is read using the IMU sensor. x ,ω y ,ω z ) and acceleration a = (a x ,a y ,a z The data is used to calculate the motion displacement Δt(t) and rotation angle Δθ(t) of the visible light camera in the current frame relative to the previous frame.

[0043]

[0044] Construct the camera rotation matrix R t R t =R x R y R z , where R x R y R z These are matrices that rotate about the X, Y, and Z axes, respectively.

[0045]

[0046] The coordinates P of the target point in the camera coordinate system in the previous frame. k-1 =(X k-1 ,Y k-1 Z k-1 Transform to world coordinate system Point P in the world coordinate system W Transform to the camera coordinate system P of the current frame k =R t P W +Δt t .

[0047] Using a perspective projection model, the target point P k Projected onto the image coordinate system of the current frame:

[0048]

[0049] Finally, the motion-compensated detection results are results = [x′1, y′1, x′2, y′2, score, cl], where x k ,y k (k = 1, 2) represents the original detection box coordinates, f x ,f y This refers to the camera's focal length.

[0050] S3.5: First Match: Using the Hungarian algorithm to match high-confidence target boxes B H Given an existing trajectory set T, calculate the high-confidence target box B. H The matching cost function with the existing trajectory set T is:

[0051]

[0052] in, and c(T) j d1(i,j) represents the center point of the bounding box and the trajectory, respectively. A successful match is considered achieved when d1(i,j) is less than the matching threshold. The successfully matched trajectory will be updated to its position in the next frame using an Extended Kalman Filter (EKF). And store it in the current frame trajectory set T N Unmatched trajectories are added to the first set of unmatched trajectories, T. R Unmatched detection boxes are stored in the first set of unmatched detection boxes, B. R .

[0053] Specifically, the extended Kalman filter update steps are as follows:

[0054] Based on the target's state information from the previous frame, the state of the target in the current frame is predicted using a state transition model. The motion compensation prediction formula combining IMU data is as follows: in, Let F be the bounding box position of the target in the next frame, and let p be the state transition matrix. t Let G be the position of the target in the current frame, G be the IMU motion compensation matrix, and ω be the process noise.

[0055] S3.6: Second Match: For low-confidence target box B L And the first unmatched trajectory T R A second round of matching is then performed to recover as many missed targets as possible. IMU data is used to analyze the unmatched trajectory T. R Perform motion compensation to predict its position in the current frame:

[0056]

[0057] in, It is the predicted target position after IMU motion compensation, p t-1 This is the actual position of the target at the previous moment. Calculate B. L and T R The matching cost function is:

[0058]

[0059] When d2(i,j) is less than the matching threshold, the match is considered successful, and the trajectory is stored in the current frame trajectory set T. N Trajectories that fail to match are stored in the lost trajectory set T. D Unmatched low-confidence detection boxes are deleted directly to avoid target drift caused by false detections.

[0060] S3.7: Initialize new trajectory: For high-confidence bounding boxes B that still fail to match R Create a new target trajectory for each detection box and add it to the current frame trajectory set T. N Meanwhile, for the set of lost trajectories T D If a target in the tracking list loses more than a threshold N consecutively due to a missing matching frame, it is marked as a failed target and removed from the tracking list. Finally, the current frame trajectory set is output as the existing trajectory set for the next frame image, and a new loop begins from S3.3. The loop terminates when there is no image input from the visible light camera.

[0061] S4. Use the ByteTrack tracker to track the current frame trajectory set T. N Detect targets above a certain threshold and obtain a result array (outputs) containing the tracking ID (track_id) and category (cl). The specific steps are as follows:

[0062] S4.1: The ByteTrack tracker assigns the current frame track set T N For targets detected above the threshold, the results array contains individual tracking IDs. The updated results are saved as a new array with the following form:

[0063] tracks=[x1,y1,x2,y2,score,track_id]

[0064] (x1,y1) and (x2,y2) are the coordinates of the top left and bottom right corners of the target bounding box, score is the confidence value of the target, and track_id is the unique ID of the target;

[0065] S4.2: Use the IOU algorithm to calculate the intersection-union ratio (IU) of image dimensions in the results array of S2.6 and the tracks array of S4.1. The specific steps are as follows:

[0066] First, extract the top-left and bottom-right corner (x1, y1) and (x2, y2) coordinates of the target bounding box from the current frame's results and tracks arrays, and combine them into a single array.

[0067] Then, calculate the intersection region I and the union area U.

[0068]

[0069] I = I w ×I h

[0070]

[0071] U = A r +A t -I

[0072] Finally, the Intersection over Union (IoU) ratio is calculated using the following formula:

[0073]

[0074] When the IoU is greater than the threshold, the match is considered successful. The track_id in the tracks array is added to the corresponding results array and saved as a new array with the following form:

[0075] outputs=[x1,y1,x2,y2,score,track_id,cl]

[0076] (x1,y1) and (x2,y2) are the coordinates of the top left and bottom right corners of the target bounding box, score is the confidence value of the target, track_id is the unique ID of the target, and cl is the category of the target.

[0077] S5. Based on the target detection bounding boxes obtained by the tracker, crop the video frames and upload them to the inspection system. The specific steps are as follows:

[0078] S5.1: Convert the output detection box position coordinates after the tracker is matched to an integer (int) to meet the input requirements of the clipping operation.

[0079] S5.2: Perform boundary condition checks to ensure that the cropped image area does not exceed the actual range of the original image. If it does, adjust the coordinates of the cropped area to within the following boundaries:

[0080] x1 = max(0, x1)

[0081] y1 = max(0, y1)

[0082] x2=min(w org (x2)

[0083] y2=min(h org ,y2)

[0084] Among them, w org and h org These represent the width and height of the original image, respectively. The max and min functions are used to ensure that the coordinate values ​​do not exceed the image boundaries.

[0085] S5.3: Create an empty dictionary `frame_count` to count frames for each target. Initialize this dictionary when a target is detected, and increment the count by 1 for each frame the target is present.

[0086] S5.4: Dynamically set the frame capture threshold based on the detected target type and current vehicle speed. When the count reaches the threshold, determine the cropping area based on the adjusted coordinate values, and crop the image region containing the target from the original image, naming it with ID and category, and uploading it to the inspection system.

[0087] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

[0088] It should be understood that, in order to simplify the present invention and help those skilled in the art understand its various aspects, in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes described in a single embodiment or with reference to a single figure. However, the present invention should not be construed as including all features in the exemplary embodiments as essential technical features of the claims of this patent.

[0089] It should be understood that the modules, units, components, etc., included in the device of one embodiment of the present invention can be adaptively changed to be placed in a device different from that embodiment. Different modules, units, or components included in the device of the embodiment can be combined into a single module, unit, or component, or they can be divided into multiple sub-modules, sub-units, or sub-components.

Claims

1. A method for traffic asset identification and tracking based on YOLO and ByteTrack algorithms, characterized in that, include: S1. Collect traffic sign and road information data; S2. At the edge computing end, the improved YOLOv8 model using the small target detection enhancement module is used to detect real-time video captured by the vehicle camera and obtain target detection boxes. S3. The ByteTrack tracking algorithm, which uses a fusion motion compensation mechanism, tracks the detected target at the edge computing end; S4. Use the ByteTrack tracker to detect targets above the threshold in the current frame trajectory set and obtain a result array containing tracking ID and category; S5. Cropping video frames based on the target detection boxes obtained by the tracker.

2. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 1, characterized in that, The YOLOv8 model comprises a backbone network, a neck network, and a detection head. The backbone network includes a set of basic convolutional modules, four sets of sliding window modules, four sets of C2f modules, a spatial pyramid fast pooling module, and a channel attention module. The image first passes through basic convolutional layers to extract low-level edge features, and then sequentially passes through four sets of alternately stacked sliding window modules and C2f modules to gradually enhance the modeling of contextual semantic information. After the last set of C2f modules, a channel attention module and a spatial pyramid fast pooling module are embedded to finally output highly discriminative features. The neck network adopts an improved bidirectional cross-layer feature fusion structure, which enhances the feature interaction capability between different layers through multi-scale semantic and spatial information collaborative modeling; The detection head is used to generate confidence scores for bounding box coordinates and target category.

3. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 2, characterized in that, The image processing procedure of the neck network is as follows: The deep semantic features output at the end of the backbone network are upsampled by 2 times and then channel-aligned and fused with the medium-resolution features output by the third C2f module of the backbone network to integrate deep semantic and medium-level structural information. Subsequently, the fused features are upsampled by 2 times again and then fused with the high-resolution features output by the second C2f module of the backbone network to form a fine-grained feature map for small and medium-sized target perception. The fused features are upsampled by 2 times again and concatenated with the shallow high-resolution features output by the first C2f module of the backbone network to generate a new 160×160 resolution feature path, which serves as a high-resolution branch dedicated to small object detection. The high-resolution branch is compressed by 1×1 convolution. The compressed high-resolution features are then input into the CBAM attention module to enhance the effective features in the spatial and channel dimensions and suppress background redundancy. Simultaneously, the compressed high-resolution features are introduced into the fusion process of the neck network, and spatially stitched and fused with the feature maps output by the second and third C2f modules after deep upsampling of the backbone network. Then, they are finally stitched with the deep features output by the spatial pyramid fast pooling module of the backbone network. In each fusion process, the spatially effective features after fusion are enhanced by the CBAM attention module. Finally, the neck network outputs a four-level multi-scale feature map enhanced by the CBAM module.

4. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 2, characterized in that, The target detection bounding box is specifically as follows: results=[x1,y1,x2,y2,score,cl] In the formula, (x1,y1) and (x2,y2) are the coordinates of the top left and bottom right corners of the target bounding box, score is the confidence value of the target, and cl is the target category.

5. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 1, characterized in that, The specific method for tracking detected targets using the ByteTrack tracking algorithm with a fusion motion compensation mechanism at the edge computing end is as follows: Set an empty trajectory set T N It is used to store the tracking trajectory information of each target in the video stream; The tracking trajectories and bounding boxes are classified. All tracking trajectories are divided into two categories: active tracking trajectories and inactive tracking trajectories. All current frame bounding boxes are divided into two categories: high confidence and low confidence based on the tracking confidence threshold. Activating the tracking trajectory means that there is a detection box in the current frame that is higher than the tracking confidence threshold, and the intersection-union ratio of the detection box and the corresponding detection box in the previous frame is higher than the set tracking matching threshold. An inactive tracking track is a tracking track that fails to match any detection box in the current frame and whose frame count since the last successful match has reached or exceeded the inactivation threshold, but has not yet met the conditions for being completely removed. The angular velocity ω = (ω_o) of the visible light camera is read using an IMU sensor. x ,ω y ,ω z ) and acceleration a = (a x ,a y ,a z The data is used to calculate the motion displacement Δt(t) and rotation angle Δθ(t) of the visible light camera in the current frame relative to the previous frame. Construct the camera rotation matrix R t R t =R x R y R z , where R x R y R z These are the matrices for rotation about the X, Y, and Z axes, respectively; The coordinates P of the target point in the camera coordinate system in the previous frame. k-1 =(X k-1 ,Y k-1 Z k-1 Transform to world coordinate system Point P in the world coordinate system W Transform to the camera coordinate system P of the current frame k =R t P W +Δt t ; Using a perspective projection model, the target point P k Projected onto the image coordinate system of the current frame: Ultimately, the detection results after motion compensation are: results = [x ′ 1, y ′ 1,x′2,y′2,score,cl], where x k ,y k f represents the original detection box coordinates. x ,f y Let k be the focal length of the camera, where k = 1, 2; The Hungarian algorithm was used to match high-confidence target boxes B. H and the existing trajectory set T; The successfully matched trajectory will be updated to the next frame position using an extended Kalman filter. And store it in the current frame trajectory set T N Unmatched trajectories are added to the first set of unmatched trajectories, T. R Unmatched detection boxes are stored in the first set of unmatched detection boxes, B. R ; For the low-confidence target box B L And the first unmatched trajectory T R Then proceed to the second round of matching; For the high-confidence target box B that still cannot be matched R Create a new target trajectory for each detection box and add it to the current frame trajectory set T. N ; Meanwhile, for the set of lost trajectories T D If a target in the tracking list loses more than a threshold N consecutively, it is marked as a failed target and removed from the tracking list. Finally, the current frame trajectory set is output as the existing trajectory set for the next frame image.

6. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 5, characterized in that, The Hungarian algorithm was used to match high-confidence target boxes B. H The specific method for using the existing trajectory set T is as follows: Calculate the high-confidence target box B H The matching cost function with the existing trajectory set T is as follows: in, and c(T) j These are the center points of the bounding box and the trajectory, respectively. Represents the i-th high-confidence target box With the j-th trajectory T j The intersection-union ratio between the current boxes, GΔt(t) is the trajectory movement distance within the time interval Δt, and λ1 and λ2 are weight parameters; A match is successful when the matching cost function is less than the matching threshold.

7. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 5, characterized in that, For the low-confidence target box B L And the first unmatched trajectory T R The specific method for performing the second round of matching is as follows: Using IMU data to analyze the unmatched trajectory T R Perform motion compensation to predict its position in the current frame: in, It is the predicted target position after IMU motion compensation, p t-1 This is the actual position of the target at the previous moment. Calculate B. L and T R The matching cost function is: This represents the intersection-union ratio (IoU) between the i-th low-confidence bounding box and the current bounding box of the j-th unmatched trajectory. The center point coordinates of the target bounding box, λ1 and λ2 are weight parameters. When d2(i,j) is less than the matching threshold, it is considered a successful match, and the trajectory is stored in the current frame trajectory set T. N Trajectories that fail to match are stored in the lost trajectory set T. D Unmatched low-confidence detection boxes are deleted directly.

8. The traffic asset identification and tracking method based on YOLO and ByteTrack algorithms according to claim 1, characterized in that, The specific method for detecting targets above a threshold in the current frame trajectory set using the ByteTrack tracker and obtaining a result array containing tracking IDs and categories is as follows: S4.1: The ByteTrack tracker assigns the current frame track set T N For targets detected above the threshold, the results array contains individual tracking IDs. The updated results are saved as a new array with the following form: tracks=[x1,y1,x2,y2,score,track_id] (x1,y1) and (x2,y2) are the coordinates of the top left and bottom right corners of the target bounding box, score is the confidence value of the target, and track_id is the unique ID of the target; S4.2: Use the IOU algorithm to calculate the intersection-union ratio of image dimensions in the array obtained in S2 and the tracks array in S4.1; When the IoU is greater than the threshold, the match is considered successful. The track_id in the tracks array is added to the corresponding array obtained in S2, and saved as a new array with the following form: outputs=[x1,y1,x2,y2,score,track_id,cl] (x1,y1) and (x2,y2) are the coordinates of the top left and bottom right corners of the target bounding box, score is the confidence value of the target, track_id is the unique ID of the target, and cl is the category of the target.

Citation Information

Cited By

  • Multi-target tracking method and system for small target of unmanned aerial vehicle

    CN121661545A

  • A multi-target tracking method and system for small targets of unmanned aerial vehicles

    CN121661545B