Multi-target tracking method based on YOLOv8 model and Byte Track algorithm
By optimizing the YOLOv8 detection model and using symmetric quantization algorithm, the problems of large delay and insufficient computing power during deployment on edge devices in the existing technology are solved, and real-time multi-objective tracking effect on resource-constrained devices are achieved.
Patent Information
- Application Number
- CN202510175799.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
When the existing target tracking system based on YOLOv8+Byte Track is deployed on edge devices with limited resources, it has a large delay and insufficient computing power, so it cannot achieve the ideal effect of real-time multi-target tracking.
By optimizing the YOLOv8 detection model, replacing its detection head with a dynamic detection head, and adding a convolutional layer to the neck network to achieve cross-scale feature fusion, forming a DC-YOLOv8 network structure. At the same time, a symmetric quantization algorithm is used to quantize the model to reduce the computing power consumption.
The optimized YOLOv8+Byte Track object detection and tracking model is implemented on edge devices, improving the lightweight and accuracy of the model, ensuring real-time multi-objective tracking on resource-constrained devices.
Smart Images

Figure CN120107313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and edge device technology, and in particular to a multi-target tracking method based on a YOLOv8 model and a Byte Track algorithm. Background Art
[0002] In recent years, in the field of artificial intelligence, especially computer vision, the development of target detection and target tracking technologies has been rapid. At the same time, there is a large demand in intelligent driving, drone detection, medical imaging, etc. Major companies have launched a series of AI developer kits, namely edge devices. These developer kits are products for AI developers' knowledge learning, algorithm verification and application development scenarios. They have the characteristics of surging computing power, complete interfaces, out-of-the-box use, and rich samples. They are suitable for individual developers, university teachers and students, industry engineers and other user groups, and meet the development needs of developers in video image analysis, natural language processing, robotics and other fields. At the same time, various systems and applications accompanying the corresponding development boards have also been born. In the system that needs to achieve target tracking, how to select the detection + tracking algorithm according to the actual situation and further optimize and deploy it is a problem that must be considered.
[0003] The Byte Track algorithm is an advanced algorithm for multiple object tracking (MOT), which achieves continuous tracking of multiple moving targets in a video by combining a deep learning detector and an efficient online association method. The Byte Track algorithm aims to improve the accuracy and efficiency of tracking tasks while maintaining good real-time performance. The YOLO (You Only Look Once) series of algorithms are a set of deep learning models for real-time target detection in the field of computer vision. These models are popular in multiple applications such as autonomous driving, security monitoring, and smart retail due to their fast speed and good accuracy. Since its first release in 2016, the YOLO series has undergone multiple iterations, and each version has introduced new improvements and technologies to improve performance and efficiency. Among them, YOLOv8 is the eighth version of the YOLO series of target detection algorithms proposed in 2023. Compared with previous versions, it has been optimized in various aspects. It currently supports image classification, object detection, and instance segmentation tasks. In view of the good performance of previous versions, the v8 version has attracted much attention, and once it was born, it has generated countless application cases in related fields. For example, the Chinese patent "CN119049305A A vehicle flow statistics method and related equipment based on YOLOv8 algorithm" is a target tracking system application solution based on YOLOv8+Byte Track, and its basic process is: obtain monitoring video data, the monitoring video data includes a vehicle to be detected area and a vehicle must pass through the departure area; detect the vehicle in the vehicle to be detected area according to the YOLOv8 algorithm to obtain a detection result; calculate the detection result according to the improved Byte Track tracking algorithm to obtain a tracked trajectory; obtain the vehicle flow statistics result of the vehicle to be detected area according to the vehicle must pass through the departure area and the tracked trajectory.
[0004] However, the existing target tracking systems based on YOLOv8+Byte Track have not optimized the detection model or have not optimized it properly, and are generally not quantified during deployment. As a result, when the target tracking system is deployed on resource-constrained edge devices, there will be large delays and insufficient computing power. In the scenario of real-time multi-target tracking, the ideal effect is often not achieved. Summary of the invention
[0005] In view of the above-mentioned shortcomings of the prior art, the present invention optimizes the YOLOv8 detection model and uses the Byte Track algorithm to further optimize the tracking process, and proposes a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm. The purpose is to provide a method for deploying the optimized YOLOv8+Byte Track target detection and tracking model on the edge device.
[0006] The present invention proposes a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm, the method comprising the following steps:
[0007] Step 1: Obtain several image data and convert their formats to construct an image dataset;
[0008] Step 2: Build a DC-YOLOv8 network structure and use the image data set for training, and then use the optimal DC-YOLOv8 network structure obtained through training to build a target detection model based on the DC-YOLOv8 network structure;
[0009] Step 3: Obtain the video stream or image sequence to be detected and extract continuous image frames from it, preprocess the extracted image frames, and initialize the target detection model based on the DC-YOLOv8 network structure;
[0010] Step 4: Input the preprocessed image frame into the initialized target detection model based on the DC-YOLOv8 network structure for target detection. If the target is detected, execute step 5; if the target is not detected, use the image frame output by the target detection model to generate a video stream as the final video output;
[0011] Step 5: Use the Byte Track algorithm to track the targets detected in all image frames and generate tracking trajectories of the tracked targets in all image frames;
[0012] Step 6: Generate a video stream according to the tracking trajectory of the tracked target in all image frames and output it as the final video;
[0013] The specific content of step 1 is: obtaining a plurality of image data, wherein each image data includes: an image and annotation information of the image; the annotation information includes: bounding box information and category; converting the annotation information of all image data into YOLO format to ensure that each image file includes a corresponding .txt file, thereby constructing an image data set; and then dividing the image data set into a training set, a validation set, and a test set according to a preset ratio;
[0014] The specific content of step 2 is: improving the existing YOLOv8 network structure; the improvement is: replacing the detection head in the head network of the YOLOv8 network structure with a dynamic detection head Dynamic Head;
[0015] In the neck network of the YOLOv8 network structure, for each upsampling layer, a convolution layer is added before and after the upsampling layer and the number of channels of the convolution layer is reduced, thereby transforming the unidirectional feature pyramid network FPN in the YOLOv8 network structure into a cross-scale feature fusion network CCFF, and the improved YOLOv8 network structure is called the DC-YOLOv8 network structure;
[0016] The DC-YOLOv8 network structure is iteratively trained using the training set, and the trained DC-YOLOv8 network structure is evaluated using the validation set. The DC-YOLOv8 network structure is optimized by adjusting the network structure parameters according to the evaluation results, and then the optimized DC-YOLOv8 network structure is tested using the test set to obtain the optimal DC-YOLOv8 network structure.
[0017] The symmetric quantization algorithm is used to symmetrically quantize the optimal DC-YOLOv8 network structure to obtain a target detection model based on the DC-YOLOv8 network structure;
[0018] Furthermore, the dynamic detection head Dynamic Head includes a series of π L , π S and π C Three parts; among them π L It includes the following: average pooling layer, convolution layer, ReLU activation function and hard sigmoid activation function connected in series, and multiplying the output of the hard sigmoid activation function by the feature map input in the average pooling layer element by element to obtain π L The output of S It includes the index feature channel and convolution layer connected in series, and the output of the convolution layer is processed by the offset layer and the sigmoid activation function respectively and then merged as π S The output of C It includes the following series: average pooling layer, two fully connected layers and normalization layer; a ReLU activation function is added between the two fully connected layers, and a bias term is set to adjust the output of the normalization layer, and the adjusted output is used as π C Output:
[0019] Further, the process of symmetrically quantizing the optimal DC-YOLOv8 network structure using the symmetric quantization algorithm is as follows: converting the model file of the optimal DC-YOLOv8 network structure into a model file format that meets the requirements of the actual target detection task, and symmetrically quantizing the converted optimal DC-YOLOv8 network structure, converting the network parameters in the converted optimal DC-YOLOv8 network structure from floating point numbers to signed integers, calculating the proportional factor using the maximum value of the absolute values of all weights in the network parameters, and quantizing the weights of all network parameters using the proportional factor, and then saturating the quantized weights to complete the symmetrical quantization of the optimal DC-YOLOv8 network structure, and obtaining a target detection model based on the DC-YOLOv8 network structure;
[0020] The specific content of step 4 is: input the preprocessed image frame into the initialized target detection model based on the DC-YOLOv8 network structure, for any image frame, use the backbone network in the target detection model to extract features of the image frame, and use the extracted feature map to detect whether there is a target in the image frame, if there is, generate the detection results of all targets in the image frame, including: the coordinates and size of the bounding box, the category to which the target belongs and the category confidence, and continue to perform target detection on the next image frame; if there is no target, directly perform target detection on the next image frame until the detection of all image frames is completed; if no target is detected in all image frames, use the image frame output by the target detection model to generate a video stream as the final video output, otherwise execute step 5;
[0021] The specific contents of step 5 are as follows: for any image frame, each target detected in the image frame is taken as a tracking target, a tracking trajectory is initialized for each tracking target in the image frame, and a unique ID is generated for each tracking trajectory; for any tracking target, a Kalman filter is initialized, an initial trajectory state is generated for the tracking target, and the tracking trajectory of the tracking target is added, wherein the initial trajectory state includes an initial position and an initial velocity of the tracking target; based on the initial trajectory state of the tracking target, the position and velocity of the tracking target in the current image frame are predicted using the initialized Kalman filter, and the predicted position and velocity are used as the existing trajectory of the tracking target in the current image frame, thereby obtaining the existing trajectories of all tracking targets in the current image frame;
[0022] Preliminarily matching each bounding box in the current image frame with the existing trajectories of all tracked targets in the current image frame in turn to generate matching trajectories of each bounding box;
[0023] The Kalman filter is used to update the matching trajectory of each bounding box, and the position and speed of the tracking target corresponding to the matching trajectory of each bounding box are corrected, and the updated matching trajectory is added to the tracking trajectory of the tracking target; at the same time, the updated matching trajectory is used as the existing trajectory of the tracking target for tracking the target in the next image frame;
[0024] Initialize a death list with a capacity of N, and after each preliminary match between each bounding box in an image frame and the existing tracks of all the tracked targets in the image frame, add the existing tracks in the image frame that have not been successfully matched to the death list, and while performing preliminary matching between each bounding box in the next image frame and the existing tracks of all the tracked targets in the image frame, perform secondary matching between each existing track in the death list and each bounding box in the next image frame, and remove the existing tracks that have been successfully matched for the secondary matching from the death list; if any existing track remains in the death list after N consecutive frames, the tracked target corresponding to the existing track is regarded as a missing target; wherein the secondary matching method is the same as the method for performing preliminary matching between each bounding box in the current image frame and the existing tracks of all the tracked targets in the current image frame;
[0025] Furthermore, the method for preliminarily matching each bounding box in the current image frame with the existing trajectories of all tracked targets in the current image frame is as follows: for any bounding box in the current image frame, the IoU values between the bounding box and the predicted trajectories of all tracked targets in the current image frame are calculated respectively; if the IoU value between the bounding box and the existing trajectory of a tracked target in the current image frame is greater than a set IoU threshold, the bounding box is considered to be associated with the existing trajectory; otherwise, the bounding box is considered to be unassociated with the existing trajectory, and the target corresponding to the bounding box is taken as a new tracked target;
[0026] For all existing tracks associated with the bounding box, delete all existing tracks that are different from the category corresponding to the bounding box. Among all the retained existing tracks, when the difference between the IoU value of the bounding box and an existing track and the IoU threshold is within a preset range, select the matching track of the bounding box using the category confidence of the bounding box;
[0027] Further, the method of selecting the matching track of the bounding box using the category confidence of the bounding box is: using the category confidence of the bounding box and the IoU value between the bounding box and the existing track, respectively calculating the similarity score between each existing track associated with the bounding box and the bounding box, and selecting the existing track with the highest similarity score as the matching track of the bounding box;
[0028] The similarity score is expressed as:
[0029] S total =w1 ×IoU+w 2 ×confidence
[0030] Where S total represents the similarity score between the existing track and the bounding box; w 1 and w 2 Both represent weights; IoU represents the IoU value between the bounding box and the existing track; confidence represents the category confidence of the bounding box.
[0031] The beneficial effects of adopting the above technical solution are:
[0032] The method of the present invention changes the original detection head in the YOLOv8 network structure into Dyhead (Dynamic Head), and adds conv convolution at the appropriate position in the neck network, thereby achieving lightweight reduction and improved accuracy of the detection model. At the same time, in the deployment stage, the method of the present invention quantizes the detection model through the amct tool, which reduces the computing pressure of the edge device and makes the video smoother when tracking more targets.
[0033] The method of the present invention performs lightweight or optimization operations on various aspects of data set selection, model network structure optimization, model training, quantization and conversion in the multi-target tracking process, thereby ensuring that the task of real-time multi-target tracking can be achieved on resource-constrained devices. When the method of the present invention is used to deploy a target detection and tracking system on an edge device, the method of the present invention minimizes the occupation of computer resources, including CPU, GPU or NPU, while being able to complete the reasoning task and achieve the ideal effect, so as to solve the problem that the conventional direct deployment effect is not ideal. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is an improved schematic diagram of the multi-target tracking method in this embodiment;
[0035] Figure 2 This is a flowchart of a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm in this embodiment;
[0036] Figure 3 Schematic diagram of the Dynamic Head network structure in the YOLOv8 network structure in this implementation mode;
[0037] Figure 4 Schematic diagram of the improved head network structure in the YOLOv8 network structure in this implementation mode;
[0038] Figure 5 Schematic diagram of the improved neck network structure in the YOLOv8 network structure in this implementation manner;
[0039] Figure 6 This is a flow chart of the detection model deployment process in this implementation mode;
[0040] Figure 7 is a schematic diagram of the process of automatically quantizing the detection model in this embodiment;
[0041] Figure 8 This is a schematic diagram of the process of running the detection model on the edge device in this implementation. DETAILED DESCRIPTION
[0042] In order to facilitate the understanding of the present application, the specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thoroughly understood.
[0043] like Figure 1 As shown, this embodiment improves the existing multi-tracking method, including the following three aspects: research based on the improved YOLOv8 target detection model, research based on the Byte Track algorithm lightweight target tracking algorithm and optimized deployment for edge devices. Among them, the research based on the improved YOLOv8 target detection model refers to the use of YOLO algorithm, convolution optimization and detection head lightweight technical means to optimize the existing YOLOv8 model to obtain a target detection model with higher accuracy and lower parameters. The research on lightweight target tracking algorithm based on the Byte Track algorithm refers to tracking the target obtained by the target detection model based on the technical means of the Byte Track algorithm, Kalman filtering and Hungarian algorithm to obtain an accurate and lightweight target tracking model. The content of the optimized deployment for edge devices includes model conversion, model quantization and embedded design, and its purpose is to be able to build a real-time multi-target tracking system for completing actual scene applications.
[0044] This implementation method will be dedicated to optimizing the target tracking deployment method for edge devices, especially in aerial areas such as drones. The main technologies include target detection, target tracking, model lightweighting and model conversion.
[0045] This embodiment is a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm. Figure 2 As shown, the method comprises the following steps:
[0046] Step 1: Obtain several image data and convert their formats to build an image dataset.
[0047] The specific content of step 1 is: obtaining a number of image data, wherein each image data includes: an image and annotation information of the image; the annotation information includes: bounding box information and category; converting the annotation information of all image data into YOLO format to ensure that each image file includes a corresponding .txt file, thereby constructing an image data set; and then dividing the image data set into a training set, a validation set, and a test set according to a preset ratio.
[0048] In this embodiment, it is necessary to first build the relevant software environment, including Python, PyTorch and some other auxiliary libraries, and then obtain the VisDrone2019 image dataset from the official website. In order to match the input requirements of YOLOv8, the acquired image data needs to be preprocessed, that is, the data format is converted so that each image has a corresponding .txt file, which contains bounding box information and category ID. The image dataset is divided according to a preset ratio, and the image dataset consists of 10,209 images, including 6,471 training set images, 3,190 test set images, and 548 verification set images.
[0049] Step 2: Build a DC-YOLOv8 network structure and use the image data set for training, and then use the optimal DC-YOLOv8 network structure obtained through training to build a target detection model based on the DC-YOLOv8 network structure.
[0050] The specific content of step 2 is: improving the existing YOLOv8 network structure; the improvement is: replacing the detection head in the head network of the YOLOv8 network structure with a dynamic detection head Dynamic Head.
[0051] The dynamic detection head Dynamic Head includes a series of π L , π S and π C Three parts; among them π L It includes the following: average pooling layer, convolution layer, ReLU activation function and hard sigmoid activation function connected in series, and multiplying the output of the hard sigmoid activation function by the feature map input in the average pooling layer element by element to obtain π L The output of S It includes the index feature channel and convolution layer connected in series, and the output of the convolution layer is processed by the offset layer and the sigmoid activation function respectively and then merged as π S The output of CIt includes the following series: average pooling layer, two fully connected layers and normalization layer; a ReLU activation function is added between the two fully connected layers, and a bias term is set to adjust the output of the normalization layer, and the adjusted output is used as π C Output.
[0052] In this embodiment, the YOLOv8 network structure is mainly composed of three parts: the backbone network Backbone, the neck network Neck and the head network Head; and the improvement of the existing YOLOv8 network structure is divided into three steps. First, the dynamic detection head Dynamic Head is introduced into the YOLOv8 head network. The structural principle of Dynamic Head is as follows: Figure 3 As shown, where π L The average pooling layer in is used to perform average pooling on the input feature map to reduce the spatial dimension of the feature map; then a 1x1 convolution layer is used to adjust the number of channels or integrate features, and a ReLU activation function is used to increase the nonlinearity of the network; finally, a hard sigmoid activation function is used to limit the output value to the range of [0,1]. S It includes a series of index feature channels and convolution layers, where the index feature channel is used to select or index a specific feature channel; the convolution layer is a 3x3 convolution layer for feature extraction; and the output of the convolution layer is processed by an offset layer and a sigmoid activation function, respectively, where the sigmoid activation function is used to limit the output value to the range of [0,1], and finally the output results processed by the offset layer and the sigmoid activation function are merged as π S The output of π C It includes an average pooling layer for average pooling of input feature maps; two fully connected layers for feature processing, and a ReLU activation function is applied between the two fully connected layers; a normalization layer for normalizing features; and a bias term [1,0,0,0] for adjusting the output of the normalization layer. The introduction position of Dynamic Head is as follows Figure 4 As shown in the figure, the ability to perceive small targets is increased without increasing the computational cost. The core idea of DynamicHead lies in its dynamic convolution mechanism, which allows the model to adaptively adjust the convolution operation according to the input data to better match the characteristics of different targets. The improved head network Head not only improves the overall flexibility and expressiveness of the model, but also enhances the ability to detect targets in complex scenes.
[0053] In the neck network of the YOLOv8 network structure, for each upsampling layer, a convolutional layer is added before and after the upsampling layer and the number of channels of the convolutional layer is reduced, thereby transforming the unidirectional feature pyramid network (FPN) in the YOLOv8 network structure into a cross-scale feature fusion network (CNN-based Cross-scale Feature Fusion, CCFF), and the improved YOLOv8 network structure is called DC-YOLOv8 (DyHead and CCFF YOLOv8) network structure.
[0054] In this embodiment, if Figure 5 As shown in the figure, four layers of conv are added to the neck network part of the original YOLOv8 network structure. Specifically, a convolution layer is added before and after each upsampling layer and the number of channels of the convolution layer is changed to 256, so that the original unidirectional feature pyramid network FPN becomes a partially bidirectional cross-scale feature fusion network CCFF, so that feature information can be extracted more fully. By reducing the number of channels, the performance can be improved while achieving lightweight. Through the improvements in the above two parts, the improved YOLOv8 network structure is named DC-YOLOv8 network structure.
[0055] The DC-YOLOv8 network structure is iteratively trained using the training set, and the trained DC-YOLOv8 network structure is evaluated using the validation set. The DC-YOLOv8 network structure is optimized by adjusting the network structure parameters according to the evaluation results, and then the optimized DC-YOLOv8 network structure is tested using the test set to obtain the optimal DC-YOLOv8 network structure.
[0056] The symmetric quantization algorithm is used to symmetric quantize the optimal DC-YOLOv8 network structure to obtain a target detection model based on the DC-YOLOv8 network structure.
[0057] The process of symmetrically quantizing the optimal DC-YOLOv8 network structure by using the symmetric quantization algorithm is as follows: converting the model file of the optimal DC-YOLOv8 network structure into a model file format that meets the requirements of an actual target detection task, and symmetrically quantizing the converted optimal DC-YOLOv8 network structure, converting the network parameters in the converted optimal DC-YOLOv8 network structure from floating point numbers to signed integers, calculating the proportional factor using the maximum value of the absolute values of all weights in the network parameters, and quantizing the weights of all network parameters using the proportional factor, and then saturating the quantized weights to complete the symmetrical quantization of the optimal DC-YOLOv8 network structure, and obtaining a target detection model based on the DC-YOLOv8 network structure.
[0058] In this implementation, the model is optimized and deployed for edge devices with limited computing resources. The basic process is as follows: Figure 6 As shown in the figure, the .pt model generated by the optimal DC-YOLOv8 is converted into a .onnx model file, and in the quantization stage, the symmetric quantization algorithm is used to symmetrically quantize the .onnx model file. By compressing the model, the computing power usage is reduced without reducing the accuracy. Finally, a quantized model that meets the edge device deployment is generated, and the generated quantized model is deployed on the edge device with limited computing power resources to achieve efficient target detection. The principle of the symmetric quantization algorithm is: the conversion between the original high-precision data and the quantized int8 data is expressed as:
[0059] data float =scale×data int8
[0060] Where data float Represents the original high-precision data, usually a 32-bit floating point number; scale represents the scale factor, which is a float32 floating point number in order to be able to represent positive and negative numbers; data int8 The quantized integer data is represented by the signedint8 data type. The operation of converting the original high-precision data to int8 data is as follows, that is, the floating-point data is divided by the scale factor scale and then rounded to the nearest integer to obtain the quantized integer data:
[0061]
[0062] Where round is the rounding function.
[0063] From this, we can see that the value that the symmetric quantization algorithm needs to determine is the scale factor, that is, the quantization of weights and data can be attributed to the process of finding scale. int8 For signed numbers, the symmetry of the range of positive and negative values must be ensured. Therefore, all data are firstly subjected to the absolute value operation, so that the range of the data to be quantized is transformed into [0, data max ], where data max Indicates the maximum absolute value in the data; then determine the scale. Since the value range that int8 can represent in the positive range is [0,127], the scale can be calculated as follows:
[0064]
[0065] After the scale is determined, the corresponding representation range of int8 data is [-128×scale, 127×scale]. The quantization operation is to saturate the quantized data with [-128×scale, 127×scale]. That is, the data exceeding the range is saturated to the boundary value, and then the quantization operation shown in the formula is performed. After quantization, the model performance is evaluated using the validation set to ensure that the accuracy meets the requirements. If necessary, adjust the quantization strategy based on the evaluation results and repeat the quantization process until the model performance and accuracy meet the deployment standards. Finally, convert the quantized model into a .om model file and deploy it to the edge device.
[0066] In this embodiment, if Figure 7 As shown, the FP32 model, i.e., a model represented by 32-bit floating point numbers, is converted into a quantized model and the accuracy is evaluated. The FP32 model is converted into a quantized mode, and the model is saved. The quantized model is evaluated using ModelEvaluator to check whether its performance meets the accuracy requirements. If satisfied, the process ends, and the generated model can be used for accuracy simulation or deployment; if not satisfied, the process enters the next step; determine whether the current round is the first iteration, if it is the first iteration, perform layer sensitivity analysis, and execute Strategy according to the layer sensitivity analysis results; if it is not the first iteration, execute Strategy; wherein the layer sensitivity analysis is to analyze the sensitivity of each layer in the model to quantization, and the sensitivity values include: 0.97, 0.98, 0.99, which are used to indicate the tolerance of different layers to quantization; Strategy is to adjust the quantization strategy according to the layer sensitivity analysis results. After adjusting the strategy, quantization is performed again and the model is saved. Test experiments show that the improved DC-YOLOv8 network structure in this implementation method can be used for multi-target tracking in conjunction with the Byte Track algorithm, and the constructed target detection model has improved parameters and accuracy.
[0067] In this embodiment, the original YOLOv8 network structure is optimized according to the above three improvements to form a DC-YOLOv8 network structure, and is trained on the processed data set VisDrone2019 to obtain a .pt file of the DC-YOLOv8 target detection model. The model can detect ten types of targets from a bird's-eye view, including pedestrians, people, cars, vans, buses, trucks, motorcycles, bicycles, awning-tricycles and tricycles. In the multi-target tracking task, each frame in the subsequent video will be detected, and the obtained bounding box coordinate information will be used as the input for subsequent multi-target tracking using the Byte Track algorithm.
[0068] Step 3: Obtain the video stream or image sequence to be detected and extract continuous image frames from it, preprocess the extracted image frames, and initialize the target detection model based on the DC-YOLOv8 network structure.
[0069] The preprocessing includes: size adjustment and normalization.
[0070] In this implementation, the extracted image frame is resized to 960*960; the resized image frame is then normalized to map the pixel value from the default [0,255] range to [0,1]. The trained DC-YOLOv8 model and Byte Track algorithm configuration are loaded at the same time; DC-YOLOv8 is used to detect target objects from video streams or image sequences, and the Byte Track algorithm is responsible for managing these detection results and tracking them.
[0071] Step 4: Input the preprocessed image frame into the initialized target detection model based on the DC-YOLOv8 network structure for target detection. If the target is detected, execute step 5; if the target is not detected, use the image frame output by the target detection model to generate a video stream as the final video output.
[0072] The specific content of step 4 is: input the preprocessed image frame into the initialized target detection model based on the DC-YOLOv8 network structure, for any image frame, use the backbone network in the target detection model to extract features of the image frame, and use the extracted feature map to detect whether there is a target in the image frame, if so, generate the detection results of all targets in the image frame, including: the coordinates and size of the bounding box, the category of the target and the category confidence, and continue to perform target detection on the next image frame; if not, directly perform target detection on the next image frame until the detection of all image frames is completed; if no target is detected in all image frames, use the image frames output by the target detection model to generate a video stream as the final video output, otherwise execute step 5.
[0073] In this embodiment, for each frame input, a target detection model based on the DC-YOLOv8 network structure is used to perform target detection, and a series of bounding boxes and their corresponding category confidences are output. These detection results include the position of each detected object, i.e., the bounding box coordinates, size, and category. The categories include 10 categories, namely: pedestrian, person, car, van, bus, truck, motorcycle, bicycle, awning-tricycle, and tricycle.
[0074] In this embodiment, category confidence can help ByteTrack more accurately associate detection boxes with tracks. Although the subsequent initial matching is mainly based on IoU, combining category information and its confidence can increase the accuracy of matching, especially when two objects are very close. For example, if the IoU value of a detection box is close to that of the track, but their categories are different or the category confidence differs significantly, they may not be associated. When dealing with complex situations such as occlusion and crowding, category confidence can serve as additional information to help the system make more informed decisions. For example, when there are multiple possible matches, detection results with higher category confidence will be given priority to reduce the possibility of ID switching and ensure that the ID of the same target remains consistent throughout the video.
[0075] Step 5: Use the Byte Track algorithm to track the targets detected in all image frames and generate tracking trajectories of the tracked targets in all image frames.
[0076] In this embodiment, the content of the Byte Track algorithm is: first, a high-performance target detection model is used to obtain the target bounding box in each frame. Then, these detection results, i.e., bounding boxes, are used to match targets. In order to effectively associate targets between different frames, the Byte Track algorithm uses Intersection Over Union (IoU) for preliminary matching, and combines the Kalman filter to predict the target position to cope with fast movement or sudden changes in direction. In addition, to address the problem of short-term disappearance, the Byte Track algorithm designs a backtracking mechanism to try to find the lost target again in subsequent frames. Overall, the Byte Track algorithm emphasizes the real-time nature of the algorithm, while ensuring tracking accuracy while minimizing computational complexity, ensuring real-time tracking can be achieved under ordinary hardware environments. Due to its high efficiency and precision, the Byte Track algorithm has been widely used in intelligent monitoring, autonomous driving and other fields.
[0077] The specific content of step 5 is:
[0078] For any image frame, each target detected in the image frame is taken as a tracking target, a tracking trajectory is initialized for each tracking target in the image frame, and a unique ID is generated for each tracking trajectory; for any tracking target, a Kalman filter is initialized, an initial trajectory state is generated for the tracking target, and the tracking trajectory of the tracking target is added, wherein the initial trajectory state includes the initial position and initial velocity of the tracking target; based on the initial trajectory state of the tracking target, the initialized Kalman filter is used to predict the position and velocity of the tracking target in the current image frame, and the predicted position and velocity are used as the existing trajectory of the tracking target in the current image frame, thereby obtaining the existing trajectories of all tracking targets in the current image frame.
[0079] Each bounding box in the current image frame is preliminarily matched with the existing trajectories of all tracked targets in the current image frame in turn to generate a matching trajectory for each bounding box.
[0080] The method for preliminarily matching each bounding box in the current image frame with the existing trajectories of all tracking targets in the current image frame is as follows: for any bounding box in the current image frame, the IoU values between the bounding box and the predicted trajectories of all tracking targets in the current image frame are calculated respectively; if the IoU value between the bounding box and the existing trajectory of a tracking target in the current image frame is greater than a set IoU threshold, the bounding box is considered to be associated with the existing trajectory; otherwise, the bounding box is considered to be unassociated with the existing trajectory, and the target corresponding to the bounding box is used as a new tracking target.
[0081] In this embodiment, Byte Track receives the detection results from the target detection model and matches the newly detected target with the existing predicted trajectory. This step is based on the intersection over union (IoU) for preliminary matching. An ID is assigned to the newly detected target. This ID is unique and represents a trajectory. It serves as the "identity card" of the target. If the ID of the same target changes due to occlusion and deformation, it means that the tracking is unsuccessful. The Hungarian algorithm is used to find the optimal one-to-one matching relationship to ensure that each detected object is associated with the most appropriate trajectory as accurately as possible. The core idea of the Hungarian algorithm is to minimize the cost function to find an optimal allocation scheme. In this invention, the cost is defined based on the IoU intersection over union ratio, which represents the degree of match between the bounding box and the tracking trajectory. The IoU is set to 0.5. When the high match is higher than 0.5, the ID of the current target is used, so that when the IoU is greater than 0.5, the two bounding boxes, that is, the bounding box predicted by the Kalman filter technology and the bounding box actually detected by the target detection model, have the same ID. When it is lower, the target ID is set to a new unique value, indicating that a new target has appeared.
[0082] For all existing tracks associated with the bounding box, delete all existing tracks with different categories corresponding to the bounding box. Among all the retained existing tracks, when the difference between the IoU value of the bounding box and an existing track and the IoU threshold is within a preset range, select the matching track of the bounding box using the category confidence of the bounding box.
[0083] In this embodiment, for each bounding box in any image frame, the IoU value between it and the bounding box of the predicted position in all existing tracks in the current image frame is calculated respectively; an IoU threshold is defined to determine when a detection result is considered to be associated with a track. If the IoU value between the bounding box of a newly detected object, i.e., the tracked target, and a track exceeds this threshold, the bounding box is considered to be part of the track, i.e., the ID of the bounding box is set to be the same as the track. When the IoU is too close to the threshold, i.e., the IoU value between the bounding box and an existing track is within the preset range of 0.5±0.1, the category confidence is used for association analysis, i.e., in matching, the category confidence can help Byte Track more accurately associate the bounding box with the track.
[0084] The method of selecting the matching track of the bounding box using the category confidence of the bounding box is: using the category confidence of the bounding box and the IoU value between the bounding box and the existing track, respectively calculating the similarity score between each existing track associated with the bounding box and the bounding box, and selecting the existing track with the highest similarity score as the matching track of the bounding box.
[0085] The similarity score is expressed as:
[0086] S total =w 1 ×IoU+w 2 ×confidence
[0087] Where S total represents the similarity score between the existing track and the bounding box; w 1 and w 2 Both represent weights. In this implementation, w 1 =0.6, w 2 =0.4; IoU represents the IoU value between the bounding box and the existing track; confidence represents the category confidence of the bounding box, that is, the confidence score of the model that the detected object belongs to a specific category, which is a value between 0 and 1.
[0088] In this embodiment, S total Then compare it with other bounding boxes with IoU or S greater than the IoU threshold totalThe values are compared and the bounding box with the maximum value is taken as the matching object.
[0089] The Kalman filter is used to update the matching trajectory of each bounding box respectively, and the position and speed of the tracking target corresponding to the matching trajectory of each bounding box are corrected, and the updated matching trajectory is added to the tracking trajectory of the tracking target; at the same time, the updated matching trajectory is used as the existing trajectory of the tracking target for target tracking in the next image frame.
[0090] In this embodiment, the Kalman filter is used to predict the position and speed of the target at the next moment, and adjust its trajectory in combination with the actual detection results. The workflow is divided into two main steps: prediction and update. These two steps are continuously iterated to produce the optimal state estimate. In the prediction stage, the Kalman filter predicts the position at the current moment based on the state estimate and control input at the previous moment. In the update stage, the bounding box at the current moment is used to correct the prediction result to obtain a more accurate state estimate. When a new frame arrives, the predicted bounding box position is compared with the actual detected bounding box position, and the predicted value is corrected using the Kalman filter to obtain a more accurate target position estimate and speed.
[0091] A death list with a capacity of N is initialized, and each time a preliminary match is completed between each bounding box in an image frame and the existing trajectories of all tracking targets in the image frame, the existing trajectories in the image frame that have not been successfully matched are added to the death list, and while a preliminary match is performed between each bounding box in the next image frame and the existing trajectories of all tracking targets in the image frame, a secondary match is performed between each existing trajectory in the death list and each bounding box in the next image frame, and the existing trajectories that have successfully been matched for the secondary match are removed from the death list; if any existing trajectory remains in the death list after N consecutive frames, the tracking target corresponding to the existing trajectory is regarded as a missing target; wherein the secondary matching method is the same as the method for performing a preliminary match between each bounding box in the current image frame and the existing trajectories of all tracking targets in the current image frame.
[0092] In this implementation, the value of N is generally less than or equal to 30, and the higher the number of frames, the higher the computing power required and the better the tracking effect. Therefore, in this implementation, N is set to 30, that is, if a target is not detected within 30 frames, ByteTrack will temporarily retain the target's information for a period of time and try to find it again. If it is still not found after exceeding the set time threshold, it is considered that the target has left the scene and its track is terminated.
[0093] In this embodiment, even if the target is lost, its trajectory information will still be calculated and saved for a period of time to facilitate possible rematching later. When a new frame arrives, not only will the new bounding box be matched for the first time, but also these new bounding boxes will be tried to match with the previously lost targets, that is, if the tracking trajectory of a certain tracking target is not matched successfully within N frames, Byte Track retains the information of the tracking trajectory and adds the tracking trajectory to the death list. The death list contains the trajectory information of the target that has not been detected in the past N frames, including ID, position and speed. After each frame, the death list will be refreshed to the trajectory information of the target that has not been detected in the latest N frames. If it exceeds N frames, the target is considered to be missing. The target that failed to match the first time can be matched twice in the "death list". The bounding box that successfully matched the second time is assigned the same ID as the trajectory information that appeared in the "death list" and is considered to be the same target. If the corresponding ID is not matched in the "death list", a new unique ID is generated for the target. This solves the ID switching problem caused by the occlusion of the target to a certain extent. If the lost target is successfully found again, the state of the target, including position and speed, will be updated according to the latest detection results and tracking will continue. If it is not found, the trajectory of the target will be terminated after a certain period of time, i.e. the time threshold mentioned above.
[0094] Step 6: Generate a video stream based on the tracking trajectory of the tracked target in all image frames and output it as the final video.
[0095] In this embodiment, the tracking information contained in the tracking tracks of all tracked targets in each frame output by Byte Track is: ID, position, and speed. Such tracking information is usually used for subsequent applications, such as behavior analysis, traffic statistics, etc.
[0096] In this embodiment, the operation process of configuring the target detection model based on the DC-YOLOv8 network structure on the edge device is as follows Figure 8 As shown. Compared with the prior art, this method has higher accuracy, fewer parameters, and lower computing power consumption when deploying target tracking systems for edge devices. Experiments show that under the same experimental operating environment, the deep learning framework used is PyTorch, the operating system is Ubuntu22.04, and the graphics card configuration is NVIDIA RTX A4000 GPU. The number of parameters of the multi-target tracking system deployed based on the target detection model based on the DC-YOLOv8 network structure in this implementation method is reduced by about 21.5%, and the amount of calculation is reduced by about 615MFLOPs. The ablation experiment results are shown in Table 1.
[0097] Table 1 DC-YOLOv8 ablation experiment results
[0098]
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm, characterized in that: The method comprises the following steps: Step 1: Obtain several image data and convert their formats to construct an image dataset; Step 2: Build a DC-YOLOv8 network structure and use the image data set for training, and then use the optimal DC-YOLOv8 network structure obtained through training to build a target detection model based on the DC-YOLOv8 network structure; Step 3: Obtain the video stream or image sequence to be detected and extract continuous image frames from it, preprocess the extracted image frames, and initialize the target detection model based on the DC-YOLOv8 network structure; Step 4: Input the preprocessed image frame into the initialized target detection model based on the DC-YOLOv8 network structure for target detection. If the target is detected, execute step 5; If no target is detected, the image frames output by the target detection model are used to generate a video stream as the final video output; Step 5: Use the Byte Track algorithm to track the targets detected in all image frames and generate tracking trajectories of the tracked targets in all image frames; Step 6: Generate a video stream based on the tracking trajectory of the tracked target in all image frames and output it as the final video.
2. According to claim 1, a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm is characterized in that: The specific content of step 1 is: obtaining a number of image data, wherein each image data includes: an image and annotation information of the image; the annotation information includes: bounding box information and category; converting the annotation information of all image data into YOLO format to ensure that each image file includes a corresponding .txt file, thereby constructing an image data set; and then dividing the image data set into a training set, a validation set, and a test set according to a preset ratio.
3. According to claim 2, a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm, characterized in that: The specific content of step 2 is: improving the existing YOLOv8 network structure; the improvement is: replacing the detection head in the head network of the YOLOv8 network structure with a dynamic detection head Dynamic Head; In the neck network of the YOLOv8 network structure, for each upsampling layer, a convolution layer is added before and after the upsampling layer and the number of channels of the convolution layer is reduced, thereby transforming the unidirectional feature pyramid network FPN in the YOLOv8 network structure into a cross-scale feature fusion network CCFF, and the improved YOLOv8 network structure is called the DC-YOLOv8 network structure; The DC-YOLOv8 network structure is iteratively trained using the training set, and the trained DC-YOLOv8 network structure is evaluated using the validation set. The DC-YOLOv8 network structure is optimized by adjusting the network structure parameters according to the evaluation results, and then the optimized DC-YOLOv8 network structure is tested using the test set to obtain the optimal DC-YOLOv8 network structure. The symmetric quantization algorithm is used to symmetric quantize the optimal DC-YOLOv8 network structure to obtain a target detection model based on the DC-YOLOv8 network structure.
4. According to claim 3, a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm is characterized in that: The dynamic detection head Dynamic Head includes a series of π L , π S and π C Three parts; among them π L It includes the following: average pooling layer, convolution layer, ReLU activation function and hard sigmoid activation function connected in series, and multiplying the output of the hard sigmoid activation function by the feature map input in the average pooling layer element by element to obtain π L The output of S It includes the index feature channel and convolution layer connected in series, and the output of the convolution layer is processed by the offset layer and the sigmoid activation function respectively and then merged as π S The output of C It includes the following series: average pooling layer, two fully connected layers and normalization layer; a ReLU activation function is added between the two fully connected layers, and a bias term is set to adjust the output of the normalization layer, and the adjusted output is used as π C Output.
5. According to claim 4, a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm is characterized in that: The process of symmetrically quantizing the optimal DC-YOLOv8 network structure by using the symmetric quantization algorithm is as follows: converting the model file of the optimal DC-YOLOv8 network structure into a model file format that meets the requirements of an actual target detection task, and symmetrically quantizing the converted optimal DC-YOLOv8 network structure, converting the network parameters in the converted optimal DC-YOLOv8 network structure from floating point numbers to signed integers, calculating the proportional factor using the maximum value of the absolute values of all weights in the network parameters, and quantizing the weights of all network parameters using the proportional factor, and then saturating the quantized weights to complete the symmetrical quantization of the optimal DC-YOLOv8 network structure, and obtaining a target detection model based on the DC-YOLOv8 network structure.
6. According to claim 5, a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm is characterized in that: The specific content of step 4 is: input the preprocessed image frame into the initialized target detection model based on the DC-YOLOv8 network structure, for any image frame, use the backbone network in the target detection model to extract features of the image frame, and use the extracted feature map to detect whether there is a target in the image frame, if so, generate the detection results of all targets in the image frame, including: the coordinates and size of the bounding box, the category of the target and the category confidence, and continue to perform target detection on the next image frame; if not, directly perform target detection on the next image frame until the detection of all image frames is completed; if no target is detected in all image frames, use the image frames output by the target detection model to generate a video stream as the final video output, otherwise execute step 5.
7. The multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm according to claim 6, characterized in that: The specific content of step 5 is: for any image frame, each target detected in the image frame is taken as a tracking target, a tracking trajectory is initialized for each tracking target in the image frame, and a unique ID is generated for each tracking trajectory; for any tracking target, a Kalman filter is initialized and an initial trajectory state is generated for the tracking target and added to the tracking trajectory of the tracking target, wherein the initial trajectory state includes an initial position and an initial velocity of the tracking target; Based on the initial trajectory state of the tracking target, the initialized Kalman filter is used to predict the position and speed of the tracking target in the current image frame, and the predicted position and speed are used as the current trajectory of the tracking target in the current image frame, thereby obtaining the current trajectories of all tracking targets in the current image frame; Preliminarily matching each bounding box in the current image frame with the existing trajectories of all tracked targets in the current image frame in turn to generate matching trajectories of each bounding box; The Kalman filter is used to update the matching trajectory of each bounding box, and the position and speed of the tracking target corresponding to the matching trajectory of each bounding box are corrected, and the updated matching trajectory is added to the tracking trajectory of the tracking target; at the same time, the updated matching trajectory is used as the existing trajectory of the tracking target for tracking the target in the next image frame; Initialize a death list with a capacity of N. After each preliminary match between each bounding box in an image frame and the existing tracks of all tracked targets in the image frame, add the existing tracks in the image frame that have not been successfully matched to the death list. While performing preliminary matches between each bounding box in the next image frame and the existing tracks of all tracked targets in the image frame, perform secondary matches between each existing track in the death list and each bounding box in the next image frame, and remove the existing tracks that have been successfully matched twice from the death list. If any existing track remains in the death list after N consecutive frames, the tracking target corresponding to the existing track is regarded as a missing target; wherein the secondary matching method is the same as the method of performing preliminary matching between each bounding box in the current image frame and the existing tracks of all tracking targets in the current image frame.
8. According to claim 7, a multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm is characterized in that: The method for preliminarily matching each bounding box in the current image frame with the existing trajectories of all tracked targets in the current image frame is as follows: for any bounding box in the current image frame, the IoU values between the bounding box and the predicted trajectories of all tracked targets in the current image frame are calculated respectively; if the IoU value between the bounding box and the existing trajectory of a tracked target in the current image frame is greater than a set IoU threshold, the bounding box is considered to be associated with the existing trajectory; otherwise, the bounding box is considered to be unassociated with the existing trajectory, and the target corresponding to the bounding box is used as a new tracked target; For all existing tracks associated with the bounding box, delete all existing tracks with different categories corresponding to the bounding box. Among all the retained existing tracks, when the difference between the IoU value of the bounding box and an existing track and the IoU threshold is within a preset range, select the matching track of the bounding box using the category confidence of the bounding box.
9. The multi-target tracking method based on the YOLOv8 model and the Byte Track algorithm according to claim 8, characterized in that: The method of selecting the matching track of the bounding box by using the category confidence of the bounding box is: using the category confidence of the bounding box and the IoU value between the bounding box and the existing track, respectively calculating the similarity score between each existing track associated with the bounding box and the bounding box, and selecting the existing track with the highest similarity score as the matching track of the bounding box; The similarity score is expressed as: S total =w1×IoU+w2×confidence Where S total represents the similarity score between the existing track and the bounding box; w1 and w2 both represent weights; IoU represents the IoU value between the bounding box and the existing track; confidence represents the category confidence of the bounding box.
Citation Information
Patent Citations
Traffic flow statistical method based on YOLOv8 algorithm and related equipment
CN119049305A
Cited By
Embedded video target detection tracking system and method
CN120526358A
Tracking result generation method and device based on YOLO and MixFormer models, equipment and medium
CN121190524A
Tracking result generation method and device based on YOLO and MixFormer model, equipment and medium
CN121190524B
Face recognition system and method for construction site personnel training and on-duty dynamic verification
CN121438363A
Dynamic screening and trajectory tracking method and system for accurate evaluation of sperm motility
CN121861451A