DeepSort-based multi-target tracking method, device, equipment and medium
By improving the DeepSort algorithm with the optimization of the YOLOv8 model and the EIoU loss function, the problem of small and medium-sized target recognition and occlusion of multi-objective tracking is solved, and high-precision and real-time target tracking effect is achieved.
Patent Information
- Application Number
- CN202510341433.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing DeepSort multi-objective tracking algorithm has limitations in dealing with target occlusion, fast motion, and target overlap, resulting in insufficient tracking accuracy and robustness, especially in small target recognition and tracking.
Multi-objective detection is performed using the improved YOLOv8 model, combining dynamic feature pyramid fusion and deformable convolution, bounding box regression is optimized through the EIoU loss function, and target matching is combined with the Hungarian algorithm, and one-to-one matching and trajectory update are performed using the DeepSort algorithm.
It significantly improves the detection accuracy of small targets, reduces bounding box regression error, improves target matching and tracking speed, enhances robustness in fast motion or occlusion, and achieves accurate identification and real-time tracking of small targets.
Smart Images

Figure CN120279060A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target tracking, and particularly relates to a multi-target tracking method, device, equipment and medium based on DeepSort. Background Art
[0002] Target detection and tracking technologies are important research directions in the field of computer vision and are widely applied in multiple scenarios such as security monitoring, autonomous driving, and intelligent retail. With the rapid development of deep learning technologies, target detection and tracking algorithms have gradually shifted from methods based on handcrafted features to methods based on deep learning. The powerful feature expression ability of deep learning has significantly improved the accuracy and robustness of target detection and tracking.
[0003] As a real-time target detection algorithm, YOLO has undergone multiple iterations and improvements since its proposal in 2015, continuously achieving breakthroughs in terms of speed and accuracy. For example, YOLOv8 significantly optimizes the feature extraction ability and detection performance by adopting advanced backbone networks and neck architectures; while YOLOv11 further improves the accuracy and efficiency while reducing the number of parameters. However, in multi-target tracking scenarios, YOLO still faces some challenges, such as target ID switching and occlusion problems. DeepSort is a multi-target tracking algorithm based on deep learning, which combines the traditional SORT algorithm and deep feature extraction technologies. DeepSort obtains the target information in each frame of image through a target detector (such as YOLO), and then uses the Kalman filter and Hungarian algorithm to complete target matching and trajectory prediction. Nevertheless, DeepSort still has certain limitations in dealing with target occlusion, fast movement, and target overlap. In addition, the deficiencies of traditional DeepSort in feature fusion and target matching algorithms also limit the further improvement of tracking accuracy and robustness. Therefore, there is an urgent need to develop a high-precision and real-time target tracking method to achieve precise recognition and efficient tracking of small targets. This will provide a new technical path and application value for solving the key problems in multi-target tracking. Summary of the Invention
[0004] In view of the deficiencies in the prior art, the present invention provides a multi-target tracking method, device, equipment and medium based on DeepSort, which can achieve precise recognition and real-time tracking of small targets and improve the detection accuracy.
[0005] The present invention provides the following technical solutions:
[0006] In the first aspect, a multi-target tracking method based on DeepSort is provided, including the following steps:
[0007] S1: Use the improved YOLOv8 model to perform multi-object detection on each frame of the collected video stream, and output multiple candidate bounding boxes, class labels, and class confidence scores for each object. Among them, the improved YOLOv8 model incorporates a dynamic detection head in the detection head and performs feature fusion by adding a cross-scale dynamic gating mechanism in the feature pyramid;
[0008] S2: Filter the candidate bounding boxes through a double screening method of confidence threshold and class to obtain valid target boxes, and use the valid target boxes as boundaries to crop the ROI image from the original input image;
[0009] S3: Use the DeepSort algorithm to extract appearance feature vectors from the cropped ROI image, and predict the state vectors of all current tracking objects in the current frame, and convert them into prediction boxes;
[0010] S4: Based on the parameters of the prediction boxes and valid target boxes, as well as the extracted and tracked appearance feature vectors, use the DeepSort algorithm to perform one-to-one matching on all valid target boxes and prediction boxes. If the matching is successful, update the trajectory information of the corresponding object and update the state vector and appearance features of the DeepSort algorithm. If the matching fails, initialize the unmatched valid target boxes as new tracking objects or mark the unmatched prediction boxes as lost objects.
[0011] Optionally, step S1 specifically includes:
[0012] S1.1: Perform frame-by-frame preprocessing on the collected video stream; the preprocessing includes one or more of dynamic size normalization, introducing a dynamic padding strategy, pixel normalization, and adaptive image enhancement;
[0013] S1.2: Embed a dynamic convolutional layer in the detection head, and dynamically fuse multi-level features through a cross-scale dynamic gating mechanism in the feature pyramid to improve the YOLOv8 model; during the training process of the improved YOLOv8 model, use a clustering algorithm to generate anchor box sizes for small objects. In the inference stage, adaptively adjust the anchor box density of different feature layers according to the feature map resolution in advance;
[0014] S1.3: Use the improved YOLOv8 model to perform multi-object detection on each preprocessed frame of image, and output multiple candidate bounding boxes, class labels, and class confidence scores for each object.
[0015] Optionally, step S2 specifically includes:
[0016] S2.1: Adaptively set the confidence threshold according to the scene complexity of the current frame to perform the first filtering on all candidate bounding boxes;
[0017] S2.2: Perform category filtering on the candidate bounding boxes retained after the first filtering, retain the candidate bounding boxes of the set focus categories, and use the retained candidate bounding boxes as valid target boxes;
[0018] S2.3: Perform inverse normalization on the output coordinates of the valid target boxes to obtain the coordinates x gt , y gt , w gt , h gt ;
[0019] S2.4: Use the absolute coordinates of the valid target boxes as the boundaries of the cropping region, and extract the pixel data of the current cropping region from the original input image to generate an ROI image.
[0020] Optionally, step S3 specifically includes:
[0021] S3.1: Use the DeepSort algorithm to perform normalization processing on the cropped ROI image, and use a pre-trained convolutional neural network to extract the appearance feature vectors of all targets;
[0022] S3.2: Associate the appearance feature vectors of the ROI image with the valid target boxes;
[0023] S3.3: Based on the historical frame state vectors of all targets being tracked, the DeepSort algorithm uses the Kalman filter corresponding to the target to predict the state vector x k = [x pred , y pred , w pred , h pred , v x , v y T ;
[0024] S3.4: Ignore the motion speeds v x and v y of the target in the x and y directions, and deform the state vector x k of all current targets in the current frame into a prediction box (x pred , y pred , w pred , h pred ).
[0025] Optionally, step S4 specifically includes:
[0026] S4.1: Use the EIoU cost to perform an initial matching of the prediction boxes and the valid target boxes to obtain a first candidate matching pair combination that meets the initial matching threshold;
[0027]
[0028] Among them, IoU is the interaction ratio between the effective target box and the predicted box, and ρ 2 (b pred , b gt ) is the deviation degree between the center points of the effective target box and the predicted box, and ρ 2 (w pred , w gt ) is the deviation degree between the effective target box and the predicted box in terms of width, and ρ 2 (h pred , h gt ) is the deviation degree between the effective target box and the predicted box in terms of height, x pred , y pred , w pred , h pred are respectively the center point coordinates, width, and height of the predicted box; x gt , y gt , w gt , h gt are respectively the center point coordinates, width, and height of the effective target box; c, c w and c h are respectively the diagonal length, width, and height of the minimum bounding rectangle of the effective target box and the predicted box; ρ 2 is the square of the Euclidean distance;
[0029] S4.2: For the effective target boxes that are not successfully matched during the initial matching process, perform secondary matching to obtain a second candidate matching pair combination that meets the secondary matching threshold. The method of secondary matching is to compare the similarity CosineSimilarity between the tracking appearance feature vector f1 of the DeepSort algorithm and the appearance feature vector f2 corresponding to the ineffective target box that has not been successfully matched;
[0030]
[0031] S4.3: Construct a bipartite graph from the first candidate matching pair combination and the second candidate matching pair combination, and use the Hungarian algorithm to solve the optimal matching scheme. In the optimal matching scheme, if the effective target box and the detection box correspond one by one, the matching is successful, and use the successfully matched detection box to update the trajectory information, state vector, and appearance features of the corresponding target. Otherwise, the matching fails, and the unmatched effective target boxes are marked as new targets, and the unmatched predicted boxes are retained or marked as lost.
[0032] Optionally, in step S4.3, in the process of using the Hungarian algorithm to solve the optimal matching scheme, use the EIoU cost of the initial matching and the similarity of the secondary matching to construct a cost matrix, and the goal of using the Hungarian algorithm to solve the optimal matching scheme is to minimize the total cost of the cost matrix.
[0033] Optionally, step S4.3 includes updating the state vector and appearance features of the DeepSort algorithm using the successfully matched detection box, specifically:
[0034] S4.3.1: For the successfully matched predicted box and valid target box, obtain the observation vector z of the current Kalman filter according to the parameters of the valid target box k ;
[0035] z k = [x measured , y measured , w measured , h measured T
[0036] where x measured , y measured , w measured , h measured are the center point coordinates, width, and height of the successfully matched valid target box;
[0037] S4.3.2: Use the Kalman gain K and the observation vector z of the current Kalman filter k to update the state vector of the current Kalman filter;
[0038]
[0039] where, is the updated state vector, and x k is the current state vector predicted by the Kalman filter;
[0040] S4.3.3: For the successfully matched predicted box and valid target box, update the appearance feature vector f2 extracted from the ROI image corresponding to the valid target box to the tracking appearance feature vector of the current tracking target of the DeepSort algorithm.
[0041] In a second aspect, a multi-object tracking system based on DeepSort is provided, including:
[0042] An object detection module, which is used to perform multi-object detection on each frame of the collected video stream using an improved YOLOv8 model, and output multiple candidate bounding boxes, class labels, and class confidences for each object. Among them, the improved YOLOv8 model incorporates a dynamic detection head in the detection head and performs feature fusion by adding a cross-scale dynamic gating mechanism in the feature pyramid;
[0043] An ROI image acquisition module, which is used to filter the candidate bounding boxes by means of double screening of the confidence threshold and category, obtain valid target boxes, and crop the ROI image from the original input image using the valid target boxes as the boundaries;
[0044] A prediction module, configured to extract an appearance feature vector from the cropped ROI image using the DeepSort algorithm, predict the state vectors of all current tracking targets in the current frame, and convert them into prediction bounding boxes;
[0045] A target matching and tracking module, configured to perform one-to-one matching on all valid target bounding boxes and prediction bounding boxes using the DeepSort algorithm based on the parameters of the prediction bounding boxes and valid target bounding boxes, and the extracted and tracked appearance feature vectors. If the matching is successful, update the trajectory information of the corresponding target and update the state vector and appearance features of the DeepSort algorithm. If the matching fails, initialize the unmatched valid target bounding boxes as new tracking targets or mark the unmatched prediction bounding boxes as lost targets.
[0046] In a third aspect, a computer device is provided, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the multi-object tracking method based on DeepSort according to any item in the first aspect are implemented.
[0047] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program; when the computer program is executed by a processor, the steps of the multi-object tracking method based on DeepSort according to any item in the first aspect are implemented.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] (1) The detection accuracy of small targets in the present invention is significantly improved, and the regression error of the bounding box is greatly reduced. By introducing Dynamic Feature Pyramid Network (Dynamic FPN) and Deformable Convolution, the detail capture ability of the shallow features (P2 layer) is enhanced, and the bounding box regression is optimized by combining with the EIoU loss function. In DeepSort, EIoU is used to replace IoU to calculate the matching cost matrix, and it is optimized in combination with the Hungarian algorithm. Compared with the traditional method, it can find the best matching scheme more efficiently, reduce unnecessary computational complexity, improve the speed of target matching and tracking, and enable the entire system to process video stream data in real time.
[0050] (2) By introducing penalty terms for width and height, EIoU can more accurately reflect the geometric differences of the bounding boxes; it can accelerate convergence during the training process and improve the efficiency of target detection and tracking; especially when the target is moving fast or occluded, it shows higher robustness in tracking small targets.
[0051] (3) In terms of object tracking, existing technologies often rely solely on a single feature for object matching and trajectory prediction, and are prone to losing the object when the object's appearance changes or is occluded. In the present invention, DeepSort not only extracts the apparent features (Re-ID features) of the object, but also combines motion features for data association. When calculating the similarity of objects, the traditional IoU only considers the overlapping area of the bounding boxes and cannot comprehensively reflect the true relationship between objects. The present invention introduces the EIoU distance metric method, which comprehensively considers multi-dimensional factors such as the geometric shape, size, center point distance, and side length difference of the object. Description of the Drawings
[0052] Figure 1 is a flowchart of the steps of the multi-object tracking method based on DeepSort of the present invention;
[0053] Figure 2 is a network model framework diagram of the improved YOLOv8 model of the present invention;
[0054] Figure 3 is a schematic diagram of the structure of the multi-object tracking system based on DeepSort of the present invention. Detailed Embodiments
[0055] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention. It should be noted that the terms "including" and any variations thereof in the specification and claims of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0056] Embodiment 1
[0057] As Figure 1 shown, a multi-object tracking method based on DeepSort includes the following steps:
[0058] S1: Use the improved YOLOv8 model to perform multi-object detection on each frame of the collected video stream, and output multiple candidate bounding boxes, class labels, and class confidences for each object. Among them, the improved YOLOv8 model incorporates a dynamic detection head in the detection head and performs feature fusion by adding a cross-scale dynamic gating mechanism to the feature pyramid.
[0059] Step S1 specifically includes:
[0060] S1.1: Perform frame-by-frame preprocessing on the collected video stream; the preprocessing includes one or more of dynamic size normalization, introducing a dynamic padding strategy, pixel normalization, and adaptive image enhancement.
[0061] The collected video stream can be a real-time video stream or an offline video file. After these input data are input through a camera or a storage device, they are parsed frame by frame through OpenCV or FFmpeg to ensure that the image format and quality meet the input requirements of the YOLO model, thereby improving the accuracy and efficiency of detection.
[0062] Specifically, the steps of the preprocessing include:
[0063] S1.1.1: Dynamic size normalization: To adapt to the input requirements of the YOLO model, each frame of the image needs to be scaled to a fixed size. By default, the input image is scaled to 640 pixels for the long side according to the original width-to-height ratio, and the short side is scaled proportionally to avoid stretching and deformation. This process is called dynamic size normalization, which ensures that input images of different resolutions have a consistent size before entering the model, thus avoiding fluctuations in detection performance caused by differences in image size. By unifying the image size, the model can process various input data more stably, improving the accuracy and consistency of detection.
[0064] S1.1.2: Adopt a dynamic padding strategy to automatically calculate the minimum padding area, and the padding value is optionally set to 114, which is the background value of the YOLO standard.
[0065] S1.1.3: Dynamic pixel normalization: The pixel values of an image are usually in the range of 0 to 255. To enable the improved yolov8 model to have a unified input standard when processing different images, the pixel values need to be normalized to the range [0,1]. The normalization formula is:
[0066]
[0067] where I raw is the original pixel value, and I norm is the pixel value after normalization.
[0068] Through the normalization process, the improved yolov8 model can process image data more effectively, reducing calculation errors caused by differences in pixel value ranges, thereby improving the stability and accuracy of detection.
[0069] S1.1.4: Adaptive Image Enhancement: In practical applications, the input images may be affected by adverse conditions such as low light or blur, thereby reducing the performance of object detection. To improve the image quality and enhance the robustness of the model, image enhancement techniques can be adopted. Specific methods include: Histogram Equalization: By adjusting the histogram distribution of the image, the contrast of the image is enhanced, making the details in the image more clearly visible. This method is particularly suitable for images under low light conditions and can effectively improve the detectability of the image. Gaussian Filtering: The image is smoothed through a Gaussian filter to remove noise and blur in the image, thereby improving the clarity and quality of the image. Gaussian filtering is a commonly used image smoothing technique that can reduce the interference of noise while retaining the main features of the image. Low Light Enhancement: If the average brightness of the detected image is lower than 50 (in the range of 0 - 255), contrast-limited adaptive histogram equalization is automatically triggered.
[0070] S1.2: A dynamic convolutional layer is embedded in the detection head, and multi-level features are dynamically fused through a cross-scale dynamic gating mechanism in the feature pyramid to improve the YOLOv8 model; during the training process of the improved YOLOv8 model, a clustering algorithm is used to generate anchor box sizes for small objects. In the inference stage, the anchor box density of different feature layers is adaptively adjusted according to the feature map resolution in advance.
[0071] Specifically, as Figure 2 shown, the improved YOLOv8 model is based on the YOLOv8 model. A dynamic convolutional layer is embedded in the detection head of the basic YOLOv8 model to form a dynamic detection head (DynamicHead). Deformable convolutions are embedded in the regression branch of the detection head to dynamically adjust the sampling positions of the convolution kernels and enhance the ability to capture the edge features of small objects.
[0072] The dynamic detection head (DynamicHead) applies different attention mechanisms in the hierarchical, spatial, and channel dimensions: Scale-aware Attention: Dynamically fuse features of different scales. Spatial-aware Attention: Focus on discriminative regions in spatial positions through deformable convolutions. Task-aware Attention: Dynamically switch feature channels to support different tasks, such as classification and bounding box regression.
[0073] The dynamic detection head adaptively adjusts the convolution kernel weights according to the input features to enhance the ability to capture the edge features of small objects. The specific formula is:
[0074] W dynamic =softmax(MLP(F input )).W base
[0075] where, W dynamicis the weight of the dynamic convolution kernel, MLP is the adjustment term for dynamically generating the convolution kernel weight according to the input features, and F input is the input feature, and W base is the weight of the basic convolution kernel.
[0076] In the Feature Pyramid Network (FPN), Cross-Scale Dynamic Gate is added to dynamically fuse the P2-P5 features. Cross-Scale Dynamic Gate can refer to the existing technology, and the formula for dynamic feature fusion is:
[0077]
[0078] where F fused is the multi-scale feature map after dynamic fusion, F i are the feature maps at different levels in the feature pyramid, W g are learnable parameters, and GAP is global average pooling.
[0079] Dynamic gates (such as the update gate in GRU and the forget gate in LSTM) dynamically adjust the fusion weights of features at different scales through learnable parameters. Its essence is an adaptive feature selector, which determines the utilization rate of features at each scale in real time according to the characteristics of the input data (such as the target scale, texture complexity). FPN generates P2-P5 features through Bottom-Up and Top-Down paths, capturing high-resolution details and deep semantic information respectively.
[0080] During the training process of the improved YOLOv8, the K-means++ algorithm is used to cluster the widths and heights of the targets in the training dataset to generate anchor box sizes that are more suitable for small targets, such as 4×5, 8×10. In the inference stage, the anchor box density of different feature layers is adaptively adjusted according to the feature map resolution in advance. For example, more small anchor boxes are assigned to high-resolution features.
[0081] S1.3: Use the improved YOLOv8 model to perform multi-object detection on each preprocessed frame of the image, and output multiple candidate bounding boxes, class labels, and class confidences for each target.
[0082] The parameters of the candidate bounding box include coordinate parameters (x1, y1, x2, y2). (x1, y1) is the upper left coordinate of the candidate bounding box, and (x2, y2) is the lower right coordinate of the candidate bounding box. These coordinates are absolute coordinates relative to the original image. Through these coordinates, the position of the target in the image can be accurately located.
[0083] Class confidence (score): The class confidence represents the probability that the object within the detection box belongs to a certain class. It is calculated by the product of the object confidence (indicating the probability that there is an object within the detection box) and the class confidence (indicating the probability that the object belongs to a specific class). This confidence reflects the reliability of the detection result, and the way to obtain the confidence refers to the existing technology. Class label (class_id): The class label is used to identify the class to which the detected object belongs. For example, the class ID corresponding to "target" is 0. Through the class label, the target classes of interest can be quickly filtered out.
[0084] S2: Filter the candidate bounding boxes through a dual screening method of confidence threshold and class, obtain the effective target boxes, and use the effective target boxes as boundaries to crop the ROI image from the original input image.
[0085] Step S2 specifically includes:
[0086] S2.1: Adaptively set the confidence threshold according to the scene complexity of the current frame to perform the first filtering on all candidate bounding boxes.
[0087] The first filtering is to filter the candidate bounding boxes with low-quality predictions. The way to adaptively set the confidence threshold is specifically: statistically calculate the average confidence of the objects in the continuous N frames, and update the threshold in real time. The update formula is:
[0088] conf_threshold = max(0.5, min(0.95, μ score - 0.1))
[0089] conf_threshold is the updated threshold, and μ score is the average confidence of all valid objects in the previous N frames.
[0090] S2.2: Perform class filtering on the candidate bounding boxes retained after the first filtering, retain the candidate bounding boxes of the set concerned classes, and use the retained candidate bounding boxes as the effective target boxes.
[0091] Specifically, the way of class filtering is to introduce dynamic semantic attention (DynamicSemantic Attention) in the class prediction branch and enhance the response of the channels related to the target class through the SE module (Squeeze-and-Excitation).
[0092] Filtering through the class label (class_id) can ensure that the objects for subsequent processing are the classes of interest to the user.
[0093] Through double screening based on confidence and category, the accuracy and robustness of object detection can be effectively improved. The confidence threshold filtering ensures the reliability of the detection bounding boxes, while the category filtering ensures the pertinence of the detection results. This double screening mechanism not only reduces false detections but also improves the detection efficiency of the system for specific objects.
[0094] S2.3: Perform inverse normalization on the output coordinates of the valid object bounding boxes to obtain the coordinates x of the valid object bounding boxes in the original input image gt , y gt , w gt , h gt .
[0095] In the object detection stage, the bounding box coordinates output by the YOLOv8 model are usually normalized values, that is, the coordinate range is between [0,1]. To restore these normalized coordinates to the size of the original image, coordinate inverse normalization processing is required. The specific steps are as follows: Obtain the width (W) and height (H) of the original image. For the coordinates (x1, y1, x2, y2) of each bounding box, convert them from normalized coordinates to absolute coordinates: The coordinates of the bounding box are restored to the size of the original image, thus ensuring the accuracy of subsequent processing.
[0096] S2.4: Use the absolute coordinates of the valid object bounding boxes as the boundaries of the cropping region, and extract the pixel data of the current cropping region from the original input image to generate an ROI image.
[0097] After completing the coordinate inverse normalization, the region of interest (ROI) of the target is cropped from the original image according to the coordinates of the bounding box. Specifically, use the inverse normalized coordinates of the bounding box as the boundaries of the cropping region. Extract the pixel data of this region from the original image to generate an ROI image. The cropped ROI image will be used as the input of the DeepSort tracking module for subsequent object matching and trajectory prediction.
[0098] S3: Use the DeepSort algorithm to extract the appearance feature vectors from the cropped ROI image, and predict the state vectors of all current tracking objects in the current frame and convert them into prediction bounding boxes.
[0099] Step S3 specifically includes:
[0100] S3.1: Use the DeepSort algorithm to perform normalization processing on the cropped ROI image, and use a pre-trained convolutional neural network to extract the appearance feature vectors of all objects.
[0101] Specifically, before extracting the appearance feature vectors, it is necessary to perform normalization processing on the cropped region of interest (ROI) image to ensure that the images input into the Re-ID model have consistent sizes and formats. The specific operations are as follows:
[0102] Image Scaling: Scale the cropped ROI image to a fixed resolution. In the present invention, 256×128 pixels are adopted. This size has been verified through experiments and can balance the computational efficiency and the accuracy of feature extraction.
[0103] Aspect Ratio Preservation: During the scaling process, in order to prevent the distortion of the target's shape, it is necessary to preserve the aspect ratio of the image. If the aspect ratio of the original ROI image is inconsistent with the target size, the image ratio can be adjusted by padding to ensure that the scaled image meets the model input requirements.
[0104] Normalization: Normalize the pixel values of the scaled image, mapping the pixel value range from [0, 255] to [0, 1] to improve the generalization ability and stability of the model.
[0105] After completing the standardization process of the ROI image, use a pre-trained convolutional neural network (such as OSNet) to extract the appearance features of the target. The specific steps are as follows:
[0106] Input the standardized ROI image into OSNet, and the network will output a feature vector with a fixed dimension. In the present invention, the dimension of the feature vector is 128. This dimension selection not only ensures the richness of the features but also takes into account the computational efficiency. The generated 128-dimensional feature vector can effectively represent the appearance features of the target, such as color, texture, shape, etc. These features will be used for subsequent target matching and trajectory association. In order to further improve the robustness and consistency of the feature vector, it is necessary to perform normalization on it. Specifically as follows: L2 Normalization: Perform L2 normalization on the extracted 128-dimensional feature vector. L2 normalization is a commonly used feature normalization method that can adjust the length (norm) of the feature vector to 1, thereby eliminating the scale difference between different feature vectors. The normalized feature vector: After L2 normalization, the length of the feature vector is 1, which makes the feature vector more stable and reliable in subsequent similarity calculations (cosine similarity).
[0107] Specifically, the calculation formula for L2 normalization is:
[0108]
[0109] where f and f norm are the appearance feature vectors before and after L2 normalization respectively.
[0110] S3.2: Associate the appearance feature vector of the ROI image with the valid target box.
[0111] That is, the parameters of each valid target box are (x gt , ygt , w gt , h gt , f norm ), x gt , y gt , w gt , h gt are the center point coordinates, width, and height of the valid target bounding box respectively; f norm is the appearance feature vector corresponding to the valid target bounding box.
[0112] S3.3: Based on the historical frame state vectors of all the tracked targets, the DeepSort algorithm uses the Kalman filter corresponding to the target to predict the state vector x of the target in the current frame k = [x pred , y pred , w pred , h pred , v x , v y T .
[0113] Specifically, the present invention uses a Kalman filter (Kalman Filter) to predict the motion state of the target. The prediction and update of the Kalman filter (Kalman Filter) can refer to the prior art.
[0114] Specifically, in the prediction stage of the Kalman filter, according to the state vector and state transition matrix F at the previous moment k-1, the state vector at the current moment k is predicted. The state transition matrix F describes the dynamic change of the target motion state. Usually, it is assumed that the target motion is uniform, so:
[0115]
[0116] Predicted state vector:
[0117] x k = F · x k-1 = [x pred , y pred , w pred , h pred , v x , v y T
[0118] S3.4: Ignore the motion speeds v x and v y in the x and y directions of the target, and transform the state vector x of all the current targets in the current frame k into a predicted bounding box (x pred , y pred , w pred , h pred )。
[0119] S4: Based on the parameters of the predicted bounding boxes and the valid target bounding boxes, as well as the extracted and tracked appearance feature vectors, use the DeepSort algorithm to perform one-to-one matching on all valid target bounding boxes and predicted bounding boxes. If the matching is successful, update the trajectory information of the corresponding target and update the state vector and appearance features of the DeepSort algorithm. If the matching fails, initialize the unmatched valid target bounding boxes as new tracking targets or mark the unmatched predicted bounding boxes as lost targets.
[0120] Step S4 specifically includes:
[0121] S4.1: Use the EIoU cost to perform an initial match on the predicted bounding boxes and the valid target bounding boxes to obtain a first candidate match pair combination that meets the initial matching threshold;
[0122]
[0123] Among them, IoU is the intersection ratio of the valid target bounding box and the predicted bounding box, and the calculation method can refer to the prior art. ρ 2 (b pred ,b gt ) is the deviation degree of the center points of the valid target bounding box and the predicted bounding box. ρ 2 (w pred ,w gt ) is the deviation degree of the widths of the valid target bounding box and the predicted bounding box. ρ 2 (h pred ,h gt ) is the deviation degree of the heights of the valid target bounding box and the predicted bounding box. x pred ,y pred ,w pred ,h pred are the center point coordinates, width, and height of the predicted bounding box respectively; x gt ,y gt ,w gt ,h gt are the center point coordinates, width, and height of the valid target bounding box respectively; c, c w and c h are the diagonal length, width, and height of the minimum circumscribed rectangle of the valid target bounding box and the predicted bounding box respectively; ρ 2 is the square of the Euclidean distance.
[0124] The threshold for the initial match can be set to 0.5.
[0125] S4.2: For the valid target bounding boxes that failed to be matched in the initial matching process, perform secondary matching to obtain a second candidate matching pair combination that meets the secondary matching threshold. The method of secondary matching is to compare the similarity CosineSimilarity between the tracking appearance feature vector f1 of the DeepSort algorithm and the appearance feature vector f2 corresponding to the valid target bounding boxes that failed to be matched.
[0126]
[0127] The similarity CosineSimilarity is the cosine similarity. The cosine similarity calculation is performed using the feature vectors after L2 normalization to ensure that the length of the feature vectors is 1, thereby improving the stability of the similarity calculation. The cosine similarity threshold is set to 0.7. Only when the similarity is greater than 0.7, it is considered that the two targets are successfully matched. This threshold ensures that the matched targets have a high similarity in appearance features.
[0128] Of course, in this embodiment, it is also necessary to manage the trajectory life cycle, that is, set the threshold of the number of lost frames of the trajectory to 30 frames. If a target is not detected within 30 consecutive frames, it is considered that the target has been lost and its trajectory will be deleted. For each tracked target, record the number of consecutive lost frames. Once the number of lost frames exceeds the threshold, the trajectory will be removed from the tracking list to avoid the interference of long-unseen targets on the tracking results.
[0129] S4.3: Combine the first candidate matching pair combination and the second candidate matching pair combination to form a bipartite graph, and use the Hungarian algorithm to solve the optimal matching scheme. In the optimal matching scheme, if the valid target bounding boxes and the detection bounding boxes correspond one by one, the matching is successful, and use the successfully matched detection bounding boxes to update the trajectory information, the state vector, and the appearance features of the corresponding targets. Otherwise, the matching fails, and the unmatched valid target bounding boxes are marked as new targets, and the unmatched prediction bounding boxes are retained or marked as lost.
[0130] Specifically, in using the Hungarian algorithm to solve the optimal matching scheme, use the EIoU cost of the initial matching and the similarity of the secondary matching to construct a cost matrix. The goal of using the Hungarian algorithm to solve the optimal matching scheme is to minimize the total cost of the cost matrix.
[0131] Each detection bounding box and prediction bounding box are respectively used as two sets of the bipartite graph, and the elements in the cost matrix represent the edge weights between the two sets; use the Hungarian algorithm to solve the minimum total cost matching of the bipartite graph. The Hungarian algorithm can efficiently find the globally optimal matching scheme to ensure that each detection bounding box is only matched with one prediction bounding box, and vice versa. Cost matrix construction: The matrix element is C ij= 1 - EIoU(i, j). During the initial matching stage, the cost matrix is calculated based on EIoU; during the secondary matching stage, the cost matrix is calculated based on the cosine similarity of the appearance features. In the process of solving by the Hungarian algorithm, it is ensured that each detection box only matches one prediction box, and vice versa. This one-to-one matching constraint avoids the situation where multiple detection boxes match to the same prediction box, thereby improving the accuracy and robustness of the matching. The goal of the Hungarian algorithm is to minimize the total cost of the cost matrix, that is, to maximize the matching similarity between the detection box and the prediction box. In this way, the algorithm can find the globally optimal matching scheme and improve the overall performance of object tracking.
[0132] Using the successfully matched detection boxes, update the state vector and appearance features of the DeepSort algorithm. Specifically:
[0133] S4.3.1: For the successfully matched prediction boxes and valid target boxes, obtain the observation vector z of the current Kalman filter according to the parameters of the valid target box k ;
[0134] z k = [x measured , y measured , w measured , h measured T
[0135] Among them, x measured , y measured , w measured , h measured are the center point coordinates, width and height of the successfully matched valid target box;
[0136] S4.3.2: Use the Kalman gain K and the observation vector z of the current Kalman filter k to update the state vector of the current Kalman filter;
[0137] That is, in the update process of the Kalman filter, the update formula of the state vector is:
[0138]
[0139] Among them, is the updated state vector, and x k is the current state vector predicted by the Kalman filter.
[0140] S4.3.3: For the successfully matched prediction boxes and valid target boxes, update the appearance feature vector f2 extracted from the ROI image corresponding to the valid target box to the tracking appearance feature vector of the current tracking target of the DeepSort algorithm.
[0141] Embodiment 2
[0142] As Figure 3 shown, a multi-object tracking system based on DeepSort includes:
[0143] An object detection module, which uses an improved YOLOv8 model to perform multi-object detection on each frame of the collected video stream, and outputs multiple candidate bounding boxes, class labels, and class confidences for each object. Among them, the improved YOLOv8 model incorporates a dynamic detection head in the detection head and performs feature fusion by adding a cross-scale dynamic gating mechanism in the feature pyramid;
[0144] An ROI image acquisition module, which filters candidate bounding boxes by means of double screening of confidence threshold and class, obtains valid target boxes, and crops the ROI image from the original input image using the valid target boxes as boundaries;
[0145] A prediction module, which uses the DeepSort algorithm to extract appearance feature vectors from the cropped ROI image, predicts the state vectors of all current tracking objects in the current frame, and converts them into prediction boxes;
[0146] An object matching and tracking module, which uses the DeepSort algorithm to perform one-to-one matching on all valid target boxes and prediction boxes based on the parameters of the prediction boxes and valid target boxes, and the extracted and tracked appearance feature vectors. If the matching is successful, the trajectory information of the corresponding object is updated, and the state vector and appearance features of the DeepSort algorithm are updated. If the matching fails, the unmatched valid target boxes are initialized as new tracking objects or the unmatched prediction boxes are marked as lost objects.
[0147] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0148] Embodiment 3
[0149] The present invention provides a computer device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the above multi-object tracking method based on DeepSort are implemented.
[0150] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0151] Embodiment 4
[0152] The present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the above multi-object tracking method based on DeepSort are implemented.
[0153] For a more specific process of the above method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0154] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems, devices, and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and reference can be made to the description of the method part for related parts.
[0155] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0156] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A multi-object tracking method based on DeepSort, characterized in that, It includes the following steps: S1: Perform multi-object detection on each frame of the collected video stream using an improved YOLOv8 model, and output multiple candidate bounding boxes, class labels, and class confidences for each object. Among them, the improved YOLOv8 model incorporates a dynamic detection head in the detection head and performs feature fusion by adding a cross-scale dynamic gating mechanism to the feature pyramid; S2: Filter the candidate bounding boxes through a double screening method of confidence threshold and class, obtain effective target boxes, and use the effective target boxes as boundaries to crop the ROI image from the original input image; S3: Use the DeepSort algorithm to extract the appearance feature vectors from the cropped ROI image, and predict the state vectors of all current tracking objects in the current frame, and convert them into prediction boxes; S4: Based on the parameters of the prediction boxes and effective target boxes, as well as the extracted and tracked appearance feature vectors, use the DeepSort algorithm to perform one-to-one matching on all effective target boxes and prediction boxes. If the matching is successful, update the trajectory information of the corresponding object and update the state vector and appearance features of the DeepSort algorithm. If the matching fails, initialize the unmatched effective target boxes as new tracking objects or mark the unmatched prediction boxes as lost objects.
2. The multi-object tracking method based on DeepSort according to claim 1, wherein Step S1 specifically includes: S1.1: Perform frame-by-frame preprocessing on the collected video stream; the preprocessing includes one or more of dynamic size normalization, introducing a dynamic padding strategy, pixel normalization, and adaptive image enhancement; S1.2: Embed a dynamic convolutional layer in the detection head, and dynamically fuse multi-level features through a cross-scale dynamic gating mechanism in the feature pyramid to improve the YOLOv8 model; during the training process of the improved YOLOv8 model, use a clustering algorithm to generate anchor box sizes for small objects, and during the inference stage, adaptively adjust the anchor box density of different feature layers according to the feature map resolution in advance; S1.3: Use the improved YOLOv8 model to perform multi-object detection on each preprocessed frame of the image, and output multiple candidate bounding boxes, class labels, and class confidences for each object.
3. The multi-object tracking method based on DeepSort according to claim 1, characterized in that Step S2 specifically includes: S2.1: Adaptively set the confidence threshold according to the scene complexity of the current frame to perform the first filtering on all candidate bounding boxes; S2.2: Perform class filtering on the candidate bounding boxes retained after the first filtering, retain the candidate bounding boxes of the set concerned classes, and use the retained candidate bounding boxes as effective target boxes; S2.3: Perform inverse normalization on the output coordinates of the valid target boxes to obtain the coordinates x gt , y gt , w gt , h gt ; S2.4: Use the absolute coordinates of the effective target boxes as the boundaries of the cropping area, and extract the pixel data of the current cropping area from the original input image to generate the ROI image.
4. The multi-object tracking method based on DeepSort according to claim 1, characterized in that, Step S3 specifically includes: S3.1: Use the DeepSort algorithm to perform normalization processing on the cropped ROI image, and use a pre-trained convolutional neural network to extract the appearance feature vectors of all objects; S3.2: Associate the appearance feature vectors of the ROI image with the effective target boxes; S3.3: Based on the historical frame state vectors of all the targets being tracked, the DeepSort algorithm uses the Kalman filter corresponding to the target to predict the state vector x of the target in the current frame k = [x pred , y pred , w pred , h pred , v x , v y T ; S3.4: Ignore the movement speeds v of the target in the x and y directions x and v y , and transform the state vector x of all current targets in the current frame k into a predicted bounding box (x pred , y pred , w pred , h pred ).
5. The multi-object tracking method based on DeepSort according to claim 1, wherein Step S4 specifically includes: S4.1: Use the EIoU cost to perform an initial matching between the predicted bounding boxes and the valid target bounding boxes, and obtain the first candidate matching pair combination that meets the initial matching threshold; Among them, IoU is the intersection ratio of the effective target box and the predicted box, and ρ 2 (b pred ,b gt ) is the deviation degree of the center points of the effective target box and the predicted box, and ρ 2 (w pred ,w gt ) is the deviation degree of the widths of the effective target box and the predicted box, and ρ 2 (h pred ,h gt ) is the deviation degree of the heights of the effective target box and the predicted box, x pred ,y pred ,w pred ,h pred are the center point coordinates, width, and height of the predicted box respectively; x gt ,y gt ,w gt ,h gt are the center point coordinates, width, and height of the effective target box respectively; c, c w and c h are the diagonal length, width, and height of the minimum bounding rectangle of the effective target box and the predicted box respectively; ρ 2 is the square of the Euclidean distance; S4.2: For the valid target bounding boxes that are not successfully matched during the initial matching process, perform a secondary matching to obtain the second candidate matching pair combination that meets the secondary matching threshold. The way of secondary matching is to compare the similarity CosineSimilarity between the tracking appearance feature vector f1 of the DeepSort algorithm and the appearance feature vector f2 corresponding to the valid target bounding boxes that are not successfully matched; S4.3: Construct a bipartite graph from the first candidate matching pair combination and the second candidate matching pair combination, and use the Hungarian algorithm to solve the optimal matching scheme. In the optimal matching scheme, if the valid target bounding boxes and the detection bounding boxes are in one-to-one correspondence, the matching is successful, and use the successfully matched detection bounding boxes to update the trajectory information of the corresponding targets, as well as the state vector and appearance features of the DeepSort algorithm. Otherwise, the matching fails, and the unmatched valid target bounding boxes are marked as new targets, and the unmatched predicted bounding boxes are retained or marked as lost.
6. The multi-object tracking method based on DeepSort according to claim 5, wherein In step S4.3, in the process of using the Hungarian algorithm to solve the optimal matching scheme, use the EIoU cost of the initial matching and the similarity of the secondary matching to construct a cost matrix, and the goal of using the Hungarian algorithm to solve the optimal matching scheme is to minimize the total cost of the cost matrix.
7. The multi-object tracking method based on DeepSort according to claim 5, characterized in that, Step S4.3 includes using the successfully matched detection bounding boxes to update the state vector and appearance features of the DeepSort algorithm. Specifically: S4.3.1: For the successfully matched prediction boxes and valid target boxes, obtain the observation vector z of the current Kalman filter according to the parameters of the valid target boxes k ; z k = [x measured , y measured , w measured , h measured T Among them, x measured , y measured , w measured , h measured are the center point coordinates, width, and height of the successfully matched valid target bounding box; S4.3.2: Use the Kalman gain K and the observation vector z of the current Kalman filter k to update the state vector of the current Kalman filter; Among them, is the updated state vector, and x k is the current state vector predicted by the Kalman filter; S4.3.3: For the successfully matched predicted bounding boxes and valid target bounding boxes, update the appearance feature vector f2 extracted from the ROI image corresponding to the valid target bounding box to the tracking appearance feature vector of the current tracking target of the DeepSort algorithm.
8. A multi-object tracking system based on DeepSort, characterized in that, It includes: A target detection module, which is used to perform multi-target detection on each frame of the collected video stream using an improved YOLOv8 model, and output multiple candidate bounding boxes, class labels, and class confidences for each target. Among them, the improved YOLOv8 model incorporates a dynamic detection head in the detection head and performs feature fusion by adding a cross-scale dynamic gating mechanism to the feature pyramid; An ROI image acquisition module, which is used to filter the candidate bounding boxes by means of double screening of the confidence threshold and the class, obtain the valid target bounding boxes, and crop the ROI image from the original input image using the valid target bounding boxes as the boundaries; A prediction module, which is used to use the DeepSort algorithm to extract the appearance feature vector from the cropped ROI image, and predict the state vector of all current tracking targets in the current frame, and convert it into a predicted bounding box; A target matching and tracking module, which is used to perform one-to-one matching on all valid target bounding boxes and predicted bounding boxes using the DeepSort algorithm based on the parameters of the predicted bounding boxes and valid target bounding boxes, as well as the extracted and tracked appearance feature vectors. If the matching is successful, update the trajectory information of the corresponding targets and update the state vector and appearance features of the DeepSort algorithm. If the matching fails, initialize the unmatched valid target bounding boxes as new tracking targets or mark the unmatched predicted bounding boxes as lost targets.
9. A computer device, characterized in that, It includes a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the multi-object tracking method based on DeepSort described in any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, It is used to store a computer program; when the computer program is executed by the processor, the steps of the multi-object tracking method based on DeepSort described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Target detection method and device based on multi-modal fusion, equipment and storage medium
CN120808296A
Multi-moving-target tracking counting method, device and equipment and storage medium
CN121353345A
Multi-target tracking method and system based on dynamic weight and multistage feature fusion
CN121582298A
Fry automatic identification, tracking and counting method based on machine vision
CN121982688A