A multi-object tracking method based on joint detection

By adopting PVT backbone network and Guided Anchoring technology in the QDTrack model, the problem of model interactions and cross-spatial features is solved, and more efficient detection and tracking performance is achieved.

CN118918143BActive Publication Date: 2025-05-30GUANGXI UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410975506.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-05-30
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

Existing QDTrack models are difficult to effectively capture through local convolution operations when capturing interactions between targets or key features within targets that span a larger spatial range.

Method used

Pyramid Vision Transformer (PVT) is used as the backbone network, replacing the original ResNet50, enhancing remote modeling capabilities, and introducing Guided Anchoring to guide the generation of anchor boxes in RPN results. At the same time, the two-way softmax function is optimized for target matching based on cosine similarity and IoU distance.

Benefits of technology

The detection accuracy of the model is improved, the number of missed and missed detection is reduced, and the tracking performance of the model is enhanced, especially when the target faces occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918143B_ABST
    Figure CN118918143B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-object tracking method based on joint detection, belonging to the technical field of object tracking, including: using the PVT model as the backbone network to extract the features of the image; positioning and identifying the target features in the image based on the GA-RPN structure to obtain candidate target regions; optimizing the bidirectional softmax function based on cosine similarity and IoU distance, and matching the targets in the candidate target regions through the optimized bidirectional softmax function to generate multi-object tracking results. The present invention studies the tracking model QDTrack using quasi-dense similarity learning and optimizes the model in three aspects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target tracking, and particularly relates to a multi-target tracking method based on joint detection. Background Art

[0002] With the promotion of social needs and technological progress, multi-target tracking technology has made remarkable progress, and the development of current multi-target tracking algorithms has tended to be mature. QDTrack is a two-stage tracking model based on Faster R-CNN, which densely samples hundreds of region proposals on a pair of images for contrastive learning through quasi-dense similarity learning. This tracking model does not require additional motion priors, and only calculates similarity through a simple bi-directional softmax function during the tracking matching process, and achieves efficient tracking through a simple matching method.

[0003] However, the original QDTrack uses ResNet as the feature extraction network, and it may be difficult to capture the interaction between targets or the key features within a large spatial range inside the target through local convolution operations. How to optimize the QDTrack model has become an urgent problem to be solved. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a multi-target tracking method based on joint detection to solve the problems existing in the above prior art.

[0005] To achieve the above object, the present invention provides a multi-target tracking method based on joint detection, including:

[0006] Using the PVT model as the backbone network to extract the features of the image;

[0007] Based on the GA-RPN structure, localize and identify the target features in the image to obtain candidate target regions;

[0008] Optimize the bi-directional softmax function based on cosine similarity and IoU distance, and match the targets in the candidate target regions through the optimized bi-directional softmax function to generate multi-target tracking results.

[0009] Preferably, the process of using the PVT model as the backbone network to extract the features of the image includes:

[0010] The PVT model controls the scale of the feature map through a progressive shrinking strategy;

[0011] For the i-th stage, the feature map F from the input of the previous stage i-1 (H i-1 ×Wi-1 ×C i-1 ) is divided into patches, and then each patch is flattened and mapped into a feature embedding of dimension C i . After passing through a linear mapping, the feature size can be regarded as

[0012] Preferably, the PVT model adopts a spatial scaling attention mechanism;

[0013] The expression of the spatial scaling attention mechanism is:

[0014]

[0015] where Concat() is the concatenation operation in the Transformer, is a weight parameter, N i is the number of attention heads in the i-th stage. Therefore, the dimension of each head is LSR() is a linear dimensionality reduction operation on the spatial dimension of the input sequence.

[0016] Preferably, in the process of localizing and recognizing the features of the target in the image based on the GA-RPN structure and obtaining the candidate target regions, it further includes: constructing a quasi-dense similarity learning model based on a joint learning model with sparse ID loss, and sampling the surrounding regions of the candidate target regions based on the quasi-dense similarity learning model to obtain positive and negative samples of the information regions.

[0017] Preferably, the expression for the GA-RPN structure to predict the best shape at each position is:

[0018] w = σ · s · e dw ;

[0019] h = σ · s · e dh ;

[0020] where w and h are the best shapes at each position, that is, the shapes that may result in the highest coverage rate with the nearest true label bounding box, s is the stride, and σ is the scale factor.

[0021] Preferably, the GA-RPN structure is also used to add a 1×1 convolution on the output of the shape prediction branch to predict the offset field, and then map the offset amount in a 3×3 convolution to obtain an adaptive feature map.

[0022] Preferably, in the process of matching the target of the candidate target region through an optimized bidirectional softmax function to generate multi-target tracking results, it includes: constructing an IoU threshold β iou , when the IoU threshold of the candidate target is greater than βiou When matching, cosine similarity is preferentially considered, and the remaining candidates less than β iou with a detection confidence higher than β obj , then consider matching the detected target with the existing trajectory. If the matching score is greater than β m , then determine the match. For the unmatched candidate targets, if the detection confidence is higher than β new , then create a new trajectory.

[0023] Preferably, the expression of the loss function of the overall network is:

[0024] L = L det + γ 1 L embed + γ 2 L aux ;

[0025] Among them, L det is the original Faster R-CNN loss function, including RPN loss, class loss, and regression loss. L aux is the auxiliary loss.

[0026] Compared with the prior art, the present invention has the following advantages and technical effects:

[0027] The present invention studies the tracking model QDTrack using quasi-dense similarity learning and optimizes the model in three aspects. First, the present invention replaces the ResNet50 backbone network in the original tracking model with Pyramid Vision Transformer, improving the long-range modeling ability of the model. Second, the present invention adopts GA-RPN with position guidance to improve the model's positioning ability for information in the feature map. Experiments prove that these two improvement methods effectively improve the detection accuracy of the model, reduce the number of missed detections and false detections, and thus improve the tracking performance of the model. Finally, the present invention proposes a two-way softmax matching integrating cosine similarity to optimize the original matching method and selects the optimal parameter values through experiments. The improved model achieves the best results in MOTA, MT, and ML metrics compared with mainstream algorithms and has good continuous tracking ability when the tracked target is occluded. Brief Description of the Drawings

[0028] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0029] Figure 1 is the optimized model structure diagram of the embodiment of the present invention;

[0030] Figure 2 Schematic diagram of the PVT network module structure according to an embodiment of the present invention;

[0031] Figure 3 Schematic diagram of the GA-RPN network structure according to an embodiment of the present invention;

[0032] Figure 4 Schematic diagram of the comparison of MOTA metrics for each sequence in the MOT16 test set according to an embodiment of the present invention;

[0033] Figure 5 Schematic diagram of the comparison of MOTA metrics for each sequence in the MOT17 test set according to an embodiment of the present invention. Detailed implementation manners

[0034] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0035] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0036] Embodiment 1

[0037] In this embodiment, a multi-object tracking method based on joint detection is provided, including:

[0038] The original QDTrack uses ResNet as the feature extraction network, and it may be difficult to capture the interaction between targets or the key features within a target that span a large spatial range through local convolution operations. Therefore, in this embodiment, Pyramid Vision Transformer is used as the backbone network to enhance the long-range modeling ability of the network, and Guided Anchoring is introduced into the RPN results to guide the generation of anchor boxes. The optimized network structure is as Figure 1 shown;

[0039] Transformer-based detection models have been widely applied in computer vision tasks and achieved advanced performance as the backbone network. However, most models only propose attention mechanisms similar to Transformer and rarely use convolution-free Transformer for feature extraction. For dense prediction tasks, the long-range modeling ability of Transformer can better extract feature information and provide better support for subsequent processing. Pyramid Vision Transformer (PVT) overcomes the problems of low output resolution and high computational cost. As the network deepens like a traditional convolutional neural network, it can achieve feature scaling and channel deepening, thus constructing a feature pyramid-like structure to generate multi-scale feature maps.

[0040] As Figure 2 shown, similar to the CNN backbone network, PVT contains 4 stages, and each stage controls the scale of the feature map through a progressive shrinking strategy. For the i-th stage, the feature map F i-1 (H i-1 ×W i-1 ×C i-1 ) from the input of the previous stage is divided into patchs, and then each patch is flattened and mapped into a C i -dimensional feature embedding. After linear mapping, the feature size can be regarded as The size of the feature map is reduced by P i times compared with the previous stage.

[0041] Each stage contains L i Transformer encoders, and each encoder consists of an attention layer and a feed-forward layer. The traditional Transformer adopts the multi-head attention mechanism (MHA), which greatly increases the computational cost. PVT adopts the spatial reduction attention mechanism (SRA) to replace MHA. Similar to MHA, SRA receives a query Q, a key K, and a value V as inputs and outputs a refined feature. The difference is that SRA reduces the spatial scale of K and V before the attention operation, and the specific operation is shown in equations (1)-(2):

[0042]

[0043] where Concat() is the concatenation operation in Transformer. are weight parameters. N iis the number of attention heads in the i-th stage. Therefore, the dimension of each head is LSR() performs a linear dimensionality reduction operation on the spatial dimension of the input sequence, as shown in Equation (3):

[0044] LSR(x) = Norm(Pooling(x, P)W S ) (3)

[0045] where represents the length of the input sequence, P represents the reduction size of each stage attention layer, with a default value of 7. W S is a linear projection operation that reduces the input sequence to C i dimensions. Similar to the attention calculation of the original Transformer, the final calculation is shown in Equation (4):

[0046]

[0047] By adopting the PVT structure, the improved network obtains more local continuity of the feature images, the input variable resolution is more flexible, and it has the same linear complexity as CNN.

[0048] Anchors are an important part of the object detection model. Most detectors rely on generating dense anchor schemes. By uniformly sampling in the image, a set of predefined scale anchors are generated. However, this method is very inefficient. In this embodiment, the RPN structure in the original tracking model is optimized, and the Guided Anchoring is introduced to form the GA-RPN structure to guide the generation of anchors.

[0049] GA-RPN uses semantic features to guide the positioning of anchors. This method simultaneously predicts the possible positions where the centers of the objects of interest may exist and the proportional sizes at different positions. As Figure 3 shown, this module is divided into two parts: anchor generation and feature adaptation. Among them, anchor generation is divided into position prediction and shape prediction. In the position prediction branch, the feature map obtains the score map of each position through a 1×1 convolution, and then is converted into the corresponding probability value through the sigmoid function. By selecting the probability threshold, most of the background regions can be filtered out.

[0050] After determining the possible positions of the objects, the next step is to determine the shapes of the objects that may exist at each position through the shape prediction branch. Different from traditional bounding box regression, the shape prediction branch does not change the position of the anchor. Specifically, for a given feature map, this branch will predict the optimal shapes w and h at each position, that is, the shapes that may result in the highest coverage rate with the nearest ground truth bounding box. However, directly predicting the length and width from practical experience is unstable because the value range is large. Simplified prediction can be carried out through equations (5) and (6).

[0051] w = σ · s · e dw (5)

[0052] h = σ · s · e dh (6)

[0053] Where s is the stride and σ is the scale factor (default set to 8). The shape prediction branch will output dw and dh, and then map them to w and h. This transformation can map the output values to the interval [-1, 1], making it easier to learn stably. In the shape prediction branch, a 1×1 convolution is used to generate a two-channel map containing dw and dh to achieve the transformation.

[0054] In the traditional RPN network, the positions of the anchors are uniform on the entire feature map, with the same shape and ratio at each position, and the feature map can learn a consistent representation. However, through the position and shape prediction branches, the sizes of the anchors are not fixed. Larger anchors should encode the content of larger regions, and smaller anchors should encode the content of smaller regions. Therefore, the network can be guided by the prediction of the anchor shapes to achieve adaptive feature transformation. A 1×1 convolution is added to the output of the shape prediction branch to predict the offset field, and then the offset is mapped in the 3×3 convolution to obtain an adaptive feature map.

[0055] The original tracking model uses a bidirectional softmax function to calculate the similarity for tracking matching. In this embodiment, the matching process is further optimized on this basis. By introducing the cosine similarity function and combining with IoU, more accurate association can be achieved. Although using appearance similarity is a reliable metric in most scenarios, for dense target objects, combining with the IoU distance can improve the matching accuracy. Under conditions of blurred images and some fast-moving targets, appearance similarity is not reliable. If the search and matching are carried out within a certain IoU distance range, the interference of similar targets can be effectively avoided. Therefore, this embodiment re-optimizes and proposes a new association matching method.

[0056] Cosine Similarity is a measure method to measure the similarity between two vectors, and is usually used to calculate the similarity in fields such as text, images, user preferences, etc. Cosine Similarity measures the cosine value of the angle between two vectors. The closer the value is to 1, the higher the similarity; the closer the value is to -1, the lower the similarity. For two vectors a and b, the calculation formula of Cosine Similarity is:

[0057]

[0058] Cosine Similarity mainly focuses on the direction of the vectors rather than the length. Therefore, it is insensitive to the changes in scale and contrast in images. This means that when the image undergoes scaling or brightness changes, Cosine Similarity can better capture the similarity between images.

[0059] In trajectory matching, in this embodiment, the IoU threshold β is introduced iou , when the IoU threshold of the candidate target is greater than β iou , the Cosine Similarity will be preferentially considered for matching. For the remaining candidate targets less than β iou , if the detection confidence is higher than β obj , then consider matching the detected target with the existing trajectory. If the matching score is greater than β m , then determine the match. If the detection confidence of the unmatched candidate target is higher than β new , then create a new trajectory. The remaining unprocessed candidate targets will be placed in the background. The activated trajectory means that the target in the previous frame has completed the matching with the current candidate target, and the unmatched trajectory will be in an unactivated state. If the trajectory is not activated within K frames, it will be removed and no longer considered for matching. Generally, K is defaulted to 30. The pseudo-code of the optimized tracking and matching process is as follows.

[0060]

[0061] The configuration information of the hardware and software used in the experiment of this embodiment is shown in Table 1.

[0062] Table 1

[0063]

[0064] To ensure fairness, the experiments in this embodiment followed the parameter settings of the original QDTrack model training. 128 RoIs were selected from the key frames as training samples, and 256 RoIs were selected from the reference frames with a positive-negative ratio of 1 as the comparison targets. RoIs were sampled using IoU-balanced sampling. 4 convolutional layers and 1 fully connected layer were used as the Embedding Head for feature extraction, and the size of the feature dimension was 256. The experiments in this embodiment used AdamW as the optimizer, with an initial learning rate of 1e-4, a weight decay coefficient of 1e-4, a batch size of 2, and warm-up training was performed by linearly increasing the learning rate in the first 1000 iterations. A total of 12 epochs were trained, and after the 8th epoch and the 11th epoch, the learning rate was reduced by 0.1 respectively.

[0065] In the experiments, the original scale of the images was used for training and inference. Except for random horizontal flipping, no other data augmentation methods were used. The experiments used pre-trained weights on the COCO dataset for training. When performing online joint object detection and tracking, if the detection confidence is greater than 0.8, a new trajectory is initialized. Only one frame of the background is retained, and objects can be associated only when they are classified as the same category.

[0066] In addition to the detection loss and similarity loss, L2 loss was additionally introduced as auxiliary training in model training, as shown in Equation (8).

[0067]

[0068] If two samples are matched as positive, c is 1, otherwise it is 0. The auxiliary loss L aux The purpose is to limit the logarithm size in similarity learning rather than improve performance. Therefore, the loss of the entire network is as shown in Equation (9):

[0069] L = L det + γ 1 L embed + γ 2 L aux (9)

[0070] where L det is the original Faster R-CNN loss function, including RPN loss, class loss, and regression loss. To maintain fairness, the experiments referred to the parameter settings of the original QDTrack training and set the weights of γ 1 and γ 2 to 0.25 and 0.1 respectively.

[0071] In addition to conducting experiments on the MOT16 and MOT17 datasets, this embodiment additionally uses the CrowdHuman dataset for training.

[0072] The CrowdHuman dataset is a large-scale dataset for pedestrian detection. This dataset focuses on crowds in dense scenes and has images with complex occlusions, diverse poses, and different perspectives. The dataset contains approximately 15K images and 340K pedestrian instances. In this embodiment, only the training set in the CrowdHuman dataset is used as a supplementary dataset for training on the basis of the original MOT dataset.

[0073] To verify the effectiveness of the improved tracking model in this embodiment, ablation experiments were conducted on the MOT dataset. Since the test set of MOT needs to be submitted for evaluation on the official website and is restricted by the submission time and number of times, in the ablation experiments of this embodiment, half of the MOT17 training set was divided into the training set and the other half was divided into the validation set to verify the tracking performance of the model. This embodiment mainly conducted three experiments, including ablation experiments on each module, matching effect experiments of different similarity functions, and optimized tracking matching performance experiments under different thresholds.

[0074] In this embodiment, on the basis of the original tracking model QDTrack, the PVT and GA-RPN modules are introduced respectively to verify the detection and tracking performance of the model. As shown in Table 2, after replacing the backbone network of QDTrack with PVT, the MOTA and IDF1 metrics of the model are improved by 3% and 2.8%, the number of identity switches IDs is reduced by 63, and the AP and Recall are increased by 2.1% and 1.3% respectively. It is verified that PVT without CNN structure as the backbone network has stronger feature extraction ability. After replacing the original RPN with GA-RPN, the MOTA and IDF1 metrics of the model are improved by 2.6% and 1.2%, the number of identity switches IDs is reduced by 132 times, and the AP and Recall are increased by 2.3% and 1.9% respectively. It shows that GA-RPN can optimize the anchor position through the feature map to improve the detection accuracy of the model, and then improve the overall tracking performance. When the PVT and GA-RPN modules are introduced simultaneously, the performance of the model is further improved. Compared with the original model, the MOTA and IDF1 metrics are improved by 3.3% and 3.8%, the number of identity switches IDs is reduced by 177 times, and the AP and Recall are increased by 2.4% and 1.8% respectively. Although the Recall of the model is reduced by 0.1%, the overall tracking performance is not reduced, especially the number of identity switches during the tracking process is significantly reduced. It shows that while improving the feature extraction ability of the model, introducing more targeted anchor positioning has better discrimination of the features of the detection target, making it easier to distinguish different identities when calculating the similarity in tracking matching, and effectively reducing the number of identity switches.

[0075] Table 2

[0076]

[0077] When performing similarity matching, calculating similarity through different similarity functions will lead to different matching results. To verify the impact of various similarity functions on the tracking results, in this embodiment, the softmax function, cosine function, and bidirectional softmax function are respectively selected for comparative experiments. As shown in Table 3, the bidirectional softmax function can effectively reduce the number of identity switches IDs, verifying that the bidirectional softmax function can better calculate the similarity relationship between the detected object and the candidate target, and thus can accurately find the corresponding target during matching. When using R50 - RPN as the backbone network of the detection model, the difference in the MOTA and IDF1 metrics for different similarity functions is only 0.2%. However, when using PVT - GA - RPN as the backbone network of the detection model, the bidirectional softmax function makes both the MOTA and IDF1 metrics reach the highest, increasing by 0.4% and 0.7% compared to the softmax function, and increasing by 0.3% and 1.5% compared to the cosine function. This verifies that the bidirectional softmax function can better utilize the feature information extracted by the PVT - GA - RPN network for similarity matching.

[0078] Table 3

[0079]

[0080] In the previous experiments, it can be observed that using different similarity functions for matching has certain advantages in different aspects. Therefore, in this embodiment, a fusion of cosine similarity and bidirectional softmax matching is proposed. Based on the bidirectional softmax function, the IoU distance and cosine similarity are introduced. By preferentially calculating the similarity between the detected object and the candidate target within a certain IoU distance, and then using bidirectional softmax to match the remaining objects. As shown in Table 4, this embodiment selects different IoU distance thresholds to verify the effect of tracking matching. It can be seen that the matching algorithm proposed in this embodiment has achieved further improvement on both the original model and the improved model. When β iou is 0.8, the best tracking effect is achieved in all metrics. Analyzing in combination with the experimental results in Table 3, after the original tracking model adopts the new matching method, the number of identity switches IDs is reduced by 74 times, and the MOTA and IDF1 metrics are improved by 0.7% and 0.5%. After the improved tracking model adopts the new matching method, the number of identity switches IDs is reduced by 22 times, and the MOTA and IDF1 metrics are improved by 0.1% and 0.3%.

[0081] Table 4

[0082]

[0083] In this embodiment, the MOTA metrics of the algorithm before and after improvement are first compared on the MOT16 and MOT17 datasets. As Figure 4 shown, the MOTA metrics of each video sequence in the MOT16 test set are presented. It can be seen that the tracking performance of the improved algorithm has a relatively obvious improvement in other video sequences, except that there is no obvious change in the MOT16-03 video sequence.

[0084] Figure 5 Similar results can also be found in the MOT17 test set results in . The tracking performance of the algorithm before and after improvement has a significant improvement on the MOT17-01, MOT17-06, MOT17-07, and MOT17-08 sequences. The tracking performance remains unchanged on other video sequences. It verifies that the improved algorithm in this embodiment has a certain generalization ability and is superior to the original QDTrack model.

[0085] The tracking model proposed in this embodiment and the mainstream multi-object tracking models are respectively compared on the MOT16 and MOT17 benchmarks. Two types of tracking methods are selected for comparison in this embodiment, including detection-based tracking models: DeepSort, TAP, POI, TubeTK, CTracker, and joint-detection-based tracking models: JDE, Tracktor, CenterTrack, FairMOT, GSDT, Swin_JDE, QDTrack. As shown in Tables 5 and 6, * indicates detection-based tracking models, and the metrics of some methods are empty, indicating that the specific values cannot be found in the relevant papers or the MOT official website.

[0086] In the MOT16 dataset, the improved algorithm in this embodiment has a 3.4-point increase in the MOTA metric compared to the original QDTrack. Although there are decreases in the IDF1 and HOTA metrics, there is a 25-point increase in the MT metric, reflecting an increase in the number of continuous tracking trajectories of the model, and a corresponding 53-point decrease in the ML metric. During the tracking process, the IDs metric for target identity switching decreases by 70, indicating a certain improvement in the continuous tracking ability of the improved model. In addition, the FP and FN metrics of the improved model decrease by 19.3% and 9.6% respectively, indicating that the model has significantly reduced the number of false detections and has better detection accuracy. The Tracktor algorithm has some similarities in results with the algorithm proposed in this embodiment, and both use Faster-RCNN as the detector. However, the algorithm in this embodiment uses quasi-dense similarity learning to learn instance similarity from dense image pairs, and has stronger discriminative ability compared to Tracktor, which learns the similarity of a single image pair. Therefore, the algorithm proposed in this embodiment has a significant performance improvement compared to Tracktor. The TAP algorithm has significant advantages in the IDF1 and IDs metrics, but its comprehensive performance metric MOTA is 8.4 points lower than that of the algorithm proposed in this embodiment. These gaps are reflected in the relatively high false detections and missed detections of TAP. This also shows that the performance of the detection model has a greater impact on the performance of the tracking model.

[0087] Table 5

[0088]

[0089] Table 6

[0090]

[0091] In the MOT17 dataset, compared with the original QDTrack, the improved algorithm in this embodiment has certain improvements in all other performance metrics, although there is no change in the IDF1 metric. Among them, the MOTA metric has increased by 4.3 points, the HOTA metric has increased by 0.8 points, and the IDs metric has decreased by 222. The MT metric has increased by 141, reaching the best, and the corresponding ML metric has decreased by 240, further verifying that the improved model has a stronger continuous tracking ability for targets. The number of false detections and missed detections has decreased by 11.5% and 14.4% respectively, further verifying that the improved model has good detection performance. Similar to the MOT16 dataset, Tracktor has good results in the number of false detections and IDs, but its MOTA metric is still the lowest. Swin_JDE adopts the idea of the JDE algorithm and uses a more powerful Transformer as the detector, so its performance has been greatly improved compared with JDE. The algorithm proposed in this embodiment also uses a Transformer as the feature extraction network to improve the detection performance of the model. In addition, this embodiment proposes to fuse cosine similarity and bidirectional softmax matching to improve the data association ability of the model, which is better than Swin_JDE in the MT and ML metrics, so the overall performance of the model is leading.

[0092] To further verify the effectiveness of the method proposed in this embodiment in improving the performance of the tracking model, this section shows the tracking results of some sequences in the MOT17 test set. In the 614th frame image of this embodiment, a pedestrian with ID 71 marked by a red arrow is selected as the research target. In the image of the 633rd frame, it can be seen that the pedestrian is blocked by a street lamp, and QDTrack cannot continuously track the target during the occlusion. On the contrary, the improved algorithm can still continuously track the target under the occlusion of the obstacle. Therefore, after the target completely appears in the frame at the 650th frame, the improved algorithm still maintains the original ID. However, QDTrack switches the ID at this time, and the ID becomes 177, judging the target as a new target.

[0093] At the same time, a pedestrian with ID 59 is selected as the research target. It can be seen that this target is about to intersect with the pedestrian in front. In the image of the 145th frame, the target is completely blocked by the pedestrian, and neither algorithm can continuously track the target. But in the image of the 150th frame, the target reappears, and the improved algorithm correctly judges that the ID of the target is 59 and realizes continuous tracking of the target. On the contrary, QDTrack judges this target as a newly emerging target. Through the above analysis, it is verified that the improved algorithm has a strong continuous tracking ability, and after the target blocked by an obstacle or a pedestrian reappears, the number of ID switches can be greatly reduced.

[0094] Although the improved model has achieved good results on the MOT dataset, quasi-dense similarity learning has great potential. Therefore, this section selects the more challenging DanceTrack dataset for further analysis. The DanceTrack dataset contains more than 100K image frames, almost 10 times that of the MOT17 dataset. This data mainly collects the recognition of group dances, which pays more attention to the motion discrimination ability of the tracking model. Therefore, the DanceTrack dataset brings two challenges: firstly, the appearance of people in the dataset is extremely similar, and it is difficult for a simple ReID model to distinguish them; secondly, the body posture movement amplitude of people in the dataset varies greatly, and there are frequent intersections and occlusions between people.

[0095] The DanceTrack dataset uses HOTA as the main evaluation metric in the evaluation metrics because the MOTA metric pays more attention to the detection quality of the model, while the DanceTrack dataset has fewer targets compared to the MOT dataset, and it pays more attention to the association ability of the model. The AssA and IDF1 metrics are used to measure the association performance, and the DetA and MOTA metrics are used to measure the detection quality. The experiment refers to the configuration trained on the MOT17 dataset to train the DanceTrack dataset. The test set results are submitted on the official website of DanceTrack for evaluation and compared with mainstream tracking models. The test results are shown in Table 7.

[0096] Table 7

[0097]

[0098] It can be found that all indicators of the improved algorithm have been improved compared with QDTrack, and the HOTA indicator has been improved by 3 points. ByteTrack leads in the MOTA and IDF1 indicators because its detection model and appearance discrimination model are independent. Although ByteTrack based on YOLOX as a detector has good detection ability, when the target features extracted by the ReID model face targets with similar appearances, its HOTA and AssA indicators do not reach the ideal results. Relying solely on the ReID algorithm is not the optimal solution for multi-object tracking models. Quasi-dense similarity learning can better learn the differences between target objects and the surrounding environment compared to sparse feature learning.

[0099] The sequences of the DanceTrack test set were selected for the experiment to compare with the tracking effect of the QDTrack algorithm. In the image of the 178th frame, the ID of a target in QDTrack changed, while the ID of the improved algorithm did not change. Then, in the image of the 300th frame, the IDs of four targets in QDTrack all changed, while the ID of the improved algorithm still remained unchanged. This shows that the improved algorithm has better discriminative ability for targets in the face of complex motion scenes.

[0100] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-target tracking method based on joint detection, characterized in that: The method is implemented based on the optimized tracking model QDTrack and includes the following steps: Use the PVT model as the backbone network to extract image features; Based on the GA-RPN structure, the target features in the image are located and identified to obtain a candidate target area; the GA-RPN structure includes anchor generation and feature adaptation, wherein the anchor generation is divided into position prediction and shape prediction, in the position prediction branch, the feature map is subjected to 1×1 convolution to obtain a score map for each position, and then converted into a corresponding probability value through a sigmoid function; the shape prediction branch predicts the best shape for each position for a given feature map; a 1×1 convolution is added to the output of the shape prediction branch to predict the offset domain, and the offset is mapped in a 3×3 convolution to obtain an adaptive feature map; The bidirectional softmax function is optimized based on the cosine similarity and the IoU distance, and the targets in the candidate target area are matched by the optimized bidirectional softmax function to generate a multi-target tracking result; The target in the candidate target area is matched by the optimized bidirectional softmax function, and the process of generating multi-target tracking results includes: constructing an IoU threshold β iou , when the IoU threshold of the candidate target is greater than β iou When , cosine similarity will be given priority for matching, and the remaining value is less than β iou The candidate target with detection confidence higher than β obj , then consider matching the detected target with the existing trajectory, if the matching score is greater than β m , then the match is determined, and the unmatched candidate target is determined if the detection confidence is higher than β new , a new trajectory is created.

2. The multi-target tracking method based on joint detection according to claim 1, characterized in that: The process of using the PVT model as the backbone network to extract image features includes: The PVT model controls the scale of the feature map through a progressive shrinkage strategy; For the i-th stage, the feature map F from the previous stage input i-1 Divided into patches, where F i-1 Size H i-1 ×W i-1 ×C i-1 , and then flatten each patch and map it to C i dimensional feature embedding, and then after linear mapping, the feature size can be regarded as 3. The multi-target tracking method based on joint detection according to claim 2, characterized in that: The PVT model adopts a spatial scaling attention mechanism; The expression of the spatial scaling attention mechanism is: Among them, Concat() is the concatenation operation in Transformer. is the weight parameter, N i is the number of attention heads in the i-th stage, so the dimension of each head is LSR() performs a linear dimensionality reduction operation on the spatial dimension of the input sequence. Q, K, and V represent a query Q, a key K, and a value V, respectively. Attention() is the attention function.

4. The multi-target tracking method based on joint detection according to claim 1, characterized in that: The features of the target in the image are located and identified based on the GA-RPN structure, and the process of obtaining the candidate target area also includes: constructing a quasi-dense similarity learning model based on the joint learning model of sparse ID loss, sampling the surrounding area of ​​the candidate target area based on the quasi-dense similarity learning model, and obtaining positive and negative samples of the information area.

5. The multi-target tracking method based on joint detection according to claim 1, characterized in that: The GA-RPN structure predicts the expression of the optimal shape for each position: w=σ·s·e dw ; h=σ·s·e dh ; where w and h are the best shapes at each location, i.e., the shapes that may result in the highest coverage with the nearest true label bounding box, s is the stride, σ is the scaling factor, and the shape prediction branch will output dw and dh, which are then mapped to w and h.

6. The multi-target tracking method based on joint detection according to claim 1, characterized in that: The expression of the loss function of the overall network is: L=L det +γ1L embed +γ2L aux ; If the two samples match positively, c is 1, otherwise it is 0. det is the original Faster R-CNN loss function, including RPN loss, category loss and regression loss, L aux Auxiliary loss.

Citation Information

Patent Citations

  • Multi-target tracking and segmentation system and method

    CN115880330A

  • Method, device and system for detecting target object in video data

    CN116994181A