Discriminant target tracking method based on fusion backbone and effective attention
The fusion of ResNet50 and SwinT-B with GhostAtFeat and attention mechanisms in visual object tracking improves feature representation and online update strategies, addressing generalization issues in Transformer-based trackers and achieving superior tracking performance.
Patent Information
- Application Number
- CN202510401964.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
The existing visual target tracking technology lacks generalization ability in complex and variable scenarios, lacks feature extraction and fusion discriminant enough, and lacks online update modules, resulting in the tracker being easily drifted when the target changes.
A fusion backbone network and effective attention module are adopted, combined with the GhostAtFeat module and the online update structure, a target tracking model based on the Transformer model is built, and a more discriminant feature representation is generated through feature extraction, attention enhancement and online update.
Improve the robustness and accuracy of target tracking, enhance the adaptability to complex scenarios, reduce model drift, and achieve better tracking performance.
Smart Images

Figure CN120318275A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual object tracking, and particularly relates to a discriminative object tracking method based on a fusion backbone and effective attention. Background Art
[0002] Visual object tracking is a notable direction for tracking single and arbitrary target objects. This task processes video sequences and only provides the initial target state. In the past few decades, trackers based on discriminative correlation filters (DCF) and Siamese networks have led the development trend of object tracking. In recent years, deep learning technology has been an important driving force for the advancement of visual tracking. However, the complexity and variability of the object tracking scenario affect the generalization ability of the tracker in the actual scenario. Therefore, it is necessary to develop more discriminative feature representations to achieve stable tracking.
[0003] Generally, the choice of the feature extraction backbone determines the upper limit of the tracking performance. Existing trackers either adopt CNN or vision Transformer as the feature extraction backbone. Since the feature expression ability of freezing the pre-trained network weights is not fully applicable to the object tracking task, the current feature extraction network is initialized with pre-trained weights and fine-tuned in the tracking dataset. In terms of feature fusion, the DCF-based tracker adopts the concept of DCF, integrates the training and testing branches, trains the model predictor and obtains the DCF target model, and then uses the DCF target module to generate the Gaussian response map. Nowadays, most mainstream trackers adopt the Transformer structure, which greatly improves the tracking performance. Specifically, ViT and SwinT, as competitive vision backbones, can replace CNN in obtaining the representative features of the target. In fact, feature extraction is to obtain better representative features, and the purpose of feature fusion is to better obtain discriminative target features. After obtaining the feature representation, the next step is to accurately predict the tracking state of the target through the prediction head, whose main function is to perform bounding box regression. However, visual tracking can be regarded as a one-shot learning task because this task only provides the target information of the first frame and makes subsequent predictions through the proposed tracker. Since the tracked target has uncertain changes, it is difficult to obtain robust tracking results only based on the feature information of the first frame. Therefore, an effective online update component is particularly important. The online update part is a standard configuration of the DCF-based tracker, while the Siamese tracker usually does not have the online update ability. In the current era of the popularity of Transformer, recent trackers unify the essence of the Siamese network and the DCF-based tracker into the Transformer tracker, such as ToMP, TrDiMP, STARK, MixFormer, TransT, and OSTrack. In particular, TransT and OSTrack do not have an online update component. For the tracking task, the online update module is essential, which can mitigate the impact of changes in the tracked target, but incorrect updates are also fatal because they can lead to model drift. The current online update methods mainly involve keeping the target of the first frame unchanged and adding dynamic sample frames, which well alleviates the online tracking problem.
[0004] Since the introduction of Transformer into the tracking field in 2021, new tracking records have been continuously broken, and Transformer-based trackers have achieved gratifying results. We found that their current tracking performance largely owes to the feature learning ability of Transformer. In particular, the single-stream Transformer tracking process adopts efficient feature extraction and sufficient feature fusion, but it mainly involves the extension of siamese trackers. In addition, ToMP uses Transformer to learn the DCF target model, which has achieved significant improvements compared with the optimization-based DiMP. However, there is still a performance gap compared with the single-stream Transformer tracking performance.
[0005] Based on the above analysis of current trackers, visual tracking consists of two branches. The input of DCF-like trackers is called the training and test branches, while the input of siamese trackers is called the template and search branches. In addition, the training branch of DCF trackers contains more background information, aiming to distinguish foreground targets from the background and making it more discriminative. On the other hand, the template of siamese trackers contains little background information and mainly uses template matching to find the most similar target. With the help of Transformer, the effectiveness of feature extraction and feature fusion has been significantly improved. However, we found that the existing ToMP still uses ResNet as the backbone for feature extraction, while other Transformer-based trackers (such as OSTrack) adopt ViT as the backbone to jointly learn feature extraction and fusion. Therefore, a superior backbone network is the key to improving tracking performance. In addition, further transforming the backbone features into more discriminative feature representations is also crucial. Various convolutional or attention modules can effectively enhance the acquisition of representative target features. Nowadays, some siamese-based Transformer trackers also use a score prediction head to obtain a dynamic template and a fixed initial template, and achieve feature interaction under the Transformer architecture, effectively alleviating the disadvantages of the fixed template in early siamese trackers. However, this method requires additional branches and secondary training.
[0006] Therefore, it is an urgent problem for those skilled in the art to propose a discriminative object tracking method based on fused backbone and effective attention to solve the difficulties existing in the prior art. Summary of the Invention
[0007] In view of this, the present invention provides a discriminative object tracking method based on fused backbone and effective attention. By means of a feature extraction structure, an attention feature enhancement structure, and an online update structure, an object tracking model based on a Transformer model predictor is constructed. In the feature extraction structure stage, the fused backbone and the spatial and channel fusion attention module with Ghost can generate more discriminative feature representations for training a robust object tracking model.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A discriminative object tracking method based on fused backbone and effective attention, comprising the following steps:
[0010] Obtain a visual object tracking dataset and perform preprocessing to obtain a training sample set;
[0011] Train an object tracking model to obtain a trained object tracking model. The object tracking model includes a feature extraction structure, an attention feature enhancement structure, and an online update structure;
[0012] Input the object to be tracked into the trained object tracking model, and output the predicted tracking object position to achieve object tracking.
[0013] Optionally, the feature extraction structure includes a fused backbone network, an object encoding module, and a GhostAtFeat module;
[0014] The fused backbone network includes ResNet50 and SwinT-B;
[0015] The object encoding module is used to combine the Gaussian label and the ltrb encoding as training features, and the object encoding based on the bounding box will be integrated into the fused backbone network as training features;
[0016] The GhostAttFeat module plays a role in the feature conversion layer and is used to fuse and enhance the feature expression of the fused backbone network.
[0017] Optionally, the fused backbone network adopts the layer3 feature layer of the pre-trained ResNet50 as the CNN backbone SwinT-B as the ViT backbone Extract the CNN and ViT features of the image patch x im as follows:
[0018]
[0019] Adopt the stages [1, 2, 3] of the pre-trained SwinT-B to generate the feature map of stage 3, so Connect in the channel dimension [x c , x v , and adopt GhostAtFeat G att to obtain the final discriminative target feature, as follows:
[0020] e b = G att (cat([x c , x v ))
[0021] where e b is the discriminative backbone feature of the training frame image, x c is the feature map extracted by the CNN backbone, x v is the feature map extracted by the ViT backbone, and cat is the channel-wise concatenation operation.
[0022] Optionally, the target encoding module includes Gaussian labels and ltrb encoding. Given the target bounding box [x, y, w, h], the bounding box is converted into an ltrb distance map d r and a Gaussian classification label y g ;
[0023] The multi-layer perceptron layer θ maps the ltrb distance map d r to obtain the ltrb encoding, i.e., e d = θ(d r );
[0024] The Gaussian classification label y g is multiplied element-wise by the target token t g to produce the label encoding, i.e.,
[0025] The ltrb encoding and the Gaussian label are integrated into the fusion backbone network through element-wise addition to supervise the training feature learning, as follows:
[0026]
[0027] e test = e b '
[0028] where e b is the discriminative backbone feature of the training frame image, e d is the ltrb encoding, e g is the Gaussian label encoding, e b ' is the discriminative backbone feature of the test frame image, is the element-wise summation operation.
[0029] Optionally, the GhostAttFeat module functions in the feature transformation layer. First, it performs tracking by adjusting the classification features. Second, it reduces the feature channel dimension to decrease the number of parameters, and adopts the Ghost module to reduce the number of parameters, as follows:
[0030]
[0031] And further uses the Softmax function to ensure that the sum of the weighted graphs of the separated features [x1, x2] is 1.
[0032] Optionally, the attention feature enhancement structure integrates the attention module into the feature transformation layer to enhance the feature expression. The attention module includes channel attention and spatial attention;
[0033] For channel attention and spatial attention the weight graphs are as follows:
[0034]
[0035] Apply the weight graphs to the separated features [x1, x2] through element-wise product and addition to obtain the representative target features Adopt cascade and convolution fusion to reduce the dimension of the channel features, and also introduce residual connections to enhance the feature expression. Finally, use convolution and instance normalization operations to obtain the final discriminative backbone feature e b , as follows:
[0036]
[0037] Among them, is Conv3×3 with batch normalization, σ is the ReLU function, η is instance normalization, is Conv1×1 with batch normalization, is the dual-channel attention fusion feature, is the channel attention feature, is the spatial attention feature.
[0038] Optionally, the online update structure, based on ToMP, adopts the tracking concept of DCF, and uses the response map obtained by the target classifier and the target appearance features generated by the fusion backbone to jointly determine whether to update the samples.
[0039] Through the above technical solutions, compared with the prior art, the present invention provides a discriminative target tracking method based on a fusion backbone and effective attention, having the following beneficial effects:
[0040] (1) The present invention constructs an object tracking model based on a Transformer model predictor through a feature extraction structure, an attention feature enhancement structure, and an online update structure. In the feature extraction structure stage, fusing the backbone and the spatial and channel fusion attention module with Ghost can generate more discriminative feature representations for training a robust object tracking model.
[0041] (2) An online update structure is proposed, which combines the response map and the object appearance to maintain discriminative representation tracking. Brief Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0043] Figure 1 Schematic diagram of the object tracking model provided by the present invention;
[0044] Figure 2 Schematic diagram of the fusion backbone network and the target encoding module in the feature extraction structure provided by the present invention;
[0045] Figure 3 Schematic diagram of the GhostAttFeat module in the feature extraction structure provided by the present invention;
[0046] Figure 4 Schematic diagram of the online update structure provided by the present invention;
[0047] Figure 5 Success diagrams of LaSOT and LaSOText provided by the present invention, where 5a is the success diagram of LaSOT and 5b is the success diagram of LaSOText;
[0048] Figure 6 Comparison diagrams of DRTT proposed by the present invention and mainstream Transformer trackers on the LaSOT dataset, where 6a is the precision plot of the comparison of DRTT and mainstream Transformer trackers on the LaSOT dataset, 6b is the normalized precision plot of the comparison of DRTT and mainstream Transformer trackers on the LaSOT dataset, and 6c is the success rate plot of the comparison of DRTT and mainstream Transformer trackers on the LaSOT dataset;
[0049] Figure 7The figure shows the comparison of the visualization results of the DRTT tracker proposed in the present invention and the mainstream Transformer trackers ToMP50, OSTrack, and MixCVT on several challenging sequences in the LaSOT dataset. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] Refer to Figure 1 As shown, the present invention discloses a discriminative object tracking method based on a fused backbone and effective attention, including the following steps:
[0052] Obtain a visual object tracking dataset and perform preprocessing to obtain a training sample set;
[0053] Train an object tracking model to obtain a trained object tracking model. The object tracking model includes a feature extraction structure, an attention feature enhancement structure, and an online update structure;
[0054] Input the object to be tracked into the trained object tracking model, and output the predicted tracking object position to achieve object tracking.
[0055] Further, the feature extraction structure includes a fused backbone network, an object encoding module, and a GhostAtFeat module;
[0056] The fused backbone network includes ResNet50 and SwinT-B;
[0057] The object encoding module is used to combine the Gaussian label and ltrb encoding as training features, and the object encoding based on the bounding box will be integrated into the fused backbone network as training features;
[0058] The GhostAttFeat module plays a role in the feature transformation layer and is used to fuse and enhance the feature expression of the fused backbone network.
[0059] Further, the fused backbone network uses the layer3 feature layer of the pre-trained ResNet50 as the CNN backbone SwinT-B as the ViT backbone Extract the CNN and ViT features of the image patch x im as follows:
[0060]
[0061] Adopt the stages [1, 2, 3] of the pre-trained SwinT-B to generate the feature map of stage 3. Therefore Connect [x c , x v on the channel dimension, and adopt GhostAtFeat G att to obtain the final discriminative target feature, as follows:
[0062] e b = G att (cat([x c , x v ))
[0063] where e b is the discriminative backbone feature of the training frame image, x c is the feature map extracted by the CNN backbone, x v is the feature map extracted by the ViT backbone, and cat is the channel-wise concatenation operation.
[0064] Specifically, select ResNet50 and SwinT-B (representing the basic version of SwinT) as the fusion objects of the backbone network. Mainly use ResNet50 and SwinT-B as supplements, and only fine-tune layer3 of ResNet50. This not only utilizes the feature extraction capabilities of CNN and ViT, but also fine-tunes layer3 to adjust its tracking ability. As Figure 2 shown, the fusion backbone network consists of ResNet and SwinT. For the training frame, the target encoding module aims to encode the Gaussian label and the ltrb expression into a combination of training features to supervise network training, and finally use the GhostAttFeat layer to fuse and enhance the fusion backbone feature expression.
[0065] Furthermore, the target encoding module includes Gaussian label and ltrb encoding. Given the target bounding box [x, y, w, h], convert the bounding box into the ltrb distance map d r and the Gaussian classification label y g ;
[0066] The multi-layer perceptron layer θ maps the ltrb distance map d r to obtain the ltrb encoding, that is, e d = θ(d r );
[0067] The Gaussian classification label y g is multiplied element-wise by the target token t g to generate the label encoding, that is
[0068] Integrate the ltrb encoding and Gaussian labels into the fusion backbone network by element addition to supervise the training of feature learning, as follows:
[0069]
[0070] e test = e b ′
[0071] where e b is the discriminative backbone feature of the training frame image, e d is the ltrb encoding, e g is the Gaussian label encoding, e b ′ is the discriminative backbone feature of the test frame image, is the element-wise summation operation.
[0072] Furthermore, as Figure 3 shown, the GhostAttFeat module plays a role in the feature transformation layer. First, it performs tracking by adjusting the classification features. Second, it reduces the feature channel dimension to reduce the number of parameters, using the Ghost module g to reduce the number of parameters, as follows:
[0073] [x1,x2] = G(cat[x c ,x v )
[0074] And further use the Softmax function to ensure that the sum of the weighted graphs of the separated features [x1,x2] is 1.
[0075] Furthermore, the attention feature enhancement structure integrates the attention module into the feature transformation layer to enhance feature expression. The attention module includes channel attention and spatial attention;
[0076] For channel attention and spatial attention the weight graphs are as follows:
[0077]
[0078] Apply the weight graphs to the separated features [x1,x2] through element-wise product and addition to obtain the representative target features Adopt cascade and convolution fusion to reduce the dimension of the channel features. Also introduce residual connections to enhance feature expression. Finally, use convolution and instance normalization operations to obtain the final discriminative backbone feature e b , as follows:
[0079]
[0080]
[0081] Among them, is a Conv3×3 with batch normalization, σ is the ReLU function, η is instance normalization, is a Conv1×1 with batch normalization, is the dual-channel attention fusion feature, is the channel attention feature, is the spatial attention feature.
[0082] Furthermore, for the online update structure, based on ToMP, adopting the tracking concept of DCF, the response map obtained by the target classifier and the target appearance feature generated by the fusion backbone are used to jointly determine whether to update the sample.
[0083] Specifically, as Figure 4 shown, an online update sample mechanism is constructed using the inference result of the basic tracker, mainly including the peak of the response map and the target appearance, without training parameters to obtain appearance similarity Since the response map is only the result of the target classifier, appearance matching is introduced to use the final tracking bounding box and directly act on the fusion backbone features extracted during the tracking process. Based on the feature map obtained by the fusion backbone network in feature extraction, PrPool 4×4 and AvgPool 1×1 are used to convert the fusion backbone features into target appearance features, as follows:
[0084] T a = α1(ρ4(cat[x c ,x v ))
[0085] Among them, ρ4 is PrPool 4×4, and 4 is determined based on the size of the fusion backbone feature of the image patch x im and the search area factor, that is, 18 / 5; the average pooling α1 compresses the feature into a one-dimensional feature vector; the reference appearance feature T a ′ is obtained through the first-frame sample, and the cosine similarity is further used to measure the appearance similarity between T a and T a ′ As follows:
[0086]
[0087] Among them, || || is the L2 norm;
[0088] Combined with the peak of the response map to determine whether to update the dynamic sample, as follows:
[0089]
[0090] Among them, the fractional threshold π is default set to 0.9, and the target similarity of the second frame is used and the fractional threshold π to set the appearance threshold γ, that is When the condition is satisfied, the sample predicted by the current frame replaces the dynamic sample. The proposed online update mechanism does not require additional training costs. A ready-made feature extraction network is used to obtain the target appearance feature expression to calculate the appearance similarity, and combined with the response map to jointly determine whether to update the dynamic training sample.
[0091] In a specific embodiment, it includes the following content:
[0092] Overview of the Effective Attention Discriminative Representation Transformer Tracker (DRTT) of the present invention. The DRTT tracker uses a Transformer model predictor to obtain a DCF target classifier and an ltrb bounding box regressor. A ResNet and SwinT fusion backbone for image feature extraction is proposed, combined with Gaussian labels and ltrb encoding from the training frame bounding boxes. Specifically, lightweight Ghost and attention modules are introduced to enhance the fusion of backbone features and improve the expression of discriminative features.
[0093] To verify the effectiveness of the present invention, extensive experimental verifications were carried out on eight challenging benchmarks, including LaSOT, LaSOTex, TrackingNet, GOT-10k, OTB100, NFS, UAV123, and TNL2K.
[0094] Training experiment details:
[0095] The tracker is trained on the training sets of LaSOT, TrackingNet, GOT-10k, and COCO2014. 40k subsequences are sampled on two NVIDIA GeForce RTX 3090 GPUs and trained for 300 epochs. Specifically, the image patch size is 288, the search scale factor is 5, the encoder-decoder of the transformer has 6 layers, and the AdamW optimizer is used. The learning rate of AdamW is 1e-4, while the learning rate of the backbone network is 2e-5. After 150 and 250 epochs and a weight decay of 1e-4, the learning rate decays by 0.2 times. Two training frames and a test frame with an interval of 200 frames are randomly selected from a video sequence to construct a training subsequence. Then, image patches are extracted after randomly translating and scaling the images. In addition, random image flipping and color jittering are used to augment the data.
[0096] Test experiment details: To verify the practicality of the present invention in the discriminative tracking framework, performance comparisons were conducted on 8 challenging benchmarks, including LaSOT, LaSOText, TrackingNet, GOT-10k, OTB100, NFS, UAV123, and TNL2K. Implemented in Python using PyTorch, all test experiments were run on a GPU processor of GeForce RTX 4070.
[0097] Ablation experiments:
[0098] GhostAttFeat: The original motivation of the present invention was to introduce an attention mechanism to enhance the discriminative representation ability of backbone features. Since the present invention is a study on a DCF-based discriminative tracker, SuperDiMP was a highly competitive tracker before introducing the tracker in Transformer, and it used ResNet50 as the backbone network. Based on the optimized model predictor, SuperDiMP only needed to be trained for 50 epochs to converge. Therefore, the attention module was first integrated into the classification feature layer based on SuperDiMP to verify its impact on the tracking performance, as shown in Table 1. The attention method mainly follows the design concepts of CBAM and SE. Therefore, the tracking results of these two classic attention modules were compared, and the GhostAttFeat module was proposed in Figure 3 It. GhostAttFeat without residual connections was also tested. It can be seen from the experimental results that adding the attention module affects the tracking results, but the unimproved CBAM and SE perform poorly in the LaSOT benchmark test. Finally, it was determined that the GhostAtFeat module with residual connections was generally the optimal one.
[0099] After testing the effectiveness of the proposed GhostAtFeat module on SuperDiMP, it was further integrated into the ToMP framework for further ablation experiments, successively called AttDiMP and AttToMP. As shown in Table 2, comparative experiments on seven datasets were provided. It can be seen from Table 2 that AttDiMP is almost comprehensively superior to the baseline method, and AttToMP performs significantly in the first four datasets. For example, the AO gain of GOT-10k is 1.7%. The experiments show that the proposed GhostAtFeat module can enhance the learning of discriminative representation features.
[0100] Table 1 Analyze different attention modules and their impact on the baseline tracker (i.e., SuperDiMP) according to the area under the curve (AUC) score
[0101]
[0102]
[0103] Table 2 analyzes different attention modules and their impact on the AUC or average overlap (AO) scores of the SuperDiMP and ToMP50 trackers.
[0104] LaSOT GOT-10k TrackingNet TNL2K OTB100 NFS UAV123 SuperDiMP 63.1 67.0 78.1 50.7 70.1 64.8 67.7 AttDiMP <![CDATA[63.4 +0.3 > 66.9 <![CDATA[78.3 +0.2 > 50.5 <![CDATA[71.1 +1.0 > <![CDATA[66.5 +1.7 > <![CDATA[69.1 +1.4 > ToMP50 67.6 72.0 81.2 54.1 70.1 66.9 69.0 AttToMP <![CDATA[68.3 +0.7 > <![CDATA[73.7 +1.7 > <![CDATA[81.9 +0.7 > <![CDATA[54.7 +0.6 > 69.6 66.4 67.8
[0105] Fusion backbone: Experiments were conducted using the mainstream SwinT and ResNet backbones. ToMP has two versions that use ResNet50 and ResNet101 respectively. During actual testing, it was found that ResNet50 had the best performance, so ToMP50 was chosen as the baseline. Since the training cost of SwinT is very high, the basic version of SwinT was used without updating its pre-trained weights. To prove that the fusion backbone is more effective, ResNet50 was also replaced with SwinT-base. As shown in Table 3, the weights of SwinT were completely frozen, and only SwinT was used to extract features, which led to a significant drop in performance. Compared with ToMP50, this approach reduced the AUC score of LaSOT by 2.2%. Referring to the fine-tuning of the third layer of ResNet50 in ToMP50, the comparison results of the fine-tuning of the second stage of the SwinT layer were tested. As can be seen from Table 3, the performance was comparable to that of ResNet50, but the fine-tuning cost of SwinT was higher and no significant results had been achieved. It is speculated that the combination of CNN + Transformer might be more suitable. Further attempts were made to extract features in parallel with ResNet and SwinT. However, fine-tuning SwinT consumed too much GPU memory. Therefore, ResNet50 was regarded as the main backbone and SwinT as a supplement, which is the origin of the proposed fusion backbone module. As can be seen from Table 3, the proposed fusion backbone achieved 69.1% and 68.7% on NFS and LaSOT respectively, with the AUC increasing by 2.2% and 1.1% respectively. Finally, it was determined that the fusion backbone could significantly improve the tracking performance.
[0106] Table 3 analyzes different backbones and their impact on the ToMP50 tracker according to the AUC score.
[0107] OTB100 NFS LaSOT ToMP50 70.1 66.9 67.6 + Frozen SwinT 67.7 64.6 65.4 + SwinT 64.8 66.4 67.7 + Integrated backbone 68.4 69.1 68.7
[0108] In particular, GhostAtFeat and the fusion backbone are further used to enhance discriminative feature learning. As shown in Table 4, ablation experiments are conducted using FB (Fusion Backbone), GA (GhostAttFeat), and the combination of these two. Compared with ToMP50, the improved tracker has made great progress in the LaSOT, GOT-10k, TrackingNet, and TNL2K benchmark tests. ToMP50+FB+GA achieves 69.2, 74.2, 82.4, 55.7 (%) with performance improvements of 1.6, 2.2, 1.2, and 1.6% respectively. In addition, better performance improvements are also achieved for NFS and UAV123.
[0109] It can be seen from the quantitative results that both GhostAtFeat and the fusion backbone can improve the performance of the basic tracker. In addition, the combination of the two can achieve stronger performance improvement.
[0110] Table 4 analyzes the fusion backbone and GhostAttFeat module, and their impact on the ToMP50 tracker in terms of AUC or AO score
[0111] FB GA LaSOT GOT-10k TrackingNet TNL2K OTB100 NFS UAV123 - - 67.6 72.0 81.2 54.1 70.1 66.9 69.0 √ - <![CDATA[68.7 +1.1 > <![CDATA[74.1 +2.1 > <![CDATA[82.4 +1.2 > <![CDATA[54.6 +0.5 > 68.4 <![CDATA[69.1 +2.2 > 69.0 √ <![CDATA[68.3 +0.7 > <![CDATA[73.7 +1.7 > <![CDATA[81.9 +0.7 > <![CDATA[54.7 +0.6 > 69.6 66.4 67.8 √ √ <![CDATA[69.2 +1.6 > <![CDATA[74.2 +2.2 > <![CDATA[82.4 +1.2 > <![CDATA[55.7 +1.6 > 69.2 <![CDATA[67.4 +0.5 > <![CDATA[69.3 +0.3 >
[0112] Online update: It has been demonstrated above that the fusion backbone and effective attention modules can enhance the discriminative tracking representation. Since the tracking target is complex and variable in the wild, it is wise to adopt an online update strategy. Different from the mainstream Siamese Transformer tracking methods that require secondary training, the online update module designed in the present invention is simpler. Only the response map of the DCF object classifier and the feature extraction module of the original tracking network are used to obtain the classification score and the target appearance similarity score respectively, and jointly determine whether to update. As Figure 5 shown, an online update method based on calculating the similarity of the target appearance is proposed, which has a block effect. Among them, 5a is the success map of LaSOT, and 5b is the success map of LaSOText. Compared with the online update mechanism that does not use the target appearance module (i.e., ToMP50+FB+GA), the method of the present invention achieves AUC gains of 0.4 and 0.7 (%) on LaSOT and LaSOText respectively. In addition, ToMP50 using only the FB module can achieve an AUC score of 49.9% and a relative gain of 3.2%.
[0113] A series of ablation experiments have been conducted on multiple datasets to verify the usability of the method proposed by the present invention. The experimental results show that the method proposed by the present invention has good improvement effects. A novel Transformer tracker (i.e., the object tracking model) is constructed based on the feature extraction structure, the attention feature enhancement structure, and the online update structure, aiming to achieve and maintain the Transformer tracking of discriminative representations, so it is called the DRTT tracker. The DRTT tracker proposed by the present invention is compared with the state-of-the-art (SOTA) trackers on four challenging tracking benchmarks.
[0114] LaSOT: Table 5 shows the comparison results of the precision, normalized precision, and AUC scores of various trackers. Among them, the evaluation metrics include precision (P), normalized precision (NP), AUC, success rate (SR), and AO (%). The LaSOT dataset contains most of the longer video sequences, with an average of 2500 frames per video sequence. The feature extraction structure, the attention feature enhancement structure, and the online update structure proposed by the present invention can enhance the discriminative ability of the target features. The recent DiMP, TransT, STARK, and ToMP50 are considered for comparison. The method proposed by the present invention integrated into ToMP50 achieves competitive performance, exceeding ToMP50 with a relative gain of 2.0%. In addition, compared with the pure Transformer-based trackers, such as MixFormer, OSTrack, ARTrack, and MixCVT, the DRTT of the present invention shows competitive performance. Figure 6 The comparison graph of the DRTT proposed by the present invention and the mainstream Transformer trackers on the LaSOT dataset is also provided. Among them, 6a is the precision plot of the comparison between the DRTT and the mainstream Transformer trackers on the LaSOT dataset, 6b is the normalized precision plot of the comparison between the DRTT and the mainstream Transformer trackers on the LaSOT dataset, and 6c is the success rate plot of the comparison between the DRTT and the mainstream Transformer trackers on the LaSOT dataset. Since SuperDiMP has an important relationship with the tracker of the present invention, it is also included in the comparison. Generally speaking, the method proposed by the present invention significantly improves the tracking performance.
[0115] Table 5 Comparison results of the DRTT of the present invention with the state-of-the-art trackers on the LaSOT, LaSOText, TrackingNet, and GOT-100k datasets
[0116]
[0117]
[0118] LaSOText: LaSOText is an extended subset of LaSOT, which includes 150 additional videos and 15 new categories. There are many similar interfering objects and fast-moving small objects in these videos, which greatly increases the difficulty of tracking. As shown in Table 5, the DRTT performance of the present invention is higher than that of state-of-the-art trackers. It can be observed that the DRTT of the present invention has 2.4% and 3.4% higher AUC scores than OSTrack and ARTrack, respectively. In addition, compared with ToMP50, it can be seen that the DRTT of the present invention promotes ToMP50 with a relative AUC gain of 3.1%.
[0119] TrackingNet: The test set of the TrackingNet dataset is also used to test the method of the present invention, which consists of 511 video sequences. The experimental results are shown in Table 5. The AUC score of the DRTT tracker of the present invention is 82.4%, which is better than state-of-the-art trackers such as TransT, STARK, and ToMP50. In addition, compared with ToMP50, it can be seen that the DRTT of the present invention promotes ToMP50, with a relative AUC gain of 1.2%, and the precision and normalized precision increase by 1.9 and 1.0 (%) respectively. The competitive results on this dataset further prove the effectiveness of the method proposed by the present invention. However, it can also be seen that there is still room for improvement in the present invention compared with single-stream converter trackers (such as ARTrack).
[0120] GOT-10k: This dataset is a large-scale and highly diverse benchmark for general object tracking in the wild. The test set of GOT-10k is used to test the tracker of the present invention. In addition, the tracker is trained on the GOT-10k train. As shown in Table 5. The DRTT of the present invention obtains an AO score of 74.7%, exceeding all comparison methods. More importantly, the DRTT of the present invention performs better than ToMP50, with a relative AO gain of 2.7%. The experiment further proves that the DRTT of the present invention is feasible.
[0121] Qualitative analysis:
[0122] The above ablation and SOTA experiments quantitatively verify the effectiveness of the method proposed by the present invention, and the DRTT tracker can effectively improve the tracking results. In addition, the tracking effects of representative video sequences in Figure 7 are qualitatively provided. The tracking results of ToMP50, OSTrack, MixCVT, and the DRTT tracker of the present invention on six challenging sequences in the LaSOT dataset are visualized. As Figure 7As shown, the number of video frames of the six video sequences are 1575, 3120, 1721, 5009, 2967, and 2964 respectively. The present invention shows four representative frames, and in each frame, the ground truth of the tracked target is shown using a green bounding box. From Figure 7 It can be seen that despite occlusion, deformation, and background clutter, the DRTT tracker of the present invention can still accurately predict the target object. In the 1250th frame of the bird-15 sequence, the 1707th frame of coin-6, and the 0662nd frame of person-5, MixCVT misidentifies similar distractors as the target, but the DRTT of the present invention can also obtain an appropriate bounding box. In the person-5 sequence, it is noted that the tracker of the present invention can track the actual dancer in the presence of distractors, but ToMP50, OSTrack, and MixCVT are interfered by the distractors and background clutter. In the 3002nd frame of the dog-15 sequence, ToMP misidentifies the dog in the mirror as the target, resulting in inaccurate tracking results, while the tracker of the present invention has a more powerful discrimination function and can effectively prevent inaccurate tracking. In the #1030th frame of the crab-12 sequence, when the crab target is about to disappear, MixCVT will be severely affected. In addition, the color of the sand is similar to that of the crab, which also affects the tracking accuracy. In the zebra-17 sequence, ToMP50 is not robust enough and may be affected by similar objects and self-deformation. In the 0263rd frame, most trackers gradually track the background information, while ToMP50 only tracks a part of other targets. However, although the tracker of the present invention is closer to the tracking area of the ground truth, updating the model as usual at this time will also lead to contamination of the tracking model. Therefore, the online update module is particularly important. As can be seen from the 2032nd frame, the compared trackers still maintain target tracking, but the tracking bounding box is not accurate enough, and in complex scenarios, surrounding objects are easily misidentified as targets. In the 2463rd frame, ToMP50 only tracks part of the target, while the DRTT, MixCVT, and OSTrack of the present invention can accurately track the target. In summary, although the tracker of the present invention has excellent robustness against complex situations such as deformation, distractors, and background clutter, a gap with the marked bounding box is also found.
[0123] Although the present invention is verified by examples based on the ToMP tracker, it is equally applicable to other discriminative trackers, deep learning trackers such as Transformer trackers. The target tracking technologies formed by adopting similar practices of fusing the backbone, effective attention, and online update of the present invention are all within the protection scope of the present invention.
[0124] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0125] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A discriminative object tracking method based on fused backbone and effective attention, characterized in that, It includes the following steps: Obtain a visual object tracking dataset and perform preprocessing to obtain a training sample set; Train an object tracking model to obtain a trained object tracking model, where the object tracking model includes a feature extraction structure, an attention feature enhancement structure, and an online update structure; Input the object to be tracked into the trained object tracking model, and output the predicted tracking object position to achieve object tracking.
2. The discriminative object tracking method based on fused backbone and effective attention according to claim 1, wherein The feature extraction structure includes a fused backbone network, an object encoding module, and a GhostAtFeat module; The fused backbone network includes ResNet50 and SwinT-B; The object encoding module is used to combine the Gaussian label and ltrb encoding as training features, and the object encoding based on the bounding box will be integrated into the fused backbone network as training features; The GhostAttFeat module plays a role in the feature transformation layer and is used to fuse and enhance the feature expression of the fused backbone network.
3. The discriminative object tracking method based on fused backbone and effective attention according to claim 2, wherein The fusion backbone network uses the layer3 feature layer of the pre-trained ResNet50 as the CNN backbone SwinT-B is used as the ViT backbone Extract the image patch x im The CNN and ViT features are as follows: Adopt the stages [1, 2, 3] of the pre-trained SwinT-B to generate the feature maps of stage 3. Therefore, Connect [x c , x v on the channel dimension, and adopt GhostAtFeat G att to obtain the final discriminative target features as follows: e b = G att (cat([x c , x v )) Among them, e b is the discriminative backbone feature of the training frame image, x c is the feature map extracted by the CNN backbone, x v is the feature map extracted by the ViT backbone, and cat is the channel-wise concatenation operation.
4. The discriminative object tracking method based on fused backbone and effective attention according to claim 2, wherein The target encoding module includes Gaussian labels and ltrb encoding. Given the target bounding box [x, y, w, h], the bounding box is converted into an ltrb distance map d r and a Gaussian classification label y g ; The multi-layer perceptron layer θ maps the ltrb distance map d r to obtain the ltrb encoding, namely e d = θ(d r ); Gaussian classification label y g Element-wise multiply with target label t g To generate label encoding, i.e., Integrate the ltrb encoding and the Gaussian label into the fused backbone network by element-wise addition to supervise the training feature learning, as follows: e test = e b ′ Among them, e b is the discriminative backbone feature of the training frame image, e d is the ltrb encoding, e g is the Gaussian label encoding, e b ' is the discriminative backbone feature of the test frame image, is the element-wise summation operation.
5. The discriminative object tracking method based on fused backbone and effective attention according to claim 2, wherein The GhostAttFeat module plays a role in the feature transformation layer. First, it tracks by adjusting the classification features. Second, it reduces the number of parameters by decreasing the feature channel dimension and adopts the Ghost module to reduce the number of parameters, as shown below: And further use the Softmax function to ensure that the sum of the weighted graphs of the separated features [x1, x2] is 1.
6. A discriminative object tracking method based on fused backbone and effective attention according to claim 1, It is characterized in that The attention feature enhancement structure integrates the attention module into the feature transformation layer to enhance the feature expression, and the attention module includes channel attention and spatial attention; For channel attention and spatial attention the weight maps are as follows: Apply the weight map to the separated features [x1, x2] through element-wise product and addition to obtain the representative target features Adopt cascade and convolutional fusion to reduce the dimension of channel features, also introduce residual connections to enhance feature representation, and finally use convolutional and instance normalization operations to obtain the final discriminative backbone feature e b , as follows: Among them, is a Conv3×3 with batch normalization, σ is the ReLU function, η is instance normalization, is a Conv1×1 with batch normalization, is the dual-channel attention fusion feature, is the channel attention feature, is the spatial attention feature.
7. The discriminative object tracking method based on fused backbone and effective attention according to claim 1, wherein The online update structure, based on ToMP, adopts the tracking concept of DCF, and uses the response map obtained by the object classifier and the object appearance features generated by the fused backbone to jointly determine whether to update the sample.