An end-to-end single-object tracking method and device based on a hybrid attention mechanism

By constructing a Transformer tracking framework MixFormer based on a hybrid attention mechanism, unified feature extraction and information fusion, the problem of insufficient feature extraction capabilities in the existing technology is solved, and a more efficient target tracking effect is achieved.

CN114550040BActive Publication Date: 2025-07-25NANJING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202210152336.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-07-25
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

The existing end-to-end target trackers rely on convolutional networks for feature extraction, resulting in limited feature extraction capabilities and complex design of information fusion modules, making it difficult to adapt to challenges such as target deformation and occlusion.

Method used

A Transformer tracking framework based on a hybrid attention mechanism is constructed. Through the hybrid attention module MAM unified feature extraction and information fusion, self-attention and mutual attention operations are adopted, combined with regression heads and classification heads, to achieve end-to-end target tracking.

Benefits of technology

It improves the accuracy and robustness of target tracking, can better adapt to target deformation, and improves the accuracy and tracking success rate of object regression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550040B_ABST
    Figure CN114550040B_ABST
Patent Text Reader

Abstract

An end-to-end single-object tracking method based on a hybrid attention mechanism constructs a tracking framework MixFormer based on Transformer tracking for object tracking. The construction of the tracking framework includes the following steps: 1) data preparation stage; 2) network configuration stage; 3) offline training stage; 4) online tracking stage. The present invention adopts a backbone network based on hybrid attention to simultaneously perform feature extraction and target information fusion, obtaining a concise and clear tracking framework and effectively improving performance. In addition, the tracking method of the present invention has better adaptability to object deformation during the tracking process, effectively improving the accuracy of target regression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer software, relates to single-object tracking technology, and specifically relates to an end-to-end single-object tracking method and device based on a hybrid attention mechanism. Background Art

[0002] As a basic task in computer vision, visual object tracking aims to estimate the spatial position of an arbitrary general object in a video at each frame and mark the object's bounding box. Although significant progress has been made in object tracking, how to design a simple and effective end-to-end tracker remains a challenge. The main challenges come from scale changes, object deformation, occlusion, and confusion from similar objects.

[0003] Current popular trackers usually include three components to complete the tracking task: (1) a CNN backbone network for extracting general features of the target to be tracked and the search region; (2) a fusion module for communicating information between the tracking target and the search region for subsequent target-aware localization; (3) a precise localization bounding box module for generating the final tracking result. Among them, the fusion module is the most critical part of the tracking algorithm, which is responsible for integrating the target information to be tracked into the features of the search region, so as to produce a specific box according to a specific tracking target. Traditional information fusion methods include correlation-based operations and online model update algorithms. Recently, due to the global and dynamic modeling capabilities of Transformer, it has been introduced into the tracking field for information interaction between the tracking target and the tracking region, and good tracking performance has been achieved. It mainly uses the Transformer model to perform feature fusion on the target features and the search region, and then predicts the fused features to achieve tracking. However, these Transformer-based trackers still rely on a convolutional backbone network for feature extraction and only apply attention operations in a relatively high-level and abstract representation space. However, the representation ability of the convolutional backbone network is limited. First, it is usually pre-trained based on general object recognition tasks, and second, it may ignore more fine-grained structural information for tracking. Summary of the Invention

[0004] The problem to be solved by the present invention is: how to design a concise end-to-end object tracking framework that does not rely on a convolutional network for feature extraction and can further unify the feature extraction and information fusion modules.

[0005] The technical solution of the present invention is as follows: An end-to-end single-object tracking method based on a hybrid attention mechanism, constructing a tracking framework MixFormer for object tracking. The tracking framework MixFormer is a Transformer tracking network trained end-to-end, including a backbone network and a tracking head. The construction of the tracking framework MixFormer includes the following stages:

[0006] 1) Data preparation stage: Crop the target search area from all video frames in the training dataset. Extract two frames from the first half of the frame sequence of each video as template frames, and extract one frame from the second half as a test frame. Label the target box for the test frame as a validation frame, and use the diagonal coordinates of the target box in each validation frame as the ground truth label during the offline training process.

[0007] 2) Network configuration stage: The backbone network is a feature extractor based on a hybrid attention module, which unifies feature extraction and information fusion through the Transformer structure. The tracking head is a regression head implemented using a convolutional network. Input the template frame and the test frame into the backbone network simultaneously to generate the test frame features fused with the template information, and then generate the diagonal coordinates of the target through the regression head as the final target box generated by the test frame.

[0008] Among them, the backbone network is based on the hybrid attention mechanism, performing self-attention and cross-attention operations on the features of the template frame and the test frame. Self-attention is used to extract the self-features of the template frame and the test frame, and cross-attention is used for feature information interaction between the target frame and the test frame to obtain the test frame features fused with the template information.

[0009] 3) Offline training stage: For the training of the regression head target box, use the L1 loss function and the GIoU loss function for supervision. Combine the ground truth label obtained from the validation frame, use the AdamW optimizer, and update the entire network parameters through the backpropagation algorithm. Continuously train the configured network until the number of iterations is reached to obtain the tracking framework MixFormer.

[0010] Online tracking: Label the target search area of the first frame of the video to be tracked as the template frame, and the subsequent frames as test frames. Input them into the trained tracking framework MixFormer, and output the target box on the test frame to achieve object tracking.

[0011] Furthermore, the cross-attention operation of the backbone network only performs one-way cross-attention from the template frame to the test frame, without performing cross-attention from the test frame to the template frame, to obtain the test frame features fused with the template information.

[0012] Further, the tracking head further includes a classification head which is used to obtain the classification target confidence of the test frame. The classification head has a preset learnable confidence vector, which performs attention operations on the test frame features and the template frame's own features respectively, perceives the information of both, and predicts the classification target confidence of the current test frame. During the tracking process, frames with qualified confidence are selected from the sequence of already tracked video frames and supplemented as template frames.

[0013] During online tracking, first, the target search area in the first frame image of the video to be tracked is cropped as the template frame F train , and the frame to be tracked is used as the test frame F test , and the target box on the test frame F is obtained through the tracking framework MixFormer. During the tracking process, from the sequence of already tracked frames, one frame with the highest confidence and its tracked target box are selected every N frames as labels and supplemented as the template frame F test . train .

[0014] The present invention constructs a neat and effective tracking framework, which only includes a backbone network for simultaneous feature extraction and information fusion and a tracking head. This coupling paradigm of the tracking framework of the present invention has the following advantages. First, it will make our feature extraction more adaptable to specific tracking targets and capture more discriminative features related to the target. In addition, it also allows for the fusion of target information at more scales, thus better capturing the correlation between the target and the search area.

[0015] Based on the above tracking method, the present invention also provides an end-to-end single-object tracking device based on a hybrid attention mechanism, which has a computer storage medium. A computer program is configured in the computer storage medium, and the computer program is used to implement the above tracking framework MixFormer. When the computer program is executed, the above tracking method is implemented.

[0016] The present invention has the following advantages compared with the prior art.

[0017] The present invention proposes an end-to-end single-object tracking method based on a hybrid attention mechanism, constructs a tracking framework MixFormer based on Transformer, and adopts a specially designed transformer backbone network, that is, a feature extractor based on the hybrid attention module MAM to simultaneously perform feature extraction and target information fusion. As Figure 2 shown, first, the concatenated vector of the target frame and the test frame is split and respectively reshaped into a 2D vector, then passed through a multi-head attention function, the two generated 2D vectors are concatenated and passed through a linear layer to obtain the test frame features fused with the template information. Finally, as Figure 1As shown, through two simple regression heads and classification heads, the tracking target box is obtained, and the tracking label is further updated by supplementing the online tracking results, resulting in a concise and clear tracking framework that can effectively improve the tracking accuracy.

[0018] The present invention designs an online-updatable template sample space, and during the tracking process, a confidence prediction module is used to screen the template samples that are more suitable for the current tracking, thereby improving the robustness of the model. Compared with existing tracking methods, the online tracking method of the present invention has better adaptability to the deformation of objects during the tracking process and effectively improves the accuracy of target regression.

[0019] The present invention has achieved good accuracy in the visual object tracking task and improved the accuracy of object regression. Compared with existing methods, the MixFormer tracking method proposed by the present invention has achieved the best tracking success rate and positioning accuracy in multiple visual tracking test benchmark datasets (LaSOT, TrackingNet, GOT-10k, VOT2020, UAV123). Brief Description of the Drawings

[0020] Figure 1 It is a schematic diagram of the tracking framework MixFormer of the present invention.

[0021] Figure 2 It is a schematic diagram of the hybrid attention module MAM of the backbone network in the present invention.

[0022] Figure 3 It is a schematic diagram of the confidence prediction module SPM of the classification head of the present invention. Detailed Implementation Manner

[0023] The present invention proposes a tracking framework MixFormer, aiming to unify the feature extraction and information fusion modules through the Transformer structure. The attention module is a very flexible architecture building block with dynamic and global modeling capabilities, with few assumptions about the data structure, and can be generally applied to different types of relationship modeling. The core idea of the present invention is to utilize the flexibility of this attention operation and propose a hybrid attention module MAM, as Figure 2As shown in the figure, the module first splits the concatenated vectors of the target template frame and the test frame and reshapes them into a 2D vector respectively, and then passes through a multi-head attention function. The two generated 2D vectors are concatenated and passed through a linear layer. Repeating the MAM module multiple times can obtain a feature extractor based on MAM. Through multiple serial MAM modules, the network depth is deepened. This module simultaneously performs feature extraction and information interaction on the target template and the search area. In MAM, the present invention designs a hybrid interaction scheme, which performs self-attention and mutual-attention operations on the features from the target template and the search area, that is, the template frame and the test frame. Self-attention is responsible for extracting the self-features of the target template or the search area, while mutual-attention ensures the communication between them to mix the target and search area information. In addition, in order to reduce the computational cost of MAM and allow the use of multiple templates to handle problems such as online target deformation, we further propose a specific asymmetric attention mechanism, that is, during the mutual-attention process of MAM, only one-way mutual-attention from the template frame to the test frame is performed, and no mutual-attention from the test frame to the template frame is performed.

[0024] An end-to-end single-object tracking method based on a hybrid attention mechanism proposed by the present invention is offline trained on four training data sets, namely TrackingNet-Train, LaSOT-Train, COCO-Train, and GOT-10k-Train, and tested on test sets such as UAV123, VOT2020, LaSOT -Test , TrackingNet -Test, GOT-10k Five to achieve high accuracy and tracking success rate, and is specifically implemented using the Python 3.7 programming language and the Pytorch 1.7 deep learning framework.

[0025] Figure 1 is a schematic diagram of the tracking framework of the method of the present invention. Through the designed end-to-end trained Transformer tracking network, the target box of the object to be tracked is directly obtained, thereby realizing the target tracking task. The whole method includes a data preparation stage, a network configuration stage, an offline training stage, and an online tracking stage. The specific implementation steps are as follows:

[0026] 1) The data preparation stage, that is, the stage of generating training examples. During the offline training process, training examples are generated. First, jitter processing is performed on the target area of each frame image in the offline training data set, and then the target search area after jitter processing is cropped. Three frames are extracted from the first half of each video frame sequence as training frames, and one frame is extracted from the second half of each video frame sequence as a test frame. The target box is labeled for the test frame as a verification frame. For the coordinates of the upper left corner point and the lower right corner point of the target box in each verification frame, they are used as the true labels during the offline training process.

[0027] 2) Network configuration stage, that is, the configuration stage of tracking the network. The overall structure and process of the network of the present invention are very concise compared with other trackers and are divided into three parts. The entire framework and process are as Figure 1 shown, and the specific operations are as follows.

[0028] 2.1) Extract test frame features dependent on the tracking template: Given T-frame tracking templates and a test frame, the T-frame tracking templates are concatenated into a template frame. The input sizes of the template frame and the test frame are T×128×128×3 and 320×320×3 respectively. First, a convolutional layer with a kernel size of 7 and a stride of 4 is used to generate overlapping block vectors. Next, the obtained block vectors are flattened and concatenated to produce a concatenated vector F token , and then input into the Mixed Attention Module (MAM) to generate a mixed vector that fuses the tracking target information and the test frame information. The specific structure of MAM is as Figure 2 shown. First, the concatenated vector F token is split and Reshape operations are performed. Self-attention operations are carried out to obtain the self-features of the template frame and the test frame. Then, the self-features of the two are respectively passed through a common multi-head attention function to obtain their respective query, key, and value. Then, as Figure 2 shown, parallel attention operations are performed. Finally, after concatenating the two and passing through a linear layer to achieve interaction, it is added to the initial vector F token to obtain a once-mixed vector. By repeating the above operations M times, the mixed vector is split and Reshape, and then self-attention and mutual-attention operations are performed to deepen the network depth. The final mixed feature is obtained. Then, only the test frame corresponding features in the mixed feature need to be segmented and Reshape operations are performed to obtain the test frame features F test fused with template information, with a size of 20×20×384.

[0029] 2.2) Obtain the tracking box of the target in the test frame: The tracking regression head uses a convolutional network. Five simple convolutional layers act on the F test obtained in step 2.1). The number of input channels is 384, and the number of output channels is 2. Heatmaps of the upper left corner and the lower right corner of the target are obtained respectively, each with a size of 20×20×1. Finally, the coordinates of the upper left corner and the lower right corner are obtained by taking the maximum value of the heatmap, thereby obtaining the target box, that is, the tracking box of the target.

[0030] 2.3) Obtain the classification confidence of the test sample: The tracking framework of the present invention also sets a classification head in the tracking head for online tracking. The F testThrough a classification confidence prediction module SPM (Score Prediction Module), the classification confidence of each test frame can be obtained, that is, whether each test frame is a positive sample. The structure of the SPM is as Figure 3 shown. Through a preset learnable confidence vector, attention operations are respectively performed on the test frame features and the target template frame's own features to perceive the information of both, so as to predict the final classification confidence, which is used to select higher-quality online samples as the template frame for tracking in the online tracking stage.

[0031] The following specifically illustrates the network configuration stage with an embodiment. Using the above-mentioned backbone network based on MAM, the parameters of the ImageNet pre-trained model are loaded into the network, and features dependent on the target template are extracted from the test frame. The size of the feature map is 20×20×384, which respectively represents the length, width, and number of channels of the feature map. Next, this feature map is input into the regression head and the classification head SPM to obtain the final target box and the classification confidence of this target box respectively. This target box can be used as the tracking result, and the classification confidence is used to select online samples.

[0032] 3) In the offline training stage, cross-entropy is used as the loss function for the offline training of the classification branch, and the GIoU loss function and the L1 loss function are used for the regression branch. The AdamW optimizer is used, the single-card BatchSize is set to 32, the total number of training epochs is set to 500, the learning rate is divided by 10 at the 400th epoch, and it is trained on 8 Nvidia Tesla V100s. The entire network parameters are updated through the backpropagation algorithm, and steps 2.1) to 2.3) are continuously repeated until the number of iterations is reached.

[0033] 4) In the online tracking stage, the first frame of the video to be tracked is labeled with the target search area as the template frame, and the subsequent frames are used as test frames, which are input into the trained network to obtain the tracking target box on the basis of the initial parameters.

[0034] As a preferred method, first, the target search area in the first frame image of the video to be tracked is cropped as the template frame, which is used as the initial target template. The frame to be tracked in the video to be tracked is used as the test frame. The target template and the test frame are input into the network in step 2) to obtain the target box on the test frame. During the tracking process, from the sequence of frames that have been tracked, a frame with the highest classification confidence obtained by the SPM and its tracked target box are selected every 20 frames as labels and added to the online target template set as the template frame to achieve online target tracking with self-adaptability to the deformation of the video target.

[0035] On the test dataset, the tracking efficiency is 22fps (Nvidia GTX 1080Ti). In terms of tracking accuracy, the Auc reaches 70.5% on the GOT-10k dataset; on the LaSOT dataset, the Auc reaches 69.5%; on the TrackingNet dataset, the Auc reaches 83.6% and the Pre reaches 82.8%; on the VOT2020 dataset, the EAO reaches 0.550, the Robustness reaches 0.843, and the Accuracy reaches 0.760. On the UAV123 dataset, the Auc reaches 70.4%. The metrics on the above test datasets exceed the current best-performing method.

Claims

1. An end-to-end single-object tracking method based on a hybrid attention mechanism, characterized in that Construct a tracking framework MixFormer for object tracking. The tracking framework MixFormer is an end-to-end trained Transformer tracking network, including a backbone network and a tracking head. The construction implementation of the tracking framework MixFormer includes the following stages: 1) Data preparation stage: Crop the target search area from all video frames in the training dataset. Extract two frames from the first half of the frame sequence of each video as template frames, and extract one frame from the second half as a test frame. Label the target bounding box for the test frame as a validation frame, and use the diagonal coordinates of the target bounding box in each validation frame as the ground truth label during the offline training process; 2) Network configuration stage: The backbone network is a feature extractor based on a hybrid attention module, which unifies feature extraction and information fusion through the Transformer structure. The tracking head is a regression head implemented using a convolutional network. Input the template frame and the test frame into the backbone network simultaneously to generate the test frame features fused with template information, and then generate the diagonal coordinates of the target through the regression head as the final target bounding box generated by the test frame; Among them, the backbone network is based on a hybrid attention mechanism, which performs self-attention and cross-attention operations on the features of the template frame and the test frame. Self-attention is used to extract the self-features of the template frame and the test frame, and cross-attention is used for the feature information interaction between the target frame and the test frame to obtain the test frame features fused with the template information. Specifically, the backbone network generates block vectors for the template frame and the test frame respectively, performs self-attention operations to obtain the self-features of the template and the test frame, passes them through a common multi-head attention function to obtain their respective queries, keys, and values, and then performs cross-attention operations. After concatenating the queries, keys, and values of the two and passing them through a linear layer, the block vectors of the template frame and the test frame are flattened and concatenated to obtain a concatenated vector F token , the output of the linear layer is added to F token to obtain a once-mixed vector. After splitting the mixed vector, self-attention and cross-attention operations are performed again to obtain a new mixed vector. This process is repeated M times to obtain the final mixed feature, and after splitting and reshaping, the test frame features fused with the template information are obtained; 3) Offline training stage: For the training of the regression head target bounding box, use the L1 loss function and the GIoU loss function for supervision. Combine the ground truth label obtained from the validation frame, use the AdamW optimizer, and update the entire network parameters through the backpropagation algorithm. Continuously train the configured network until the number of iterations is reached to obtain the tracking framework MixFormer; Online tracking: Label the target search area of the first frame of the video to be tracked as a template frame, and the subsequent frames as test frames. Input the trained tracking framework MixFormer, and output the target bounding box on the test frame to achieve object tracking.

2. The end-to-end single-object tracking method based on a hybrid attention mechanism according to claim 1, characterized in that The cross-attention operation of the backbone network only performs one-way cross-attention from the template frame to the test frame, and does not perform cross-attention from the test frame to the template frame, to obtain the test frame features fused with template information.

3. An end-to-end single-object tracking method based on a hybrid attention mechanism according to claim 1 or 2, characterized in that The tracking head also includes a classification head, which is used to obtain the classification target confidence of the test frame. The classification head has a preset learnable confidence vector, which performs attention operations with the test frame features and the template frame's own features respectively, perceives the information of both to predict the classification target confidence of the current test frame. During the tracking process, select frames with confidence meeting the conditions from the already tracked video frame sequence to supplement as template frames.

4. An end-to-end single-object tracking method based on a hybrid attention mechanism according to claim 3, characterized in that During online tracking, first, crop the target search area in the first frame image of the video to be tracked as the template frame F train , and the frame to be tracked is used as the test frame F test . Through the tracking framework MixFormer, obtain the target bounding box on the test frame F test . During the tracking process, select a frame with the highest confidence and its tracked target bounding box from every N frames in the already tracked frame sequence as labels, and supplement them as the template frame F train .

5. An end-to-end single-object tracking method based on a hybrid attention mechanism according to claim 1, characterized in that In the data preparation stage, perform target area jitter processing on each frame image of each video in the training dataset, and then crop the target search area after jitter processing.

6. An end-to-end single-object tracking device based on a hybrid attention mechanism, characterized in that There is a computer storage medium, in which a computer program is configured. The computer program is used to implement the tracking framework MixFormer according to any one of claims 1-5, and when the computer program is executed, it implements the tracking method according to any one of claims 1-5.

Citation Information

Cited By

  • Single-target tracking method based on multi-modal language self-updating

    CN121883540A