A lightweight object tracking data annotation method based on Transformer
Through the lightweight video automatic labeling method based on the Transformer model, combined with the lightweight tracking algorithm HCAT and the quality evaluation network, the problem of high labor and time costs in the video labeling process is solved, and efficient and accurate video target labeling is achieved, especially in complex scenarios to improve the labeling quality.
Patent Information
- Application Number
- CN202211543612.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing single-objective visual tracking algorithm relies on large-scale manual labeling during the video labeling process, which consumes a lot of human resources and time. The existing automatic labeling method is not of high marking quality in complex scenarios, making it difficult to meet the needs of efficiently generating high-quality video data sets.
The lightweight video automatic labeling method based on the Transformer model is adopted. Through manual sparse labeling combined with the lightweight tracking algorithm HCAT and the quality evaluation network, difficult frames are filtered and manually corrected after the initial labeling is generated, and the target bounding box is optimized using bidirectional timing information fusion.
Efficient and accurate video target labeling is achieved, which reduces labor costs and improves labeling quality, especially in complex scenarios for prediction success rate and bounding box regression accuracy.
Smart Images

Figure CN115908496B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of machine learning, single-object visual tracking, and video object annotation, and relates to a lightweight single-object tracking algorithm HCAT, a cross-correlation algorithm, and an attention mechanism. Background Art
[0002] As one of the basic researches in the field of computer vision, single-object visual tracking has made remarkable progress in recent years. It requires the tracking algorithm to determine the coordinate position of the tracking target in a series of video frames. Existing types of tracking algorithms have shown robust tracking performance on various tracking datasets, such as SiamRPN++, Ocean, BTCF, ATOM, etc. Most of these tracking algorithms are based on deep models and require large-scale video data with accurate target box annotations for training to ensure high-performance tracking of the model. However, manually annotating target boxes frame by frame consumes a large amount of human resources and time costs. There is still a large gap in the existing large-scale video datasets available for tracking model training, which has become one of the bottlenecks for further improving tracking performance. Therefore, efficiently generating high-quality large-scale video data annotations is still an urgent problem to be solved in this field.
[0003] To reduce the human and time costs of video annotation, several methods have been tried to achieve automatic target box annotation of videos. The basic idea of the existing solutions is usually to first sparsely annotate several key frames of the video sequence, and then automatically complete the target boxes of other frames by interpolation. Among them, the interpolation methods are mainly divided into three categories. The first category is linear interpolation based on geometric information, such as LabelMe (Jenny Yuen et al., Labelme video: Building a video database with human annotations). This type of method assumes a single target motion pattern and only obtains the target annotations of other frames based on the geometric cues of the target. The second category is complex interpolation based on visual information, such as VATIC (Carl Vondrick et al., Efficiently scaling up video annotation with crowdsourced marketplaces) which generates target annotations by extracting the visual features of the target entity and using a more complex dynamic interpolation method. The third category is interpolation based on existing tracking algorithms. For example, TrackingNet annotates 1 frame every 1 second, and on this basis, the tracker STAPLECA (Matthias Mueller et al., Context-aware correlation filter tracking) is used to obtain the final video target annotation.
[0004] None of the above methods consider the correction of annotations. When faced with complex tracking scenarios, such as complex target motion patterns, the presence of interference objects, complex backgrounds, or partial occlusion of the target, existing trackers and other interpolation methods may all lead to unreliable annotation results. Designing an effective annotation quality assessment module and manually correcting the automatically generated annotations will further improve the accuracy and reliability of automatic annotation. To address this issue, VASR (Kenan Dai et al., Video annotation for visual tracking via selection and refinement) proposed a brand-new automatic annotation process based on selection and refinement. The selection module is used to evaluate the quality of the forward and backward tracking results, select the final target annotation based on the scores of the tracking results, and screen out the frames with tracking errors for manual correction. The refinement module then introduces a geometric parameter prediction model to generate more accurate target bounding box annotations, which can effectively improve the annotation quality.
[0005] Temporal information is an important part that cannot be ignored in video annotation tasks. To obtain high-quality target bounding boxes, annotators or automatic annotation algorithms need to pay attention to the appearance and position changes of the target between adjacent frames. Due to its powerful global semantic information capture ability, the Transformer model has been widely used in sequence tasks. Currently, the Transformer model has been found to have the potential to process visual information and is gradually replacing convolutional neural networks and being applied to various fields of computer vision. Summary of the Invention
[0006] The present invention aims to provide a lightweight video automatic annotation method based on the Transformer model with strong generalization ability, which solves the problem that existing automatic annotation algorithms deeply rely on specific trackers and, to a certain extent, solves problems such as low annotation quality and slow annotation speed.
[0007] The method of the present invention can perform simple and efficient automatic video target annotation during the annotation process of large-scale datasets.
[0008] The technical solution of the present invention is as follows:
[0009] A lightweight video automatic annotation method based on the Transformer model with strong generalization ability, the steps are as follows:
[0010] Step 1: Perform manual sparse annotation on the video sequence to be annotated, that is, manually annotate the target bounding box every 30 frames to obtain partial manual initial annotations (target bounding box coordinates), accounting for 3.3% of the total number of frames;
[0011] Step 2: Use the lightweight tracking algorithm HCAT for forward and backward tracking. The tracking results include the target bounding box coordinates of the remaining 96.7% of the frames except for the 3.3% manually annotated frames. The bounding boxes of the 3.3% manually annotated frames and the bounding boxes of the 96.7% frames recognized by the tracker are used as the complete initial annotations of the video sequence to be annotated. Specifically:
[0012] The lightweight tracking algorithm HCAT is mainly composed of a feature extraction network, a feature fusion network, and a prediction network. The basic module of the feature extraction network refers to ResNet18. Remove the last stage of ResNet18, stack convolutional modules to deepen the network depth, use a convolutional layer with a stride of 2 for feature extraction and downsampling, and construct a feature map with a downsampling factor of 16. During tracking, first crop the search area from the frame to be tracked (crop based on the target position in the previous frame of the frame to be tracked, where the target position of the previous frame has been obtained when the previous frame is tracked), and then crop the template area from the corresponding image according to the manually annotated bounding box of the first frame in every 30 frames. Input the template area and the search area into the feature extraction network respectively to obtain the feature maps corresponding to the template area and the search area, and then use the feature fusion network to fuse the two feature maps to obtain a fused feature map carrying the target appearance information and position information. Based on the fused feature map, use the prediction network to predict the confidence score and the regression coordinates of the target box to obtain the target bounding box to be tracked in the current frame.
[0013] During the process of using the lightweight tracking algorithm HCAT for forward and backward tracking, use the 3.3% manually sparse annotated frames in Step 1 as template frames, and use HCAT to track the subsequent 29 frames to obtain the target bounding box coordinates of the remaining 96.7% of the frames. The obtained forward tracking results, backward tracking results, and the 3.3% manual annotations are used as the initial annotations together. The annotation algorithm in the subsequent steps performs difficult frame selection and normal frame re-optimization based on the initial annotations.
[0014] Step 3: Crop the images to be annotated according to the forward and backward initial annotations to obtain the forward and backward search areas.
[0015] Step 4: Take the cropped forward and backward images to be annotated and the corresponding initial annotations in groups of 20 frames in length, and input them into the quality score evaluation network for difficult frame screening. Specifically:
[0016] The above quality score evaluation network is composed of a target multi-dimensional feature extraction module, a Transformer temporal feature fusion module, and a prediction module. The evaluation process of the quality score is as follows:
[0017] (1) Input a set of (20 frames) of forward, backward search regions and template regions into the relatively lightweight backbone network ResNet18 respectively, and perform 8-fold downsampling to obtain forward, backward feature maps and template feature maps:
[0018]
[0019] Among them, represents the forward and backward feature maps corresponding to the j-th frame of the image to be labeled, and T represents that each group of inputs contains continuous T frames of search regions. Perform cross-correlation operations between the forward and backward feature maps and the template feature respectively to obtain forward and backward response maps M f / b . And reduce the computational amount to achieve lightweight design. The response maps will be input into the response map network. After being processed by the response map network composed of three convolutional layers, the forward and backward response maps obtain forward and backward visual features:
[0020]
[0021] Among them, d v represents the dimension of the visual feature vector, and R represents a real number. At the same time, the target bounding box coordinates corresponding to the forward and backward search regions are processed by the motion linear layer to obtain forward and backward motion features
[0022] (2) Connect the forward visual feature and the forward motion feature to obtain the forward target multi-dimensional feature Similarly, the backward target multi-dimensional feature can be obtained Input the forward and backward target multi-dimensional features into the Transformer temporal feature fusion module simultaneously for forward and backward feature fusion to obtain the bidirectional temporal fusion feature. The Transformer temporal feature fusion module is designed with reference to the TransT encoder structure and is composed of a self-attention module and a cross-attention module. The core operation of both, the attention mechanism operation, is defined as follows:
[0023]
[0024] Among them, Q, K, and V represent the input query, key value, and value item respectively, and d k represents the dimension of the feature vector.
[0025] (3) Based on the bidirectional temporal fusion feature, use the prediction module to predict the initial annotation quality score of this group of 20 frames of images. The frames with quality scores lower than the quality score threshold of 0.5 are used as difficult frames and sent for manual re-annotation.
[0026] Step 5: Send the frames with quality scores higher than the quality score threshold of 0.5 in Step 4 to the target position optimization network to further optimize the initial annotation. Specifically:
[0027] The network structure and input of the target location optimization network, except for the prediction module, are the same as those of the quality score evaluation network. The prediction module of the quality score evaluation network outputs a one-dimensional vector representing the quality score for each frame. In the target location optimization network, after obtaining the bidirectional temporal fusion feature, the prediction module predicts the target bounding box coordinates of 20 frames in this group (the bounding box coordinates of each frame are represented by the output four-dimensional vector). Since this feature fuses bidirectional temporal information as well as visual and motion features, the accuracy of the output target location will be significantly improved compared to the initial annotation.
[0028] Advantages of the present invention:
[0029] (1) The annotation method proposed by the present invention is decoupled from the specific tracking algorithm. The generation of the initial annotation can use any lightweight tracking algorithm and does not depend on the tracking results of the tracking algorithm, having strong generalization. At the same time, the use of the lightweight tracking algorithm and the lightweight design of the annotation network significantly reduce the annotation time, and the manual annotation ratio within 10% greatly reduces the labor cost.
[0030] (2) By adopting the method of combining bidirectional temporal information fusion, motion and appearance information, the prediction success rate of the network for difficult frames and the accuracy of the target bounding box coordinate regression are improved, and the quality of automatic annotation is significantly improved. Description of the drawings
[0031] Figure 1 It is a schematic diagram of the structure of the lightweight tracker HCAT.
[0032] Figure 2 It is the target location optimization network and the quality score evaluation network.
[0033] Figure 3 It is the flow chart of the proposed lightweight video annotation method. Detailed implementation manners
[0034] The following further illustrates the detailed implementation manners of the present invention in combination with the drawings and technical solutions.
[0035] Figure 1It is a schematic diagram of the structure of the lightweight tracker HCAT. The network consists of a feature extraction module, a feature fusion module, and a prediction module. The feature extraction module removes the last stage of the ResNet18 network while keeping the stride of other parts unchanged. Therefore, this module downsamples the image features by 16 times while extracting the image features. The feature fusion module includes an encoder and a decoder. The encoder is composed of N sequentially connected cross-attention Transformer modules, and the decoder is a single cross-attention module. Compared with the commonly used correlation operations, the application of the Transformer module further improves the fusion effect of the template features and the search region features. The fused features output by the feature fusion module are used to extract appearance and position information in the prediction module, and the coordinates and confidence scores of the target bounding box in the input image are obtained.
[0036] Figure 2 It is the target position optimization network and the quality score evaluation network. Figure 3 It is a flowchart of the lightweight video annotation method. The lightweight tracker HCAT generates forward and backward target bounding boxes, crops the search region in the original image according to the bounding box coordinates at twice the size, and sends it to the quality score evaluation network in groups of twenty frames. The forward and backward temporal search features first go through the extraction of visual features and motion features in the quality score evaluation network. After connecting the forward and backward visual and motion features, the forward and backward multi-dimensional features are obtained. The forward and backward multi-dimensional features are used for feature extraction and fusion in the Transformer temporal feature fusion module. The fused features contain the target motion information and appearance information based on the initial annotation. Through the prediction module, the initial annotation quality score of this group of pictures can be obtained. This quality score represents the comprehensive score of the forward and backward initial annotations of these 20 frames of pictures in this group, and therefore represents the difficulty level of annotating this frame. The frames with lower quality scores are identified as difficult frames and directly sent for manual annotation. The target bounding boxes of the remaining frames are further optimized by the target position optimization network.
[0037] The input and data processing of the target position optimization network are the same as those of the quality score evaluation network. However, in the prediction module, a secondary prediction will be made on the target bounding box coordinates. Due to the combination of the bidirectional temporal motion information and appearance information in the fused features, the prediction accuracy of the target position optimization network will be higher than that of the initial annotation, so as to achieve the purpose of further optimizing the initial bounding box coordinates.
[0038] HCAT or other lightweight trackers can be selected arbitrarily. The training of the quality score evaluation network and the target location optimization network needs to be carried out separately to ensure high performance gains. The training set uses all video sequences of the LaSOT and Got10k datasets. The optimizer is selected as AdamW, the initial learning rate is set to 0.0001, and the learning rate is reduced by 10 times every 20 training epochs, with a total of 70 epochs of training. In addition to the initial annotation box coordinates, the inputs of the quality score evaluation network and the target location optimization network also need to crop the search area and the template area at 2 times the size in the original image according to the initial bounding box and scale them to a unified size of 128*128. The score threshold of the quality score evaluation network is set to 0.5.
[0039] The feature extraction network structures in the quality score evaluation network and the target location optimization network are as follows:
[0040] Sequence Operation type Input size Output size 1 Input 3*128*128 Null 2 Conv2d 3*128*128 64*64*64 3 BasicBlock 64*64*64 64*32*32 4 BasicBlock 64*32*32 128*16*16 5 BasicBlock 128*16*16 256*16*16
Claims
1. A lightweight object tracking data annotation method based on Transformer, characterized in that, The steps are as follows: Step 1: Manually sparsely annotate the frames of the video sequence to be annotated. That is, in a manual way, perform the annotation of the target bounding box every 30 frames to obtain some initial manual annotations, namely the coordinates of the target bounding box, accounting for 3.3% of the total number of frames; Step 2: Use the lightweight tracking algorithm HCAT for forward and backward tracking. The tracking results include the coordinates of the target bounding box for the remaining 96.7% of the frames except for the 3.3% of the manually initially annotated frames; The bounding boxes of the 3.3% of the manually initially annotated frames and the bounding boxes of the 96.7% of the frames identified by the tracker are used as the complete initial annotation of the video sequence to be annotated. Specifically: The lightweight tracking algorithm HCAT is mainly composed of a feature extraction network, a feature fusion network, and a prediction network; The basic module of the feature extraction network refers to ResNet18. Remove the last stage of ResNet18, stack convolutional modules to deepen the network depth, use a convolutional layer with a stride of 2 for feature extraction and downsampling to construct a feature map with a downsampling factor of 16; During tracking, first crop the search area from the frame to be tracked: crop based on the target position in the previous frame of the frame to be tracked, where the target position in the previous frame has been obtained when the previous frame is tracked, and then crop the template area from the corresponding image according to the manually annotated bounding box of the first frame in every 30 frames; Input the template area and the search area into the feature extraction network respectively to obtain the feature maps corresponding to the template area and the search area, and then use the feature fusion network to fuse the two feature maps to obtain a fused feature map carrying the target appearance information and position information; Based on the fused feature map, use the prediction network to predict the confidence score and the regression coordinates of the target box to obtain the target bounding box to be tracked in the current frame; During the process of using the lightweight tracking algorithm HCAT for forward and backward tracking, use the 3.3% of the manually initially annotated frames in Step 1 as template frames, and use the lightweight tracking algorithm HCAT to track the subsequent 29 frames to obtain the coordinates of the target bounding box for the remaining 96.7% of the frames; The obtained forward and backward tracking results and the 3.3% of the manually initial annotations are used as the initial annotation together. The annotation algorithm in the subsequent steps performs difficult frame selection and normal frame re-optimization based on the initial annotation; Step 3: According to the forward and backward initial annotations, crop the pictures to be annotated to obtain the forward and backward search areas; Step 4: Take the forward and backward search areas obtained by cropping and the corresponding initial annotations in groups of 20 frames in length and input them into the quality score evaluation network for difficult frame screening. Specifically: The quality score evaluation network is mainly composed of a target multi-dimensional feature extraction module, a Transformer temporal feature fusion module, and a prediction module; The evaluation process of the quality score is as follows: (4.1) Input a group of forward and backward search areas and the template area into the backbone network ResNet18 respectively, and perform 8-fold downsampling to obtain the forward, backward feature maps and the template feature map: Among them, represents the forward and backward feature maps corresponding to the j-th frame of the image to be annotated, and T represents that each group of inputs contains consecutive T frames of search regions; the forward and backward feature maps are respectively subjected to cross-correlation operations with the template feature to obtain the forward and backward response maps M f / b ; the forward and backward response maps are processed by a response map network composed of three convolutional layers to obtain forward and backward visual features: where d v represents the dimension of the visual feature vector; R represents real numbers; meanwhile, the target bounding box coordinates corresponding to the forward and backward search regions are processed by the motion linear layer to obtain forward and backward motion features (4.2) Connect the forward visual features and forward motion features to obtain the forward target multi-dimensional features Similarly, the reverse target multi-dimensional features can be obtained Input the forward and reverse target multi-dimensional features into the Transformer temporal feature fusion module simultaneously for the fusion of forward and reverse features, obtaining the bidirectional temporal fusion features; the Transformer temporal feature fusion module is designed with reference to the TransT encoder structure and mainly consists of a self-attention module and a cross-attention module. The core operation of both, the attention mechanism operation, is defined as follows: Among them, Q, K, and V respectively represent the input query, key value, and value item, and d k represents the dimension of the feature vector; (4.3) Based on the bidirectional temporal fusion features, use the prediction module to predict the initial annotation quality scores of 20 frames in this group; frames with quality scores lower than the quality score threshold of 0.5 are sent as difficult frames for manual re-annotation; Step 5: Send the frames with quality scores higher than the quality score threshold of 0.5 in Step 4 to the target location optimization network to further optimize the initial annotation; specifically: The network structure and input of the target location optimization network are the same as those of the quality score evaluation network except for the prediction module; the prediction module of the quality score evaluation network outputs a one-dimensional vector representing the quality score for each frame, while in the target location optimization network, after obtaining the bidirectional temporal fusion features, the prediction module will predict the target bounding box coordinates of 20 frames in this group, that is, the bounding box coordinates of each frame are represented by the output four-dimensional vector. Since this feature fuses bidirectional temporal information as well as visual and motion features, the accuracy of the output target location will be significantly improved compared to the initial annotation.