A single-object tracking method based on a fusion feature decoding structure

By introducing a cross attention enhancement module in target tracking, the problem of discrimination difficulties in traditional cross feature enhancement operations under similar interference is solved, and the accuracy and robustness of target tracking are improved.

CN117218156BActive Publication Date: 2025-07-25HUAIBEI NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311140657.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-06
Publication Date
2025-07-25
Estimated Expiration
2043-09-06

AI Technical Summary

Technical Problem

Traditional cross-feature enhancement operations are difficult to distinguish targets from similar objects when dealing with interference from similar objects, resulting in inaccurate target tracking.

Method used

The intersection attention enhancement module (ECA) in the fusion feature decoding structure is introduced to enhance the attention and modeling capabilities of the network by enhancing the feature information of the target area, especially to improve the robustness of target tracking in complex scenarios.

Benefits of technology

It significantly improves the target tracking performance, especially when dealing with similar objects and complex scenarios, and has strong robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218156B_ABST
    Figure CN117218156B_ABST
Patent Text Reader

Abstract

A single-object tracking method based on a fusion feature decoding structure, which uses a cross-attention module that fuses decoded features to enhance the feature information of the target region, belongs to the field of computer vision technology, and solves the technical problem that traditional cross-feature enhancement operations are difficult to distinguish the target from similar objects when dealing with the interference of similar objects. The solution includes the following steps: First, a training set is obtained from the video dataset after data augmentation, and then a pair of template images and search region images are cropped from the training set, and features of this pair of images are extracted; Then, the extracted features are fed into a feature fusion network, where the self-attention module is used to perform self-attention enhancement on the template features and search region features, and the improved ECA module is used to fuse these two feature maps and perform feature interaction; Finally, the fused features are fed into the prediction head for classification tasks and regression tasks, and the target is accurately located by combining the classification result and the regression result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a single-object tracking method based on a fusion feature decoding structure. Background Art

[0002] Object tracking is a computer vision task whose goal is to track the position of a specific target object in real time in a given video sequence. This target object may move, deform, be occluded, or be subject to other external interferences in different frames of the video. The challenge lies in how to accurately track the position of the target under these complex conditions.

[0003] The process of object tracking includes first selecting an initial target region (usually called the template frame) in the video, and then determining the new position of the target in the subsequent frames by comparing the similarity or distance between the current frame and the template frame. The Siamese network is a commonly used feature fusion method in object tracking, which matches the target by calculating the cross-correlation relationship between two features. However, due to the complexity of the object tracking task, traditional Siamese networks may lead to suboptimal tracking results, especially when the target is occluded or partially visible. Therefore, to solve these problems, a Transformer-based structure is proposed to improve the robustness and accuracy of object tracking by capturing global context information and preserving the semantic information of the target.

[0004] Currently, in the field of computer vision, DETR and ViT have successfully introduced the Transformer model, bringing breakthrough progress to object detection and image classification tasks. In terms of object tracking, algorithms such as TransT, Stark, and SwinTrack all utilize the attention idea in Transformer to replace the correlation operation in the traditional Siamese network for feature fusion between the template and the search region, thus significantly improving the tracking performance. However, the traditional cross-feature enhancement operation is difficult to distinguish the target from similar objects when dealing with similar object interference. Summary of the Invention

[0005] The main purpose of the present invention is to overcome the deficiencies in the prior art and solve the technical problem that it is difficult to distinguish the target from similar objects in the traditional cross-feature enhancement operation when dealing with similar object interference. The present invention provides a single-object tracking method based on a fusion feature decoding structure.

[0006] The design concept of this application is as follows: A novel Enhanced Cross-Attention (ECA) module is proposed. By introducing the ECA module, it is possible to enhance the attention of the tracked target, strengthen the target information in the fused features, and enable the network to have better attention and modeling capabilities when processing similar objects. Its highlight lies in embedding the ECA module into the traditional Vision Transformer structure, providing features with better attention information for the decoder, thereby effectively enhancing the target information in the feature map and further improving the network's modeling ability. The key to this method is to introduce an enhanced cross-attention module in object tracking, explicitly enhancing the interaction relationship between elements in the modeling sequence, improving the modeling ability of features and the attention to the target, enabling the network to better identify and track the target object, especially having stronger robustness in complex scenarios, bringing significant performance improvement to the object tracking task, and having broad application prospects.

[0007] The present invention uses TransT as the benchmark for tracking and is realized through the following technical solutions: A single-object tracking method based on a fused feature decoding structure, using a cross-attention module that fuses decoded features to enhance the feature information of the target area, including the following steps:

[0008] S1. Crop a pair of pictures from the enhanced training set, extract the features of the template and the search area through the Siamese network, and obtain the template feature f z and the search area feature f x ;

[0009] S2. Use a Transformer-based feature fusion method to respectively fuse the template feature f z and the search area feature f x obtained in step S1, enhance the features through the first self-attention module and the second self-attention module respectively, and then simultaneously receive the template feature map and the search area feature map through the first cross-attention enhancement module and the second cross-attention enhancement module, perform feature interaction after fusing the template feature map and the search area feature map, repeat the feature fusion layer N times and then input it into the third cross-attention enhancement module, and fuse the feature maps of the two branches through the third cross-attention enhancement module;

[0010] S3. Input the data output by the third cross-attention enhancement module into the prediction head. The prediction head consists of a classification branch and a regression branch, and each branch includes a three-layer perceptron with a hidden dimension of d and a ReLU activation function; for the feature map generated by the feature fusion network the prediction head predicts each vector to obtain n = H z W zForeground / background classification results and n = H z W z Normalized coordinates of the search areas to complete single-object tracking based on the fusion feature decoding structure.

[0011] Further, the step S1 includes the following steps:

[0012] S1-1. Sample images from any video sequence in the video dataset, collect training samples, and then use conventional data augmentation methods (such as translation or brightness jitter) to expand the training set to obtain an augmented training set;

[0013] S1-2. In the augmented training set, first, define the area enclosed by doubling the side lengths outward with the target as the center in the first frame of the video sequence as the template image The template image includes the target information and the appearance information of the local scene around it; then, define the area enclosed by quadrupling the side lengths outward with the target in the previous frame as the center in the augmented training set as the search area image The search area image can cover the potential moving range of the target; finally, crop the template image and the search area image into squares;

[0014] S1-3. Input the cropped template image and search area image into the ResNet-50 network pre-trained on the ImageNet dataset for feature extraction to obtain the template feature and the search area feature

[0015] Further, the step S2 includes the following steps:

[0016] S2-1. Pass the template feature f z and the search area feature f x through a 1×1 convolution to reduce the channel dimension, and obtain two corresponding low-dimensional feature maps and

[0017] S2-2. Since the input of the Transformer is a set of feature vectors, flatten the template feature f z and the search area feature f x in the spatial dimension d respectively to obtain the corresponding template branch and the search area branch That is, the template branch f z1 and the search area branch f x1 are sets of feature vectors with a length of d;

[0018] S2-3. Pass the template branch fz1 and the search area branch f x1 are respectively input into the first self-attention module and the second self-attention module, where the self-attention calculation is as follows:

[0019]

[0020] In the formula, Softmax is a normalization function, Q and K are two sets of feature vectors with dimension d k , and V is the weighted value;

[0021] The attention score between Q and K is obtained through scaled dot product operation, then the attention map is generated through Softmax operation, and the weighted value V is recalculated according to the attention map; in this way, the attention mechanism adaptively focuses on the useful positions in V according to the correlation between Q and K;

[0022] The extended formula for multi-head attention is as follows:

[0023] MultiHead(Q, K, V) = Concat(H1, …, H nh )W o ;

[0024]

[0025] In the formula, is the parameter matrix; H i is the attention matrix of different heads;

[0026] Self-attention adaptively integrates the information of different positions of the feature map by using multi-head attention in residual form, and generates the spatial position encoding P x by introducing the sine function, and then determines the position information of the feature sequence. The output X1 of the self-attention module is:

[0027] X1 = X + MultiHead(X + P x , X + P x , X);

[0028] In the formula, P x is the spatial position encoding, X1 is the output of the self-attention module,

[0029] S2-4. Take the output X1 of the self-attention module as the input of the cross-attention enhancement module. The cross-attention enhancement module uses the residual form of multi-head cross-attention to fuse the feature vectors of the two inputs; similarly to the self-attention module, the spatial position encoding P xis also used in the cross-attention enhancement module; on this basis, a feed-forward network module (FFN) is adopted to enhance the fitting ability of the model. The model consists of two linear transformations with a ReLU in the middle. The calculation formula of the feed-forward network module is as follows:

[0030] FFN(x) = max(0, xW1 + b1)W2 + b2;

[0031] In the formula, W is the weight matrix, b is the bias vector, x is the feature vector output by the self-attention module, and the subscripts 1 and 2 represent the first layer and the second layer respectively;

[0032] In summary, the calculation formula of the cross-attention enhancement module is:

[0033]

[0034]

[0035]

[0036] In the formula, X q is the input of the branch, P q is the spatial position encoding corresponding to X q ; X kv is the input of another branch, P kv is the position encoding of X kv ;

[0037] When the attention score is greater than 0.7, the attention score is multiplied by a preset coefficient to complete attention enhancement. The principle is as follows: First, calculate the average attention degree of each position in the search sequence to obtain the average attention weight; Second, sort the positions according to the attention degree and find the indexes Topk_idx of the top 512 positions with the highest attention degree; Then, perform weighted processing on the elements in the search sequence, increase the weights of the positions with high attention degree by 30%, and keep the weights of other positions unchanged; Finally, splice the weighted search sequence and the template sequence together to form the target sequence to improve the matching accuracy. is the output of the cross-attention enhancement module.

[0038] Further, the step S3 includes the following steps:

[0039] S3-1. Locate the position of the target according to the position with the largest value in the classification result, and map this position back to the search area to obtain the center position of the target; Select the offset values of the corresponding upper, lower, left, and right of the target relative to the center position in the regression result according to the position with the largest value in the classification result;

[0040] S3-2. If tracking fails or the classification score is too low, update the entire model using the gradient descent method;

[0041] S3-3. Draw the tracking coordinate box based on the center position and offset value of the target to obtain the tracking result.

[0042] The beneficial effects of the present invention are as follows: The present invention introduces a cross-attention mechanism that fuses decoded features in the target tracking technology, enhances the interaction relationship between elements in the modeling sequence, and enables the network to better identify and track target objects. The attention enhancement module proposed by the present invention improves the feature modeling ability and target attention, significantly enhances the target tracking performance, especially performs well in dealing with similar objects and complex scenarios, and has broad application prospects. At the same time, the technology shows strong robustness in complex scenarios, can effectively handle interference, and brings positive effects to the field of target tracking. Description of the Drawings

[0043] Figure 1 is the overall framework diagram of the present invention;

[0044] Figure 2 is the flowchart of the self-attention module based on Transformer of the present invention;

[0045] Figure 3 is the ECA module flowchart of the cross-attention enhancement module that fuses decoded features in the present invention;

[0046] Figure 4 is the comparison chart of the success rate of the tracking effect of the present invention and other existing trackers on the OTB100 general dataset;

[0047] Figure 5 is the comparison chart of the precision rate of the tracking effect of the present invention and other existing trackers on the OTB100 general dataset;

[0048] Figure 6 is the comparison chart of the tracking effect of the present invention and other trackers on the OTB100 dataset. Detailed Embodiments

[0049] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0050] As Figure 1A single-object tracking method based on a fusion feature decoding structure is shown. First, a training set is obtained after data augmentation from a video dataset. Subsequently, a pair of template images and search region images are cropped from the training set. Then, this pair of images is input into a ResNet-50 network for feature extraction. Secondly, the extracted features are fed into a feature fusion network, where the self-attention module is used to perform self-attention enhancement on the template features and search region features, and the cross-attention enhancement module is used to fuse these two feature maps and perform feature interaction. Finally, the fused features are fed into a prediction head for classification tasks and regression tasks; and the target is accurately located by combining the classification result and the regression result, which specifically includes the following steps:

[0051] S1. Crop a pair of images from the augmented training set, and extract the features of the template and the search region through a Siamese network, respectively obtaining the template feature f z and the search region feature f x ; The step S1 includes the following steps:

[0052] S1-1. Sample images from any video sequence in the video dataset, collect training samples, and then use conventional data augmentation methods to expand the training set to obtain an augmented training set;

[0053] S1-2. In the augmented training set, first, define the region enclosed by expanding each side length of the first frame of the video sequence outward by two times with the target as the center as the template image The template image includes the target information and the appearance information of the local scene around it; then, define the region enclosed by expanding each side length of the target in the previous frame outward by four times in the augmented training set as the search region image The search region image can cover the potential moving range of the target; finally, crop the template image and the search region image into squares;

[0054] S1-3. Input the cropped template image and search region image into a ResNet-50 network pre-trained on the ImageNet dataset for feature extraction, respectively obtaining the template feature and the search region feature

[0055] S2. Use a Transformer-based feature fusion method to fuse the template feature f z obtained in step S1 and the search region feature f xFeature fusion is performed separately, and the features are enhanced by the first self-attention module and the second self-attention module respectively. Then, through the first cross-attention enhancement module and the second cross-attention enhancement module, the template feature map and the search region feature map are received simultaneously. After fusing the template feature map and the search region feature map, feature interaction is carried out. After the feature fusion layer is repeated N times, it is input into the third cross-attention enhancement module to fuse the feature maps of the two branches; step S2 includes the following steps:

[0056] S2-1. The template feature f z and the search region feature f x are passed through a 1×1 convolution to reduce the channel dimension, and two corresponding low-dimensional feature maps and

[0057] S2-2. Since the input of the Transformer is a set of feature vectors, the template feature f z and the search region feature f x are flattened in the spatial dimension d, and the corresponding template branch and the search region branch are obtained respectively. That is, the template branch f z1 and the search region branch f x1 are sets of feature vectors of length d;

[0058] S2-3. The template branch f z1 and the search region branch f x1 are respectively input into the first self-attention module and the second self-attention module. As shown in Figure 2 , the self-attention module is used to enhance the attention of the template branch f z1 and the search region branch f x1 . The self-attention calculation is as follows:

[0059]

[0060] In the formula, Softmax is a normalization function, Q and K are two sets of feature vectors with dimension d k , and V is the weighted value;

[0061] The attention score between Q and K is obtained through the scaled dot product operation, and then the attention map is generated through the Softmax operation. The weighted value V is recalculated according to the attention map;

[0062] Extended to the multi-head attention calculation formula as follows:

[0063] MultiHead(Q,K,V)=Concat(H1,…,Hnh )W o ;

[0064]

[0065] In the formula, is a parameter matrix; H i is the attention matrix of different heads;

[0066] Self-attention adaptively integrates the information of different positions of the feature map by using multi-head attention in the residual form. By introducing a sine function, the spatial position encoding P x is generated, and then the position information of the feature sequence is determined. The X1 output by the self-attention module is:

[0067] X1 = X + MultiHead(X + P x , X + P x , X);

[0068] In the formula, P x is the spatial position encoding, X1 is the output of the self-attention module,

[0069] S2-4. As Figure 3 shown, the X1 output by the self-attention module is used as the input of the cross-attention enhancement module. The cross-attention enhancement module uses the residual form of multi-head cross-attention to fuse the feature vectors of the two inputs; similarly to the self-attention module, the spatial position encoding P x is also used in the cross-attention enhancement module; on this basis, a feed-forward network module (FFN) is used to enhance the fitting ability of the model. The calculation formula of the feed-forward network module is as follows:

[0070] FFN(x) = max(0, xW1 + b1)W2 + b2;

[0071] In the formula, W is the weight matrix, b is the bias vector, x is the feature vector output by the self-attention module, and the subscripts 1 and 2 represent the first layer and the second layer respectively;

[0072] In summary, the calculation formula of the cross-attention enhancement module is:

[0073]

[0074]

[0075]

[0076] In the formula, X q is the input of the branch, P p is for Xq The corresponding spatial position encoding, X kv is the input of another branch, P kv is the position encoding of X kv ;

[0077] When the attention score is greater than 0.7, the attention score is multiplied by a preset coefficient to complete attention enhancement;

[0078] S3. Input the data output by the third cross-attention enhancement module into the prediction head. The prediction head consists of a classification branch and a regression branch. Each branch includes a three-layer perceptron with a hidden dimension of d and a ReLU activation function; for the feature map generated by the feature fusion network the prediction head predicts each vector to obtain n = H z W z foreground / background classification results and n = H z W z normalized coordinates of the search regions, completing single-object tracking based on the fusion feature decoding structure; The step S3 includes the following steps:

[0079] S3-1. Locate the position of the target according to the position with the largest value in the classification result, and map this position back to the search region to obtain the center position of the target; Select the offset values of the target above, below, left, and right relative to the center position in the regression result according to the position with the largest value in the classification result;

[0080] S3-2. If there is tracking failure or the classification score is too low, update the entire model by the gradient descent method;

[0081] S3-3. Draw the tracking coordinate box according to the center position and offset value of the target to obtain the tracking result.

[0082] In this specific embodiment, preliminary test experiments are conducted on the fusion feature decoding structure scheme proposed by the present invention, as follows.

[0083] Experiments are conducted on a server with an NVIDIA TITAN RTX4090 using the Pytorch framework. The software platform is Pycharm, and the fusion feature decoding structure scheme is implemented in Python. In the experiment, first, the success rate and precision on the OTB100 dataset are tested, and then the comparison with the true tracking results in the OTB100 dataset is tested.

[0084] Such as Figure 4 , Figure 5As shown, where Ours is the tracker proposed in the present invention, and SiamCAR, SiamBAN, SimRPN++, TransT, OcECAn, DaSiamRPN, SiamRPN, TCTrack, SiamFC are the names of trackers proposed by other scholars in recent years. The numbers corresponding to each tracker name respectively represent the average precision and average success rate of the tracker.

[0085] As Figure 4 shown, the precision graph represents the Euclidean distance between the center point of the prediction box of the tracking algorithm and the center point of the Ground Truth box. Usually, the threshold is 20 pixels. That is, if their Euclidean distance is within 20 pixels, it is considered a successful tracking. As Figure 5 shown, the success rate graph represents the percentage of the overlap rate (Overlap Score, OS) between the tracking box drawn by the tracking algorithm and the manually annotated tracking box being greater than a given threshold. From Figure 4 and Figure 5 it can be seen that when compared with advanced trackers in recent years on the general dataset OTB100, the tracking method Ours of the present invention has significantly improved in both precision and success rate.

[0086] Figure 6 This is a comparison of the actual tracking results of the present invention and advanced trackers MDNet, SiamRPN++, TransT, OcECAn, DaSiamRPN, SiamRPN, Staple in the OTB100 dataset. From Figure 5 it can be seen that when there are similar targets in the search area and during long-term tracking, the bounding boxes generated by the tracking method of the present invention are more accurate than those of other trackers.

[0087] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.

Claims

1. A single-object tracking method based on a fusion feature decoding structure, characterized in that It includes the following steps: S1. Crop a pair of images from the enhanced training set, extract the template and search region features through the Siamese network, and obtain the template feature f z and the search region feature f x ; S2. Use the Transformer-based feature fusion method to fuse the template feature f obtained in step S1 z and the search region feature f x respectively for feature fusion, and enhance the features through the first self-attention module and the second self-attention module respectively. Then, through the first cross-attention enhancement module and the second cross-attention enhancement module, simultaneously receive the template feature map and the search region feature map, fuse the template feature map and the search region feature map for feature interaction. After repeating the feature fusion layer N times, input it into the third cross-attention enhancement module to fuse the feature maps of the two branches through the third cross-attention enhancement module; The calculation formula of the cross-attention enhancement module is: where X q is the input of the branch, P q is the spatial position encoding corresponding to X q , X kv is the input of another branch, P kv is the position encoding of X kv , When the attention score is greater than 0.7, the attention score is multiplied by a preset coefficient to complete attention enhancement; S3. Input the data output by the third cross-attention enhancement module into the prediction head. The prediction head consists of a classification branch and a regression branch. Each branch includes a three-layer perceptron with a hidden dimension of d and a ReLU activation function; For the feature map generated by the feature fusion network The prediction head makes predictions for each vector to obtain n = H z W z foreground / background classification results and n = H z W z normalized coordinates of the search regions, completing single-object tracking based on the fusion feature decoding structure.

2. The single-object tracking method based on a fusion feature decoding structure according to claim 1, characterized in that: The step S1 includes the following steps: S1-1. Sample images from any video sequence in the video dataset, collect training samples, and then use conventional data augmentation methods to expand the training set to obtain an enhanced training set; S1-2. In the enhanced training set, first, define the region enclosed by expanding the sides of the region centered on the target in the first frame of the video sequence by two times as the template image. The template image includes the target information and the appearance information of the local scene around it; then, define the region enclosed by expanding the sides of the region centered on the target in the previous frame in the enhanced training set by four times as the search region image. The search region image can cover the potential movement range of the target; finally, crop the template image and the search region image into squares respectively. S1-3. Input the cropped template image and the search area image into the ResNet-50 network pre-trained on the ImageNet dataset for feature extraction, and obtain the template feature and the search area feature respectively. and the search area feature 3. A single-object tracking method based on a fusion feature decoding structure according to claim 1, characterized in that: The step S2 includes the following steps: S2-1. Convolve the template feature f z and the search area feature f x through a 1×1 convolution to reduce the channel dimension, respectively obtaining two corresponding low-dimensional feature maps and S2-2. Since the input of the Transformer is a set of feature vectors, the template feature f z and the search region feature f x are flattened in the spatial dimension d to obtain the corresponding template branch and the search region branch That is, the template branch f z1 and the search region branch f x1 are sets of feature vectors of length d; S2-3. Input the template branch f z1 and the search area branch f x1 into the first self-attention module and the second self-attention module respectively, where the self-attention calculation is as follows: where Softmax is the normalization function, Q and K are two sets of feature vectors with dimension d k , and V is the weighted value; Obtain the attention score between Q and K through scaled dot-product operation, then generate an attention map through Softmax operation, and then recalculate the weighted value V according to the attention map; The extended formula for multi-head attention is as follows: MultiHead(Q,K,V)=Concat(H1,…,H nh )W O ; In the formula, is the parameter matrix; H i is the attention matrix for different heads; Self-attention adaptively integrates information at different positions of the feature map using multi-head attention in residual form, and generates spatial position encoding P by introducing a sine function x , and then determines the position information of the feature sequence. The output X1 of the self-attention module is as follows: X1 = X + MultiHead(X + P x , X + P x , X); where P x is the spatial position encoding, X1 is the output of the self-attention module, S2-4. Use X1 output by the self-attention module as the input of the cross-attention enhancement module. The cross-attention enhancement module fuses the feature vectors of the two inputs in the residual form of multi-head cross-attention. Similarly to the self-attention module, the spatial position encoding P x is also used in the cross-attention enhancement module. On this basis, a feed-forward network module is used to enhance the fitting ability of the model. The calculation formula of the feed-forward network module is as follows: FFN(x) = max(0, xW1 + b1)W2 + b2; In the formula, W is the weight matrix, b is the bias vector, x is the feature vector output by the self-attention module, and the subscripts 1 and 2 represent the first layer and the second layer respectively.

4. The single-object tracking method based on a fusion feature decoding structure according to claim 1, characterized in that: The step S3 includes the following steps: S3-1. Locate the position of the target according to the position with the largest value in the classification result, and map this position back to the search area to obtain the center position of the target; select the offset values of the corresponding target above, below, left, and right relative to the center position in the regression result according to the position with the largest value in the classification result; S3-2. If there is tracking failure or the classification score is too low, update the entire model by the gradient descent method; S3-3. Draw the tracking coordinate box according to the center position and offset value of the target to obtain the tracking result.

Citation Information

Patent Citations

  • Transform-based single target tracking method

    CN114266996A

  • Three-dimensional point cloud single target tracking method based on regional self-attention mechanism

    CN115909010A