A target tracking algorithm based on complementary feature fusion and key frame template updating
By adopting the strategies of feature complementary fusion and keyframe template update in the target tracking algorithm, the problem of local information loss in the Transformer model under occlusion or appearance changes is solved, and higher tracking accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411942481.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The target tracking method based on the Transformer model will cause local information to be lost when facing occlusion or appearance changes, which will affect the accuracy of the tracking results.
The goal tracking algorithm based on feature complementary fusion and keyframe template update is adopted to enhance the robustness and accuracy of target tracking by building a complementary feature extraction module and using a keyframe-based template update strategy.
This algorithm improves the accuracy of target tracking by taking into account the local characteristics and global information of the template frame and the search frame at the same time, reduces the risk of tracking failure, and improves speed and accuracy in long-term tracking.
Smart Images

Figure CN119380045B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a target tracking algorithm based on feature complementary fusion and key frame template updating. Background Art
[0002] Visual object tracking is a fundamental task in computer vision and pattern recognition. It aims to predict the unknown position of an object in a video that has a well-defined position in the first frame. Currently, this technology has been widely used in video processing applications such as autonomous driving and visual surveillance. However, achieving high-precision and robust tracking under real-world challenges such as rotation, deformation, scale change, and occlusion remains a difficult problem.
[0003] With the widespread application of Transformer models in computer vision, trackers based on Transformer models have been proposed. When a tracker based on the Transformer model performs visual target tracking, the global feature representation of the image is enhanced through the Transformer encoder, and the template frame and the search frame are compared through the Transformer decoder to obtain the target tracking result. The implementation process includes: using the Transformer encoder to encode the image in the template frame, capturing the global context information in the image through the self-attention mechanism, and generating the feature representation of the template frame; inputting the search frame into the Transformer decoder, the decoder compares the feature representation of the template frame and the current input of the search frame through the cross-attention mechanism, gradually focusing on the parts of the search frame that are similar to the template frame, and generating the corresponding tracking results.
[0004] The defect of the above-mentioned prior art is that when the target tracking method based on the Transformer model faces the situation that the tracking target is occluded or the appearance changes, the tracking target is occluded or the appearance changes, which will cause the problem of local information loss, making the Transformer model unable to obtain complete target features, thereby causing target tracking failure or inaccurate target tracking results. Summary of the invention
[0005] Based on this, it is necessary to provide a target tracking algorithm based on feature complementary fusion and key frame template updating to address the above technical problems.
[0006] The embodiment of the present invention provides a target tracking algorithm based on feature complementary fusion and key frame template update, including:
[0007] A complementary feature extraction module is constructed. The complementary feature extraction module consists of two branches, each of which includes a convolutional network in the channel direction and a residual attention mechanism. A convolutional neural network is added before the input of each branch of the complementary feature extraction module. The output of the complementary feature extraction module is connected to the input of the residual cross attention mechanism. The output of the residual cross attention mechanism is connected to a bounding box prediction module including a classification branch and a regression branch to construct a target tracking model. The target tracking model uses a keyframe-based template update strategy.
[0008] Obtain a video with a tracking target, divide the video into frames, use the first frame of the video with the tracking target as a template frame, and use subsequent frames of the first frame of the video as search frames;
[0009] The template frame and the search frame are input into the target tracking model, and the template features of the template frame are extracted through the convolutional neural network of the first branch. The convolution network in the channel direction of the first branch is used to independently perform convolution on each channel of the template features, and the dynamic weight of the convolution is smoothed using the maximum value to obtain the local features corresponding to the template features; the residual attention mechanism in the first branch is used to determine the proportion of the local features corresponding to the template features and the results of the previous layer attention mechanism in the subsequent attention operation, and the global information corresponding to the template features is obtained;
[0010] The convolutional neural network of the second branch is used to extract the search features of the search frame. The convolutional network in the channel direction of the second branch is used to independently perform convolution on each channel of the search features, and the dynamic weight of the convolution is smoothed using the maximum value to obtain the local features corresponding to the search features. The residual attention mechanism in the second branch is used to determine the proportion of the local features corresponding to the search features and the results of the previous layer attention mechanism in the subsequent attention operation, and the global information corresponding to the search features is obtained.
[0011] The similarity between the global information corresponding to the template feature and the global information corresponding to the search feature is obtained through the cross-attention mechanism, and the template feature, search feature and similarity are fused to obtain the fused feature; the regression branch in the bounding box prediction module is used to confirm whether the position of each pixel in the fused feature is the foreground or background, and the classification branch in the bounding box prediction module is used to predict the bounding box position of the tracked target in the fused feature.
[0012] Optionally, the target tracking model uses a keyframe-based template update strategy, including:
[0013] The selected frames with a confidence threshold greater than the high confidence threshold are taken as key frames, and the confidence of target tracking is the maximum probability of the foreground of the classification result, and the formula is:
[0014] ;
[0015] in, Indicates the probability that the position is foreground or background, is the confidence of target tracking;
[0016] When a selected frame with a value greater than the high confidence threshold appears again, it is used as a new keyframe for target tracking; if a keyframe does not appear for a long time, the classification branch updates the keyframe at a set frame interval.
[0017] Optionally, the processing process of the convolutional network in the channel direction specifically includes:
[0018] Dynamic channel convolution is used to extract local features from each channel of the feature. The formula is:
[0019] ;
[0020] in, represents the convolution in the channel direction, Indicates dynamic weight Linear mapping of ;
[0021] The result of dynamic channel convolution is used as the local feature.
[0022] Optionally, the processing of the residual attention mechanism specifically includes:
[0023] Determine the weight of local features and the previous layer attention mechanism results in subsequent attention operations. The calculation formula is:
[0024] ;
[0025] in, represents the local feature matrix after dynamic convolution processing, represents the transposed matrix of the local feature matrix, Representing Template Features The dimension of represents the autocorrelation result of the previous layer, represents the local feature matrix of the previous layer, Represents the transposed matrix of the local feature matrix of the previous layer;
[0026] The weight obtained by the residual attention mechanism is used as global information.
[0027] Optionally, the similarity between the global information corresponding to the template feature and the global information corresponding to the search feature is obtained through a cross attention mechanism, specifically including:
[0028] Determine the similarity between the global information corresponding to the search feature and the global information corresponding to the template feature. The specific formula is:
[0029] ;
[0030] in, Represents the template feature, Representing Template Features The dimension of Represents the search feature.
[0031] Optionally, the regression branch in the bounding box prediction module is used to confirm whether the location of each pixel in the fused feature is the foreground or the background, and the classification branch in the bounding box prediction module is used to predict the bounding box position of the tracked target in the fused feature. The formula is expressed as:
[0032] ;
[0033] in, represents the forward propagation network for classification, represents the forward propagation network for bounding box regression, indicating that the length and width are The classification probability of each pixel in the search image represents the length and width of The regression probability for each pixel in the search image, Indicates the probability that the position is foreground or background, Indicates the distance from the pixel position to the four sides of the bounding box.
[0034] Compared with the prior art, the target tracking algorithm based on feature complementary fusion and key frame template updating provided by the embodiment of the present invention has the following beneficial effects:
[0035] The present invention constructs a complementary feature extraction module through a convolutional network in the channel direction and a residual attention mechanism. The module can simultaneously consider the local features and global information of the template frame and the search frame, and analyze the frame content more accurately. The residual attention mechanism not only considers the local features in the process of extracting global information, but also considers the influence of the results of the previous layer attention mechanism on the subsequent attention operations. It can solve the problem of local information loss caused by occlusion or appearance change of the tracking target in the prior art, making the target tracking in the interference scene more robust, thereby improving the accuracy of target tracking and reducing the risk of target tracking failure.
[0036] In addition, the target tracking model uses a keyframe-based template update strategy to improve the speed and accuracy of long-term tracking and reduce the drift of tracking results during long-term tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A model structure diagram of a target tracking algorithm based on feature complementary fusion and key frame template update provided in one embodiment;
[0038] Figure 2A schematic diagram of a residual attention mechanism component of a target tracking algorithm based on feature complementary fusion and key frame template update provided in one embodiment;
[0039] Figure 3 The success rate curve diagram of a target tracking algorithm based on feature complementary fusion and key frame template update provided in one embodiment. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0041] In one embodiment, a target tracking algorithm based on feature complementary fusion and key frame template update is provided, the method comprising:
[0042] A complementary feature extraction module is constructed. The complementary feature extraction module consists of two branches, each of which includes a convolutional network and a residual attention mechanism in the channel direction. A convolutional neural network is added before the input of each branch of the complementary feature extraction module, and the output of the complementary feature extraction module is connected to the input of the residual cross-attention mechanism. The output of the residual cross-attention mechanism is connected to a bounding box prediction module including a classification branch and a regression branch to construct a target tracking model. The target tracking model uses a keyframe-based template update strategy.
[0043] A video with a tracking target is obtained, the video is divided into frames, the first frame of the video with the tracking target is used as a template frame, and subsequent frames of the first frame of the video are used as search frames.
[0044] The template frame and the search frame are input into the target tracking model, and the template features of the template frame are extracted through the convolutional neural network of the first branch. The convolution network in the channel direction of the first branch is used to independently perform convolution on each channel of the template features, and the dynamic weights of the convolution are smoothed using the maximum value to obtain the local features corresponding to the template features. The residual attention mechanism in the first branch is used to determine the proportion of the local features corresponding to the template features and the results of the previous layer attention mechanism in the subsequent attention operations to obtain the global information corresponding to the template features.
[0045] The convolutional neural network of the second branch is used to extract the search features of the search frame. The convolutional network in the channel direction of the second branch is used to independently perform convolution on each channel of the search features, and the dynamic weight of the convolution is smoothed using the maximum value to obtain the local features corresponding to the search features. The residual attention mechanism in the second branch is used to determine the proportion of the local features corresponding to the search features and the results of the previous layer attention mechanism in the subsequent attention operations to obtain the global information corresponding to the search features.
[0046] The similarity between the global information corresponding to the template feature and the global information corresponding to the search feature is obtained through the cross-attention mechanism, and the template feature, search feature and similarity are fused to obtain the fused feature; the regression branch in the bounding box prediction module is used to confirm whether the position of each pixel in the fused feature is the foreground or background, and the classification branch in the bounding box prediction module is used to predict the bounding box position of the tracked target in the fused feature.
[0047] A specific embodiment of the present invention is provided:
[0048] S1. Build a target tracking model.
[0049] like Figure 1 As shown, the target tracking model includes a complementary feature extraction module, a residual attention mechanism, and a bounding box prediction model.
[0050] S2. The target tracking model uses a keyframe-based template update strategy. It determines the keyframe as the update template by setting a high confidence threshold, and compares the results below the threshold with the results of the first frame as the template frame. This ensures the lower limit of the result when the template is unstable after long-term tracking, reduces the drift of the results during long-term tracking, and is faster than the template update strategy learned by neural network.
[0051] Specifically include:
[0052] A key frame is one that is greater than a high confidence threshold. For the selected frame, the confidence of target tracking is the maximum probability that the classification result is the foreground, and the formula is as follows:
[0053] ;
[0054] in, Indicates the probability that the position is foreground or background, is the confidence of target tracking.
[0055] When a selected frame with a value greater than the high confidence threshold appears again, it is used as a new key frame for target tracking; if a key frame does not appear for a long time, the classification branch updates the key frame at a set frame interval to ensure the accuracy of the key frame.
[0056] During long-term tracking, due to challenges such as target deformation, the subsequent tracking results will not be accurate enough. The most representative result is a decrease in confidence. When the result of the key frame is less than the low confidence threshold When the template is unstable, the first frame is used as the template to generate the result again, and the results of the first frame and the key frame are compared, and the better one is selected as the final result to ensure the lower limit of the result when the template is unstable.
[0057] S3. Get template frame and search frame
[0058] Get a video with a tracking target, split the video into frames, use the first frame of the video with the tracking target as the template frame, and the subsequent frames of the first frame of the video as the search frame. Send the template frame and the search frame to the target tracking model, and the template features extracted by the convolutional neural network and the search features of the search frame are sent to the complementary feature mixing module. Then, the two branches in the complementary feature extraction module process the template features and search features through the convolutional network in the channel direction and the residual attention mechanism.
[0059] Specifically include:
[0060] The input template frame and search frame are fed into the convolutional neural network to extract features and then fed into the complementary feature mixing module. This module consists of a channel-wise convolutional network and a residual attention mechanism. The channel-wise convolutional network performs convolutions independently on each channel of the image and has the advantages of consistent sparse connectivity with the self-attention mechanism, which enables the extraction of local features.
[0061] Dynamic channel convolution performs local feature extraction on each channel of the feature, and its formula is:
[0062] ;
[0063] in, represents the convolution in the channel direction, Indicates dynamic weight The linear mapping of can make the convolution in the channel direction have dynamic weights.
[0064] At the same time, the residual attention mechanism is used to obtain global information. Figure 2 As shown in Figure 1, the residual attention mechanism can retain the results of the previous layer. This mechanism can determine the weight of local features and the results of the previous layer attention mechanism in subsequent attention operations, obtain global information, and make global information more robust. Incorporating it into the complementary feature extraction module, the formula is as follows:
[0065] ;
[0066] in, represents the local feature matrix after dynamic convolution processing, represents the transposed matrix of the local feature matrix, Representing Template Features The dimension of represents the autocorrelation result of the previous layer, represents the local feature matrix of the previous layer, Represents the transposed matrix of the local feature matrix of the previous layer.
[0067] S4. For the search feature in S3, the global information of the search feature is extracted by the residual attention mechanism just like the template feature. The similarity between the global information corresponding to the template feature and the global information corresponding to the search feature is obtained through the cross attention mechanism, and the template feature, search feature and similarity are fused to obtain the fused feature. This module can make the entire tracking process more stable and realize the interaction between the template frame and the search frame.
[0068] Specifically, in the case of a search frame with spatial position encoding, a residual attention mechanism is used to extract global search features, which is consistent with the template feature processing. After addition and normalization, enter the cross attention mechanism to obtain the search features With template features The similarity between them is as follows:
[0069] ;
[0070] in, Represents the template feature, Representing Template Features The dimension of Represents the search feature.
[0071] S5. Use the regression branch in the bounding box prediction module to confirm whether each pixel in the fusion feature is located in the foreground or background, and use the classification branch in the bounding box prediction module to predict the bounding box position of the tracked target in the fusion feature. Under this premise, target tracking will become very easy.
[0072] Specifically include:
[0073] After the similarity comparison between the search feature and the template feature, the similarity result will have the same dimension size as the search feature. Each position in can be mapped to the corresponding position in the search frame .
[0074] The regression branch is used to confirm whether the location of each pixel in the fusion feature belongs to the foreground of the target to be tracked or the background that does not belong to the tracked target. The classification branch uses a three-layer perceptron and a ReLU activation function to predict the bounding box of the tracked target in the fusion feature. The formula is expressed as:
[0075] ;
[0076] in, represents the forward propagation network for classification, represents the forward propagation network for bounding box regression, indicating that the length and width are The classification probability of each pixel in the search image represents the length and width of The regression probability for each pixel in the search image, Indicates the probability that the position is foreground or background, Indicates the distance from the pixel position to the four sides of the bounding box.
[0077] The embodiment of the present invention also provides a comparative experiment to prove it, using a method based on a twin network, a method based on a Transformer, and a target tracking algorithm provided by the present invention for illustration. Figure 3 As shown, the three curves are respectively the success rate curves of the method based on the twin network, the method based on the Transformer and the target tracking algorithm provided by the present invention on the GOT-10k test data set. It can be seen that:
[0078] The success rate of the twin network-based method is lower than that of the Transformer-based method and the target tracking algorithm provided by the present invention, and from the overall curve, the average success rate of the target tracking algorithm provided by the present invention is higher than that of the Transformer-based method. This shows that compared with the twin network-based method and the Transformer-based method, the target tracking algorithm provided by the present invention has obvious advantages in tracking effect. Therefore, the target tracking algorithm proposed by the present invention can improve the accuracy of target tracking and reduce the risk of target tracking failure.
[0079] The above-mentioned embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A target tracking algorithm based on feature complementary fusion and key frame template update, characterized in that: include: A complementary feature extraction module is constructed, wherein the complementary feature extraction module consists of two branches, each branch includes a convolutional network in a channel direction and a residual attention mechanism; a convolutional neural network is added before the input end of each branch of the complementary feature extraction module, and the output end of the complementary feature extraction module is connected to the input end of the residual cross attention mechanism; the output end of the residual cross attention mechanism is connected to a bounding box prediction module including a classification branch and a regression branch to construct a target tracking model, wherein the target tracking model uses a keyframe-based template update strategy; Obtain a video with a tracking target, divide the video into frames, use the first frame of the video with the tracking target as a template frame, and use subsequent frames of the first frame of the video as search frames; The template frame and the search frame are input into the target tracking model, and the template features of the template frame are extracted through the convolutional neural network of the first branch. The convolution network in the channel direction of the first branch is used to independently perform convolution on each channel of the template features, and the dynamic weight of the convolution is smoothed using the maximum value to obtain the local features corresponding to the template features; the residual attention mechanism in the first branch is used to determine the proportion of the local features corresponding to the template features and the results of the previous layer attention mechanism in the subsequent attention operation, and the global information corresponding to the template features is obtained; The convolutional neural network of the second branch is used to extract the search features of the search frame. The convolutional network in the channel direction of the second branch is used to independently perform convolution on each channel of the search features, and the dynamic weight of the convolution is smoothed using the maximum value to obtain the local features corresponding to the search features. The residual attention mechanism in the second branch is used to determine the proportion of the local features corresponding to the search features and the results of the previous layer attention mechanism in the subsequent attention operation, and the global information corresponding to the search features is obtained. The similarity between the global information corresponding to the template feature and the global information corresponding to the search feature is obtained through the cross attention mechanism, and the template feature, the search feature and the similarity are fused to obtain the fused feature; the regression branch in the bounding box prediction module is used to confirm whether the position of each pixel in the fused feature is the foreground or the background, and the classification branch in the bounding box prediction module is used to predict the bounding box position of the tracked target in the fused feature; Among them, the regression branch in the bounding box prediction module is used to confirm whether the location of each pixel in the fusion feature is the foreground or the background, and the classification branch in the bounding box prediction module is used to predict the bounding box position of the tracked target in the fusion feature. The formula is expressed as: in, represents the forward propagation network for classification, represents the forward propagation network for bounding box regression, Indicates length and width (H x ,W x ) searches for the classification probability of each pixel in the image, Represents length and width (H x ,W x ) searches for the regression probability of each pixel in the image, (p f ,p b ) represents the probability of the position being the foreground or background, and (l, t, r, b) represents the distance from the pixel position to the four sides of the bounding box.
2. The target tracking algorithm based on feature complementary fusion and key frame template updating as claimed in claim 1, characterized in that: The target tracking model uses a keyframe-based template update strategy, which specifically includes: The selected frames with a confidence threshold greater than the high confidence threshold are taken as key frames, and the confidence of target tracking is the maximum probability that the classification result is the foreground, and the formula is: Among them, (p f ,p b ) indicates the probability that the position is foreground or background, conf f is the confidence of target tracking, Indicates length and width (H x ,W x ) Search the classification probability of each pixel in the image; When a selected frame with a value greater than the high confidence threshold appears again, it is used as a new keyframe for target tracking; if a keyframe does not appear for a long time, the classification branch updates the keyframe at a set frame interval.
3. The target tracking algorithm based on feature complementary fusion and key frame template updating as claimed in claim 1, characterized in that: The processing of the convolutional network in the channel direction specifically includes: Dynamic channel convolution is used to extract local features from each channel of the feature. The formula is: ChannelConv(X)=Conv(X,f(X)); Among them, Channel Conv represents the convolution in the channel direction, f(X)=WX represents the linear mapping with dynamic weight W; The result of dynamic channel convolution is used as the local feature.
4. The target tracking algorithm based on feature complementary fusion and key frame template updating as claimed in claim 1, characterized in that: The processing of the residual attention mechanism specifically includes: Determine the weight of local features and the previous layer attention mechanism results in subsequent attention operations. The calculation formula is: Among them, Q represents the local feature matrix after dynamic convolution processing, represents the transposed matrix of the local feature matrix, d k represents the dimension of the template feature K, represents the autocorrelation result of the previous layer, Q p represents the local feature matrix of the previous layer, Represents the transposed matrix of the local feature matrix of the previous layer; The weight obtained by the residual attention mechanism is used as global information.
5. The target tracking algorithm based on feature complementary fusion and key frame template updating as claimed in claim 1, characterized in that: The obtaining of the similarity between the global information corresponding to the template feature and the global information corresponding to the search feature through the cross attention mechanism specifically includes: Determine the similarity between the global information corresponding to the search feature and the global information corresponding to the template feature. The specific formula is: Among them, K = V represents the template feature, d k represents the dimension of the template feature K, Q n Represents the search feature.
Citation Information
Patent Citations
Visual target tracking method based on multi-modal large language model
CN118314169A