Foreground-aware Transformer Object Tracking Method and System Based on Re-tracking Strategy

By introducing a foreground-aware Transformer encoder, a foreground-enhancing multi-head self-attention and background screening module into the Transformer target tracking method, combining an adaptive re-tracking strategy and an online update template strategy, the problem of reduced target tracking accuracy in complex scenarios in the existing technology is solved, and a more efficient and stable target tracking effect is achieved.

CN119832027BActive Publication Date: 2025-05-27NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510316889.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-27
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing Transformer target tracking method is susceptible to background clutter interference, target occlusion and target deformation in complex scenarios, resulting in a decrease in tracking accuracy.

Method used

The foreground-aware Transformer target tracking method based on retracking strategy is adopted, and the template and search area features are extracted and fused through the foreground-aware Transformer encoder, combining the foreground enhancement of multi-head self-attention and background screening modules are enhanced to enhance perception of the foreground targets and reduce background interference. At the same time, an adaptive re-tracking strategy and online update template strategy are adopted to improve the adaptability and stability of tracking.

Benefits of technology

It significantly improves the accuracy and adaptability of target tracking, reduces computational complexity, and avoids tracking failures caused by template aging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832027B_ABST
    Figure CN119832027B_ABST
Patent Text Reader

Abstract

The present invention discloses a foreground-aware Transformer object tracking method and system based on a re-tracking strategy, which relates to the technical fields of image recognition and object tracking, and includes: inputting the most recent frame template image and the search area image into a pre-constructed preprocessing layer to respectively generate a template token sequence and a search token sequence, and caching them in a re-tracking buffer; then splicing the two sequences to obtain a mixed sequence, and inputting the mixed sequence into a pre-established foreground-aware encoder to obtain an output sequence, filling the search area feature sequence in the output sequence and inputting it into a prediction head to obtain a prediction result; if the prediction results of consecutive t frames are all lower than a preset expected value, the re-tracking strategy is enabled, the re-tracking prediction result is obtained through a general encoder, and a comprehensive evaluation is performed in the re-tracking evaluation area to obtain the best tracking result. This method can not only effectively reduce background interference, but also timely correct tracking deviations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition and target tracking, specifically a foreground-aware Transformer target tracking method and system based on a re-tracking strategy. Background Art

[0002] In the field of computer vision, visual target tracking has become a highly regarded research hotspot, showing broad and promising application prospects in many key fields such as military, autonomous driving, and medical. Specifically, visual target tracking aims to accurately define the position and shape contour of a target of interest marked with a bounding box in the first frame and then automatically and accurately determine the position and shape of the target in subsequent consecutive frames of the video. However, in real-world complex scenarios, visual target tracking faces many difficulties, such as background clutter interference, frequent target occlusion, and target deformation. These adverse factors will significantly reduce the tracking accuracy.

[0003] At present, mainstream visual target tracking methods focus on innovative exploration based on the Transformer architecture, giving full play to the excellent advantages of the Transformer in processing local and global information, greatly enhancing the modeling potential of the algorithm, and thus significantly improving the overall performance of the tracking algorithm. According to the difference in working mode, current Transformer-based target tracking strategies are mainly divided into two categories: two-stream two-stage tracking method and single-stream single-stage tracking method. Among them, the two-stream two-stage tracking method uses two completely identical network branches, which are respectively dedicated to extracting template features and search region features. In terms of process design, feature extraction and feature fusion are carefully divided into two successive independent stages. In contrast, the single-stream single-stage tracking method relies on only a single network branch, synchronously realizing feature extraction and fusion through a single-stage process, greatly simplifying the overall architecture of the algorithm and effectively improving the operation efficiency.

[0004] However, most existing Transformer tracking methods have certain limitations. The self-attention mechanism lacks attention to the most relevant foreground information in the search region, is prone to giving too much attention to background information and being interfered. Moreover, the self-attention calculation amount of the whole image is too large, and the inference speed is slow. Secondly, in real dynamic scenarios, the tracker not only needs to cope with complex and changeable background interference and frequent occlusion problems, but also as the tracking time span continues to extend, such challenges show an exponential amplification trend. Summary of the Invention

[0005] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a foreground-aware Transformer target tracking method and system based on a re-tracking strategy.

[0006] In a first aspect, the object of the present invention can be achieved by the following technical solutions: A foreground-aware Transformer object tracking method based on a re-tracking strategy, the method comprising the following steps:

[0007] Receive the most recent frame template image and the search area image, input the most recent frame template image and the search area image into a pre-constructed preprocessing layer to generate a template token sequence and a search token sequence respectively, and cache the template token sequence and the search token sequence into a re-tracking buffer;

[0008] Concatenate the template token sequence and the search token sequence to obtain a mixed sequence, input the mixed sequence into a pre-established foreground-aware Transformer encoder to obtain an output sequence, fill the search area feature sequence in the output sequence, and input the filled output sequence into a prediction head to obtain a prediction result;

[0009] Compare the prediction result with a preset expected value. If the prediction results of consecutive t frames are all lower than the preset expected value, store the prediction result in the re-tracking evaluation area, and input the template token sequence and the search token sequence in the re-tracking buffer into a general Transformer encoder for tracking prediction to obtain a re-tracking prediction result. In the re-tracking evaluation area, comprehensively evaluate the prediction result and the re-tracking prediction result to obtain the best tracking result, and update the template image online.

[0010] Combined with the first aspect, in certain implementation manners of the first aspect, the method further includes: the most recent frame template image , where the height of the template image is H z , the width is W z , and the number of channels is 3. represents the set of real numbers. This image is divided into multiple image patches with side length P and flattened into a template block sequence , and the number of template blocks is . The search area image , where the height of the template image is H x , the width is W x , and the number of channels is 3. This image is divided into multiple image patches with side length P and flattened into a search block sequence , and the number of search blocks is .

[0011] Combined with the first aspect, in certain implementation manners of the first aspect, the method further includes: the generation process of the template token sequence and the search token sequence:

[0012] Using a linear projection layer with input parameter L, map the module block sequence and the search block sequence to a D-dimensional space, and add learnable position embeddings to each block embedding respectively to obtain the template token embedding Z and the search token embedding X:

[0013]

[0014] where z p i is the i-th block in the module block sequence, x p i is the i-th block in the search block sequence, P z is the position embedding of the template token, P x is the position embedding of the search token, and L represents the linear projection layer with parameter L.

[0015] Combined with the first aspect, in some implementations of the first aspect, the method further includes: obtaining the hybrid sequence M by concatenating Z and X along the channel dimension for the hybrid sequence 0 ;

[0016] Input the hybrid sequence into a pre-established foreground-aware Transformer encoder, and obtain the query q z , key k z , value v z of the template and the query q x , key k x , value v x of the search region through the normalization layer.

[0017] Combined with the first aspect, in some implementations of the first aspect, the method further includes: after inputting the hybrid sequence into the pre-established foreground-aware Transformer encoder, using a foreground-enhanced multi-head self-attention module to extract and fuse the features of the hybrid sequence, and using a background screening module to screen and discard background tokens.

[0018] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the calculation process of using the foreground-enhanced multi-head self-attention module to extract and fuse the features of the hybrid sequence is as follows:

[0019]

[0020] where FEA represents the calculation of a single foreground-enhanced self-attention head, q, k, and v represent the input query vector, key vector, and value vector respectively, Softmax is the activation function, Maxj represents the top K maximum values in each row of the attention weight matrix, and k T represents the transpose of the hybrid sequence key k. is the dimension of the mixed token sequence represents a zero-value matrix, k z T and k x T are the transposes of the template sequence key and the search sequence key respectively. The formula for foreground-enhanced multi-head self-attention is as follows:

[0021]

[0022] In the formula, FEMSA represents the result of foreground-enhanced multi-head attention, and Concat represents the concatenation operation, and head i represents the calculation result of the i-th foreground-enhanced self-attention head. The output of the foreground-enhanced multi-head attention module is added to the mixed token sequence M 1 with a residual, and then input into the background screening module.

[0023] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the process of using the background screening module to screen and discard background tokens:

[0024] Select the similarity results between the template center position and all positions in the search area as the basis for screening the background. The expression is:

[0025]

[0026] In the formula, w i represents the similarity calculated by the i-th head. Since there are h attention heads, h similarity results can be calculated for the center position of the template. The background screening threshold is the average value of the h similarity results. The expression is

[0027]

[0028] According to the background screening threshold, the search tokens with similarity lower than the threshold are deleted, and the remaining search tokens need to record their original positions and then be re-concatenated with the template tokens to obtain the mixed token sequence M 2 ;

[0029] M 2 After normalization and passing through the feed-forward network, it is added to its own residual to obtain the output M of the Transformer block 1 ;

[0030] After passing through N Transformer blocks, the final output M of the encoder is obtained N .

[0031] In combination with the first aspect, in some implementations of the first aspect, the method further includes: taking the search token sequence of M N and extracting it, restoring it to the initial order according to the initial positions of the tokens, and using a 0 filling strategy for the missing positions to obtain a complete search token sequence X';

[0032] Reshape X' into a two-dimensional feature map and input it into a three-branch prediction head to obtain classification scores, position offsets, and object scales respectively, and generate predicted bounding boxes;

[0033] Compare the prediction result with a preset expected value τ. If it is lower than τ for consecutive t frames, it indicates that there is a deviation in the current prediction result. At this time, start the re-tracking strategy and temporarily store the classification result and the bounding box result in the re-tracking evaluation area;

[0034] Take out the template token sequence and the search token sequence in the re-tracking buffer and input them into a general Transformer encoder. Different from the foreground-aware Transformer encoder, the general Transformer encoder replaces the foreground-enhanced multi-head self-attention module with a multi-head self-attention module and reduces the number of Transformer blocks. Among them, the calculation process of the multi-head self-attention module is as follows:

[0035]

[0036] In the formula, Attention(q,k,v) represents the calculation result of a single self-attention head for the mixed sequence input, and Softmax is the activation function, is the dimension of the mixed token sequence, and head i represents the calculation result of the i-th self-attention head, and MHSA represents the multi-head self-attention result.

[0037] Input the output result of the general Transformer encoder into the prediction head to obtain the re-tracking prediction result and store it in the re-tracking evaluation area;

[0038] Take the bounding box with the highest classification score in the re-tracking evaluation area as the final prediction result;

[0039] If the classification score of the final prediction result is higher than the update threshold θ, where θ > the preset expected value τ, then crop a new template image from the current frame image. Specifically, take the center of the predicted bounding box as the center and crop it according to the specification of size H z ×W z to achieve online update of the template image.

[0040] Second aspect, to achieve the above object, the present invention discloses a foreground-aware Transformer object tracking system based on a re-tracking strategy, including:

[0041] An image processing module, configured to receive the most recent frame template image and the search area image, input the most recent frame template image and the search area image into a pre-constructed preprocessing layer to generate a template token sequence and a search token sequence respectively, and cache the template token sequence and the search token sequence into a re-tracking buffer;

[0042] A sequence prediction module, configured to splice the template token sequence and the search token sequence to obtain a hybrid sequence, input the hybrid sequence into a pre-established foreground-aware Transformer encoder to obtain an output sequence, fill the search area feature sequence in the output sequence, and input the filled output sequence into a prediction head to obtain a prediction result;

[0043] A prediction evaluation module, configured to compare the prediction result with a preset expected value. If the prediction results of consecutive t frames are lower than the preset expected value, the template token sequence and the search token sequence in the re-tracking buffer are re-predicted by a general Transformer encoder to obtain the prediction result in the re-tracking evaluation area, and the prediction result is comprehensively evaluated with the prediction result in the re-tracking evaluation area to obtain the best tracking result.

[0044] In yet another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, the foreground-aware Transformer object tracking method based on the re-tracking strategy as described above is adopted.

[0045] Advantages of the present invention:

[0046] The present invention adopts a single-stream single-stage tracking method, which greatly simplifies the tracking architecture; uses a foreground-aware Transformer encoder to extract and fuse template and search area features, while enhancing the perception of foreground objects and suppressing background interference, greatly reducing the computational complexity; uses an adaptive re-tracking strategy to correct tracking deviations in a timely manner, and has stronger adaptability to complex dynamic scenarios; and the online template update strategy avoids tracking failures caused by template aging. Description of the Drawings

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;

[0048] Figure 1 is a flowchart of the target tracking method of the present invention;

[0049] Figure 2 is a working flowchart of the foreground-aware Transformer module of the present invention;

[0050] Figure 3 is a schematic diagram of the working process of the present invention;

[0051] Figure 4 is a schematic diagram of the system structure of the present invention. Specific embodiments

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] Embodiment 1:

[0054] As Figure 1 shown, a foreground-aware Transformer target tracking method based on a re-tracking strategy, the method includes the following steps:

[0055] S101: Receive the most recent frame template image and the search area image, input the most recent frame template image and the search area image into a pre-constructed preprocessing layer to generate a template token sequence and a search token sequence respectively. If the most recent frame template image is the initial frame template image, cache the template token sequence and the search token sequence into the re-tracking buffer, otherwise only update the search token sequence in the re-tracking buffer;

[0056] During the first tracking, the most recent frame template image is the initial frame template image. Input it and the search area image into the preprocessing layer to generate the most recent frame template token sequence (abbreviated as the template token sequence) and the search token sequence, and then cache the template token sequence and the search token sequence into the re-tracking buffer; in subsequent tracking, after the image passes through the preprocessing layer, update the template token sequence and the search token sequence in the re-tracking buffer;

[0057] The most recent frame template image , where the height of the template image is H z and the width is W z and the number of channels is 3. represents the set of real numbers. The image is divided into multiple image patches with side length P and flattened into a sequence of template patches , and the number of template patches is . The search region image , where the height of the template image is H x and the width is W x and the number of channels is 3. The image is divided into multiple image patches with side length P and flattened into a sequence of search patches , and the number of search patches is .

[0058] Using a linear projection layer with a parameter of L, map the sequence of template patches and the sequence of search patches to a D-dimensional space, and add learnable position embeddings to each block embedding respectively to obtain the template token embedding Z and the search token embedding X:

[0059]

[0060] In the formula, z p i is the i-th block in the sequence of template patches, x p i is the i-th block in the sequence of search patches, P z is the position embedding of the template token, P x is the position embedding of the search token, and L represents the linear projection layer with a parameter of L.

[0061] If the template has been updated, store both the template token sequence and the search token sequence of the current frame in the re-tracking buffer; otherwise, only replace the search token sequence in the re-tracking buffer with the search token sequence of the current frame.

[0062] S102: Concatenate the template token sequence and the search token sequence to obtain a mixed sequence, input the mixed sequence into the pre-established foreground-aware Transformer encoder to obtain an output sequence, fill the search region feature sequence in the output sequence, and input the filled output sequence into the prediction head to obtain a prediction result;

[0063] Concatenate Z and X by channel to obtain a mixed token sequence M 0 ;

[0064] Input the mixed token sequence into the foreground-aware Transformer block of the encoder, and obtain the query q of the template z , key kz , value v z and query q of the search area x , key k x , value v x ;

[0065] Feature extraction and feature fusion are performed on the mixed token sequence using foreground-enhanced multi-head self-attention to enhance the foreground area and reduce the attention to the background. Different from the previous multi-head attention, the foreground-enhanced self-attention performs Softmax calculation on the first j maximum values in each row of the attention weight matrix, so as to ensure that the attention is concentrated on the foreground area and reduce the influence of the background on the attention weight. Then, the remaining positions in each row are filled with 0 to discard the weights with low attention and highlight the foreground. The formula is as follows:

[0066]

[0067] In the formula, FEA represents the calculation of a single foreground-enhanced self-attention head, q, k, and v represent the input query vector, key vector, and value vector respectively, and Softmax is the activation function, Maxj represents the first K maximum values in each row of the attention weight matrix, k T represents the transpose of the key k of the mixed sequence, is the dimension of the mixed token sequence, represents a matrix of 0 values, k z T and k x T are the transpose of the key of the template sequence and the transpose of the key of the search sequence respectively. The formula for foreground-enhanced multi-head self-attention is as follows:

[0068]

[0069] In the formula, FEMSA represents the result of foreground-enhanced multi-head attention, Concat represents the concatenation operation, and head i represents the calculation result of the i-th foreground-enhanced self-attention head. The output of the foreground-enhanced multi-head attention module is added to the mixed token sequence M 1 with residual and input to the background screening module.

[0070] Since there is also background in the template, and the center of the template is mostly the target area, the similarity result between the center position of the template and all positions in the search area is selected as the basis for screening the background, and the expression is

[0071]

[0072] In the formula, w iIndicates the similarity calculated by the i-th head. Since there are h attention heads, the center position of the template can calculate h similarity results. In order to obtain a fairer result, the background screening threshold is the average value of h similarity results, expressed as

[0073]

[0074] According to the background screening threshold, all search tokens with a similarity lower than this threshold are deleted, and the retained search tokens need to record their original positions and then reassemble them with the template token to obtain a mixed token sequence M. 2 .

[0075] M 2 After normalization and feedforward network, it is added to its own residual to obtain the output M of the Transformer block. 1 ;

[0076] After N Transformer blocks, the final output M of the encoder is obtained N .

[0077] M N The search token sequence is extracted and restored to its original order according to the initial position embedding of the token. For the missing positions, a 0-filling strategy is adopted to obtain the complete search token sequence X'.

[0078] Reshape X' into a two-dimensional feature map and input it into the three-branch prediction head to obtain the classification score, position deviation, target scale, and generate a predicted bounding box.

[0079] S103: Compare the prediction result with the preset expected value. If the prediction results of t consecutive frames are all lower than the preset expected value, the prediction result is stored in the re-tracking evaluation area, and the template token sequence and the search token sequence in the re-tracking cache are input into the universal Transformer encoder for tracking prediction to obtain the re-tracking prediction result. The prediction result and the re-tracking prediction result are comprehensively evaluated in the re-tracking evaluation area to obtain the best tracking result, and the template image is updated online.

[0080] When the classification scores of consecutive t frames are all lower than the credible threshold τ, it indicates that the current prediction result may be biased. At this time, the re-tracking strategy is decisively enabled to obtain more accurate tracking results.

[0081] Temporarily store the classification results and bounding box results in the re-tracking evaluation area to provide key data reserves for subsequent result evaluation and optimization;

[0082] Take out the template token sequence and the search token sequence in the re-tracking buffer and input them into the general Transformer encoder. Different from the foreground-aware Transformer encoder, the general Transformer encoder replaces the foreground-enhanced multi-head self-attention module with a multi-head self-attention module and reduces the number of Transformer blocks to prevent discarding too many important clues that may be related to the target, improve the tracking accuracy, and reduce the computational burden of re-tracking. Among them, the calculation process of the multi-head self-attention module is as follows:

[0083]

[0084] In the formula, Attention(q,k,v) represents the calculation result of a single self-attention head for the mixed sequence input, and Softmax is the activation function, is the dimension of the mixed token sequence, and head i represents the calculation result of the i-th self-attention head, and MHSA represents the result of the multi-head self-attention.

[0085] Input the output result of the general Transformer encoder into the prediction head to obtain the prediction result of re-tracking and store it in the re-tracking evaluation area;

[0086] Take the bounding box with the highest classification score in the re-tracking evaluation area as the final prediction result;

[0087] If the classification score of the final prediction result is higher than the update threshold θ (θ > the confidence threshold τ), then crop a new template image from the current frame image. Specifically, centered on the center of the predicted bounding box, crop it according to the specification of size H z ×W z to achieve the online update of the template image.

[0088] Specifically, the solution of the present invention will be further elaborated through the following embodiments:

[0089] Such as Figure 2As shown, first, the most recent frame template image and the search area image are input into the preprocessing layer to generate the corresponding template token sequence and search token sequence, and the cache is updated, that is, the cached template token sequence and search token sequence in the re-tracking buffer are updated. Then, the two sequences are concatenated and input into the foreground-aware Transformer encoder for feature extraction and fusion, and the obtained search sequence is input into the prediction head. The prediction head outputs the classification score and the predicted bounding box. If the classification scores of consecutive t frames are all less than the confidence threshold τ, the re-tracking strategy is enabled. At this time, the classification score and the bounding box are cached during re-tracking, that is, after the re-tracking strategy is enabled, the classification score and the bounding box are cached in the re-tracking evaluation area; otherwise, direct prediction is performed and the prediction result is output. After the re-tracking strategy is enabled, the template token and the search area token in the re-tracking buffer are taken out, and both are reprocessed using the general Transformer encoder during re-tracking. The obtained search sequence is input into the prediction head to obtain the classification score and the bounding box. The classification score and the bounding box are input into the re-tracking evaluation area and comprehensively evaluated with the previous classification score and bounding box to output the best prediction result. Finally, if the classification score of the prediction result is greater than the update threshold θ, the current search image is cropped to update the template image.

[0090] As Figure 3 shown, first, the mixed token sequence is input, and layer normalization is performed on it, which helps to stabilize the model training process and accelerate the convergence speed. Then, the mixed token sequence enhances the attention to foreground information through the foreground-enhanced self-attention module, uses the foreground-enhanced self-attention mechanism to capture the dependencies between tokens, and highlights the foreground part. The output of the foreground-enhanced self-attention module is added to the original input residually to retain the original information and prevent gradient disappearance. Then, the background screening module is used to remove the tokens belonging to the background in the search token sequence, making the model more focused on the foreground. Layer normalization is performed again on the output of the background screening module to standardize the data distribution. Subsequently, high-level features are extracted through the feed-forward network, and the output result is added to the output of the background screening module residually to fuse the information. Finally, the mixed token sequence after background screening is output to improve the model's perception and processing ability of foreground objects and improve performance and accuracy.

[0091] Embodiment 2: Second aspect, as Figure 4 shown, to achieve the above object, the present invention discloses a foreground-aware Transformer object tracking system based on a re-tracking strategy, including:

[0092] The image processing module 11 is configured to receive the most recent frame template image and the search area image, input the most recent frame template image and the search area image into a pre-constructed preprocessing layer to generate a template token sequence and a search token sequence respectively, and cache the template token sequence and the search token sequence into the re-tracking buffer;

[0093] The sequence prediction module 12 is configured to splice the template token sequence and the search token sequence to obtain a hybrid sequence, input the hybrid sequence into a pre-established foreground-aware Transformer encoder to obtain an output sequence, fill the search area feature sequence in the output sequence, and input the filled output sequence into a prediction head to obtain a prediction result;

[0094] The prediction evaluation module 13 is configured to compare the prediction result with a preset expected value. If the prediction results of consecutive t frames are lower than the preset expected value, the template token sequence and the search token sequence in the re-tracking buffer are re-predicted by a general Transformer encoder to obtain the prediction result in the re-tracking evaluation area, and the prediction result is comprehensively evaluated with the prediction result in the re-tracking evaluation area to obtain the best tracking result.

[0095] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically to load and execute one or more instructions in the computer storage medium to implement the above method.

[0096] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and when the computer program is run by a processor, the above method is executed. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0097] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0098] The above shows and describes the basic principles, main features, and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and the above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

Claims

1. A foreground-aware Transformer target tracking method based on a re-tracking strategy, characterized in that: The method comprises the following steps: Receiving the most recent frame template image and the search area image, inputting the most recent frame template image and the search area image into a pre-built preprocessing layer to generate a template token sequence and a search token sequence respectively, and caching the template token sequence and the search token sequence into a re-tracking buffer area; The template token sequence and the search token sequence are concatenated to obtain a mixed sequence, and the mixed sequence is input into the pre-established foreground-aware Transformer encoder to obtain an output sequence, and the search area feature sequence in the output sequence is padded, and the padded output sequence is input into the prediction head to obtain a prediction result; After the mixed sequence is input into the pre-established foreground-aware Transformer encoder, a foreground-enhanced multi-head self-attention module is used to extract the features of the fused mixed sequence, and a background-screening module is used to screen and discard background tokens; The calculation process of extracting the features of the fused mixed sequence using the foreground enhanced multi-head self-attention module is as follows: In the formula, FEA Represents the calculation of a single foreground enhancement self-attention head, q, k, v represent the input query vector, key vector, value vector, Softmax is the activation function, Maxj represents the K largest values ​​in each row of the attention weight matrix, k T represents the transpose of the mixed sequence key k, is the dimension of the mixed token sequence, represents a zero-value matrix, k z T and k x T They are the transposition of the template sequence key and the transposition of the search sequence key respectively. The foreground enhanced multi-head self-attention calculation formula is as follows: Where FEMSA represents the foreground enhanced multi-head attention result, Concat Indicates the splicing operation, head i represents the calculation result of the i-th foreground enhancement self-attention head. The output of the foreground enhancement multi-head attention module is added to the residual of the mixed token sequence M1 and input into the background removal module; The prediction result is compared with the preset expected value. If the prediction results of t consecutive frames are lower than the preset expected value, the prediction result is stored in the re-tracking evaluation area, and the template token sequence and search token sequence in the re-tracking buffer are input into the universal Transformer encoder for tracking prediction to obtain the re-tracking prediction result. The prediction result and the re-tracking prediction result are comprehensively evaluated in the re-tracking evaluation area to obtain the best tracking result, and the template image is updated online.

2. The foreground-aware Transformer target tracking method based on re-tracking strategy according to claim 1 is characterized in that: The most recent frame template image , where the template image height is H z , width is W z , Represents a real number set. The template image of the latest frame is divided into multiple image blocks with a side length of P and flattened into a template block sequence , the number of template blocks is , search area image , where the template image height is H x , width is W x , the search area image is divided into multiple image blocks with a side length of P and flattened into a search block sequence , the number of search blocks is .

3. The foreground-aware Transformer target tracking method based on re-tracking strategy according to claim 2 is characterized in that: The generation process of the template token sequence and the search token sequence: Using a linear projection layer with an input parameter of L, the template block sequence and the search block sequence are mapped to a D-dimensional space, and a learnable position embedding is added to each block embedding to obtain the template token embedding Z and the search token embedding X: In the formula, z p i is the i-th block in the template block sequence, x p i is the i-th block in the search block sequence, P z is the position embedding of the template token, P x To search for the position embedding of the token, L represents a linear projection layer with parameter L.

4. The foreground-aware Transformer target tracking method based on re-tracking strategy according to claim 3 is characterized in that: The mixed sequence is obtained by splicing Z and X according to the channel. 0 ; The mixed sequence is input into the pre-established foreground-aware Transformer encoder and the query q of the template is obtained after the normalization layer. z , key k z 、value v z and the query q for the search area x , key k x 、value v x .

5. The foreground-aware Transformer target tracking method based on re-tracking strategy according to claim 4 is characterized in that: The process of using the background screening module to filter and discard background tokens: The similarity between the center position of the template and all positions in the search area is selected as the basis for screening the background. The expression is: In the formula, w i It represents the similarity calculated by the i-th head. Since there are h attention heads, the center position of the template can calculate h similarity results. The background removal threshold is the average value of h similarity results, expressed as: According to the background screening threshold, the search tokens with a lower similarity are deleted, and the retained search tokens need to record their original positions and then reassembled with the template token to obtain the mixed token sequence M2; After normalization and feedforward network, M2 is added to its own residual to obtain the output M of the Transformer block. 1 ; After N Transformer blocks, the final output M of the encoder is obtained N .

6. The foreground-aware Transformer target tracking method based on re-tracking strategy according to claim 5 is characterized in that: M N The search token sequence is extracted and restored to the initial order according to the initial position embedding of the token. For the missing positions, a 0-filling strategy is adopted to obtain the complete search token sequence X'; Reshape X' into a two-dimensional feature map and input it into the three-branch prediction head to obtain the classification score, position deviation, and target scale respectively, and generate the predicted bounding box; Compare the prediction result with the preset expected value τ. If the value of consecutive t frames is lower than τ, it indicates that there is a deviation in the current prediction result. At this time, the re-tracking strategy is started, and the classification result and the bounding box result are temporarily stored in the re-tracking evaluation area. The template token sequence and search token sequence in the re-tracking buffer are taken out and input into the universal Transformer encoder. The universal Transformer encoder replaces the foreground enhancement multi-head self-attention module with a multi-head self-attention module. The calculation process of the multi-head self-attention module is as follows: In the formula, Attention(q,k,v) represents the calculation result of a single self-attention head for mixed sequence input, Softmax is the activation function, is the dimension of the mixed token sequence, head i represents the calculation result of the i-th self-attention head, and MHSA represents the multi-head self-attention result; The output result of the general Transformer encoder is input into the prediction head to obtain the prediction result of the re-tracking and store it in the re-tracking evaluation area; Take the bounding box with the highest classification score in the re-tracking evaluation area as the final prediction result; If the classification score of the final prediction result is higher than the update threshold θ, where θ> the preset expected value τ, a new template image is cropped from the current frame image: centered on the predicted bounding box, with a size of H z ×W z Cut to specifications.

7. A foreground-aware Transformer target tracking system based on a re-tracking strategy, which adopts the foreground-aware Transformer target tracking method based on a re-tracking strategy as claimed in any one of claims 1 to 6, characterized in that: include: An image processing module is used to receive a most recent frame template image and a search area image, input the most recent frame template image and the search area image into a pre-built pre-processing layer to generate a template token sequence and a search token sequence respectively, and cache the template token sequence and the search token sequence in a re-tracking buffer area; The sequence prediction module is used to concatenate the template token sequence and the search token sequence to obtain a mixed sequence, input the mixed sequence into the pre-established foreground-aware Transformer encoder to obtain an output sequence, fill the search area feature sequence in the output sequence, and input the filled output sequence into the prediction head to obtain a prediction result; The prediction evaluation module is used to compare the prediction result with the preset expected value. If the prediction result of t consecutive frames is lower than the preset expected value, the template token sequence and the search token sequence in the re-tracking buffer area are predicted again through the universal Transformer encoder to obtain the prediction result of the re-tracking evaluation area. The prediction result is comprehensively evaluated with the prediction result of the re-tracking evaluation area to obtain the best tracking result.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the foreground-aware Transformer target tracking method based on the re-tracking strategy described in any one of claims 1 to 6 is adopted.

Citation Information

Patent Citations

  • Vehicle tracking method and system

    CN115797410A

  • Single target tracking method and tracking system based on background weakening mechanism

    CN116721130A