Visual Object Tracking Method, Device and Electronic Device Based on Efficient Sequence Generation

By adopting an efficient sequence generation method in visual target tracking, using the visual transformer framework and tracking prediction head, the problems of high computing burden and memory utilization in the prior art are solved, and a faster and more efficient tracking process is achieved.

CN119785355BActive Publication Date: 2025-06-10NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510291919.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Existing visual target tracking algorithms have high demands in terms of computational burden and memory usage, resulting in tracking latency and deployment challenges, especially on resource-constrained edge devices.

Method used

The visual target tracking method based on efficient sequence generation is adopted, and by obtaining search images, template images and randomly initialized tracking marks, blocking, linear projection and position encoding are performed. The encoder based on the visual transformer framework is used to perform global self-attention operations, and the target tracking results are generated by combining the tracking prediction head and transformer decoder.

Benefits of technology

Increases the speed of forward inference, reduces computational latency, improves tracking speed, and demonstrates state-of-the-art performance in multiple tracking benchmarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785355B_ABST
    Figure CN119785355B_ABST
Patent Text Reader

Abstract

The present application relates to a visual object tracking method, device and electronic device based on efficient sequence generation. The method includes: processing the acquired search image, template image and randomly initialized tracking marker to obtain a search visual embedding and a template visual embedding; concatenating the search and template visual embeddings and the tracking marker for position encoding, then using an encoder based on a visual transformer framework to encode the position encoding result, and processing the encoded feature of the tracking marker with a tracking prediction head to obtain a target tracking prediction result; if the target tracking prediction result meets a preset condition, then using the target tracking prediction result as the target tracking result; otherwise, using a transformer decoder to decode the encoding structure; processing according to the decoding result with a tracking prediction head to obtain the target tracking result. Using this method improves the forward inference speed and the tracking speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of visual object tracking, and particularly to a visual object tracking method, device, and electronic device based on efficient sequence generation. Background Art

[0002] Object tracking has always been a hot and difficult problem in the field of computer vision. Its definition is to estimate the state information (position, size, etc.) of an object in the entire video given the position and size of the object in the first frame of the video (usually a rectangular box tightly enclosing the object). Existing state-of-the-art tracking algorithms usually adopt powerful feature extraction backbones, complex feature fusion modules, and prediction head networks to achieve high accuracy in public benchmark tests. However, these trackers have high computational burdens and memory usage rates, which lead to tracking delays and deployment challenges, especially on resource-constrained edge devices. Therefore, how to achieve a good balance between tracking accuracy and efficiency remains a key issue faced by the tracking community.

[0003] In recent years, tracking methods represented by SeqTrack model the tracking task as a sequence generation task and are widely popular for their simple network architecture (transformer encoder-decoder network architecture) and loss function (cross-entropy loss function). The SeqTrack tracking process is mainly divided into two steps: 1) Joint feature extraction and feature fusion are performed from the template and search region using a transformer encoder; 2) A transformer decoder is used to generate a sequence of bounding box values (start X Y W H) (where X and Y are the upper left coordinates of the object; W and H are the width and height of the object) in an autoregressive manner from a randomly initialized start token and search image features. Although successful in both single-modal and multi-modal tracking tasks, SeqTrack has the following two problems: 1) The four values (X Y W H) of the object bounding box sequence are generated one by one, which means that the decoder must run four times per frame during forward inference, inevitably resulting in high computational latency; 2) The tracking tokens are randomly initialized, which poses a challenge for the decoder to accurately predict the bounding box sequence. Summary of the Invention

[0004] Based on this, it is necessary to provide a visual object tracking method, device, and electronic device based on efficient sequence generation for the above technical problems.

[0005] A visual object tracking method based on efficient sequence generation, the method comprising:

[0006] Obtain a search image, a template image, and randomly initialized tracking tokens.

[0007] The search image and the template image are respectively segmented and linearly projected to obtain a search visual embedding and a template visual embedding.

[0008] After connecting the search visual embedding, the template visual embedding, and the randomly initialized tracking token, position encoding is performed to obtain the input tokens after position encoding.

[0009] An encoder based on the vision transformer framework is used to perform global self-attention operations on all the input tokens after position encoding to obtain the tracking token encoded features and the search image encoded features.

[0010] The tracking token encoded features are processed by a tracking prediction head to obtain the target tracking prediction result.

[0011] If the target tracking prediction result meets the preset conditions, the target tracking prediction result is used as the target tracking result.

[0012] If the target tracking prediction result does not meet the preset conditions, a transformer decoder is used to decode the tracking token encoded features and the search image encoded features; according to the decoding result, it is processed by a tracking prediction head to obtain the target tracking result; the target tracking result is a sequence of bounding boxes.

[0013] In one embodiment, the process of randomly initializing the tracking token includes: using the method of initializing the token in VIT to randomly initialize the tracking token.

[0014] In one embodiment, after connecting the search visual embedding, the template visual embedding, and the randomly initialized tracking token, position encoding is performed, and the input tokens after position encoding are:

[0015]

[0016] Among them, is the input token after position encoding, is the randomly initialized tracking token, is the visual embedding of the search image, is the visual embedding of the template image; is the position encoding, is the set of real numbers, is the dimension, is the number of blocks.

[0017] In one embodiment, using an encoder based on the vision transformer framework to perform global self-attention operations on all the input tokens after position encoding to obtain the tracking token encoded features and the search image encoded features includes:

[0018] The position-encoded input tokens are input into an encoder based on a vision Transformer framework to achieve the interaction between the template features and the search features, the interaction between the tracking tokens and the search features, and the interaction between the tracking tokens and the template features. The encoding results corresponding to the tracking tokens and the search image in the output of the last layer of the encoder are output to a decoder.

[0019] In one embodiment, the Transformer decoder includes multiple decoder layers; each decoder layer contains an unmasked multi-head self-attention block, a multi-head attention block, and a feed-forward network block; in the first decoder layer:

[0020] The encoded features of the tracking tokens interact with each other through an unmasked multi-head attention mechanism and are used as queries. Based on the multi-head attention block, an interaction operation is performed between the queries and the search image features, and then the output is processed by the feed-forward network block and output to a tracking prediction head to obtain the bounding box of the target.

[0021] In one embodiment, the Transformer decoder is used to decode the encoded features of the tracking tokens and the encoded features of the search image; according to the decoding results, the tracking prediction head is used for processing to obtain the target tracking result; the target tracking result is a sequence of bounding boxes, including:

[0022] The encoded features of the tracking tokens and the encoded features of the search image are input into the Transformer decoder for decoding.

[0023] The obtained decoding results are input into the tracking prediction head to obtain a sequence of bounding boxes; the tracking prediction head is used to map the output of the decoder layer to an integer in the vocabulary V by using an embedding-to-word network, then calculate the SOFTMAX probability through the softmax function, and finally sample words from the vocabulary V according to the SOFTMAX probability to predict and obtain the sequence of bounding boxes.

[0024] In one embodiment, the Transformer decoder is used to decode the encoded features of the tracking tokens and the encoded features of the search image; according to the decoding results, the tracking prediction head is used for processing to obtain the target tracking result; the target tracking result is a sequence of bounding boxes, including:

[0025] The encoded features of the tracking tokens and the encoded features of the search image are input into the first encoder layer of the Transformer decoder, and the output of the first encoder layer is processed by the tracking prediction head to obtain the SOFTMAX probability of the four bounding box values of the target.

[0026] If the SOFTMAX probability of the four bounding box values of the target is not greater than the preset threshold, the output of the first-layer decoder layer is used as the input of the second-layer decoder, and decoding and tracking result judgment are continued.

[0027] If the SOFTMAX probability of the four bounding box values of the target is greater than the preset threshold, the forward propagation process of the transformer decoder is terminated and the tracking result is output.

[0028] A visual object tracking device based on efficient sequence generation, the device includes:

[0029] An input module for obtaining a search image, a template image, and a randomly initialized tracking marker.

[0030] A position encoding module for respectively partitioning and linearly projecting the search image and the template image to obtain a search visual embedding and a template visual embedding; connecting the search visual embedding, the template visual embedding, and the randomly initialized tracking marker and performing position encoding to obtain the input marker after position encoding.

[0031] An encoding module for performing global self-attention operations on all the input markers after position encoding by using an encoder based on a visual transformer framework to obtain a tracking marker encoding feature and a search image encoding feature.

[0032] A decoding module for processing the tracking marker encoding feature by using a tracking prediction head to obtain a target tracking prediction result; if the target tracking prediction result meets the preset condition, the target tracking prediction result is used as the target tracking result; if the target tracking prediction result does not meet the preset condition, the tracking marker encoding feature and the search image encoding feature are decoded by using a transformer decoder; the target tracking result is obtained by processing according to the decoding result by using the tracking prediction head; the target tracking result is a bounding box sequence.

[0033] The above-mentioned visual object tracking method, device, and electronic device based on efficient sequence generation, the method includes: obtaining a search image, a template image, and a randomly initialized tracking mark; respectively performing block division and linear projection on the search image and the template image to obtain a search visual embedding and a template visual embedding; connecting the search and template visual embeddings and the tracking mark and performing position encoding; using an encoder based on the visual transformer framework to encode the obtained position encoding result, and processing the encoded feature of the tracking mark using a tracking prediction head to obtain a target tracking prediction result; if the target tracking prediction result meets a preset condition, then using the target tracking prediction result as the target tracking result; if the target tracking prediction result does not meet the preset condition, then using a transformer decoder to decode the encoded feature of the tracking mark and the encoded feature of the search image; processing according to the decoding result using a tracking prediction head to obtain a target tracking result; the target tracking result is a bounding box sequence. Using this method improves the forward inference speed and the tracking speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 FIG. is a schematic flowchart of a visual object tracking method based on efficient sequence generation in one embodiment;

[0035] Figure 2 FIG. is a flowchart of a visual object tracking method based on efficient sequence generation in another embodiment;

[0036] Figure 3 FIG. is a schematic structural diagram of a decoder layer in another embodiment;

[0037] Figure 4 FIG. is an internal structural diagram of an electronic device in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0039] In one embodiment, as Figure 1 shown, a visual object tracking method based on efficient sequence generation is provided, and the method includes the following steps:

[0040] Step 100: Obtain a search image, a template image, and a randomly initialized tracking mark.

[0041] Specifically, obtain a piece of search image , a piece of template image , and four randomly initialized tracking marks .

[0042] Step 102: Perform block division and linear projection on the search image and the template image respectively to obtain the search visual embedding and the template visual embedding.

[0043] Specifically, divide the search image and the target image into blocks respectively to obtain search image blocks and template image blocks , where represents the block size, represents the number of blocks.

[0044] After the action of the linear projection matrix , project the search image blocks and the template image blocks onto and respectively. Use the method of initializing tokens in VIT to randomly initialize the tracking tokens as .

[0045] Step 104: Connect the search visual embedding, the template visual embedding, and the randomly initialized tracking tokens, and then perform position encoding to obtain the input tokens after position encoding.

[0046] Specifically, along the spatial dimension, connect the tracking tokens, the template, and the search image features, add position encoding, and then input them into the encoder. The input of the encoder is expressed as: , where is the position encoding.

[0047] Step 106: Use the encoder based on the visual transformer framework to perform global self-attention operations on all the input tokens after position encoding to obtain the tracking token encoding features and the search image encoding features.

[0048] Specifically, in the forward propagation process, the encoder performs global self-attention operations on all the input tokens, which can be decomposed into three different cross-correlation operations, namely , and . Among them, realizes the interaction between the template features and the search features, and realize the interaction between the tracking tokens and the template and search features. Output the tracking tokens output by the last layer of the encoder and the search image features to the decoder.

[0049] Embed the four learnable tracking tokens, templates, and search regions corresponding to the target four bounding box values (X, Y, W, H), and input the embedded template tokens and search tokens into the encoder together. While extracting features, refine the tracking tokens. These tracking tokens are similar to the anchor concept in object detection tasks and can initially indicate the target location.

[0050] The encoder based on the vision transformer framework can adopt but is not limited to the ViT-B structure (An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.).

[0051] Step 108: Process the encoded features of the tracking tokens using a tracking prediction head to obtain the target tracking prediction result; if the target tracking prediction result meets the preset conditions, use the target tracking prediction result as the target tracking result; if the target tracking prediction result does not meet the preset conditions, use the transformer decoder to decode the encoded features of the tracking tokens and the encoded features of the search image; process according to the decoding result using the tracking prediction head to obtain the target tracking result; the target tracking result is a sequence of bounding boxes.

[0052] Specifically, after the encoder, the encoded features of the tracking tokens will be directly input into the embedding-to-word network to predict the target sequence. When the tracking result meets the preset conditions, the entire subsequent decoder structure is directly omitted. When the tracking effect does not meet the requirements, it is input into the decoder for layer-by-layer operations.

[0053] Use the transformer decoder to regressively predict four bounding box values from the four tracking tokens and the search image features at one time. The decoder only needs to perform one forward propagation operation to obtain the four bounding box values of the target, which greatly improves the speed of forward inference.

[0054] This method is used for an efficient sequence generation tracking framework for single-object tracking. This framework better initializes the tracking tokens and generates a sequence of bounding boxes at one time.

[0055] The process of using the trained network model (the network model proposed in the method of this application) to track the target online includes: during online tracking, the encoder receives the search image, the initial template image, the dynamic template image, and four randomly initialized tracking tokens, and inputs the calculated tracking tokens and search image features into the decoder to predict the target sequence in parallel. At the same time, following the SeqTrack operation, the online template update interval is set to 1, and the judgment threshold in Algorithm 1 Set to 1.6; for the target sequence predicted by the encoder, based on the Hamming window combined position prior, the final tracking result is obtained.

[0056] This method is mainly improved on the classic sequence generation method SeqTrack. After connecting the search feature, template feature, and randomly initialized tracking token and performing position encoding, it can be distinguished from SeqTrack when input into the encoder, and it can better initialize the tracking token and improve the tracking accuracy. Compared with the sequence generation method that generates the target sequence one by one each time (first generate x, then y, then w, then h), this method can predict and output at one time.

[0057] In the above visual object tracking method based on efficient sequence generation, the method includes: obtaining a search image, a template image, and a randomly initialized tracking token; respectively performing block division and linear projection on the search image and the template image to obtain a search visual embedding and a template visual embedding; connecting the search and template visual embeddings and the tracking token and performing position encoding; using an encoder based on the visual transformer framework to encode the obtained position encoding result, and processing the encoded feature of the tracking token using a tracking prediction head to obtain a target tracking prediction result; if the target tracking prediction result meets the preset condition, then using the target tracking prediction result as the target tracking result; if the target tracking prediction result does not meet the preset condition, then using a transformer decoder to decode the encoded feature of the tracking token and the encoded feature of the search image; processing according to the decoding result using a tracking prediction head to obtain a target tracking result; the target tracking result is a bounding box sequence. Using this method improves the speed of forward inference and the tracking speed.

[0058] In one embodiment, the process of randomly initializing the tracking token in step 100 includes: using the method of initializing the token in VIT to randomly initialize the tracking token.

[0059] In one embodiment, step 104 includes: connecting the search visual embedding, the template visual embedding, and the randomly initialized tracking token and performing position encoding, and the input token after position encoding is:

[0060]

[0061] Where, is the input token after position encoding, is the randomly initialized tracking token, is the visual embedding of the search image, is the visual embedding of the template image; is the position encoding, is the set of real numbers, is the dimension, is the number of blocks.

[0062] In one embodiment, step 106 includes: inputting the position-encoded input tokens into an encoder based on a vision Transformer framework to achieve the interaction between the template feature and the search feature, the interaction between the tracking tokens and the search feature, and the interaction between the tracking tokens and the template feature, and outputting the encoding results corresponding to the tracking tokens and the search image in the output of the last layer of the encoder to the decoder.

[0063] In one embodiment, the Transformer decoder in step 108 includes M decoder layers; M is an integer greater than 0 and less than or equal to 6; as Figure 3 shown, each decoder layer contains an unmasked self-attention block, a multi-head attention block, and a feed-forward network (FFN) block; in the first decoder layer: the encoded features of the tracking tokens are used as queries after interacting through an unmasked multi-head attention mechanism, perform the interaction operation between the queries and the search image features based on the multi-head attention block, and then output to the tracking prediction head after being processed by the feed-forward network block to obtain the bounding box of the target.

[0064] Specifically, the Transformer decoder is composed of M decoder layers, and each decoder layer contains an unmasked self-attention block, a multi-head attention block, and a feed-forward network (FFN) block. The tracking tokens output by the encoder are used as queries, and perform the interaction operation based on the multi-head attention block to predict the bounding box of the target in one-time learning. Different from SeqTrack that randomly initializes the tracking tokens, the tracking tokens of this method are initialized in advance in the encoder, which helps the decoder to accurately predict the bounding box of the target. Preferably, the number of decoder layers of the Transformer decoder is set to 2.

[0065] Starting from the second decoder layer, the output of the previous decoder layer is the input of the next decoder layer.

[0066] In one embodiment, a Transformer decoder is used to decode the encoded features of the tracking markers and the encoded features of the search image; the tracking prediction head is used to process the decoding result to obtain the target tracking result; the target tracking result is a sequence of bounding boxes, including: inputting the encoded features of the tracking markers and the encoded features of the search image into the Transformer decoder for decoding; inputting the obtained decoding result into the tracking prediction head to obtain a sequence of bounding boxes; the tracking prediction head is used to map the output of the decoder layer to an integer in the vocabulary V by using an embedding-to-word network, then calculate the SOFTMAX probability through the softmax function, and finally sample words from the vocabulary V according to the SOFTMAX probability to predict the sequence of bounding boxes.

[0067] Specifically, the tracking prediction head includes an embedding-to-word network. The embedding-to-word network is added to the last layer of the decoder as the tracking prediction head. The embedding-to-word network includes a fully connected network (FCN) and a softmax function, and can predict the probabilities of the four bounding box values of the target based on the tracking markers output by the decoder, and then output the target bounding box.

[0068] Each coordinate of the target bounding box [x, y, w, h] is uniformly discretized into an integer within the vocabulary The embedding-to-word network maps the high-dimensional tracking markers output by the decoder to an integer between [1, n bins], corresponding to the words in V. The embedding-to-word network is implemented through a fully connected network (FCN) and a softmax function (softmax). The FCN maps the tracking markers to an integer between in the vocabulary After calculation by the softmax function, the probability is obtained. Finally, the target box b = [x, y, w, h] is predicted by sampling words from the vocabulary according to the SOFTMAX probability. The SOFTMAX probability here is also the basis for judging whether the tracking effect is reliable later.

[0069] In one embodiment, a Transformer decoder is used to decode the tracking token encoded feature and the search image encoded feature; the tracking prediction head is used to process the decoding result to obtain the target tracking result; the target tracking result is a bounding box sequence, including: inputting the tracking token encoded feature and the search image encoded feature into the first encoder layer of the Transformer decoder, and processing the output of the first encoder layer through the tracking prediction head to obtain the SOFTMAX probability of the four target bounding box values; if the SOFTMAX probability of the four target bounding box values is not greater than the preset threshold, the output of the first decoder layer is used as the input of the second decoder, and decoding and tracking result judgment are continued; if the SOFTMAX probability of the four target bounding box values is greater than the preset threshold, the forward propagation process of the Transformer decoder is terminated and the tracking result is output.

[0070] Specifically, in addition, in order to further improve the tracking efficiency, an early exit mechanism is introduced in the decoder part of the network. The tracking effect is judged layer by layer after each decoder layer. When the tracking result is reliable enough, the forward propagation process of the decoder is terminated in advance and the tracking result is output.

[0071] The early exit mechanism is seamlessly integrated into the Transformer decoder, and the forward propagation process is terminated in time according to the prediction accuracy, thereby further improving the tracking speed.

[0072] The visual object tracking method based on efficient sequence generation is as Figure 2 shown. An embedding-to-word network is added after each decoder layer to calculate the average softmax probability of the four bounding box values. If the average probability exceeds a specific threshold , it is determined that the predicted bounding box is reliable enough. At this time, the forward propagation process of the decoder is terminated and the tracking result is output. The pseudo code of the early exit mechanism is shown in Algorithm 1. It should be noted that the parameters of the embedding-to-word network (embedding-to-word network) used after each decoder layer are shared, avoiding an increase in the parameters of the tracking model. The embedding-to-word network includes multiple layers of perceptrons and a softmax layer.

[0073] Algorithm 1: Early Exit Mechanism in Decoder

[0074] Input: Search image encoded feature , tracking token encoded feature ;

[0075] Parameter: The l-th layer decoder ;

[0076] Embedding-to-word network;

[0077] Output: target bounding box ;

[0078] Step 1: Let the input ; early exit flag = False

[0079] Step 2: For each decoder layer from 1 to N , loop the following operations;

[0080] Step 3: Calculate the output of the th layer ;

[0081] Step 4: Calculate the target bounding box and the confidence of the target bounding box based on embedding-to-word ;

[0082] Step 5: If the confidence is greater than the threshold, i.e., , then:

[0083] Step 6: early exit flag = True;

[0084] Step 7: Return the target bounding box ;

[0085] Step 8: End;

[0086] Step 9: End;

[0087] Step 10: If the early exit flag is False after all decoder layers are calculated, then:

[0088] Step 11: Return the target bounding box ;

[0089] Step 12: End.

[0090] The tracking network architecture (abbreviation: FastSeqTrack model) adopted in the method proposed in this application consists of four parts: a) an encoder for joint feature extraction, feature fusion, and refinement of tracking markers; b) a decoder for searching for interactions between image features and tracking markers; c) a tracking prediction head based on the embedding-to-word network; d) an early exit mechanism.

[0091] 1) The loss function during the training process of the FastSeqTrack model combines cross-entropy loss and generalized IoU loss:

[0092]

[0093] Among them, denotes the target ground truth bounding box, denotes the bounding box predicted by the decoder layer, is a hyperparameter of the cross - entropy and generalized IOU loss functions, is the set of real numbers, denotes the number of training samples.

[0094] 2) Training settings. The training data includes the training data in the COCO, LaSOT, GOT - 10k, TrackingNet, and VastTrack datasets. The training parameters follow the parameter settings in the SeqTrack tracking method. The FastSeqTrack model is trained with 60k images for a total of 500 epochs, and the learning rate is reduced by 10 times after 400 epochs. Preferably, the vocabulary in is set to 4000, and are set to 1 and 5 respectively.

[0095] The running speed of this method is 3 times faster than SeqTrack, exceeding 100 fps, and demonstrates state - of - the - art performance in multiple tracking benchmarks.

[0096] The early - exit mechanism is an optimization measure proposed in this application to improve the tracking speed; the early - exit mechanism is seamlessly integrated into the transformer decoder, and the forward propagation process is terminated in a timely manner according to the prediction accuracy, thereby further improving the tracking speed.

[0097] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0098] At least a part of the steps in

[0099] The input module is used to obtain the search image, the template image, and a randomly initialized tracking token.

[0100] A position encoding module, which is used to perform chunking and linear projection on the search image and the template image respectively to obtain a search visual embedding and a template visual embedding; after connecting the search visual embedding, the template visual embedding, and a randomly initialized tracking token, perform position encoding to obtain the input tokens after position encoding.

[0101] An encoding module, which is used to perform global self-attention operations on all the input tokens after position encoding by using an encoder based on the vision transformer framework to obtain a tracking token encoding feature and a search image encoding feature.

[0102] A decoding module, which is used to process the tracking token encoding feature by using a tracking prediction head to obtain a target tracking prediction result; if the target tracking prediction result meets a preset condition, then use the target tracking prediction result as the target tracking result; if the target tracking prediction result does not meet the preset condition, then use a transformer decoder to decode the tracking token encoding feature and the search image encoding feature; perform processing by using the tracking prediction head according to the decoding result to obtain the target tracking result; the target tracking result is a sequence of bounding boxes.

[0103] In one embodiment, the process of randomly initializing the tracking token in the input module includes: using the method of initializing tokens in VIT to randomly initialize the tracking token.

[0104] In one embodiment, the position encoding module is further used to perform position encoding after connecting the search visual embedding, the template visual embedding, and the randomly initialized tracking token to obtain the input tokens after position encoding; the input tokens are as shown in the expression of the input tokens after position encoding described above.

[0105] In one embodiment, the encoding module is further used to input the input tokens after position encoding into an encoder based on the vision transformer framework to implement the interaction between the template feature and the search feature, the interaction between the tracking token and the search feature, and the interaction between the tracking token and the template feature, and output the encoding results corresponding to the tracking token and the search image in the output of the last layer of the encoder to the decoder.

[0106] In one embodiment, the transformer decoder in the decoding module includes layer decoder layers; wherein, each decoder layer contains an unmasked self-attention block, a multi-head attention block, and a feed-forward network (FFN) block; in the first decoder layer: the tracking token encoding feature is used as a query after interacting through an unmasked multi-head attention mechanism, perform an interaction operation between the query and the search image feature based on the multi-head attention block, and then output to the tracking prediction head after being processed by the feed-forward network block to obtain the bounding box of the target.

[0107] In one embodiment, the decoding module is further configured to input the tracking marker encoding feature and the search image encoding feature into a Transformer decoder for decoding; input the obtained decoding result into a tracking prediction head to obtain a bounding box sequence; the tracking prediction head is used to map the output of the decoder layer to an integer in the vocabulary V by using an embedding-to-word network, then calculate the SOFTMAX probability through a softmax function, and finally sample words from the vocabulary V according to the SOFTMAX probability to predict and obtain the bounding box sequence.

[0108] In one embodiment, the decoding module is further configured to input the tracking marker encoding feature and the search image encoding feature into the first encoder layer of the Transformer decoder, and process the output of the first encoder layer through the tracking prediction head to obtain the SOFTMAX probability of the four target bounding box values; if the SOFTMAX probability of the four target bounding box values is not greater than a preset threshold, then use the output of the first decoder layer as the input of the second decoder to continue decoding and tracking result judgment; if the SOFTMAX probability of the four target bounding box values is greater than the preset threshold, then terminate the forward propagation process of the Transformer decoder and output the tracking result.

[0109] For the specific limitations of the visual object tracking device based on efficient sequence generation, reference can be made to the limitations of the visual object tracking method based on efficient sequence generation in the above text, which will not be elaborated here. Each module in the above visual object tracking device based on efficient sequence generation can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0110] In one embodiment, an electronic device is provided. The electronic device can be a terminal, and its internal structure diagram can be as Figure 4As shown in the figure. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a visual target tracking method based on efficient sequence generation. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0111] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0112] In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above method embodiment.

[0113] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0114] The above-described embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A visual target tracking method based on efficient sequence generation, characterized in that: The method comprises: Get the search image, template image, and randomly initialized tracking markers; The search image and the template image are divided into blocks and linearly projected to obtain a search visual embedding and a template visual embedding; The search visual embedding, the template visual embedding and the randomly initialized tracking mark are connected and position-encoded to obtain a position-encoded input mark; The encoder based on the visual transformer framework is used to perform global self-attention operations on all input tags after position encoding to obtain tracking tag encoding features and search image encoding features; The tracking mark coding feature is processed by a tracking prediction head to obtain a target tracking prediction result; If the target tracking prediction result meets the preset condition, the target tracking prediction result is used as the target tracking result; If the target tracking prediction result does not meet the preset conditions, a transformer decoder is used to decode the tracking mark coding feature and the search image coding feature; a tracking prediction head is used to process according to the decoding result to obtain a target tracking result; the target tracking result is a bounding box sequence; Among them, an encoder based on the visual transformer framework is used to perform global self-attention operations on all input tags after position encoding to obtain tracking tag encoding features and search image encoding features, including: The position-encoded input mark is input into the encoder based on the visual transformer framework to realize the interaction between template features and search features, the interaction between tracking marks and search features, and the interaction between tracking marks and template features. The encoding results corresponding to the tracking marks and search images in the last layer output of the encoder are output to the decoder.

2. The visual target tracking method based on efficient sequence generation according to claim 1, characterized in that: The process of randomly initializing the tracking mark includes: using the VIT The token initialization method randomly initializes the tracking token.

3. The visual target tracking method based on efficient sequence generation according to claim 1, characterized in that: The search visual embedding, the template visual embedding and the randomly initialized tracking mark are connected and position encoded to obtain the position encoded input mark: ; in, is the position-encoded input token, are randomly initialized tracking markers, is the set of real numbers, For the dimension, To search for visual embeddings of images, is the visual embedding of the template image, is the number of blocks, is a positional encoding.

4. The visual target tracking method based on efficient sequence generation according to claim 1, characterized in that: The transformer decoder includes M decoder layers; each decoder layer includes an unmasked multi-head self-attention block, a multi-head attention block and a feedforward network block; in the first decoder layer: The tracking marker encoding features interact with each other through an unmasked multi-head attention mechanism as a query, and the interaction between the query and search image features is performed based on the multi-head attention block, and then processed by the feedforward network block and output to the tracking prediction head to obtain the target's bounding box.

5. The visual target tracking method based on efficient sequence generation according to claim 4 is characterized in that: Decoding the tracking mark encoding feature and the search image encoding feature using a transformer decoder; The decoding result is processed using a tracking prediction head to obtain a target tracking result; the target tracking result is a bounding box sequence, including: Inputting the tracking mark encoding feature and the search image encoding feature into a transformer decoder for decoding; The obtained decoding result is input into the tracking prediction head to obtain a bounding box sequence; the tracking prediction head is used to map the output of the decoder layer to an integer in the vocabulary V using an embedding-to-word network, and then calculate the SOFTMAX probability through a softmax function, and finally sample words from the vocabulary V according to the SOFTMAX probability to predict the bounding box sequence.

6. The visual target tracking method based on efficient sequence generation according to claim 4, characterized in that: Decoding the tracking mark encoding feature and the search image encoding feature using a transformer decoder; The decoding result is processed using a tracking prediction head to obtain a target tracking result; the target tracking result is a bounding box sequence, including: Inputting the tracking mark coding features and the search image coding features into the first encoder layer of the transformer decoder, and processing the output of the first encoder layer through the tracking prediction head to obtain the SOFTMAX probability of the four bounding box values ​​of the target; If the SOFTMAX probability of the four bounding box values ​​of the target is not greater than the preset threshold, the output of the first decoder layer is used as the input of the second decoder layer to continue decoding and tracking result judgment; If the SOFTMAX probability of the four bounding box values ​​of the target is greater than the preset threshold, the forward propagation process of the transformer decoder is terminated and the tracking result is output.

7. A visual target tracking device based on efficient sequence generation, characterized in that: The device comprises: Input module, used to obtain the search image, template image and randomly initialized tracking markers; A position encoding module is used to perform block division and linear projection on the search image and the template image respectively to obtain a search visual embedding and a template visual embedding; the search visual embedding, the template visual embedding and a randomly initialized tracking mark are connected and then position encoded to obtain an input mark after position encoding; The encoding module is used to perform global self-attention operations on all input tags after position encoding using an encoder based on the visual transformer framework to obtain tracking tag encoding features and search image encoding features; A decoding module, used for processing the tracking mark coding feature by using a tracking prediction head to obtain a target tracking prediction result; if the target tracking prediction result meets a preset condition, the target tracking prediction result is used as the target tracking result; if the target tracking prediction result does not meet the preset condition, a transformer decoder is used to decode the tracking mark coding feature and the search image coding feature; according to the decoding result, the tracking prediction head is used to process the tracking mark coding feature to obtain a target tracking result; the target tracking result is a bounding box sequence; The encoding module is also used to input the position-encoded input mark into the encoder based on the visual transformer framework, realize the interaction between template features and search features, the interaction between tracking marks and search features, and the interaction between tracking marks and template features, and output the encoding results corresponding to the tracking marks and search images in the last layer output of the encoder to the decoder.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the visual target tracking method based on efficient sequence generation according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Dynamic reasoning path target tracking method based on conditional early leaving mechanism

    CN115861374A

  • Satellite video single target tracking method and device

    CN117197192A