Long-term target tracking method based on time sequence query propagation

Through timing-intensive video frame sampling and timing propagation methods of two types of propagable query vectors, the problem of insufficient robustness and reliability in long-term target tracking in the prior art is solved, and more efficient target tracking performance is achieved.

CN120014299AActive Publication Date: 2025-05-16UNIV OF SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510122327.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-16
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The existing deep learning target tracking methods have problems such as target appearance changes, background interference, occlusion and model drift in long-term tracking, resulting in insufficient robustness and reliability.

Method used

A long-term target tracking method based on timing query propagation is proposed. Time-series propagation is performed through timing-series dense video frame sampling and two types of propagated query vectors (target query vector and background query vector). The target tracking task is redescribed as a sequence propagation task of query vectors, and the timing information of the video sequence is effectively utilized.

Benefits of technology

It significantly improves the robustness and accuracy of long-term target tracking, avoids the shortcomings of complex hyperparameter settings and dynamic template update strategies, realizes the decoupling learning of targets and background features, and reduces background interference and tracking drift.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014299A_ABST
    Figure CN120014299A_ABST
Patent Text Reader

Abstract

The invention discloses a long-term target tracking method based on time sequence query propagation, and the method comprises the steps: 1, obtaining a video sequence frame comprising a multi-frame template frame set and a multi-frame search frame set, and flattening the video sequence frame to obtain template patch embedding and search patch embedding; step 2, adding position embedding; 3, designing a target query vector and a background query vector, embedding the first frame template patch obtained in the step 2, inputting the embedded search patch into a plurality of visual Transform layers together with the designed target query vector and background query vector, and carrying out attention calculation; 4, inputting the target query vector and the background query vector updated in the step 3 and template patch embedding and search patch embedding of subsequent frames into a plurality of visual Transform layers, and performing attention calculation again; and 5, embedding the finally generated search frame patch into an input prediction head to obtain a target position coordinate. According to the invention, the robustness and accuracy of long-term tracking can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of pattern recognition and computer vision, and in particular to a long-term target tracking method based on temporal query propagation. Background Art

[0002] Object tracking is a basic and important task in the field of computer vision. Its goal is to predict the state of the object in the subsequent frames, usually including the position and size, given the position or area of ​​the object in the initial frame of the video sequence. With its core role, object tracking technology has been widely used in many key fields such as video surveillance, human-computer interaction, autonomous driving, and robot navigation.

[0003] In recent years, benefiting from the rapid development of deep learning technology, deep learning-based target tracking methods have achieved significant performance improvements and become the mainstream of research. However, despite the great progress, existing deep learning target tracking methods still face many challenges in practical applications, especially when dealing with long-term tracking tasks, which restrict their robustness and reliability. These challenges are mainly reflected in the following aspects:

[0004] Target appearance changes: During the movement, the target will inevitably experience appearance changes such as rotation, scaling, deformation, lighting changes and even changes in its own posture. These changes may cause the target features extracted by the tracker to differ greatly from the initial template, thus causing tracking failure.

[0005] Background interference: In a complex and dynamic background environment, there are often objects that look similar to the target, or there are interference factors such as sudden changes in light and scene switching. These can easily cause the tracker to drift and misjudge the background as the target.

[0006] Occlusion and target disappearance: The target may be completely or partially occluded by other objects during its movement, or even temporarily leave the field of view. Traditional trackers often have difficulty in resuming tracking after the target reappears due to the lack of effective use of the target's historical information.

[0007] Model drift in long-term tracking: Although some trackers based on online learning strategies have certain adaptive capabilities, if they encounter inaccurate predictions during long-term operation, they will use incorrect tracking results as positive samples to update the model. Over time, errors are easily accumulated, leading to model drift and ultimately reducing tracking performance.

[0008] As shown in Figure 1(a), at present, most mainstream deep learning target tracking methods are trained based on single-frame image pairs, that is, using an image pair consisting of a template image containing the target and a search image of the target to be searched to learn the feature representation of the target. This training paradigm simplifies the target tracking task into a similarity matching problem between image pairs, and its core lies in learning the visual similarity between the template image and the candidate area in the search image. However, this training method based on single-frame image pairs essentially treats video frames as independent static images, ignoring the inherent time dimension of the video sequence and the rich temporal information between frames. This neglect of temporal information makes it difficult for the tracker to effectively model the dynamic changes of the target during motion, and thus it is unable to cope with the challenges of the above-mentioned target appearance changes, background interference, etc.

[0009] Although some target tracking methods attempt to introduce temporal information, their main approach is to dynamically update the template image, that is, to replace the template image of the initial frame with the image block cropped according to the prediction results in the subsequent tracking frame. The original intention of this dynamic update strategy is to hope that the model can learn target features that are more similar to the current frame. However, this strategy has inherent defects: on the one hand, when the tracking prediction results are unreliable, such as drift or occlusion of the target, the updated template image is likely to no longer contain the target to be tracked, but instead introduces background noise, thereby accelerating tracking failure. On the other hand, methods based on dynamic template updates often require the careful design and adjustment of complex hyperparameters, such as the time interval for template updates, the confidence threshold for updates, and the fusion weights. The settings of these hyperparameters have a significant impact on the tracking performance and lack universality. They need to be adjusted for different scenarios, which increases the complexity of the algorithm and the difficulty of deployment.

[0010] In addition, in actual tracking scenarios, the change trends of the target area and the background area are often significantly different. For example, the target may deform or move rapidly, while the background is relatively stable, or there may be dynamic interference in the background that is unrelated to the target. However, most existing tracking algorithms often use the same learning strategy when learning target and background features, without considering the difference in their change characteristics, resulting in the inability to effectively decouple the learning of target and background features, thereby limiting the further improvement of tracking performance. This is like using exactly the same standard to measure two things when learning to distinguish them, while ignoring their respective characteristics.

[0011] There is a wealth of temporal information such as motion patterns and appearance evolution between consecutive frames in a video sequence. Especially when dealing with challenges such as drastic changes in target appearance and complex background interference, temporal information can serve as an effective supplement to appearance representation learning, providing additional discriminative clues, thereby significantly improving the robustness and accuracy of long-term target tracking. Existing technologies either ignore temporal information or use temporal information in a complex and error-prone way (such as dynamic template updates), and lack effective modeling of the difference between target and background changes. Therefore, there is an urgent need for a target tracking method that can effectively and concisely utilize the temporal information of video sequences while avoiding the introduction of complex hyperparameters and can distinguish between learning target and background change characteristics, so as to overcome the shortcomings of existing technologies and achieve more robust and accurate long-term target tracking. Summary of the invention

[0012] In view of the problems existing in the prior art of target tracking methods, such as ignoring the temporal information of video sequences, the defects of dynamic template updating strategy, and the failure to effectively decouple target and background learning, the present invention aims to propose a novel long-term target tracking method based on temporal query propagation, aiming to:

[0013] 1) Solve the problem that existing methods do not fully utilize the temporal information of video sequences. By proposing a novel temporal dense video frame sampling method for target tracking algorithm training, the model input is expanded from traditional single-frame image pairs to continuous, temporally associated video sequences. This is intended to enable the model to effectively learn and model the contextual information and cross-frame correlation between consecutive video frames, making up for the inherent deficiencies of single-frame image pair training in temporal modeling.

[0014] 2) Construct a more robust temporal target representation and achieve decoupled learning of target and background features: A target tracking algorithm based on two types of propagable query vectors is proposed, namely, a target query vector used to accurately characterize the dynamic change characteristics of the target salient area over time, and a background query vector used to finely characterize the dynamic change characteristics of the background area over time. By explicitly introducing and propagating these two types of query vectors in the video sequence, the target tracking task is restated as a sequence propagation task of these two types of query vectors, so that the dynamic evolution of the target in the time domain and the trajectory change in the spatial domain can be accurately captured by online propagation of the query vector, and the robustness of the model to target appearance changes and background interference is improved.

[0015] By introducing the target query vector and the background query vector respectively and performing time-series propagation independently, the aim is to decouple the learning of the changing characteristics of the target and the background, overcome the suboptimal situation caused by the existing technology of not distinguishing between the two, improve the model's ability to distinguish between the target and the background, and reduce the tracking drift caused by background interference.

[0016] 3) Avoid introducing complex hyperparameters and unstable dynamic template update strategies: The method based on time series query propagation proposed in the present invention does not require the design and adjustment of complex hyperparameters, nor does it rely on dynamic template update strategies that may introduce errors, thereby providing a more stable, reliable and easy-to-deploy target tracking solution.

[0017] In summary, the present invention aims to provide a target tracking method that can effectively utilize the timing information of video sequences, improve the long-term tracking robustness and accuracy, overcome the shortcomings of the existing technology, and provide new ideas for the development of target tracking technology in the field of computer vision.

[0018] The present invention proposes a long-term target tracking method based on temporal query propagation, the core of which is to use the target query vector and the background query vector to propagate between consecutive video frames to implicitly learn and model the temporal dynamic changes of the target and the background. The specific steps of the method are as follows:

[0019] Step 1: Video sequence segmentation and patch embedding: Get a set of multi-frame template frames and multi-frame search frame sets The video sequence frames are divided and flattened into a series of image blocks. Then, each image block is flattened into a one-dimensional vector and a trainable linear projection network is used. The flattened template frame image patch and the search frame image patch are mapped to a high-dimensional latent space respectively to obtain the template patch embedding and search patch embedding ;

[0020] Step 2: Add position embedding: For the template patch embedding obtained in step 1 and search patch embedding , respectively adding learnable position embeddings and ;

[0021] Step 3: Initial Query and Transformer Interaction: Designing the Target Query Vector , background query vector The target query vector aims to learn and characterize the salient features of the target object, while the background query vector aims to learn and characterize the features of the background area. The first frame template patch obtained in step 2 is embedded into , search patch embedding Target query vector with design , background query vector Input multiple visual Transformer layers together to perform attention calculation;

[0022] Step 4: Time series query propagation and Transformer iteration: Update the target query vector after step 3 , background query vector Embedded with template patches of subsequent frames and search patch embedding Input into multiple visual Transformer layers together, and perform attention calculation again; repeat the target query vector , background query vector Update and input the visual Transformer layer, that is, continuously use the Transformer layer to perform attention calculations between consecutive video frames;

[0023] Step 5: Embed the search frame patch generated in step 4 after multi-layer Transformer iterations Input the prediction head to obtain the target position coordinates.

[0024] Beneficial effects of the technical solution of the present invention:

[0025] 1) This paper proposes a new temporal dense video sequence sampling method for target tracking algorithm training, which expands the model input from single-frame image pairs to continuous video sequences, thereby modeling the contextual information and cross-frame correlation of continuous video frames.

[0026] The present invention innovatively proposes a temporally dense video sequence sampling method for training target tracking algorithms, changing the previous method's mode of relying only on single-frame image pairs for training. As shown in Figure 1(a), previous target tracking algorithms perform temporally sparse sampling training, that is, only sampling isolated single-frame image pairs, and simplifying the target tracking task to visual similarity matching between template images and search image pairs. Although this method can learn a certain visual representation ability, its inherent temporal sparsity is destined to have fundamental defects in utilizing the rich temporal information in video sequences, so that the model can only focus on the appearance similarity of the target, but cannot fully explore and utilize the correlation between frames. In contrast, the temporally dense video frame sampling method proposed in the present invention, as shown in Figure 1(b), samples continuous template frames and search frame images as model input, so that the model can be exposed to richer temporal context information. Under this new input paradigm, the model can effectively learn and model rich temporal information such as motion patterns and appearance evolution contained in the previous and next frames of the video sequence, thus overcoming the problem of poor long-term tracking robustness of traditional methods due to insufficient utilization of temporal information, and significantly improving the stability and accuracy of target tracking.

[0027] 2) This paper proposes a target tracking algorithm based on two types of query vectors for temporal propagation, namely, a target query vector used to characterize the temporal variation of the target's salient region and a background query vector used to characterize the temporal variation of the background region. Target tracking is reformulated as a sequential propagation task of these two types of query vectors, capturing the temporal and spatial trajectory relationship of the target by online propagation of the query vector. This effectively decouples the temporal information learning of the target and the background, and improves the ability to resist background interference.

[0028] The present invention innovatively proposes a target tracking algorithm based on temporal propagation of two types of query vectors, namely, a target query vector used to characterize the temporal variation characteristics of the target salient region, and a background query vector used to characterize the temporal variation characteristics of the background region. Different from the previous view that target tracking is a simple visual similarity matching task, the present invention reformulates target tracking as a propagation task of these two types of query vectors in a video frame sequence, and captures the temporal and spatial trajectory relationship of the target through the online propagation of the query vector. Specifically, as shown in FIG1(b), based on continuous video frame input, the present invention models target tracking as an autoregressive propagation process of the query vector in the video frame. The video frame input provides a basis for visual representation learning, while the propagation of the query vector focuses on the learning of temporal representation. The two work together to achieve better tracking results. More importantly, in actual tracking scenarios, there are often significant differences in the change trends of the target region and the background region, that is, the rate and pattern of change between the two are usually different. However, existing tracking algorithms often take an indiscriminate approach to the learning of the target and background, resulting in the inability to effectively decouple the learning of the features of the two, thus creating a bottleneck in performance optimization. To address this problem, the present invention further designs query vectors into two categories, namely target query vectors and background query vectors, and makes these two types of query vectors independently perform time series autoregressive learning. This design cleverly achieves the decoupled learning of target and background time series information, avoids mutual interference between the two in the feature learning process, and thus significantly improves the model's ability to distinguish between target and background, effectively suppresses background interference, and reduces the occurrence of tracking drift. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] FIG. 1( a ) is a schematic flow chart of a target tracking method in the prior art that uses only a single frame image for training;

[0030] FIG1( b ) is a schematic flow chart of a long-term target tracking method based on temporal query propagation using multi-frame video sequence training proposed in the present invention;

[0031] Figure 2 A schematic flow chart of attention calculation for introducing temporal query updates according to the present invention is shown. DETAILED DESCRIPTION

[0032] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms, and the present disclosure should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0033] In order to fully understand the temporal dense video sequence sampling method proposed in the present invention, it is necessary to first review the temporal sparse image pair sampling method.

[0034] As shown in Figure 1(a), the traditional temporal sparse image pair sampling method independently samples a template frame image from the video sequence. and a search frame image , in , are the height and width of the template frame image respectively, , is the height and width of the search frame image. Then, the pair of images and Enter the target tracking algorithm Processing to predict the bounding box coordinates of the target in the current search frame ,Right now:

[0035] ,

[0036] if For a twin tracking network, it needs to go through three stages: visual feature extraction, feature interaction, and bounding box prediction. For a single-stream tracking network, it only contains a visual Transformer backbone network and a prediction head network, where the backbone network performs visual feature extraction and interaction steps simultaneously.

[0037] against In the case of a single-stream tracking network, the visual Transformer backbone network receives a series of image patch embeddings as input. Specifically, the reference frame image and search frame image Divide into The image patch is then passed through a trainable linear projection network To generate a template image tag sequence and search image tag sequences .in is the label dimension, , These image tags are then concatenated and fed into In each visual Transformer layer, feature extraction and interaction are performed simultaneously. It includes a multi-head attention and a multi-layer perceptron. The forward propagation process of the visual Transformer layer is expressed as:

[0038] ,

[0039] in, Indicated by The layer vision Transformer layer generates a labeled sequence of template frame-search frame image pairs.

[0040] Although the above-mentioned temporal sparse image pair sampling method can construct a relatively concise target tracking framework, its fundamental flaw is that the tracking model only focuses on the visual similarity of the target within a single frame and lacks the ability to establish temporal associations between previous and subsequent frames. This neglect of temporal information greatly hinders the robustness of target tracking in long video sequences. Since the model cannot effectively utilize the target's temporal motion trajectory and appearance change rules, it is prone to tracking failure or drift when facing complex situations such as target occlusion, rapid motion, and appearance deformation.

[0041] Sampling of time-dense video sequences

[0042] Different from the traditional temporal sparse image pair sampling method, the present invention expands the input of the target tracking algorithm from temporally isolated image pairs to temporally continuous video sequences, so that temporal information can be directly modeled. In addition, the present invention also innovatively introduces two types of specially designed query vectors: target query vector and the background query vector Based on the time-intensive video sequence input and the dual query vector mechanism, the tracking process of the present invention can be expressed as:

[0043] ,

[0044] in, Indicates the length Template frame sequence, Indicates the length The search frame sequence, is the predicted target box coordinate of the current search frame, It is a single-stream tracking network.

[0045] When sampling a time-intensive video sequence, we do not simply select adjacent frames, but use a larger sampling interval than traditional methods. Within this larger sampling interval, we randomly sample multiple video frames to construct a template frame sequence of any length. and search frame sequence The core reason for adopting a larger sampling interval is that we hope that the model can learn the temporal relationship contained in the video sequence over a longer time scale. By sampling video frames with temporal sequence relationships over a longer time span, it can effectively help the tracking model learn stronger and more robust temporal correlation information, thereby better coping with challenges such as target appearance changes and occlusions.

[0046] Target query and background query design

[0047] As pointed out in the background technology section, in actual tracking scenarios, the change trends of the target area and the background area are often significantly different. The target may deform or move rapidly, while the background is relatively stable, or there may be dynamic interference unrelated to the target. In order to solve the problem of not distinguishing between target and background learning in the prior art and realize the decoupled learning of target and background time series information, the present invention innovatively designs the target query vector and the background query vector , which are used to model the temporal dynamic changes of the target and background respectively.

[0048] Target query vector is randomly initialized as a one-dimensional vector. In the subsequent tracking process, the target query vector will propagate in the continuous video sequence and interact with the image features of each frame. Through this propagation and interaction mechanism, the target query vector can gradually aggregate and compress the key trajectory information and appearance change information of the target in each frame, and finally form an effective representation of the dynamic changes of the target time series.

[0049] Different from the target query vector, the background query vector The design of fully considers the characteristics of the background area. A multi-scale convolutional network is used to process the label sequence of the background area. Specifically, the multi-scale convolutional network takes the label sequence of the background area as input, and gradually reduces the resolution of the feature map through layer-by-layer convolution and pooling operations, and finally learns a two-dimensional vector with a sampling step of 16. Then, the two-dimensional vector is flattened into a one-dimensional vector as the final background query vector. The reason why the multi-scale convolutional network can be used to downsample the label sequence of the background area is based on the following two considerations: First, the redundancy of background information, that is, the background area of ​​the image usually contains a lot of redundant information, such as repeated textures, similar colors, etc. Therefore, fewer labels can be used to effectively represent the information of the background area without losing too much discrimination ability. The second is to reduce computational complexity: Since the attention computational complexity of the visual Transformer is quadratically proportional to the number of input labels, reducing the number of labels in the background area can significantly reduce the computational complexity of the model and improve tracking speed and efficiency.

[0050] By designing the target query vector and the background query vector separately and adopting different modeling strategies, the present invention can effectively decouple the temporal information learning of the target and the background. The target query vector focuses on aggregating the trajectory and appearance change information of the target, while the background query vector reduces the amount of calculation as much as possible while ensuring the ability to represent the background information. This differentiated design is one of the key factors that enable the present invention to effectively deal with complex background interference.

[0051] Introducing attention calculation for temporal query updates

[0052] The present invention uses the commonly used two-dimensional visual Transformer network as the basic architecture of the backbone network. However, in order to adapt to the introduction of temporal query vectors and realize the autoregressive update of query vectors between frames, the key innovation of the present invention is to expand the original spatial two-dimensional attention calculation method and design a new attention calculation mechanism that introduces temporal query updates.

[0053] Traditional Transformer-based object tracking methods usually use a single-frame image pair as input, and the attention calculation process of its backbone network can be summarized as: , in, An image tag representing the template frame, The image tag representing the search frame, represents the traditional two-dimensional attention calculation operation, The visual feature tags are obtained by attention calculation. This attention calculation method can only model the visual similarity within a single frame image, and cannot capture the temporal correlation information between the previous and next frames in the video sequence.

[0054] like Figure 2 As shown in FIG. 1 , in order to solve the above problem, the present invention expands the input into a multi-frame video sequence and propagates the temporal query vector between the previous and next frames in an autoregressive manner, thereby implicitly modeling the temporal information. Specifically, for the first Frame, the attention calculation method of the present invention can be expressed as:

[0055] ,

[0056] ,

[0057] in, For the The template frame tag of the frame, For the Search frame marker for frames, is a trainable linear projection layer that maps image tags to query, key, and value spaces. Respectively Visual query, key, and value tags for frames; is the visual feature tag obtained by attention calculation, They are the introduced target query vector and background query vector respectively.

[0058] The key point of the present invention is to introduce a target query vector in each video frame and a background query vector , which is used to compress and store target information and background information in time-intensively sampled video sequences. In the current frame, the attention mechanism is used to map the visual features. and the time series query vector More importantly, the updated target query vector is autoregressively transformed into and the background query vector From The frame is propagated to frame and compare it with the Visual query, key, value tagging of frames Together as the The input of frame attention calculation. The specific process is as follows Figure 2 shown.

[0059] In this target tracking paradigm based on time-series query update, the present invention can convert the target query vector and the background query vector As a hint for the next frame’s attention calculation, the past information is used to guide future reasoning, which is consistent with the intuition of object tracking in continuous video sequences. Moreover, the target query vector and the background query vector The propagation of is equivalent to implicitly propagating the positioning and appearance information of the target and the background, which can form a complementary enhancement with the image visual features provided by the template frame and the search frame, thereby effectively improving the performance of target tracking.

[0060] The long-term target tracking method based on time series query propagation of the present invention comprises the following steps:

[0061] Step 1: Video sequence segmentation and patch embedding: Get a set of multi-frame template frames and multi-frame search frame sets The video sequence frames are divided and flattened into a series of image blocks. Then, each image block is flattened into a one-dimensional vector. This operation can convert image information into sequence data that is easier to process by models such as Transformer. Use a trainable linear projection network The flattened template frame image patches and search frame image patches are respectively mapped to a high-dimensional latent space to obtain the plate patch embedding and search patch embedding The role of the linear projection network is to transform the original pixel information into a more discriminative feature representation.

[0062] Trainable Linear Projection Network It can be a linear function, and the dimension of the high-dimensional latent space can be 384.

[0063] Step 2: Add position embedding: In order to make the Transformer model aware of the position information of the patch in the original image, the template patch embedding obtained in step 1 is and search patch embedding , respectively adding learnable position embeddings and Position embeddings encode the spatial position of patches in an image, allowing the model to consider spatial relationships when performing attention calculations.

[0064] Adding position embedding to template patch embedding and search patch embedding is to add template patch embedding and search patch embedding to position embedding element by element. Position embedding is a learnable position embedding vector. The way to obtain position embedding is to directly define a tensor, and then the deep learning network autonomously learns to obtain the position encoding information.

[0065] Step 3: Initial query and Transformer interaction: Design two learnable query vectors, namely the target query vector , background query vector The target query vector aims to learn and characterize the salient features of the target object, while the background query vector aims to learn and characterize the features of the background area. The template patch of the first frame obtained in step 2 is embedded into , search patch embedding Target query vector with design , background query vector The attention calculation is performed by inputting multiple visual Transformer layers together. The attention calculation method includes: in multiple visual Transformer layers, the target query vector and the background query vector will be respectively and Each patch embedding in performs cross-attention calculation to extract feature information related to the target and background. Internal and The internal patch embeddings are also self-attentioned to learn contextual relationships within the frame.

[0066] The number of multiple vision Transformer layers can be 24 layers.

[0067] Step 4: Time series query propagation and Transformer iteration: The target query vector calculated by the visual Transformer layer in step 3 is , background query vector As the updated query vector. , background query vector Embedded with template patches of subsequent frames and search patch embedding The query vector is input into multiple visual Transformer layers together, and the attention calculation is performed again. Repeated query update and input into the visual Transformer layer, that is, between consecutive video frames, the Transformer layer is continuously used to perform attention calculation. The above calculation process is applicable to the template patch embedding and search patch embedding of frames other than the first and second frames. In this process, the target query vector and the background query vector It will be continuously updated and propagated in the process of interacting with the image information of the previous and next frames, thereby implicitly learning and modeling the temporal dynamic change characteristics of the target and background.

[0068] Step 5: Embed the search frame patch generated in step 4 after multi-layer Transformer iterations Input the prediction head to obtain the target position coordinates. The prediction head consists of a fully connected layer that is used to embed the patch into the target position coordinate information, including the center point coordinates and width and height of the target.

[0069] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0070] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is primarily selected for readability and instructional purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention.

Claims

1. A long-term target tracking method based on temporal query propagation, characterized in that: include: Step 1: Video sequence segmentation and patch embedding: Get a set of multi-frame template frames and multi-frame search frame sets The video sequence frames are divided and flattened into a series of image blocks. Then, each image block is flattened into a one-dimensional vector and a trainable linear projection network is used. The flattened template frame image patch and the search frame image patch are mapped to a high-dimensional latent space respectively to obtain the template patch embedding and search patch embedding ; Step 2: Add position embedding: For the template patch embedding obtained in step 1 and search patch embedding , respectively adding learnable position embeddings and ; Step 3: Initial Query and Transformer Interaction: Designing the Target Query Vector , background query vector The target query vector aims to learn and characterize the salient features of the target object, while the background query vector aims to learn and characterize the features of the background area. The first frame template patch obtained in step 2 is embedded into , search patch embedding Target query vector with design , background query vector Input multiple visual Transformer layers together to perform attention calculation; Step 4: Time series query propagation and Transformer iteration: Update the target query vector after step 3 , background query vector Embedded with template patches of subsequent frames and search patch embedding Input into multiple visual Transformer layers together, and perform attention calculation again; repeat the target query vector , background query vector Update and input the visual Transformer layer, that is, continuously use the Transformer layer to perform attention calculations between consecutive video frames; Step 5: Embed the search frame patch generated in step 4 after multi-layer Transformer iterations Input the prediction head to obtain the target position coordinates.

2. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: Target query vector is randomly initialized to a one-dimensional vector, the background query vector The acquisition method includes: using a multi-scale convolutional network to take the labeled sequence of the background area as input, gradually reducing the resolution of the feature map through layer-by-layer convolution and pooling operations, and finally learning a two-dimensional vector with a sampling step of 16. Then, the two-dimensional vector is flattened into a one-dimensional vector as the final background query vector.

3. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: Trainable Linear Projection Network is a linear function.

4. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: The dimension of the high-dimensional latent space is 384.

5. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: Adding position embedding to template patch embedding and search patch embedding is to perform element-wise addition of template patch embedding, search patch embedding and position embedding respectively.

6. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: Position embedding is a learnable position embedding vector. The way to obtain position embedding is to directly define a tensor, and then the deep learning network autonomously learns to obtain the position encoding information.

7. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: In multiple visual Transformer layers, the target query vector and the background query vector will be respectively and Each patch embedding in performs cross-attention calculation to extract feature information related to the target and background; at the same time, Internal and The internal patch embeddings also undergo self-attention computation to learn contextual relationships within the frame.

8. The long-term target tracking method based on time series query propagation according to claim 1 is characterized in that: The number of multiple vision Transformer layers is 24.

9. The long-term target tracking method based on time series query propagation according to claim 1, characterized in that: The prediction head consists of a fully connected layer to map the patch embedding to the location coordinate information of the target.

10. The long-term target tracking method based on time series query propagation according to claim 1, characterized in that: The target's location coordinate information includes the target's center point coordinates and width and height.

Citation Information

Patent Citations

  • Video multi-target tracking method based on multi-scale channel feature aggregation

    CN117173217A

  • Single-target tracking network based on combination of Vmamba and Transform

    CN118537369A

  • Visual target tracking method based on mask contrast learning pre-training

    CN118674749A

  • High-performance visual object tracking for embedded vision systems

    US20190304105A1