Single-target tracking model based on single-flow framework and aggregating time sequence information by using knowledge tokens
By introducing historical valid frames and timing knowledge tokens into the single-stream tracking model and aggregating timing information, the problem of the single-stream tracking model ignoring spatiotemporal information is solved, and the tracking performance is significantly improved and the best level of current technology is achieved.
Patent Information
- Application Number
- CN202510041720.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The existing single-stream tracking model ignores the spatiotemporal information in continuous video frames, resulting in insufficient tracking accuracy and success rate.
Design a single-objective tracking model based on a single-stream framework, and achieve online update and dissemination of target appearance and position information by introducing historical valid frames and timing knowledge tokens.
By aggregating timing information, the model has significantly improved its tracking performance, and its performance has reached the best level of current technology, and the effectiveness of timing information has been verified through ablation experiments.
Smart Images

Figure CN119963607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision target tracking, and in particular to a single target tracking model based on a single stream framework and using knowledge tokens to aggregate temporal information. Background Art
[0002] In recent years, with the improvement of science and technology and the release of scientific research results, target tracking technology has developed and iterated rapidly, and has been widely used in video surveillance, mobile robots, autonomous driving, scene understanding, human-computer interaction and other fields. As an important research direction in the field of computer vision, target tracking technology calibrates the target to be tracked through a bounding box, and visually tracks the target frame by frame in subsequent video frames.
[0003] In order to improve the accuracy and robustness of tracking, scholars have studied and designed various theoretical models, among which the models in the direction of deep learning have been studied more deeply due to their outstanding performance. In the past two years, the single-stream tracking framework with a pure Transformer model as the network backbone has become a new research direction. So far, there have been many theoretical and application results in the research of single-stream tracking models, but they all ignore the spatiotemporal information of the target appearance and position contained in continuous video frames. The temporal context information of continuous frames has been used in other tracking models, and it has been proven to have an excellent improvement effect on tracking accuracy and success rate. Therefore, whether the timing information contained in continuous frames can also perform well in the single-stream framework, and whether the already excellent single-stream tracking model can be further improved in tracking performance, has become a topic worthy of study. Based on this, the present invention proposes a single target tracking model based on a single-stream framework and using knowledge tokens to aggregate timing information. Summary of the invention
[0004] The purpose of the present invention is to provide a single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information, so as to solve the problems raised in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solution: a single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information, comprising the following steps:
[0006] Step S1: redesign the model input of the single-stream framework based on the existing single-stream framework model, replace the hybrid template in the improved single-stream framework with the historical valid frame, and realize the online update of target tracking;
[0007] Step S2: Design a training strategy, reasoning strategy and update method for the historical valid frames in step S1;
[0008] Step S3: Design an online updated temporal knowledge token, which is spliced with other tokens to enter and exit the model backbone network, and aggregate the appearance information of the target in the model;
[0009] Step S4: Design training strategies and reasoning strategies for temporal knowledge tokens, as well as their initialization methods;
[0010] Step S5: The designs of step S1 and step S3 are combined to form a backbone network model, which is fully trained, verified and tested on an authoritative data set to obtain a single target tracking model that meets the requirements.
[0011] Preferably: the model input proposed in step S1 includes an inherent template frame, a historical valid frame, a search frame and a timing token, and the model output is a historical valid frame, a search frame and a timing token; the core of the model is to calculate the mixed attention of the online updated valid frame and the current search frame.
[0012] Preferably: the valid frame updating method proposed in step S2 is a simple updating algorithm, which specifically comprises the following steps: firstly judging the maximum score of the tracking result, and updating is possible only if it is higher than a set threshold; if the score is greater than the score of the valid frame, directly updating the valid frame to the result of the current search frame; if the score is less than the score of the valid frame, calculating the distance between the valid frame and the search frame, setting a bonus point based on the conditions, but not exceeding a set bonus point threshold; if the current score is higher than the valid frame score after the bonus point is added, updating is performed, otherwise not updating.
[0013] Preferably: the temporal knowledge token proposed in step S3 is a tensor of shape (B, 1, C), where B is the batch size and C is the feature dimension.
[0014] Preferably: the knowledge token initialization method proposed in step S4 is to first initialize a neutral token with all values 0.5 Expanding in the batch dimension, where dim is the feature dimension, the attention calculation is performed on the token t0 and the template T, as follows:
[0015]
[0016] in Indicates that the neutral token has learned the characteristics of the template, then the initialized knowledge token
[0017] Preferably, the attention calculation process of the single target tracking model backbone network based on the single stream framework and the use of knowledge tokens to aggregate timing information proposed in step S5 is described as follows:
[0018] Q=[q V ,q S ,q K ],
[0019] K=[k T , k V , kS , k K ],
[0020] V=[v T , v V ,v S ,v K ]
[0021]
[0022] Among them A S =M S,T v T +M S,V v V +M S,S v S +M S,K v K Represents the attention result of the search frame, that is, the feature fusion result of the search frame, template frame and valid frame.
[0023] Preferably: the training strategy proposed in step S5 divides the model training into three stages: the first stage uses templates and search frames to perform 300 rounds of image matching ability training; the second stage uses template frames, valid frames and search frames to perform 100 rounds of valid frame and search frame feature fusion training; the third stage uses template frames, valid frames, search frames and timing tokens to perform 200 rounds of timing information aggregation training; the initial learning rate of the second and third stages is set to 10% of the learning rate of the first stage.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] The model designed by the present invention aggregates temporal information in two ways: 1) using historical valid frames and search frames to perform the core image matching process of the model; 2) introducing a temporal knowledge token to propagate the implicit appearance and position change information of the target over time during the tracking process of consecutive frames. The model proposed by the present invention has been experimentally shown to have the best performance at the current technical level, and the effectiveness of the two types of temporal information has been proved through ablation experiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the overall structure of the model of the present invention;
[0027] Figure 2 This is a schematic diagram of online updating of valid frames;
[0028] Figure 3 Initialize the time series knowledge token and transfer the information logic diagram;
[0029] Figure 4 Schematic diagram of the backbone network attention calculation method;
[0030] Figure 5 This is the result of the ablation experiment;
[0031] Figure 6 This is a comparison chart of SOTA model performance. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] Example
[0034] See also Figure 1-Figure 6 ,The single target tracking model based on a single stream framework and using ,knowledge tokens to aggregate temporal information in the figure includes the following steps,Step S1: redesign the model input of the single stream framework ,based on the existing single stream framework model, replace the hybrid template in the ,improved single stream framework with historical valid frames to achieve online ,update of target tracking;
[0035] Step S2: Design a training strategy, reasoning strategy and update method for the historical valid frames in step S1;
[0036] Step S3: Design an online updated temporal knowledge token, which is spliced with other tokens to enter and exit the model backbone network, and aggregate the appearance information of the target in the model;
[0037] Step S4: Design training strategies and reasoning strategies for temporal knowledge tokens, as well as their initialization methods;
[0038] Step S5: The designs of step S1 and step S3 are combined to form a backbone network model, which is fully trained, verified and tested on an authoritative data set to obtain a single target tracking model that meets the requirements.
[0039] Furthermore, step S1 of the present invention refers to human vision and simplifies the effective target tracking mechanism to locking the current target through the initial target impression and the most recent target appearance. Therefore, the historical valid frame is introduced as an effective reference for the most recent target appearance, and the initial target template is retained to participate in feature fusion. The existing improved single-stream framework model is analyzed. Its core is still the image matching ability of the training template and the search frame, and the training set sampling method is still the image pair of the template frame and the search frame. Therefore, replacing the original reference template of the model with the online updated valid frame does not affect the image matching ability trained by the image pair. Finally, a new target tracking model based on the single-stream framework is designed. The image matching of the online updated valid frame and the search frame is the core task of the model, and the initial template is used to effectively guide the image matching process through mixed attention.
[0040] The core part of the model lies in the design of its attention calculation method, such as Figure 4 As shown, the model implementation principle is as follows:
[0041] The feature extraction method of the initial template's self-attention can be expressed as follows:
[0042]
[0043] Among them A T represents the self-attention calculation result of the template, Softmax is the normalization function, d is the dimension of the feature, Q T , K T and V T They represent the query vector, key vector, and value vector of the template in the attention calculation respectively.
[0044] The hybrid attention module uses the keys and values of the template frame, valid frame, and search frame to perform attention calculations on the valid frame and search frame. The calculation process can be described as follows:
[0045] Q=[q V ,q S ],
[0046] K=[k T , k V , k S ],
[0047] V=[v T , v V , v S ]
[0048]
[0049] Among them A S =M S,T v T +M S,Vv V +M S,S v S Represents the attention result of the search frame, that is, the feature fusion result of the search frame, template frame and valid frame.
[0050] In order to effectively use the model in step S2, a matching training strategy and inference strategy must be designed, and a method for online updating of valid frames must also be designed.
[0051] As for the training strategy, in the first stage of model training, the original sampling and model framework are kept unchanged to obtain a stable image matching ability of the model. In the second stage of training, three image frames with a certain frame interval are sampled as inherent template frames, valid frames and search frames respectively. The interval between the first two frames is larger, and the interval between the last two frames is smaller. After image processing, the corresponding positions of the model are input for training.
[0052] For the inference strategy, both the template frame and the valid frame are initialized to the first frame of the video. In the subsequent frame-by-frame tracking process, the template frame is kept unchanged, and the valid frame is updated after the tracking is completed.
[0053] For the effective frame update method, the update process is as follows: Figure 2 As shown, a simple algorithm is designed. First, the maximum score of the tracking result is determined. It can only be updated if the score is greater than 0.92. If the score is greater than the score of the valid frame, the valid frame is directly updated to the result of the current search frame. If the score is less than the score of the valid frame, the distance between the valid frame and the search frame is calculated, and the score is set at 0.0001 points per frame, but not more than 0.03. If the current score is higher than the valid frame score after the score is added, it is updated, otherwise it is not updated.
[0054] In step S3, after studying the tokens related to the Transformer model, a temporal knowledge token is designed. The image frame will be segmented into patches and inserted into the position code to form the token form of the model input. The temporal knowledge token of the present invention will be spliced with the token of the image frame to participate in the mixed attention calculation. After the calculation is completed, the token is segmented separately and then participates in the image matching of the next frame in the form of a token. By analogy, the attention information is transmitted frame by frame, that is, the historical and current appearance information of the target is transmitted. Therefore, the attention calculation process in step S1 is updated, which is expressed as follows:
[0055] Q=[q V ,q S q K ],
[0056] K=[k T , k V , k S , k K ],
[0057] V=[v T , v V , v S , v K ]
[0058]
[0059] Among them A S =M S,T v T +M S,V v V +M S,S v S +M S,K v K The token representing the temporal knowledge provides the temporal context information as a reference for the feature fusion process of the search frame.
[0060] In order to effectively use the token, step S4 requires designing a matching training strategy and inference strategy, including an initialization method.
[0061] For the training strategy, in the third stage of model training, the initialization method for the temporal knowledge token is: first initialize a neutral token with all values 0.5 Expanding in the batch dimension, where dim is the feature dimension, the attention calculation is performed on the token t0 and the template T, as follows:
[0062]
[0063] in Indicates that the neutral token has learned the characteristics of the template, then the initialized knowledge token
[0064] The present invention samples a new frame before the search frame, and performs attention calculation again on the frame image together with the template frame and the knowledge token, so that it has the attention information of the previous frame. Then, the token is used together with the template, the valid frame and the search frame as the input of the model for model training.
[0065] For the inference strategy, the token transfer process is as follows Figure 3 As shown, in the process of initializing the inherent template in the first frame of the video, the knowledge token is initialized at the same time, which is similar to the training strategy, but here only the token results separated after the attention calculation of the neutral token and the inherent template are needed. In the subsequent frame-by-frame tracking process, the knowledge token separated from the output of the previous frame model is used as the input of the next frame model to realize the temporal information transmission.
[0066] Step S5 integrates the designs of step S1 and step S3. The overall structure diagram of the single target tracking model based on the single stream framework and the use of knowledge tokens to aggregate timing information according to the present invention is as follows: Figure 1 as shown.
[0067] The model training is divided into three stages: the first stage uses templates and search frames for 300 rounds of image matching ability training; the second stage uses template frames, valid frames and search frames for 100 rounds of valid frame and search frame feature fusion training; the third stage uses template frames, valid frames, search frames and time series tokens for 200 rounds of temporal information aggregation training; the initial learning rate of the second and third stages is set to 10% of the learning rate of the first stage, and the model is verified every 20 rounds. The comparison results of the verification set model loss and the model intersection over union (IoU) in the last 100 rounds of the three stages are shown as follows: Figure 5 As shown in the figure, the models obtained in the three stages are tracked and tested on the LaSOT dataset, and the performance comparison results are as follows Figure 6 As shown, according to Figure 5 and Figure 6 It can be concluded that the effective frame technology and timing knowledge token technology of the present invention are practical and effective, and the tracking performance of the single target tracking model based on the single stream framework and the use of knowledge tokens to aggregate timing information is at the forefront of the technology.
[0068] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0069] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate temporal information, characterized in that: The steps include: Step S1: redesign the model input of the single-stream framework based on the existing single-stream framework model, replace the hybrid template in the improved single-stream framework with the historical valid frame, and realize the online update of target tracking; Step S2: Design a training strategy, reasoning strategy and update method for the historical valid frames in step S1; Step S3: Design an online updated temporal knowledge token, which is spliced with other tokens to enter and exit the model backbone network, and aggregate the appearance information of the target in the model; Step S4: Design training strategies and reasoning strategies for temporal knowledge tokens, as well as their initialization methods; Step S5: The designs of step S1 and step S3 are combined to form a backbone network model, which is fully trained, verified and tested on an authoritative data set to obtain a single target tracking model that meets the requirements.
2. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information according to claim 1, characterized in that: The model input proposed in step S1 includes an inherent template frame, a historical valid frame, a search frame and a timing token, and the model output is a historical valid frame, a search frame and a timing token; The core of the model is to calculate the mixed attention of the online updated valid frame and the current search frame.
3. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information according to claim 2, characterized in that: The valid frame update method proposed in step S2 is a simple update algorithm, which specifically first determines the maximum score of the tracking result, and it is possible to update if it is higher than the set threshold; if the score is greater than the score of the valid frame, the valid frame is directly updated to the result of the current search frame; if the score is less than the valid frame score, the distance between the valid frame and the search frame is calculated, and the bonus is set according to the conditions, but it does not exceed the set bonus threshold. If the current score is higher than the valid frame score after the bonus, it is updated, otherwise it is not updated.
4. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information according to claim 3, characterized in that: The temporal knowledge token proposed in step S3 is a tensor of shape (B, 1, C), where B is the batch size and C is the feature dimension.
5. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information according to claim 4, characterized in that: The knowledge token initialization method proposed in step S4 is to first initialize a neutral token with all values 0.5 Expanding in the batch dimension, where dim is the feature dimension, the attention calculation is performed on the token t0 and the template T, as follows: in Indicates that the neutral token has learned the characteristics of the template, then the initialized knowledge token 6. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information according to claim 5, characterized in that: The single target tracking model backbone network proposed in step S5 based on a single stream framework and using knowledge tokens to aggregate timing information, and its attention calculation process is described as follows: Q=[q V ,q S ,q K ], K=[k T ,k V ,k S ,k K ], V=[v T ,v V ,v S ,v K ] Among them A S =M S,T v T +M S,V v V +M S,S v S +M S,K v K Represents the attention result of the search frame, that is, the feature fusion result of the search frame, template frame and valid frame.
7. A single target tracking model based on a single stream framework and using knowledge tokens to aggregate timing information according to claim 6, characterized in that: The training strategy proposed in step S5 divides the model training into three stages: the first stage uses the template and the search frame to perform 300 rounds of image matching ability training; the second stage uses the template frame, the valid frame and the search frame to perform 100 rounds of valid frame and search frame feature fusion training; the third stage uses the template frame, the valid frame, the search frame and the timing token to perform 200 rounds of timing information aggregation training; The initial learning rate of the second and third stages is set to 10% of the learning rate of the first stage.
Citation Information
Cited By
Multi-mode visual single-target tracking method and device based on memory prompt, equipment and medium
CN121304734A
Single-target tracking method and system based on spatial-temporal characteristic memory online updating
CN121883537A