Uniflow lightweight target tracking method based on knowledge distillation

By introducing knowledge distillation technology into the single-stream target tracking algorithm, using the weight initialization of the DeiT model and optimizing the Token similarity distillation algorithm, the problem of excessive calculation and parameters of the single-stream target tracking algorithm on edge computing devices is solved, and high-performance lightweight target tracking is achieved.

CN119963928AActive Publication Date: 2025-05-09DALIAN UNIV OF TECH

Patent Information

Application Number
CN202510435465.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing single-stream target tracking algorithm is difficult to deploy on edge computing devices, mainly due to the large amount of computing and parameters, making it difficult to achieve a balance between speed and accuracy.

Method used

Using a single-flow lightweight target tracking method based on knowledge distillation, the single-flow target tracking model is initialized by designing an interactive distillation module using the weight of the DeiT image classification model, and the model performance is optimized through the Token similarity distillation algorithm.

Benefits of technology

While reducing the calculation amount by about 46% and the parameters by about 15%, the performance similar to the original model is maintained, which significantly reduces the computing power requirement of the lightweight tracking model and achieves a better balance of speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963928A_ABST
    Figure CN119963928A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of machine learning, computer vision and target tracking, and provides a uniflow lightweight target tracking method based on knowledge distillation. According to the invention, a light-weight single-flow target tracking network model based on a pure Transform backbone network is realized, and the light-weight tracking model is distilled by using an OSTrack algorithm as a teacher model. According to the method, the similarity matrix of the template and the feature region is provided, the similarity matrix can efficiently represent interaction information of the template and the search region in a single-flow tracking algorithm, the information serves as a supervision signal in knowledge distillation, and the performance of a lightweight single-flow target tracking model can be improved. According to the invention, the performance of a lightweight target tracking network is improved, and a target tracking model with stronger performance is provided for an edge computing scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of machine learning, computer vision and target tracking, and relates to knowledge distillation, a single-stream target tracking algorithm, a single-stream target tracking algorithm OSTrack and an image classification algorithm DeiT; specifically, it is a single-stream lightweight target tracking method based on knowledge distillation. Background Art

[0002] The target tracking task aims to accurately locate the position, scale and other information of a target in the subsequent frames given an initial frame of interest in a video sequence. Target tracking is one of the basic tasks of computer vision and is the application basis for many computer vision-related problems such as action recognition, event detection, behavior understanding and video object detection. In addition, visual tracking has extremely broad practical application prospects in many fields such as national defense, medical treatment, transportation, entertainment and photography.

[0003] The target tracking method cuts the target information input by the user as a template in the initial frame, and cuts the search area according to the target position and size of the previous frame in the subsequent frame, thereby converting the tracking task into a matching task between the template and the search area. With the development of deep learning technology, target tracking has gradually transitioned from dual-stream target tracking methods to single-stream target tracking methods. Dual-stream target tracking methods (such as SiamFC, SiamRPN++, TransT) usually use a shared parameter backbone network for feature extraction, and then use a cross-correlation module to match the template and the search area. This type of method separates the feature extraction and feature interaction modules, and the model tracking effect is relatively poor. Single-stream tracking methods (such as OSTrack, SeqTrack, ARTrack, etc.) use the Transformer structure to perform feature extraction and feature interaction at the same time. Although the tracking performance has been improved, the Transformer structure itself has a large amount of computation and parameters, which makes it difficult to deploy such algorithms on edge devices.

[0004] It should be pointed out that all the algorithms mentioned above are designed around GPU computing resources. In actual reasoning applications, a large amount of computing resources are often required as support. However, in devices such as autonomous driving, drones, and intelligent tracking gimbals, computing power is greatly limited. The HiT algorithm specifically designs a single-stream tracking framework with a hybrid architecture of convolution and Transformer for edge computing devices. This method enhances the performance of the tracking method by fusing visual features at different levels. Unfortunately, this method has a large number of parameters and a relatively complex model structure, which makes it difficult to deploy and optimize.

[0005] Therefore, how to design a high-performance tracking algorithm that is suitable for edge computing devices and achieve a balance between speed and accuracy is one of the challenges that need to be urgently solved in current tracking tasks. Summary of the invention

[0006] The present invention aims to provide a single-stream lightweight target tracking method based on knowledge distillation, and improves the performance of the existing Transformer-based lightweight tracking algorithm by designing an interactive distillation module. The algorithm proposed in the present invention can achieve performance similar to that of other models while reducing the amount of calculation to about 46% of other models and the amount of parameters to about 15% of other models, greatly reducing the demand for computing power for lightweight tracking models.

[0007] The technical solution of the present invention:

[0008] A single-stream lightweight target tracking method based on knowledge distillation, the steps are as follows:

[0009] Step 1: Establish a single-stream target tracking model;

[0010] The single-stream object tracking model mainly consists of an image patch embedding module, a Transformer encoder module, and a tracking head;

[0011] The image patch embedding module consists of a layer of convolution kernels: , span is The input of the image block embedding module is an image pair, which includes two images: the template and the search area. The template is obtained by cutting out the target position in the initial frame of the video stream, and the search area is obtained by cutting out the target position in the previous frame in the current frame of the video stream. The template and the search area are first downsampled 16 times by a convolutional layer with shared parameters, and then transformed into word-meta sequences by deformation, and position encoding is added. Classification words and distillation words are also introduced here. The word-meta sequences generated by the template and the search area are then spliced ​​with the classification words and distillation words into a new word-meta sequence as the output of the image block embedding module.

[0012] The Transformer encoder module is mainly composed of multiple stacked multi-head attention modules and multi-layer perceptrons. Its input is the new word-meta sequence output by the image block embedding module, and its output is the word-meta sequence features that complete feature extraction and feature interaction. After completing feature extraction and feature interaction, the word-meta sequence features are re-split into template features and search area features. Classification words and distillation words are only used to enhance feature extraction.

[0013] The tracking head mainly consists of a classification branch and a regression branch. The inputs of both branches are search area features. The classification branch outputs the score of each feature point in the downsampled search area to obtain the approximate position of the center point of the target. The regression branch calibrates the center point position of the target and outputs the size of the target. The outputs of the classification head and the regression head are combined to obtain the accurate position of the target.

[0014] Step 2: Load the initialization weights, and use the weights of the DeiT image classification model based on knowledge distillation as the initialization weights of the backbone network, as follows: The image block embedding module and the Transformer encoder module in the single-stream target tracking model need the weights of the DeiT image classification model for initialization. The structure of the image block embedding module in the single-stream target tracking model is different from that of the DeiT image classification model. The following processing needs to be performed during the weight loading process: Since the input format of the DeiT image classification model is an image, the input in the single-stream target tracking model is two images of the template and the search area, and the corresponding position encoding formats are different; in order to adapt the format, first separate the classification words and distillation words of the position encoding in the weights, transform the remaining sequence-form position encoding in the position encoding into a two-dimensional image format, and then adapt the two-dimensional image format to the different image sizes of the template and the search area through bilinear interpolation to realize the loading of the initialization weights of the position encoding part; the Transformer encoder in the single-stream target tracking model has the same structure as the Transformer encoder of the DeiT image classification model, and the weights can be directly loaded;

[0015] Step 3: Build the computational framework of knowledge distillation, as follows: The knowledge distillation part includes the information expression method of distillation, the configuration of the teacher model, and the loss function;

[0016] In the information expression method of distillation, the similarity matrix of template features and search area features is introduced, and the cosine similarity of each word in the template features and the search area features is calculated pairwise. The formula is:

[0017]

[0018] in, and They are word-units in the template features and word-units in the search area features, respectively. The similarity matrix is ​​used to describe the relationship between the template features and the search area features;

[0019] The teacher model uses the OSTrack model for single-stream target tracking. Since the output features of all word units are required, the model without candidate word unit elimination is used. The student model is the single-stream target tracking model in step 1. During the training process, the teacher model inputs the same image data as the student model, adopts the evaluation mode, does not perform gradient calculation and parameter update, and only retains the similarity matrix calculated by the template features and search area features after feature extraction of the teacher model. The similarity matrix output by the student model uses the similarity matrix of the teacher model as a supervision signal.

[0020] In terms of loss function, the classification head of the student model uses the Focal loss function, the regression head uses the GIOU loss function and the L1 loss function, and the loss function of the similarity matrix distillation uses the L1 loss function; the total loss function is:

[0021]

[0022] in, , , , They represent the weights of the Focal loss function, L1 loss function, GIOU loss function, and similarity matrix distillation loss function respectively.

[0023] Beneficial effects of the present invention:

[0024] (1) The Token Similarity Distillation Algorithm proposed in this paper can effectively utilize the features of the single-stream target tracking teacher model and achieve high-performance tracking on the edge computing platform.

[0025] (2) The lightweight target tracking algorithm proposed in this invention better balances speed and accuracy, and significantly improves the running speed of the target tracking algorithm in the edge computing platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a framework diagram of the model structure and distillation of the present invention.

[0027] Figure 2 Schematic diagram of similarity matrix calculation. DETAILED DESCRIPTION

[0028] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.

[0029] First, a single-stream target tracking model is built. The structure of the single-stream target tracking model is as follows: Figure 1As shown in the figure, the ViT-Tiny model mentioned in DeiT is used as the backbone network of the target tracking model, that is, a 12-layer, 192-channel Transformer encoder for feature extraction and feature fusion. The head uses classification plus regression to obtain the target position. The classification head uses 5 convolutional layers, and finally uses sigmoid calculation to obtain the probability that each downsampled position is the target; the regression head uses two groups of 5-layer convolutional layers, one for regressing the offset of the target center point, and the other for regressing the size of the target.

[0030] The process of image block embedding can be expressed as:

[0031]

[0032]

[0033] In the formula, is the search area, For the template, is the convolutional layer in the image patch embedding, represents the operation of flattening the image block, and is the word sequence converted from the search area and template. The input of the Transformer encoder block can be expressed as:

[0034]

[0035] In the formula, is the feature of the input Transformer encoder, Represents a splicing operation, Represents a classification term, represents the distilled word, and Represent the position codes corresponding to the template and the search area respectively.

[0036] The process of the Transformer encoder block can be expressed as:

[0037]

[0038] In the formula, represents the input feature of the nth layer, represents the attention module, represents a multi-layer perceptron module, Represents the features of the nth layer output.

[0039] During the pre-training weight loading process, the position encoding is first split:

[0040]

[0041] In the formula, are the position codes corresponding to the image, classification word unit and distillation word unit respectively. Represents a split operation, is the position code in the pre-trained weights. Then the position code of the image part is processed as follows:

[0042]

[0043] In the formula, is the position code corresponding to the search area, represents the flattening operation, represents bilinear interpolation, To represent the deformation to a two-dimensional shape, the position encoding corresponding to the template is obtained in the same way.

[0044] The distillation similarity matrix is ​​calculated as Figure 2 As shown in the figure, firstly, the image encoder is used to extract features from the template and the search area image, and then the cosine similarity between the features of the template and the features of the search area is calculated to obtain the similarity matrix. At the same time, the similarity matrix is ​​also calculated in the single-stream target tracking teacher model, and knowledge is transferred through this similarity matrix. The similarity calculation formula is:

[0045]

[0046] In the formula, and They are the word-grams in the template features and the word-grams in the search area features respectively.

[0047] During the training process, the classification head uses the Focal loss function to balance positive and negative samples. After the regression head calculates the target box, it uses the L1 loss function and the GIOU loss function to jointly optimize it, and the distillation calculation of the similarity matrix is ​​optimized by L1 Loss. The calculation process of the distillation loss function is as follows:

[0048]

[0049] In the formula, and The subscripts in the teacher model similarity matrix and the student model similarity matrix are data, and represents the shape of the similarity matrix, Represents the loss function corresponding to the similarity matrix. The total loss function is expressed as

[0050]

[0051] in, , , , They represent the weights of the Focal loss function, L1 loss function, GIOU loss function, and similarity matrix distillation loss function respectively.

[0052] Table 1 is a comparison of the lightweight tracking algorithm proposed in the present invention with other efficient trackers. It can be seen that the target tracking algorithm proposed in the present invention achieves tracking performance close to that of HiT while doubling the speed of the HiT algorithm. It reaches 66.5% AO on the GOT10K dataset, exceeding HiT by 1.5% AO. The result on the TrackingNet dataset is only 0.1% AUC worse, and the result on the LaSOT dataset is only 0.5% AUC worse. It is faster and has higher tracking accuracy than other high-performance target tracking algorithms, achieving a better balance between speed and accuracy.

[0053] Table 1 Performance comparison

[0054]

[0055] In addition, as shown in Table 2, the parameter amount and theoretical calculation amount of the present invention are also greatly compressed. Compared with the current HiT model with better performance, the parameter amount is reduced by 6 times and the calculation amount is reduced by 2 times, realizing a faster target tracker.

[0056] Table 2 Comparison of parameter calculation amount

[0057] .

Claims

1. A single-stream lightweight target tracking method based on knowledge distillation, characterized in that: Here are the steps: Step 1: Establish a single-stream target tracking model; The single-stream object tracking model mainly consists of an image patch embedding module, a Transformer encoder module, and a tracking head; The image patch embedding module consists of a layer of convolution kernels: , span is The input of the image block embedding module is an image pair, which includes two images: the template and the search area. The template is obtained by cutting out the target position in the initial frame of the video stream, and the search area is obtained by cutting out the target position in the previous frame in the current frame of the video stream. The template and the search area are first downsampled 16 times by a convolutional layer with shared parameters, and then transformed into word-meta sequences by deformation, and position encoding is added. Classification words and distillation words are also introduced here. The word-meta sequences generated by the template and the search area are then spliced ​​with the classification words and distillation words into a new word-meta sequence as the output of the image block embedding module. The Transformer encoder module is mainly composed of multiple stacked multi-head attention modules and multi-layer perceptrons. Its input is the new word-meta sequence output by the image block embedding module, and its output is the word-meta sequence features that complete feature extraction and feature interaction. After completing feature extraction and feature interaction, the word-meta sequence features are re-split into template features and search area features. Classification words and distillation words are only used to enhance feature extraction. The tracking head consists of a classification branch and a regression branch, and the input of both branches is the search area features; The classification branch outputs the score of each feature point in the downsampled search area to obtain the approximate position of the center point of the target; the regression branch calibrates the center point position of the target and outputs the size of the target; the outputs of the classification head and the regression head are combined to obtain the accurate position of the target; Step 2: Load the initialization weights, and use the weights of the DeiT image classification model based on knowledge distillation as the initialization weights of the backbone network, as follows: The image block embedding module and the Transformer encoder module in the single-stream target tracking model need the weights of the DeiT image classification model for initialization. The structure of the image block embedding module in the single-stream target tracking model is different from that of the DeiT image classification model. The following processing needs to be performed during the weight loading process: Since the input format of the DeiT image classification model is an image, the input in the single-stream target tracking model is two images of the template and the search area, and the corresponding position encoding formats are different; in order to adapt the format, first separate the classification words and distillation words of the position encoding in the weights, transform the remaining sequence-form position encoding in the position encoding into a two-dimensional image format, and then adapt the two-dimensional image format to the different image sizes of the template and the search area through bilinear interpolation to realize the loading of the initialization weights of the position encoding part; the Transformer encoder in the single-stream target tracking model has the same structure as the Transformer encoder of the DeiT image classification model, and the weights can be directly loaded; Step 3: Build the computational framework of knowledge distillation, as follows: The knowledge distillation part includes the information expression method of distillation, the configuration of the teacher model, and the loss function; In the information expression method of distillation, the similarity matrix of template features and search area features is introduced, and the cosine similarity of each word in the template features and the search area features is calculated pairwise. The formula is: in, and They are word elements in the template features and word elements in the search area features, respectively. The similarity matrix is ​​used to describe the relationship between the template features and the search area features; The teacher model uses the OSTrack model for single-stream target tracking. Since the output features of all word units are required, the model without candidate word unit elimination is used. The student model is the single-stream target tracking model in step 1. During the training process, the teacher model inputs the same image data as the student model, adopts the evaluation mode, does not perform gradient calculation and parameter update, and only retains the similarity matrix calculated by the template features and search area features after feature extraction of the teacher model. The similarity matrix output by the student model uses the similarity matrix of the teacher model as a supervision signal. In terms of loss function, the classification head of the student model uses the Focal loss function, the regression head uses the GIOU loss function and the L1 loss function, and the loss function of the similarity matrix distillation uses the L1 loss function; the total loss function is: in, , , , They represent the weights of the Focal loss function, L1 loss function, GIOU loss function, and similarity matrix distillation loss function respectively.

Citation Information

Patent Citations

  • Satellite video multi-target tracking method based on knowledge distillation

    CN115797794A

  • Running event detection method based on large model distillation and electronic equipment

    CN117831116A

  • Knowledge distillation-based lightweight method for pedestrian re-identification in video scene and method thereof

    CN118351415A

  • Lightweight unmanned aerial vehicle single target tracking method based on distillation

    CN118379332A

  • Landing tracking control method and system based on lightweight twin network and unmanned aerial vehicle

    US20220332415A1

Cited By

  • Visual target tracking dynamic calculation and distribution method based on scene complexity perception

    CN121414787A