Knowledge distillation method and system for target tracking
By combining mask feature distillation, context adaptive learning and related distance supervision, the problems of insufficient versatility and poor performance in target tracking are solved, and efficient tracking performance in edge devices and complex scenarios are achieved.
Patent Information
- Application Number
- CN202510163448.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The prior art has problems of insufficient universality and poor performance in the target tracking process, especially when adapting to edge devices and complex scenarios.
Using a knowledge distillation method, the spatial and feature knowledge of the teacher model is extracted and applied to the student model through a combination of mask-based feature distillation, context adaptive learning, and related distance supervision to improve its performance in different depths and complex scenarios.
It achieves improved versatility in the target tracking process and significantly better than existing baseline performance, especially in edge devices and complex scenarios.
Smart Images

Figure CN119648748B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of tracking video technology, and in particular, to a knowledge distillation method and system for object tracking. Background Art
[0002] Video is a typical form of multimedia data. Tracking important objects in video content is a key task. With the development of intelligent embedded devices, object tracking through edge devices has become increasingly important in various fields, including intelligent transportation systems and disaster relief operations. This reduces the potential latency associated with transmitting large amounts of data, thus helping to capture time-sensitive targets more effectively. In addition, it helps the tracking model to operate effectively in the case of limited bandwidth and unstable communication.
[0003] Currently, there are two main trends in the automatic real-time object tracking technology of edge devices. The first method focuses on simplifying the model at the structural level, which includes compressing large models and designing lightweight alternatives. However, as the compression ratio increases, this compensation often leads to a significant decrease in performance. On the other hand, developing a network architecture that is both lightweight and high-performance remains a considerable challenge. In contrast, another method emphasizes using knowledge distillation (KD) technology to integrate the knowledge obtained from large models into smaller models. This method has an advantage over the previous methods because it can provide more accurate adaptive supervision information. However, when adapting KD to object tracking on edge devices, some challenges still exist.
[0004] Currently, there are two dominant tracking model architectures: Siamese network-based and ViT-based, and their tracking pipelines are significantly different. In addition, for their respective backbone networks: convolutional neural network (CNN) and vision transformer (ViT), existing research shows that KD algorithms that are beneficial to CNN may have an adverse effect on ViT. Therefore, it is urgent to study knowledge distillation methods that can be effectively applied to both tracking architectures.
[0005] Secondly, the object tracking task is only divided into foreground and background regions. This problem is similar to a binary classification problem, which limits the enhancement through logit distillation. In terms of feature extraction, the background may contain interfering objects that are similar to or belong to the same category as the target. However, most existing methods ignore the inherent conflict between classification and regression.
[0006] Thirdly, a significant challenge related to the current KD technology applied to tracking models is that when the difference between the teacher model and the student model is overly significant, the benefits obtained from KD are significantly reduced. A common strategy to mitigate this problem includes adding an intermediate teacher model to facilitate the transition, and this method greatly increases the training overhead.
[0007] Therefore, how to improve the versatility in the target tracking process and outperform the existing baseline performance has become a technical problem that needs to be solved urgently. SUMMARY OF THE INVENTION
[0008] In order to accurately predict carbon prices, this application provides a knowledge distillation method and system for target tracking.
[0009] In the first aspect, the present application provides a knowledge distillation method for target tracking using the following technical solution:
[0010] A knowledge distillation method for target tracking, comprising:
[0011] Mask-based feature distillation extracts spatial attention and channel attention from the feature fusion layer of the teacher model, generates an attention mask, and multiplies the attention mask with the feature map of the student model;
[0012] Based on context-adaptive learning, the spatial and feature knowledge of each layer of the teacher model is adaptively integrated through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to the different depths of the network;
[0013] Distillation is performed through correlation distance supervision, using a positive affine transformation to project the regression box predictions of the teacher model and the student model into the probability distribution space, and the ability gap between the teacher model and the student model is supervised by evaluating the correlation distance of the probability distribution.
[0014] Optionally, the step of extracting spatial attention and channel attention from the feature fusion layer of the teacher model, generating an attention mask, and multiplying the attention mask with the feature map of the student model by the mask-based feature distillation comprises:
[0015] Get the teacher model and the student model, align the feature maps of the teacher model and the student model, and ensure that the size of the feature map of the student model is consistent with that of the teacher model;
[0016] Extract spatial attention and channel attention from the feature fusion layer of the teacher model to generate an attention mask;
[0017] Multiply the attention mask with the feature maps of the teacher model and the student model.
[0018] Optionally, the context-based adaptive learning method adaptively integrates the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to the different depths of the network, including:
[0019] Perform multi-scale pooling on the feature maps of the teacher model and the student model in a context-adaptive learning manner to divide them into shallow, middle, and deep features;
[0020] Adaptively integrate the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron;
[0021] Adjust the learning focus of the student model so that the student model can be adaptively adjusted according to the spatial and feature knowledge of each layer in the teacher model.
[0022] Optionally, the step of distilling through correlation distance supervision, which projects the regression box predictions of the teacher model and the student model into the probability distribution space using a positive affine transformation and supervises the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distributions, includes:
[0023] Use the Pearson distance as a metric to evaluate the correlation distance between the teacher model and the student model in the probability distribution space;
[0024] Combine the correlation distance and project the regression box predictions of the teacher model and the student model into the probability distribution space using a positive affine transformation;
[0025] Supervise the ability gap between the teacher model and the student model through inter-class position correlation and intra-class position correlation in the probability distribution space.
[0026] Optionally, the method is applicable to object tracking models based on Siamese networks and Vision Transformers.
[0027] Optionally, the method enhances the focusing ability of the student model on the target area and reduces the influence of low resolution and intra-class interference by combining spatial attention and channel attention.
[0028] Optionally, the method reduces the overfitting of the student model to the output of the teacher model through correlation distance supervision and improves the tracking performance of the student model in complex scenarios.
[0029] In a second aspect, the present application provides a knowledge distillation system for object tracking, and the knowledge distillation system for object tracking includes:
[0030] An attention mask generation module, which is used for feature distillation based on a mask, extracts spatial attention and channel attention from the feature fusion layer of the teacher model, generates an attention mask, and multiplies the attention mask with the feature map of the student model;
[0031] A learning adjustment module, which is used in a context-adaptive learning manner to adaptively integrate the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to different depths of the network;
[0032] An ability gap evaluation module is used for distillation through correlation distance supervision. It projects the regression box predictions of the teacher model and the student model into the probability distribution space by using positive affine transformation, and supervises the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution.
[0033] In a third aspect, the present application provides a computer device, which includes: a memory and a processor. When the processor runs the computer instructions stored in the memory, it executes the method described above.
[0034] In a fourth aspect, the present application provides a computer-readable storage medium, including instructions. When the instructions run on a computer, the computer is made to execute the method described above.
[0035] In summary, the present application includes the following beneficial technical effects:
[0036] Through mask-based feature distillation, the present application extracts spatial attention and channel attention from the feature fusion layer of the teacher model to generate an attention mask, and multiplies the attention mask with the feature map of the student model; in a context-adaptive learning manner, the spatial and feature knowledge of each layer of the teacher model is adaptively integrated through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to different depths of the network; through correlation distance supervision for distillation, the regression box predictions of the teacher model and the student model are projected into the probability distribution space by using positive affine transformation, and the ability gap between the teacher model and the student model is supervised by evaluating the correlation distance of the probability distribution. The technical effect of improving generality during the target tracking process and having better performance than existing baselines is achieved. Description of the Drawings
[0037] Figure 1 is a schematic structural diagram of a computer device in the hardware operating environment related to the solution of the embodiment of the present application;
[0038] Figure 2 is a schematic flowchart of an embodiment of the knowledge distillation method for target tracking of the present application;
[0039] Figure 3 is a process concept diagram in the knowledge distillation method for target tracking of the present application;
[0040] Figure 4 is a feature conversion diagram attached to the vit-based model in the knowledge distillation method for target tracking of the present application;
[0041] Figure 5 is a structural block diagram of an embodiment of the knowledge distillation system for target tracking of the present application. Detailed Embodiments
[0042] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application through the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0043] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of a computer device for the hardware operating environment involved in the solution of the embodiment of this application.
[0044] As Figure 1 shown, the computer device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0045] Those skilled in the art can understand that Figure 1 the structure shown in
[0046] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 1 As
[0047] In Figure 1In the computer device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in this application can be arranged in the computer device, and the computer device calls, through the processor 1001, the knowledge distillation program stored in the memory 1005 for target tracking, and executes the knowledge distillation method for target tracking provided by the embodiments of this application.
[0048] An embodiment of this application provides a knowledge distillation method for target tracking. Refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the knowledge distillation method for target tracking in this application.
[0049] In this embodiment, the knowledge distillation method for target tracking includes the following steps:
[0050] Step S10: Mask-based feature distillation, extracting spatial attention and channel attention from the feature fusion layer of the teacher model, generating an attention mask, and multiplying the attention mask by the feature map of the student model.
[0051] It should be noted that the step of mask-based feature distillation, extracting spatial attention and channel attention from the feature fusion layer of the teacher model, generating an attention mask, and multiplying the attention mask by the feature map of the student model includes: obtaining the teacher model and the student model, aligning the feature maps of the teacher model and the student model to ensure that the feature map of the student model is the same size as that of the teacher model; extracting spatial attention and channel attention from the feature fusion layer of the teacher model to generate an attention mask; multiplying the attention mask by the feature maps of the teacher model and the student model.
[0052] In specific implementation, this embodiment provides a comprehensive description of the OTKD pipeline related to the target tracking task. It consists of three stages, as Figure 3 shown. In the context of a viti-based model, the process includes performing the operations shown in Figure 4 . Extracting features and classifying them into four different categories. Subsequently, these features are reconstructed into a two-dimensional format, and then the OTKD method is applied. For the curve in Figure 3 , it is the probability distribution formed by four numerical values of the regression box, and the coordinates can be x, y (equivalent to each frame of each video sequence having these four numerical values, and their projection onto a two-dimensional space will form a probability distribution, and the function curve is just a fitting schematic of the probability distribution).
[0053] In specific implementations, compared with tasks such as target classification or detection, the tracking task only differentiates between background and foreground regions. However, there are interfering elements in the background that are similar to the target, which requires the tracker to have both feature extraction and region prediction capabilities (corresponding to classification and regression respectively).
[0054] Extract knowledge from the feature fusion layer of the teacher model through multi-layer feature distillation. First, fix the feature maps of each layer in the teacher model. Then, align the feature maps of the corresponding layers in the student model to ensure that their sizes are the same as those of the teacher model. Subsequently, modify the dimensions of the deep feature maps in the student model to conform to the structure of the shallow feature maps.
[0055] On this basis, mask-based features (MBF) are developed to extract spatial attention and channel attention from the feature maps generated by each layer of the teacher model. Assume that the feature fusion layers of the teacher model are m1T, m2T, …, mNT. N is the number of feature fusion layers. And the student model is m1S, m2S, …, mNs. These two forms of attention are used as attention masks and multiplied by each feature map respectively. Subsequently, the features obtained through the above operations are combined to form the final output. These two attention masks can be described as:
[0056]
[0057]
[0058]
[0059]
[0060] Among them, represents the channel attention mask of the i-th layer of the student model, represents the spatial attention mask of the i-th layer of the student model, is the feature map of the i-th layer of the student model, , , represent the number of channels, height, and width of the i-th layer of the student model respectively. Among them, represents the feature map of the student after adjustment and aligned with the feature map of the teacher model in size. is the loss after channel attention masking, is the loss after spatial attention masking. D represents the distance function used to quantify the difference between the student and teacher models, and ⊙ represents the multiplication broadcast operation.
[0061] Step S20: In a context - adaptive learning manner, adaptively integrate the spatial and feature knowledge of each layer of the teacher model through a multi - layer perceptron so that the student model can adaptively adjust the learning focus according to the different depths of the network.
[0062] It can be understood that the step of adaptively integrating the spatial and feature knowledge of each layer of the teacher model through a multi - layer perceptron in a context - adaptive learning manner so that the student model can adaptively adjust the learning focus according to the different depths of the network includes: performing multi - scale pooling on the feature maps of the teacher model and the student model in a context - adaptive learning manner to divide them into shallow, middle, and deep features; adaptively integrating the spatial and feature knowledge of each layer of the teacher model through a multi - layer perceptron; and adjusting the learning focus of the student model so that the student model can adaptively adjust according to the spatial and feature knowledge of each layer in the teacher model.
[0063] It should be noted that through the above mechanism, the student model can benefit from the guidance provided by each layer of the teacher model. However, in the matching transformation and multi - layer fusion process, the hierarchical structure information of the feature map is destroyed. On the other hand, although this study uses the MBF method to divide knowledge into spatial information knowledge and target appearance knowledge, the importance of feature information and spatial information varies in neural networks at different layers. This embodiment introduces a context - adaptive learning mechanism (CAL). When solving the first problem, the feature map is pooled at three scales, divided into shallow, middle, and deep, denoted as P(). Therefore, the knowledge is divided into context information at different levels.
[0064] Before performing multi - scale pooling, the spatial and feature knowledge of each layer of the teacher model is adaptively integrated using a three - layer MLP. This gives the student model enhanced adaptive adjustment ability. Subsequently, the L2 loss between the teacher model and the student model is calculated. The above comprehensive loss can be transformed into the following form:
[0065]
[0066]
[0067]
[0068]
[0069] Among them, is the feature of the j - th layer of the student model after passing through the channel attention mask and the spatial attention mask, is the feature of the j - th layer of the teacher model after passing through the channel attention mask and the spatial attention mask. P() is the pooling function. MLP() is the multi - layer perceptron operation. , , They are the features of three scales after multi-scale pooling of the features of the j-th layer of the teacher model after masking. , , They are the features of three scales after multi-scale pooling of the features of the j-th layer of the student model after masking. It is the loss function.
[0070] In this embodiment, the parameters of the MLP are incorporated into the iterative training of the tracking model, aiming to enable the student model to autonomously determine whether the focus of learning at different scales and levels of the network is spatial knowledge or feature knowledge.
[0071] Step S30: Perform distillation through correlation distance supervision. Use positive affine transformation to project the regression box predictions of the teacher model and the student model into the probability distribution space, and supervise the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution.
[0072] In specific implementation, the step of performing distillation through correlation distance supervision, using positive affine transformation to project the regression box predictions of the teacher model and the student model into the probability distribution space, and supervising the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution includes: using Pearson distance as a metric to evaluate the correlation distance between the teacher model and the student model in the probability distribution space; combining the correlation distance and using positive affine transformation to project the regression box predictions of the teacher model and the student model into the probability distribution space; supervising the ability gap between the teacher model and the student model through inter-class position correlation and intra-class position correlation in the probability distribution space.
[0073] It should be noted that the method is applicable to object tracking models based on Siamese networks and Vision Transformers.
[0074] It can be understood that the method enhances the focusing ability of the student model on the target area by combining spatial attention and channel attention, and reduces the influence of low resolution and intra-class interference.
[0075] In specific implementation, the method reduces the overfitting of the student model to the output of the teacher model through correlation distance supervision, and improves the tracking performance of the student model in complex scenarios.
[0076] It should be noted that the regression bounding boxes depict the teacher's prediction of the target position for each frame. Classical KD requires exact correspondence between each frame of the teacher's output, and the best performance can only be obtained when all values show exact consistency. This brings two related problems. First, the student model shows a high sensitivity to overfitting, resulting in suboptimal generalization in specific cases. This phenomenon will be further elucidated through subsequent analysis of attention visualization. On the other hand, the regression bounding boxes generated by the teacher model are distributed discretely in space, which poses a challenge for the student model to determine the optimal fitting plane. Therefore, this embodiment adopts a positive affine transformation to project the video sequence into the probability distribution space. It supervises the ability gap between the teacher and the student by evaluating the relevant distance of the probability distribution within the specified space, and refers to this concept as inter-class position correlation.
[0077] In specific implementation, similar or interfering objects of the same category may also appear in each video sequence, and the teacher model shows different confidence levels in different regions. Represent the relationship between these confidence levels as intra-class position correlation, which can then be transferred to the student model:
[0078]
[0079] The Pearson correlation coefficient representing the teacher-student relationship within the i-th video sequence, is the number of picture frames of the i-th video sequence.
[0080] Therefore, the total training loss can be expressed as:
[0081] .
[0082] represents the label loss, while , and are the coefficients of the balanced loss components.
[0083] It should be noted that this embodiment proposes a new knowledge distillation method for object tracking, called OTKD, which is applicable to two mainstream tracking architectures. The method of this embodiment combines an attention mask to promote the distillation of feature knowledge and introduces an MLP to achieve the learning ability of adaptive distillation. In addition, it also enables the student to focus on understanding the probability distribution of the regression bounding boxes generated by the teacher model, rather than striving for exact correspondence with the teacher's output.
[0084] In this embodiment, one of the implementation environments is as follows: The input is a video stream collected by edge devices such as drones and the initial position of the selected target to be tracked (manually drawing a bounding box or initializing a box through object detection). Then, the target image and the search area image are input into the tracking model. The tracking model extracts the features of the target image and the search area image respectively and compares them. By selecting the region most relevant to the target (with the most similar appearance features and the spatial position features should maintain temporality) as the next frame position for target tracking, and outputting the regression box coordinates of the position (the algorithm automatically draws a box).
[0085] In this embodiment, through mask-based feature distillation, spatial attention and channel attention are extracted from the feature fusion layer of the teacher model to generate an attention mask, and the attention mask is multiplied by the feature map of the student model; in the way of context adaptive learning, the spatial and feature knowledge of each layer of the teacher model is adaptively integrated through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to the different depths of the network; through correlation distance supervision for distillation, the regression box predictions of the teacher model and the student model are projected into the probability distribution space by positive affine transformation, and the ability gap between the teacher model and the student model is supervised by evaluating the correlation distance of the probability distribution. The technical effect of improving generality during the target tracking process and having better performance than the existing baseline is achieved.
[0086] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a program for knowledge distillation for target tracking is stored. When the program for knowledge distillation for target tracking is executed by a processor, the steps of the method for knowledge distillation for target tracking as described above are implemented.
[0087] Refer to Figure 5 , Figure 5 which is a structural block diagram of an embodiment of the knowledge distillation system for target tracking of the present application.
[0088] As Figure 5 shown, the knowledge distillation system for target tracking proposed in the embodiment of the present application includes:
[0089] An attention mask generation module 10, configured to extract spatial attention and channel attention from the feature fusion layer of the teacher model through mask-based feature distillation, generate an attention mask, and multiply the attention mask by the feature map of the student model;
[0090] A learning adjustment module 20, configured to adaptively integrate the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron in the way of context adaptive learning so that the student model can adaptively adjust the learning focus according to the different depths of the network;
[0091] The ability gap evaluation module 30 is used for distillation through correlation distance supervision. It projects the regression box predictions of the teacher model and the student model into the probability distribution space by using positive affine transformation, and supervises the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution.
[0092] It should be understood that the above is only an example, and does not constitute any limitation to the technical solution of the present application. In specific applications, those skilled in the art can set according to needs, and the present application does not make any restrictions in this regard.
[0093] In this embodiment, through mask-based feature distillation, spatial attention and channel attention are extracted from the feature fusion layer of the teacher model to generate an attention mask, and the attention mask is multiplied by the feature map of the student model; based on the context adaptive learning method, the spatial and feature knowledge of each layer of the teacher model is adaptively integrated through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to different depths of the network; through correlation distance supervision for distillation, the regression box predictions of the teacher model and the student model are projected into the probability distribution space by using positive affine transformation, and the ability gap between the teacher model and the student model is supervised by evaluating the correlation distance of the probability distribution. The technical effect of improving generality during the target tracking process and having better performance than existing baselines is achieved.
[0094] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present application. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no restrictions are made here.
[0095] In addition, for technical details not described in detail in this embodiment, reference can be made to the method for knowledge distillation for target tracking provided in any embodiment of the present application, which will not be elaborated here.
[0096] In addition, it should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0097] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as Read Only Memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0099] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A knowledge distillation method for target tracking, characterized in that: include: Mask-based feature distillation in tracking video content, extracting spatial attention and channel attention from the feature fusion layer of the teacher model, generating an attention mask, and multiplying the attention mask with the feature map of the student model to enhance the student model's ability to focus on the target area and reduce low-resolution interference; Based on context-adaptive learning, the spatial and feature knowledge of each layer of the teacher model is adaptively integrated through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to the different depths of the network; Distillation through correlation distance supervision, using positive affine transformation to project the regression box predictions of the teacher model and the student model into the probability distribution space, and supervising the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution; The mask-based feature distillation extracts spatial attention and channel attention from the feature fusion layer of the teacher model, generates an attention mask, and multiplies the attention mask by the feature map of the student model, including: Obtain the teacher model and the student model, align the feature maps of the teacher model and the student model, and ensure that the size of the feature map of the student model is consistent with that of the teacher model; Extracting spatial attention and channel attention from the feature fusion layer of the teacher model to generate an attention mask; Multiplying the attention mask with the feature maps of the teacher model and the student model; The context-based adaptive learning method adaptively integrates the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron so that the student model can adaptively adjust the learning focus according to the different depths of the network, including: Based on context-adaptive learning, the feature maps of the teacher model and the student model are pooled at multiple scales to be divided into shallow, medium and deep features. Adaptively integrate the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron; Adjusting the learning focus of the student model so that the student model can be adaptively adjusted based on the spatial and feature knowledge of each layer in the teacher model; The step of performing distillation through correlation distance supervision, projecting the regression box predictions of the teacher model and the student model into the probability distribution space using a positive affine transformation, and supervising the ability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution, includes: Pearson distance is used as a metric to evaluate the correlation distance between the teacher model and the student model in the probability distribution space; Combining the correlation distance and using a positive affine transformation to project the regression box predictions of the teacher model and the student model into a probability distribution space; The ability gap between the teacher model and the student model is supervised by the inter-class position correlation and the intra-class position correlation in the probability distribution space.
2. The knowledge distillation method for target tracking according to claim 1, characterized in that: The method is applicable to both Siamese network-based and visual Transformer-based object tracking models.
3. A knowledge distillation system for target tracking, characterized in that: The system executes the method according to claim 1, wherein the knowledge distillation system for target tracking comprises: An attention mask generation module for mask-based feature distillation, extracting spatial attention and channel attention from the feature fusion layer of the teacher model, generating an attention mask, and multiplying the attention mask with the feature map of the student model; The learning adjustment module is used to adaptively integrate the spatial and feature knowledge of each layer of the teacher model through a multi-layer perceptron based on context-based adaptive learning so that the student model can adaptively adjust the learning focus according to the different depths of the network; The capability gap assessment module, used for distillation through correlation distance supervision, adopts positive affine transformation to project the regression box predictions of the teacher model and the student model into the probability distribution space, and supervises the capability gap between the teacher model and the student model by evaluating the correlation distance of the probability distribution.
4. A computer device, characterized in that: The device comprises: a memory and a processor, wherein the processor executes the method according to any one of claims 1 to 2 when running computer instructions stored in the memory.
5. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Pedestrian re-identification method and device based on mask guided double-flow network
CN118799923A
Sparsity constraints and knowledge distillation based learning of sparser and compressed neural networks
US20200387782A1