Multi-stage decoupling knowledge distillation method and system based on output normalization

Through a multi-level decoupled knowledge distillation method based on output normalization, the problems of large dataset annotation requirements and low hardware configuration in visual recognition models are solved, and the model memory usage is reduced and the accuracy is improved. It is suitable for image classification, object detection and instance segmentation tasks.

CN120635660APending Publication Date: 2025-09-12ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510474404.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing visual recognition models cannot be effectively deployed when the dataset annotation requirements are large and the hardware configuration is low. In addition, the model complexity is high, resulting in large memory usage and limited accuracy improvement.

Method used

A multi-level decoupled knowledge distillation method based on output normalization is adopted. By designing instance-level, batch-level and category-level decoupled alignment components, combined with the temperature pool mechanism and optimal normalization strategy, the knowledge distillation process is optimized and the performance and adaptability of the student model are improved.

Benefits of technology

Significantly reduce model memory usage, improve the accuracy of small-scale models, adapt to different hardware environments, and improve the performance of image classification, object detection, and instance segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635660A_ABST
    Figure CN120635660A_ABST
Patent Text Reader

Abstract

The invention discloses a multistage decoupling knowledge distillation method and system based on output normalization, and the method comprises the following steps: S1, inputting an image into a teacher model and a student model, carrying out the multiple convolution processing of the image through neural networks in the teacher model and the student model, carrying out the full-link layer processing, and obtaining the logits output by the models, the logit of the student model and the logit of the teacher model extract decoupled target class and non-target class probability distribution characteristics through a softmax function; s2, introducing a temperature pool mechanism, and performing multi-temperature distillation after extracting probability distribution characteristics of the teacher model and the student model by designing a multi-stage decoupling knowledge distillation assembly; s3, designing a preferable normalization strategy component, and distributing different normalization strategies for the knowledge distillation framework according to different visual identification tasks; and S4, combining the loss calculated by the multi-stage decoupling knowledge distillation assembly to obtain the total loss of model training. According to the method, the universality of knowledge distillation in the fields of image classification, target detection and instance segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual recognition model optimization, and in particular relates to a multi-level decoupling knowledge distillation method and system based on warm output normalization. Background Art

[0002] With the continuous development of modern visual recognition technology, the past few decades have witnessed a boom in deep learning for computer vision tasks. General models are trained on numerous datasets, and accuracy is often the most important metric for evaluating model performance. As model accuracy continues to improve, the memory usage also increases. To address this challenge, researchers have introduced knowledge distillation to reduce model capacity.

[0003] However, current image classification algorithms and their applications face two major challenges: First, labeling massive datasets requires significant manpower and resources. Furthermore, few companies are willing to make their datasets public, resulting in a lack of datasets and ineffective validation of most image classification models. Second, as model performance continues to improve, the complexity of these models increases, making them infeasible to deploy on certain devices with lower hardware configurations, thus failing to meet the demands of practical applications.

[0004] Given the above situation, considering how to improve the knowledge distillation strategy to optimize the model to reduce model memory loss while improving model accuracy has become a technical problem that needs to be solved in this field. Based on this, the present invention proposes a multi-level decoupling knowledge distillation method and system based on output normalization, which can significantly improve the performance of small-scale models and greatly reduce the model's memory usage. Summary of the Invention

[0005] In response to the above-mentioned status quo of the prior art, the present invention proposes a multi-level decoupling knowledge distillation method and system based on output normalization.

[0006] The present invention adopts the following technical solutions:

[0007] A multi-level decoupling knowledge distillation method based on output normalization includes the following steps:

[0008] S1: The image is input into the given teacher model and student model. The neural networks in the teacher model and the student model respectively perform several convolution operations on the image, and then process it through the fully connected layer to obtain the logistic regression logit output of the model. The logit of the student model and the teacher model respectively use the softmax function to extract the decoupled target class and non-target class probability distribution features;

[0009] S2: Introducing a temperature pool mechanism, by designing a multi-level decoupled knowledge distillation component, and performing multi-temperature distillation after extracting the probability distribution features of the teacher model and the student model;

[0010] S3: Design an optimal normalization strategy component to assign different normalization strategies to the knowledge distillation framework according to different visual recognition tasks.

[0011] S4: Combine the losses calculated by the multi-level decoupled knowledge distillation components to obtain the total model training loss.

[0012] Preferably, in step S2, a temperature pool mechanism is used to expand the output of a single temperature into outputs of multiple temperatures. Specifically, prediction enhancement is performed through temperature calibration, and prediction decoupling is performed through the coupling characteristics of knowledge distillation (KD). The temperature calibration method used can be expressed as:

[0013]

[0014] Among them, C represents the number of categories, z i,t and z i,c They represent the logit of category t and c output in the i-th input sample of the student model, p i,t,k is the temperature hyperparameter T k The probability of the next i-th input sample about the target class t, p i,\t,k It is defined as the sum of the probabilities of samples on all non-target classes under the same temperature parameter. Specifically, this method treats the probability distribution of non-target classes as a differentiable competitive space while retaining the independent modeling of the target class. Its expression can be written as:

[0015]

[0016] in, The i-th sample input has a temperature hyperparameter T on the non-target category n k In this temperature pool mechanism, T k (k∈[1,K]) forms a pool so that K outputs with different smoothness can be obtained.

[0017] Preferably, in step S2, the multi-level decoupled knowledge distillation component includes an instance-level decoupled alignment component (IDAD), a batch-level decoupled alignment component (BDAD), a class-level decoupled alignment component (CDAD) and a cross-entropy loss calculation component.

[0018] As a preference, in step S2, the instance-level decoupling alignment component is designed. First, a threshold parameter m is introduced. i To minimize the KL divergence (Kullback-Leibler divergence) of binary prediction enhancements between target and non-target categories of the teacher and student models, as well as the KL divergence between the enhanced predictions of non-target categories:

[0019]

[0020] Among them, L IDAD represents the instance-level disentangled alignment loss, i.e., IDAD component, m i represents the threshold parameter under sample i, KL represents KL divergence loss, B represents batch size, K represents the total number of temperatures, and C represents the number of categories. and Represents the distillation temperature T of the teacher and the student at the input sample i k The binary output of target class t and non-target class \t. and Represents the distillation temperature T of the teacher and the student at the input sample i k Instance-level disentangled alignment forces the student model to imitate the teacher's prediction for each instance. In addition, prediction enhancement via temperature scaling is employed to transfer knowledge at multiple smoothness levels to the student model.

[0021] Preferably, in step S2, a batch-level decoupled alignment component is designed, which first uses the target class input correlation and the non-target class input correlation to align the instance-level predictions at the batch level, rather than aligning the predictions only at the instance level. In this method, the Gram matrix is ​​used to quantify it:

[0022]

[0023] Among them, the matrix G i,k and It is defined as a square matrix of dimension B, where scalar B represents the batch size of sample i. In addition, when the temperature T k When acting on sample i, p i,k and Quantify the model's predicted probabilities for the target class and non-target class respectively. Then, calculate the loss based on the two given matrices and introduce the self-supervised threshold parameter m i The corresponding loss can be expressed as:

[0024]

[0025] Among them, L BDADis the batch-level decoupled alignment loss, i.e., the BDAD component, where α and β represent the hyperparameters of the target class and the non-target class, respectively. and are the input correlation matrices of teachers and students in the target and non-target classes, respectively, and the predicted temperature T of sample i k The following calculation is obtained.

[0026] Preferably, in step S2, a class-level decoupling alignment component is designed to first decouple the predictions corresponding to the target class and reinforce the correct class. This class correlation can be modeled by predicting the following data:

[0027]

[0028] Among them, M i,n,k is the transpose of the binary probability distribution of the target class and itself i,n,k The 2×2 matrix obtained by multiplication is is the transpose of the probability distribution of the non-target class and itself The (C-1)×(C-1) matrix obtained by multiplication.

[0029] As a preference, the category confidence threshold parameter h is introduced i , and let the student model learn this part of knowledge through the following loss:

[0030]

[0031] Among them, L CDAD is the class-level decoupled alignment loss, i.e., the CDAD component, h i represents the threshold parameter under sample i, and are the input correlation matrices for target and non-target classes and teacher and student, respectively, and the predicted temperature T k Calculated at sample i.

[0032] Preferably, in step S2, the cross entropy loss is calculated as follows:

[0033]

[0034] Among them, y represents the true label, CE represents the cross entropy loss, and p stu represents the probability distribution of student output.

[0035] As a preference, in step S3, a mechanism for prioritizing normalization of logit is proposed. Specifically, a function method for normalization operation is first defined:

[0036]

[0037] Here, x is the input logit and N(x) is the regularization function applied to the probability distribution of each prediction. Then p i,t,k 、p i,\t,k and can be rewritten as P i,t,k 、P i,\t,k and

[0038]

[0039] The normalization method is applied in different situations according to different implementation tasks. When the corresponding task is an image classification task, the IDAD component is applied with the normalization strategy. When the corresponding task is an object detection and instance segmentation task, the BDAD component and the CDAD component are applied with the normalization strategy. The specific implementation method of this strategy is to apply the p contained in each component in step 2 to the i,t,k 、p i,\t,k and Replace with P i,t,k 、P i,\t,k and

[0040] Preferably, in step S4, the above-obtained loss of each type is combined with the classification loss. The classification loss can be obtained by using the cross entropy calculation formula. ce Combining with each type of loss gives the total training loss:

[0041] L total =λ ce L ce +λ kd (L IDAD +L BDAD +L CDAD ).

[0042] Among them, λ ce represents the classification loss weight, λ kd Represents the distillation loss weight, which is used to balance the training process.

[0043] The present invention also discloses a multi-level decoupling knowledge distillation system based on output normalization, which includes the following modules based on the above method:

[0044] Probability distribution feature extraction module: The image is input into the given teacher model and student model. The neural networks in the teacher model and student model respectively perform several convolution operations on the image, and then process it through the fully connected layer to obtain the logistic regression logit output of the model. The logit of the student model and the teacher model are respectively extracted with the multi-temperature softmax function.

[0045] Multi-temperature distillation module: This module introduces a temperature pool mechanism and designs a multi-level decoupled knowledge distillation component to perform multi-temperature distillation after extracting the probability distribution features of the teacher model and the student model.

[0046] Optimal Normalization Module: Designs an optimal normalization strategy component to assign different normalization strategies to the knowledge distillation framework based on different visual recognition tasks;

[0047] Total loss calculation module: Combines the losses calculated by the multi-level decoupling knowledge distillation components to obtain the total loss of model training.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] (1) The present invention proposes a multi-level decoupled knowledge distillation method and system based on output normalization. Under the original knowledge distillation system, the coupling characteristics of the probability distribution of the student model and the teacher model are fully considered, and three components, instance-level decoupled alignment (IDAD), batch-level decoupled alignment (BDAD), and class-level decoupled alignment (CDAD), are designed to more effectively transfer sample diversity knowledge from teachers to students, thereby improving students' ability to grasp input relevance and category relevance.

[0050] (2) In order to solve the problem of forced logit matching in the traditional logit-based KD process, the present invention proposes an optimized logit distillation preprocessing mechanism and further improves the original method to adaptively distribute logical knowledge and numerical matching knowledge, so that the present invention can better adapt to different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work. Figure 1 It is a flowchart of a multi-stage decoupling knowledge distillation method based on output normalization in a preferred embodiment of the present invention.

[0052] Figure 2 is a flow chart of a preferred normalization strategy of a preferred embodiment of the present invention.

[0053] Figure 3 It is a flow chart of the teacher neural network and the student network of the present invention outputting logistic regression logit after passing through the full link layer.

[0054] Figure 4 It is a diagram of the preferred strategy selection of the preferred normalization method of the preferred embodiment of the present invention.

[0055] Figure 5 This is a block diagram of a multi-stage decoupled knowledge distillation system based on output normalization in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0057] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0058] The preferred embodiment of the present invention discloses a multi-level decoupling knowledge distillation method based on output normalization. The process of the method is as follows: Figure 1 As shown, the preferred normalization strategy is as follows Figure 2 As shown, the following steps are included:

[0059] Step 1: Input the image into the given teacher model and student model. The neural network in the model performs several convolution operations on the image, and then processes it through the full connection layer to obtain the logistic regression logit output by the model. The process of this step is as follows: Figure 3 As shown, the logit of the student and teacher models extracts probability distribution features through the softmax function;

[0060] Step 2: After extracting the probability distribution features, a temperature pool mechanism is introduced to perform multi-temperature distillation. Specifically, after extracting the probability distributions of the teacher model and the student model, a multi-level decoupled knowledge distillation component is introduced to extract the instance-level decoupled alignment loss, batch-level decoupled alignment loss, category-level decoupled alignment loss, and classification loss to calculate the total loss.

[0061] Step S3: Design an optimal normalization strategy component and assign different normalization strategies to the knowledge distillation framework according to different visual recognition tasks;

[0062] In step S4, the losses calculated by the multi-level decoupling knowledge distillation components are combined to obtain the total loss of model training.

[0063] This embodiment uses a multi-level decoupling knowledge distillation method based on output normalization to achieve performance enhancement of image classification, target detection, and instance segmentation models. The student model enhances the knowledge distillation performance by learning diverse alignment and matching knowledge.

[0064] Specifically, in step S2, a decoupling prediction enhancement mechanism is used to expand a single decoupled output into multiple decoupled outputs. The prediction enhancement is performed through temperature calibration and the prediction decoupling is performed through the coupling characteristics of KD. The temperature calibration method can be expressed as:

[0065]

[0066] Among them, C represents the number of categories, z represents the logit output of the student model, and p i,t,k is the temperature hyperparameter T k The probability of the next i-th input sample about the target class t, p i,\t,k It is defined as the sum of the probabilities of samples on all non-target classes under the same temperature parameter. Specifically, this method treats the probability distribution of non-target classes as a differentiable competitive space while retaining the independent modeling of the target class. Its expression can be written as:

[0067]

[0068] in, The i-th sample input has a temperature hyperparameter T on the non-target category n k In this mechanism, T k (k∈[1,K]) forms a pool so that K outputs with different smoothness can be obtained.

[0069] Filter high-confidence and low-confidence samples and categories to improve the self-supervised learning efficiency of the model:

[0070]

[0071] Among them, Con i Represents the maximum predicted probability of the teacher model for sample i, Thr represents the label confidence threshold, and median() is used to calculate the 50% percentile (ie, median). The label confidence threshold Thr is used to distinguish between high confidence and low confidence predictions, and the threshold parameter m i It can be expressed as:

[0072]

[0073] Among them, m iIt is used to mark all samples whose label confidence values ​​are less than or equal to the position confidence, and filter out pseudo labels with lower confidence (i.e., the category corresponding to the maximum probability), thus achieving instance-level decoupled alignment. Specifically, this method introduces a threshold parameter to minimize the KL divergence of binary prediction enhancement between the target and non-target categories of the teacher and student models, as well as the KL divergence between the enhanced predictions of non-target categories:

[0074]

[0075] Among them, L IDAD represents the instance-level disentangled alignment loss, p tea and p stu Represents the binary outputs of target and non-target classes for teacher and student. and Represents the non-target class outputs of the teacher and student. Instance-level disentangled alignment forces the student model to imitate the teacher's predictions for each instance. Furthermore, prediction enhancement via temperature scaling is employed to transfer knowledge at multiple smoothness levels to the student model.

[0076] Instead of aligning predictions only at the instance level, we use the target class input correlation and non-target class input correlation to align instance-level predictions at the batch level. To do this, we use the Gram matrix to quantify it:

[0077]

[0078] Among them, the matrix G i,k and It is defined as a square matrix of dimension B, where scalar B represents the batch size of sample i. In addition, when the temperature T k When acting on sample i, p i,k and Quantify the model's predicted probabilities for the target class and non-target class respectively. Then, calculate the loss based on the two given matrices and introduce the self-supervised threshold parameter m i The corresponding loss can be expressed as:

[0079]

[0080] Among them, L BDAD is the batch-level decoupled alignment loss, and are the input correlation matrices of teachers and students in the target and non-target classes, respectively, and the predicted temperature T of sample i k The following calculation is obtained.

[0081] Decouple the predictions corresponding to the target class and reinforce the correct class. This class dependency can be modeled by predicting the following data:

[0082]

[0083] Among them, M i,n,k is a 2×2 matrix, It is a (C-1)×(C-1) matrix. After quantifying the category correlation, the method includes a category-specific threshold parameter to implement the self-supervision mechanism:

[0084]

[0085] Among them, con n Represents the sum of the predicted probabilities of the teacher model for all samples of class n, and thr represents the category confidence threshold. The category confidence threshold thr is used to distinguish high confidence and low confidence predictions, and the threshold parameter h i It can be expressed as:

[0086]

[0087] Among them, h i It is used to mark categories whose category confidence values ​​are less than or equal to the median confidence of all categories, and to filter out categories with low confidence, thereby improving the effect of self-supervised learning. After quantifying the category relevance and defining the category confidence threshold parameters, the student model is allowed to learn this part of knowledge through the following loss:

[0088]

[0089] Among them, L CDAD is the class-level decoupled alignment loss, and are the input correlation matrices of teachers and students in target and non-target classes, respectively, and the predicted temperature T k Calculated at sample i. Using multiple temperatures for enhanced prediction can alleviate the overconfidence phenomenon in neural networks.

[0090] Cross entropy loss calculation formula:

[0091]

[0092] Among them, y represents the true label, CE represents the cross entropy loss, and p stu represents the probability distribution of student output.

[0093] In step S3, the optimal normalization strategy of the logit can be specifically expressed in the following way. First, the function method of the normalization operation is defined:

[0094]

[0095] Here, x is the input logit and N(x) is the regularization function applied to the probability distribution of each prediction. Then p i,t,k 、p i,\t,k and can be rewritten as P i,t,k 、P i,\t,k and

[0096]

[0097] Normalization methods are applied in different situations according to different implementation tasks. Figure 4 As shown in Figure 2, when the corresponding task is an image classification task, the IDAD component is normalized. When the corresponding task is an object detection and instance segmentation task, the BDAD component and the CDAD component are normalized. The specific implementation method of this strategy is to apply the p contained in each component in step 2 to the i,t,k 、p i,\t,k and Replace with P i,t,k 、P i,\t,k and

[0098] Combine each type of loss obtained above with the classification loss. The classification loss can be calculated using the cross entropy formula.

[0099] The classification loss L ce Combining with each type of loss gives the total training loss:

[0100] L total =λ ce L ce +λ kd (L IDAD +L BDAD +L CDAD ).

[0101] Among them, λ ce represents the classification loss weight, λ kd Represents the distillation loss weight, which is used to balance the training process.

[0102] The method of the present invention is compared with the prior art in the following experiments:

[0103] Multi-level Decoupled Knowledge Distillation (MDKD) represents the method described in the present invention. "-" indicates experimental results missing from the prior art. "*" indicates experimental results used to complete comparative experiments. Image classification tasks were performed on the CIFAR-100 dataset, as shown in Table 1. Among them, FitNet, AT, RKD, CRD, OFD, CAT-KD, ReviewKD, SimKD, DPK, FAM-KD, KD, TAKD, CTKD, DKD, LSKD, WTTM, SKD, and MLKD are prior art.

[0104] Table 1: Image classification accuracy results on the CIFAR-100 dataset (%)

[0105]

[0106] The target detection task was performed on the MS-COCO2017 dataset, as shown in Table 2. ReviewKD+MDKD represents a combination of the proposed method MDKD and the prior art ReviewKD. mAP represents class recognition accuracy; AP50 represents the average detection accuracy when the IOU threshold is greater than 0.5; AP75 represents the average detection accuracy when the IOU (Intersection over Union) threshold is greater than 0.75; AP1 represents the average detection accuracy when the image is a large target; APm represents the average detection accuracy when the image is a medium-sized target; and APs represents the average detection accuracy when the image is a small target.

[0107] Table 2: Object detection accuracy results on the MS-COCO2017 dataset (%)

[0108]

[0109] The MS-COCO2017 dataset is used for instance segmentation tasks, as shown in Table III.

[0110] Table 3: Instance segmentation accuracy results on the MS-COCO2017 dataset (%)

[0111]

[0112] The above experimental results demonstrate the significant advancement of the method of the present invention. Specifically, as shown in Table 1, in image classification tasks, compared with the prior art, the distillation effect of the present invention between dissimilar category models is superior to the distillation effect between similar category models, surpassing most existing logit-based research methods, proving that the distillation effect of the present invention between different models is stronger. As shown in Tables 2 and 3, the present invention has high generalization and can also be applied to the distillation tasks of target detection and instance segmentation, and both show high performance improvement effects, which are better than existing research results.

[0113] like Figure 5 As shown, this embodiment discloses a multi-level decoupling knowledge distillation system based on output normalization, which is used to execute the above method embodiment and includes the following modules:

[0114] Probability distribution feature extraction module: The image is input into the given teacher model and student model. The neural networks in the teacher model and student model respectively perform several convolution operations on the image, and then process it through the fully connected layer to obtain the logistic regression logit output of the model. The logit of the student model and the teacher model are respectively extracted with the multi-temperature softmax function.

[0115] Multi-temperature distillation module: This module introduces a temperature pool mechanism and designs a multi-level decoupled knowledge distillation component to perform multi-temperature distillation after extracting the probability distribution features of the teacher model and the student model.

[0116] Optimal Normalization Module: Designs an optimal normalization strategy component to assign different normalization strategies to the knowledge distillation framework based on different visual recognition tasks;

[0117] Total loss calculation module: Combines the losses calculated by the multi-level decoupling knowledge distillation components to obtain the total loss of model training.

[0118] For other contents of this embodiment, please refer to the above method embodiment.

[0119] In summary, the present invention discloses a multi-level decoupling knowledge distillation method and system based on output normalization, which realizes multiple smoothness fitting of teacher model and student model based on multi-temperature knowledge distillation algorithm, performs instance-level decoupling alignment of outputs through instance-level decoupling alignment component (IDAD); performs multi-batch decoupling alignment of outputs through batch-level decoupling alignment component (BDAD). The category decoupling alignment of outputs is performed through category-level decoupling alignment component (CDAD); and the method is adapted to different image recognition tasks through the optimal normalization module. The present invention fully considers the probability distribution characteristics of the student model and the teacher model, extracts the diversified knowledge of the two, expands the decoupling characteristics of knowledge distillation, solves the problem of knowledge simplification of knowledge distillation, and greatly improves the versatility of knowledge distillation in the fields of image classification, target detection and instance segmentation.

[0120] The above are merely preferred embodiments of the present application and are not intended to limit the present application. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

[0121] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A multi-level decoupling knowledge distillation method based on output normalization, characterized by: The following steps are involved: S1: The image is input into the given teacher model and student model. The neural networks in the teacher model and the student model respectively perform several convolution operations on the image, and then process it through the fully connected layer to obtain the logistic regression logit output of the model. The logit of the student model and the teacher model respectively use the softmax function to extract the decoupled target class and non-target class probability distribution features; S2: Introducing a temperature pool mechanism, by designing a multi-level decoupled knowledge distillation component, and performing multi-temperature distillation after extracting the probability distribution features of the teacher model and the student model; S3: Design an optimal normalization strategy component to assign different normalization strategies to the knowledge distillation framework according to different visual recognition tasks; S4: Combine the losses calculated by the multi-level decoupled knowledge distillation components to obtain the total model training loss.

2. The multi-level decoupling knowledge distillation method based on output normalization according to claim 1, characterized in that: In step S2, the temperature pool mechanism is used to expand the output of a single temperature into the output of multiple temperatures, as follows: Prediction enhancement is performed through temperature calibration, and prediction decoupling is performed through the coupling characteristics of knowledge distillation; the temperature calibration method used is expressed as: Among them, C represents the number of categories, z i,t and z i,c They represent the logit of category t and c output in the i-th input sample of the student model, p i,t,k is the temperature hyperparameter T k The probability of the next i-th input sample about the target class t, p i,\t,k It is defined as the sum of the probabilities of samples on all non-target classes under the same temperature parameter; the prediction decoupling method regards the probability distribution of non-target classes as a differentiable competitive space while retaining the independent modeling of the target class. Its expression is as follows: in, The i-th sample input has a temperature hyperparameter T on the non-target category n k probability.

3. The multi-level decoupling knowledge distillation method based on output normalization according to claim 2, characterized in that: The multi-level decoupled knowledge distillation component includes the instance-level decoupling alignment component IDAD, the batch-level decoupling alignment component BDAD, the class-level decoupling alignment component CDAD and the cross-entropy loss calculation component, which are used to extract instance-level decoupling alignment loss, batch-level decoupling alignment loss, class-level decoupling alignment loss and cross-entropy loss respectively.

4. The multi-level decoupling knowledge distillation method based on output normalization according to claim 3, characterized in that: In step S2, the instance-level decoupling alignment component is designed. First, a threshold parameter m is introduced. i To minimize the KL divergence of binary prediction enhancement between target and non-target categories of the teacher and student models, as well as the KL divergence between the enhanced predictions of non-target categories: Among them, L IDAD represents the instance-level disentangled alignment loss component, m i represents the threshold parameter under sample i, KL represents KL divergence loss, B represents batch size, K represents the total number of temperatures, and C represents the number of categories. and Represents the distillation temperature T of the teacher and the student at the input sample i k The binary output of target class t and non-target class\t; and Represents the distillation temperature T of the teacher and the student at the input sample i k Non-target class output under .

5. The multi-level decoupling knowledge distillation method based on output normalization according to claim 4, characterized in that: In step S2, a batch-level decoupled alignment component is designed to align instance-level predictions at the batch level using the target class input correlation and the non-target class input correlation, and quantized using the Gram matrix: Among them, the matrix G i,k and are defined as a square matrix of dimension B, where B represents the batch size of sample i; when the temperature T k When acting on sample i, p i,k and Quantify the model's predicted probabilities for the target class and non-target class respectively; calculate the loss based on the two given matrices and introduce the self-supervised threshold parameter m i ; The corresponding loss is expressed as: Among them, L BDAD is the batch-level decoupled alignment loss component, α and β represent the hyperparameters of the target class and non-target class, respectively. and are the input correlation matrices of teachers and students in the target and non-target classes, respectively, and the predicted temperature T of sample i k The following calculation is obtained.

6. The multi-level decoupling knowledge distillation method based on output normalization according to claim 5, characterized in that: In step S2, a class-level decoupling alignment component is designed; predictions corresponding to the target class are decoupled and reinforced for the correct class; and the model is built by predicting the following data: Among them, M i,n,k is the transpose of the binary probability distribution of the target class and itself i,n,k The 2×2 matrix obtained by multiplication is is the transpose of the probability distribution of the non-target class and itself The (C-1)×(C-1) matrix obtained by multiplication; Introducing category confidence threshold parameter h i , and let the student model learn this part of knowledge through the following loss: Among them, L CDAD Decouple the alignment loss component for the class level, h i represents the threshold parameter under sample i, and are the input correlation matrices for target and non-target classes and teacher and student, respectively, and the predicted temperature T k Calculated at sample i.

7. The multi-level decoupling knowledge distillation method based on output normalization according to claim 6, characterized in that: In step S2, the cross entropy loss calculation formula is: Among them, y represents the true label, CE represents the cross entropy loss, and p stu represents the probability distribution of student output.

8. The multi-level decoupling knowledge distillation method based on output normalization according to claim 7, characterized in that: In step S3, a mechanism for prioritizing normalization of logit is proposed; first, the function method for normalization operation is defined: Where x is the input logit, N(x) is the regularization function, and the regularization function is applied to the probability distribution of each prediction; then p i,t,k 、p i,\t,k and Rewrite them as P i,t,k 、P i,\t,k and When the corresponding task is an image classification task, the IDAD component is applied with a normalization strategy; when the corresponding task is an object detection and instance segmentation task, the BDAD component and the CDAD component are applied with a normalization strategy; the normalization strategy applies the p contained in the components in step 2 to the i,t,k 、p i,\t,k and Replace with P i,t,k 、P i,\t,k and 9. The multi-level decoupling knowledge distillation method based on output normalization according to claim 8, characterized in that: In step S4, the losses are combined to obtain the total training loss: L total =λ ce L ce +λ kd (L IDAD +L BDAD +L CDAD ). Among them, λ ce represents the classification loss weight, λ kd represents the distillation loss weight.

10. A multi-level decoupled knowledge distillation system based on output normalization, for executing the method according to any one of claims 1 to 9, characterized in that: Includes the following modules: Probability distribution feature extraction module: The image is input into the given teacher model and student model. The neural networks in the teacher model and student model respectively perform several convolution operations on the image, and then process it through the fully connected layer to obtain the logistic regression logit output of the model. The logit of the student model and the teacher model are respectively extracted with the multi-temperature softmax function. Multi-temperature distillation module: This module introduces a temperature pool mechanism and designs a multi-level decoupled knowledge distillation component to perform multi-temperature distillation after extracting the probability distribution features of the teacher model and the student model. Optimal Normalization Module: Designs an optimal normalization strategy component to assign different normalization strategies to the knowledge distillation framework based on different visual recognition tasks; Total loss calculation module: Combines the losses calculated by the multi-level decoupling knowledge distillation components to obtain the total loss of model training.