Knowledge distillation method and device, electronic equipment and computer storage medium
By introducing an auxiliary output head with the same structural capacity as the student model output head in knowledge distillation and using the prediction results of the teacher model as soft labels, the problem of feature map differences between the teacher and student models is solved, achieving efficient knowledge distillation and improving the performance of the student model.
Patent Information
- Application Number
- CN202510875973.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-10
AI Technical Summary
Due to the large differences in structure and parameter capacity between the teacher model and the student model, the feature maps of the teacher model and the student model have large differences in feature dimension, semantics and amplitude, which affects the effect of knowledge distillation and thus affects the processing effect of the trained student model.
An auxiliary output head is introduced to make it consistent with the output head structure capacity of the student model, and the prediction results output by the output head of the teacher model are used as soft labels for the auxiliary prediction results. The knowledge distillation of the student model is performed by calculating the distillation loss.
It avoids the differences in feature maps between the teacher model and the student model, ensures the effectiveness and efficiency of knowledge distillation, and improves the performance of the student model on various tasks.
Smart Images

Figure CN120766006A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a knowledge distillation method and device, an electronic device and a computer storage medium. BACKGROUND
[0002] In the technical field of image processing, in order to improve the training efficiency of an image processing model, a trained model meeting the expected processing effect is usually used as a teacher model, a model to be trained is used as a student model, and the student model is trained by using a knowledge distillation method to improve the performance of the student model.
[0003] Due to the large difference in structure and parameter capacity between the teacher model and the student model, the feature maps of the teacher model and the student model have large differences in feature dimension, semantics and amplitude, which affects the effect of knowledge distillation and further affects the processing effect of the trained student model. SUMMARY
[0004] The present application aims to provide a knowledge distillation method and device, an electronic device and a computer storage medium, which can avoid the difference between the feature maps of the teacher model and the student model in knowledge distillation while ensuring the effect of knowledge distillation, so as to ensure the processing efficiency and processing effect when using the trained student model for image processing.
[0005] Embodiments of the present application can be implemented as follows: In a first aspect, the present application provides a knowledge distillation method, comprising: obtaining a sample image; inputting the sample image into a pre-constructed student model and a trained teacher model respectively to obtain an intermediate feature map, a first prediction result and a second prediction result, the intermediate feature map being obtained by feature extraction of the student model, the first prediction result being output by a student output head of the student model, and the second prediction result being output by a teacher output head of the teacher model; inputting the intermediate feature map into a pre-constructed auxiliary output head to obtain an auxiliary prediction result, the auxiliary output head and the student output head having the same structure capacity; performing knowledge distillation on the student model based on the first prediction result and taking the second prediction result as a soft label of the auxiliary prediction result to obtain a trained student model.
[0006] In an optional embodiment, the step of performing knowledge distillation on the student model based on the first prediction result and taking the second prediction result as a soft label of the auxiliary prediction result to obtain a trained student model comprises: Calculating a distillation loss based on the soft label and the auxiliary prediction result; Calculating a main loss based on the first prediction result and the label of the sample image; Performing knowledge distillation on the student model according to the distillation loss and the main loss to obtain a trained student model.
[0007] In an optional embodiment, the student output head is used to perform a mask segmentation task, and the step of calculating the distillation loss according to the soft label and the auxiliary prediction result includes: Obtaining a probability distribution of the soft label and a probability distribution of the auxiliary prediction result; Calculate the forward KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain the global distillation loss; Calculating the inverse KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain a local distillation loss; The distillation loss is calculated according to the global distillation loss and the local distillation loss.
[0008] In an optional embodiment, the student output head is used to perform a classification task, and the step of calculating the distillation loss according to the soft label and the auxiliary prediction result includes: A binary cross entropy between the soft label and the auxiliary prediction result is calculated to obtain the distillation loss.
[0009] In an optional embodiment, the student output head is used to perform a positioning task, and the step of calculating the distillation loss according to the soft label and the auxiliary prediction result includes: The complete intersection-over-union (IoU) between the soft label and the auxiliary prediction result is calculated to obtain the distillation loss.
[0010] In an optional embodiment, the student output head is used as a classification output head for a classification task and a positioning output head for a positioning task, and the method further includes: Obtain the image to be detected; Inputting the image to be detected into the trained student model to obtain a first classification result output by the classification output head and a first positioning result output by the positioning output head; Target detection is performed on the image to be detected according to the first classification result and the first positioning result.
[0011] In an optional embodiment, the student output head includes a prototype head, a classification output head for performing a classification task, a positioning output head for performing a positioning task, and a segmentation output head for performing a mask segmentation task, and the method further includes: Obtain the image to be segmented; Inputting the image to be segmented into the trained student model to obtain the prototype mask output by the prototype head, the second classification result output by the classification output head, the second positioning result output by the positioning output head, and the mask coefficient output by the segmentation output head; Performing instance detection on the image to be segmented according to the second classification result and the second positioning result to obtain a detected target instance; The target instance is segmented according to the prototype mask and the mask coefficient to obtain an instance mask of the target instance.
[0012] In a second aspect, the present invention provides a knowledge distillation device, comprising: An acquisition module, used for acquiring a sample image; a distillation module, configured to input the sample image into a pre-built student model and a trained teacher model, respectively, to obtain an intermediate feature map, a first prediction result, and a second prediction result, wherein the intermediate feature map is obtained by feature extraction by the student model, the first prediction result is output by the student output head of the student model, and the second prediction result is output by the teacher output head of the teacher model; The distillation module is further configured to input the intermediate feature map into a pre-built auxiliary output head to obtain an auxiliary prediction result, wherein the auxiliary output head and the student output head have the same structural capacity; The distillation module is further used to perform knowledge distillation on the student model based on the first prediction result and use the second prediction result as a soft label for the auxiliary prediction result to obtain a trained student model.
[0013] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a program, and the processor is used to implement the knowledge distillation method as described in the first aspect when executing the program.
[0014] In a fourth aspect, the present invention provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the knowledge distillation method as described in the first aspect.
[0015] Compared with the prior art, the present invention has the following beneficial effects: In this example of the present invention, an auxiliary output head is pre-constructed, and the intermediate feature map output by the student model is used as the input of the auxiliary output head to obtain an auxiliary prediction result. Since the structural capacity of the auxiliary output head and the student output head are consistent, the difference in the feature maps of the teacher model and the student model is avoided, thereby ensuring the effect of knowledge distillation. At the same time, the second prediction result output by the teacher output head of the teacher model is used as the soft label of the auxiliary prediction result, thereby fully learning the knowledge of the teacher model and ensuring the efficiency of knowledge distillation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is an example diagram of the output results of the real-time instance segmentation task provided in this embodiment.
[0018] Figure 2 This is an example diagram of the real-time instance segmentation model framework based on YOLOv8 provided in this embodiment.
[0019] Figure 3 This is a diagram illustrating an example of the structure of the existing knowledge distillation method provided in this embodiment.
[0020] Figure 4 This is an example diagram of the structure of the knowledge distillation method provided in this embodiment.
[0021] Figure 5 This is an example flow chart of the knowledge distillation method provided in this embodiment.
[0022] Figure 6 This is a block diagram of an example of the knowledge distillation device provided in this embodiment.
[0023] Figure 7 This is a block diagram of an example of an electronic device provided in this embodiment.
[0024] Icons: 10-electronic device; 11-processor; 12-memory; 13-bus; 100-knowledge distillation device; 110-acquisition module; 120-distillation module; 130-target detection module; 140-instance segmentation module. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0028] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or is the orientation or position relationship in which the product of the invention is usually placed when in use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be understood as a limitation on the present invention.
[0029] In addition, the terms "first", "second", etc., if used, are merely used to distinguish and describe, and should not be understood as indicating or implying relative importance.
[0030] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.
[0031] Instance segmentation is a very important application branch in the field of image processing technology. It requires accurate pixel-level segmentation of each object in the image and outputs classification labels and bounding boxes at the same time. Figure 1 , Figure 1 This is an example diagram of the output results of the real-time instance segmentation task provided in this embodiment. Figure 1In the figure, two instances are segmented: the instance corresponding to the classification label "bicycle", with a probability of 0.63, and the instance corresponding to the classification label "fire hydrant", with a probability of 0.88. In addition to the classification label, each instance also corresponds to a bounding box, as shown in the two bright blue rectangles in the figure. Inside each rectangle, an example mask of the corresponding instance is also displayed. Therefore, real-time segmentation of the corresponding instance can be achieved based on the classification label, bounding box, and instance mask of each instance.
[0032] To achieve this in real-time applications, complex deep learning models are often used. Figure 2 , Figure 2 This diagram shows an example of a real-time instance segmentation model framework based on YOLOv8. YOLOv8 supports a variety of tasks, including but not limited to: object detection (identifying objects and their locations in images); instance segmentation (not only detecting objects but also providing pixel-level segmentation masks); pose estimation (detecting key points of people or other objects); and classification (supporting separate image classification tasks). This multi-task support allows users to complete different computer vision tasks using the same code and model structure, significantly reducing development costs.
[0033] Figure 2 In , the output head includes a decoupling head and a prototype head. Each decoupling head includes a classification head for classifying the identified targets, a localization head for locating the identified targets, and a segmentation head for generating mask coefficients of the identified targets (also called instances). The prototype head is used to output prototype masks. The prototype masks can capture various shapes and structural information in the image and can be regarded as a shared knowledge representation. The image is first input into the backbone network (Backbone), where "1" in (1,3,640,640) represents the number of images, "3" represents the number of channels of the image, and "640,640" represent the height and width of the image respectively. The backbone network is responsible for extracting multi-scale features from the input image. After the multi-scale features extracted by the backbone network enter the neck network (Neck) for fusion and enhancement, feature maps of different scales and different numbers of channels are obtained: P5, P4, and P3, where the number of channels of P5 is 512 and the size is (20,20), and P4 and P3 are shown in Figure 2. Figure 2The corresponding annotations are shown in Figure 2. P5, P4, and P3 each input the three decoupling heads. The final classification score is the concatenation of the outputs of the three classification heads, the final bounding box is the concatenation of the outputs of the three localization heads, and the final mask coefficients are the concatenation of the outputs of the three segmentation heads. The classification score (1, nc, h, w) indicates that the identified object category is nc and the object size is (h, w). The bounding box (1, 4*reg_max, h, w) indicates the four vertex positions (4*reg_max) of the identified object and the size of the bounding box is (h, w). The mask coefficients (1, 32, h*w) indicate that the number of channels is 32 and the size is (h, w). Furthermore, to obtain the instance mask, the largest feature map P3 output by the neck network is also input to the prototype head. The prototype head generates a prototype mask based on this, encoding general knowledge related to the segmentation task. The segmentation head predicts the mask coefficients for each instance. Finally, the mask coefficients are linearly combined with the prototype mask. The final instance mask is obtained by cropping the bounding box and binarizing it.
[0034] However, models such as YOLOv8 are often computationally intensive and time-consuming, making them difficult to meet real-time requirements. Therefore, knowledge distillation technology was introduced to improve the performance of a small, simple model (student model) by transferring knowledge from a large, complex model (the teacher model).
[0035] There are generally two approaches to knowledge distillation for training student models: one based on the feature distillation paradigm, and the other based on the prediction distillation paradigm. The former transfers mixed knowledge across all tasks by aligning the feature maps of the teacher and student models, thereby improving overall performance. However, this coupled information transfer is not the most efficient and reasonable approach. The latter, based on decoupled prediction information for distillation, can efficiently and reasonably transfer task-specific information, making it a knowledge distillation approach more suitable for real-time instance segmentation. However, the existing prediction-based distillation paradigm uses an allocator for dynamic label allocation, which conflicts with the static teacher prediction labels, affecting the efficiency of knowledge distillation.
[0036] To alleviate the problem of target conflict, one way of knowledge distillation is to pass the intermediate features of the student output head to the teacher output head, and then force the obtained cross-head predictions to imitate the predictions of the teacher model. Please refer to Figure 3 , Figure 3 This is a diagram showing an example of the structure of the existing knowledge distillation method provided in this embodiment. Figure 3 In the student model (e.g. Figure 3 The yellow box in the figure) and the teacher model (shown as Figure 3The blue box in the figure shows that both are backbone networks + feature pyramids. The scale and specific structure of the student model and the teacher model are different. The teacher model is already trained and its model parameters are fixed. The main task loss is calculated by the true label of the input sample, and the intermediate features of the output head of the student model are directly passed to the output head of the teacher model with fixed weights, and we get , the distillation loss of the knowledge distillation task is based on And the output of the teacher model output This approach alleviates the conflict of objectives to a certain extent, but it also introduces the feature alignment problem that exists with feature distillation. However, the teacher and student models differ significantly in their architecture and parameter capacity. The intermediate features output by the student model's output head do not match those output by the teacher model's output head. Therefore, this approach introduces the feature alignment problem.
[0037] In light of this, this embodiment provides a knowledge distillation method, apparatus, electronic device, and computer storage medium. The core improvement lies in the introduction of an auxiliary output head. The auxiliary output head has the same structural capacity as the output head of the student model, mitigating differences in the feature maps of the teacher and student models. Furthermore, the prediction results output by the teacher model's output head serve as soft labels for the auxiliary prediction results. This allows for full learning of the teacher model's knowledge and ensures efficient knowledge distillation. This is described in detail below.
[0038] For the sake of convenience, this embodiment provides a method based on Figure 3 For an improved example of the structure, please refer to Figure 4 , Figure 4 This is an example diagram of the structure of the knowledge distillation method provided in this embodiment. Figure 4 The student model and teacher model in Figure 3 Similar to, but different from, Figure 4 An auxiliary output head is introduced. The input of the auxiliary output head is the intermediate features output by the student model. Since the structure capacity of the auxiliary output head and the output head of the student model are consistent, there is no feature mismatch problem in the intermediate features of the auxiliary output head. The main task loss is based on the prediction results output by the student model. And the true label gt of the input sample is calculated. The distillation loss of the knowledge distillation task is to convert As a soft label, the prediction result output by the soft label and the auxiliary output head is The calculation solves the problem of conflicting goals between dynamic label assignment and static teacher prediction labels, which can fully learn the knowledge of the teacher model and improve the efficiency of knowledge distillation.
[0039] exist Figure 4 Based on this, the following describes the process of knowledge distillation method, please refer to Figure 5 , Figure 5 This is an example flow chart of the knowledge distillation method provided in this embodiment. Figure 5 In [1], the knowledge distillation method includes the following steps: Step S101: Acquire a sample image.
[0040] In this embodiment, sample images provide an input data source for the student model and the teacher model. For a student model solving an instance segmentation task, the sample images can be image samples from a dataset used to train the student model for instance segmentation tasks. These images typically contain multiple instance objects, and each sample image typically has a corresponding true label. When a sample image includes multiple instance objects, its corresponding true label includes annotation information corresponding to each instance object, such as a classification label, bounding box coordinates, and mask coefficients.
[0041] The sample images must meet the input data requirements of both the student and teacher models, such as image resolution, number of color channels (e.g., RGB channels), and other preprocessing specifications. Furthermore, to ensure the effectiveness of the knowledge distillation process, the sample images should cover a wide range of possible scenarios and object types in object detection and instance segmentation tasks, allowing the teacher model to fully exploit the knowledge and transfer it to the student model.
[0042] Since the quality and diversity of sample images directly affect the learning effect of the subsequent student model and the final distillation performance, in practical applications, necessary preprocessing operations such as normalization, resizing, or data augmentation can also be performed on the sample images to improve the generalization ability and robustness of the model.
[0043] In step S102, the sample image is input into the pre-built student model and the trained teacher model respectively to obtain an intermediate feature map, a first prediction result, and a second prediction result. The intermediate feature map is obtained by feature extraction by the student model, the first prediction result is output by the student output head of the student model, and the second prediction result is output by the teacher output head of the teacher model.
[0044] In this embodiment, the sample image is used as input data and fed into the student model and the teacher model for processing. The student model contains a backbone network structure for feature extraction, such as the CSPDarknet backbone network in YOLOv8-seg. When the sample image enters the student model, the backbone network performs multi-level feature extraction operations on it to generate intermediate feature maps. These intermediate feature maps not only contain the spatial information of the sample image, but also integrate multi-scale semantic information.
[0045] The intermediate feature maps are then passed to the student model's output head, which ultimately outputs the first prediction. Simultaneously, the sample image is fed into the trained teacher model. Because the teacher model typically has a larger parameter size and stronger performance, it can generate a more accurate second prediction. This second prediction serves as the learning objective for the student model during the knowledge distillation process.
[0046] In step S103, the intermediate feature map is input into the pre-built auxiliary output head to obtain an auxiliary prediction result. The structure capacity of the auxiliary output head is consistent with that of the student output head.
[0047] In this embodiment, structural capacity characterizes the capabilities of the output head and determines the complexity of the tasks it can handle. This includes at least one of the following aspects: the number of network layers, the number of channels, the number of parameters, and the computational complexity. The number of network layers can be the number of convolutional or fully connected layers in the output head; the number of channels can be the number of convolution kernels or the dimension of hidden units in each layer; the number of parameters can be the sum of the trainable parameters in the entire output head module; and the computational complexity can be the computational resources required by the output head for forward and backward propagation. Different output head designs are typically required for different types of tasks (such as classification, localization, and segmentation). For example, a classification task may only require a simple fully connected layer, while a segmentation task may require complex convolution operations and upsampling.
[0048] In this embodiment, as a specific implementation method, the student output header can be directly copied and the copied copy can be used as the auxiliary output header.
[0049] In step S104, knowledge distillation is performed on the student model based on the first prediction result and using the second prediction result as a soft label for the auxiliary prediction result to obtain a trained student model.
[0050] In this embodiment, although the student model and the teacher model receive the same sample image input, their internal feature extraction methods and the presentation of their prediction results differ due to differences in their network structures and parameter capacities. However, this difference is precisely the core value of knowledge distillation. Therefore, using the second prediction result output by the teacher output head as a soft label for the auxiliary prediction result output by the auxiliary output head can efficiently transfer the knowledge of the teacher model to the student model, improving the latter's performance on various tasks.
[0051] The above method provided in this embodiment avoids the problem of feature mismatch in the intermediate features of the auxiliary output head by constructing an auxiliary output head with the same structural capacity as the student output head. The prediction results of the teacher output head are used as soft labels to calculate the distillation loss, which solves the problem of target conflict between dynamic label allocation and static teacher prediction labels. It can fully learn the knowledge of the teacher model and improve the efficiency of knowledge distillation.
[0052] In this embodiment, in order to more reasonably adjust the model parameters of the student model during the training process, the gap between the prediction result and the reference label, that is, the loss, is usually calculated. In this embodiment, the loss includes two parts: the main loss and the distillation loss. The main loss is calculated based on the first prediction result and the label of the sample image. The distillation loss uses the second prediction result as the soft label of the auxiliary prediction result and is calculated based on the soft label and the auxiliary prediction result. A specific implementation method is as follows: First, the distillation loss is calculated based on the soft labels and auxiliary prediction results; In this embodiment, soft labels refer to the secondary predictions generated by the teacher model's output head for the sample image. These secondary predictions, after probabilistic processing, reflect the teacher model's detailed prediction distribution for various tasks (such as classification, localization, or segmentation) in the sample image, rather than simply the final hard decision. By comparing the soft labels with the auxiliary predictions and calculating the corresponding distillation loss, the student model can inherit more detailed task-related knowledge from the teacher model, thereby improving its own performance.
[0053] Soft labels are represented differently for different tasks. For example, for classification tasks, soft labels can be represented as a vector where each element corresponds to the probability of a class prediction. For localization tasks, soft labels may be represented as a continuous distribution of bounding box coordinates. For segmentation tasks, soft labels can be represented as a probability distribution of prototype masks.
[0054] Secondly, the main loss is calculated based on the first prediction result and the label of the sample image; In this embodiment, as an implementation method, the main loss and the distillation loss can be calculated in different ways. For example, the two correspond to their own loss functions, and their respective losses are calculated through their respective loss functions. The loss functions of the two can be different.
[0055] Finally, knowledge distillation is performed on the student model according to the distillation loss and the main loss to obtain the trained student model.
[0056] In this embodiment, the distillation loss and the main loss together constitute the total loss used to train the student model. Taking the teacher-student model pair based on YOLOv8-seg, the second prediction result and the first prediction result are respectively recorded as and , the auxiliary output head generates the auxiliary prediction result and is recorded as , the mathematical expression of the total loss can be:
[0057] in, is the total loss, and Represent the region selection criterion and normalization factor respectively. The region selection criterion is used to filter out the "important" parts of the teacher model's predictions and guide the student model to focus on learning these parts. In knowledge distillation, the teacher model's predictions may contain a lot of information, but not all of this information is helpful for improving the student model's performance. For example, in a classification task, the background area may contribute less to target detection, so these unimportant areas can be ignored through the region selection criterion. In order to simplify the calculation, this embodiment does not design a complex , but rather on the entire prediction graph (i.e. the entire domain of prediction results) and Specifically, is a constant function that is always equal to 1. The normalization factor is used to balance the contributions of different tasks, ensuring that the various losses during the knowledge distillation process are not unbalanced due to large numerical differences. In knowledge distillation, different types of tasks (such as classification, localization, and segmentation) may produce predictions or losses of varying sizes. Without an appropriate normalization factor, a single task may dominate the overall loss due to its larger value, neglecting the learning effects of other tasks. Represents the entire prediction graph, Indicates the part of the prediction graph that needs attention. In this embodiment, in order to simplify the calculation, that is . is the loss function corresponding to the task type, Indicates The auxiliary prediction results on Indicates The second prediction result above.
[0058] In this embodiment, according to the different types of tasks implemented by the student model, the student output head may include output heads for different types of tasks. For example, the classification task corresponds to the classification output head, the positioning task corresponds to the positioning output head, and the segmentation task corresponds to the segmentation output head. In order to make the calculation of the distillation loss more targeted, the method of calculating the distillation loss for different types of tasks is also different. For the above formula, according to different task types, adopt different forms to more effectively transfer task-specific knowledge to the student model.
[0059] The following describes how to calculate distillation loss for three different types of tasks.
[0060] For the classification task, since the student output head includes a classification output head for the classification task, the auxiliary output head also includes a classification output head with the same structure and capacity. The distillation loss can be calculated as follows: The binary cross entropy between the soft labels and the auxiliary prediction results is calculated to obtain the distillation loss.
[0061] As a way of expressing It can be expressed as: ,in, Represents the activation function, for example, the activation function can be the Sigmoid function, represents the binary cross entropy function, represents the auxiliary prediction result, Indicates a soft label.
[0062] For the positioning task, since the student output head includes a positioning output head for the positioning task, the auxiliary output head also includes a positioning output head with the same structure and capacity. The method for calculating the distillation loss can be: The distillation loss is obtained by calculating the complete intersection-over-union between the soft labels and the auxiliary prediction results.
[0063] As a way of expressing It can be expressed as: ,in, Represents the complete intersection-union ratio, which is used to measure the degree of overlap between the auxiliary prediction results and the soft labels. represents the auxiliary prediction result, Indicates a soft label.
[0064] For the segmentation task, the student output head includes the prototype head and the segmentation output head for the mask segmentation task. Therefore, the auxiliary output head also includes the prototype head and the segmentation output head with the same structural capacity. The distillation loss can be calculated as follows: First, obtain the probability distribution of soft labels and the probability distribution of auxiliary prediction results; Secondly, the forward KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result is calculated to obtain the global distillation loss; In this embodiment, the forward KL divergence is also called FKL. Its goal is to make the probability distribution of the auxiliary prediction result as close as possible to the probability distribution of the soft label. FKL is suitable for capturing global information or significant features, such as high-response areas such as geometric shapes and spatial distributions in the teacher model.
[0065] Third, calculate the inverse KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain the local distillation loss; In this embodiment, the reverse KL divergence is also called RKL. Its goal is to make the probability distribution of the auxiliary prediction result guide the probability distribution of the soft label as much as possible. RKL is suitable for completing detailed information or local features, such as low-response areas such as texture and contour in the teacher model.
[0066] Finally, the distillation loss is calculated based on the global distillation loss and the local distillation loss.
[0067] In this example, the global distillation loss can be expressed as:
[0068] in, represents the global distillation loss, Indicates the number of channels, represents the temperature coefficient, Indicates the channels, and Respectively represent the height and width of the image output by the prototype head, The first image in the prototype head output pixels, The prototype mask of the image output by the prototype head is shared, where the prototype mask of the image output by the prototype head of the teacher model and the prototype mask of the image output by the prototype head of the auxiliary output head are recorded as and , The image output by the prototype head of the teacher model shares the prototype mask in the channel No. The characteristic value of each pixel, Express The result after normalization is Represents a function that converts the eigenvalue into a probability distribution divided by the channel to highlight the significant area in the channel. The image output by the prototype head of the auxiliary output head shares the prototype mask in channel No. The characteristic value of each pixel, Indicates use Convolution pair Perform linear mapping and normalize the result of linear mapping. Convolution pair The purpose of linear mapping is to adjust Channel semantics, implementation and The purpose of normalization is to help alleviate the amplitude difference of the image output by the prototype head and improve the distribution consistency.
[0069] The global distillation loss prioritizes the alignment of high-response regions in the prototype head of the teacher model to learn global information such as the geometry and spatial distribution of the teacher prototype.
[0070] The local distillation loss can be expressed as:
[0071] in, represents the local distillation loss, where the letters have the same meaning as in the global distillation loss formula. Local distillation loss prioritizes aligning low-response regions, which typically contain local information such as instance-specific texture and contours. Aligning low-response regions helps complete the detailed representation of semantic features, thereby improving the model's generalization and segmentation accuracy in complex scenarios.
[0072] The distillation loss can be expressed as:
[0073] in, Represents distillation loss.
[0074] In order to make the above global distillation loss, local distillation loss and the global distillation loss and local distillation loss have a basis, the derivation process of the above formula is introduced below.
[0075] For the global distillation loss, let The prototype mask is shared by the image output by the prototype head, where the output of the prototype head of the teacher model and the output of the prototype head of the auxiliary output head are respectively recorded as and , in, and W represent the height and width of the prototype mask shared by the prototype head output image, respectively, and It indicates that the image output by the prototype head shares the same number of channels as the prototype mask. and All have the same size ,However, due to the differences in the structure and capacity of the prototype heads, ,directly comparing the teacher prototype and the student prototype faces the ,problem of inconsistent feature semantics and amplitude.
[0076] In order to achieve and Semantic alignment using 1 Convolution pair Perform linear mapping, the mathematical expression is:
[0077] in, For The result of linear mapping is express 1 convolution kernel weight, is the bias vector, is a standard convolution operation. This linear transformation adjusts Channel semantics, mitigation and The problem of inconsistent feature semantics.
[0078] To alleviate and The difference in the magnitude of Normalize it to have zero mean and unit variance and follow the convolution characteristics, that is, the elements of the same channel in different spatial positions in the same feature map are normalized in the same way. Define the normalization function:
[0079] in and Input Then, and Normalize them separately and get:
[0080] This unified normalization operation helps alleviate prototype amplitude differences and improve prototype distribution consistency.
[0081] To achieve accurate alignment of semantic information, it is necessary to design an alignment loss that focuses on semantics rather than spatial information.
[0082] First, for any channel Feature activation on , construct probability distribution:
[0083] in, Indicates channel In spatial position The eigenvalues of is the corresponding temperature coefficient.
[0084] Then, based on the softened prototype probability distribution, FKL (Forward Kullback-Leibler Divergence) is used as the global distillation loss:
[0085] Consider the definition in the common sample space Two probability distributions on: teacher distribution and student distribution , its FKL (Forward Kullback-Leibler Divergence) is defined as:
[0086] In this formula, As regards student distribution function, and assuming that the student distribution is the student model parameter A differentiable function of Therefore, we can Calculate the gradient to guide the optimization process of the model. Find the partial derivative and we get:
[0087] because Does not depend on parameters or student distribution , so its partial derivative is zero, and the above formula is simplified to:
[0088] In order to avoid the situation where the denominator is zero or the gradient overflows, in actual optimization, Introducing a minimum cutoff constant (For example ), considering FKL relative to the student distribution The absolute value of the gradient:
[0089] Don't consider it and situation, The value range is ( , 1), then:
[0090] The above Replace with , Replace with , and we get the formula for the global distillation loss mentioned above.
[0091] Based on the above characteristics, FKL will provide greater gradient supervision to the high response areas in the teacher prototype distribution, give priority to matching the high response areas in the teacher prototype, and assign higher probability to the low response areas. In other words, under the constrained optimization of FKL, Tend to Higher probabilities are assigned to multiple peaks, thus forming the "mean-seeking" characteristic of FKL.
[0092] Based on the above analysis, in order to synchronously align the low-response areas in the teacher prototype distribution, RKL (Reverse Kullback-Leibler Divergence) is considered as the local distillation loss:
[0093] Consider RKL relative to student distribution The absolute value of the gradient is obtained by using the product derivative rule:
[0094]
[0095] This means that when Smaller hour, should also be small (i.e. not too large a value of 1). When the distribution is Gaussian, RKL tends to cover only This is the "mode-seeking" feature of RKL.
[0096] Based on the above analysis, the FKL of the global loss function is replaced by the RKL of the local loss function, and the local distillation loss function is designed:
[0097] Finally, the global distillation loss and the local distillation loss are combined to obtain the distillation loss:
[0098] The weight term of the global distillation loss is , under the synergistic effect of global distillation loss and local distillation loss, its weight term is , shifting the focus of the global distillation loss from the high-response area to the area where the prototype head of the teacher model and the prototype head of the student model are inconsistent, thereby significantly improving the overall performance of the segmentation head of the student model.
[0099] In this embodiment, the student output head may include different types of output heads according to different processing requirements. For target detection requirements, the student output head may include a classification output head for classification tasks and a positioning output head for positioning tasks. After obtaining the trained student model, one implementation method for target detection using the trained student model may be: First, obtain the image to be detected; Secondly, the image to be detected is input into the trained student model to obtain the first classification result output by the classification output head and the first positioning result output by the positioning output head; In this embodiment, the first classification result may include a category and a classification score belonging to the category, and the first positioning result may include the predicted target box coordinates or the offset of the key point.
[0100] Third, target detection is performed in the image to be detected based on the first classification result and the first positioning result.
[0101] For instance segmentation requirements, the student output head can include a prototype head, a classification output head for classification tasks, a localization output head for localization tasks, and a segmentation output head for mask segmentation tasks. After obtaining the trained student model, one way to implement instance segmentation using the trained student model can be: First, obtain the image to be segmented; Secondly, the image to be segmented is input into the trained student model to obtain the prototype mask output by the prototype head, the second classification result output by the classification output head, the second positioning result output by the positioning output head, and the mask coefficient output by the segmentation output head; In this embodiment, the second classification result may include a category and a classification score belonging to the category, and the second positioning result may include the predicted target box coordinates or the offset of the key point.
[0102] The prototype head is responsible for generating a set of universal prototype masks that capture various shapes and structural information in the image and can be regarded as a shared knowledge representation. The segmentation output head generates corresponding mask coefficients for each detected instance. These coefficients reflect the relationship between the specific instance and the prototype mask. By linearly combining the prototype mask and the mask coefficient, the final mask of the instance can be obtained. The role of the mask coefficient is to interact with the prototype mask through linear combination to generate a specific mask for each target instance. In this embodiment, the mask coefficient plays a key role in connecting universal knowledge (prototype mask) with instance-specific information, ensuring that the instance segmentation task can maintain high accuracy while having good generalization ability.
[0103] Third, instance detection is performed in the image to be segmented based on the second classification result and the second positioning result to obtain the detected target instance; In this embodiment, the target instance detection process is similar to the aforementioned target detection process.
[0104] Finally, the target instance is segmented according to the prototype mask and the mask coefficient to obtain the instance mask of the target instance.
[0105] In this embodiment, with respect to the prototype mask and the mask coefficient of each target instance, each target instance may be segmented to obtain an instance mask of each target instance.
[0106] In order to execute the corresponding steps in the above embodiment and each possible implementation method, a method for implementing the knowledge distillation device 100 is given below. Figure 6 , Figure 6 This is a block diagram of the knowledge distillation device provided in this embodiment. It should be noted that the basic principles and technical effects of the knowledge distillation device 100 provided by the present invention are the same as those of the corresponding above-mentioned embodiments. For the sake of brief description, they are not mentioned in this embodiment.
[0107] The knowledge distillation device 100 includes an acquisition module 110 , a distillation module 120 , an object detection module 130 and an instance segmentation module 140 .
[0108] An acquisition module 110 is configured to acquire a sample image; A distillation module 120 is configured to input a sample image into a pre-built student model and a trained teacher model, respectively, to obtain an intermediate feature map, a first prediction result, and a second prediction result. The intermediate feature map is obtained by feature extraction by the student model, the first prediction result is output by the student output head of the student model, and the second prediction result is output by the teacher output head of the teacher model. The distillation module 120 is further configured to input the intermediate feature map into a pre-built auxiliary output head to obtain an auxiliary prediction result, wherein the auxiliary output head and the student output head have the same structural capacity; The distillation module 120 is further configured to perform knowledge distillation on the student model based on the first prediction result and using the second prediction result as a soft label for the auxiliary prediction result to obtain a trained student model.
[0109] In an optional embodiment, the distillation module 120 is specifically configured to: Calculate the distillation loss based on the soft labels and auxiliary prediction results; Calculate the main loss based on the first prediction result and the label of the sample image; According to the distillation loss and the main loss, the student model is subjected to knowledge distillation to obtain the trained student model.
[0110] In an optional implementation, the student output head is used for a mask segmentation task, and the distillation module 120 is specifically used for: obtaining a probability distribution of the soft label and a probability distribution of the auxiliary prediction result; calculating a forward KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain a global distillation loss; calculating a reverse KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain a local distillation loss; calculating the distillation loss according to the global distillation loss and the local distillation loss.
[0111] In an optional implementation, the student output head is used for a classification task, and the distillation module 120 is specifically used for: calculating a binary cross-entropy between the soft label and the auxiliary prediction result to obtain the distillation loss.
[0112] In an optional implementation, the student output head is used for a positioning task, and the distillation module 120 is specifically used for: calculating a complete intersection-over-union between the soft label and the auxiliary prediction result to obtain the distillation loss.
[0113] In an optional implementation, the student output head includes a classification output head used for a classification task and a positioning output head used for a positioning task, and the knowledge distillation apparatus 100 further includes a target detection module 130, which is specifically used for: obtaining a to-be-detected image; inputting the to-be-detected image into the trained student model to obtain a first classification result output by the classification output head and a first positioning result output by the positioning output head; performing target detection on the to-be-detected image according to the first classification result and the first positioning result.
[0114] In an optional implementation, the student output head includes a prototype head, a classification output head used for a classification task, a positioning output head used for a positioning task, and a segmentation output head used for a mask segmentation task, and the knowledge distillation apparatus 100 further includes an instance segmentation module 140, which is specifically used for: obtaining a to-be-segmented image; inputting the to-be-segmented image into the trained student model to obtain a prototype mask output by the prototype head, a second classification result output by the classification output head, a second positioning result output by the positioning output head, and mask coefficients output by the segmentation output head; Perform instance detection on the image to be segmented according to the second classification result and the second positioning result to obtain a detected target instance; The target instance is segmented according to the prototype mask and the mask coefficient to obtain the instance mask of the target instance.
[0115] The embodiment of the present invention also provides a block diagram of an electronic device 10, which implements the knowledge distillation method of the above embodiment. Figure 7 , Figure 7 This is a block diagram of an electronic device 10 provided in this embodiment. The electronic device 10 includes a processor 11 , a memory 12 , and a bus 13 . The processor 11 and the memory 12 are connected via the bus 13 .
[0116] Processor 11 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the knowledge distillation method of the above embodiment can be completed by hardware integrated logic circuits or software instructions in processor 11. The above-mentioned processor 11 can be a general-purpose processor, including a CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Logic Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0117] The memory 12 is used to store a program for implementing the knowledge distillation method. The program may be a software function module stored in the memory 12 in the form of software or firmware or solidified in the OS (Operating System) of the electronic device 10 .
[0118] After receiving the execution instruction, the processor 11 executes the program to implement the knowledge distillation method of the aforementioned embodiment.
[0119] This embodiment provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, the knowledge distillation method described in the above embodiment is implemented.
[0120] In summary, the embodiments of the present invention provide a knowledge distillation method, device, electronic device and computer storage medium, the method comprising: obtaining a sample image; inputting the sample image into a pre-constructed student model and a trained teacher model respectively to obtain an intermediate feature map, a first prediction result and a second prediction result, wherein the intermediate feature map is obtained by feature extraction by the student model, the first prediction result is output by the student output head of the student model, and the second prediction result is output by the teacher output head of the teacher model; inputting the intermediate feature map into a pre-constructed auxiliary output head to obtain an auxiliary prediction result, and the structural capacity of the auxiliary output head and the student output head are consistent; based on the first prediction result and with the second prediction result as a soft label of the auxiliary prediction result, knowledge distillation is performed on the student model to obtain a trained student model. Compared with the prior art, this embodiment has at least the following advantages: (1) By pre-constructing an auxiliary output head, the intermediate feature map output by the student model is used as the input of the auxiliary output head to obtain an auxiliary prediction result. Since the structural capacity of the auxiliary output head and the student output head are consistent, the difference between the feature maps of the teacher model and the student model is avoided, and the effect of knowledge distillation is guaranteed. At the same time, the second prediction result output by the teacher output head of the teacher model is used as the soft label of the auxiliary prediction result, thereby fully learning the knowledge of the teacher model and ensuring the efficiency of knowledge distillation; (2) For the instance segmentation task, the distillation loss is determined by focusing on the area with large distribution differences, and the forward KL divergence is used to obtain the global distillation loss to fully capture the global information or significant features. The reverse KL divergence is used to obtain the local distillation loss to fully complete the detailed information or local features. This knowledge distillation method can better meet the requirements of real-time instance segmentation with high efficiency.
[0121] The above descriptions are merely various embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A knowledge distillation method, characterized in that: The method comprises: Get a sample image; Inputting the sample image into a pre-built student model and a trained teacher model respectively to obtain an intermediate feature map, a first prediction result, and a second prediction result, wherein the intermediate feature map is obtained by feature extraction by the student model, the first prediction result is output by the student output head of the student model, and the second prediction result is output by the teacher output head of the teacher model; Inputting the intermediate feature map into a pre-built auxiliary output head to obtain an auxiliary prediction result, wherein the auxiliary output head has the same structural capacity as the student output head; Based on the first prediction result and using the second prediction result as a soft label for the auxiliary prediction result, knowledge distillation is performed on the student model to obtain a trained student model.
2. The method according to claim 1, characterized in that The step of performing knowledge distillation on the student model based on the first prediction result and using the second prediction result as a soft label for the auxiliary prediction result to obtain a trained student model includes: Calculating a distillation loss based on the soft label and the auxiliary prediction result; Calculating a main loss based on the first prediction result and the label of the sample image; Performing knowledge distillation on the student model according to the distillation loss and the main loss to obtain a trained student model.
3. The method according to claim 2, characterized in that The student output head is used to perform a mask segmentation task, and the step of calculating the distillation loss according to the soft label and the auxiliary prediction result includes: Obtaining a probability distribution of the soft label and a probability distribution of the auxiliary prediction result; Calculate the forward KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain the global distillation loss; Calculating the inverse KL divergence of the probability distribution of the soft label and the probability distribution of the auxiliary prediction result to obtain a local distillation loss; The distillation loss is calculated according to the global distillation loss and the local distillation loss.
4. The method according to claim 2, characterized in that The student output head is used to perform a classification task, and the step of calculating the distillation loss according to the soft label and the auxiliary prediction result includes: A binary cross entropy between the soft label and the auxiliary prediction result is calculated to obtain the distillation loss.
5. The method according to claim 2, characterized in that The student output head is used to perform a positioning task, and the step of calculating the distillation loss according to the soft label and the auxiliary prediction result includes: The complete intersection-over-union (IoU) between the soft label and the auxiliary prediction result is calculated to obtain the distillation loss.
6. The method according to claim 1, characterized in that The student output heads include a classification output head for performing a classification task and a positioning output head for performing a positioning task, and the method further includes: Obtain the image to be detected; Inputting the image to be detected into the trained student model to obtain a first classification result output by the classification output head and a first positioning result output by the positioning output head; Target detection is performed on the image to be detected according to the first classification result and the first positioning result.
7. The method according to claim 1, characterized in that The student output head includes a prototype head, a classification output head for performing a classification task, a positioning output head for performing a positioning task, and a segmentation output head for performing a mask segmentation task. The method further includes: Obtain the image to be segmented; Inputting the image to be segmented into the trained student model to obtain the prototype mask output by the prototype head, the second classification result output by the classification output head, the second positioning result output by the positioning output head, and the mask coefficient output by the segmentation output head; Performing instance detection on the image to be segmented according to the second classification result and the second positioning result to obtain a detected target instance; The target instance is segmented according to the prototype mask and the mask coefficient to obtain an instance mask of the target instance.
8. A knowledge distillation device, characterized in that: The device comprises: An acquisition module, used for acquiring a sample image; a distillation module, configured to input the sample image into a pre-built student model and a trained teacher model, respectively, to obtain an intermediate feature map, a first prediction result, and a second prediction result, wherein the intermediate feature map is obtained by feature extraction by the student model, the first prediction result is output by the student output head of the student model, and the second prediction result is output by the teacher output head of the teacher model; The distillation module is further configured to input the intermediate feature map into a pre-built auxiliary output head to obtain an auxiliary prediction result, wherein the auxiliary output head and the student output head have the same structural capacity; The distillation module is further used to perform knowledge distillation on the student model based on the first prediction result and use the second prediction result as a soft label for the auxiliary prediction result to obtain a trained student model.
9. An electronic device, characterized in that: It includes a processor and a memory, the memory is used to store a program, and the processor is used to implement the knowledge distillation method according to any one of claims 1 to 7 when executing the program.
10. A computer storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the knowledge distillation method as described in any one of claims 1 to 7.