Self-distillation model compression method and device, electronic product and medium

Through self-distillation technology and deep supervision technology, combined with knowledge distillation and attention map mapping, the problems of large model accuracy loss and high training cost in existing model compression tools are solved, and efficient and automated model compression and acceleration are achieved.

CN120068974APending Publication Date: 2025-05-30SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311620315.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing model compression tools have problems such as large loss of model accuracy, high two-stage training costs and high professional requirements for personnel, and it is difficult to ensure the compression rate and accuracy of the model at the same time.

Method used

The self-distillation technology is used to realize a one-stage training strategy, provide automatic model compression functions through deep supervision technology, and enhance multiple types of distillation knowledge through knowledge distillation and attention map mapping, and comprehensive distillation is used to improve the distillation effect.

Benefits of technology

It reduces training costs, realizes the trade-off between model accuracy and efficiency, improves the accuracy of algorithms on the test set, and provides simple, highly automated, user-friendly model compression and acceleration tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068974A_ABST
    Figure CN120068974A_ABST
Patent Text Reader

Abstract

The invention provides a self-distillation model compression method and device, an electronic product and a medium. The method comprises the steps that a teacher model needing model compression is constructed, the basic structure of the teacher model is divided into a plurality of residual blocks, and student models are constructed according to the residual blocks; through knowledge distillation of each model, enhancing result knowledge; enhancing process knowledge through attention map mapping of feature maps of the teacher model and each student model; according to the enhancement result of the effective knowledge and the enhancement result of the process knowledge, calculating a loss function of the teacher model and each student model; according to the loss function obtained through calculation, a student model with the classification effect closest to the teacher model is extracted to serve as a compression model of the teacher model. According to the self-distillation model compression method and device, the electronic product and the medium provided by the invention, a method technical support is provided for forming a simple, convenient, high-automation and user-friendly model compression and acceleration tool.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep neural networks, and particularly relates to a self-distillation model compression method, device, electronic product, and medium. Background Art

[0002] In recent years, deep neural networks (DNNs) have achieved remarkable success in the field of computer vision, such as image classification, object recognition, and semantic segmentation. To adapt to various complex situations in the real world, various deep learning tasks usually require deeper neural networks, resulting in increasing computational complexity and storage resource requirements. At the same time, with the development of the industrial Internet under cloud-edge collaboration, the applications of portable edge devices and embedded devices are becoming increasingly popular. However, these devices with limited memory and low computing power cannot support the online calculation of deep networks. Therefore, model compression in deep learning has become the focus of current research. How to deploy complex deep neural network models to devices with limited resources and complete inference calculations has become a new challenge.

[0003] In response to the challenges in the field of model compression mentioned above, some model compression tools have been formed in the industrial community. Most of them use traditional model compression techniques such as network pruning algorithms, low-rank decomposition algorithms, parameter quantization, knowledge distillation algorithms, etc. However, these methods have disadvantages such as large model accuracy loss, high two-stage training cost, and being difficult to interpret. In addition, most model compression tools adopt the method of manually selecting compression algorithms and parameters, which requires careful manual trade-off between the compression rate and accuracy loss, has high requirements for personnel expertise, and it is difficult to ensure both the compression rate and accuracy of the model simultaneously. Summary of the Invention

[0004] The purpose of the present invention is to provide a self-distillation model compression method, device, electronic product, and medium to solve the problems existing in the above background art.

[0005] To achieve the above purpose, in the first aspect, the present invention provides a self-distillation model compression method, including:

[0006] Construct a teacher model that needs model compression, divide the basic structure of the teacher model into several residual blocks, and construct a student model according to each residual block;

[0007] Enhance the result knowledge through knowledge distillation of the teacher model and each student model;

[0008] Enhance the process knowledge through attention map mapping of the feature maps of the teacher model and each student model;

[0009] Calculate the loss functions of the teacher model and each student model according to the enhancement results of the effective knowledge and the enhancement results of the process knowledge;

[0010] According to the calculated loss function, extract a student model whose classification effect is closest to the teacher model as the compressed model of the teacher model.

[0011] In a second aspect, the present invention provides a self-distillation model compression device, including:

[0012] A basic framework construction module, configured to construct a teacher model that needs model compression, divide the basic structure of the teacher model into several residual blocks, and construct a student model according to each residual block;

[0013] A result knowledge enhancement module, configured to enhance the result knowledge through knowledge distillation of the teacher model and each student model;

[0014] A process knowledge enhancement module, configured to enhance the process knowledge through attention map mapping of the feature maps of the teacher model and each student model;

[0015] A loss function calculation module, configured to calculate the loss functions of the teacher model and each student model according to the enhancement result of the effective knowledge and the enhancement result of the process knowledge;

[0016] A compressed model acquisition module, configured to extract a student model whose classification effect is closest to the teacher model as the compressed model of the teacher model according to the calculated loss function.

[0017] In a third aspect, the present invention provides an electronic product, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the self-distillation model compression method.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program realizes the self-distillation model compression method when executed by a processor.

[0019] On the one hand, the method provided by the embodiment of the present invention realizes a one-stage training strategy through self-distillation technology to reduce the training cost. On the other hand, it adopts deep supervision technology to provide the function of automatic model compression to meet the trade-off between model accuracy and efficiency. In addition, it enhances and integrates various types of distilled knowledge for comprehensive distillation to improve the distillation effect and the accuracy of the algorithm on the test set. The proposed invention provides an effective solution to the practical problems such as large model accuracy loss of the above traditional model compression tools, high two-stage training cost, and high professional requirements for personnel, and provides method technical support for forming a simple, highly automated, and user-friendly model compression and acceleration tool, which is of great significance in the field of deep learning model compression. Description of the Drawings

[0020] Figure 1 It is a flowchart of the self-distillation model compression method provided by the embodiment of the present invention;

[0021] Figure 2 It is a flowchart of the basic framework construction provided by the embodiment of the present invention;

[0022] Figure 3 It is an overall architecture diagram of the basic framework provided by the embodiment of the present invention;

[0023] Figure 4 It is a flowchart of the result knowledge enhancement provided by the embodiment of the present invention;

[0024] Figure 5 It is a flowchart of the process knowledge enhancement provided by the embodiment of the present invention;

[0025] Figure 6 It is a flowchart of the loss function calculation provided by the embodiment of the present invention;

[0026] Figure 7 It is a structural diagram of the self-distillation model compression device provided by the embodiment of the present invention;

[0027] Figure 8 It is a schematic diagram of the composition of an electronic product provided by the embodiment of the present invention. Detailed Embodiments

[0028] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0029] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0030] As Figure 1 shown, the self-distillation model compression method in an embodiment of the present invention includes:

[0031] S11, constructing a teacher model that needs model compression, dividing the basic structure of the teacher model into several residual blocks, and constructing a student model according to each residual block.

[0032] S12, enhancing the result knowledge through knowledge distillation of the teacher model and each student model.

[0033] S13, enhancing the process knowledge through attention map mapping of the feature maps of the teacher model and each student model.

[0034] S14, calculating the loss functions of the teacher model and each student model according to the enhancement results of the effective knowledge and the enhancement results of the process knowledge.

[0035] S15, extracting a student model with the classification effect closest to the teacher model according to the calculated loss function as the compressed model of the teacher model.

[0036] By executing the above series of steps, on the one hand, the one-stage training strategy is realized through the self-distillation technology to reduce the training cost. On the other hand, the deep supervision technology is adopted to provide the function of automatic model compression to meet the trade-off between model accuracy and efficiency. Here, various types of distilled knowledge are also enhanced and fused for comprehensive distillation to improve the distillation effect and the accuracy of the algorithm on the test set. The proposed method provides an effective solution to the problems such as large model accuracy loss of the above traditional model compression tools, high two-stage training cost, and high professional requirements for personnel, and provides method technical support for forming a simple, highly automated, and user-friendly model compression and acceleration tool.

[0037] As Figure 2 shown, the basic framework construction process in an embodiment of the present invention includes:

[0038] S21, selecting the ResNet convolutional network as the object of model compression, and dividing it into 4 ResBlocks according to the basic structure of the ResNet network.

[0039] First, select the ResNet convolutional network as the object of model compression. According to the basic structure of the ResNet network itself, it is divided into 4 ResBlocks, which are called ResBlock1, ResBlock2, ResBlock3, and ResBlock4 from shallow to deep, that is, residual block 1, residual block 2, residual block 3, and residual block 4.

[0040] S22. After the first three divided ResBlock blocks, an auxiliary classification network is connected respectively to output the classification results of the input samples.

[0041] Code is written to connect an auxiliary classification network after the first three divided ResBlock blocks in S21 to output the classification results of the input samples, which are called Classifier1, Classifier2, and Classifier3. Each auxiliary classification network is composed of a Bottleneck module, a pooling layer, and a fully connected layer.

[0042] S23. Construct a basic framework for self-distillation based on deep supervision.

[0043] Construct a basic framework for self-distillation based on deep supervision. Define "ResBlock1 + Classifier1" in S21 and S22 as the No. 1 student model, denoted as S1, and the output class probability distribution is denoted as z S1 ; "ResBlock1 + ResBlock2 + Classifier2" as the No. 2 student model, denoted as S2, and the output class probability distribution is denoted as z S2 ; "ResBlock1 + ResBlock2 + ResBlock3 + Classifier3" as the No. 3 student model, denoted as S3, and the output class probability distribution is denoted as z S3 ; The complete ResNet classification network is the teacher model, denoted as T, and the output class probability distribution is denoted as z T .

[0044] More specifically, the operation of S22, that is, after the first three divided ResBlock blocks, an auxiliary classification network is connected respectively to output the classification results of the input samples, includes:

[0045] S221. Build the Bottleneck modules in 3 auxiliary classification networks.

[0046] Write code to build the Bottleneck modules in the three auxiliary classification networks, which are called Bottleneck1, Bottleneck2, and Bottleneck3 respectively. The structures of the three Bottlenecks are similar. Each Bottleneck module is composed of three groups of "convolution layer + batch normalization + activation function" connected in series. Among them, the three convolution layers of Bottleneck1 use convolution kernels with sizes of 1×1, 8×8, and 1×1 respectively, the three convolution layers of Bottleneck2 use convolution kernels with sizes of 1×1, 4×4, and 1×1 respectively, and the three convolution layers of Bottleneck3 use convolution kernels with sizes of 1×1, 2×2, and 1×1 respectively for convolution operations. The number of each group of 3 convolution kernels is set to 128x, 128x, and 512x respectively, where x is determined according to the basic structure of ResNet itself and is generally set to 1 or 4. In addition, batch normalization uses the Batch Normalization module to normalize the features of each batch of input samples, and the activation function is set to the ReLU function.

[0047] S222. Build the pooling layers in the three auxiliary classification networks.

[0048] Write code to build the pooling layers in the three auxiliary classification networks. After the three Bottlenecks built in S221, average pooling layers are connected respectively. The average pooling layer is used to perform downsampling operations on the feature maps output from the Bottlenecks to obtain new feature map sizes of (512x, 1, 1).

[0049] S223. Build the fully connected layers in the three auxiliary classification networks.

[0050] Write code to build the fully connected layers in the three auxiliary classification networks. After each average pooling layer built in S222, a fully connected layer is connected to classify the input samples and output the class distribution probability results. The dimension of the output of the fully connected layer is set according to the actual number of classification classes. For example, for a 10-classification task on the public dataset CIFAR10, the output dimension of the fully connected layer is set to 10.

[0051] As Figure 3 shown, the self-distillation basic framework constructed from S21 to S23 includes: 4 ResBlocks divided from the basic structure of the ResNet convolutional network. Among them, the three ResBlocks at deeper levels are respectively constructed with their own Bottleneck, fully connected layer, and output layer. These three ResBlocks, together with the corresponding Bottleneck, fully connected layer, and output layer, jointly constitute three independent student models.

[0052] AsFigure 4 As shown in the figure, the result knowledge enhancement process in an embodiment of the present invention includes:

[0053] S41. Perform temperature coefficient processing on the output z of the 4-category probability distributions of 3 student models and 1 teacher model respectively.

[0054] First, perform temperature coefficient processing on the output z of the 4-category probability distributions of 3 student models and 1 teacher model in S23. The 4 processed probability distributions are respectively denoted as p S1 , p S2 , p S3 and p T .

[0055]

[0056] Among them, z i represents the probability value that the model outputs that the sample belongs to the i-th category. T is the temperature coefficient, which is set to 4 and is used to smooth the result of the probability distribution, so that the probability distribution contains more "dark knowledge" than the result before processing, laying a foundation for improving the distillation effect of subsequent self-distillation.

[0057] S42. Classify the 4-category probability distribution outputs of 3 student models and 1 teacher model that have undergone temperature coefficient processing respectively. The classification results include: the target category probability and the non-target category probability.

[0058] Classify the 4-category probability distribution outputs of 3 student models and 1 teacher model in S41 that have undergone temperature coefficient processing respectively. They are all classified into the target category probability, denoted as (teacher model) and (student model), that is, the probability value that the model outputs the actual category of an input sample; and the non-target category probability, that is, the probability distribution of all other categories except the actual category of an input sample output by the model.

[0059] S43. For the teacher model and each student model, perform knowledge distillation calculation on the probability distributions of the two parts of the divided target category and non-target category respectively to obtain the KL divergence logits distillation loss.

[0060] Define to represent the ratio of the probability of a certain non-target category to the probabilities of all non-target categories, where p \t represents the sum of the probabilities of all non-target categories, represents the binary probability of the target category (p t ) and all other non-target categories (p \t ).

[0061] For the teacher model and each student model, the KL divergence is calculated for the probability distributions of the two parts of the target category and non-target category divided in S42 to obtain the logits distillation loss for knowledge distillation.

[0062]

[0063] Among them, p T and p S represent the probability distributions output by the teacher and student models after being processed by the temperature coefficient. represent the predicted values of the teacher and student models for the target category after being processed by the temperature coefficient. and represent the predicted values of the teacher and student models for the i-th non-target category after being processed by the temperature coefficient.

[0064] S44. Decouple the simplified distillation loss formula, and enhance the effective knowledge distillation knowledge by flexibly matching the weights of TCKD and NCKD.

[0065] According to the definition in S43 The distillation loss formula in S44 can be simplified to:

[0066]

[0067] Among them, and are the sums of the probability values of all non-target categories for the teacher and student models after being processed by the temperature coefficient.

[0068] According to the KL divergence and the definition in S43 The distillation loss formula in S44 can be simplified to:

[0069]

[0070] Among them, b T and b S represent a binary probability distribution of the teacher and student networks that includes the sum of the predicted values of the target category and the predicted values of the non-target category. and represent the proportion distribution of a certain non-target category in the total probability sum of non-target categories. Here represents a fixed coupling coefficient.

[0071] Regarding KL(b T ||b S) regarded as the distillation loss of the probability distribution of the binary tuple composed of the predicted values of the teacher-student model for the target class and the sum of all non-target classes, denoted as TCKD (Target Class Knowledge Distillation); and ) regarded as the distillation loss of the probability distribution of the teacher-student model for each non-target class, denoted as NCKD (NoneTarget Class Knowledge Distillation). Therefore, the above distillation loss formula can be simplified as:

[0072]

[0073] Through experimental research, it is found that the TCKD module in S44 has little improvement on the distillation effect, while NCKD contains a large amount of "dark knowledge" and can significantly improve the distillation effect. Decouple the coupled distillation loss formula in S44, and enhance the effective knowledge distillation knowledge by flexibly matching the weights of TCKD and NCKD. The formula is as follows:

[0074] L kd = α·TCKD + β·NCKD

[0075] Among them, α and β are the weight hyperparameters of TCKD and NCKD in the distillation loss. In this method, α is set to 1 and β is set to 8.

[0076] As Figure 5 shown, the process knowledge enhancement process in an embodiment of the present invention includes:

[0077] S51, using differentiable soft attention to calculate the attention map, and extracting the feature maps of the last layer from ResBlock1, ResBlock2, ResBlock3, and ResBlock4 in the network.

[0078] First, use differentiable soft attention to calculate the attention map, and extract the feature maps of the last layer from ResBlock1, ResBlock2, ResBlock3, and ResBlock4 in the network in S21, that is, a 3D tensor, denoted as where C is the number of tensors, that is, the number of channels, and H and W represent the height and width of the tensor respectively.

[0079] S52, perform attention map mapping on the feature map, take the feature map as the input, and then use the activated mapping function F to compress the three-dimensional tensor into a planar two-dimensional vector in the spatial dimension, and output a spatial-based attention map.

[0080] Secondly, perform attention map mapping on the feature map, take the three-dimensional tensor (feature map) in S51 as the input, and then use the activated mapping function Compress the three-dimensional tensor into a planar two-dimensional vector in the spatial dimension, and output a spatial-based attention map, that is where C is the number of tensors, that is, the number of channels, and H and W represent the height and width of the tensor respectively. In this method, the mapping function The specific formula is as follows:

[0081]

[0082] where A represents the tensor of the feature map in S51, and A i represents the tensor of the i-th channel in the 3D feature map. In this method, p takes the value of 2, then this formula means squaring each planar two-dimensional feature map slice of a 3D feature map tensor and adding them according to the channel dimension, and converting it into a two-dimensional attention map.

[0083] Such as Figure 6 shown, the loss function calculation process in an embodiment of the present invention includes:

[0084] S61, calculate the true value loss part.

[0085] Calculate the true value loss part, that is, calculate the cross entropy (CE) loss loss1 for the true label ground truth and the output of each ResBlock of the network in S21:

[0086]

[0087] where C = 4 represents the number of auxiliary classifiers Classifier, and z i represents the predicted probability distribution obtained by the i-th classifier for the sample without temperature coefficient processing, and y represents the true label, that is, the sum of the differences between the class probability distributions of the output samples of the teacher-student model and the true label is calculated.

[0088] S62, calculate the distillation loss part.

[0089] Calculate the distillation loss part, that is, calculate the KL divergence (Kullback-Leibler Divergence) loss2 for the enhanced result knowledge of the teacher model (original network) for each student model (Classifier1, Classifier2, and Classifier3 in S21):

[0090]

[0091] where C = 4 represents the number of auxiliary classifiers Classifier, TCKD i and NCKDi Denote the distilled loss of the target class and the non - target class between the output of the \(i\) - th classifier and the teacher model as \(\alpha\) dkd and \(\beta\) dkd represent the weight ratios of TCKD and NCKD. As shown in S44, \(\alpha\) dkd is set to 1 and \(\beta\) dkd is set to 8.

[0092] S63, calculate the intermediate feature information loss.

[0093] Calculate the intermediate feature information loss, which is the \(L\) loss obtained by calculating the inter - layer attention map obtained in S52. The KED method selects, in the distillation path strategy of self - attention distillation, to make each self - attention map of the shallow layer simulate and learn from the self - attention map of the next intermediate layer, and calculate the \(L\) 2 loss loss3. 2 Here, \(AT\)

[0094]

[0095] denotes the output of the attention map of the \(i\) - th ResBlock. i S64, adopt a combination of multiple knowledge sources, combine the ground - truth loss, the decoupled logits distilled loss of deep self - supervision, and the attention map distilled loss, and calculate the total loss function.

[0096] Adopt a combination of multiple knowledge sources, combine the ground - truth loss, the decoupled logits distilled loss of deep self - supervision, and the attention map distilled loss, and construct a self - distillation framework KED based on deep supervision and knowledge enhancement. The total loss function Loss is as follows:

[0097] Loss=(1 - \(\alpha\))·loss1+\(\alpha\)·loss2+\(\lambda\)·loss3

[0098] where \(\alpha\) and \(\lambda\) are weight hyperparameters respectively, which can be manually adjusted to balance each loss part. In this method, \(\alpha\) takes 0.1 and \(\lambda\) takes \(1\times10\)

[0099] , loss1, loss2, and loss3 correspond to the loss functions of S61, S62, and S63 respectively. -6 S65, perform backpropagation of network training through the total loss function to update network parameters, and finally obtain a network trained by the self - distillation method based on deep supervision and knowledge enhancement.

[0100] Perform backpropagation of network training through the loss function of S64 to update network parameters, and finally obtain a network trained by the self - distillation method based on deep supervision and knowledge enhancement.

[0101] ​

[0102] As Figure 7 shown, the self-distillation model compression device in an embodiment of the present invention includes: a basic framework construction module, a result knowledge enhancement module, a process knowledge enhancement module, a loss function calculation module, and a compressed model acquisition module.

[0103] The basic framework construction module is used to construct a teacher model that needs model compression, divide the basic structure of the teacher model into several residual blocks, and construct a student model according to each residual block.

[0104] The result knowledge enhancement module is used to enhance the result knowledge through knowledge distillation of the teacher model and each student model.

[0105] The process knowledge enhancement module is used to enhance the process knowledge through attention map mapping of the feature maps of the teacher model and each student model.

[0106] The loss function calculation module is used to calculate the loss functions of the teacher model and each student model according to the enhancement results of the effective knowledge and the enhancement results of the process knowledge.

[0107] The compressed model acquisition module is used to extract a student model with the classification effect closest to the teacher model as the compressed model of the teacher model according to the calculated loss function.

[0108] In some embodiments, the basic framework construction module includes: a selection unit, an access unit, and a construction unit.

[0109] The selection unit is used to select the ResNet convolutional network as the object of model compression, and divide it into 4 ResBlocks according to the basic structure of the ResNet network.

[0110] The access unit is used to connect an auxiliary classification network to the first three divided ResBlock blocks respectively to output classification results for the input samples.

[0111] The construction unit is used to construct a self-distillation basic framework based on deep supervision.

[0112] In some embodiments, the access unit is specifically used to: build the Bottleneck module in 3 auxiliary classification networks; build the pooling layer in 3 auxiliary classification networks; build the fully connected layer in 3 auxiliary classification networks.

[0113] In some embodiments, the result knowledge enhancement module includes: a temperature coefficient processing unit, a classification unit, a calculation unit, and a decoupling unit.

[0114] The temperature coefficient processing unit is used to perform temperature coefficient processing on the 4 category probability distribution outputs z of 3 student models and 1 teacher model respectively.

[0115] The classification unit is used to classify the four temperature coefficient processed categorical probability distribution outputs of three student models and one teacher model respectively. The classification results include: the target category probability and the non-target category probability.

[0116] The calculation unit is used to calculate the KL divergence of the probability distributions of the two parts of the target category and the non-target category divided for the teacher model and each student model respectively to obtain the logits distillation loss.

[0117] The decoupling unit is used to decouple the simplified distillation loss formula and enhance the effective knowledge distillation knowledge by flexibly matching the weights of TCKD and NCKD.

[0118] In some embodiments, the process knowledge enhancement module includes: an extraction unit and a mapping unit.

[0119] The extraction unit is used to calculate the attention map using differentiable soft attention and extract the feature maps of the last layer from ResBlock1, ResBlock2, ResBlock3, and ResBlock4 in the network.

[0120] The mapping unit is used to perform attention map mapping on the feature map. Taking the feature map as the input, and then using the activated mapping function F, compress the three-dimensional tensor into a planar two-dimensional vector in the spatial dimension and output a spatial-based attention map.

[0121] In some embodiments, the loss function calculation module includes: a first loss calculation unit, a second loss calculation unit, a third loss calculation unit, a fusion unit, and an update unit.

[0122] The first loss calculation unit is used to calculate the true value loss part.

[0123] The second loss calculation unit is used to calculate the distillation loss part.

[0124] The third loss calculation unit is used to calculate the intermediate feature information loss.

[0125] The fusion unit is used to fuse multiple kinds of knowledge, combine the true value loss, the decoupled logits distillation loss of deep self-supervision, and the attention map distillation loss to calculate the total loss function.

[0126] The update unit is used to perform backpropagation of network training through the total loss function to update the network parameters, and finally obtain a network trained by the self-distillation method based on deep supervision and knowledge enhancement.

[0127] In some embodiments, the compressed model acquisition module includes: a determination unit and a storage unit.

[0128] The determination unit is used to determine the student model whose classification effect is closest to the teacher model.

[0129] The storage unit is used to store the weight parameters of this student model.

[0130] In one embodiment of the present invention, an electronic product is provided, such as Figure 8 shown, the electronic device includes at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned self-distillation model compression method.

[0131] Wherein, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, etc., through an interface, which are well known in the art. The interface provides an interface between the bus and the transceiver, such as a communication interface and a user interface. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium through an antenna. Further, the antenna also receives data and transmits the data to the processor.

[0132] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. And the memory can be used to store the data used by the processor when executing operations.

[0133] In one embodiment of the present invention, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the above-mentioned method embodiments are implemented.

[0134] Those skilled in the art can understand through the above description that all or part of the steps in implementing the method of the above embodiments can be completed by a program instructing relevant hardware. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. And the foregoing storage medium includes, but is not limited to, various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic memories, and optical memories.

[0135] In several embodiments provided in this application, it should be understood that the disclosed system, apparatus, or method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices, modules, or units can be in electrical, mechanical, or other forms.

[0136] The modules / units described as separate components may or may not be physically separated. The components shown as modules / units may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of this application. For example, in each embodiment of this application, the functional modules / units can be integrated into one processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated into one module / unit.

[0137] Those of ordinary skill in the art should also further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0138] The descriptions of the processes or structures corresponding to the above-mentioned respective drawings each have their own focuses. For the parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.

[0139] The above embodiments are only illustrative of the principles and effects of this application and are not used to limit this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.

Claims

1. A self-distillation model compression method, including: Construct a teacher model that needs model compression, divide the basic structure of the teacher model into several residual blocks, and construct a student model according to each residual block; Enhance the result knowledge through knowledge distillation of the teacher model and each student model; Enhance the process knowledge through attention map mapping of the feature maps of the teacher model and each student model; Calculate the loss functions of the teacher model and each student model according to the enhancement results of the effective knowledge and the enhancement results of the process knowledge; Extract a student model with the classification effect closest to the teacher model as the compression model of the teacher model according to the calculated loss function.

2. The self-distillation model compression method according to claim 1, characterized in that Construct a teacher model that needs model compression, divide the basic structure of the teacher model into several residual blocks, and construct a student model according to each residual block, including: Select the ResNet convolutional network as the object of model compression, and divide it into 4 ResBlocks according to the basic structure of the ResNet network; Connect auxiliary classification networks to the first three divided ResBlock blocks respectively to output classification results for the input samples; Construct a self-distillation basic framework based on deep supervision.

3. The self-distillation model compression method according to claim 2, characterized in that Connect auxiliary classification networks to the first three divided ResBlock blocks respectively to output classification results for the input samples, including: Build the Bottleneck modules in 3 auxiliary classification networks; Build the pooling layers in 3 auxiliary classification networks; Build the fully connected layers in 3 auxiliary classification networks.

4. The self-distillation model compression method according to claim 1, characterized in that Enhance the result knowledge through knowledge distillation of the teacher model and each student model, including: Perform temperature coefficient processing on the 4 category probability distribution outputs z of 3 student models and 1 teacher model respectively; Classify the 4 category probability distribution outputs after temperature coefficient processing of 3 student models and 1 teacher model respectively, and the classification results include: target category probability and non-target category probability; For the teacher model and each student model, calculate the KL divergence of the probability distributions of the two parts of the divided target category and non-target category respectively to obtain the logits distillation loss; Decouple the simplified distillation loss formula, and enhance the effective knowledge distillation knowledge by flexibly matching the weights of TCKD and NCKD.

5. The self-distillation model compression method according to claim 1, characterized in that Enhance the process knowledge through attention map mapping of the feature maps of the teacher model and each student model, including: Use differentiable soft attention to calculate the attention map, and extract the feature maps of the last layer from ResBlock1, ResBlock2, ResBlock3 and ResBlock4 in the network; Perform attention map mapping on the feature map. Take the feature map as the input, and then adopt the activated mapping function F to compress the three-dimensional tensor into a planar two-dimensional vector in the spatial dimension, and output a spatial-based attention map.

6. The self-distillation model compression method according to claim 1, wherein, According to the enhancement result of effective knowledge and the enhancement result of process knowledge, calculate the loss functions of the teacher model and each student model, including: Calculate the true value loss part; Calculate the distillation loss part; Calculate the intermediate feature information loss; Adopt a combination of multiple types of knowledge, combine the true value loss, the decoupled logits distillation loss of deep self-supervision, and the attention map distillation loss, and calculate the total loss function; Perform backpropagation of network training through the total loss function to update the network parameters, and finally obtain a network trained by the self-distillation method based on deep supervision and knowledge enhancement.

7. The self-distillation model compression method according to claim 1, wherein, According to the calculated loss function, extract a student model whose classification effect is closest to the teacher model as the compression model of the teacher model, including: Determine the student model whose classification effect is closest to the teacher model; Save the weight parameters of this student model.

8. A self-distillation model compression device, including: A basic framework construction module for constructing a teacher model that needs model compression, dividing the basic structure of the teacher model into several residual blocks, and constructing student models according to each residual block; A result knowledge enhancement module for enhancing the result knowledge through knowledge distillation of the teacher model and each student model; A process knowledge enhancement module for enhancing the process knowledge through attention map mapping of the feature maps of the teacher model and each student model; A loss function calculation module for calculating the loss functions of the teacher model and each student model according to the enhancement result of effective knowledge and the enhancement result of process knowledge; A compression model acquisition module for extracting a student model whose classification effect is closest to the teacher model as the compression model of the teacher model according to the calculated loss function.

9. An electronic product, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the self-distillation model compression method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the self-distillation model compression method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Learning type image compression method and system based on stage module distillation

    CN120835152A

  • A learning-based image compression method and system based on stage-wise modular distillation

    CN120835152B