A model distillation method, device and medium

By constructing comparative learning sample pairs and sorting differential sample pairs, a distillation loss function is constructed, and the student model is iteratively trained, which solves the problems of deep learning models in taking into account both computational accuracy and running speed, and achieves efficient training and performance improvement of student model.

CN119692430BActive Publication Date: 2025-06-13ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510205236.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In practical applications, existing deep learning models are difficult to take into account both the calculation accuracy and the running speed, which makes it difficult to run on mobile or embedded devices.

Method used

By obtaining the similarity matrix of the teacher model and the student model and the sorting difference matrix, constructing a comparative learning sample pair and sorting difference sample pair, constructing a distillation loss function, and improving the performance of the student model through iterative training.

Benefits of technology

At lower computing costs, the student model can learn the knowledge of high-precision teacher model, achieve performance close to the teacher model, thereby improving the accuracy and computing efficiency of the student model and expanding the scope of application of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119692430B_ABST
    Figure CN119692430B_ABST
Patent Text Reader

Abstract

The present application discloses a model distillation method, apparatus, and medium. The method includes: obtaining a teacher similarity matrix of a teacher model, a student similarity matrix of a student model, and a ranking difference matrix representing the difference in similarity rankings between the teacher model and the student model; the calculation accuracy of the teacher model is higher than that of the student model. Constructing contrastive learning sample pairs according to the teacher similarity matrix, and constructing ranking difference sample pairs according to the ranking difference matrix; constructing a distillation loss function based on the contrastive learning sample pairs and the ranking difference sample pairs; and iteratively training the student model through the distillation loss function to obtain a target student model. Thus, by constructing contrastive learning sample pairs and ranking difference sample pairs based on the teacher model, constructing a distillation loss function, and iteratively training the student model through this distillation loss function, the student model can learn the knowledge of the teacher model, improve the accuracy of the student model, and thus achieve both computational efficiency and computational accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model distillation method, apparatus, and medium. Background Art

[0002] With the rapid development of deep learning technology, neural network models are widely used in various fields such as autonomous driving, image processing, natural language processing, and computer vision. In the actual application of neural network models, the higher the model calculation accuracy and complexity, the slower the model running speed, and it may even be unable to run on mobile or embedded devices.

[0003] Therefore, how to balance the model calculation accuracy and running speed and expand the model usage range is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, one aspect of this application provides a model distillation method, and the method includes:

[0005] Obtain a teacher similarity matrix of a teacher model, a student similarity matrix of a student model, and a sorting difference matrix representing the difference in similarity sorting between the teacher model and the student model; the calculation accuracy of the teacher model is higher than that of the student model;

[0006] Construct contrast learning sample pairs according to the teacher similarity matrix; and construct sorting difference sample pairs according to the sorting difference matrix;

[0007] Construct a distillation loss function based on the contrast learning sample pairs and the sorting difference sample pairs;

[0008] Iteratively train the student model through the distillation loss function to obtain a target student model.

[0009] Optionally, constructing a distillation loss function based on the contrast learning sample pairs and the sorting difference sample pairs includes:

[0010] Obtain the task type of the student model;

[0011] Determine the label loss function of the student model according to the task type;

[0012] Determine a learning contrast loss function according to the contrast learning sample pairs; and determine a sorting difference loss function according to the sorting difference sample pairs;

[0013] Determine the distillation loss function according to the label loss function, the learning contrast loss function, and the sorting difference loss function.

[0014] Optionally, determining a sorting difference loss function according to the sorting difference sample pairs includes:

[0015] Determining a teacher sorting matrix, a student sorting matrix, and a similarity difference matrix according to the teacher similarity matrix and the student similarity matrix;

[0016] Determining a sorting difference weight according to the teacher sorting matrix, the student sorting matrix, and the sorting difference sample pairs;

[0017] Determining a similarity difference weight according to the similarity difference matrix and the sorting difference sample pairs;

[0018] Determining a teacher model prior weight according to the teacher similarity matrix and the sorting difference sample pairs;

[0019] Determining a difference loss weight through the sorting difference weight, the similarity difference weight, and the teacher model prior weight;

[0020] Determining the sorting difference loss function according to the difference loss weight and the sorting difference sample pairs.

[0021] Optionally, the greater the sorting difference between the teacher sorting matrix and the student sorting matrix for the same sorting difference sample pairs, the greater the sorting difference weight;

[0022] The greater the similarity difference between the teacher similarity matrix and the student similarity matrix for the same sorting difference sample pairs, the greater the similarity difference weight;

[0023] The smaller the similarity of the sorting difference sample pairs in the teacher similarity matrix, the greater the teacher model prior weight.

[0024] Optionally, determining a difference loss weight through the sorting difference weight, the similarity difference weight, and the teacher model prior weight includes:

[0025] Assigning a first weight to the sorting difference weight; assigning a second weight to the similarity difference weight; and assigning a third weight to the teacher model prior weight;

[0026] Determining the difference loss weight according to the first weight, the second weight, the third weight, the sorting difference weight, the similarity difference weight, and the teacher model prior weight; wherein, as the current iteration number of the iterative training increases, the first weight increases, and the second weight and the third weight decrease.

[0027] Optionally, constructing contrastive learning sample pairs according to the teacher similarity matrix includes:

[0028] Use the samples in the teacher similarity matrix with similarity greater than the first threshold as positive sample pairs, and the samples with similarity less than the second threshold as negative sample pairs; wherein, the first threshold is greater than the second threshold, and as the current iteration number of the iterative training increases, the larger the first threshold is, the smaller the second threshold is;

[0029] The positive samples and the negative samples constitute the contrast learning sample pairs.

[0030] Optionally, construct sorting difference sample pairs according to the sorting difference matrix, including:

[0031] Determine the target objects in the sorting difference matrix that meet the preset conditions; wherein, the preset conditions include that the similarity sorting difference is greater than the preset value, the first preset quantile in the sorting result of the similarity sorting differences from large to small, and any one of the first preset in the sorting result.

[0032] Use the samples corresponding to the target objects as the sorting difference sample pairs.

[0033] Another aspect of the present application provides a model distillation device, and the device includes:

[0034] A matrix acquisition module, configured to acquire the teacher similarity matrix of the teacher model, the student similarity matrix of the student model, and the sorting difference matrix representing the similarity sorting difference between the teacher model and the student model; the calculation accuracy of the teacher model is higher than that of the student model;

[0035] A sample pair construction module, configured to construct contrast learning sample pairs according to the teacher similarity matrix; and construct sorting difference sample pairs according to the sorting difference matrix;

[0036] A distillation loss function construction module, configured to construct a distillation loss function based on the contrast learning sample pairs and the sorting difference sample pairs;

[0037] A student model training module, configured to iteratively train the student model through the distillation loss function to obtain a target student model.

[0038] Another aspect of the present application provides a model distillation device, including a memory and a processor, and a computer program that can run on the processor is stored on the memory, and when the processor executes the program, the steps of the model distillation method are implemented.

[0039] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the model distillation method are implemented.

[0040] A model distillation method, apparatus and medium provided by the present application have the following beneficial effects: based on the teacher model, contrast learning sample pairs and ranking difference sample pairs are constructed, a distillation loss function is built, and the student model is iteratively trained through this distillation loss function, so that the student model with lower accuracy can learn the knowledge of the high-accuracy teacher model, and achieve performance close to that of the teacher model at a lower computational cost, thereby achieving the purpose of improving the accuracy of the student model, enabling the student model to take into account both computational efficiency and computational accuracy, and expanding the scope of application of the model. Description of the Drawings

[0041] Figure 1 It is a schematic flowchart of a model distillation method provided by an embodiment of the present application;

[0042] Figure 2 It is a schematic diagram of the principle of a model distillation method provided by an embodiment of the present application;

[0043] Figure 3 It is a schematic flowchart of a model distillation method provided by another embodiment of the present application;

[0044] Figure 4 It is a schematic structural diagram of a model distillation apparatus provided by an embodiment of the present application;

[0045] Figure 5 It is a schematic structural diagram of a model distillation apparatus provided by another embodiment of the present application.

[0046] The reference numerals are as follows: 50 is a memory, 51 is a processor, 52 is a display screen, 53 is an input / output interface, 54 is a communication interface, 55 is a power supply, 56 is a communication bus, 501 is a computer program, 502 is an operating system, and 503 is data. Detailed Embodiments

[0047] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0048] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0049] Figure 1 The flowchart of a model distillation method provided by an embodiment of this application is shown as Figure 1 shown, and this method includes:

[0050] S10: Obtain the teacher similarity matrix of the teacher model, the student similarity matrix of the student model, and the ranking difference matrix representing the difference in similarity ranking between the teacher model and the student model; the calculation accuracy of the teacher model is higher than that of the student model;

[0051] In a specific embodiment, the calculation accuracy of the teacher model is higher than that of the student model, that is, the teacher model has a higher model complexity and significantly better performance than the student model. Therefore, in practical applications, the teacher model usually requires relatively high computing resources and is usually difficult to be directly deployed on resource-constrained devices, while the student model requires lower computing resources and can be directly deployed. In order to enable the student model to have high running speed and high-precision characteristics at the same time, the student model can learn the knowledge of the teacher model so as to approach the performance of the teacher model.

[0052] Figure 2 The schematic diagram of the principle of a model distillation method provided by an embodiment of this application. Specifically, input samples with a sample quantity of are respectively input into the teacher model and the student model, so as to obtain the teacher output result of the teacher model and the student output result of the student model. Among them, . Among them, , , is used to represent the input sample quantity and also represents the output result quantity.

[0053] Furthermore, according to the teacher output result , the teacher similarity matrix of the teacher model can be generated, and according to the student output result , the student similarity matrix of the student model can be generated. Among them, . It should be noted that the similarity matrix is composed of similarity values between output results and is used to characterize the degree of difference between output samples. For example, in the teacher similarity matrix , it describes the degree of difference between the th input sample and the th input sample. When , it means that the similarity value between the 1st input sample and the 3rd input sample is 0.7, indicating a relatively high degree of closeness. As Figure 2 shows, after inputting the input samples into the teacher model and the student model, the teacher similarity matrix and the student similarity matrix can be constructed respectively.

[0054] In an optional embodiment, the teacher similarity matrix and the student similarity matrix are obtained. Specifically, by inputting in sequence, taking the output result corresponding to the current input sample as the base vector and the output results of other samples as the target vectors, calculating the similarity values between the base vector and the target vectors in sequence, and after traversing all input samples, constructing the similarity matrix through all similarity values. In fact, it can be understood as calculating the similarity values between pairwise output results to form the similarity matrix.

[0055] It should be noted that when calculating the similarity value between the base vector and the target vector, it can be calculated through, but not limited to, metric functions such as cosine similarity and Euclidean distance.

[0056] Furthermore, by taking the difference between the teacher similarity matrix and the student similarity matrix , a similarity difference matrix and for describing the similarity difference between the teacher similarity matrix can be obtained. Among them, , the similarity matrix is symmetric, , and the diagonal elements are 0. It should be noted that in the similarity difference matrix , it describes the degree of similarity difference between the teacher model and the student model for input samples at the same position (i.e., between the th input sample and the th input sample). That is, it describes the degree of similarity difference between the teacher similarity matrix and the student similarity matrix for input samples at the same position.

[0057] In addition, as Figure 2 shows, the teacher ranking matrix can also be obtained according to the teacher similarity matrix, and the student ranking matrix can be obtained according to the student similarity matrix. In the teacher ranking matrix It describes that, taking the th input sample as the reference sample, the similarity ranking position of the th input sample relative to the th input sample. For example, = 1 describes that for the 1st sample, the 3rd sample is the closest to it, that is, the similarity ranking position is the 1st. If = 3, it describes that taking the 1st sample as the reference sample, the similarity ranking position of the 3rd sample to it is the 3rd, that is, the 3rd most similar. The student ranking matrix is the same. It can be understood that since the ranking matrix is used to describe the similarity ranking between two input samples, all elements in the ranking matrix are integers representing the ranking positions.

[0058] Furthermore, after taking the difference between the teacher ranking matrix and the student ranking matrix , a ranking difference matrix can be obtained, which is used to characterize the similarity ranking difference between the teacher model and the student model. Among them, , the diagonal elements of the batch sample ranking matrix are 0. In the ranking difference matrix , it describes the similarity ranking difference degree between the teacher model and the student model for the input samples at the same position (that is, between the th input sample and the th input sample). That is, it describes the similarity ranking difference degree between the teacher ranking matrix and the student ranking matrix for the input samples at the same position.

[0059] S11: Construct contrastive learning sample pairs according to the teacher similarity matrix; and construct ranking difference sample pairs according to the ranking difference matrix;

[0060] S12: Based on the contrastive learning sample pairs and the ranking difference sample pairs, construct a distillation loss function;

[0061] S13: Iteratively train the student model through the distillation loss function to obtain the target student model.

[0062] To learn the performance of the teacher model, contrastive learning sample pairs are constructed according to the teacher similarity matrix. Specifically, the samples corresponding to the preset similarity conditions in the teacher similarity matrix are used as the contrastive learning sample pairs. In addition, in an optional embodiment, ranking difference sample pairs are constructed through the samples corresponding to the preset conditions in the ranking difference matrix.

[0063] Further, a distillation loss function is constructed based on the constructed contrastive learning sample pairs and ranking difference sample pairs. It can be understood that model distillation is a technique for transferring the knowledge of a complex model (i.e., the teacher model) to a small and simple model (i.e., the student model). By this method, the student model can achieve performance close to that of the teacher model while maintaining a low computational cost. Therefore, it is necessary to construct a distillation loss function to iteratively train the student model to obtain a target student model with high running speed and high accuracy. It should be noted that the iterative training of the student model can be trained until a preset number of iterations, and the application does not limit the iterative termination condition.

[0064] That is, the target student model trained by the constructed distillation loss function has a good running speed, performance close to that of the teacher model, and the characteristics of high accuracy and high speed, enabling a large number of mobile terminals and embedded devices to run the target student model, achieving the purpose of expanding the scope of model use.

[0065] Thus, the model distillation method provided by the embodiments of the present application constructs contrastive learning sample pairs and ranking difference sample pairs based on the teacher model, constructs a distillation loss function, and iteratively trains the student model through this distillation loss function, enabling the student model with low accuracy to learn the knowledge of the high-accuracy teacher model, obtaining performance close to that of the teacher model at a low computational cost, thereby achieving the purpose of improving the accuracy of the student model, realizing that the student model takes both computational efficiency and computational accuracy into account, and expanding the scope of model application.

[0066] In an alternative embodiment, constructing a distillation loss function based on the contrastive learning sample pairs and ranking difference sample pairs includes:

[0067] Obtain the task type of the student model;

[0068] Determine the label loss function of the student model according to the task type;

[0069] Determine the learning contrast loss function according to the contrastive learning sample pairs; and determine the ranking difference loss function according to the ranking difference sample pairs;

[0070] Determine the distillation loss function according to the label loss function, the learning contrast loss function, and the ranking difference loss function.

[0071] In a specific embodiment, as Figure 2 shown, the distillation loss function is a comprehensive loss function, including the learning contrast loss function determined according to the contrastive learning sample pairs , and the ranking difference loss function determined according to the ranking difference sample pairs .

[0072] Among them, the learning contrast loss function is mainly used for contrastive learning and is a self-supervised learning method. This function encourages the representations of similar samples to be closer and the representations of dissimilar samples to be farther apart. By introducing contrastive learning during the model distillation process, the student model can not only mimic the prediction results of the teacher model but also mimic the internal representations or the feature space structure of the teacher model, thus enabling the student model to approach the performance of the teacher model. And the ranking difference loss function can be used to ensure that the student model maintains a similarity ranking similar to that of the teacher model, thereby further ensuring that the student model is as close as possible to the teacher model.

[0073] It should be noted that the learning contrast loss function can include but is not limited to the normalized temperature-scaled cross-entropy loss and the conventional contrastive learning loss. In an optional embodiment, the learning contrast loss function can select the normalized temperature-scaled cross-entropy loss, which is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs in the contrastive learning sample pairs. According to the contrastive learning sample pairs, the specific calculation formula for the learning contrast loss function is:

[0074] (1)

[0075] where is the learning contrast loss function, is the positive sample pair in the contrastive learning sample pairs, is the negative sample pair in the contrastive learning sample pairs, is the number of positive sample pairs, , and are the output results corresponding to the th, th, and th input samples of the student model respectively. is the similarity function. In an optional embodiment, it can be represented by the cosine similarity and is directly obtained from the student similarity matrix. is used to calculate the similarity of positive sample pairs, is used to calculate the similarity of positive sample pairs and negative sample pairs. is the temperature parameter, which is used to control the smoothness of the distribution.

[0076] It should be noted that the positive sample pair in the contrastive learning sample pairs refers to the input sample corresponding to the similarity greater than the first threshold in the teacher similarity matrix, and the negative sample pair refers to the input sample corresponding to the similarity less than the second threshold in the teacher similarity matrix, where the first threshold is greater than the second threshold.

[0077] In an alternative embodiment, to ensure that the student model does not deviate too far from the basic information provided by the true labels when learning the knowledge of the teacher model, therefore, as Figure 2 shown, a label loss function is introduced. The label loss function can be understood as a task loss function, that is, different loss functions are selected according to different task types. In a specific embodiment, when constructing the distillation loss function , the task type of the student model is obtained, and the label loss function is selected according to the task type.

[0078] It can be understood that the label loss function mainly depends on the specific task type of the student model. In an alternative embodiment, when the task type is a classification task, cross-entropy loss can be used as the label loss function , and the specific calculation formula is formula (2):

[0079] (2)

[0080] where, is the number of classification categories, is the true label (if represented by one-hot encoding, the target category is 1, and other categories are 0), is the prediction probability of the student model.

[0081] Furthermore, as Figure 2 shown, according to the sorted difference sample pairs, a sorted difference loss function is determined. In fact, the distillation loss function can be understood as including two parts. One part is the loss describing the student model and the true label, that is, the supervision loss of the student model itself, that is, the label loss function . The other part is the imitation loss function, that is, the loss function imitating the output of the teacher model, specifically including the learning contrast loss function and the sorted difference loss function . Thus, the formula for the distillation loss function is:

[0082] (3)

[0083] In a specific embodiment of the CIFAR100 image classification task, a 26-layer residual network is used as the teacher model, and a 14-layer residual network is used as the student model. The accuracy rate of the teacher model is 61.72%, and the accuracy rate of the student model trained by conventional training is 58.15%. Through the distillation loss function provided by the embodiments of the present application After training the student model, the accuracy rate of the target student model obtained can reach 61.08%, which is only 0.64% lower than that of the teacher model. That is, while maintaining a high running speed, the accuracy is greatly improved by learning the performance of the teacher model.

[0084] Figure 3 The flowchart of a model distillation method provided by another embodiment of the present application is shown in Figure 3 As shown, in an alternative embodiment, determining a sorting difference loss function according to sorting difference sample pairs includes:

[0085] S30: Determine a teacher sorting matrix, a student sorting matrix, and a similarity difference matrix according to the teacher similarity matrix and the student similarity matrix;

[0086] S31: Determine a sorting difference weight according to the teacher sorting matrix, the student sorting matrix, and the sorting difference sample pairs;

[0087] In a specific embodiment, a teacher sorting matrix describing the similarity sorting relationship between different input samples can be generated according to the teacher similarity matrix, that is, the similarity sorting relationship between the input samples is recorded in the teacher sorting matrix. At the same time, a student similarity matrix can be generated according to the student similarity matrix.

[0088] Further, according to the teacher sorting matrix and the student sorting matrix, the similarity sorting difference between the teacher sorting matrix and the student sorting matrix at the same position input samples can be obtained. The specific obtaining method can refer to the description of the above embodiments and will not be elaborated here.

[0089] In an alternative embodiment, the sorting difference loss function Adaptive weights are used. Specifically, the difference loss weight of the sorting difference loss function Includes a sorting difference weight And a similarity difference weight And a teacher model prior weight And .

[0090] Among them, the sorting difference weight is determined through step S31 . Specifically, the calculation formula for determining the sorting difference weight according to the teacher sorting matrix, the student sorting matrix, and the sorting difference sample pairs is:

[0091] (4)

[0092] Among them, is the sorting difference weight. In the teacher sorting matrix, taking the th input sample as the reference sample, the th input sample is the similarity sorting position relative to the th input sample. In the student sorting matrix, taking the th input sample as the reference sample, the th input sample is the similarity sorting position relative to the th input sample. And the th input sample and the th input sample form a sorting difference sample pair. is the number of input samples.

[0093] S32: Determine the similarity difference weight according to the similarity difference matrix and the sorting difference sample pair;

[0094] In an optional embodiment, the specific calculation formula for determining the similarity difference weight according to the similarity difference matrix and the sorting difference sample pair is:

[0095] (5)

[0096] Among them, is the similarity difference weight. is the similarity value corresponding to the sorting difference sample pair formed by the th input sample and the th input sample in the similarity difference matrix, which is used to describe the difference between the similarity of the th input sample and the th input sample in the teacher similarity matrix and the similarity of the th input sample and the th input sample in the student similarity matrix. is the maximum similarity value in the similarity difference matrix.

[0097] S33: Determine the teacher model prior weight according to the teacher similarity matrix and the sorting difference sample pair;

[0098] In an optional embodiment, the specific calculation formula for determining the teacher model prior weight according to the teacher similarity matrix and the sorting difference sample pair is:

[0099] (6)

[0100] Among them, is the prior weight of the teacher model, is the similarity value corresponding to the sorting difference sample pair composed of the th input sample and the th input sample in the teacher similarity matrix, is the maximum similarity value in the teacher similarity matrix.

[0101] S34: Determine the difference loss weight through the sorting difference weight, similarity difference weight, and prior weight of the teacher model;

[0102] S35: Determine the sorting difference loss function according to the difference loss weight and sorting difference sample pairs.

[0103] Furthermore, based on the sorting difference sample pairs, determine the sorting difference weight , similarity difference weight and prior weight of the teacher model . After that, the difference loss weight of the sorting difference loss function can be calculated, and then the sorting difference loss function can be determined according to formula (7):

[0104] (7)

[0105] where is the difference between the output results corresponding to the th input sample and the th input sample in the sorting difference sample pairs of the student model.

[0106] Based on the above embodiments, as an alternative embodiment, when the sorting difference between the teacher sorting matrix and the student sorting matrix for the same sorting difference sample pairs is larger, the sorting difference weight is larger.

[0107] In a specific embodiment, the sorting difference weight is used to measure the sorting difference between the student model and the teacher model for input sample pairs at the same position. For example, for the th input sample and the th input sample of the input sample pair at the same position, when the sorting difference between the teacher sorting matrix and the student sorting matrix for the th input sample and the th input sample is larger, it means that the student model needs to pay more attention to the th input sample and the th input sample. Therefore, a larger sorting difference weight should be set.

[0108] In another alternative embodiment, the greater the similarity difference between the teacher similarity matrix and the student similarity matrix for the same sorted-difference sample pairs, the greater the similarity-difference weight.

[0109] In a specific embodiment, the similarity-difference weight is used to measure the similarity difference between the student model and the teacher model for the input sample pairs at the same position. When the similarity difference is greater, it indicates that the student model has learned the corresponding knowledge, and at this time, a higher similarity-difference weight should be set.

[0110] In an alternative embodiment, the smaller the similarity in the teacher similarity matrix for the sorted-difference sample pairs, the greater the prior weight of the teacher model.

[0111] In a specific embodiment, the prior weight of the teacher model is determined according to the teacher model similarity matrix. The smaller the similarity in the teacher similarity matrix, that is, the greater the difference between the sample pairs, a higher prior weight of the teacher model should be given at this time.

[0112] In an alternative embodiment, the difference loss weight is determined by the sorted-difference weight, the similarity-difference weight, and the prior weight of the teacher model, including:

[0113] Assign a first weight to the sorted-difference weight; assign a second weight to the similarity-difference weight; and assign a third weight to the prior weight of the teacher model;

[0114] Determine the difference loss weight according to the first weight, the second weight, the third weight, the sorted-difference weight, the similarity-difference weight, and the prior weight of the teacher model; wherein, as the current iteration number of the iterative training increases, the first weight increases, and the second weight and the third weight decrease.

[0115] To further improve the training effect of the target student model, in a specific embodiment, the sorted-difference weight, the similarity-difference weight, and the prior weight of the teacher model are linearly fused to obtain the difference loss weight. Specifically, for the sorted-difference weight assign a first weight ; for the similarity-difference weight assign a second weight ; and for the prior weight of the teacher model assign a third weight . Thus, the specific calculation formula for determining the difference loss weight is:

[0116] (8)

[0117] In an alternative embodiment, the sorted-difference weight , the similarity-difference weight , and the prior weight of the teacher model , it can also be achieved by adjusting the first weight , the second weight and the third weight . Specifically, as the current iteration number of iterative training increases continuously, the first weight is larger, and the second weight and the third weight are smaller. In a specific embodiment, increasing the first weight can strengthen the learning of difficult samples. While the second weight and the third weight can make the student model rely on its own learning results.

[0118] In an alternative embodiment, according to the teacher similarity matrix, contrastive learning sample pairs are constructed, including:

[0119] Samples with similarity greater than the first threshold in the teacher similarity matrix are used as positive sample pairs, and samples with similarity less than the second threshold are used as negative sample pairs; the first threshold is greater than the second threshold, and as the current iteration number of iterative training increases continuously, the first threshold is larger and the second threshold is smaller;

[0120] The positive samples and negative samples form contrastive learning sample pairs.

[0121] In a specific embodiment, as shown in Figure 2 , when constructing contrastive learning sample pairs, they are constructed based on the teacher similarity matrix and a preset first screening rule. Among them, the contrastive learning sample pairs include positive sample pairs and negative sample pairs. The positive sample pairs refer to the sample pairs corresponding to the similarity greater than the first threshold in the teacher similarity matrix, and the negative sample pairs refer to the samples corresponding to the similarity less than the second threshold in the teacher similarity matrix as negative sample pairs.

[0122] It can be understood that in the teacher similarity matrix , it describes the degree of difference between the th input sample and the th input sample. Therefore, when is greater than the first threshold, the th input sample and the th input sample can be used as a positive sample pair, and when is less than the second threshold, the th input sample and the th input sample can be used as a negative sample pair.

[0123] In an alternative embodiment, the first threshold is greater than the second threshold, and as the distillation loss function iteratively trains the student model, the first threshold and the second threshold can be dynamically adjusted. Specifically, as the current iteration number increases, the first threshold becomes larger and the second threshold becomes smaller. In a specific embodiment, at the initial stage of iterative training, relatively loose thresholds can be set to ensure that the student model has sufficient learning signals. As the training progresses, gradually increase the first threshold and decrease the second threshold, thereby strengthening the learning difficulty of the student model in order to improve the performance of the student model.

[0124] In an alternative embodiment, a rule for adjusting the first threshold and the second threshold can be preset. Specifically, preset the initial first threshold, the adjustment amplitude of the first threshold, the initial second threshold, and the adjustment amplitude of the second threshold, and obtain the current iteration number of the iterative training of the student model by the distillation loss function in real time, and calculate the target first threshold and the target second threshold for the next round of training according to the current iteration number. The specific calculation formulas are formulas (9) and (10):

[0125] (9)

[0126] (10)

[0127] Wherein, is the first threshold for the th round of iterative training, is the second threshold for the th round of iterative training, is the initial first threshold, is the initial second threshold, is the total number of iterations, is the adjustment amplitude of the first threshold, is the adjustment amplitude of the second threshold.

[0128] For example, in an alternative embodiment, the initial first threshold can be set to 0.7, and the initial second threshold can be set to 0.3. For each iterative training, different positive sample pairs and negative sample pairs can be obtained according to different first thresholds and second thresholds, that is, different contrast learning sample pairs including positive sample pairs and negative sample pairs can be obtained.

[0129] In an alternative embodiment, according to the sorting difference matrix, construct sorting difference sample pairs, including:

[0130] Determine the target object in the sorting difference matrix that meets the preset conditions; wherein, the preset conditions include that the similarity sorting difference is greater than the preset value, the first preset quantile in the sorting result of the similarity sorting differences from large to small, and any one of the first preset in the sorting result.

[0131] Use the samples corresponding to the target object as the sorted difference sample pairs.

[0132] In a specific embodiment, as Figure 2 shown, when constructing the sorted difference sample pairs, it is necessary to first determine the sorted difference matrix according to the teacher sorting matrix and the student sorting matrix. Further, select the target objects that meet the preset conditions from the sorted difference matrix. Specifically, it can be understood that the sorted difference matrix describes the similarity sorting difference degree between the teacher sorting matrix and the student sorting matrix for the input samples at the same position. Therefore, the objects with a similarity sorting difference value greater than the preset value in the sorted difference matrix can be used as the target objects.

[0133] In another alternative embodiment, the objects corresponding to the top preset quantiles in the sorting result of the similarity sorting difference values from large to small can also be used as the target objects. For example, the top 30% in the sorting result are used as the target objects. Of course, the top preset number in the sorting result can also be used as the target objects. For example, the top 40 are used as the target objects.

[0134] Further, use the samples corresponding to the selected target objects as the sorted difference samples. It should be noted that any one of the above preset conditions can be selected.

[0135] In the above embodiment, the model distillation method is described in detail. The present application also provides an embodiment corresponding to a model distillation device.

[0136] Figure 4 As shown in the structural schematic diagram of a model distillation device provided by an embodiment of the present application, Figure 4 shown, the device includes:

[0137] A matrix acquisition module 40, configured to acquire the teacher similarity matrix of the teacher model, the student similarity matrix of the student model, and the sorted difference matrix representing the similarity sorting difference between the teacher model and the student model; the calculation accuracy of the teacher model is higher than that of the student model;

[0138] A sample pair construction module 41, configured to construct contrast learning sample pairs according to the teacher similarity matrix; and construct sorted difference sample pairs according to the sorted difference matrix;

[0139] A distillation loss function construction module 42, configured to construct a distillation loss function based on the contrast learning sample pairs and the sorted difference sample pairs;

[0140] A student model training module 43, configured to iteratively train the student model through the distillation loss function to obtain the target student model.

[0141] In addition, the model distillation device provided by the embodiment of the present application further includes:

[0142] A task type acquisition module for acquiring the task type of a student model;

[0143] A label loss function determination module for determining the label loss function of the student model according to the task type;

[0144] A first determination module for determining a learning contrast loss function according to a contrast learning sample pair; and determining a sorting difference loss function according to a sorting difference sample pair;

[0145] A second determination module for determining a distillation loss function according to the label loss function, the learning contrast loss function, and the sorting difference loss function.

[0146] A sorting matrix determination module for determining a teacher sorting matrix, a student sorting matrix, and a similarity difference matrix according to a teacher similarity matrix and a student similarity matrix;

[0147] A sorting difference weight determination module for determining a sorting difference weight according to the teacher sorting matrix, the student sorting matrix, and the sorting difference sample pair;

[0148] A similarity difference weight determination module for determining a similarity difference weight according to the similarity difference matrix and the sorting difference sample pair;

[0149] A teacher model prior weight determination module for determining a teacher model prior weight according to the teacher similarity matrix and the sorting difference sample pair;

[0150] A difference loss weight determination module for determining a difference loss weight through the sorting difference weight, the similarity difference weight, and the teacher model prior weight;

[0151] The second determination module is further used for determining a sorting difference loss function according to the difference loss weight and the sorting difference sample pair.

[0152] A weight assignment model for assigning a first weight to the sorting difference weight; assigning a second weight to the similarity difference weight; and assigning a third weight to the teacher model prior weight;

[0153] The difference loss weight determination module is further used for determining a difference loss weight according to the first weight, the second weight, the third weight, the sorting difference weight, the similarity difference weight, and the teacher model prior weight; wherein, as the current iteration number of the iterative training increases, the first weight increases, and the second weight and the third weight decrease.

[0154] A contrastive learning sample pair acquisition module, configured to use the samples corresponding to the similarity greater than the first threshold in the teacher similarity matrix as positive sample pairs, and the samples corresponding to the similarity less than the second threshold as negative sample pairs; the first threshold is greater than the second threshold, and as the current iteration number of the iterative training increases, the greater the first threshold, the smaller the second threshold; the positive samples and the negative samples constitute the contrastive learning sample pairs.

[0155] A sorting difference sample pair acquisition module, configured to determine a target object that meets a preset condition in a sorting difference matrix; wherein, the preset condition includes any one of a similarity sorting difference greater than a preset value, the first preset quantile in the sorting result of the similarity sorting differences from large to small, and the first preset number in the sorting result.

[0156] Use the sample corresponding to the target object as the sorting difference sample pair.

[0157] Figure 5 The following is a schematic structural diagram of a model distillation device provided in another embodiment of the present application, as Figure 5 shown, the model distillation device includes: a memory 50, configured to store a computer program;

[0158] A processor 51, configured to implement the steps of the model distillation method mentioned in the above embodiment when executing the computer program.

[0159] The model distillation device provided in this embodiment may include, but is not limited to, a laptop computer or a desktop computer, etc.

[0160] Among them, the processor 51 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 51 may be implemented in at least one of the following hardware forms: a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 51 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 51 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 51 may further include an artificial intelligence (AI) processor, and the AI processor is used to process the computational operations related to machine learning.

[0161] The memory 50 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 50 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 50 is at least used to store the following computer program 501. After the computer program is loaded and executed by the processor 51, it can implement the relevant steps of the model distillation method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 50 may further include an operating system 502 and data 503, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 502 may include Windows, Unix, Linux, etc. The data 503 may include, but is not limited to, the relevant data involved in the model distillation method.

[0162] In some embodiments, the model distillation device may further include a display screen 52, an input / output interface 53, a communication interface 54, a power supply 55, and a communication bus 56.

[0163] Those skilled in the art can understand that Figure 5 the structure shown in

[0164] The model distillation device provided by the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the model distillation device in the above embodiment.

[0165] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

Claims

1. A model distillation method, characterized in that: The method comprises: Obtaining a teacher similarity matrix of the teacher model, a student similarity matrix of the student model, and a ranking difference matrix representing the ranking difference of similarity between the teacher model and the student model; the calculation accuracy of the teacher model is higher than the calculation accuracy of the student model; the teacher model and the student model are both residual networks applied to image classification tasks, and the number of network layers of the teacher model is greater than the number of network layers of the student model; the similarity matrix is ​​composed of similarity values ​​between output results, and is used to represent the degree of difference between output samples; According to the teacher similarity matrix, constructing comparative learning sample pairs; and according to the sorting difference matrix, constructing sorting difference sample pairs; Constructing a distillation loss function based on the contrastive learning sample pair and the sorting difference sample pair; Iteratively training the student model through the distillation loss function to obtain a target student model; Based on the contrastive learning sample pair and the sorting difference sample pair, a distillation loss function is constructed, including: Obtaining the task type of the student model; Determining a label loss function of the student model according to the task type; Determine a learning contrast loss function based on the contrast learning sample pair; and determine a sorting difference loss function based on the sorting difference sample pair; Determining the distillation loss function according to the label loss function, the learning contrast loss function and the sorting difference loss function; According to the sorting difference matrix, constructing sorting difference sample pairs, including: Determine the target object that meets the preset conditions in the sorting difference matrix; wherein the preset conditions include any one of the similarity sorting difference being greater than a preset value, the first preset quantile of the sorting result from large to small of the similarity sorting difference, and the first preset number of the sorting result; The samples corresponding to the target object are used as the sorting difference sample pairs.

2. The model distillation method according to claim 1, characterized in that Determining a sorting difference loss function according to the sorting difference sample pair includes: Determine a teacher ranking matrix, a student ranking matrix and a similarity difference matrix according to the teacher similarity matrix and the student similarity matrix; Determining a ranking difference weight according to the teacher ranking matrix, the student ranking matrix and the ranking difference sample pairs; Determining a similarity difference weight according to the similarity difference matrix and the ranked difference sample pairs; Determining a priori weights of a teacher model according to the teacher similarity matrix and the ranked difference sample pairs; Determine a difference loss weight by using the ranking difference weight, the similarity difference weight and the teacher model prior weight; The sorting difference loss function is determined according to the difference loss weight and the sorting difference sample pair.

3. The model distillation method according to claim 2, characterized in that: The greater the ranking difference between the teacher ranking matrix and the student ranking matrix between the same ranking difference sample pairs, the greater the ranking difference weight; When the similarity difference between the teacher similarity matrix and the student similarity matrix in the same sorted difference sample pair is greater, the similarity difference weight is greater; The smaller the similarity of the sorted difference sample pair in the teacher similarity matrix, the greater the prior weight of the teacher model.

4. The model distillation method according to claim 2, characterized in that: Determining a difference loss weight by using the ranking difference weight, the similarity difference weight and the teacher model prior weight includes: Assigning a first weight to the ranking difference weight; assigning a second weight to the similarity difference weight; and assigning a third weight to the teacher model prior weight; The difference loss weight is determined according to the first weight, the second weight, the third weight, the sorting difference weight, the similarity difference weight and the teacher model prior weight; wherein, as the current iteration number of the iterative training continues to increase, the larger the first weight is, the smaller the second weight and the third weight are.

5. The model distillation method according to claim 1, characterized in that: According to the teacher similarity matrix, a comparative learning sample pair is constructed, including: The samples corresponding to the similarity greater than the first threshold in the teacher similarity matrix are taken as positive sample pairs, and the samples corresponding to the similarity less than the second threshold are taken as negative sample pairs; wherein the first threshold is greater than the second threshold, and as the current iteration number of the iterative training increases, the greater the first threshold, the smaller the second threshold; The positive sample and the negative sample constitute the contrastive learning sample pair.

6. A model distillation apparatus, characterized in that: The device comprises: A matrix acquisition module, used to obtain a teacher similarity matrix of a teacher model, a student similarity matrix of a student model, and a ranking difference matrix representing the ranking difference of similarity between the teacher model and the student model; the calculation accuracy of the teacher model is higher than the calculation accuracy of the student model; the teacher model and the student model are both residual networks applied to image classification tasks, and the number of network layers of the teacher model is greater than the number of network layers of the student model; the similarity matrix is ​​composed of similarity values ​​between output results, and is used to represent the degree of difference between output samples; A sample pair construction module, used to construct comparative learning sample pairs according to the teacher similarity matrix; and to construct sorting difference sample pairs according to the sorting difference matrix; A distillation loss function construction module, used to construct a distillation loss function based on the contrastive learning sample pair and the sorting difference sample pair; A student model training module, used for iteratively training the student model through the distillation loss function to obtain a target student model; A task type acquisition module, used to acquire the task type of the student model; A label loss function determination module, used to determine the label loss function of the student model according to the task type; A first determination module is used to determine a learning contrast loss function according to the contrast learning sample pair; and to determine a sorting difference loss function according to the sorting difference sample pair; A second determination module is used to determine the distillation loss function according to the label loss function, the learning contrast loss function and the sorting difference loss function; A sorting difference sample pair acquisition module is used to determine the target object that meets the preset conditions in the sorting difference matrix; wherein the preset conditions include the similarity sorting difference being greater than the preset value, the first preset quantile in the similarity sorting difference sorting result from large to small, and any one of the first preset quantiles in the sorting result; and the samples corresponding to the target object are used as the sorting difference sample pair.

7. A model distillation device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the model distillation method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the model distillation method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Video text cross-modal retrieval method and device, equipment and medium

    CN117725261A

  • Rank Distillation for Training Supervised Machine Learning Models

    US20230206134A1