Multi-Teacher Knowledge Distillation Method and Device Based on Adversarial Examples

By generating boundary sample pairs and optimizing the weight allocation of teacher model based on boundary distance, the problem of unreasonable weight allocation in knowledge distillation of multiple teachers is solved, and the knowledge transfer efficiency and classification accuracy of student models are improved.

CN114219043BActive Publication Date: 2025-07-08HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111568528.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-07-08
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

The existing multi-teacher knowledge distillation method lacks effective indicators when measuring the knowledge transfer efficiency of the teacher model, resulting in unreasonable weight allocation, affecting the knowledge transfer efficiency and the accuracy of the student model.

Method used

By generating boundary sample pairs located on both sides of the decision boundary of the teacher model, using the point-to-line distance vector algorithm for iterative modification, calculate the weight allocation of the teacher model on the boundary sample, and optimize student model training.

Benefits of technology

It improves the efficiency of knowledge transfer, shortens the knowledge distillation process, accelerates the improvement of the classification accuracy of the student model, and has practical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114219043B_ABST
    Figure CN114219043B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-teacher knowledge distillation method, device and computer storage medium based on adversarial samples. The method includes: selecting an original sample to be modified based on the principle of maximizing the difference between the teacher probability output and the student probability output for the sample; taking the classification with the highest classification probability of the original sample to be modified on the teacher model as the target classification for adversarial attack and the corresponding original sample to be modified as the original sample that can be modified; obtaining the decision boundary of the teacher model based on the classification probability of the class of the original sample that can be modified by the teacher model, and using the vector algorithm of the distance from a point to a line, taking the original sample that can be modified just crossing the decision boundary and just not crossing the decision boundary as the goal, iteratively modifying the original sample that can be modified to generate a pair of boundary samples located on both sides of the decision boundary; using the generated boundary samples to train the student model using multi-teacher weight allocation based on the boundary distance. The present invention can improve the classification accuracy of the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to knowledge distillation of deep network models, and specifically relates to a multi-teacher knowledge distillation method, device, and computer storage medium based on adversarial samples. Background Art

[0002] With the in-depth research on deep neural networks (DNNs), deep networks are applied to more and more complex problems, and the depth and breadth of the networks are becoming larger and larger. However, the huge parameter scale not only causes difficulties in training, but also greatly increases the time spent in the inference stage, making the network model unable to be deployed on devices with weak computing power such as personal computers. Therefore, many recent works are dedicated to researching how to compress huge deep networks into lighter networks, and one of the methods is knowledge distillation (KD).

[0003] Knowledge distillation is a way to achieve knowledge transfer. It trains a simple model by using the outputs of a trained complex model, so as to achieve the effects of model compression and improvement of the accuracy of the simple model. In this process, the complex model is called the teacher model and the simple model is called the student model. Multi-teacher knowledge distillation is a branch of knowledge distillation, which refers to fusing the outputs of multiple teacher models and then applying them in knowledge distillation to improve the accuracy of the student model. At the same time, since knowledge distillation is actually a process in which the student model learns the decision boundary of the teacher model, the samples closer to the decision boundary have higher learning efficiency.

[0004] At present, the vast majority of multi-teacher distillation methods adopt the practice of averaging the distillation losses of each teacher when measuring the proportion of each teacher. This is because there is a lack of an indicator to judge the level of the role played by teachers in knowledge distillation. Whether it is single-teacher distillation or multi-teacher distillation, more attention should be paid to the amount and transfer efficiency of the dark knowledge contained in the soft labels of the teacher models, rather than whether the classification results of the teacher models are correct. Some conventional indicators, such as classification accuracy or prediction probability on correct classifications, etc., cannot measure the learning value of teachers in knowledge distillation. Even if the teacher model makes a classification error, the predicted probability it outputs still contains a lot of dark knowledge worth learning. Therefore, in the existing knowledge distillation methods, there is no indicator that can show which teacher is more worthy of learning, and the weights of each teacher in distillation have to be considered equal. However, from the perspective of knowledge transfer efficiency, due to the different distances of samples from the decision boundaries of each teacher, the knowledge transfer efficiency of the same sample to each teacher is different. Obviously, this equal-treatment approach makes the possible advantages of some teachers disappear compared with other teachers. The outputs of teachers with different knowledge transfer efficiencies are given the same weight, resulting in the inability to fully carry out knowledge transfer. Therefore, exploring how to reasonably allocate the weights of each teacher according to the knowledge transfer efficiency of different samples is one of the keys to improving the performance of multi-teacher knowledge distillation. Summary of the Invention

[0005] In view of the above problems, the present invention provides a multi-teacher knowledge distillation method, device and computer storage medium based on adversarial samples.

[0006] In the first aspect of the present invention, a multi-teacher knowledge distillation method based on adversarial samples is provided. The method includes the following steps:

[0007] Based on the principle of maximizing the difference between the teacher probability output and the student probability output for each batch of samples, select a part of the samples as the original samples to be modified;

[0008] Take the classification with the highest classification probability of the original samples to be modified on the teacher model as the target classification for adversarial attack, and the original samples to be modified corresponding to the target classification as the original samples that can be modified;

[0009] Based on the classification probability of the teacher model for the categories of the original samples that can be modified, obtain the decision boundary of the teacher model. Using the vector algorithm of the distance from a point to a line, with the goal that the original samples that can be modified just cross the decision boundary and just do not cross the decision boundary, iteratively modify the original samples that can be modified to generate boundary sample pairs located on both sides of the decision boundary;

[0010] Use the generated boundary samples to train the student model using multi-teacher weight allocation based on boundary distance.

[0011] Further, the specific method for selecting the original samples to be modified is as follows: the samples need to satisfy that the classification results of the teacher model for the samples are the same as those of the student model for the samples. When there are an excessive number of samples with the same classification results, the original samples to be modified are selected in the order of the highest priority of the difference between the classification probabilities of the teacher model and the student model among the samples with the same classification results.

[0012] Further, the specific steps for generating the boundary sample pairs on both sides of the decision boundary include:

[0013] The modification result where the modifiable original sample just crosses the boundary is called the outer sample, and the modification result where the modifiable original sample just does not cross the boundary is called the inner sample;

[0014] The iteration formula for the outer sample is:

[0015]

[0016] Among them, is the vector differential operator, η is the learning rate less than 1, ε represents the hyperparameter, and the initial value of the outer sample is the modifiable original sample. respectively represent the probabilities of the teacher model f for the sample on the original class c0 and other classes c;

[0017] The iteration of the outer sample ends when it satisfies one of (1) and (2), where: (1): and (2): i + 1 > I max , i is the number of iterations, and I max is the preset maximum number of iterable times.

[0018] Further, the specific method for obtaining the inner sample includes:

[0019] If after the iteration of the outer sample ends, the previous sample of the final outer sample satisfies end represents the final number of iterations of the outer sample, and the inner sample directly takes the value Otherwise, the initial value of the inner sample is and the inner sample is iteratively calculated. The iteration formula is:

[0020]

[0021] Among them, η j is the variable learning rate, and its initial value η0 is the same as η. If it satisfies after the (i + 1)-th iteration, then η j+1 = η j / 2, and recalculate Until Perform the next iteration, where j represents the number of times of learning rate decay;

[0022] The iteration ends when the inner samples satisfy one of (3) to (5), where: (3): (4): i + 1 > I max ; (5): j + 1 > J max , where x out represents the outer samples obtained after the iteration of the outer samples J max is the preset maximum number of times the learning rate can decay.

[0023] Furthermore, using the generated boundary samples, train the student model using multi-teacher weight assignment based on the boundary distance, and the specific steps include:

[0024] For each generated boundary sample, calculate the weight of each teacher in the training of the student model according to the ratio of the classification probabilities of each teacher model on the target classification and the original classification.

[0025] Furthermore, the weight of each teacher model in the training of the student model is specifically expressed as:

[0026]

[0027] where N represents the number of teacher models, h n (x) represents the score of the nth teacher model f n (·) when the student is learning the boundary sample x,

[0028]

[0029] where and are the classification probabilities of the teacher model for the boundary sample x in categories c0 and c.

[0030] Furthermore, the method also includes using the weight assignment of each teacher model in the training of the student model to calculate the proportion of the loss generated by each teacher model in the training of the student model.

[0031] In the second aspect of the present invention, a multi-teacher knowledge distillation device based on adversarial samples is provided, and the device includes:

[0032] An original sample acquisition model to be modified, which is used to select a part of the samples as the original samples to be modified based on the principle of maximizing the difference between the teacher probability output and the student probability output for each batch of samples;

[0033] The original sample modification acquisition model can be used to take the classification with the highest classification probability of the original sample to be modified on the teacher model as the target classification of the adversarial attack, and the original sample to be modified corresponding to the target classification is used as the original sample to be modified;

[0034] The boundary sample pair generation module obtains the decision boundary of the teacher model based on the classification probability of the teacher model for the category of the original sample to be modified. Using the vector algorithm of the distance from a point to a line, with the goal that the original sample to be modified just crosses the decision boundary and just does not cross the decision boundary, the original sample to be modified is iteratively modified to generate boundary sample pairs located on both sides of the decision boundary;

[0035] The student model training module is used to utilize the generated boundary samples to train the student model using multi-teacher weight assignment based on the boundary distance.

[0036] In the third aspect of the present invention, a multi-teacher knowledge distillation device based on adversarial samples is provided, including: a processor; and a memory, wherein, computer-executable programs are stored in the memory, and when the computer-executable programs are executed by the processor, the above-mentioned multi-teacher knowledge distillation method based on adversarial samples is executed.

[0037] In the fourth aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned multi-teacher knowledge distillation method based on adversarial samples.

[0038] A multi-teacher knowledge distillation method, device and computer storage medium based on adversarial samples provided by the present invention adopt a method similar to the generation of adversarial samples, adding subtle changes to the original sample to create samples as close as possible to the decision boundary of a certain teacher model. To make the samples more adaptable to non-average weight multi-teacher distillation, a new type of sample group called boundary sample pairs is proposed, and the existing adversarial sample generation method is improved to obtain boundary sample pairs. Compared with previous adversarial samples, boundary sample pairs have better effects in knowledge distillation. This method uses the teacher model to classify samples, calculates the distance-based score based on the classification probabilities of the teacher for the boundary samples in the original classification and the adversarial attack target classification, and assigns weights by the score. The finally achieved beneficial effects: Compared with the existing multi-teacher distillation methods, the multi-teacher knowledge distillation method, device and computer storage medium based on adversarial samples provided by the present invention improve the efficiency of knowledge transfer, thereby accelerating the process of knowledge distillation, improving the classification accuracy of the student model, and having great practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flowchart of the multi-teacher knowledge distillation method based on adversarial samples according to an embodiment of the present invention;

[0040] Figure 2 Schematic diagram of the structure of the multi-teacher knowledge distillation device based on adversarial samples according to an embodiment of the present invention;

[0041] Figure 3 Architecture of the computer device according to an embodiment of the present invention;

[0042] Figure 4 Comparison chart of the classification accuracy of the student model on the CIFAR-10 dataset with other methods according to an embodiment of the present invention;

[0043] Figure 5 Comparison chart of the classification accuracy of the student model on the ImageNet dataset with other methods according to an embodiment of the present invention;

[0044] Figure 6 Comparison chart of the loss function curves during distillation between the multi-teacher knowledge distillation method based on adversarial samples and the ordinary multi-teacher distillation method according to an embodiment of the present invention. Detailed implementation manners

[0045] To further elaborate on the technical solution of the present invention in detail, this embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific steps are given.

[0046] Based on Embodiment 1 of the present invention

[0047] The specific steps of a multi-teacher knowledge distillation method based on adversarial samples in this embodiment are as follows Figure 1 Shown as the flowchart of the multi-teacher knowledge distillation method based on adversarial samples according to an embodiment of the present invention:

[0048] S1. Based on the principle of maximizing the difference between the teacher probability output and the student probability output for each batch of samples, select a part of the samples as the original samples to be modified;

[0049] Furthermore, the specific way to select the original samples to be modified is: the samples need to satisfy that the classification results of the teacher model for the samples are the same as those of the student model for the samples. When there are too many samples with the same classification results, the original samples to be modified are selected in the order of the highest priority of the difference between the classification probabilities of the teacher model and the student model among the samples with the same classification results. This can make the spatial distance between the selected classification probability vectors of the teacher model and the student model as large as possible.

[0050] During the specific implementation process, in order to control the total number M bs of the boundary samples to be equal to or slightly less than the number M of each batch of data in batch training batch . For this purpose, when the number of teachers is N, at most Generate boundary sample pairs from a single original sample. The selection condition for these at most M samples is that the classification results of the teacher for the samples are the same as those of the student. Of course, there may be more than M samples in a batch that meet this condition. In this case, select the top M largest samples.

[0051] S2. Use the classification with the highest classification probability of the original sample to be modified on the teacher model as the target classification for adversarial attacks, and use the original sample to be modified corresponding to the target classification as the original sample to be modified.

[0052] In the specific implementation process, for each selected original sample to be modified, select the target classification for adversarial attacks according to the classification probabilities of each class on the teacher model; the selected target classification for adversarial attacks is the classification with the highest classification prediction probability of the teacher model except for the original classification.

[0053] S3. Obtain the decision boundary of the teacher model based on the classification probabilities of the classes of the original sample to be modified. Using the vector algorithm of the distance from a point to a line, with the goal that the original sample to be modified just crosses the decision boundary and just does not cross the decision boundary, iteratively modify the original sample to be modified to generate boundary sample pairs located on both sides of the decision boundary.

[0054] In the specific implementation process, based on the latent space of the data manifold, denote the probability scores of the teacher model f for the sample x in the original class c0 and a certain other class c as and f c (x), use the difference between the two to measure the distance of the sample to the boundary. The surface represented by F c (x)=0 is the decision boundary surface.

[0055] Furthermore, denote the modification result where the original sample to be modified just crosses the boundary as the outer sample, and denote the modification result where the original sample to be modified just does not cross the boundary as the inner sample.

[0056] S31. The iterative formula for the outer sample is:

[0057]

[0058] where is the vector differential operator, η is the learning rate less than 1 to prevent the iteration step size from being too large when the estimated gradient is greater than the actual gradient, and ε represents the hyperparameter used to ensure that the outer sample can cross the decision boundary. The initial value of the outer sample is the original sample to be modified, respectively represent the probabilities of the teacher model f for the sample in the original class c0 and the other class c;

[0059] The iteration ends if the outer sample satisfies one of (1) and (2), where: (1): and (2): i + 1 > I max , where i is the iteration number and I max is the preset maximum number of iterations.

[0060] S32. The specific method for obtaining the inner sample includes:

[0061] If, after the iteration of the outer sample ends, the previous sample of the final outer sample satisfies where end represents the final iteration number of the outer sample, the inner sample directly takes the value that is which may be closer to the boundary compared to the outer sample. Otherwise, the initial value of the inner sample is and iterative calculation is performed on the inner sample. The iterative formula is:

[0062]

[0063] where η j is the variable learning rate, whose initial value η0 is the same as η. If, after the (i + 1)-th iteration, it satisfies then η j+1 = η j / 2, and recalculate until and then perform the next iteration. j represents the number of times the learning rate decays;

[0064] The iteration of the inner sample ends if it satisfies one of (3) to (5), where: (3): (4): i + 1 > I max ; (5): j + 1 > J max , where x out represents the outer sample obtained after the iteration of the outer sample ends J max is the preset maximum number of times the learning rate can decay. This can achieve that the outer sample does not cross the boundary, and the inner sample is closer to the decision boundary than the outer sample.

[0065] Based on Embodiment 2 of the present invention

[0066] This embodiment is used to execute S4 on the basis of Embodiment 1 and train the student model using the generated boundary sample pairs with a multi-teacher weight assignment based on the boundary distance. A weight assignment method based on the boundary distance is provided for the training of the student model, including: using the ratio of the probabilities of two classes in binary classification to quantify the distance of a sample to the boundary. When the boundary sample completely falls on the decision boundary, the ratio of the two classifications will be 1, and the farther the boundary sample is from the boundary, the larger the ratio of the larger probability to the smaller probability will be; when this ratio gradually increases, the weight of the teacher model should rapidly decrease. Normalize the scores of each teacher model on this boundary sample to obtain their respective weights, and assign the coefficients of the distillation losses of each teacher model in multi-teacher knowledge distillation according to the weights.

[0067] In the specific implementation process, normalize the scores of each teacher model on the boundary sample to obtain their respective weights. The specific expression of the weight of each teacher model in the training of the student model is:

[0068]

[0069] where N represents the number of teacher models, h n (x) represents the score of the nth teacher model f n (·) when the student is learning the boundary sample x,

[0070]

[0071] where and are the classification probabilities of the teacher model for the boundary sample x in classes c0 and c.

[0072] Based on Embodiment 3 of the present invention

[0073] Hereinafter, with reference to Figure 2 to describe the Figure 1The apparatus corresponding to the method shown, a multi-teacher knowledge distillation apparatus based on adversarial samples, the apparatus 100 includes: an original sample to be modified acquisition model 101, which is used to select a part of the samples as the original samples to be modified based on the principle of maximizing the difference between the teacher probability output and the student probability output for each batch of samples; a modifiable original sample acquisition model 102, which is used to take the classification with the highest classification probability of the original samples to be modified on the teacher model as the target classification for adversarial attacks, and the original samples to be modified corresponding to the target classification as the modifiable original samples; a boundary sample pair generation module 103, which obtains the decision boundary of the teacher model based on the classification probability of the modifiable original samples by the teacher model, and uses the vector algorithm of the distance from a point to a line to iteratively modify the modifiable original samples with the goal of exactly crossing the decision boundary and exactly not crossing the decision boundary of the modifiable original samples, and generates boundary sample pairs located on both sides of the decision boundary; a student model training module 104, which is used to train the student model by using the boundary sample pairs generated by the boundary sample pair generation module 103. In addition to these 4 units, the apparatus 100 may also include other components. However, since these components are not related to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.

[0074] The specific working process of a multi-teacher knowledge distillation apparatus 100 based on adversarial samples refers to the descriptions of Embodiment 1 and Embodiment 2 of the above-mentioned multi-teacher knowledge distillation method based on adversarial samples, and will not be elaborated here.

[0075] Based on Embodiment 4 of the present invention

[0076] The apparatus according to the embodiments of the present invention can also be realized by means of Figure 3 the architecture of the computing device shown. Figure 3 The architecture shown is only exemplary. When implementing different devices, adjust one or more components in Figure 3 according to actual needs. Figure 3 The architecture shown is only exemplary. When implementing different devices, adjust one or more components in Figure 3 according to actual needs.

[0077] Based on Embodiment 5 of the present invention

[0078] The embodiments of the present invention can also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium according to Embodiment 5. When the computer-readable instructions are run by a processor, the methods according to Embodiments 1-2 of the present invention described with reference to the above drawings can be executed.

[0079] Embodiment 1 - Embodiment 5 of the multi - teacher knowledge distillation method based on adversarial examples, as well as the device embodiment and computer storage medium embodiment of the present invention. When comparing the results of the above 5 embodiments with the current optimal multi - teacher knowledge distillation methods Ensemble, Triplet, and FEED in terms of the classification accuracy of the student model, the embodiments are carried out on two real - world datasets CIFAR - 10 and ImageNet. The introductions of the two example datasets are as follows:

[0080] CIFAR - 10 dataset: A color image dataset containing 10 common categories such as airplanes, cars, birds, dogs, etc. Each image in CIFAR - 10 has a size of 32 pixels * 32 pixels and consists of 3 channels in RGB mode. The total dimension size of the samples is 3072. The training set contains 50,000 images in total, and the test set contains 10,000 images in total.

[0081] ImageNet dataset: A high - resolution image dataset with a tree - like classification structure and a sample size reaching tens of millions. The dataset used in the verification of the present invention is its most widely used subset ISLVRC2012. Each image has a size of 299 pixels * 299 pixels, consists of 3 channels, and the total dimension size is 268,203. The training set contains nearly 1.3 million images in total, and the test set contains 50,000 images in total.

[0082] The classification accuracy and loss of the attack algorithm of the embodiments of the present invention on the two datasets are as Figure 4 、 Figure 5 and Figure 6 shown.

[0083] From Figure 4 and Figure 5 , under the same attack settings of other adversarial attack methods, the experimental results prove that the present method has the best performance on the CIFAR - 10 dataset, outperforming the best - performing multi - teacher distillation method FEED last year. For the ImageNet dataset, although the initial method of this project performs slightly worse than the FEED method, since the present invention does not modify the teacher - student framework, it can be freely combined with other multi - teacher methods, which is one of the advantages of this method. When the method of the present invention is combined with the FEED architecture to form the Our + method, it further shows a greater performance advantage. The Ours + method still has the best performance among various knowledge distillation methods. From Figure 6 , it can be seen that the method of the present invention converges faster than the ordinary ensemble method and the FEED method. This indicates that the present method has an advantage over ordinary methods in terms of the overall knowledge transfer efficiency.

[0084] Based on the multi-teacher knowledge distillation method, device and computer storage medium for adversarial examples provided in the above embodiments, boundary sample pairs more suitable for multi-teacher knowledge distillation can be generated, and the efficiency of knowledge transfer can be improved through the boundary sample pairs and the weight assignment method based on boundary distance, accelerating the process of knowledge distillation and improving the classification accuracy of the student model.

[0085] In this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a step or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such a step or method.

[0086] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A multi-teacher knowledge distillation method based on adversarial examples, characterized in that, The method includes the following steps: Based on the principle of maximizing the difference between the teacher probability output and the student probability output for each batch of samples, select a part of the samples as the original samples to be modified; Take the classification with the highest classification probability of the original samples to be modified on the teacher model as the target classification of the adversarial attack, and the original samples to be modified corresponding to the target classification as the original samples that can be modified; Based on the classification probability of the teacher model for the category of the original samples that can be modified, obtain the decision boundary of the teacher model. Using the vector algorithm of the distance from a point to a line, with the goal that the original samples that can be modified just cross the decision boundary and just do not cross the decision boundary, iteratively modify the original samples that can be modified to generate boundary sample pairs on both sides of the decision boundary; Using the generated boundary sample pairs, train the student model according to the multi-teacher weight allocation based on the boundary distance; Among them, the samples are all image data.

2. The multi-teacher knowledge distillation method according to claim 1, wherein The specific way to select the original samples to be modified is as follows: The samples need to satisfy that the classification result of the teacher model for the samples is the same as the classification result of the student model for the samples. When there are too many samples with the same classification result, select the original samples to be modified in the order of the highest priority of the difference between the classification probability of the teacher model and the classification probability of the student model among the samples with the same classification result.

3. The multi-teacher knowledge distillation method according to claim 1, characterized in that The specific steps to generate boundary sample pairs on both sides of the decision boundary include: Call the modification result where the original samples that can be modified just cross the boundary the outer samples, and call the modification result where the original samples that can be modified just do not cross the boundary the inner samples; The iterative formula for the outer samples is: Among them, is the vector differential operator, η is the learning rate less than 1, ε represents the hyperparameter, and the initial value of the outer sample is the modifiable original sample. respectively represent the probabilities of the teacher model f for the sample on the original class c0 and other classes c; The iteration ends if the outer sample satisfies one of (1) and (2), where: (1): and (2): i + 1 > I max , where i is the number of iterations and I max is the preset maximum number of iterations.

4. The multi-teacher knowledge distillation method according to claim 3, wherein The specific way to obtain the inner samples includes: If after the iteration of the outer sample ends, the final outer sample and the previous sample meet end represents the final number of iterations of the outer sample, and the inner sample directly takes the value Otherwise, the initial value of the inner sample is and iterative operations are performed on the inner sample. The iterative formula is: Among them, η j is a variable learning rate, whose initial value η0 is the same as η. If the following condition is satisfied after the (i + 1)-th iteration then η j+1 = η j / 2, and recalculate until Then perform the next iteration. j represents the number of times of learning rate decay; The iteration ends if the inner sample satisfies one of (3) to (5), where: (3): (4): i + 1 > I max ; (5): j + 1 > J max , where x out represents the outer sample obtained after the iteration of the outer sample ends J max is the maximum number of times the preset learning rate can decay.

5. The multi-teacher knowledge distillation method according to claim 1, wherein Using the generated boundary samples, train the student model using the multi-teacher weight allocation based on the boundary distance. The specific steps include: For each generated boundary sample, calculate the weight of each teacher in the training of the student model according to the ratio of the classification probabilities of each teacher model for the target classification and the original classification.

6. The multi-teacher knowledge distillation method according to claim 5, wherein The weight of each teacher model in the training of the student model, the specific expression is: Among them, N represents the number of teacher models, h n (x) represents the score of the nth teacher model f n (·) when the student is learning the boundary sample x, where and are the classification probabilities of the teacher model for the boundary sample x in classes c0 and c, respectively.

7. The multi-teacher knowledge distillation method according to claim 6, wherein, The method also includes using the weight allocation of each teacher model in the training of the student model to allocate the proportion of the loss generated by each teacher model in the training of the student model.

8. A multi-teacher knowledge distillation device based on adversarial examples, characterized in that The device includes: A model for obtaining the original samples to be modified, which is used to select a part of the samples as the original samples to be modified based on the principle of maximizing the difference between the teacher probability output and the student probability output for each batch of samples; A model for obtaining the original samples that can be modified, which is used to take the classification with the highest classification probability of the original samples to be modified on the teacher model as the target classification of the adversarial attack, and the original samples to be modified corresponding to the target classification as the original samples that can be modified; A boundary sample pair generation module, which obtains the decision boundary of the teacher model based on the classification probability of the teacher model for the category of the original samples that can be modified, and uses the vector algorithm of the distance from a point to a line. With the goal that the original samples that can be modified just cross the decision boundary and just do not cross the decision boundary, iteratively modify the original samples that can be modified to generate boundary sample pairs on both sides of the decision boundary; A student model training module, which is used to use the generated boundary sample pairs to train the student model using the multi-teacher weight allocation based on the boundary distance; Among them, the samples are all image data.

9. A multi-teacher knowledge distillation device based on adversarial examples, characterized in that, Includes: A processor; and a memory, wherein the memory stores a computer-executable program, and when the computer-executable program is executed by the processor, the method according to any one of claims 1-7 is executed.

10. A computer-readable storage medium having instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 1-7.