A model distillation training method, device, electronic device and storage medium

By combining data enhancement, loss optimization and knowledge distillation in neural network models, the problem of low accuracy of long-tail category samples recognition is solved, and the high accuracy recognition effect under unbalanced data distribution is achieved.

CN113935234BActive Publication Date: 2025-06-24ZHONGKE DINGFU BEIJING TECH DEV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111186029.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2025-06-24
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

During the training of neural network models, the sample recognition accuracy of long-tail categories is low, and the existing methods have limited effects on increasing the number of data strips or reducing the number of data strips in non-long-tail categories.

Method used

By obtaining training data sets including long-tail categories, using multiple data augmentation methods to perform data augmentation, multiple data sets are obtained, and different types of loss optimization training are performed on multiple teacher models to obtain the trained teacher model. Then, the student model is distilled and trained to screen out the student model with the highest accuracy rate to improve the sample recognition accuracy of the long-tail category.

Benefits of technology

Through the combination of data enhancement, loss optimization and knowledge distillation, the accuracy of recognition of long-tail category samples by student models under uneven data distribution is effectively improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113935234B_ABST
    Figure CN113935234B_ABST
Patent Text Reader

Abstract

The present application provides a model distillation training method, apparatus, electronic device and storage medium, which are used to improve the problem that the improvement of the recognition accuracy rate of samples of long-tailed categories is very limited. The method includes: obtaining a training data set including long-tailed categories, and performing data augmentation on the training data set by using a variety of data augmentation means to obtain a plurality of data sets; using the plurality of data sets to perform different types of loss optimization training on a plurality of teacher models respectively to obtain a plurality of trained teacher models, where one teacher model is obtained by using one data set to perform one type of loss optimization training; selecting a preset number of teacher models from the plurality of teacher models according to the accuracy rate; obtaining a preset number of student models, and using the teacher models to perform distillation training on the student models to obtain a preset number of distilled student models; screening out the student model with the highest accuracy rate from the preset number of distilled student models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of neural networks, deep learning, data augmentation, and knowledge distillation. Specifically, it relates to a model distillation training method, apparatus, electronic device, and storage medium. Background Art

[0002] Long-tail classes, also known as few-shot classes, refer to classes with a small number of samples in the training dataset of a model. Specifically, for example: if the evaluations on a food delivery platform are used as the training dataset, then the number of positive evaluations in this training dataset is usually much larger than the number of negative evaluations. Therefore, the negative evaluation class here can be understood as the above-mentioned long-tail class.

[0003] Currently, during the training process of a neural network model, when the number of samples of a specific class in the training dataset used to train the neural network model is small, the recognition accuracy of this specific class of samples is much lower than that of other classes. This phenomenon is called the data imbalance problem. To increase the recognition accuracy of long-tail class samples, the common approach is to increase the number of long-tail class data or reduce the number of non-long-tail class data. Specifically, for example: manually collecting more long-tail class samples as training data, etc. However, in the process of practice, it is found that this approach has very limited improvement in the recognition accuracy of long-tail class samples, and sometimes there is even no improvement. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a model distillation training method, apparatus, electronic device, and storage medium, which are used to improve the problem that the improvement of the recognition accuracy of long-tail class samples is very limited.

[0005] The embodiments of this application provide a model distillation training method, including: obtaining a training dataset including long-tail classes, and using a variety of data augmentation means to perform data augmentation on the training dataset to obtain multiple data sets; using the multiple data sets to perform different types of loss optimization training on multiple teacher models respectively to obtain multiple trained teacher models, where one teacher model is obtained by using one data set to perform one type of loss optimization training; selecting a preset number of teacher models from the multiple teacher models according to the accuracy; obtaining a preset number of student models, using the teacher models to perform distillation training on the student models to obtain a preset number of distilled student models; screening out the student model with the highest accuracy from the preset number of distilled student models. In the above implementation process, by organically combining data augmentation, loss optimization, and knowledge distillation means, multiple student models learn data representations from different perspectives from the teacher models, and then screening out the student model with the highest accuracy from the multiple student models effectively improves the recognition accuracy of long-tail class samples of the student model under the unbalanced data distribution.

[0006] Optionally, in the embodiments of the present application, after screening out the student model with the highest accuracy rate from the preset number of distilled student models, it further includes: obtaining the data to be processed; using the student model with the highest accuracy rate to perform classification prediction on the data to be processed, and obtaining a classification result. In the above implementation process, by using the student model with the highest accuracy rate that learns different perspectives from the teacher model to perform classification prediction on the data to be processed, the correct recognition rate of the student model for samples of long-tail categories under the unbalanced data distribution is effectively improved.

[0007] Optionally, in the embodiments of the present application, the training dataset is a text dataset; various data augmentation means are used to augment the training dataset, including: using data augmentation means such as synonym replacement, back translation, dynamic masking, random insertion, random swapping, and / or random deletion to augment the text dataset. In the above implementation process, by using data augmentation means such as synonym replacement, back translation, dynamic masking, random insertion, random swapping, and / or random deletion to augment the text dataset, the correct recognition rate of samples of long-tail categories in the text dataset is effectively improved.

[0008] Optionally, in the embodiments of the present application, the training dataset is an image dataset and / or a video dataset; various data augmentation means are used to augment the training dataset, including: using data augmentation means such as image scaling, image rotation, horizontal flipping, and vertical flipping to augment the image dataset and / or the video dataset. In the above implementation process, by using data augmentation means such as image scaling, image rotation, horizontal flipping, and vertical flipping to augment the image dataset and / or the video dataset, the correct recognition rate of samples of long-tail categories in the image dataset and / or the video dataset is effectively improved.

[0009] Optionally, in the embodiments of the present application, the teacher model includes: a first embedding layer, a first transformer layer, and a first prediction layer, and the student model includes: a second embedding layer, a second transformer layer, and a second prediction layer; using the teacher model to perform distillation training on the student model, including: using the Earth Mover's Distance (EMD) to calculate the data labels in the data set, the output of the first prediction layer, and the output of the second prediction layer to obtain a first distillation loss, and respectively calculating a second distillation loss between the output of the first transformer layer and the output of the second transformer layer, and a third distillation loss between the output of the first embedding layer and the output of the second embedding layer; training and optimizing the parameters in the student model according to the first distillation loss, the second distillation loss, and / or the third distillation loss to obtain a trained student model. In the above implementation process, by using the teacher model to perform distillation training on the student model, multiple student models can learn data representations from different perspectives from the teacher model, and the correct recognition rate of the student model for samples of long-tail categories under the unbalanced data distribution is effectively improved.

[0010] Optionally, in the embodiments of the present application, the Earth Mover's Distance (EMD) between the bulldozer and the data in the data set is used to calculate the data labels, the output of the first prediction layer, and the output of the second prediction layer, and a first distillation loss is obtained, including: calculating the first EMD distance between the output of the first prediction layer and the output of the second prediction layer, and calculating the second EMD distance between the data labels in the data set and the output of the second prediction layer; calculating the first distillation loss according to the first EMD distance and the second EMD distance. In the above implementation process, the first distillation loss is calculated based on the first EMD distance calculated from the soft target loss and the second EMD distance calculated from the hard target loss, thereby effectively improving the recognition accuracy of the student model for samples of long-tailed categories under unbalanced data distributions.

[0011] Optionally, in the embodiments of the present application, different types of loss optimizations include: CeLoss, FocalLoss, GHM, and / or DiceLoss.

[0012] The embodiments of the present application also provide a model distillation training device, including: a data set acquisition module, configured to obtain a training data set including long-tailed categories, and perform data augmentation on the training data set using a variety of data augmentation means to obtain multiple data sets; a teacher model acquisition module, configured to use the multiple data sets to perform different types of loss optimization training on multiple teacher models respectively to obtain multiple trained teacher models, where one teacher model is obtained by performing one type of loss optimization training using one data set; a teacher model selection module, configured to select a preset number of teacher models from the multiple teacher models according to the accuracy rate; a student model acquisition module, configured to obtain a preset number of student models, and perform distillation training on the student models using the teacher models to obtain a preset number of distilled student models; a student model selection module, configured to screen out the student model with the highest accuracy rate from the preset number of distilled student models.

[0013] Optionally, in the embodiments of the present application, the model distillation training device includes: a processed data acquisition module, configured to acquire data to be processed; a target data prediction module, configured to perform classification prediction on the data to be processed using the student model with the highest accuracy rate to obtain a classification result.

[0014] Optionally, in the embodiments of the present application, the training data set is a text data set; the data set acquisition module includes: a text data augmentation module, configured to perform data augmentation on the text data set using data augmentation means such as synonym replacement, back translation, dynamic masking, random insertion, random swapping, and / or random deletion.

[0015] Optionally, in the embodiments of the present application, the training data set is an image data set and / or a video data set; the data set obtaining module includes: an image / video enhancement module, which is used to perform data enhancement on the image data set and / or the video data set by means of data enhancement such as image scaling, image rotation, horizontal flipping, and vertical flipping.

[0016] Optionally, in the embodiments of the present application, the teacher model includes: a first embedding layer, a first transformer layer, and a first prediction layer, and the student model includes: a second embedding layer, a second transformer layer, and a second prediction layer; the student model obtaining module includes: a distillation loss obtaining module, which is used to calculate the first distillation loss by using the Earth Mover's Distance (EMD) for the data labels in the data set, the output of the first prediction layer, and the output of the second prediction layer, and calculate the second distillation loss between the output of the first transformer layer and the output of the second transformer layer, and the third distillation loss between the output of the first embedding layer and the output of the second embedding layer respectively; the student model training module is used to train and optimize the parameters in the student model according to the first distillation loss, the second distillation loss, and / or the third distillation loss to obtain a trained student model.

[0017] Optionally, in the embodiments of the present application, the distillation loss obtaining module includes: an EMD distance calculation module, which is used to calculate the first EMD distance between the output of the first prediction layer and the output of the second prediction layer, and calculate the second EMD distance between the data labels in the data set and the output of the second prediction layer; the distillation loss calculation module calculates the first distillation loss according to the first EMD distance and the second EMD distance.

[0018] Optionally, in the embodiments of the present application, different types of loss optimizations include: CeLoss, FocalLoss, GHM, and / or DiceLoss.

[0019] The embodiments of the present application also provide an electronic device, including: a processor and a memory, the memory stores machine-readable instructions executable by the processor, and when the machine-readable instructions are executed by the processor, the methods described above are executed.

[0020] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the methods described above are executed. Description of the Drawings

[0021] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0022] Figure 1 Flow schematic diagram of the model distillation training method provided by the embodiments of the present application shown;

[0023] Figure 2 Flow schematic diagram of using the teacher model to perform distillation training on the student model provided by the embodiments of the present application shown;

[0024] Figure 3 Flow schematic diagram of using the student model for classification prediction provided by the embodiments of the present application shown;

[0025] Figure 4 Structural schematic diagram of the model distillation training device provided by the embodiments of the present application shown;

[0026] Figure 5 Structural schematic diagram of the electronic device provided by the embodiments of the present application shown. Specific embodiments

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the embodiments of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the embodiments of the present application.

[0028] Before introducing the model distillation training method provided by the embodiments of the present application, some concepts involved in the embodiments of the present application will be introduced first:

[0029] Distillation training, also known as Knowledge Distillation, model distillation, dark knowledge extraction, or distillation learning, refers to migrating knowledge from an old machine learning model to a new machine learning model. The network structures of the old machine learning model and the new machine learning model can be the same or different.

[0030] Back Translation means translating from Language A to another language, obtaining the translation B of Language A, and then translating the translation B of Language A back to Language A.

[0031] The Easiest Data Augmentation (EDA) refers to performing data augmentations such as random synonym replacement, random insertion, random swapping, and / or random deletion on data to create new data and achieve the purpose of expanding the number of data entries.

[0032] The Earth Mover Distance (EMD), also known as the EMD distance or Wasserstein distance, is a measure of the distance between two probability distributions and can be used to describe the similarity between two multi-dimensional distributions.

[0033] It should be noted that the model distillation training method provided in the embodiments of this application can be executed by an electronic device. Here, the electronic device refers to a device terminal with the function of executing computer programs or the above-mentioned server. Examples of device terminals include: smart phones, personal computers (PCs), tablet computers, personal digital assistants (PDAs), or mobile Internet devices (MIDs), etc. Examples of servers include: x86 servers and non-x86 servers. Non-x86 servers include: mainframes, minicomputers, and UNIX servers.

[0034] The following introduces the application scenarios applicable to the model distillation training method. Here, the application scenarios include but are not limited to: using the model distillation training method to perform distillation training on a model, thereby improving the recognition accuracy of the long-tail class samples in a dataset with an unbalanced class distribution and effectively improving problems such as unbalanced distribution in the training dataset.

[0035] Please refer to Figure 1 the schematic flowchart of the model distillation training method provided in the embodiments of this application shown; the main idea of this model distillation training method is to organically combine data augmentation, loss optimization, and knowledge distillation means, enabling multiple student models to learn data representations from different perspectives from the teacher model, and then selecting the student model with the highest accuracy from multiple student models, effectively improving the recognition accuracy of the student model for long-tail class samples under an unbalanced data distribution. The above-mentioned model distillation training method may include:

[0036] Step S110: Obtain a training data set including long-tail categories, and use a variety of data augmentation methods to augment the training data set to obtain multiple data sets.

[0037] Among them, the above training data set can be a text data set, an image data set, and / or a video data set, etc. There are many ways to obtain the above training data set including long-tail categories, including: The first acquisition method is to receive the training data set sent by other terminal devices and store the training data set in a file system, a database, or a removable storage device; The second acquisition method is to obtain a pre-stored training data set. Specifically, for example: obtain the training data set from a file system, or obtain the training data set from a database, or obtain the training data set from a removable storage device; The third acquisition method is to use software such as a browser to obtain the training data set on the Internet, or use other applications to access the Internet to obtain the training data set. Specifically, for example: download the THUCNews text data set or the ImageNet image data set from the Internet, etc. For the convenience of understanding and explanation, the THUCNews text data set is a Chinese news data set, and the following will take the THUCNews text data set as an example for explanation.

[0038] There are many implementation methods for the above step S110, including but not limited to the following several:

[0039] The first implementation method, if the training data set is a text data set, then a variety of data augmentation methods such as synonym replacement, back translation augmentation, EDA, dynamic masking (DM), random insertion, random swapping, and / or random deletion can be used to augment the text data set to obtain multiple data sets. Specifically, for example, eleven categories can be selected from the THUCNews text data set, and a preset number (for example, 5000) of data can be selected from each category, and each category can be cut into a training set and a validation set according to a preset ratio (for example, 8:2). Then, most of the training sets of four of the eleven categories (for example, 3 / 4) are removed, so that these four categories become the data sets of long-tail categories; Finally, the above seven data augmentation methods of synonym replacement, back translation augmentation, EDA, DM, random insertion, random swapping, and / or random deletion are respectively used to augment the training sets and / or validation sets of the remaining seven categories, so that these seven categories become the data sets of non-long-tail categories. Therefore, the above multiple data sets can include: data sets of four categories of long-tail categories and data sets of seven categories of non-long-tail categories.

[0040] In the second implementation, if the training data set is an image data set and / or a video data set, data augmentation means such as changing the background color or brightness, translational cropping, contrast adjustment, noise addition, image scaling, image rotation, horizontal flipping, and vertical flipping can be used to perform data augmentation on the image data set and / or the video data set to obtain multiple data sets.

[0041] In the third implementation, if the training data set is an audio data set, multiple data augmentation means such as adding random noise after speech recognition, randomly swapping sentences in splitting, and / or randomly deleting sentences can be used to perform data augmentation on the audio data set to obtain multiple data sets.

[0042] In addition, in the specific implementation process, a combination and / or aggregation method can also be used to perform data augmentation on the above text data set, audio data set, image data set, and / or video data set. Here, the combination means using two or more data augmentation means to perform data augmentation on the same sample data, and the aggregation means dividing the data set into two or more partial sample data, and each partial sample data is processed using different data augmentation means. Taking the text data set as an example, assuming that there are a total of 7 data augmentation means, if any two means are combined, then there are types of data augmentation means. Similarly, if any three means are combined, then there are types of data augmentation means, and so on.

[0043] After step S110, step S120 is executed: using multiple data sets to perform different types of loss optimization training on multiple teacher models respectively to obtain multiple trained teacher models. A teacher model is obtained by using a data set to perform one type of loss optimization training.

[0044] In the specific implementation process, the network structures of the above teacher models and student models can be the same or different. The focus of the distillation training here is that as long as the teacher models can teach the student models more data representations from different perspectives, rather than making the volume or complexity of the student models smaller than that of the teacher models. Among them, the basic idea of the above loss optimization is to assign higher weights to the long-tail classes and relatively lower weights to the non-long-tail classes during the training process, so that although the number of samples of the long-tail classes is small, a relatively good recognition accuracy can still be achieved.

[0045] It can be understood that the above different types of loss optimization methods may include: CrossEntropy Loss (CeLoss), FocalLoss, Gradient Harmonizing Mechanism (GHM), and / or DiceLoss. Among them, CeLoss is the cross-entropy loss function commonly used in classification problems, which assigns corresponding different weights to different classes, but it does not assign greater weights to long-tail classes itself. FocalLoss is modified from CeLoss and is a loss function for solving class imbalance and classification difficulty differences in classification problems. This function itself assigns lower weights to easily learned samples and higher weights to difficult-to-learn samples. GHM, like FocalLoss, is also modified from CeLoss and is a loss optimization function for a long-tail problem in object detection. DiceLoss is a set similarity metric function. Its general idea is that the loss of the current pixel point is not only related to the predicted value of the current pixel, but also related to other points within the area of this pixel point. Therefore, if the intersection of the predicted pixel point area and the true pixel point area is larger, it is considered that the two are closer.

[0046] After step S120, step S130 is executed: Select a preset number of teacher models from multiple teacher models according to the accuracy rate.

[0047] There are many implementation manners of the above step S130, including but not limited to the following several types:

[0048] The first implementation manner is to select teacher models with three different data augmentation methods from multiple teacher models respectively. Specifically, for example: select teacher models trained with single data augmentation, single loss optimization, and combined data augmentation and loss optimization from multiple teacher models. The teacher models trained with these three different data augmentation methods are respectively denoted as T(dataAug), T(bestLoss), and T(group). Among them, the single data augmentation here means using only one of the data augmentation means to perform data augmentation on the same sample data. The single loss optimization means using only one of the loss optimizations to perform data augmentation on the same sample data. And the combined data augmentation and loss optimization means using different combination methods of data augmentation means combined with loss optimizations to perform data augmentation on the same sample data.

[0049] The second implementation manner is to directly select teacher models with an accuracy rate greater than a preset threshold from multiple teacher models trained in a single data augmentation, single loss optimization, and combined loss optimization of data augmentation manner. For example, in this implementation manner, assume that there are six teacher models, denoted as T(dataAug1), T(bestLoss1), T(group1), T(dataAug2), T(bestLoss2), and T(group3), and their accuracy rates are 91%, 89%, 88%, 84%, 83%, and 84% respectively. If the preset threshold is set to 85%, then the selected teacher models are T(dataAug1), T(bestLoss1), and T(group1). It is also possible to select a preset number of teacher models with the top-ranked accuracy rates. The above-mentioned preset number and preset threshold can be set according to specific circumstances. For example, the preset number can be set to 3, or the preset threshold can be set to 87% and so on.

[0050] After step S130, step S140 is executed: Obtain a preset number of student models, and use the teacher model to perform distillation training on the student models to obtain a preset number of distilled student models.

[0051] Please refer to Figure 2 The schematic flowchart of using the teacher model to perform distillation training on the student model provided by the embodiment of the present application is shown; wherein, the above-mentioned teacher model may include: a first embedding layer, a first prediction layer, and a plurality of first transformer layers, and the student model includes: a second embedding layer, a second prediction layer, and a plurality of second transformer layers. The connection relationship between each layer is shown by the solid arrows in the figure. During the distillation training process, the knowledge distillation relationship between the teacher model and the student model is shown by the dotted lines in the figure (the top horizontal dotted line is the prediction layer distillation, the bottom horizontal dotted line is the embedding layer distillation, and the middle dotted line is the transformer layer distillation). The specific distillation training process will be described in detail below; wherein, the number of the first transformer layers and the number of the second transformer layers can be set according to specific circumstances, for example, set to: 3, 10, or 50 layers, etc.

[0052] After step S140, step S150 is executed: Select the student model with the highest accuracy rate from the preset number of distilled student models.

[0053] In the above implementation process, first, after using a variety of data augmentation means to augment the training data set, multiple data sets are used to perform different types of loss optimization training on multiple teacher models respectively, and some teacher models with higher accuracy are selected from the trained multiple teacher models. Then, these teacher models with higher accuracy are used to teach a preset number of distilled student models through distillation. Finally, the student model with the highest accuracy is selected from the preset number of distilled student models. By organically combining data augmentation, loss optimization, and knowledge distillation means, multiple student models learn data representations from different perspectives from the teacher models, and then the student model with the highest accuracy is selected from multiple student models, effectively improving the correct recognition rate of the student model for long-tail category samples under the unbalanced data distribution.

[0054] The implementation manner of the above step S140 may include:

[0055] Step S141: Calculate the first distillation loss by using the Earth Mover's Distance (EMD) for the data labels, the output of the first prediction layer, and the output of the second prediction layer in the data set.

[0056] Step S142: Calculate the second distillation loss between the output of the first transformer layer and the output of the second transformer layer, and the third distillation loss between the output of the first embedding layer and the output of the second embedding layer, respectively.

[0057] There are many distillation training methods for the above steps S141 to S142, including but not limited to the following several:

[0058] The first distillation training method combines soft targets (i.e., calculating the loss between model outputs) and hard targets (i.e., calculating the loss between data labels and model outputs) to calculate the distillation loss. This implementation manner includes: calculating the first EMD distance between the output of the first prediction layer and the output of the second prediction layer, and calculating the second EMD distance between the data labels in the data set and the output of the second prediction layer, and calculating the first distillation loss based on the first EMD distance and the second EMD distance. Or, calculating the first EMD distance between the output of the first transformer layer and the output of the second transformer layer, and the second EMD distance between the data labels in the data set and the output of the second transformer layer, and calculating the first distillation loss based on the first EMD distance and the second EMD distance. Then, calculate the second distillation loss between the output of the first transformer layer and the output of the second transformer layer, and the third distillation loss between the output of the first embedding layer and the output of the second embedding layer, respectively.

[0059] The second distillation training method simply uses soft targets (i.e., calculates the loss between the model outputs) to calculate the distillation loss. This implementation includes: calculating the EMD distance between the output of the first prediction layer and the output of the second prediction layer, and using this EMD distance as the first distillation loss above. Or, calculating the EMD distance between the output of the first transformer layer and the output of the second transformer layer, and using this EMD distance as the first distillation loss above. Then, calculate the second distillation loss between the output of the first transformer layer and the output of the second transformer layer, and the third distillation loss between the output of the first embedding layer and the output of the second embedding layer, respectively.

[0060] In the specific practice process, if the THUCNews dataset is used for training and selection, then a neural network model with single back-translation data augmentation can be selected as the teacher model, and the dataset with back-translation augmentation plus EDA augmentation during distillation training can be used as the training dataset, and the student model obtained by the above second distillation training method can be selected as the final neural network model, which can obtain the student model with the highest accuracy.

[0061] Step S143: Train and optimize the parameters in the student model according to the first distillation loss, the second distillation loss, and / or the third distillation loss to obtain a trained student model.

[0062] The implementation of the above step S143 includes: updating the network weight parameters in the student model according to at least one of the first distillation loss, the second distillation loss, and the third distillation loss until the loss value is less than a preset ratio or the number of iterations (epochs) is greater than a preset threshold, then the trained student model can be obtained. Among them, the above preset ratio can be set according to specific circumstances, such as set to 5% or 10%, etc.; the above preset threshold can also be set according to specific circumstances, such as set to 100 or 1000, etc.

[0063] It should be noted that the data augmentation scheme using hierarchical models (i.e., first selecting teacher models that learn different perspectives, and then screening out the model with the highest accuracy from the student models that learn different perspectives from the teacher models) considers the problem of data imbalance from a more comprehensive perspective, and uses the means of distillation learning to shield the impact of data imbalance on the final model as much as possible. Compared with the data augmentation means or loss optimization means that only consider from one angle, the embodiments of the present application use two generations of models (i.e., including teacher models and student models) instead of a single generation of models, allowing multiple student models to learn data representations from different perspectives from the teacher models, and then screening out the student model with the highest accuracy from multiple student models, effectively improving the recognition accuracy of the student model for long-tail category samples under the unbalanced data distribution.

[0064] Please refer to Figure 3 the schematic flowchart of classification prediction using the student model provided by the embodiment of the present application shown; optionally, in the embodiment of the present application, after screening out the student model with the highest accuracy from the pre-set number of distilled student models, the student model can also be used for classification prediction, and this process may include:

[0065] Step S210: Obtain the data to be processed.

[0066] The implementation manners of the above step S210 include: the first obtaining manner, receiving the data to be processed sent by other terminal devices and storing the data to be processed in a file system, a database or a mobile storage device; the second obtaining manner, obtaining the pre-stored data to be processed, specifically, for example, obtaining the data to be processed from a file system, or obtaining the data to be processed from a database, or obtaining the data to be processed from a mobile storage device; the third obtaining manner, using software such as a browser to obtain the data to be processed on the Internet, or using other application programs to access the Internet to obtain the data to be processed.

[0067] Step S220: Use the student model with the highest accuracy to perform classification prediction on the data to be processed and obtain a classification result.

[0068] The implementation manner of the above step S220 is, for example: if the data to be processed is text data, then Figure 2 it can be known that the student model may include an embedding layer, a prediction layer and a plurality of transformer layers. At this time, the embedding layer in the student model can be used to perform embedding layer data representation technology on the data to be processed to obtain an input representation vector; the plurality of transformers in the student model are used to perform attention vector operations on the input representation vector to obtain an attention vector, and finally, the prediction layer in the student model is used to perform classification prediction on the attention vector to obtain the classification result of the data to be processed. The classification result here may be a long-tail category or may not be a long-tail category. Since the student model is the model that learns different perspectives from the teacher model and has the highest accuracy, the student model has a higher recognition accuracy for long-tail categories compared with ordinary classification models.

[0069] In the above implementation process, by using the student model that learns different perspectives from the teacher model and has the highest accuracy to perform classification prediction on the long-tail category samples in the data to be processed, the correct recognition rate of the student model for long-tail category samples under the unbalanced data distribution is effectively improved.

[0070] Please refer to Figure 4 the schematic structural diagram of the model distillation training device provided by the embodiment of the present application shown; the embodiment of the present application provides a model distillation training device 300, including:

[0071] A data set acquisition module 310, configured to acquire a training data set including long-tail categories, and perform data augmentation on the training data set by using a variety of data augmentation means to obtain a plurality of data sets.

[0072] A teacher model acquisition module 320, configured to perform different types of loss optimization training on a plurality of teacher models respectively by using a plurality of data sets, to obtain a plurality of trained teacher models, where one teacher model is obtained by performing one type of loss optimization training by using one data set.

[0073] A teacher model selection module 330, configured to select a preset number of teacher models from a plurality of teacher models according to the accuracy rate.

[0074] A student model acquisition module 340, configured to acquire a preset number of student models, and perform distillation training on the student models by using the teacher models to obtain a preset number of distilled student models.

[0075] A student model selection module 350, configured to screen out the student model with the highest accuracy rate from a preset number of distilled student models.

[0076] Optionally, in an embodiment of the present application, the model distillation training device includes:

[0077] A processed data acquisition module, configured to acquire data to be processed.

[0078] A target data prediction module, configured to perform classification prediction on the data to be processed by using the student model with the highest accuracy rate to obtain a classification result.

[0079] Optionally, in an embodiment of the present application, the training data set is a text data set; the data set acquisition module includes:

[0080] A text data augmentation module, configured to perform data augmentation on the text data set by using data augmentation means such as synonym replacement, back translation, dynamic masking, random insertion, random swapping, and / or random deletion.

[0081] Optionally, in an embodiment of the present application, the training data set is an image data set and / or a video data set; the data set acquisition module includes:

[0082] An image / video augmentation module, configured to perform data augmentation on the image data set and / or the video data set by using data augmentation means such as image scaling, image rotation, horizontal flipping, and vertical flipping.

[0083] Optionally, in an embodiment of the present application, the teacher model includes: a first embedding layer, a first transformer layer, and a first prediction layer, and the student model includes: a second embedding layer, a second transformer layer, and a second prediction layer; the student model acquisition module includes:

[0084] The distillation loss acquisition module is used to calculate the data labels in the data set, the output of the first prediction layer and the output of the second prediction layer using the bulldozer distance EMD to obtain the first distillation loss, and respectively calculate the second distillation loss between the output of the first converter layer and the output of the second converter layer, and the third distillation loss between the output of the first embedding layer and the output of the second embedding layer.

[0085] The student model training module is used to train and optimize the parameters in the student model according to the first distillation loss, the second distillation loss and / or the third distillation loss to obtain a trained student model.

[0086] Optionally, in an embodiment of the present application, a distillation loss obtaining module includes:

[0087] The EMD distance calculation module is used to calculate the first EMD distance between the output of the first prediction layer and the output of the second prediction layer, and calculate the second EMD distance between the data label in the data set and the output of the second prediction layer.

[0088] The distillation loss calculation module calculates the first distillation loss according to the first EMD distance and the second EMD distance.

[0089] Optionally, in an embodiment of the present application, different types of loss optimization methods may include: CeLoss, FocalLoss, GHM and / or DiceLoss.

[0090] It should be understood that the device corresponds to the above-mentioned model distillation training method embodiment and can perform the various steps involved in the above-mentioned method embodiment. The specific functions of the device can be found in the above description. To avoid repetition, the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the operating system (OS) of the device.

[0091] See also Figure 5 An electronic device 400 provided in an embodiment of the present application includes: a processor 410 and a memory 420, wherein the memory 420 stores machine-readable instructions executable by the processor 410, and when the machine-readable instructions are executed by the processor 410, the above method is executed.

[0092] The embodiment of the present application further provides a computer-readable storage medium 430 , on which a computer program is stored. When the computer program is executed by the processor 410 , the above method is executed.

[0093] Among them, the computer-readable storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0094] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may also occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, which mainly depends on the functions involved.

[0095] In addition, in each embodiment of the embodiments of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0096] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0097] The above description is only an optional implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the embodiments of the present application.

Claims

1. A method for distillation training of a classification model, characterized in that, Including: Obtain a training data set including long-tail categories, and perform data augmentation on the training data set using a variety of data augmentation means to obtain multiple data sets. The training data set includes: a text data set, an image data set, and / or a video data set; Use the multiple data sets to perform different types of loss optimization training on multiple teacher models respectively to obtain multiple trained teacher models. One teacher model is obtained by performing one type of loss optimization training using one data set; Select a preset number of teacher models from the multiple teacher models according to the accuracy rate; Obtain the preset number of student models, and use the teacher models to perform distillation training on the student models to obtain the preset number of distilled student models; Screen out the student model with the highest accuracy rate from the preset number of distilled student models; Use the student model with the highest accuracy rate to perform classification prediction on the data to be processed to obtain the classification result of the data to be processed. The data to be processed includes text data, image data, and / or video data; Wherein, both the teacher model and the student model are neural network models, and the different types of loss optimization training methods include: cross-entropy loss, FocalLoss, gradient coordination mechanism, and / or DiceLoss used in classification problems.

2. The method according to claim 1, characterized in that The training data set is a text data set; the performing data augmentation on the training data set using a variety of data augmentation means includes: Performing data augmentation on the text data set using data augmentation means such as synonym replacement, back translation, dynamic masking, random insertion, random swapping, and / or random deletion.

3. The method according to claim 1, wherein The training data set is an image data set and / or a video data set; The performing data augmentation on the training data set using a variety of data augmentation means includes: Performing data augmentation on the image data set and / or the video data set using data augmentation means such as image scaling, image rotation, horizontal flipping, and vertical flipping.

4. The method according to claim 1, wherein The teacher model includes: a first embedding layer, a first transformer layer, and a first prediction layer. The student model includes: a second embedding layer, a second transformer layer, and a second prediction layer; the performing distillation training on the student model using the teacher model includes: Calculating the earth mover's distance EMD for the data labels in the data set, the output of the first prediction layer, and the output of the second prediction layer to obtain a first distillation loss, and respectively calculating a second distillation loss between the output of the first transformer layer and the output of the second transformer layer, and a third distillation loss between the output of the first embedding layer and the output of the second embedding layer; Training and optimizing the parameters in the student model according to the first distillation loss, the second distillation loss, and / or the third distillation loss to obtain the trained student model.

5. The method according to claim 4, wherein The calculating the earth mover's distance EMD for the data labels in the data set, the output of the first prediction layer, and the output of the second prediction layer to obtain a first distillation loss includes: Calculate the first EMD distance between the output of the first prediction layer and the output of the second prediction layer, and calculate the second EMD distance between the data labels in the data set and the output of the second prediction layer; Calculate the first distillation loss according to the first EMD distance and the second EMD distance.

6. A classification model distillation training device, characterized in that, It includes: A data set acquisition module, configured to obtain a training data set including long-tail categories, and perform data augmentation on the training data set using various data augmentation means to obtain multiple data sets. The training data set includes: a text data set, an image data set, and / or a video data set; A teacher model acquisition module, configured to perform different types of loss optimization training on multiple teacher models respectively using the multiple data sets to obtain multiple trained teacher models. One teacher model is obtained by performing one type of loss optimization training using one data set; A teacher model selection module, configured to select a preset number of teacher models from the multiple teacher models according to the accuracy rate; A student model acquisition module, configured to obtain the preset number of student models, and perform distillation training on the student models using the teacher models to obtain the preset number of distilled student models; A student model selection module, configured to screen out the student model with the highest accuracy rate from the preset number of distilled student models; The model distillation training device is further configured to use the student model with the highest accuracy rate to perform classification prediction on the data to be processed, and obtain the classification result of the data to be processed. The data to be processed includes text data, image data, and / or video data; Wherein, both the teacher model and the student model are neural network models, and the different types of loss optimization training methods include: cross-entropy loss, FocalLoss, gradient coordination mechanism, and / or DiceLoss used in classification problems.

7. An electronic device, characterized in that, It includes: A processor and a memory. The memory stores machine-readable instructions executable by the processor. When the machine-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is executed.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the method according to any one of claims 1 to 5 is executed.

Citation Information

Patent Citations

  • Model distillation method and device, and text retrieval method and device

    CN111553479A

  • Model training method, text classification method, electronic equipment and storage medium

    CN111898707A