Multi-Self-Supervised Task Fusion Method, Device and Storage Medium Based on Knowledge Distillation
By introducing knowledge distillation technology into supervised learning and integrating multiple self-supervised tasks to train neural network models, the problems of difficulty in obtaining training data and overfitting in the existing technology are solved, and the improvement of model performance and complementary advantages are achieved.
Patent Information
- Application Number
- CN202210737555.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The existing supervised learning and self-supervised learning each have problems such as high cost of obtaining training data, difficulty in obtaining large-scale data, and overfitting, and have failed to effectively integrate their advantages.
Using a multi-self-supervised task fusion method based on knowledge distillation, self-supervised training is carried out by establishing multiple first neural network models and self-supervised tasks, and these models are fused through knowledge distillation, and the second neural network model is trained using classification tasks to achieve the improvement of model performance.
Through self-supervised learning, the model is pre-trained using labelless data, which improves the generalization performance of the model, avoids the overfitting problem of supervised learning, and realizes the complementary advantages of supervised learning and self-supervised learning.
Smart Images

Figure CN115205586B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device and storage medium for fusing multiple self-supervised tasks based on knowledge distillation. Background Art
[0002] Deep learning technology can be divided into three major branches: supervised learning, unsupervised learning, and self-supervised learning. Among them, supervised learning refers to training a neural network using a large amount of labeled data, and optimizing the network by minimizing the loss between the network output and the label, so that the model has intelligence. Unsupervised learning algorithms have no labels, so training models often have no clear goals, and the training results may not be certain. In essence, unsupervised learning algorithms are a method of probability statistics to discover some potential structures in data. In the semi-supervised task scenario with a large amount of unlabeled data and a small amount of labeled data, due to the lack of labels, pure supervised learning cannot achieve good results. At this time, self-supervised technology can be applied to enhance the model performance using unlabeled data. Self-supervised learning means that the model optimizes the network by solving proxy tasks and minimizing the loss between the model output and the pseudo-labels automatically generated by the proxy tasks, so that the model can have better performance in the fine-tuning of downstream tasks of self-supervised proxy tasks.
[0003] To sum up, supervised learning, unsupervised learning, and self-supervised learning all have their respective advantages, but they also have their respective disadvantages when viewed separately. For example, the acquisition cost of training data for supervised learning is relatively high, it is difficult to obtain a large amount of data to train a neural network, and the advantages of various learning methods are not integrated. Summary of the Invention
[0004] Aiming at the technical disadvantages of current supervised learning and self-supervised learning respectively, the purpose of the present invention is to provide a method, device and storage medium for fusing multiple self-supervised tasks based on knowledge distillation.
[0005] On the one hand, an embodiment of the present invention includes a method for fusing multiple self-supervised tasks based on knowledge distillation, including:
[0006] Establishing a plurality of first neural network models and a plurality of self-supervised tasks; each of the first neural network models corresponds to one of the self-supervised tasks;
[0007] Respectively using each of the self-supervised tasks to perform self-supervised training on the corresponding first neural network model;
[0008] Establishing a classification task and a second neural network model;
[0009] Fusing each of the first neural network models and the second neural network model through knowledge distillation, and using the classification task to train the second neural network model.
[0010] Further, fusing each of the first neural network models with the second neural network model and training the second neural network model using the classification task includes:
[0011] Obtaining sample data and true labels corresponding to the classification task;
[0012] Inputting the sample data into each of the first neural network models respectively;
[0013] Obtaining output results respectively generated by each of the first neural network models for processing the sample data;
[0014] Fusing the output results of each of the first neural network models to obtain soft labels;
[0015] Inputting the sample data into the second neural network model;
[0016] Obtaining prediction results generated by the second neural network model for processing the sample data;
[0017] Obtaining soft prediction results generated by the second neural network model for processing the sample data;
[0018] Determining a first loss function value according to the prediction results and the true labels;
[0019] Determining a second loss function value according to the soft prediction results and the soft labels;
[0020] Determining a third loss function value according to the first loss function value and the second loss function value;
[0021] Optimizing and updating network parameters of the second neural network model according to the third loss function value.
[0022] Further, fusing the output results of each of the first neural network models to obtain soft labels includes:
[0023] Weighted summing the output results of each of the first neural network models to obtain the soft labels.
[0024] Further, obtaining the soft prediction results generated by the second neural network model for processing the sample data includes:
[0025] Dividing the prediction results generated by the second neural network model for processing the sample data by a distillation temperature, then inputting the result into an activation function, and obtaining the output result of the activation function as the soft prediction results.
[0026] Further, determining a first loss function value according to the prediction result and the true label includes:
[0027] Calculating the cross-entropy loss function of the prediction result and the true label, and using the obtained result as the first loss function value.
[0028] Further, determining a second loss function value according to the soft prediction result and the soft label includes:
[0029] Calculating the K1 divergence of the soft prediction result and the soft label, and using the obtained result as the second loss function value.
[0030] Further, the multi-self-supervised task fusion method based on knowledge distillation further includes:
[0031] After completing the self-supervised training of each of the first neural network models, before training the second neural network model using the classification task, use the classification task to train each of the first neural network models that have undergone self-supervised training respectively.
[0032] Further, the multiple self-supervised tasks include an image rotation direction recognition task, an image relative direction recognition task, a jigsaw task, and a rotated jigsaw task.
[0033] On the other hand, an embodiment of the present invention further includes a computer device, including a memory and a processor, where the memory is used to store at least one program, and the processor is used to load the at least one program to execute the multi-self-supervised task fusion method based on knowledge distillation in the embodiment.
[0034] On the other hand, an embodiment of the present invention further includes a storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the multi-self-supervised task fusion method based on knowledge distillation in the embodiment when executed by the processor.
[0035] The beneficial effects of the present invention are as follows: The multi-self-supervised task fusion method based on knowledge distillation in the embodiment applies the technology of knowledge distillation for model compression and acceleration, and can utilize the first neural network model trained by self-supervised tasks as the teacher model to improve the performance of the second neural network model as the student model, enabling the training of the second neural network model to integrate the advantages of supervised learning, which is easy to train to obtain a network model with high accuracy, and self-supervised learning, which is easy to perform large-scale training. Specifically, through self-supervised learning, a large amount of unlabeled data is used for model pre-training to make the model reach a good initial parameter, so as to improve the generalization performance of the model's target task and avoid the shortcoming of overfitting easily occurring in pure supervised learning, realizing the effect of using multiple self-supervised learning technologies to improve supervised learning technology. Description of the Drawings
[0036] Figure 1 It is a flowchart of the multi-self-supervised task fusion method based on knowledge distillation in the embodiment;
[0037] Figure 2 It is a schematic diagram of the multi-self-supervised task fusion method based on knowledge distillation in the embodiment;
[0038] Figure 3 It is a schematic diagram of the image relative direction recognition task in the embodiment;
[0039] Figure 4 It is a schematic diagram of using the classification task to train each first neural network model after self-supervised training in the embodiment. Detailed Embodiment
[0040] In this embodiment, referring to Figure 1 , the multi-self-supervised task fusion method based on knowledge distillation includes the following steps:
[0041] S1. Establish multiple first neural network models and multiple self-supervised tasks;
[0042] S2. Respectively use each self-supervised task to perform self-supervised training on the corresponding first neural network model;
[0043] S3. Establish a classification task and a second neural network model;
[0044] S4. Through knowledge distillation, fuse each first neural network model with the second neural network model, and use the classification task to train the second neural network model.
[0045] In steps S1 - S4, each of the established first neural network models and the second neural network model can have basically the same network structure. Specifically, since the training tasks to be performed by each first neural network model and the second neural network model during the training process are different, the dimensions of the fully connected layers in different neural network models are different. Except for the fully connected layers, the structures of the other parts of each first neural network model and the second neural network model are the same.
[0046] The principles of steps S1 - S4 are as Figure 2 shown. In step S1, it is possible to establish as Figure 2Four first neural network models such as the first neural network model 1, the first neural network model 2, the first neural network model 3, and the first neural network model 4 as shown are provided, and four self-supervised tasks including an image rotation direction recognition task, an image relative direction recognition task, a jigsaw task, and a rotated jigsaw task are established. Among them, the image rotation direction recognition task corresponds to the first neural network model 1, that is, in step S2, the first neural network model 1 is self-supervised trained using the image rotation direction recognition task; the image relative direction recognition task corresponds to the first neural network model 2, that is, in step S2, the first neural network model 2 is self-supervised trained using the image relative direction recognition task; the jigsaw task corresponds to the first neural network model 3, that is, in step S2, the first neural network model 3 is self-supervised trained using the jigsaw task; the rotated jigsaw task corresponds to the first neural network model 4, that is, in step S2, the first neural network model 4 is self-supervised trained using the rotated jigsaw task.
[0047] In this embodiment, the self-supervised task model of the image rotation direction recognition task is specifically as follows: the input image is randomly rotated by 90°, 180°, 270°, and 360°, corresponding labels 0, 1, 2, and 3 are generated according to the rotation direction, the rotated image is used as the input of the first neural network model 1, the first neural network model 1 outputs the predicted rotation angle of the image, and the output dimension of the last fully connected layer of the first neural network model 1 is 4.
[0048] In this embodiment, the self-supervised task model of the image relative direction recognition task is specifically as follows: as Figure 3 shown, a small area (the solid box in the image) of the randomly selected Figure 3 image is taken, and 8 patches of the same scale (the dashed boxes in the image) are selected in the eight directions of up, down, left, right, upper left, lower left, upper right, and lower right around it. The direction of each patch relative to the solid box corresponds to 8 labels in sequence. One of the 8 patches and the blue patch are input into the first neural network model 2, and the first neural network model 2 outputs the predicted relative direction of the patch relative to the blue patch. The output dimension of the last fully connected layer of the first neural network model 2 is 8.
[0049] In this embodiment, the self-supervised task model of the jigsaw task is specifically as follows: the picture is evenly divided into four pieces, randomly shuffled and then put together. In this way, there are 4! = 24 permutation ways for the 4 pieces. The 4-piece jigsaw is input into the first neural network model 3, and the first neural network model 3 outputs the permutation way of these four pieces of the jigsaw, that is, the dimension of the last fully connected layer of the first neural network model 3 is 24.
[0050] In this embodiment, the self-supervised task model for the rotating puzzle task is specifically as follows: Based on the puzzle task, each piece of the puzzle is also rotated randomly in 4 directions, the same as the image rotation direction recognition task. The rotation directions of the four pieces of the puzzle are the same. In addition to predicting the arrangement, the first neural network model 4 also predicts the rotation direction of the puzzle, that is, the output dimension of the first neural network model 4 is 24 * 4 = 96.
[0051] In step S3, the established classification task can be an image content recognition task. In this embodiment, after step S2 is executed, before step S4, that is, before using the classification task to train the second neural network model, the classification task established in step S3 can also be used to train each first neural network model that has undergone self-supervised training. Specifically, when using the classification task to train each first neural network model that has undergone self-supervised training, the structure of each first neural network model can be adjusted. For example, referring to Figure 4 , for the first neural network model 1, its structure includes a feature extraction module and a fully connected layer matching the image rotation direction recognition task. The fully connected layer can be replaced with a fully connected layer matching the classification task established in step S3. Similarly, the fully connected layers of the first neural network model 2, the first neural network model 3, and the first neural network model 4 are respectively replaced with fully connected layers matching the classification task established in step S3.
[0052] In this embodiment, when step S4 is executed, that is, when fusing each first neural network model and the second neural network model through knowledge distillation and using the classification task to train the second neural network model, the following steps can be specifically executed:
[0053] S401. Obtain the sample data and the true labels corresponding to the classification task;
[0054] S402. Input the sample data into each first neural network model respectively;
[0055] S403. Obtain the output results respectively generated by each first neural network model for processing the sample data;
[0056] S404. Fuse the output results of each first neural network model to obtain soft labels;
[0057] S405. Input the sample data into the second neural network model;
[0058] S406. Obtain the prediction results generated by the second neural network model for processing the sample data;
[0059] S407. Obtain the soft prediction results generated by the second neural network model for processing the sample data;
[0060] S408. Determine the first loss function value based on the prediction result and the true label;
[0061] S409. Determine the second loss function value based on the soft prediction result and the soft label;
[0062] S410. Determine the third loss function value based on the first loss function value and the second loss function value;
[0063] S411. Optimize and update the network parameters of the second neural network model according to the third loss function value.
[0064] The principles of steps S401 - S411 are as Figure 2 shown. In steps S401 - S411, each of the first neural network models such as the first neural network model 1, the first neural network model 2, the first neural network model 3, and the first neural network model 4 is used as the teacher model, and the second neural network model is used as the student model for knowledge distillation.
[0065] In step S401, the classification task is a supervised learning task. Obtain the sample data corresponding to the classification task and the true label, where the true label is used to annotate the content of the sample data. The true label in step S401 can be called the hard label relative to the soft label in step S404.
[0066] In step S402, referring to Figure 2 , input the sample data into each of the first neural network models such as the first neural network model 1, the first neural network model 2, the first neural network model 3, and the first neural network model 4 respectively. In step S403, each first neural network model processes the sample data respectively and generates its own output result.
[0067] In step S404, referring to Figure 2 , the output result of each first neural network model is used as the soft label respectively, and the Kullback - Leibler divergence can be calculated between each soft label and the soft prediction result output by the second neural network model. The various Kullback - Leibler divergences can be fused to obtain the second loss function value loss2.
[0068] In step S405, referring to Figure 2 , input the sample data into the second neural network model. In step S406, the second neural network model processes the sample data, and the direct output result of the second neural network model (the logits output by the last fully - connected layer of the second neural network model) is the Figure 2 prediction result in
[0069] In step S407, the prediction result generated by the second neural network model for the sample data, such as the prediction result in step S406, is divided by the distillation temperature T, and then the obtained result is input into the activation function to obtain the output result of the activation function as the soft prediction result. Among them, softmax can be used as the activation function.
[0070] In step S408, referring to Figure 2 , calculate the cross-entropy loss function of the prediction result obtained in step S406 and the true label corresponding to the classification task, and the obtained result is used as the first loss function value loss1.
[0071] In step S409, referring to Figure 2 , calculate the K1 divergence between the soft prediction result obtained in step S407 and the soft label obtained in step S404, and the obtained result is used as the second loss function value loss2.
[0072] In step S410, according to the first loss function value loss1 and the second loss function value loss2, determine the third loss function value loss3. Specifically, coefficients α and β can be set, and through the formula loss3 = αloss1 + βloss2, the third loss function value loss3 can be calculated.
[0073] In step S411, according to the third loss function value loss3, use the optimizer to perform backward update on the network parameters of the second neural network model, so as to realize the training of the second neural network model.
[0074] In step S411, the network parameters of each first neural network model can be maintained unchanged, that is, only the network parameters of the second neural network model are updated.
[0075] The multi-self-supervised task fusion method based on knowledge distillation in this embodiment applies the technology of knowledge distillation for model compression and acceleration. Specifically, multiple first neural network models are used as teacher models, and a second neural network model is used as a student model, thus forming a teacher-student structure. First, a teacher model is trained using self-supervised tasks, and the teacher model is fine-tuned using sample data and true labels (hard labels) in the classification task. Then, the output of the teacher model is used as soft labels. The loss during the training of the student model consists of two parts: the loss between the soft prediction result output by the student model and the soft labels, and the loss between the prediction result output by the student and the hard labels. Thus, the output distribution of the student model fits the output distribution of the teacher model. It is possible to leverage the first neural network model trained with self-supervised tasks as the teacher model to improve the performance of the second neural network model as the student model, enabling the training of the second neural network model to integrate the advantages of supervised learning, which is easy to train to obtain a highly accurate network model, and self-supervised learning, which is easy to perform large-scale training, making the advantages and disadvantages of supervised learning and self-supervised learning complementary.
[0076] It is possible to write a computer program that executes the multi-self-supervised task fusion method based on knowledge distillation in this embodiment, write this computer program into a computer device or a storage medium. When the computer program is read and run, it executes the multi-self-supervised task fusion method in this embodiment, thereby achieving the same technical effects as the multi-self-supervised task fusion method in the embodiment.
[0077] It should be noted that, unless otherwise specified, when a certain feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. In addition, the up, down, left, right, etc. descriptions used in this disclosure are only relative to the mutual positional relationship of the various components of this disclosure in the drawings. The singular forms "a", "the", and "said" used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all the technical and scientific terms used in this embodiment have the same meaning as commonly understood by those skilled in the technical field of this application. The terms used in the description of this embodiment are only for describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any combination of one or more of the related listed items.
[0078] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided in this embodiment is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.
[0079] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose the program is capable of running on a programmed application-specific integrated circuit.
[0080] Furthermore, the operations of the processes described in this embodiment can be performed in any suitable order, unless this embodiment otherwise indicates or is otherwise clearly inconsistent with the context. The processes described in this embodiment (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.
[0081] Further, the method can be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and, when read by the storage medium or device, can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention as described in this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0082] A computer program can be applied to input data to perform the functions described in this embodiment, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0083] As described above, only the preferred embodiments of the present invention are given, and the present invention is not limited to the above-described embodiments. As long as it achieves the technical effects of the present invention by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.
Claims
1. A multi-self-supervised task fusion method based on knowledge distillation, characterized in that, the multi-self-supervised task fusion method based on knowledge distillation includes: establishing a plurality of first neural network models and a plurality of self-supervised tasks; each of the first neural network models corresponds to each of the self-supervised tasks; the plurality of self-supervised tasks include an image rotation direction recognition task, an image relative direction recognition task, a jigsaw task, and a rotated jigsaw task; respectively using each of the self-supervised tasks to perform self-supervised training on the corresponding first neural network model; establishing a classification task and a second neural network model; fusing each of the first neural network models and the second neural network model through knowledge distillation, and using the classification task to train the second neural network model; the fusing each of the first neural network models and the second neural network model through knowledge distillation, and using the classification task to train the second neural network model includes: obtaining sample data and true labels corresponding to the classification task; inputting the sample data into each of the first neural network models respectively; obtaining output results respectively generated by each of the first neural network models for processing the sample data; fusing the output results of each of the first neural network models to obtain a soft label; inputting the sample data into the second neural network model; obtaining a prediction result generated by the second neural network model for processing the sample data; obtaining a soft prediction result generated by the second neural network model for processing the sample data; determining a first loss function value according to the prediction result and the true label; determining a second loss function value according to the soft prediction result and the soft label; determining a third loss function value according to the first loss function value and the second loss function value; optimizing and updating the network parameters of the second neural network model according to the third loss function value; the fusing the output results of each of the first neural network models to obtain a soft label includes: weighted summing the output results of each of the first neural network models to obtain the soft label; the obtaining a soft prediction result generated by the second neural network model for processing the sample data includes: after dividing the prediction result generated by the second neural network model for processing the sample data by the distillation temperature, inputting it into an activation function, and obtaining the output result of the activation function as the soft prediction result.
2. The multi-self-supervised task fusion method based on knowledge distillation according to claim 1, characterized in that, the determining a first loss function value according to the prediction result and the true label includes: calculating the cross-entropy loss function of the prediction result and the true label, and using the obtained result as the first loss function value.
3. The multi-self-supervised task fusion method based on knowledge distillation according to claim 1, characterized in that, the determining a second loss function value according to the soft prediction result and the soft label includes: calculating the K1 divergence of the soft prediction result and the soft label, and using the obtained result as the second loss function value.
4. The multi-self-supervised task fusion method based on knowledge distillation according to claim 1, wherein, the multi-self-supervised task fusion method based on knowledge distillation further includes: after completing the self-supervised training of each of the first neural network models, before training the second neural network model using the classification task, using the classification task to train each of the first neural network models that have undergone self-supervised training respectively.
5. A computer device, wherein, it includes a memory and a processor, the memory is used to store at least one program, and the processor is used to load the at least one program to execute the multi-self-supervised task fusion method based on knowledge distillation according to any one of claims 1-4.
6. A storage medium storing a program executable by a processor, wherein, the program executable by the processor is used to execute the multi-self-supervised task fusion method based on knowledge distillation according to any one of claims 1-4 when executed by the processor.
Citation Information
Patent Citations
Cross-sample federal classification modeling method and device, storage medium and electronic equipment
CN113408209A
Network model training method and device and computer readable storage medium
CN113947196A