Output Regularization Method Based on the Weights of the Classification Layer of the Teacher Model

By converting the classification layer weights of the teacher model into soft labels and participating in student model training, the problem of resource occupation and training time is solved, and higher classification accuracy and broader model applicability is achieved.

CN114782742BActive Publication Date: 2025-07-22ZHEJIANG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210357826.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-07-22
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

In the prior art, the problem of excessive resource use and excessive training time when training student models using teacher models is used.

Method used

The classification layer weights of the teacher model are transformed into a correlation matrix between categories, and participate in the training of the student model as a soft label, and the student model is optimized through the cross entropy loss function and the restart of the random walk algorithm.

Benefits of technology

It reduces the training resource requirements, improves the training speed, and improves the classification accuracy and applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782742B_ABST
    Figure CN114782742B_ABST
Patent Text Reader

Abstract

The present invention relates to an output regularization method based on the weights of the classification layer of a teacher model. The weights of the classification layer of the teacher model that has completed supervised training are converted into a correlation matrix between categories, and each row in the matrix is used as the soft label for the corresponding category to provide additional information for the student model and participate in the training of the student model; the student model with the highest accuracy is selected as the final target model. The present invention makes full use of the information provided by the teacher model, reduces the problems of excessive training resources occupied by the teacher model and too long overall training time during the training process. Even if some neural network models can only provide the weights of the classifier layer of the teacher model, the student model can still be trained by this method, which has higher classification accuracy, wider model applicability, faster training speed, and only requires less training resources, and can further regularize the network model under the condition of less resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computing; inferring or counting, and particularly to an output regularization method based on the weights of the classification layer of a teacher model in the field of image classification in deep learning. Background Art

[0002] In deep neural networks, neural network models with a large number of parameters have achieved excellent performance in the supervised learning task of image classification. However, such neural network models usually overfit the labeled training samples, resulting in poor generalization ability. This overfitting phenomenon is the main and common problem existing after training modern deep neural network models with millions of parameters using a labeled data set. To solve this overfitting problem, researchers at home and abroad have proposed different solutions, including regularization methods.

[0003] Regularization methods include regularization processing at the input end and the output end; at the input end, the generalization ability of the network model to data is improved by increasing the diversity of training samples. However, due to the limited nature of the labeled data set, researchers have proposed methods such as data augmentation and label mixing to increase virtual samples in the limited data to improve the generalization ability of the model; the regularization methods at the output end include the maximum entropy model, label smoothing, and knowledge distillation, etc. Among them, the maximum entropy model and the label smoothing technology optimize the loss function of model training from the perspective of information entropy to achieve the purpose of regularization. Knowledge distillation, which was initially used in the field of model compression, usually has two neural network models in normal training. The model with a larger number of parameters is called the teacher model, and the model with a smaller number of parameters is called the student model. Since the teacher usually has higher generalization ability and has a better prediction output for the classification recognition of a single sample, it can be used as a kind of implicit knowledge to provide to the student model during training to improve the generalization ability of the student model.

[0004] However, for the output regularization method, knowledge distillation has wider applicability and stronger regularization ability compared with the maximum entropy model and the label smoothing technology. However, it requires the provision of the entire teacher model during training, and the training time is also much longer than the latter. Summary of the Invention

[0005] The present invention solves the problems existing in the prior art and provides an optimized output regularization method based on the weights of the classification layer of a teacher model, overcoming the problems of excessive resources required and long training time when using a teacher model to train a student model in the prior art.

[0006] The technical solution adopted by the present invention is an output regularization method based on the weights of the classification layer of a teacher model. The method converts the weights of the classification layer of the teacher model that has completed supervised training into a correlation matrix between categories, uses each row in the matrix as the soft label corresponding to the category, provides additional information for the student model and participates in the training of the student model; selects the student model with the highest accuracy as the final target model.

[0007] Preferably, the method includes the following steps:

[0008] Step 1: Obtain a labeled data set and divide it into a training set and a validation set;

[0009] Step 2: Preprocess all data in the data set;

[0010] Step 3: Obtain a preset teacher model or construct a teacher model;

[0011] Step 4: Based on the weights of the classification layer of the teacher model, obtain a soft label matrix;

[0012] Step 5: Based on the soft label matrix, train the target model;

[0013] Step 6: Adjust the parameters several times, and repeat Steps 4 and 5, select the student model with the highest accuracy as the final target model.

[0014] Preferably, in Step 1, the number of elements in the validation set is min(0.1×n, 5000), where n is the number of elements in the data set.

[0015] Preferably, in Step 2, the preprocessing includes normalization processing and enhancement operations on the data, and the sizes of the preprocessed data are the same.

[0016] Preferably, in Step 3, the constructed teacher model is trained on the training set prepared in Step 2 through a cross-entropy loss function, and the cross-entropy loss function is shown in Equation (1).

[0017]

[0018] Among them, y represents the data label, p represents the predicted distribution of the teacher model for the data, C is the size of the current data set category, c is a positive integer between 1 and C, y c is the data label of the current c-th category, p c is the predicted distribution of the teacher model for the data in the current c-th category.

[0019] Preferably, Step 4 includes the following steps:

[0020] Step 4.1: For the teacher model, retain the weight matrix W of its classification layer, whose dimension is k×C, where k is the feature dimension of the input data of the classification layer and C is the size of the current dataset categories;

[0021] Step 4.2: Transform the weight matrix W into a category-related soft label matrix Q through Equation (2),

[0022] Q = σ(W T W) (2)

[0023] where σ(.) is the softmax function and the dimension of Q is C×C;

[0024] Step 4.3: Use the restart random walk algorithm to iteratively calculate the current soft label matrix according to Equation (3) to obtain S,

[0025] s = (1 - u)Qs + uq (3)

[0026] where s represents the probability distribution of each row of S, q represents the probability distribution of each row of Q, the dimension size is 1×C, u is the probability of restoring to the original vector, u is adjustable, and u ∈ (0, 1);

[0027] Step 4.4: If the number of iterations reaches the preset maximum value or satisfies ∈ < 1e - 6, stop the iteration and output the stabilized matrix S; otherwise, increment the number of iterations by 1 and return to Step 4.3;

[0028] where ∈ = norm L1 (S i - S i-1 ), which is the L1 norm size of the difference between matrix S in each iteration and S in the previous iteration. i is used for counting and i is not greater than the preset maximum value of the number of iterations. S i is the S calculated in the current iteration, S i-1 is the S calculated in the previous iteration. In the first iteration, S i-1 is Q.

[0029] Preferably, the said Step 5 includes the following steps:

[0030] Step 5.1: Select a network model with fewer parameters than the teacher model and train it on the current dataset; set the loss function according to Equation (6),

[0031] L tsr = αL ce + βL3 (6)

[0032] where α and β are weight coefficients, α ∈ (0, 3], β ∈ (0, 3], L ce is the cross - entropy loss function, L tis the KL divergence function, as shown in Equation (5).

[0033]

[0034] where s l is the label vector related to the l-th row category in the soft label matrix S corresponding to the current training data, and s c,l is the label vector related to the c-th category in the l-th row of the soft label matrix S corresponding to the current training data;

[0035] Step 5.2: Select the random gradient descent algorithm optimizer, perform a preset number of iterations on the model, and store the model of the last iteration.

[0036] Preferably, the classification is image classification.

[0037] The present invention relates to an optimized output regularization method based on the weights of the classification layer of the teacher model, which converts the weights of the classification layer of the teacher model that has completed supervised training into a correlation matrix between categories, uses each row in the matrix as the soft label corresponding to the category, provides additional information for the student model and participates in the training of the student model; selects the student model with the highest accuracy as the final target model.

[0038] The beneficial effects of the present invention are as follows:

[0039] (1) Make full use of the information provided by the teacher model, and reduce the problems of excessive training resources occupied by the teacher model and too long overall training time during the training process;

[0040] (2) Even if some neural network models can only provide the weights of the classification layer of the teacher model, the student model can be trained by this method;

[0041] (3) Compared with the label smoothing technique, the present invention has higher classification accuracy and wider model applicability;

[0042] (4) Compared with the knowledge distillation technique, the present invention has a faster training speed and only requires less training resources, and can further regularize the network model under the condition of less resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the flowchart of the method of the present invention;

[0044] Figure 2 is the schematic diagram of the model in the present invention;

[0045] Figure 3 is the schematic diagram of taking the soft label corresponding to the training data when the soft label matrix in the present invention trains the student model. DETAILED DESCRIPTION OF THE INVENTION

[0046] The present invention will be further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.

[0047] The present invention relates to an output regularization method based on the weights of the classification layer of a teacher model. The method converts the weights of the classification layer of the teacher model that has completed supervised training into a correlation matrix between categories, uses each row in the matrix as the soft label corresponding to the category, provides additional information for the student model and participates in the training of the student model; selects the student model with the highest accuracy as the final target model.

[0048] The classification is image classification.

[0049] In the present invention, the output regularization method based on the weights of the classification layer in the teacher model consists of six parts: dataset preparation, data preprocessing, teacher model preparation, obtaining soft labels based on the weights of the classification layer of the teacher model, constructing a network training model, and parameter adjustment. Among them, after each parameter adjustment, the acquisition of soft labels and the construction of the network training model will be repeated.

[0050] In the present invention, dataset preparation is the primary step before the training of the target network. The dataset is usually a labeled dataset within the current task domain, and then the current dataset is divided into a training set and a validation set; data preprocessing is to crop, normalize the pictures and perform a certain degree of data augmentation.

[0051] In the present invention, in the teacher model preparation stage, the teacher model usually selects a model with a larger number of parameters than the target network model, that is, the number of parameters of the student model. If the teacher model can be directly obtained, that is, it already exists or can be directly obtained by downloading, then there is no need to train it. Otherwise, the teacher model is trained through the cross-entropy loss; in the stage of obtaining soft labels based on the weights of the classification layer of the teacher model, the weight matrix of the classification layer in the teacher model is retained, and the soft label matrix is obtained through the restart random walk algorithm; in the stage of constructing the network training model, the network model is trained by combining the soft label matrix with the proposed cross-entropy loss function; in the parameter adjustment stage, the model is adjusted through the validation set, the pictures are predicted and classified on the validation set, and its classification accuracy is calculated, and the model with the highest classification accuracy is retained.

[0052] The present invention is applicable to any dataset for image classification.

[0053] Taking the cifar100 dataset as an example, the method includes the following steps:

[0054] Step 1: Obtain a labeled dataset and divide it into a training set and a validation set;

[0055] In the above step 1, the number of elements in the validation set is min(0.1×n, 5000), where n is the number of elements in the dataset.

[0056] In the present invention, a labeled dataset for the current task domain is prepared. An image dataset on any task domain can be used. The currently adopted CIFAR-100 dataset is an open-source dataset, so there is no need for the manual labeling of samples. If the dataset used has no labeling information, the image samples in the current dataset need to be labeled, generally manually. Subsequently, the dataset is divided into a training set and a validation set. The size of the validation set is min(0.1×n, 5000), which means selecting the smaller value from 0.1×n and 5000 as the number of elements in the validation set. The remaining image data is used as the training set.

[0057] Step 2: Preprocess all the data in the dataset;

[0058] In the said Step 2, the preprocessing includes the normalization processing and augmentation operations of the data, and the size of the preprocessed data is consistent.

[0059] In the present invention, the normalization processing generally normalizes all the image sizes to the same size, such as a×b, where a and b can be specified by humans and are usually approximate to the image size to avoid overcropping and damaging the core content of the image. The data augmentation of the image data includes, but is not limited to, operations such as horizontal flipping and random cropping of the image. Since random cropping will cause the change of the image size, after data augmentation, the method of filling is required to supplement the image space, usually filling black into the image to keep the image size as a×b.

[0060] In the embodiments of the present invention, after image cropping, the images in the CIFAR-100 dataset are uniformly adjusted to color images of 32×32 pixels, or all the images are cropped to the same size according to the need for network training; the data augmentation is to perform random edge cropping on the images and then randomly fill them with black to keep the image size as 32×32 pixels, and at the same time, perform a certain amount of data augmentation using random horizontal flipping.

[0061] In the present invention, the above operations can be performed on the dataset by using the PyTorch deep learning framework. The torchvision package in PyTorch already provides operations such as image cropping and a certain amount of data augmentation, that is, image normalization, etc. After completion, through the provided framework interface, the images are input and loaded into the DataLoader for subsequent training use.

[0062] Step 3: Obtain a preset teacher model or construct a teacher model;

[0063] In the said Step 3, constructing the teacher model is to train the teacher model on the training set prepared in Step 2 through the cross-entropy loss function, and the cross-entropy loss function is shown in Equation (1).

[0064]

[0065] Among them, y represents the data label, p represents the predicted distribution of the teacher model for the data, C is the size of the current dataset category, c is a positive integer between 1 and C, and y c is the data label of the current c-th category, and p c is the predicted distribution of the teacher model for the data in the current c-th category.

[0066] In the present invention, there are two ways to obtain the teacher model:

[0067] If there is a corresponding teacher model in the current data domain, such as being prepared in advance or there is a corresponding teacher model on the network, it can be directly downloaded;

[0068] If not, select and construct the network structure of the teacher model, and train the teacher model through cross-entropy loss on the training set prepared in step 2; as shown in Equation (1), assuming the number of categories in the current image classification task is 3, then the values of y are {0, 1, 2}, where 0, 1, or 2 represents which category the current image belongs to, and p is a 3-dimensional vector representing the predicted output distribution of the current model for the image.

[0069] In the embodiment of the present invention, the resnet101 network model is selected as the teacher model. According to Equation (1), the stochastic gradient descent optimizer provided by pytorch is used to train the teacher model, and after 200 iterations, the model of the last iteration is used as the teacher model.

[0070] Step 4: Obtain the soft label matrix based on the classification layer weights of the teacher model;

[0071] The steps in step 4 include the following steps:

[0072] Step 4.1: For the teacher model, retain the weight matrix W of its classification layer, whose dimension is k×C, where k is the feature dimension of the input data of the classification layer, and C is the size of the current dataset category;

[0073] Step 4.2: Transform the weight matrix W into a category-related soft label matrix Q through Equation (2),

[0074] Q = σ(W T W) (2)

[0075] Among them, σ(.) is the softmax function, and the dimension of Q is C×C;

[0076] Step 4.3: Use the restart random walk algorithm to iteratively calculate the current soft label matrix according to Equation (3) to obtain S,

[0077] s = (1 - u)Qs + uq (3)

[0078] Among them, s represents the probability distribution of each row of S, q represents the probability distribution of each row of Q, the dimension size is 1×C, u is the probability of restoring to the original vector, u is adjustable, and u ∈ (0, 1);

[0079] Step 4.4: If the number of iterations reaches the preset maximum value or satisfies ∈ < 1e-6, stop the iteration and output the stabilized matrix S; otherwise, increment the number of iterations by 1 and return to Step 4.3;

[0080] Among them, ∈ = norm L1 (S i - S i-1 ), which is the L1 norm of the difference between matrix S in each iteration and S in the previous iteration. i is used for counting and i is not greater than the preset maximum value of the number of iterations. S i is the S calculated in the current iteration, and S i-1 is the S calculated in the previous iteration. In the first iteration, S i-1 is Q.

[0081] In the present invention, after obtaining the teacher model, the weight matrix W of the classification layer in the teacher model is retained, and its dimension is k×C. It is transformed into a class-related soft label matrix through Equation (2). Among them, W T The W matrix is obtained by transpose multiplication of the W matrix, and its dimension is C×C. The softmax function σ(·) is used for each row to map the weight correlation to the probability space of (0, 1), and a probability matrix Q with the sum of each row being 1 is obtained, which is the required soft label matrix, and its dimension is C×C.

[0082] In the present invention, as the soft label matrix, the probability of each row of Q is not stable. Therefore, through the restarted random walk algorithm, the current soft label matrix is calculated according to Equation (3), and finally the probability of each row is stabilized to obtain S, which is a probability matrix tending to be stable and is closer to the correlation between classes in the real environment. Generally, u = 0.2. The purpose of the restarted random walk algorithm is to obtain the proportion of the original vector after one iteration through iteration, and its iteration termination condition is that the number of iterations exceeds 40 or ∈ < 1e-6.

[0083] In an embodiment of the present invention, the weight matrix W of the feature mapping layer in the teacher model is selected for the next operation. In the resnet101 model, the size of the W matrix is 512×100 dimensions, where 512 is the feature dimension after the feature extractor, and 100 is the class dimension of the cifar100 dataset; the probability matrix Q can be obtained through Equation (2), and the restart random walk algorithm is executed through the recurrence formula (3) to obtain the probability matrix S, where the probability of each row tends to be stable; set u = 0.2, the number of iterations is 40 times, and ∈ is set to 1e-6.

[0084] Step 5: Train the target model based on the soft label matrix;

[0085] The said Step 5 includes the following steps:

[0086] Step 5.1: Select a network model with fewer parameters than the teacher model and train it on the current dataset; set the loss function according to Equation (6).

[0087] L 34r =αL ce +βL t (6)

[0088] Where, 6 and β are weight coefficients, α∈(0,3], β∈(0,3], L ce is the cross-entropy loss function, L3 is the KL divergence function, as shown in Equation (5),

[0089]

[0090] Where, s l is the label vector related to the category in the l-th row of the soft label matrix S corresponding to the current training data, and s c,l is the label vector related to the c-th category in the l-th row of the soft label matrix S corresponding to the current training data;

[0091] Step 5.2: Select the stochastic gradient descent algorithm optimizer, perform a preset number of iterations on the model, and store the model of the last iteration.

[0092] In the present invention, after obtaining the soft label matrix S, the target model, that is, the student model, is trained.

[0093] In the present invention, s l is the label vector related to the category in the l-th row of the soft label matrix S corresponding to the current training data, which is the l-th row in the matrix S and is the prediction probability vector of the model for the current data.

[0094] In the present invention, generally, α = 1 and β = 2 are set.

[0095] In the present invention, after setting the loss function, through the selected stochastic gradient descent algorithm optimizer, the model is iterated 200 times to obtain the model of the last iteration, which is stored in the hard disk.

[0096] In an embodiment of the present invention, a network training model is constructed: after obtaining the soft label matrix S, start training the target model, that is, the student model; first, select a network model resnet18 with a relatively small number of parameters and train it on the cifar100 dataset; the loss function required for training consists of two parts, including the cross-entropy loss function L ce , where y is the label in the cifar100 training set, p is the predicted probability vector of the model for the current data, and also includes the KL divergence in Equation (5), where s l is the label vector related to the category in the soft label matrix S corresponding to the current training data, and p is the predicted probability vector of the model for the current data; thus, the loss function in Equation (6) can be obtained as the loss function L for training the current student model tsr , where α and β are the weight coefficients before the two loss functions, usually set α = 1 and β = 2; after setting the loss function, through the stochastic gradient descent algorithm optimizer in pytorch, the model is iterated 200 times to obtain the model of the last iteration, and it is saved through persistence.

[0097] Step 6: Adjust the parameters several times, and repeat Steps 4 and 5, and select the student model with the highest accuracy as the final target model.

[0098] In the present invention, by adjusting the parameter u from 0.1 to 0.5 at an interval of 0.1, calculate and save the classification layer in the teacher model, then train multiple different student models through Steps 4 and 5, and finally use Dataset and DataLoader in pytorch to load the validation set, and use each target model to predict and classify the data in the validation set, compare the prediction and classification results with the labels of the validation set itself, finally obtain the accuracy of the classification prediction, and finally select the model with the highest accuracy as the final target model.

[0099] To achieve the above content, the present invention also relates to a computer-readable storage medium, on which an output regularization program based on the weight of the classification layer of the teacher model is stored. When the program is executed by a processor, the above output regularization method based on the weight of the classification layer of the teacher model is implemented, thereby solving the problems of excessive resources and long training time required for training the student model using the teacher model in the prior art.

[0100] To achieve the above, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method for output regularization based on the weights of the classification layer of the teacher model is implemented.

[0101] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0105] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0106] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An output regularization method based on the weights of the classification layer of a teacher model, characterized in that: The method includes the following steps: Step 1: Obtain a labeled dataset and divide it into a training set and a validation set; Step 2: Preprocess all data in the dataset; Step 3: Obtain a preset teacher model or construct a teacher model; Step 4: Based on the classification layer weights of the teacher model, obtain a soft label matrix, including the following steps: Step 4.1: For the teacher model, retain the weight matrix W of its classification layer, whose dimension is k×C, where k is the feature dimension of the input data of the classification layer and C is the size of the current dataset categories; Step 4.2: Transform the weight matrix W into a category-related soft label matrix Q through Equation (2), Q = σ(W T W) (2) where σ(.) is the softmax function and the dimension of Q is C×C; Step 4.3: Use the restart random walk algorithm to iteratively calculate the current soft label matrix according to Equation (3) to obtain S, s = uQs+(1 - u)s (3) where s represents the probability distribution of each row of S, with a dimension size of 1×C, u is the probability of restoring to the original vector, u is adjustable, and u∈(0,1); Step 4.4: If the number of iterations reaches the preset maximum value or satisfies ϵ<1e - 6, stop the iteration and output the stabilized matrix S; otherwise, increment the number of iterations by 1 and return to Step 4.3; where ϵ = norm L1 (S i - S i-1 ), which is the L1 norm of the difference between matrix S and u in the previous iteration in each iteration. i is used for counting and i is not greater than the preset maximum value of the number of iterations. S i is the S calculated in the current iteration, and S i-1 is the S calculated in the previous iteration. In the first iteration, S i-1 is Q; Step 5: Based on the soft label matrix, train the target model, including the following steps: Step 5.1: Select a network model with fewer parameters than the teacher model and train it on the current dataset; set the loss function according to Equation (6), L tsr = αL ce + βL t (6) where α and β are weight coefficients, α ∈ (0, 3], β ∈ (0, 3], and L ce is the cross-entropy loss function, and L t is the KL divergence function, as shown in Equation (5). where s l is the label vector related to the category in the l-th row of the soft label matrix S corresponding to the current training data, and s c,l is the label vector related to the c-th category in the l-th row of the soft label matrix S corresponding to the current training data; Step 5.2: Select a stochastic gradient descent algorithm optimizer to perform a preset number of iterations on the model and store the model of the last iteration; Step 6: Adjust the parameters several times and repeat Steps 4 and 5, and select the student model with the highest accuracy as the final target model.

2. The output regularization method based on the weights of the classification layer of the teacher model according to claim 1, characterized in that: In Step 1, the number of elements in the validation set is min(0.1×n, 5000), where n is the number of elements in the dataset.

3. An output regularization method based on the weights of the classification layer of a teacher model according to claim 1, characterized in that: In Step 2, the preprocessing includes normalization processing and augmentation operations on the data, and the sizes of the preprocessed data are the same.

4. The output regularization method based on the weights of the classification layer of the teacher model according to claim 1, characterized in that: In Step 3, the constructed teacher model is trained on the training set prepared in Step 2 through the cross-entropy loss function, and the cross-entropy loss function is as shown in Equation (1), Among them, y represents the data label, p represents the predicted distribution of the teacher model for the data, C is the size of the current dataset category, c is a positive integer between 1 and C, and y c is the data label of the current c-th category, and p c is the predicted distribution of the teacher model for the data in the current c-th category.

5. A method for output regularization based on the weights of the classification layer of a teacher model according to claim 1, characterized in that: The classification is image classification.

Citation Information

Patent Citations

  • Image classification method and system based on isomorphic multi-teacher guidance knowledge distillation

    CN113627545A