Model training method and device based on decoupling knowledge distillation, equipment and medium
By using the diffusion model to denoising the features output by the student model in decoupled knowledge distillation, the problem of difficulty in aligning the features of the teacher model and the student model is solved, and the training effect of the student model is significantly improved.
Patent Information
- Application Number
- CN202411922274.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-13
AI Technical Summary
When there are architectural differences between the teacher model and the student model, decoupled knowledge distillation cannot effectively align the features output from the teacher model and the features output from the student model, resulting in the student model being unable to fully learn the deep-level features of the teacher model.
The noise data in the features output by the student model is denoised by the preset diffusion model, so as to achieve accurate alignment between the features output by the teacher model and the features output by the student model.
It significantly improves the training effect of the student model, allowing the student model to fully learn the deep-seated features in the teacher model.
Smart Images

Figure CN119990254A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing technology, and in particular to a model training method, device, equipment and medium based on decoupled knowledge distillation. Background Art
[0002] Knowledge Distillation (KD) is a commonly used model compression technology that transfers the knowledge trained by a complex teacher model to a simple student model, so that the student model can achieve performance close to that of the teacher model while maintaining a small size. Traditional knowledge distillation methods usually rely on directly imitating the prediction output of the teacher model. Common methods include response-based distillation and distillation based on intermediate feature layers.
[0003] Decoupled Knowledge Distillation is a response-based distillation method. Compared with traditional knowledge distillation, the biggest difference between decoupled knowledge distillation is that the KD loss is reformulated into two parts: target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD), and hyperparameters are introduced to balance the two parts of the loss. Therefore, the knowledge in the teacher model can be captured more accurately while maintaining the simplicity of the student model. However, when there are obvious architectural differences between the teacher model and the student model, decoupled knowledge distillation cannot effectively align the features output by the teacher model and the features output by the student model, and cannot effectively narrow the representation gap between the two, resulting in the student model being unable to fully learn the deep features of the teacher model. Summary of the invention
[0004] The present application provides a model training method, apparatus, equipment and medium based on decoupled knowledge distillation, which can achieve precise alignment between the features output by the teacher model and the features output by the student model, thereby narrowing the feature representation gap between the teacher model and the student model, so that the student model can fully learn the deep-level features in the teacher model, which can significantly improve the training effect of the student model.
[0005] This application provides a model training method based on decoupled knowledge distillation, comprising the following steps: Acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; Inputting a first sample image into the teacher model and the student model respectively, obtaining a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object; De-noising the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; Obtaining a KL divergence loss between the first feature and the third feature; According to the KL divergence loss, with the goal of improving the accuracy of the output result of the student model, the student model is trained through a back propagation algorithm until a preset stop condition is met.
[0006] According to the model training method based on decoupled knowledge distillation of the present application, the obtaining of the KL divergence loss between the first feature and the third feature includes: Obtaining a KL divergence loss based on the target category according to a difference between a feature representing a target category in the first feature and a feature representing the target category in the third feature, where the target category is a correct category of the target object in the first sample image; Obtaining a KL divergence loss based on the non-target category according to a difference between a feature representing a non-target category in the first feature and a feature representing the non-target category in the third feature, wherein the non-target category is an erroneous category of the target object in the first sample image; A KL divergence loss between the first feature and the third feature is determined according to the KL divergence loss based on the target category and the KL divergence loss based on the non-target category.
[0007] According to the model training method based on decoupled knowledge distillation of the present application, obtaining the KL divergence loss based on the non-target category according to the difference between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature includes: Determine a Pearson correlation coefficient between a feature in the first feature representing the non-target category and a feature in the third feature representing the non-target category; The KL divergence loss based on the non-target category is determined according to the Pearson correlation coefficient.
[0008] According to the model training method based on decoupled knowledge distillation of the present application, the determining of the Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature includes: The Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature is determined by the following formula: in, represents the Pearson correlation coefficient, represents the number of non-target categories, Represents the output of the teacher model The prediction scores of non-target categories, represents the average of the prediction scores of each non-target category output by the teacher model, The student model determines the The prediction scores of non-target categories, Represents the average of the prediction scores for each non-target class determined by the student model.
[0009] According to the model training method based on decoupled knowledge distillation of the present application, the obtaining of the KL divergence loss between the first feature and the third feature includes: Compressing the number of channels of the first feature to obtain a fourth feature; A KL divergence loss between the fourth feature and the third feature is obtained.
[0010] According to the model training method based on decoupled knowledge distillation of the present application, the obtaining of the KL divergence loss between the fourth feature and the third feature includes: Performing feature reconstruction on the fourth feature to obtain a feature-reconstructed fourth feature, wherein the feature reconstruction is used to reduce redundant information in the fourth feature; Obtain a KL divergence loss between a fourth feature after the feature reconstruction and the third feature.
[0011] According to the model training method based on decoupled knowledge distillation of the present application, the obtaining of the KL divergence loss between the first feature and the third feature includes: aligning the dimensions of the third feature with the dimensions of the first feature; Obtain the KL divergence loss between the aligned first feature and the third feature.
[0012] The present application also provides a model training device based on decoupled knowledge distillation, comprising the following modules: A first acquisition module is used to acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; An input module, used to input a first sample image into the teacher model and the student model respectively, to obtain a first feature output by the teacher model and a second feature output by the student model, wherein the first feature and the second feature both include a predicted score of the category of the target object; a denoising module, configured to perform denoising processing on the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; A second acquisition module, used to acquire a KL divergence loss between the first feature and the third feature; The training module is used to train the student model through a back propagation algorithm according to the KL divergence loss with the goal of improving the accuracy of the output result of the student model until a preset stop condition is met.
[0013] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a model training method based on decoupled knowledge distillation as described in any one of the above is implemented.
[0014] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a model training method based on decoupled knowledge distillation as described in any one of the above.
[0015] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements a model training method based on decoupled knowledge distillation as described in any one of the above.
[0016] The present application provides a model training method, device, equipment and medium based on decoupled knowledge distillation. To implement the model training method based on decoupled knowledge distillation of the present application, first obtain the teacher model and the student model to be trained as the reference target in the training process. Then, the first sample image is input into the teacher model and the student model respectively to obtain the first feature output by the teacher model and the second feature output by the student model, and then the noise data in the second feature is removed by the preset diffusion model to obtain the third feature. Then, the KL divergence loss between the first feature and the third feature is obtained, and according to the KL divergence loss, the student model is trained by the back propagation algorithm with the goal of improving the accuracy of the output result of the student model until the preset stop condition is met. In the present application, the noise data in the features output by the student model is processed by the diffusion model, so that the precise alignment between the features output by the teacher model and the features output by the student model can be achieved, thereby narrowing the feature representation gap between the teacher model and the student model, so that the student model can fully learn the deep features in the teacher model, which can significantly improve the training effect of the student model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 is a flow chart of a model training method based on decoupled knowledge distillation shown in an embodiment of the present application; Figure 2 is a schematic diagram of a process for obtaining distillation loss shown in one embodiment of the present application; Figure 3 is a schematic diagram of the principle of decoupled knowledge distillation shown in an embodiment of the present application; Figure 4 It is a structural block diagram of a model training device based on decoupled knowledge distillation shown in an embodiment of the present application; Figure 5 It is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0020] Before introducing the method of the present application in detail, the traditional knowledge distillation is briefly described below.
[0021] The core of traditional knowledge distillation is to use the logits of the teacher model as soft labels to guide the student model, so that the student model can imitate the output of the teacher model and learn the knowledge in the teacher model. During the training process, the relative entropy loss (Kullback-Leibler Divergence, KL), also known as KL divergence, is usually used to measure the difference between the logits distribution of the student model and the logits distribution of the teacher model, and by optimizing the difference, the logits distribution of the student model is made as similar as possible to the logits distribution of the teacher model.
[0022] In machine learning and deep learning, especially in classification tasks, logits are usually the output of the last layer of a neural network (usually a fully connected layer or a linear layer), which represents the predicted score for each category. Suppose a neural network is used to identify three types of objects: cats, dogs, and cars. The last layer of the neural network may have three neurons, each corresponding to a category (cat, dog, car). The output of these three neurons is logits, which represent the network's original prediction score for each category. For example, for an input image, the logits output by the neural network can be: cat: 2.0, dog: 1.5, car: -0.5. These logits represent the strength or confidence that the neural network believes that the input image belongs to each category.
[0023] The general training process of traditional knowledge distillation is as follows: Step 1: Train the teacher model A large and complex neural network, called the teacher model, is trained using a large amount of labeled data.
[0024] Step 2: Generate soft targets After the teacher model is trained, its output is the probability distribution of each category, which is called soft target. The soft target can be obtained by predicting the training data using the teacher model.
[0025] Suppose there is an image classification task, the goal is to identify whether the object in the image is a cat, dog or car, and suppose there is a trained teacher model for predicting the input image. After inputting an image containing a cat into the teacher model, the teacher model processes the image and outputs the following probability distribution: cat: 0.85 (indicates that the image has an 85% probability of being a cat), dog: 0.10 (indicates that the image has a 10% probability of being a dog), car: 0.05 (indicates that the image has a 5% probability of being a car). These probability distributions are the soft targets output by the teacher model, and are also the logits of the teacher model.
[0026] Step 3: Set the distillation temperature In order to make the output of the teacher model smoother, the distillation temperature T is introduced. By adjusting the value of T, the smoothness of the probability distribution output by the teacher model can be controlled.
[0027] Step 4: Train the student model A small and simple neural network is used as the student model to imitate the behavior of the teacher model. The input of the student model is the same as the teacher model, but the structure is simpler and has fewer parameters.
[0028] Step 5: Calculate distillation loss The difference between the features output by the student model and the features output by the teacher model is calculated through a loss function, which is usually a modified version of the cross entropy loss called distillation loss.
[0029] Step 6: Optimize The loss function representing the distillation loss is used as the final loss function. Through optimization algorithms such as gradient descent, the loss value is back-propagated to converge to the minimum value, completing the training of the student model.
[0030] However, traditional knowledge distillation has at least the following defects: (1) There is a high degree of coupling in the logit, that is, the information of the target category and the non-target category are mixed together, which makes it difficult for the student model to imitate the teacher model; (2) When the feature representation gap between the teacher model and the student model is too large, the features output by the student model (hereinafter referred to as student features) cannot effectively align with the features output by the teacher model (hereinafter referred to as teacher features), thus affecting the distillation effect.
[0031] In order to solve the above problems, a decoupled knowledge distillation is also proposed in the related technology, which achieves the decoupling effect by splitting the information in the logit into two parts: the target category and the non-target category, so that the student model can learn the knowledge in the teacher model more clearly. Specifically, decoupled knowledge distillation changes the classic knowledge distillation into two parts: target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD). Target class knowledge distillation (TCKD) mainly focuses on transferring the knowledge of the target category in the training samples, while non-target class knowledge distillation (NCKD) transfers a large amount of dark knowledge, that is, the knowledge of the non-target category. During the training process, the losses of TCKD and NCKD are calculated separately, and the logits distribution of the student model is made closer to the logits distribution of the teacher model by optimizing these two losses.
[0032] The general training process of decoupled knowledge distillation is as follows: Step 1: Train the teacher model A large and complex neural network, called the teacher model, is trained using a large amount of labeled data.
[0033] Step 2: Generate soft targets After the teacher model is trained, its output is the probability distribution of each category, which is called soft target. The soft target can be obtained by predicting the training data using the teacher model.
[0034] Step 3: Re-formulate KD loss as TCKD loss and NCKD loss The classic KD loss is reformulated into two parts: target class knowledge distillation (TCKD) loss and non-target class knowledge distillation (NCKD) loss.
[0035] Step 4: Train the student model (introducing decoupling mechanism) A small and simple neural network is used as the student model. During the training of the student model, the TCKD loss and NCKD loss are calculated respectively, and the hyperparameters are introduced. and To balance these two parts of loss, we can get the final KD loss function.
[0036] Step 5: Optimize Based on the KD loss function in step 4, the loss value is back-propagated through optimization algorithms such as gradient descent to converge to the minimum value, completing the training of the student model.
[0037] It can be seen that the biggest difference between decoupled knowledge distillation and traditional knowledge distillation is that the KD loss is reformulated into two parts, TCKD and NCKD, and the hyperparameters α and β are introduced to balance the two parts of the loss. Therefore, the knowledge in the teacher model can be captured more accurately while maintaining the simplicity of the student model.
[0038] However, in actual use, decoupled knowledge distillation also has the following defects: (1) When there are obvious architectural differences between the teacher model and the student model, decoupled knowledge distillation still cannot effectively narrow the feature representation gap between the features output by the teacher model and the features output by the student model when dealing with the alignment problem between the features output by the teacher model and the features output by the student model, resulting in the student model being unable to fully learn the deep features of the teacher model.
[0039] In decoupled knowledge distillation, if there are significant architectural differences between the teacher model and the student model, the student model will have difficulty trying to align the features of the teacher model and will not be able to fully learn the deep features of the teacher model. For example, there is a complex teacher model (such as Residual Network-50, or ResNet-50) and a simple student model (such as ShuffleNet, a lightweight deep convolutional neural network architecture). Since the architectures of these two models are significantly different, their feature representation spaces may be very different. When using decoupled knowledge distillation for knowledge distillation, although the learning of the student model is optimized by decoupling target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD), if the feature representation space of the student model is too different from that of the teacher model, then the student model still cannot effectively imitate the deep features of the teacher model.
[0040] (2) The impact of noise on the student model is not addressed Since the student model is usually small, it may contain more noise during the training process. This noise comes from the data itself, the model structure, or the training process. In traditional knowledge distillation, this noise will interfere with the learning effect of the student model. Although decoupled knowledge distillation optimizes the distillation process by decoupling TCKD and NCKD, it does not directly address the noise problem.
[0041] (3) Ignoring non-target relationships In decoupled knowledge distillation, the non-target knowledge distillation parts of the teacher model and the student model are often regarded as irrelevant information, and the inter-class relationships in this information are not fully utilized. In fact, the relative ranking relationships between these classes are crucial to the learning process of the student model. Ignoring these relationships will lead to a decrease in the distillation effect.
[0042] For example, in a multi-category classification task, suppose there is a teacher model and a student model. The output probability distribution of the teacher model for a certain input sample is [0.7, 0.2, 0.1], corresponding to three categories respectively. When using decoupled knowledge distillation for knowledge distillation, if you only focus on the target class knowledge distillation TCKD, then the student model may only learn knowledge about the target category. However, the non-target class knowledge distillation NCKD contains information about the relative ranking relationship between non-target categories (such as category 2 is more likely than category 3). This information is very important for improving the generalization ability of the student model. If the information in the target class knowledge distillation NCKD is ignored, the student model cannot fully utilize all the knowledge of the teacher model, resulting in a decrease in the distillation effect.
[0043] In order to solve the three problems existing in the above-mentioned decoupled knowledge distillation, this application proposes a decoupled knowledge distillation method that combines diffusion models and inter-class ranking consistency: Diffusion Models (DM) are used to remove noise data in the student model to achieve a more accurate alignment of student features with teacher features to solve the above-mentioned problems (1) and (2). In addition, by introducing the method of inter-class ranking consistency, the inter-class ranking relationship of non-target categories between the teacher model and the student model is maintained during the distillation process, thereby improving the learning effect of the student model.
[0044] Figure 1 This is a flowchart of a model training method based on decoupled knowledge distillation shown in an embodiment of the present application. Figure 1 , a model training method based on decoupled knowledge distillation provided by the present application is described in detail. The method of the present application mainly includes steps 101 to 105. Steps 101 to 105 are used to solve the problems (1) and (2) existing in the above-mentioned decoupled knowledge distillation.
[0045] Step 101: obtain a teacher model used as a reference target in the training process and a student model to be trained. Both the teacher model and the student model are used to identify the category of the target object in the image. The accuracy of the output result of the teacher model is higher than the accuracy of the output result of the student model.
[0046] In this embodiment, in order to facilitate the understanding of the teacher model and the student model, in the subsequent embodiments, the teacher model and the student model are used to identify the category of the target object in the image as an example for explanation, wherein the target object can be the main object in the image or the object in the area to be identified, and the area to be identified and the target object can be defined according to actual needs.
[0047] Among them, the teacher model is a model with a higher degree of training, and its prediction results for the category of the target object are more accurate, while the student model is a model with a lower degree of training, and its prediction results for the category of the target object are less accurate than those of the teacher model.
[0048] In this embodiment, the teacher model can be obtained by training with ResNet32×4, and the student model can be obtained by training with ShuffleNetV2. Secondly, the student model can also be obtained by training with ResNet, VGG (Visual Geometry Group), MobileNet, and Wide ResNet (Wide Residual Network).
[0049] This application does not impose any specific restrictions on the network structure used by the teacher model and the student model, and can be selected according to actual needs.
[0050] Step 102: input the first sample image into the teacher model and the student model respectively to obtain a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object.
[0051] In this embodiment, the first feature output by the teacher model is the logit part of the teacher model, and the second feature output by the student model is the logit part of the student model. Both the first feature and the second feature contain the predicted score of the category of the target object.
[0052] For example, the first feature can be: cat: 0.8, dog: 0.6, fox: 0.4, car: 0.2, where 0.8, 0.6, 0.4 and 0.2 are all prediction scores obtained by the teacher model. The second feature can be: cat: 0.7, dog: 0.5, fox: 0.4, car: 0.1, where 0.7, 0.5, 0.4 and 0.1 are all prediction scores obtained by the student model. The higher the prediction score, the higher the probability of the corresponding category, that is, both the teacher model and the student model believe that the probability that the target object is a cat is the highest.
[0053] Step 103, denoising the noise data in the second feature through a preset diffusion model to obtain a third feature, where the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the features output by the student model and the features output by the teacher model.
[0054] In this embodiment, the diffusion model can be obtained by pre-training according to the second sample image. The second sample image can be the same as the first sample image or different. The number of the second sample images can be set according to actual needs.
[0055] Among them, the diffusion model is a nonlinear processing method for image enhancement and image restoration. Its core is to gradually convert data into noise by simulating the physical diffusion process, and then learn the inverse process to gradually restore the original data from the noise, thereby achieving high-quality generation effects.
[0056] The diffusion model includes two processes: forward diffusion and reverse diffusion. Forward diffusion refers to the process of gradually adding Gaussian noise from the original data until the data becomes pure Gaussian noise. Each step of this process controls the amount of noise added according to the preset variance schedule. Reverse diffusion is the process of starting from pure Gaussian noise and gradually removing the noise to restore the original data. This process relies on a parameterized neural network (such as a noise prediction network) that learns to predict and remove the noise added at each step. The forward diffusion process can be described as a Markov chain that transforms the data distribution into a Gaussian distribution by gradually adding noise. The reverse diffusion process is also a Markov chain, but in the opposite direction, and the original data is restored from the noise by gradually removing the noise.
[0057] Due to the gap between the teacher model and the student model, these noises cannot be eliminated by simply imitating the teacher model in distillation, but the teacher model and the student model can focus more on valuable information. Inspired by the diffusion model, this application regards the student features as the noisy version of the teacher features. First, a noise prediction network is trained using the teacher features, and then the trained noise prediction network (which can remove noise from the noisy student features and restore purer data that is closer to the teacher features) is used as a diffusion model to denoise the student features.
[0058] Specifically, the process of training the diffusion model may include: Step 1: Get the second sample image, input the second sample image into the teacher model and the student model respectively, obtain the feature Ft output by the teacher model, and the feature Fs output by the student model. The feature Ft will be used as the target for training the diffusion model.
[0059] Step 2: Select a neural network as the noise prediction network, and train the noise prediction network based on the principle of the diffusion model using the feature Ft as the target. During the training process, the noise prediction network will try to remove the noise from the noisy feature Fs to produce a purer feature. This step can usually be achieved by minimizing the loss function, which is used to measure the difference between the output of the diffusion model and the feature Ft. The loss function can specifically use the mean square error loss function.
[0060] Step 3: During or after training, evaluate the performance of the noise prediction network. Specifically, this can be done by comparing the output of the noise prediction network with the output of the teacher model on the validation set. If the performance is not ideal, the architecture, training strategy, or loss function of the noise prediction network can be adjusted for retraining.
[0061] Step 4: After the noise prediction network is trained, it can be used as a diffusion model to denoise the student features with noise to extract more valuable features. The denoised features will be closer to the teacher features and can improve the performance of the student model.
[0062] Step 104: Obtain the KL divergence loss between the first feature and the third feature.
[0063] Specifically, step 104 may include: Step 1041: Obtain a KL divergence loss based on the target category according to the difference between the feature representing the target category in the first feature and the feature representing the target category in the third feature, where the target category is a correct category of the target object in the first sample image; Step 1042: Obtain a KL divergence loss based on the non-target category according to the difference between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature, where the non-target category is an erroneous category of the target object in the first sample image; Step 1043: Determine the KL divergence loss between the first feature and the third feature according to the KL divergence loss of the target category and the KL divergence loss based on the non-target category.
[0064] In this example, the TCKD loss and NCKD loss are calculated separately, and the hyperparameters and To balance these two losses, we get the final KD loss function. That is, KD loss function = KL divergence loss between the first feature and the third feature = *TCKD Loss+ *NCKD loss.
[0065] Wherein, the target category is the true category of the target object in the input image. For example, if an image containing a cat is input into the teacher model, the first feature output by the teacher model is: cat: 0.8, dog: 0.6, fox: 0.4, car: 0.2, where cat is the target category, and dog, fox and car are non-target categories. For example, if an image containing a dog is input into the teacher model, the first feature output by the teacher model is: dog: 0.8, cat: 0.6, fox: 0.4, car: 0.2, where dog is the target category, and cat, fox and car are non-target categories.
[0066] Step 105: According to the KL divergence loss, the student model is trained by a back propagation algorithm with the goal of improving the accuracy of the output result of the student model until a preset stop condition is met.
[0067] Execute step 105, based on the KD loss function, back propagate the loss value through optimization algorithms such as gradient descent, so that it converges to the minimum value, and complete the training of the student model. Reaching the preset stop condition indicates that it has converged to the minimum value.
[0068] The preset stop condition may be: the KD loss function obtains a minimum value, or reaches a preset number of iterations, or the value of the KD loss function is less than a preset threshold. The preset stop condition may be set according to actual needs.
[0069] To implement the model training method based on decoupled knowledge distillation of the present application, first obtain the teacher model and the student model to be trained as reference targets during the training process. Then, input the first sample image into the teacher model and the student model respectively to obtain the first feature output by the teacher model and the second feature output by the student model, and then remove the noise data in the second feature through the preset diffusion model to obtain the third feature. Then, obtain the KL divergence loss between the first feature and the third feature, and according to the KL divergence loss, with the goal of improving the accuracy of the output result of the student model, train the student model through the back propagation algorithm until the preset stop condition is met. In the present application, the noise data in the features output by the student model is processed by the diffusion model, so that the precise alignment between the features output by the teacher model and the features output by the student model can be achieved, thereby narrowing the feature representation gap between the teacher model and the student model, so that the student model can fully learn the deep features in the teacher model, which can significantly improve the training effect of the student model.
[0070] Although decoupled knowledge distillation proves the importance of NCKD, it and traditional knowledge distillation both use the method of using the student model to accurately match the prediction score output by the teacher model, which cannot reflect the relationship structure between the various classes in the teacher model and the student model. A network with a deeper scale architecture will achieve better performance, so in actual implementation, a large teacher network model is usually selected, which also leads to a large difference in scale between the teacher model and the student model. Under this gap, using the traditional KL divergence loss to improve the accuracy of the student model's prediction results is relatively weak.
[0071] Taking the above problems into consideration, the present application proposes a method for relationship matching using the Pearson correlation coefficient. This method no longer cares about the predicted scores of the NCKD part in the teacher model, but pays attention to the relative ranking relationship between the non-target classes predicted by the teacher model. Specifically, the teacher model and the student model can be ranked for the predicted scores of each non-target class respectively, and then the Pearson correlation coefficient between the two rankings is calculated. In the distillation process, the present application expects that the ranking relationship of the NCKD part output by the student model is as consistent as possible with the ranking relationship of the NCKD part output by the teacher model. Therefore, the Pearson correlation coefficient can be maximized by optimizing the pre-designed objective function, thereby realizing the transfer of knowledge from the teacher model to the student model. The method will be described in detail below.
[0072] In combination with the above embodiments, in one implementation, obtaining the KL divergence loss based on the non-target category according to the difference between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature may include: Determine the Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature; Based on the Pearson correlation coefficient, the KL divergence loss based on the non-target category is determined.
[0073] Wherein, determining the Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature includes: The Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature is determined by the following formula: in, represents the Pearson correlation coefficient, represents the number of non-target categories, Represents the output of the teacher model The prediction scores of non-target categories, represents the average of the prediction scores of each non-target category output by the teacher model, The student model determines the The prediction scores of non-target categories, Represents the average of the prediction scores for each non-target class determined by the student model.
[0074] The range of Pearson's correlation coefficient is [-1, 1]. r = 1 indicates a perfect positive correlation with identical rankings, r = -1 indicates a perfect negative correlation with opposite rankings, and r = 0 indicates no correlation.
[0075] For example, the prediction scores of the non-target classes output by the teacher model are [0.8, 0.6, 0.4, 0.2], and the prediction scores of the non-target classes output by the student model are [0.7, 0.5, 0.3, 0.1]. The teacher score ranking is: 1>2>3>4. The student score ranking is: 1>2>3>4. Therefore, the inter-class ranking relationship of the non-target classes of the teacher model is consistent with the inter-class ranking relationship of the non-target classes of the student model, and the Pearson correlation coefficient is close to 1.
[0076] In Decoupled Knowledge Distillation (DKD), the relationship matching is performed through the Pearson correlation coefficient, focusing on the relative ranking between non-target classes rather than the exact value. Using the idea of this application, the relative size (ranking relationship) between the output prediction scores of the teacher network is preserved. The student network does not need to accurately copy the prediction scores of the teacher model, but only needs to maintain a ranking relationship consistent with the teacher model.
[0077] Figure 2 FIG. 1 is a schematic diagram of a process for obtaining distillation loss according to an embodiment of the present application. Figure 2 , this application extracts the ranking relationship of each non-target class in NCKD, and matches the ranking relationship of NCKD in the student model with the ranking relationship of NCKD in the teacher model. Compared with the one-to-one exact matching method in the traditional scheme, the method based on inter-class ranking correlation in this application only needs to keep the student model in a ranking relationship similar to that of the teacher model, and then match the relationship between the student model and the teacher model, which can better express the composition of knowledge.
[0078] After introducing the inter-class ranking consistency method, the process of determining the KL divergence loss (ie, KD loss) based on the non-target category according to the Pearson correlation coefficient in this application may include: Step 1: Get the prediction score output by the teacher model, that is, get the prediction score of the non-target category in the first feature, expressed as , and rank these scores. For example, the predicted scores are [0.8, 0.6, 0.4, 0.2], and the inter-class ranking relationship of non-target categories is 1>2>3>4.
[0079] Step 2: Get the prediction scores output by the student model That is, get the prediction score of the non-target category in the third feature, expressed as , and rank these scores. For example, the predicted scores are [0.7, 0.5, 0.3, 0.1], and the inter-class ranking relationship of non-target categories is 1>2>3>4.
[0080] Step 3: Record the vector of predicted scores for non-target categories output by the teacher model as: , the vector of predicted scores of non-target categories output by the student model is recorded as: , and calculate the Pearson correlation coefficient between the teacher and student non-target class scores .
[0081] Step 4: Minimize the relation matching loss First, define the relationship matching loss , the goal is to maximize the Pearson correlation coefficient , that is, minimize the relationship matching loss .
[0082] Step 5: Obtain KD loss function Matching the relationship to the loss Combined with TCKD loss, we get the final loss function .
[0083] The principle of using the Pearson correlation coefficient to achieve inter-class ranking consistency in this application will be explained in detail below.
[0084] For a training sample from the tth class, the logits of the teacher model and the student model are and ,in The classification probability can be expressed as ,in Indicates The probability of the categories, Indicates the number of categories. Classification probability Each element can be Function and Temperature Factor Conduct an assessment.
[0085] Non-target categories (excluding target categories) ,That The probability of each non-target category in can be expressed as: For disentangled knowledge distillation, it is necessary to disentangle the target category and the non-target category, so we define , represents the binary probability of the target category and all non-target categories, represents the probability of the target category, represents the probability of non-target classes.
[0086] According to the above , , as well as , it can be inferred that , so the KD loss can be expressed as: in, Represents the similarity of the binary probabilities of the target category (teacher) and the target category (student), and is named . Represents the similarity of the probability of the teacher model and the student model in the non-target class, named . Weight Too small, making It does not work. Considering the coupling of weights, two parameters are introduced. and Substitute the weights. Then It can be expressed as: For the prediction vector and ,like ,but and An exact matching distance of 0 can be obtained between them, that is, This application uses a more relaxed matching method, so the relationship mapping is introduced and , and get: The method of this application focuses on and The internal relationship between and In order not to affect The information contained in the vector, mapping and should be equal, so choose the identity transformation and get: ,in , , , are all constants.
[0087] In this application, the Pearson correlation coefficient has one most important feature: when the variable and When the position of changes, the Pearson correlation coefficient will not change. and Move to and (where a, b, c and d are all constants) does not cause and The change in the correlation coefficient is exactly in line with the characteristics pursued by this application. In this example, the Pearson correlation coefficient It is expressed as: in yes and The covariance of and express and The mean of Represents standard deviation.
[0088] Therefore, the Pearson distance and Pearson correlation coefficient The relationship can be expressed as: Therefore, the final KD loss function of this application is: In combination with the above embodiments, in one implementation, obtaining the KL divergence loss between the first feature and the third feature may include: Compress the number of channels of the first feature to obtain the fourth feature; Get the KL divergence loss between the fourth feature and the third feature.
[0089] Among them, compressing the number of channels of the first feature and realizing the dimensionality reduction processing of the first feature can effectively save computing resources and speed up the model training process. Secondly, when training the diffusion model, the number of channels of the teacher feature can also be compressed to speed up the training process.
[0090] In combination with the above embodiments, in one implementation, obtaining the KL divergence loss between the fourth feature and the third feature includes: Performing feature reconstruction on the fourth feature to obtain a feature-reconstructed fourth feature, wherein the feature reconstruction is used to reduce redundant information in the fourth feature; Get the KL divergence loss between the fourth feature and the third feature after feature reconstruction.
[0091] In this embodiment, feature reconstruction is performed on the fourth feature to remove redundant information, unnecessary information, etc. in the fourth feature, thereby ensuring that the reconstructed fourth feature is a valuable feature, thereby improving the training effect of the student model.
[0092] In combination with the above embodiments, in one implementation, obtaining the KL divergence loss between the first feature and the third feature includes: Aligning the dimension of the third feature with the dimension of the first feature; Get the KL divergence loss between the first feature after alignment and the third feature.
[0093] In this embodiment, the third feature is projected to the same dimension as the first feature, that is, the noisy third feature output by the student model is dimensionally projected through a convolutional layer (e.g., Conv 1x1) to ensure that the dimension of the third feature is consistent with that of the first feature for input into the diffusion model.
[0094] Figure 3 Schematic diagram of a principle of decoupled knowledge distillation shown in an embodiment of the present application. Figure 3 ,The overall process of training the student model through knowledge distillation in this application is as follows: Step 1: Input the first sample image (i.e., input) to the teacher model (Tea) and the student model (Stu) respectively to obtain the teacher features and student features. Since the student model is small or not trained sufficiently, the student features have some noise.
[0095] Step 2: Input the teacher feature into an autoencoder, which consists of two convolutional neural networks (including 1x1 convolutional layers), one for reducing the number of feature channels and the other for reconstructing the teacher feature. First, the number of channels of the teacher feature is compressed through the upper 1x1 convolutional layer to reduce the feature dimension, and the latent teacher feature obtained after compression is input into the next 1x1 convolutional layer for feature reconstruction to obtain the reconstructed teacher feature.
[0096] Among them, the autoencoder uses reconstruction loss (Rec. Loss) for training. The reconstruction loss is the mean square error between the features output by the original teacher model and the reconstructed teacher features. That is, the teacher features reconstructed by the decoder after channel compression are compared with the original teacher features, and the mean square error (MSE) is calculated. The autoencoder is trained based on the mean square error.
[0097] Step 3: The student features are input into a 1x1 convolutional layer for dimension projection to ensure that the dimension of the student features is consistent with the dimension of the potential teacher features for denoising. The dimensionally projected student features are input into the noise adapter (NoiseAdapter), which treats the student features as noisy versions of the teacher features and inputs them into the diffusion model, such as Figure 3 As shown in the Diffusion Model.
[0098] Step 4: The diffusion model uses the reconstructed teacher feature as the target to remove the noise from the noisy student feature to obtain the denoised student feature.
[0099] Step 5: Based on the reconstructed teacher features and denoised student features, use KL divergence loss to perform distillation supervision on the teacher features and student features. The distillation loss of the target category part (Target Class KL Loss) is obtained according to the difference between the features representing the target category in the teacher features and the features representing the target category in the student features. The distillation loss of the non-target category part (Non-Target Class KL Loss) is obtained according to the difference between the features representing the non-target category in the teacher features and the features representing the non-target category in the student features. Hyperparameters are introduced. and Balancing the two losses yields the final loss function.
[0100] Step 6: Based on the final loss function, back-propagate the loss value through optimization algorithms such as gradient descent to converge it to the minimum value and complete the training of the student model.
[0101] exist Figure 3 In the example, the reconstruction loss (Rec. Loss) is used to ensure that the autoencoder can accurately reconstruct the teacher features. In this embodiment, the autoencoder can be trained using only the reconstruction loss, which is the original teacher and reconstructed teacher characteristics The mean square error between them is: ,in, represents the mean square error.
[0102] Diffusion loss (Diff Loss) is used to measure the denoising effect of the diffusion model. KD loss includes the distillation loss of the target category part (Target Class KL Loss) and the distillation loss of the non-target category part (Non-Target Class KL Loss).
[0103] Since the size of the teacher feature is relatively large, the denoising process of the diffusion model consumes a lot of computing resources, and it is necessary to forward the noise prediction network T times (T = 5 in this application method) to denoise the student features and train the noise prediction network with the teacher feature once. When the dimension of the teacher feature is large, T + 1 forwarding will result in high computational cost. To solve this problem, this application extracts a lightweight diffusion model composed of two bottleneck blocks superimposed together from ResNet as the noise prediction network.
[0104] exist Figure 3In the , the latent teacher feature (Latent feature) used to train the diffusion model is separated from the diffusion model (the latent teacher feature is the teacher feature compressed by the autoencoder and is not directly trained in conjunction with the diffusion model), and the diffusion model has no backward gradient (the diffusion model itself does not return gradients during training, i.e., Nogradient), which can effectively reduce the consumption of computing resources and speed up the training process.
[0105] In this application, MSE loss is used to measure the mean square error between denoised student features and reconstructed teacher features, and KL divergence loss is used to measure the distribution difference between denoised student features and teacher features. Then in the logits part, the knowledge distillation loss is decomposed into two parts: (1) binary classification prediction of target category and non-target category (emphasizing the accuracy of the target category, making the student model closer to the teacher model in the main category), (2) multi-classification prediction of non-target category (balancing the distribution of other non-target categories, making the overall output more comprehensive, and avoiding the problem of excessive bias towards the target category).
[0106] In this application, MSE loss is used for denoising supervision at the feature level, directly constraining the student features to be close to the teacher features. KL divergence loss and KD loss are used for supervision at the output probability distribution level, and the feature distribution of the student model output is made close to the feature distribution of the teacher model output through distillation. Target Class KL Loss and Non-Target ClassKL Loss are further refinements of KD loss, focusing on the distribution of target category and non-target category outputs respectively, to achieve a more detailed distillation supervision effect.
[0107] In one embodiment, the method of the present application is implemented using five NVIDIA GeForce RTX 2080 Ti GPUs under the PyTorch framework based on the Linux system, and the implementation process covers a variety of network architectures, including ResNet, VGG, ShuffleNet, MobileNet and Wide ResNet. For the CIFAR100 (Canadian Institute for Advanced Research - 100 classes) dataset, the Batchsize (the number of training samples used in one model training iteration) is set to 128, and different initial learning rates are configured for the student model. The specific values are shown in Table 1. All student models are trained for 240 epochs (the number of times the entire training dataset is forward propagated and back-propagated in the neural network. After one epoch is completed, the model will update the weights according to the back-propagation algorithm), where after 150 epochs, the learning rate decays by 0.1 times every 30 epochs. For the ImageNet-1K dataset, the standard training process is followed, the Batchsize is set to 256, and all student models are trained for 100 epochs. The learning rate is initialized to 0.1 and decays after every 30 epochs. In this application method, the hyperparameters of the total loss function are set to 1.0 and 8.0.
[0108] Table 1 CIFAR100 is a classic image classification dataset that contains 100 different categories, each of which contains 600 32x32 pixel color images. These 100 categories are divided into 20 major categories, each of which contains 5 minor categories. The size of each image is 32x32 pixels and contains three RGB channels. CIFAR100 is often used to evaluate the performance of image classification algorithms.
[0109] The ImageNet-1K dataset is a dataset released as part of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC), and the number of categories is 1000. The training set contains about 1.28 million images, the validation set contains 50,000 images, and the test set contains 100,000 images.
[0110] In this application, a diffusion model is used to denoise the student features. However, the diffusion model has a large demand for computational complexity, and its applicability may be limited to certain extent for resource-constrained devices (such as mobile devices or embedded systems). Therefore, this application can also use other types of denoising models, such as variational autoencoders (VAE) and denoising autoencoders (DA) for denoising. Specifically, the teacher features are input into the denoising autoencoder for training, and then the student features are denoised by the trained denoising model. The denoising loss can be calculated by the mean square error (MSE) or other denoising loss functions. For example, on the CIFAR-100 dataset, the teacher model uses ResNet-50, and the student model uses a smaller ResNet-18. The features output by ResNet-50 are trained by the denoising autoencoder, and then the features output by the student ResNet-18 are denoised using the trained model.
[0111] Compared with the diffusion model, the denoising autoencoder or variational autoencoder usually has lower computational complexity, can complete the denoising task more quickly, and is more suitable for real-time processing or resource-constrained application scenarios.
[0112] Contrastive Learning can extract effective features by optimizing the similarities and differences between sample pairs. Introducing contrastive learning into knowledge distillation can be used as an alternative distillation strategy, especially for scenarios without labeled data. Through contrastive learning, the student model achieves the distillation effect by narrowing the similarity of the features with the teacher model and pushing the distance between dissimilar samples farther. Specifically, in the process of knowledge distillation, contrastive loss (such as InfoNCE) is used to replace the traditional distillation loss. The student model optimizes its representation ability by comparing the feature vector with the teacher model. Secondly, a two-way contrastive learning method can be used to let the student model compare in the space of the teacher model features and optimize the feature representation of the student model. Assume that ResNet is used as the teacher model and the student model. For example, on the CIFAR-100 dataset, the student model is trained through contrastive learning to make its features similar to those of the teacher model and to make them more different from non-target categories, thereby improving the performance of the student model in classification tasks. Knowledge distillation based on contrastive learning does not require label information and has strong adaptability. Contrastive learning can perform feature learning without label information and is suitable for scenarios where data is scarce or labels cannot be obtained. This method can adaptively learn the relationship between samples and is particularly suitable for images or multimodal data.
[0113] Quantization technology is usually used for model compression. It quantizes floating-point weights and activations into smaller integers, thereby reducing the storage requirements and computational complexity of the model. In this application, quantization technology can be combined with knowledge distillation to quantize teacher features and pass them to the student model for training. Quantization can not only reduce the size of the model, but also improve computational efficiency in some cases. Specifically, a quantization strategy is used to quantize the teacher features, and the quantized features are passed to the student model. The student model is trained based on the quantized features and learns a more concise knowledge representation. The quantization error loss is added to the loss function to measure the error introduced in the quantization process and to perform corresponding optimization. For example, assuming that the teacher model is ResNet-50 and the student model is a lightweight MobileNetV2, the teacher features are quantized through quantization technology, and then the quantized teacher features are used for knowledge distillation, and finally an efficient and highly accurate student model is obtained. The knowledge distillation method based on quantization has a higher compression rate and can accelerate the reasoning process. Quantization can greatly reduce the storage requirements of model parameters, thereby adapting to resource-constrained devices. Through quantization, the amount of calculation in the reasoning process is significantly reduced, which can improve the reasoning speed.
[0114] Self-supervised learning is an unsupervised learning method that learns features by generating tasks without relying on manual labels. In this application, self-supervised learning can be combined with knowledge distillation. Specifically, the teacher model is pre-trained using a self-supervised learning method, and then the knowledge obtained through self-supervised learning is passed to the student model. In this way, the student model is not only distilled through the teacher model, but also can use self-supervised learning for feature learning. On the teacher model, a self-supervised learning task (such as a generative model task or a contrastive learning task) is first performed, and then the features obtained in the self-supervised learning process are passed to the student model. The student model is trained by combining traditional knowledge distillation loss and self-supervised learning loss. For example, suppose a self-supervised pre-trained ResNet-50 is used as the teacher model, and the student model is ResNet-18. On the CIFAR-100 dataset, the teacher model obtains the deep features of the image through self-supervised learning, and then passes these features to the student model, and the performance of the student model is further improved through knowledge distillation. Combining self-supervised learning with knowledge distillation has the advantages of applicability to unlabeled data and enhanced model generalization ability. Self-supervised learning does not rely on labeled data and is suitable for scenarios where labels are scarce. Self-supervised learning can help the model learn stronger feature expression capabilities without labels, thereby improving the effect of distillation.
[0115] In summary, the method of the present application has at least the following technical effects: (1) Effective denoising: Through the diffusion model and adaptive noise matching module, the noise in the student model is effectively removed, enabling it to learn the knowledge of the teacher model more accurately.
[0116] (2) Optimizing the learning of non-target categories: Through a method based on inter-class ranking consistency, the learning effect of the student model in the non-target category part is optimized, making the distillation process more comprehensive.
[0117] (3) Improved accuracy: On the CIFAR-100 and ImageNet datasets, compared with the traditional decoupled knowledge distillation method, this application can significantly improve the classification accuracy of the student model, especially when the architectures of the teacher and student models are very different.
[0118] The model training device based on decoupled knowledge distillation provided in the present application is described below. The model training device based on decoupled knowledge distillation described below and the model training method based on decoupled knowledge distillation described above can refer to each other.
[0119] Figure 4 1 is a structural block diagram of a model training device based on decoupled knowledge distillation shown in an embodiment of the present application. Figure 4 , the model training device 400 based on decoupled knowledge distillation of the present application includes: A first acquisition module 401 is used to acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; An input module 402, configured to input a first sample image into the teacher model and the student model respectively, to obtain a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object; A denoising module 403 is used to denoise the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; A second acquisition module 404, configured to acquire a KL divergence loss between the first feature and the third feature; The training module 405 is used to train the student model through a back propagation algorithm according to the KL divergence loss with the goal of improving the accuracy of the output result of the student model until a preset stop condition is met.
[0120] According to the model training device 400 based on decoupled knowledge distillation of the present application, the second acquisition module 404 includes: A first acquisition submodule, configured to acquire a KL divergence loss based on the target category according to a difference between a feature representing the target category in the first feature and a feature representing the target category in the third feature, wherein the target category is a correct category of the target object in the first sample image; a second acquisition submodule, configured to acquire a KL divergence loss based on the non-target category according to a difference between a feature representing a non-target category in the first feature and a feature representing the non-target category in the third feature, wherein the non-target category is an erroneous category of the target object in the first sample image; The first determination submodule is configured to determine a KL divergence loss between the first feature and the third feature according to the KL divergence loss based on the target category and the KL divergence loss based on the non-target category.
[0121] According to the model training device 400 based on decoupled knowledge distillation of the present application, the first acquisition submodule includes: A second determination submodule, configured to determine a Pearson correlation coefficient between a feature in the first feature representing the non-target category and a feature in the third feature representing the non-target category; The third determination submodule is used to determine the KL divergence loss based on the non-target category according to the Pearson correlation coefficient.
[0122] According to the model training device 400 based on decoupled knowledge distillation of the present application, the second determination submodule includes: The fourth determination submodule is used to determine the Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature by using the following formula: in, represents the Pearson correlation coefficient, represents the number of non-target categories, Represents the output of the teacher model The prediction scores of non-target categories, represents the average of the prediction scores of each non-target category output by the teacher model, The student model determines the The prediction scores of non-target categories, Represents the average of the prediction scores for each non-target class determined by the student model.
[0123] According to the model training device 400 based on decoupled knowledge distillation of the present application, the second acquisition module 404 includes: A compression submodule, used for compressing the number of channels of the first feature to obtain a fourth feature; The third acquisition submodule is used to acquire the KL divergence loss between the fourth feature and the third feature.
[0124] According to the model training device 400 based on decoupled knowledge distillation of the present application, the second acquisition module 404 includes: A reconstruction submodule, used for performing feature reconstruction on the fourth feature to obtain a feature-reconstructed fourth feature, wherein the feature reconstruction is used to reduce redundant information in the fourth feature; The fourth acquisition submodule is used to obtain the KL divergence loss between the fourth feature after the feature reconstruction and the third feature.
[0125] According to the model training device 400 based on decoupled knowledge distillation of the present application, the second acquisition module 404 includes: an alignment submodule, configured to align the dimension of the third feature with the dimension of the first feature; The fifth acquisition submodule is used to acquire the KL divergence loss between the aligned first feature and the third feature.
[0126] Figure 5 is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application, such as Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a model training method based on decoupled knowledge distillation, the method comprising: Acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; Inputting a first sample image into the teacher model and the student model respectively, obtaining a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object; De-noising the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; Obtaining a KL divergence loss between the first feature and the third feature; According to the KL divergence loss, with the goal of improving the accuracy of the output result of the student model, the student model is trained through a back propagation algorithm until a preset stop condition is met.
[0127] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0128] On the other hand, the present application also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute a model training method based on decoupled knowledge distillation provided by the above methods, the method comprising: Acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; Inputting a first sample image into the teacher model and the student model respectively, obtaining a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object; De-noising the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; Obtaining a KL divergence loss between the first feature and the third feature; According to the KL divergence loss, with the goal of improving the accuracy of the output result of the student model, the student model is trained through a back propagation algorithm until a preset stop condition is met.
[0129] On the other hand, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform a model training method based on decoupled knowledge distillation provided by the above methods, the method comprising: Acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; Inputting a first sample image into the teacher model and the student model respectively, obtaining a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object; De-noising the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; Obtaining a KL divergence loss between the first feature and the third feature; According to the KL divergence loss, with the goal of improving the accuracy of the output result of the student model, the student model is trained through a back propagation algorithm until a preset stop condition is met.
[0130] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0131] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A model training method based on decoupled knowledge distillation, characterized in that: include: Acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; Inputting a first sample image into the teacher model and the student model respectively, obtaining a first feature output by the teacher model and a second feature output by the student model, wherein both the first feature and the second feature include a prediction score of the category of the target object; Denoising the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; Obtaining a KL divergence loss between the first feature and the third feature; According to the KL divergence loss, with the goal of improving the accuracy of the output result of the student model, the student model is trained through a back propagation algorithm until a preset stop condition is met.
2. The model training method according to claim 1, characterized in that: The obtaining of the KL divergence loss between the first feature and the third feature includes: Obtaining a KL divergence loss based on the target category according to a difference between a feature representing a target category in the first feature and a feature representing the target category in the third feature, where the target category is a correct category of the target object in the first sample image; Obtaining a KL divergence loss based on the non-target category according to a difference between a feature representing a non-target category in the first feature and a feature representing the non-target category in the third feature, wherein the non-target category is an erroneous category of the target object in the first sample image; A KL divergence loss between the first feature and the third feature is determined according to the KL divergence loss based on the target category and the KL divergence loss based on the non-target category.
3. The model training method according to claim 2, characterized in that: The obtaining, according to a difference between a feature representing a non-target category in the first feature and a feature representing the non-target category in the third feature, a KL divergence loss based on the non-target category, comprises: Determine a Pearson correlation coefficient between a feature in the first feature representing the non-target category and a feature in the third feature representing the non-target category; The KL divergence loss based on the non-target category is determined according to the Pearson correlation coefficient.
4. The model training method according to claim 3, characterized in that: The determining of the Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature includes: The Pearson correlation coefficient between the feature representing the non-target category in the first feature and the feature representing the non-target category in the third feature is determined by the following formula: in, represents the Pearson correlation coefficient, represents the number of non-target categories, Represents the output of the teacher model The prediction scores of non-target categories, represents the average of the prediction scores of each non-target category output by the teacher model, The student model determines the The prediction scores of non-target categories, Represents the average of the prediction scores for each non-target class determined by the student model.
5. The model training method according to claim 1, characterized in that: The obtaining of the KL divergence loss between the first feature and the third feature includes: Compressing the number of channels of the first feature to obtain a fourth feature; A KL divergence loss between the fourth feature and the third feature is obtained.
6. The model training method according to claim 5, characterized in that: The obtaining of the KL divergence loss between the fourth feature and the third feature includes: Performing feature reconstruction on the fourth feature to obtain a feature-reconstructed fourth feature, wherein the feature reconstruction is used to reduce redundant information in the fourth feature; Obtain a KL divergence loss between a fourth feature after the feature reconstruction and the third feature.
7. The model training method according to any one of claims 1 to 6, characterized in that: The obtaining of the KL divergence loss between the first feature and the third feature includes: aligning the dimensions of the third feature with the dimensions of the first feature; Obtain the KL divergence loss between the aligned first feature and the third feature.
8. A model training device based on decoupled knowledge distillation, characterized in that: include: A first acquisition module is used to acquire a teacher model used as a reference target in a training process and a student model to be trained, wherein both the teacher model and the student model are used to identify the category of a target object in an image, and the accuracy of an output result of the teacher model is higher than the accuracy of an output result of the student model; An input module, used to input a first sample image into the teacher model and the student model respectively, to obtain a first feature output by the teacher model and a second feature output by the student model, wherein the first feature and the second feature both include a predicted score of the category of the target object; a denoising module, configured to perform denoising processing on the noise data in the second feature by using a preset diffusion model to obtain a third feature, wherein the diffusion model is pre-trained based on a noise prediction network, with the second sample image as input and with the goal of minimizing the mean square error loss between the feature output by the student model and the feature output by the teacher model; A second acquisition module, used to acquire a KL divergence loss between the first feature and the third feature; The training module is used to train the student model through a back propagation algorithm according to the KL divergence loss with the goal of improving the accuracy of the output result of the student model until a preset stop condition is met.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements a model training method based on decoupled knowledge distillation as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements a model training method based on decoupled knowledge distillation as described in any one of claims 1 to 7.