Image classification model knowledge distillation method based on text prompt
Through the knowledge distillation method of image classification model based on text prompts, the text encoder and teacher model parameters are frozen, and the learningable vectors are used for pre-training and KL divergence calculation, the calculation resource and time cost problems of traditional knowledge distillation are solved, and the learning accuracy and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510283335.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional knowledge distillation methods require multiple training, consume a lot of computing resources and time costs, and training small models may be time-consuming and may not necessarily perform well.
The knowledge distillation method of image classification model based on text prompts is adopted. By constructing a knowledge distillation model framework, the parameters of the text encoder and teacher model are frozen, and the learningable prompt vectors are used for comparison learning pre-training, and the KL divergence calculation loss is introduced in the distillation training to optimize the training process of the student model.
It reduces uncertainty in the training process of students' model, improves the learning accuracy and generalization ability of the model, optimizes the model compression performance, enables the student model to converge to key information points faster, and enhances the generalization ability of the model.
Smart Images

Figure CN120236122A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge distillation methods in computer vision, and relates to a knowledge distillation method for an image classification model based on text prompts. Background Art
[0002] On the one hand, deep network model compression and acceleration require reducing the parameters of the model. More importantly, it is necessary to reduce the model inference operation time while maintaining the performance accuracy, control the depth and total number of parameters of the model within a reasonable range, and thus meet the transplantation and deployment of actual applications. Therefore, how to minimize the computing resources while ensuring the accuracy is the key technology for deep neural network compression in multi-scenario applications.
[0003] Knowledge distillation takes the full exploration of the performance of neural networks as the research starting point, aiming to obtain a network structure with a shallower network level and a more streamlined structure while ensuring performance. Without changing the structure of the target network, knowledge distillation conducts collaborative training through a complex network (teacher network) to assist a streamlined network (student network), aiming to enhance the performance of the student network, enabling the student network to more closely approximate the performance of the teacher network, and to a certain extent realizing the operation of compressing the teacher network into the student network. From the perspective of network lightweight design, knowledge distillation aims to obtain a small network model with a shallower network layer and better performance. Traditional knowledge distillation requires multiple trainings, consuming a large amount of computing resources and time costs. Training a small model may be more time-consuming than training a large model and requires more iterations to obtain good results, and may not necessarily achieve good performance based on the time cost consumed. Summary of the Invention
[0004] The purpose of the present invention is to provide a knowledge distillation method for an image classification model based on text prompts, which solves the problem that traditional knowledge distillation requires multiple trainings and consumes a large amount of computing resources and time costs.
[0005] The technical solution adopted by the present invention is a knowledge distillation method for an image classification model based on text prompts, and the specific operation steps are as follows: Step 1, obtain a picture classification data set. Each picture in the picture classification data set corresponds to a label target, and the picture and the corresponding label target are used as text inputs; the data set is divided into a training set and a test set; Step 2, construct a knowledge distillation model framework, which includes a teacher model, a student model, a text encoder for processing text inputs, and a feature projector for aligning image and text features; Step 3: Freeze the parameters of the text encoder in the knowledge distillation model and the teacher model. Tokenize and vectorize the text information in the training set to obtain learnable prompt vectors, and then perform contrastive learning pre-training on the learnable prompt vectors. Step 4: Freeze the pre-trained learnable vectors in Step 3. At the same time, keep the text encoder in the knowledge distillation model and the teacher model frozen, and perform knowledge distillation training on the student model. Calculate the model loss through KL divergence. Step 5: Combine the model losses calculated in Step 4 to obtain the total distillation loss. 。
[0006] The features of the present invention also lie in that The image classification dataset in Step 1 is the CIFAR-100 dataset, and the training set and the test set are divided according to 5:1.
[0007] The batch size of the knowledge distillation model in Step 2 is set to 64, and the size of the text-visual feature space is 256.
[0008] The teacher model Teacher and the student model Student in Step 2 are constructed using the CNN model. The teacher model Teacher and the student model Student use two combination methods: isomorphic and heterogeneous. The isomorphic model means that the teacher model Teacher and the student model Student use the same architecture but different numbers of layers of the CNN network; the heterogeneous model means that the teacher model Teacher and the student model Student use different architectures and different numbers of layers of the CNN network.
[0009] The text encoder is constructed using the multi-modal machine learning model CLIP. The CLIP model is composed of a self-attention module and a feed-forward neural network module FFN; the input of the text encoder is text tokens. The feature projector uses the CNN model and is composed of a convolutional layer and a BN layer.
[0010] The teacher model Teacher in Step 2 includes the residual network ResNet32×4, ResNet56, and ResNet56. This network has a total of 3 stages, and the number of channels in each stage is 16, 32, and 64 respectively. ResNet110 and ResNet50; The ResNet50 network has a total of 4 stages, and the number of channels in each stage is 64, 128, 256, and 512 respectively, and is composed of a bottleneck structure. The student model Student includes ResNet8×4, ResNet32, the lightweight network ShuffleNetV1 composed of grouped convolution and channel shuffle Shuffle, ShuffleNetV2 that reduces redundant convolution operations and directly performs channel splitting and merging, and the lightweight network MobileNetV2 that uses depthwise separable convolution and inverted residual blocks.
[0011] Step 3 is specifically as follows: Input the images in the training set into the teacher model Teacher, and output image features before the fully connected layer ; For the image the label target is constructed as a text token with the text input; the part of the token except the label target is converted into a learnable vector format, and then the token is input into the text encoder to obtain text features ; Project the image features into the image-text feature space through a feature projector, align the image features and text features in this feature space, obtain the metric distance, calculate the contrastive learning loss, and learn the learnable vector part in the token. After 20 rounds of learning on the training set data, the learnable vector pre-training process is completed.
[0012] Step 4 is specifically as follows: Freeze the pre-trained learnable vectors in Step 3. At the same time, keep the text encoder in the knowledge distillation model and the teacher model frozen. Conduct knowledge distillation training on the student model, and calculate the model loss through KL divergence: Input the images in the training set into both the teacher model Teacher and the student model student at the same time, and output image features and before the fully connected layer respectively, and output probability distributions and after the fully connected layer respectively; Input the label target into the text encoder in the same way as in Step 3 to obtain text features ; Align and with respectively in the feature space to obtain the corresponding metric distances and respectively. Calculate the KL divergence for and to obtain the text-visual loss , and calculate the KL divergence for the probability distributions and Calculate the KL divergence to obtain the KD loss , calculate the KL divergence between and the label target to obtain the Target loss .
[0013] Measure the distance and :
[0014]
[0015] According to the KL divergence calculation formula:
[0016] It can be obtained that , and :
[0017]
[0018]
[0019] Step 5, combine the losses calculated in Step 4 to obtain the total distillation loss , ; wherein: , and are parameters, .
[0020] The beneficial effects of the present invention are as follows: The present invention adds text prompts to assist the teacher model to more accurately output these key information points, reduces the uncertainty in the training process of the student model, thereby making the optimization process more stable and reducing the "detour" phenomenon (i.e., the unnecessary exploration of the model when searching for effective key information points). By this method, the output probability distribution of the teacher model can be reconstructed, the knowledge quality can be improved, the knowledge type can be perfected, the learning of the student model can be more accurate, the model compression performance can be optimized, and the generalization ability of the student model can be enhanced.
[0021] The present invention realizes a method of introducing prompt learning into distillation training through a pre-trainable and optimizable prompt vector. Text prompts largely depend on the quality of the prompt words. To better adapt to different distillation tasks, by introducing an optimizable prompt vector, a certain degree of prompt word fine-tuning is performed before distillation training, making the current prompt words more suitable for the teacher model and helping the student model converge to useful key information points faster. At the same time, the output probability distribution of the teacher model encodes task-related knowledge (i.e., key information points), and the student model retains these key information by imitating the output of the teacher model. The student model will find a balance point during training, that is, find a balance between retaining task-related key information in the input data and ignoring irrelevant information. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the model training flowchart in the knowledge distillation method of the image classification model based on text prompts of the present invention; Figure 2 is the application diagram of the knowledge distillation method of the image classification model based on text prompts of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] Embodiment 1 The knowledge distillation method of the visual model based on text prompts of the present invention uses the CIFAR-100 dataset, which is divided into a training set and a test set for model training and model testing respectively. At the same time, the target labels in the dataset are extracted as the text inputs of the text model; the distillation model includes two parts: a visual distillation model and a text encoding model; the specific implementation is divided into a training process and a test process; specifically includes the following steps: Step 1, obtain a picture classification dataset, each picture in the picture classification dataset corresponds to a label target, the picture and the corresponding label target are used as text inputs; the dataset is divided into a training set and a test set; Step 2, construct a knowledge distillation model framework, the knowledge distillation model framework includes a pre-trained teacher model Teacher, an untrained student model Student, a text encoder Text-Encoder for processing text inputs, and a feature projector projector for aligning image and text features; Step 3, freeze the parameters of the text encoder Text-Encoder and the teacher model in the knowledge distillation model, tokenize and vectorize the text information in the training set to obtain a learnable prompt vector, and then perform contrastive learning pre-training on the learnable prompt vector; Step 4: Freeze the learnable vectors pre-trained in Step 3. Meanwhile, keep the text encoder (Text-Encoder) of the knowledge distillation model and the teacher model frozen, and perform knowledge distillation training on the student model. Step 5: Calculate the total distillation loss based on the training results of Step 4. 。
[0025] Example 2 Based on Example 1, The teacher model (Teacher) includes ResNet32×4 (based on ResNet32 and expanding the number of channels by 4 times), ResNet56 (including 56 layers of residual network, with a total of 3 stages, and the number of channels in each stage is 16, 32, and 64 respectively), ResNet110 (including 110 layers of residual network), and ResNet50 (including 50 layers of residual network, with a total of 4 stages, and the number of channels in each stage is 64, 128, 256, and 512 respectively, and is composed of bottleneck structures).
[0026] The student model (Student) includes ResNet8×4 (based on ResNe8 with 8 layers and expanding the number of channels by 4 times), ResNet32 (basic ResNet structure, 32 layers), ShuffleNetV1 (a lightweight network composed of grouped convolution and channel shuffle), ShuffleNetV2 (an optimized version of ShuffleNetV1 that reduces redundant convolution operations and directly splits and merges channels), and MobileNetV2 (a lightweight network using depthwise separable convolution and inverted residual blocks).
[0027] Example 3 As Figure 2 shown, the application of the visual model knowledge distillation method based on text prompts of the present invention specifically includes the following steps: Step 1: Input the images in the test set into the distilled-trained student model (student), and output the probability distribution after the fully connected layer ; Step 2: Use the probability distribution of each category in the CIFAR-100 dataset in Step 1 as the classification result. Example 4 As Figure 1 shown, the visual model knowledge distillation method based on text prompts of the present invention specifically includes the following steps: Step 1: Obtain an image classification dataset. Each image in the image classification dataset corresponds to a label target. The image and the corresponding label target are used as text inputs. The dataset is divided into a training set and a test set. Commonly used image classification datasets include: CIFAR-10, CIFAR-100, ImageNet, Tiny ImageNet, and Caltech-101 / 256, etc. Among them, CIFAR-100 (Canadian Institute for Advanced Research 100 classes) is a classic image classification dataset widely used in the research of computer vision and deep learning. Its main features are: the dataset consists of 60,000 32x32 color pictures, containing 100 categories, 600 pictures in each category, with low image resolution and blurred details, which requires higher feature extraction ability for the model; there are many categories and some categories have high similarity (such as different species of fish or insects), requiring stronger fine-grained classification ability. Therefore, CIFAR-100 is a compromise choice between CIFAR-10 (simple classification) and ImageNet (complex classification), suitable for studying fine-grained classification problems under limited research resources. The training set and the test set are divided in a ratio of 5:1 to obtain 500 training set pictures, 100 test set pictures and their corresponding label targets. At the same time, the label target corresponding to the picture is used as the text input. At the same time, set the batchsize of the distillation model to 64 and the size of the text-visual feature space to 256. Step 2: Build a knowledge distillation model framework. The entire framework includes a pre-trained teacher model Teacher, an untrained student model Student, a text encoder Text-Encoder for processing text inputs, and a feature projector projector for aligning image and text features. The teacher model Teacher and the student model Student are constructed using a CNN model. At the same time, the teacher model Teacher and the student model Student use two combination methods: isomorphic and heterogeneous. The isomorphic model means that the teacher model Teacher and the student model Student use the same architecture but CNN networks with different numbers of layers; the heterogeneous model means that the teacher model Teacher and the student model Student use different architectures and CNN networks with different numbers of layers. The teacher model Teacher and the student model Student are divided into 6 groups of teacher-student pairs of isomorphic and heterogeneous types according to the combination method in Table 1. Table 1 Isomorphic and Heterogeneous Teacher-Student Models
[0028] The text encoder Text-Encoder is constructed using the CLIP model. The CLIP model consists of a Self-Attention module and an FFN. The input to the encoder is text tokens, and its size is (blank masked except for the prompt); The feature projector projector uses a CNN model, which consists of a convolutional layer and a BN layer. The input size is the image feature size, and the output size is ; Step 3, freeze the parameters of the text encoder in the knowledge distillation model and the teacher model. Tokenize and vectorize the text information in the training set to obtain a learnable prompt vector, and then perform contrastive learning pre-training on the learnable prompt vector. Take the teacher model as ResNet32×4 as an example; In this process, the images in the training set are input into the teacher model Teacher, and image features are output before the fully connected layer , and the feature size is ; For the image , the label target is constructed as a text token with the format <a photo of a "target”>, such as Figure 1 as shown. The example text input is Convert the part of the token other than the "suffering man" into a learnable vector format. Then, input the token into the Text-Encoder to obtain text features , with a feature size of ; project the image features into the image-text feature space through a feature projector, and spatially align the image features and text features in this feature space to obtain a metric distance, calculate the contrastive learning loss, and learn the learnable vector part of the token. After 20 rounds of learning on the training set data, the learnable vector pre-training process is completed; Step 4: Freeze the pre-trained learnable vectors in Step 3. At the same time, keep the Text-Encoder in the distillation model and the teacher model frozen. Conduct a knowledge distillation training process on the student model, with the teacher model using ResNet32×4 and the student model using ResNet8×4; Step 5: Combine the losses calculated in Step 4 to obtain the total distillation loss .
[0029] Example 5 Based on Example 4, Step 4 is specifically as follows: Input the images in the training set into both the teacher model Teacher and the student model student at the same time, and output image features and before the fully connected layer, with both feature sizes being ; output probability distributions and after the fully connected layer, with both probability distribution sizes being ; Input the target into the Text-Encoder in the same way as in Step 3 to obtain text features , with a feature size of ; Align and with in the feature space respectively to obtain the corresponding metric distances and , calculate the KL divergence for and to obtain the text-visual loss , calculate the KL divergence for the probability distributions and to obtain the KD loss , calculate the KL divergence for and the target to obtain the Target loss ;
[0030]
[0031]
[0032] Total distillation loss The calculation method is as follows: ; Wherein: ; The student model is trained. After 240 rounds of learning on the training set data, the distillation training process is completed.
[0033] Example 6 Such as Figure 2 As shown, the method for knowledge distillation of a vision model based on text prompts of the present invention specifically includes the following steps: Step 1, take the images in the test set and input them into the distilled-trained student model student (ResNet8×4), and the image regarding the probability distribution of each category in the CIFAR-100 dataset is output after the fully connected layer ; Step 2, use the category represented by the item with the largest probability value in the probability distribution in Step 1 as the classification result; The classification result is one of the categories in the CIFAR-100 dataset, namely apple, aquarium fish, baby, bear, beaver, bed, bee, beetle, bicycle, bottle, bowl, boy, bridge, bus, butterfly, camel, can, castle, caterpillar, cattle, chair, chimpanzee, clock, cloud, cockroach, computer keyboard, couch, crab, crocodile, cup, dinosaur, dolphin, elephant, flatfish, forest, fox, girl, giraffe, goat, goldfish, golf ball, gorilla, grasshopper, gray fox, hamster, harp, hare, helicopter, hippopotamus, honeybee, horse, hot airballoon, hot dog, house, kangaroo, keyboard, lamp, lawn mower, leopard, lion, lizard, lobster, man, maple tree, motorcycle, mountain, mouse, mushroom, oak tree, orange, orangutan, ostrich, otter, palm tree, pear, pickup truck, pine tree, plain, plate, porcupine, possum, rabbit, raccoon, ray, road, rocket, rose, sea, seal, shark, sheep, skunk, skyscraper, snail, snake, spider, squirrel, streetcar, sunflower, sweetpepper, table, tank, telephone, television, tiger, tractor, train, trout, tulip, turtle, wardrobe, whale, willow tree, wolf, woman, worm, a total of 100 classes. The category with the highest output probability in the probability distribution is used as the classification result.
[0034] The probability distribution The category with the highest probability is used as the predicted classification result of this test and compared with the true label target, and the accuracy acc-Top1 is calculated as the test standard; The accuracy of the test results of the present invention on the CIFAR-100 dataset is compared with the accuracy of the test results of the baseline method (KD) on the CIFAR-100 dataset, that is, the accuracy acc-Top1 (%) of the present invention and the baseline method (KD) in the heterogeneous and homogeneous model teacher-student combinations, as shown in Table 2: Table 2 Experimental results on the CIFAR-100 dataset
[0035] It can be seen from the above table that the performance of the present invention on the CIFAR-100 dataset has been significantly improved compared with the baseline method, and the introduction of text prompts has made a greater improvement in the performance of heterogeneous model distillation.
[0036] The present invention includes but is not limited to the above embodiments. The above examples and descriptions are for better understanding and use of the invention. At the same time, those skilled in the art and research scholars familiar with this field can make modifications to the above embodiments according to their own understandings to improve the effect and reduce the cost. However, the modifications and improvements made without departing from the scope of the present invention are within the protection scope of the present invention.
Claims
1. A knowledge distillation method for image classification model based on textual cues, characterized in that: The specific steps are as follows: Step 1: Get an image classification dataset. Each image in the image classification dataset corresponds to a label target. The image and the corresponding label target are used as text input. Divide the dataset into a training set and a test set. Step 2: Build a knowledge distillation model framework, which includes a teacher model, a student model, a text encoder for processing text input, and a feature projector for aligning image and text features. Step 3: Freeze the parameters of the text encoder and the teacher model in the knowledge distillation model, tokenize the text information in the training set and then vectorize it to obtain a learnable prompt vector, and then perform comparative learning pre-training on the learnable prompt vector; Step 4: Freeze the pre-trained learnable vector in step 3. Meanwhile, keep the text encoder and teacher model in the knowledge distillation model frozen. Perform knowledge distillation training on the student model and calculate the model loss through KL divergence. The total distillation loss is obtained. .
2. The method for knowledge distillation of image classification model based on textual prompts according to claim 1, characterized in that: The image classification dataset described in step 1 is the CIFAR-100 dataset, and the training set and test set are divided into 5:
1.
3. The method for knowledge distillation of image classification model based on textual prompts according to claim 2, characterized in that: In step 2, the batch size of the knowledge distillation model is set to 64, and the size of the text-visual feature space is 256.
4. The method for knowledge distillation of image classification model based on textual prompts according to claim 2, characterized in that: The teacher model and student model described in step 2 are constructed using the CNN model. The teacher model Teacher and the student model Student use two combinations, homogeneous and heterogeneous. The homogeneous model is that the teacher model Teacher and the student model Student use CNN networks with the same architecture and different numbers of layers; the heterogeneous model is that the teacher model Teacher and the student model Student use CNN networks with different architectures and different numbers of layers.
5. The method for knowledge distillation of image classification model based on textual prompts according to claim 2, characterized in that: The text encoder is constructed using a multimodal machine learning model CLIP, which consists of a self-attention module and a feed-forward neural network module FFN; the text encoder input is a text token; The feature projector adopts a CNN model, which is composed of a convolutional layer and a BN layer.
6. The method for knowledge distillation of image classification model based on textual prompts according to claim 4, characterized in that: The teacher model Teacher in step 2 includes the residual network ResNet32×4, ResNet56, ResNet56, which has 3 stages, with the number of channels in each stage being 16, 32, and 64, respectively, ResNet110, and ResNet50; the ResNet50 network has 4 stages, with the number of channels in each stage being 64, 128, 256, and 512, respectively, and adopts a bottleneck structure; The student model Student includes ResNet8×4, ResNet32, a lightweight network ShuffleNetV1 composed of grouped convolution and channel shuffle, ShuffleNetV2 that reduces redundant convolution operations to directly split and merge channels, and a lightweight network MobileNetV2 that uses depth-separable convolution and inverted residual blocks.
7. The method for knowledge distillation of image classification model based on textual prompts according to claim 5, characterized in that: Step 3 is as follows: The images in the training set Input the teacher model Teacher and output image features before the fully connected layer ; for images The label target is used as text input to construct a text token; the part of the text token except the label target is converted into a learnable vector format, and then the text token is input into the text encoder to obtain the text feature ; The image features are projected into the image-text feature space through a feature projector, and the image features and text features are spatially aligned in the feature space to obtain the metric distance; the learnable vector part in the token is learned, and the learnable vector pre-training process is completed after 20 rounds of learning on the training set data.
8. The method for knowledge distillation of image classification model based on textual prompts according to claim 7, characterized in that: Step 4 is as follows: Freeze the pre-trained learnable vector in step 3, and keep the text encoder and teacher model in the distillation model frozen, and perform knowledge distillation training; The images in the training set Input the teacher model Teacher and the student model student at the same time, and output image features before the fully connected layer and , and output probability distributions after the fully connected layer and ; Input the label target into the text encoder in the same way as in step 3 to obtain the text feature ;Will and Respectively Perform alignment in feature space and obtain the corresponding metric distances and ,right and Calculate KL divergence and get text-visual loss , for the probability distribution and Calculate KL divergence and get KD loss ,right Calculate the KL divergence with the label target and get the Target loss .
9. The method for knowledge distillation of image classification model based on textual prompts according to claim 8, characterized in that: Combine the losses calculated in step 4 to get the total distillation loss , ; in: , and As parameters, .