Image classification method based on information compensation and knowledge distillation binary neural network
By introducing information compensation and knowledge distillation mechanisms into binary neural networks, parallel teacher and student models are built and iterative training is carried out, the problems of low classification accuracy and information loss in image classification are solved, and higher classification accuracy and more efficient computing resource use are achieved.
Patent Information
- Application Number
- CN202510145851.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The existing binary neural network has low classification accuracy in image classification tasks and has serious information loss during forward propagation, resulting in poor model performance.
Using a binary neural network image classification method based on information compensation and knowledge distillation, a parallel pre-trained teacher model and a binary neural network student model based on information compensation is constructed, combined with a distribution loss calculation module, iterative training is carried out to improve the classification performance of the student model.
It significantly improves the image classification accuracy of binary neural networks, reduces computing resource consumption, and improves the training efficiency of the model.
Smart Images

Figure CN120070987A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and relates to an image classification method, specifically an image classification method based on an information compensation and knowledge distillation binary neural network, which can be used in embedded systems with limited computing resources, edge computing devices, and other low-power application scenarios. Background Art
[0002] Convolutional neural network is one of the core technologies in the field of deep learning and is widely used in fields such as computer vision and natural language processing. Its basic principle is to automatically extract the spatial local features of the input data through the convolutional layer and reduce the size of the feature map through the pooling layer, thereby improving the computing efficiency and enhancing the generalization ability of the model. And the multi-layer structure enables it to gradually learn feature representations from low-level to high-level. However, convolutional neural networks usually have high redundancy, and many parameters and computations may have no effect on the final result. Therefore, using model compression techniques such as quantization and pruning to process convolutional neural networks can effectively reduce their memory occupancy, computational volume, and power consumption, making them more suitable for deployment on resource-constrained edge devices. Binary neural network is a type of convolutional neural network that quantizes weights and activations into the set {+1, -1}. Compared with traditional convolutional neural networks, binary neural networks significantly reduce the storage requirements and computational volume of the model, which gives them advantages on edge devices such as embedded systems, mobile devices, and Internet of Things devices. However, the classification performance of the binary convolutional neural network will drop sharply after binarization.
[0003] The classification performance of the binary neural network can be improved by methods such as scaling factors and quantization-aware training. For example, the patent application with the application publication number CN116681925A and the name "A Vehicle Classification Method Based on Self-Distillation Binary Neural Network" discloses a vehicle classification method based on a self-distillation binary neural network. The implementation steps of this invention are to obtain a vehicle picture dataset containing N similar labels, and obtain a training set and a test set according to the division ratio of (4 + N):1; input the training set pictures into the dynamically approximated gradient binary neural network built to obtain the class output prediction; the network trained in this round of iteration becomes the teacher, and the class output prediction is screened and averaged to obtain a soft label library indicating correctness; the network trained in the next round of iteration becomes the student, and the soft label library provides soft label teacher knowledge for self-distillation at the tail of the binary neural network; continuously iterate the teacher-student self-distillation training process, and the number of iteration rounds for distillation is 2 to SumE to improve the classification accuracy of the binary neural network; classify and predict the vehicle test set pictures according to the trained binary classification model to obtain the classification result. This invention can effectively improve the image classification accuracy of the binary neural network. However, due to activation binarization, the amount of information input into the model drops sharply during the forward propagation process, resulting in a sharp reduction in the amount of information obtained by deeper networks, and thus the classification accuracy of the model is still relatively low. Summary of the Invention
[0004] The object of the present invention is to overcome the defects existing in the above-mentioned prior art, and a method for image classification based on an information compensation and knowledge distillation binary neural network is proposed to solve the technical problem of the low classification accuracy of the binary neural network model in the prior art.
[0005] To achieve the above object, the technical solution adopted by the present invention includes the following steps:
[0006] (1) Obtain a training sample set and a validation sample set:
[0007] Obtain a total of K RGB images including N target categories, each containing M images, and label the targets in each image. Then, form a training sample set with more than half of the images and their labels in each target category, and form a validation sample set with the remaining images, where N≥2 and K≥2000;
[0008] (2) Iteratively train the teacher model:
[0009] Use the training sample set as the input of the teacher model to perform iterative training on it to obtain a pre-trained teacher model;
[0010] (3) Construct an image classification model S based on an information compensation and knowledge distillation binary neural network:
[0011] Construct an image classification model S of a knowledge distillation module and a cascaded distribution loss calculation module; among them, the knowledge distillation module includes a pre-trained teacher model and a binary neural network student model based on information compensation arranged in parallel, which are respectively used to obtain the prediction values of the teacher model and the binary neural network student model for the same training sample; the output end of the distribution loss calculation module is connected to the input of the student model, and its output result is used to update the weight parameters of the student model during the backpropagation process;
[0012] (4) Define the loss function Loss of the image classification model S:
[0013]
[0014] Among them, Σ represents the summation operation, c represents the number of training sample sets input to the model in a single iteration, x represents the training sample input to the model, p s (x), p t (x) respectively represent the prediction results of the student model and the teacher model for x;
[0015] (5) Iteratively train the image classification model S:
[0016] Use the training sample set as the input of the image classification model S and perform iterative training on it to obtain the trained image classification model S * ;
[0017] (6) Obtain the image classification result:
[0018] Use the test sample set as the input of the student model in the image classification model S * Perform forward inference to obtain the classification result of each test sample.
[0019] Compared with the prior art, the present invention has the following advantages: Specific description
[0020] 1. During the iterative training of the image classification model S, the information compensation module uses the output of the convolutional layer as supplementary feature information to compensate the output feature map of the binary depthwise separable convolutional module, significantly improving the information content of the output feature map, effectively solving the problem of serious information loss in the forward propagation process of the binary neural network, and improving the classification accuracy.
[0021] 2. The present invention introduces a knowledge distillation mechanism and designs a distribution loss function. During the training process, the pre-trained teacher model is used to guide the student model, making the output probability distribution of the student model closer to the output distribution of the pre-trained teacher model, effectively improving the classification performance of the student model, and at the same time giving the model greater flexibility, without strictly requiring the teacher model and the student model to be consistent in architecture. Compared with the prior art, the present invention not only shows better performance in classification accuracy, but also significantly reduces the consumption of computing resources and improves the training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart for the implementation of the present invention;
[0023] Figure 2 is a schematic structural diagram of the image classification model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] Referring to Figure 1 , the present invention includes the following steps:
[0026] Step 1) Obtain the training sample set and the test sample set:
[0027] Obtain a total of K RGB images including N target categories, with each category containing M images, and label the targets in each image. Then, form a training sample set with more than half of the images and their labels in each target category, and form a validation sample set with the remaining images, where N≥2 and K≥2000. In this example, the Mini-ImageNet dataset with 100 target categories is used, which contains a total of 60,000 RGB images. To improve the generalization ability of the model, data augmentation is performed on each image, specifically including random cropping and random horizontal flipping operations. Finally, all images are uniformly scaled to a size of 224×224 pixels to meet the requirements of the model input size.
[0028] Step 2) Iteratively train the teacher model:
[0029] Use the training sample set as the input of the teacher model to perform iterative training, and obtain a pre-trained teacher model. The pre-trained teacher model is fully trained and has high performance and generalization ability. It can capture more complex patterns and features in the data, thus providing more accurate knowledge guidance for the binarized student model. In this example, the teacher model uses the lightweight convolutional neural network MobileNetV1, whose structure includes stacked convolutional layers, depthwise separable convolution blocks, average pooling layers, fully connected layers, and softmax activation functions. During the training process, the cross-entropy loss function is used to calculate the loss of the teacher model, and the stochastic gradient descent method is adopted to update the model weight parameters through the calculated loss by backpropagation. The entire training process is iterated 90 times to ensure that the model converges fully and reaches the optimal performance.
[0030] Step 3) Construct the image classification model S:
[0031] Construct an image classification model S with a knowledge distillation module and a cascaded distribution loss calculation module, whose structure is as Figure 2As shown in the figure; among them, the knowledge distillation module includes a pre-trained teacher model and a binary neural network student model based on information compensation arranged in parallel; the student model includes stacked convolutional layers, a dynamic binary activation module, a binary depthwise separable convolution module, an information compensation module, an average pooling layer, a fully connected layer, and a softmax activation function; a slight distribution shift in the real-valued feature map output by the convolutional layer will result in completely different binary activation outputs, which will directly affect the information content of the features and ultimately affect the final classification accuracy. The dynamic binary activation module effectively alleviates the influence of the distribution shift on the binary activation output through an adaptive mechanism, improving the robustness of the binary activation; aiming at the information loss problem existing in the activation and weight binaryization processes of the student model, the information compensation module compensates for the information loss in the binaryization process by fusing the output feature map of the convolutional layer with the output feature map of the binary depthwise separable convolution module, thereby enhancing the feature expression ability of the model output layer; the distribution loss calculation module connects the softmax outputs of the pre-trained teacher model and the student model, and its calculation result is used to update the weight parameters of the student model during the backpropagation process;
[0032] Step 4) Define the loss function Loss of the image classification model S:
[0033]
[0034] The loss function Loss calculates the KL divergence between the output probability distributions p s (x) and p t (x) of the student model and the pre-trained teacher model, which is used to measure the difference between the two probability distributions, and optimizes the classification performance of the student model by minimizing this difference.
[0035] Step 5) Perform iterative training on the image classification model S:
[0036] (5a) Initialize the iteration number as i, the maximum iteration number as I, I≥90, and the weight parameter of the image classification model S at the i-th iteration is W i , and let i = 1;
[0037] (5b) Refer to the structural schematic diagram of the image classification model Figure 2 . The convolutional layer in the pre-trained teacher model extracts the local features of the training sample x; the depthwise separable convolution block further extracts the spatial features from the local feature map; the pooling layer downsamples the feature map output by the depthwise separable convolution block to achieve feature map dimensionality reduction, thereby reducing the computational amount and preventing overfitting; the fully connected layer flattens the downsampled feature map into a one-dimensional vector and processes it through the softmax activation function to obtain the predicted value p t (x) of x on N categories;
[0038] The convolutional layer in the student model extracts local features of the training sample x; the dynamic binary activation module performs dynamic binarization on the local feature map, and the dynamic binary activation module is implemented through an approximation function containing an adaptive bias factor:
[0039]
[0040] where a represents the adaptive bias factor;
[0041] The binary depthwise separable convolution block further extracts spatial features from the binarized feature map, significantly reducing the computational complexity and the size of weight parameters while retaining the feature extraction ability; the information compensation module compensates the feature map after extracting two-dimensional features. The specific implementation process is to use the output feature map of the convolutional layer containing rich feature information as a supplement and fuse it with the feature map output by the binary depthwise separable convolution module. The compensation process can be expressed as:
[0042] F = ReLU(BN(conv1×1(torch.cat(F 1 ,F 2 ))))
[0043] where F 1 represents the feature map output by the convolutional layer, F 2 represents the feature map output by the binary depthwise separable convolution module, torch.cat represents the process of concatenating feature maps, conv1×1, BN, and ReLU represent 1×1 convolution, batch normalization, and activation function respectively, and F represents the output of the information compensation module.
[0044] The average pooling layer reduces the dimension of the compensated feature map; the fully connected layer flattens the dimension-reduced feature map into a one-dimensional vector and processes it through the softmax activation function to obtain the predicted value p s (x) of x on N categories. Finally, the distribution loss calculation module calculates the loss value Loss of the image classification model S through p t (x) and p s (x);
[0045] (5c) Adopt the stochastic gradient descent method to update the weight W of the image classification model S through the loss value Loss to obtain the image classification model S of this iteration i , and the update formula of the weight W is:
[0046]
[0047] where W i+1 represents the updated weight parameters of the student model, W i represents the parameters of the student model of this iteration, and η represents the learning rate. Indicates the partial derivative of the loss function Loss with respect to W i Find the partial derivative;
[0048] (5d) Determine whether i = I holds. If so, obtain the trained image classification model S * , otherwise, set i = i + 1, S = S i , and execute step (5b);
[0049] Step 6) Obtain the image classification result:
[0050] Use the test sample set as the input of the student model in the image classification model S * Perform forward inference to obtain the classification result of each test sample.
[0051] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0052] 1. Simulation experiment conditions:
[0053] The hardware platform for the simulation experiment of the present invention is: The CPU model is Intel Core i7-10700, the frequency is 2.9GHz, the GPU model is NVIDIA Tesla P100, and the video memory is 12GB;
[0054] The software platform for the simulation experiment of the present invention is: Windows 11, Python 3.8, PyTorch 2.1, CUDA12.1, Brevitas 0.11.0;
[0055] The sample set used in the simulation experiment of the present invention is: Mini-ImageNet;
[0056] Conduct a comparative simulation on the classification accuracy Top-1, Top-5, and the amount of operations OPs of the present invention and the prior art. Among them, the Top-1 accuracy rate refers to the proportion of samples in which the most likely category predicted by the model is consistent with the actual label; the Top-5 accuracy rate refers to the proportion of samples in which the actual label is included in the top five categories with the highest predicted probability by the model; the amount of operations refers to the number of floating-point operations required for the model to perform a forward propagation process.
[0057] 2. Simulation content and its result analysis:
[0058] Referring to Table 1, the prior art and the present invention are simulated on three indicators of classification accuracy Top-1, Top-5, and the amount of operations OPs on the Mini-ImageNet dataset, and the results are shown in Table 1:
[0059] Table 1
[0060] Top-1(%) Top-5(%) <![CDATA[OPs(10 8 )]]> Prior art 68.1 85.5 5.69 The present invention 69.2 88.6 0.87
[0061] Comparing with the experimental results, the present invention shows better performance on the Mini-ImageNet dataset, but the amount of computation is only 15.2% of the prior art. The above parameters prove the feasibility of the present invention.
Claims
1. An image classification method based on information compensation and knowledge distillation binary neural network, characterized in that: The following steps are involved: (1) Obtain training sample set and verification sample set: Obtain a total of K RGB images including N target categories and each containing M images, and label the targets in each image. Then, more than half of the images and their labels of each target category form a training sample set, and the remaining images form a verification sample set, where N ≥ 2 and K ≥ 2000. (2) Iteratively train the teacher model: The training sample set is used as the input of the teacher model to iteratively train it to obtain a pre-trained teacher model; (3) Construct an image classification model S based on information compensation and knowledge distillation binary neural network: Construct an image classification model S of a knowledge distillation module and a distribution loss calculation module cascaded therewith; wherein the knowledge distillation module includes a pre-trained teacher model and a binary neural network student model based on information compensation arranged in parallel, which are respectively used to obtain the prediction values of the teacher model and the binary neural network student model for the same training sample; the output end of the distribution loss calculation module is connected to the input of the student model, and its output result is used to update the weight parameters of the student model during the back propagation process; (4) Define the loss function Loss of the image classification model S: Among them, Σ represents the summation operation, c represents the number of training sample sets input to the model in a single iteration, x represents the training sample of the input model, and p s (x), p t (x) represents the prediction results of x by the student model and the teacher model respectively; (5) Iteratively train the image classification model S: The training sample set is used as the input of the image classification model S and iteratively trained to obtain the trained image classification model S * ; (6) Obtain image classification results: The test sample set is used as the image classification model S * The input of the student model in is used for forward reasoning to obtain the classification result of each test sample.
2. The method according to claim 1, characterized in that: The teacher model described in step (2) adopts a lightweight convolutional neural network MobileNetV1, which includes stacked convolutional layers, depth-wise separable convolutional blocks, average pooling layers, fully connected layers, and softmax activation functions.
3. The method according to claim 2, characterized in that The binary neural network student model based on information compensation described in step (3) includes stacked convolutional layers, dynamic binary activation modules, binary depth-separable convolutional modules, information compensation modules, average pooling layers, fully connected layers and softmax activation functions, and the output end of the dynamic binary activation module is also connected to the input end of the information compensation module; the dynamic binary activation module is implemented by an adaptive approximate binary function; the information compensation module includes stacked splicing layers, 1×1 convolutional layers, batch normalization layers and activation layers.
4. The method according to claim 3, characterized in that The iterative training of the image classification model S described in step (5) is implemented as follows: (5a) The number of initialization iterations is i, the maximum number of iterations is I, I ≥ 90, and the weight parameter of the image classification model S in the i-th iteration is W i , and let i = 1; (5b) The teacher model and the student model use the softmax method to obtain the predicted value p of the training sample x on N categories respectively. t (x), p s (x), and through p t (x) and p s (x) Calculate the loss value Loss of the image classification model S; (5c) The stochastic gradient descent method is used to update the weight W of the image classification model S through the loss value Loss to obtain the image classification model S of this iteration. i ; (5d) Determine whether i=I. If so, obtain the trained image classification model S. * Otherwise, let i = i + 1, S = S i , and execute step (5b).
5. The method according to claim 4, characterized in that The predicted value p of the training sample x in step (5b) on N categories t (x), the acquisition method is: The convolution layer extracts the local features of the training sample x; the depth-separable convolution block further extracts spatial features from the local feature map; the pooling layer reduces the dimension of the feature map after further extracting spatial features; the fully connected layer flattens the feature map after dimensionality reduction into a one-dimensional vector and processes it through the softmax activation function to obtain the predicted value p of x on N categories. t (x).
6. The method according to claim 4, characterized in that The predicted value p of the training sample x in step (5b) on N categories s (x), the acquisition method is: The convolutional layer extracts local features of the training sample x; The dynamic binary activation module dynamically binarizes the local feature map through the approximate function Approx(x) containing an adaptive bias factor; the binary depth-separable convolution block further extracts two-dimensional features from the binary feature map; The information compensation module uses the feature map output by the convolution layer as a supplement to perform feature information compensation on the feature map after extracting the two-dimensional features; The average pooling layer reduces the dimension of the compensated feature map; the fully connected layer flattens the reduced feature map into a one-dimensional vector and processes it through the softmax activation function to obtain the predicted value p of x in N categories. s (x).
7. The method according to claim 6, characterized in that The approximate function Approx(x) is expressed as: Where a represents the adaptive bias factor.
8. The method according to claim 6, characterized in that The information compensation module performs information compensation on the feature map after extracting the two-dimensional features, and the implementation process is as follows: F=ReLU(BN(conv1×1(torch.cat(F1,F2)))) Among them, F1 represents the feature map output by the convolution layer, F2 represents the feature map output by the binary depth separable layer, torch.cat represents the concatenation of the feature map, conv1×1, BN, ReLU represent 1×1 convolution, batch normalization, and activation function respectively, and F represents the output of the information compensation module.
9. The method according to claim 8, characterized in that The weight W of the image classification model S described in step (5c) is updated, and the update formula is: Among them, W i+1 represents the updated student model weight parameter, W i represents the student model parameter of this iteration, η represents the learning rate, Represents the loss function Loss with respect to W i Find the partial derivative.
Citation Information
Patent Citations
Underground locomotive pedestrian and distance detection method based on binarization network
CN110837775A
Knowledge distillation method and system based on structural feature knowledge
CN115496213A
Vehicle classification method based on self-distillation binary neural network
CN116681925A