Image classification method and device, equipment and storage medium

By performing backpropagation training on the initial student network model, combining the feature extraction results of the first teacher network model and the initial student network model, the problem of insufficient accuracy of the image classification model in resource-constrained environments is solved, and high-accuracy image classification is achieved.

CN120014324APending Publication Date: 2025-05-16SHENZHEN ZTE NETVIEW TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510023449.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, when facing a resource-constrained environment, the obtained image classification model is insufficient, resulting in inaccurate image categories of images to be classified.

Method used

An image classification method is proposed, and the initial student network model is backpropagated by backpropagation training based on the first distillation loss value obtained by the first type of activation map and the second type of activation map to obtain a preset image classification model. The first type of activation map performs feature extraction of sample images through the first teacher network model, while the second type of activation map performs feature extraction of sample images through the initial student network model.

Benefits of technology

In an environment where resources are limited, a high-accuracy structured simple preset image classification model is obtained through knowledge distillation technology to achieve accurate classification of images to be classified and ensure the accuracy of image categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014324A_ABST
    Figure CN120014324A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image classification, and provides an image classification method, device and equipment and a storage medium, and the method comprises the steps: obtaining a to-be-classified image, and inputting the to-be-classified image into a preset image classification model for classification, the preset image classification model is a model obtained by performing back propagation training on an initial student network model through a first distillation loss value obtained based on a first-class activation graph and a second-class activation graph, and the first-class activation graph is obtained by performing feature extraction on a sample image through a first teacher network model; the second-class activation graph is obtained by performing feature extraction on the sample image through the initial student network model; and the image category corresponding to the to-be-classified image is obtained based on the classification result, which indicates that the method can also obtain a high-accuracy preset image classification model through back propagation training to classify the to-be-classified model under the condition that the resource environment is limited, and an accurate image classification result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image classification, and in particular to an image classification method, apparatus, device and storage medium. Background Art

[0002] Image classification technology is a key computer vision task that analyzes and identifies the content in an image and assigns it to a predefined category. Existing image classification technologies mainly rely on deep learning models, especially convolutional neural networks (CNNs), which can achieve efficient and accurate classification of images by learning features from a large number of annotated images.

[0003] In the existing methods of classifying images through classification models, parameters and complex structures are required to obtain accurate classification results. Therefore, a high-performance central processing unit (CPU) or graphics processing unit (GPU) is needed for training and inference, and a large amount of memory space is required to store model parameters and intermediate calculation results. When faced with a resource-constrained environment, the accuracy of the obtained classification model is insufficient.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of the present application is to provide an image classification method, apparatus, device and storage medium, aiming to solve the problem that the classification model obtained by the prior art in a resource-constrained environment is not accurate enough, resulting in inaccurate image categories of the images to be classified.

[0006] To achieve the above objectives, the present application proposes an image classification method, which includes:

[0007] Acquire an image to be classified, and input the image to be classified into a preset image classification model for classification, wherein the preset image classification model is a model obtained by back-propagating an initial student network model based on a first distillation loss value obtained based on a first type of activation map and a second type of activation map, wherein the first type of activation map is obtained by extracting features of a sample image using a first teacher network model, and the second type of activation map is obtained by extracting features of a sample image using the initial student network model;

[0008] The image category corresponding to the image to be classified is obtained based on the classification result.

[0009] In one embodiment, before the step of obtaining the image to be classified, the method further includes:

[0010] Obtain a sample image, input the sample image into a first teacher network model to obtain a first type of activation map, and input the sample image into an initial student network model to obtain a second type of activation map;

[0011] A first distillation loss value is obtained based on the first type of activation map and the second type of activation map, and back-propagation training is performed on the initial student network model through the first distillation loss value to obtain a preset image classification model.

[0012] In one embodiment, the step of inputting the sample image into a first teacher network model to obtain a first type of activation map, and inputting the sample image into an initial student network model to obtain a second type of activation map comprises:

[0013] Input the sample image into a first teacher network model to obtain a first feature vector, and input the sample image into an initial student network model to obtain a second feature vector;

[0014] Determine a first gradient value of the first eigenvector, and obtain a first type of activation map based on the first gradient value and the first eigenvector;

[0015] A second gradient value of the second eigenvector is determined, and a second type activation map is obtained based on the second gradient value and the second eigenvector.

[0016] In one embodiment, the step of obtaining a first distillation loss value based on the first type activation map and the second type activation map includes:

[0017] Obtaining a first batch value, a first vector height, and a first vector width corresponding to the first type of activation map;

[0018] Obtain a second batch value, a second vector height, and a second vector width corresponding to the second type of activation map;

[0019] A first distillation loss value is obtained based on the first batch value, the first vector height, the first vector width, the second batch value, the second vector height, and the second vector width.

[0020] In one embodiment, the step of performing back propagation training on the initial student network model through the first distillation loss value to obtain a preset image classification model includes:

[0021] Obtaining a sample label corresponding to the sample image, and inputting the sample image into the initial student network model to obtain a first soft label;

[0022] A cross entropy loss value is obtained based on the sample label and the first soft label, and back propagation training is performed on the initial student network model through the first distillation loss value and the cross entropy loss value to obtain a preset image classification model.

[0023] In one embodiment, the step of performing back propagation training on the initial student network model through the first distillation loss value and the cross entropy loss value to obtain a preset image classification model includes:

[0024] Inputting the sample image into a second teacher network model to obtain a second soft label;

[0025] A second distillation loss value is obtained based on the first soft label and the second soft label, and back-propagation training is performed on the initial student network model through the first distillation loss value, the cross entropy loss value and the second distillation loss value to obtain a preset image classification model.

[0026] In one embodiment, the step of obtaining a second distillation loss value based on the first soft tag and the second soft tag includes:

[0027] Obtaining a first category number and a third batch value corresponding to the first soft tag, and obtaining a second category number and a fourth batch value corresponding to the second soft tag;

[0028] Determine a first probability distribution corresponding to the first soft label based on the first soft label and the first number of categories, and determine a second probability distribution corresponding to the second soft label based on the second soft label and the second number of categories;

[0029] A second distillation loss value is obtained based on the first category number, the third batch value, the first probability distribution, the second category number, the fourth batch value, and the second probability distribution.

[0030] In addition, to achieve the above purpose, the present application also proposes an image classification device, the device comprising:

[0031] An image analysis module, used for acquiring an image to be classified, and inputting the image to be classified into a preset image classification model for classification, wherein the preset image classification model is a model obtained by back-propagating an initial student network model based on a first distillation loss value obtained from a first type of activation map and a second type of activation map, wherein the first type of activation map is obtained by extracting features from a sample image using a first teacher network model, and the second type of activation map is obtained by extracting features from a sample image using the initial student network model;

[0032] The category analysis module is used to obtain the image category corresponding to the image to be classified based on the classification result.

[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes an image classification device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image classification method described above.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image classification method described above are implemented.

[0035] The present application proposes an image classification method, apparatus, device and storage medium, the method comprising: obtaining an image to be classified, and inputting the image to be classified into a preset image classification model for classification, the preset image classification model being a model obtained by back-propagating an initial student network model based on a first distillation loss value obtained from a first type of activation map and a second type of activation map, the first type of activation map being obtained by extracting features of a sample image through a first teacher network model, and the second type of activation map being obtained by extracting features of a sample image through the initial student network model; obtaining an image category corresponding to the image to be classified based on the classification result, which shows that the present application can obtain an image to be classified by inputting the image to be classified into a preset image classification model The classification result is obtained by the type, and the image category corresponding to the image to be classified is obtained based on the classification result. Since the preset image classification model is obtained by back-propagation training of the initial student network model using the first type of activation map obtained by performing feature extraction on the sample image through the first teacher network and the second type of activation map obtained by performing feature extraction on the sample image through the initial student network model, in the case of limited resources and environment, the teacher network model result can be trained through back-propagation training knowledge distillation to the initial student model to obtain a high-accuracy preset image classification model with a simple structure to classify the model to be classified, and obtain an accurate image classification result. Based on the accurate image classification result, the image category corresponding to the image to be classified can be accurately judged. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0038] Figure 1 This is a flow chart of a first embodiment of the image classification method proposed in this embodiment;

[0039] Figure 2 It is a preset image classification model training graph in the image classification method proposed in this embodiment;

[0040] Figure 3 This is an example diagram of the calculation process of the first distillation loss value in the image classification method proposed in this embodiment;

[0041] Figure 4 This is a flow chart of a second embodiment of the image classification method proposed in this embodiment;

[0042] Figure 5 This is an example diagram of the second distillation loss value calculation process in the image classification method proposed in this embodiment;

[0043] Figure 6 A diagram of an image classification device provided in this embodiment;

[0044] Figure 7 FIG. 4 is a schematic diagram of the structure of an image classification device suitable for implementing the present embodiment.

[0045] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0048] It should be noted that all directional indications in the embodiments of the present application (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0049] It is understandable that image classification technology is a key computer vision task that analyzes and identifies the content in an image and assigns it to a predefined category. Existing image classification technologies mainly rely on deep learning models, especially convolutional neural networks (CNNs), which can achieve efficient and accurate classification of images by learning features from a large number of annotated images.

[0050] In the existing methods of classifying images through classification models, in order to obtain accurate classification results, a large number of sample images and sample labels are required to train the classification model. When faced with a resource-constrained environment, the accuracy of the obtained classification model is insufficient, resulting in inaccurate image categories of the images to be classified.

[0051] Therefore, in order to solve the problem that the classification model obtained by the prior art is not accurate enough when facing a resource-constrained environment, resulting in inaccurate image categories of the images to be classified, the present embodiment proposes an image classification method, apparatus, device and storage medium, the method comprising: obtaining the image to be classified, and inputting the image to be classified into a preset image classification model for classification, the preset image classification model is a model obtained by back-propagation training of an initial student network model based on a first distillation loss value obtained based on a first type of activation map and a second type of activation map, the first type of activation map is obtained by extracting features of a sample image through a first teacher network model, and the second type of activation map is obtained by extracting features of a sample image through an initial student network model; based on the classification result, the image category corresponding to the image to be classified is obtained. This means that in this embodiment, the classification result can be obtained by inputting the image to be classified into the preset image classification model, and the image category corresponding to the image to be classified can be obtained based on the classification result. Since the preset image classification model is obtained by back-propagation training the initial student network model using the first type of activation map obtained by extracting features from the sample image through the first teacher network and the second type of activation map obtained by extracting features from the sample image through the initial student network model, in the case of limited resources and environment, the teacher network model result can be trained through back-propagation knowledge distillation to the initial student model to obtain a high-accuracy preset image classification model with a simple structure to classify the model to be classified, and accurate image classification results can be obtained. Based on the accurate image classification results, the image category corresponding to the image to be classified can be accurately judged.

[0052] For ease of understanding, the following combination Figures 1 to 7 The image classification method provided in the embodiment of the present application and the image classification method, apparatus, device and storage medium provided in the following embodiments are specifically introduced.

[0053] The present application embodiment provides an image classification method, referring to Figure 1 , Figure 1This is a flow chart of the first embodiment of the image classification method proposed in this embodiment.

[0054] like Figure 1 As shown, the method includes:

[0055] Step S10: Obtain an image to be classified, and input the image to be classified into a preset image classification model for classification, wherein the preset image classification model is a model obtained by back-propagation training an initial student network model based on a first distillation loss value obtained based on a first type of activation map and a second type of activation map, wherein the first type of activation map is obtained by extracting features of a sample image using a first teacher network model, and the second type of activation map is obtained by extracting features of a sample image using the initial student network model.

[0056] Step S20: Obtaining the image category corresponding to the image to be classified based on the classification result.

[0057] It should be noted that the execution subject of this embodiment can be a computing service device with image classification, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, etc. The following takes an image classification device (hereinafter referred to as the device) as an example to illustrate this embodiment and the following embodiments.

[0058] It should be noted that the above-mentioned image to be classified can be an image selected by the user to be classified, which can be stored in the memory or directly uploaded to the above-mentioned device by the user. The above-mentioned classification result can be a feature label corresponding to the image (such as mountain, river, color, machinery, person), and the above-mentioned image category can be a category corresponding to the above-mentioned image to be classified (such as landscape image, character image, etc., or more specific such as man image, etc.). The above-mentioned preset classification image model can be obtained by performing a knowledge distillation process on the initial student network model through the first teacher network. It inherits the feature extraction capability of the teacher network while maintaining the lightweight characteristics and can run efficiently in a resource-constrained environment.

[0059] The knowledge distillation process involves two key networks: the first teacher network model and the initial student network model. The first teacher network model is a pre-trained large network model with strong feature extraction capabilities. The device inputs the sample image into the first teacher network model to obtain the first type of activation map. The first type of activation map reflects the characteristic response of the sample image in the teacher network and contains rich image information.

[0060] The initial student network model is a lightweight network with fewer parameters and is more suitable for running in a resource-constrained environment. The device inputs the sample image into the initial student network model to obtain the second type of activation map. The second type of activation map reflects the characteristic response of the sample image in the initial student network model and contains less information than the first type of activation map.

[0061] In order to make the initial student network model learn the knowledge of the teacher network, the device calculates the first distillation loss value based on the first type activation map and the second type activation map. The first distillation loss value measures the difference between the initial student network model and the teacher network in feature extraction. The device uses the first distillation loss value to perform back-propagation training on the initial student network model and update its parameters to make it better approximate the feature extraction capability of the teacher network.

[0062] After knowledge distillation training, the initial student network model is transformed into a preset image classification model. This model inherits the feature extraction capability of the teacher network while maintaining its lightweight characteristics and can run efficiently in a resource-constrained environment. The device inputs the image to be classified into the preset image classification model to obtain the classification result, thereby achieving efficient classification of the image.

[0063] In a specific implementation, the above-mentioned device obtains an image to be classified from the outside, and inputs it into a preset image classification model for classification, wherein the above-mentioned device may also preprocess the image to be classified before inputting it into the preset image classification model, and processes the image to be classified through data enhancement, normalization, and noise elimination to obtain an image to be classified with clearer features, and then inputs the preprocessed image to be classified into the preset image classification model for classification, obtains the label type of the image to be classified as a classification result, and obtains the image category corresponding to the image to be classified based on the classification result.

[0064] Furthermore, in order to obtain a preset image through knowledge distillation, before the step of obtaining the image to be classified, it also includes:

[0065] Step S01: obtaining a sample image, inputting the sample image into a first teacher network model to obtain a first type of activation map, and inputting the sample image into an initial student network model to obtain a second type of activation map;

[0066] Step S02: obtaining a first distillation loss value based on the first type of activation map and the second type of activation map, and performing back-propagation training on the initial student network model through the first distillation loss value to obtain a preset image classification model.

[0067] It should be noted that the sample images can be images of known image categories preset by the user, which are used to train the first teacher network and the initial student network model to obtain a preset image classification model. The first teacher network model can be a large model for classifying images, which can be a model with a large number of parameters that can accurately analyze sample images, such as the MobileNet V3 Large model and the Vision Encoder model (Contrastive Language-Image Pre-training, CLIP) in the Vision Encoder. The first type of activation map can be obtained by extracting features from the sample image by the first teacher network model, which reflects the characteristic response of the sample image in the teacher network and contains rich image information. The initial student model can be a lightweight network model with a small number of parameters, which can run normally in a resource-constrained environment. The second type of activation map can be obtained by extracting features from the sample image by the initial student network model, which can reflect the characteristic response of the sample image in the initial student network model. Compared with the first type of activation map, it contains less information. The first distillation loss value can be obtained based on the first type of activation map and the second type of activation map through a loss function (such as a cosine similarity loss function), which can be used to measure the difference between the initial student network model and the teacher network in feature extraction. The back propagation training can be a method for training a neural network, which can be achieved by calculating the gradient of the first distillation loss value to the model parameters and updating the model parameters using a gradient descent algorithm to minimize the loss function.

[0068] refer to Figure 2 , Figure 2 The image classification model training diagram is preset in the image classification method proposed in this embodiment. In the specific implementation process, before the above-mentioned device performs the image classification task, it is necessary to first train and prepare the model. Specifically, the device first obtains a sample image (i.e. Figure 2 These sample images are images with labeled categories and are used to train and optimize the network model. Next, the device inputs the sample images into the first teacher network model (i.e. Figure 2 The teacher network 1 (MobileNetV3 Large) and the initial student network model (i.e. Figure 2 Middle school student network (Mobile NetV3 Small).

[0069] The first teacher network model is a pre-trained, powerful network that can perform deep feature extraction on the input sample image and generate the first type of activation map (i.e. Figure 2 Middle class activation Figure 1(Teacher Network 1)). These activation maps represent the feature responses of the image at various levels. At the same time, the above device also inputs the same sample image into the initial student network model, which is a relatively small network whose purpose is to learn the knowledge of the teacher network. Through this process, the device obtains the second type of activation map (i.e. Figure 2 Middle class activation Figure 2 (Student Network)).

[0070] Subsequently, the above-mentioned device calculates the first distillation loss value (i.e., Figure 2 The distillation loss value is 1). This loss value measures the difference between the initial student network model activation map and the teacher network activation map, and is a key indicator in the knowledge distillation process. The above device uses this loss value to train the initial student network model through the back propagation algorithm (i.e. Figure 2 ), optimizing its parameters so that it can better mimic the behavior of the teacher network.

[0071] Through such a training process, the above-mentioned device finally obtains a preset image classification model, which is a lightweight and efficient network that can perform image classification tasks in a resource-constrained environment. For example, in an autonomous driving system, the device may need to identify various objects on the road, such as pedestrians, vehicles, traffic signs, etc. Through the above training process, the preset image classification model can accurately identify these objects, thereby helping the autonomous driving system to make correct decisions.

[0072] In summary, the above-mentioned device transfers the knowledge of the teacher network to the initial student network model through knowledge distillation technology, so that the initial student network model can achieve classification performance similar to that of the teacher network while maintaining a smaller scale. This is of great significance for achieving efficient image classification on mobile devices or embedded systems.

[0073] Furthermore, in order to obtain an accurate class activation map and obtain an accurate preset image classification model, the step of inputting the sample image into the first teacher network model to obtain a first class activation map, and inputting the sample image into the initial student network model to obtain a second class activation map includes:

[0074] Input the sample image into a first teacher network model to obtain a first feature vector, and input the sample image into an initial student network model to obtain a second feature vector;

[0075] Determine a first gradient value of the first eigenvector, and obtain a first type of activation map based on the first gradient value and the first eigenvector;

[0076] A second gradient value of the second eigenvector is determined, and a second type activation map is obtained based on the second gradient value and the second eigenvector.

[0077] It should be noted that the above feature vector can be a numerical representation of the image after being processed by a neural network, which contains abstract feature information of the image and is used for classification or recognition tasks. The above gradient value can be the derivative of the loss function with respect to the network parameters, which indicates the direction in which the loss function increases fastest in the parameter space.

[0078] refer to Figure 2 In the specific implementation, the above device takes meticulous steps to ensure that the initial student network model can effectively learn the knowledge of the teacher network during the knowledge distillation process. Specifically, the device first inputs the sample image into the first teacher network model, which processes the image through its complex network structure and outputs the first feature vector (i.e. Figure 2 This feature vector contains the high-level abstract features of the image and is the core representation of the teacher network’s understanding of the image content.

[0079] Next, the device also inputs the sample image into the initial student network model, which, despite its simple structure, is able to output the second eigenvector (i.e. Figure 2 In the teacher network, the first feature vector is the feature vector 2 (student network), which is the student's preliminary understanding of the image content. In order to further extract the activation map, the above device needs to determine the gradient values ​​of the two feature vectors, and calculate the first gradient value of the first feature vector, and then generate a first type of activation map based on this gradient value and the first feature vector. This process is actually identifying the image features in the teacher network that can most affect the classification results. Similarly, the device also calculates the second gradient value of the second feature vector, and generates a second type of activation map based on this gradient value and the second feature vector.

[0080] In one example, the class activation mapping (Gradient-weighted Class Activation Mapping, Grad-CAM) (i.e. Figure 2 The above-mentioned class activation map is obtained by the method of Grad-CAM in the convolutional layer. The Grad-CAM method is used for explanation below, but this embodiment is not specifically limited. First, the input image is forward propagated to obtain the feature map and category score of the last convolutional layer, and then the gradient of the category score to the feature map is calculated. Then, each feature map is globally averaged pooled and multiplied by the corresponding gradient to obtain a weighted feature map. Finally, all weighted feature maps are summed and upsampled to the input image size through the ReLU function and bilinear interpolation, thereby generating a heat map that highlights the most important area in the image for a specific category decision, that is, a class activation map.

[0081] Furthermore, in order to obtain an accurate first distillation loss value, the step of obtaining the first distillation loss value based on the first type activation map and the second type activation map includes:

[0082] Obtaining a first batch value, a first vector height, and a first vector width corresponding to the first type of activation map;

[0083] Obtain a second batch value, a second vector height, and a second vector width corresponding to the second type of activation map;

[0084] A first distillation loss value is obtained based on the first batch value, the first vector height, the first vector width, the second batch value, the second vector height, and the second vector width.

[0085] It should be noted that the first batch value can be the number of samples contained in the current batch when processing the first type of activation map. When processing data in batches, each batch can contain one or more samples, and the batch value indicates the size of the batch. The first vector height can be the size of the first type of activation map in the vertical direction, that is, the height of the activation map. It represents the resolution of the activation map in the image height direction. The first vector width can be the size of the first type of activation map in the horizontal direction, that is, the width of the activation map. It represents the resolution of the activation map in the image width direction.

[0086] The second batch value may be the number of samples included in the current batch when processing the second type of activation map. When processing data in batches, each batch may contain one or more samples, and the batch value indicates the size of the two batches. The second vector height may be the size of the second type of activation map in the vertical direction, that is, the height of the activation map. It represents the resolution of the activation map in the image height direction. The second vector width may be the size of the second type of activation map in the horizontal direction, that is, the width of the activation map. It represents the resolution of the activation map in the image width direction.

[0087] In the specific implementation, the device extracts the corresponding batch value, vector height, and vector width from the first type of activation map, which together define the data structure of the activation map. Specifically, the first batch value corresponding to the first type of activation map indicates the size of the data batch currently being processed, while the first vector height and the first vector width represent the size of the activation map in the vertical and horizontal directions, respectively. These parameters are critical to ensure that the activation map can be correctly aligned with the second type of activation map.

[0088] Next, the device also obtains the corresponding batch value, vector height and vector width from the second type of activation map, namely the second batch value, the second vector height and the second vector width. These parameters and the corresponding parameters of the first type of activation map together constitute the basis for comparison between the two activation maps. The second batch value ensures the consistency of the two activation maps in batch processing, while the second vector height and the second vector width ensure the compatibility of the two activation maps in the spatial dimension. Finally, based on these parameters, the above device calculates the first distillation loss value.

[0089] refer to Figure 2 as well as Figure 3 , Figure 3 This is an example diagram of the calculation process of the first distillation loss value in the image classification method proposed in this embodiment. As shown in the figure, the cosine similarity function (i.e. Figure 2 The cosine similarity loss function is used as the loss function for calculation, but this embodiment is not specifically limited. Figure 3 Class activation in Figure 1 and class activation Figure 3 First, it will go through the Reshape module to convert it into a two-dimensional vector. Then the two two-dimensional vectors will calculate the cosine similarity and get the final distillation loss value of 1. Figure 3 CAM (ClassActivationMAP) is the class activation map, and Reshape is the flattening module. The first preset loss function is:

[0090]

[0091] Where N is the batch size, H is the vector height, W is the vector width, and X 1,i represents the i-th element of matrix X1, which is a matrix composed of the vector height and vector width of the first type of activation map obtained by the first teacher network model above. 2,i Represents the i-th element of matrix X2, where matrix X2 is a matrix consisting of the vector height and vector width of the second type of activation map obtained by the above initial student network model.

[0092] This embodiment proposes an image classification method, which includes: obtaining an image to be classified, and inputting the image to be classified into a preset image classification model for classification, the preset image classification model is a model obtained by back-propagation training an initial student network model based on a first distillation loss value obtained based on a first type of activation map and a second type of activation map, the first type of activation map is obtained by extracting features of a sample image through a first teacher network model, and the second type of activation map is obtained by extracting features of a sample image through an initial student network model; based on the classification result, an image category corresponding to the image to be classified is obtained. This means that in this embodiment, the classification result can be obtained by inputting the image to be classified into the preset image classification model, and the image category corresponding to the image to be classified can be obtained based on the classification result. Since the preset image classification model is obtained by back-propagation training the initial student network model using the first type of activation map obtained by extracting features from the sample image through the first teacher network and the second type of activation map obtained by extracting features from the sample image through the initial student network model, in the case of limited resources and environment, the teacher network model result can be trained through back-propagation knowledge distillation to the initial student model to obtain a high-accuracy preset image classification model with a simple structure to classify the model to be classified, and accurate image classification results can be obtained. Based on the accurate image classification results, the image category corresponding to the image to be classified can be accurately judged.

[0093] Based on the first embodiment, in the second embodiment, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be described in detail later. Figure 4 , Figure 4 This is a flow chart of the second embodiment of the image classification method proposed in this embodiment. In order to further obtain a more accurate preset image classification model, the step of performing back propagation training on the initial student network model through the first distillation loss value to obtain the preset image classification model includes:

[0094] Step S021: obtaining a sample label corresponding to the sample image, and inputting the sample image into the initial student network model to obtain a first soft label;

[0095] Step S022: obtaining a cross entropy loss value based on the sample label and the first soft label, and performing back propagation training on the initial student network model through the first distillation loss value and the cross entropy loss value to obtain a preset image classification model.

[0096] It should be noted that the sample label can be the label corresponding to the sample image used to determine the image category of the sample image. The first soft label can be the probability distribution output by the initial student network model when classifying the sample image, which represents the model's prediction confidence for each category. The cross entropy loss value is a commonly used loss function used to measure the difference between the probability distribution predicted by the model and the true label. It reflects the accuracy of the model classification result.

[0097] refer to Figure 2 In a specific implementation, the above device first obtains the sample label corresponding to the sample image (i.e. Figure 2 The real labels (for each category) in the image are the correct classification information of the image. Then, the above device inputs the sample image into the initial student network model, and the model outputs the first soft label (i.e. Figure 2 Soft label 1 (student network) in the middle, that is, the probability distribution of the model for classifying the input image. Subsequently, the above-mentioned device calculates the cross entropy loss value (i.e. Figure 2 The device combines the first distillation loss value, which reflects the similarity between the activation map of the student network and the teacher network. Through the combination of these two loss values, the device performs back propagation training on the initial student network model, optimizes the parameters of the model, and finally obtains a preset image classification model with a high classification accuracy.

[0098] Furthermore, in order to make the preset image classification model more accurate, the present embodiment also introduces a second teacher network model. The step of performing back propagation training on the initial student network model through the first distillation loss value and the cross entropy loss value to obtain the preset image classification model includes:

[0099] Inputting the sample image into a second teacher network model to obtain a second soft label;

[0100] A second distillation loss value is obtained based on the first soft label and the second soft label, and back-propagation training is performed on the initial student network through the first distillation loss value, the cross entropy loss value, and the second distillation loss value to obtain a preset image classification model.

[0101] It should be noted that the second teacher network model can be another pre-trained network model that is more complex than the initial student network model, which is used to provide additional supervision information to help the student network learn. The second soft label is a probability distribution output by the second teacher network model, which provides another learning goal for the student network and helps improve its classification ability. The second distillation loss value can be calculated based on the difference between the soft labels output by the student network and the second teacher network model.

[0102] refer to Figure 2 In the specific implementation, first, the above device inputs the sample image into the second teacher network model (i.e. Figure 2 Teacher Network 2 (Vision Encoder) in the middle to obtain the second soft label (i.e. Figure 2 Soft label 2 (teacher network 2) in the initial student network model is the probability distribution of the second teacher network model for classifying the sample image. Next, the device calculates the second distillation loss value (i.e., Figure 2 The distillation loss value 2) measures the closeness of the student network to the second teacher network model in probability distribution.

[0103] Then, the apparatus combines the first distillation loss value, the cross entropy loss value, and the second distillation loss value (i.e. Figure 2 The initial student network model is trained by back-propagation (weighted average in the training set). This multi-loss function training strategy helps the student network learn the knowledge of the teacher network more comprehensively, including feature representation and classification decisions. Through this comprehensive training, the student network can obtain classification performance comparable to that of the teacher network while maintaining a smaller scale and lower computational complexity, and finally form a preset image classification model.

[0104] Furthermore, in order to obtain an accurate second distillation loss value, the step of obtaining the second distillation loss value based on the first soft tag and the second soft tag includes:

[0105] Obtaining a first category number and a third batch value corresponding to the first soft tag, and obtaining a second category number and a fourth batch value corresponding to the second soft tag;

[0106] Determine a first probability distribution corresponding to the first soft label based on the first soft label and the first number of categories, and determine a second probability distribution corresponding to the second soft label based on the second soft label and the second number of categories;

[0107] A second distillation loss value is obtained based on the first category number, the third batch value, the first probability distribution, the second category number, the fourth batch value, and the second probability distribution.

[0108] It should be noted that the first number of categories may be the total number of categories in the classification task represented by the first soft label, the third batch value may be the number of samples in the current batch when processing the first soft label, and the first probability distribution may be the probability distribution of the classification results of the sample image by the student network. The second number of categories may be the total number of categories in the classification task represented by the second soft label, the fourth batch value may be the number of samples in the current batch when processing the second soft label, and the second probability distribution may be the probability distribution of the classification results of the sample image by the second teacher network.

[0109] In a specific implementation, the above device first obtains the relevant parameters of the first soft label and the second soft label during the process of calculating the second distillation loss value. Specifically, the device extracts the first category number and the third batch value corresponding to the first soft label, and this information is used to determine the probability distribution of the first soft label in each category. Similarly, the device also obtains the second category number and the fourth batch value corresponding to the second soft label, so as to construct the probability distribution of the second soft label.

[0110] Next, the device determines a first probability distribution based on the first number of categories and the first soft label, which is achieved by normalizing each value in the first soft label to the corresponding category. Similarly, the device also determines a second probability distribution based on the second number of categories and the second soft label. The two probability distributions represent the confidence of the student network and the second teacher network in classifying the sample image, respectively.

[0111] Finally, the device calculates a second distillation loss value using the parameters, including the first number of categories, the third batch value, the first probability distribution, the second number of categories, the fourth batch value, and the second probability distribution. This loss value reflects the difference between the two probability distributions and is an important indicator for guiding the learning of the student network. By minimizing the second distillation loss value, the device can cause the student network to better imitate the classification behavior of the second teacher network, thereby improving its classification performance.

[0112] refer to Figure 2 as well as Figure 5 , Figure 5 This is an example diagram of the second distillation loss value calculation process in the image classification method proposed in this embodiment. Figure 5 The soft labels 1 and 2 in the distillation loss value 2 are directly calculated. The second preset loss function (i.e. Figure 2 The KL divergence loss function in ( ) is:

[0113]

[0114] Among them, N is the batch size, C is the number of categories, P1 and P2 are the results obtained by softmax function for soft label 1 and soft label 2 respectively. 1,n,c It actually represents the probability that the nth sample belongs to the cth category in the prediction of the initial student network model, P 2,n,c It actually represents the probability that the nth sample belongs to the cth category in the prediction of the second teacher network model, where

[0115]

[0116] Among them, the softmax function is a commonly used activation function in multi-class classification problems. Its function is to convert a real number vector into a probability distribution. Logits1 is soft label 1, Logits2 is soft label 2, and Logits 1,c It represents the original prediction value of the initial student network model for the cth category, Logits 2,c It represents the original prediction value of the second teacher network model for the cth category.

[0117] In addition, it should be noted that after obtaining the first distillation value, the second distillation value and the cross entropy loss value, the above-mentioned device can perform weighted average fusion on the above-mentioned first distillation value, the second distillation value and the cross entropy loss value, and use the result after the weighted average fusion to train the initial student network model to obtain a preset image classification model.

[0118] This embodiment also provides an image classification device, please refer to Figure 6 , Figure 6 The image classification device provided in this embodiment includes:

[0119] An image analysis module, used for acquiring an image to be classified, and inputting the image to be classified into a preset image classification model for classification, wherein the preset image classification model is a model obtained by back-propagating an initial student network model based on a first distillation loss value obtained from a first type of activation map and a second type of activation map, wherein the first type of activation map is obtained by extracting features from a sample image using a first teacher network model, and the second type of activation map is obtained by extracting features from a sample image using the initial student network;

[0120] The category analysis module is used to obtain the image category corresponding to the image to be classified based on the classification result. In the image classification device provided by this embodiment, it should be noted that the above-mentioned slicing module includes a video recording module and an audio recording module. The image classification method in the above-mentioned embodiment can solve the problem that the classification model obtained by the prior art in a resource-constrained environment is not accurate enough, resulting in the image category of the image to be classified being inaccurate. Compared with the prior art, the beneficial effects of the image classification device provided by this embodiment are the same as the beneficial effects of the image classification method provided by the above-mentioned embodiment, and other technical features in the image classification device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0121] This embodiment provides an image classification device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image classification method in the above-mentioned embodiment one.

[0122] Reference below Figure 7 , Figure 7 : is a schematic diagram of the structure of an image classification device suitable for implementing the present embodiment. The image classification device in the present embodiment may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players: portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The image classification device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0123] like Figure 7As shown, the image classification device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the image classification device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image classification device to communicate with other devices wirelessly or by wire to exchange data. Although the image classification device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0124] In particular, according to the present embodiment, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present embodiment includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present embodiment are executed.

[0125] The image classification device provided in this embodiment adopts the image classification method in the above embodiment, which can solve the problem that the classification model obtained in the prior art in a resource-constrained environment is not accurate enough, resulting in the image category of the image to be classified being inaccurate. Compared with the prior art, the beneficial effects of the image classification device provided in this embodiment are the same as the beneficial effects of the image classification method provided in the above embodiment, and other technical features in the image classification device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0126] It should be understood that the various parts disclosed in this embodiment can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0127] The above is only a specific implementation of this embodiment, but the protection scope of this embodiment is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in this embodiment, which should be included in the protection scope of this embodiment. Therefore, the protection scope of this embodiment should be based on the protection scope of the claims.

[0128] This embodiment provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the image classification method in the above embodiment.

[0129] The computer-readable storage medium provided in this embodiment may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0130] The computer-readable storage medium may be included in the image classification device; or may exist independently without being assembled into the image classification device.

[0131] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the image classification device, the image classification device is enabled to: perform image classification.

[0132] The computer program code for performing the operations of the present embodiment may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present embodiment. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0134] The modules involved in the description of this embodiment may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0135] The readable storage medium provided in this embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned image classification method, and is intended to solve the problem that the classification model obtained in the prior art in a resource-constrained environment is not accurate enough, resulting in the image category of the image to be classified being inaccurate. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as the beneficial effects of the image classification method provided in the above-mentioned embodiment, and are not described in detail here.

[0136] This embodiment also provides a computer program product, including a computer program, which implements the steps of the above-mentioned image classification method when executed by a processor.

[0137] The computer program product provided in this embodiment can solve the problem that the classification model accuracy obtained in the prior art in a resource-constrained environment is insufficient, resulting in inaccurate image categories of the images to be classified. Compared with the prior art, the beneficial effects of the computer program product provided in this embodiment are the same as the beneficial effects of the image classification method provided in the above embodiment, and will not be described in detail here.

[0138] The above descriptions are only some embodiments, and are not intended to limit the patent scope of this embodiment. All equivalent structural changes made using the contents of the specification and drawings of this application under the technical concept of this application, or directly / indirectly applied in other related technical fields are included in the patent protection scope of this application.

Claims

1. An image classification method, characterized in that: The method comprises: Acquire an image to be classified, and input the image to be classified into a preset image classification model for classification, wherein the preset image classification model is a model obtained by back-propagating an initial student network model based on a first distillation loss value obtained based on a first type of activation map and a second type of activation map, wherein the first type of activation map is obtained by extracting features of a sample image using a first teacher network model, and the second type of activation map is obtained by extracting features of a sample image using the initial student network model; The image category corresponding to the image to be classified is obtained based on the classification result.

2. The method according to claim 1, characterized in that Before the step of obtaining the image to be classified, the method further includes: Obtain a sample image, input the sample image into a first teacher network model to obtain a first type of activation map, and input the sample image into an initial student network model to obtain a second type of activation map; A first distillation loss value is obtained based on the first type of activation map and the second type of activation map, and back-propagation training is performed on the initial student network model through the first distillation loss value to obtain a preset image classification model.

3. The method according to claim 2, characterized in that The step of inputting the sample image into the first teacher network model to obtain a first type of activation map, and inputting the sample image into the initial student network model to obtain a second type of activation map comprises: Input the sample image into a first teacher network model to obtain a first feature vector, and input the sample image into an initial student network model to obtain a second feature vector; Determine a first gradient value of the first eigenvector, and obtain a first type of activation map based on the first gradient value and the first eigenvector; A second gradient value of the second eigenvector is determined, and a second type activation map is obtained based on the second gradient value and the second eigenvector.

4. The method according to claim 2, characterized in that The step of obtaining a first distillation loss value based on the first type activation map and the second type activation map includes: Obtaining a first batch value, a first vector height, and a first vector width corresponding to the first type of activation map; Obtain a second batch value, a second vector height, and a second vector width corresponding to the second type of activation map; A first distillation loss value is obtained based on the first batch value, the first vector height, the first vector width, the second batch value, the second vector height, and the second vector width.

5. The method according to claim 2, characterized in that The step of performing back propagation training on the initial student network model through the first distillation loss value to obtain a preset image classification model includes: Obtaining a sample label corresponding to the sample image, and inputting the sample image into the initial student network model to obtain a first soft label; A cross entropy loss value is obtained based on the sample label and the first soft label, and back propagation training is performed on the initial student network model through the first distillation loss value and the cross entropy loss value to obtain a preset image classification model.

6. The method according to claim 5, characterized in that The step of performing back propagation training on the initial student network model through the first distillation loss value and the cross entropy loss value to obtain a preset image classification model includes: Inputting the sample image into a second teacher network model to obtain a second soft label; A second distillation loss value is obtained based on the first soft label and the second soft label, and back-propagation training is performed on the initial student network model through the first distillation loss value, the cross entropy loss value and the second distillation loss value to obtain a preset image classification model.

7. The method according to claim 6, characterized in that The step of obtaining a second distillation loss value based on the first soft tag and the second soft tag includes: Obtaining a first category number and a third batch value corresponding to the first soft tag, and obtaining a second category number and a fourth batch value corresponding to the second soft tag; Determine a first probability distribution corresponding to the first soft label based on the first soft label and the first number of categories, and determine a second probability distribution corresponding to the second soft label based on the second soft label and the second number of categories; A second distillation loss value is obtained based on the first category number, the third batch value, the first probability distribution, the second category number, the fourth batch value, and the second probability distribution.

8. An image classification device, characterized in that: The device comprises: An image analysis module, used for acquiring an image to be classified, and inputting the image to be classified into a preset image classification model for classification, wherein the preset image classification model is a model obtained by back-propagating an initial student network model based on a first distillation loss value obtained from a first type of activation map and a second type of activation map, wherein the first type of activation map is obtained by extracting features from a sample image using a first teacher network model, and the second type of activation map is obtained by extracting features from a sample image using the initial student network model; The category analysis module is used to obtain the image category corresponding to the image to be classified based on the classification result.

9. An image classification device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image classification method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image classification method according to any one of claims 1 to 7 are implemented.