Image classification method and device, electronic equipment and computer readable storage medium
By combining pre-trained and to-be-trained image classification models with confidence information for image classification, the classification accuracy problem caused by incorrect labeling of image datasets is solved, and the accuracy of image classification is improved.
Patent Information
- Application Number
- CN202510677562.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The labeling error problem of image datasets in existing technologies seriously affects the model training effect and performance, making it difficult to accurately identify noisy samples and clean samples, resulting in decreased performance and generalization ability of image classification models.
Through the pre-trained image classification model and the image classification model to be trained, the first confidence and second confidence of the image to be classified in different image categories are determined respectively. The image to be classified is further classified based on the confidence, and accurate classification is performed using the more reliable confidence information provided by the pre-trained model.
The accuracy of image classification and the accuracy of the model are improved. Through the reference index of the confidence information of the pre-trained model, a more accurate classification result of the image to be classified can be determined, thereby improving the accuracy of image classification.
Smart Images

Figure CN120673125A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to an image classification method, device, electronic device, and computer-readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence and machine learning technologies, image recognition and classification tasks have been widely used in numerous fields, such as autonomous driving, medical imaging diagnosis, and security monitoring. In these applications, high-quality image datasets are crucial for training high-performance models. However, mislabeling is a common problem in image datasets, severely impacting model training and performance. For example, labeling in image datasets is often done manually. Due to factors such as the annotator's expertise, fatigue, and subjective judgment, labeling errors are inevitable. Furthermore, image datasets often contain a large number of images, covering a variety of complex scenes and objects. In some cases, even experienced annotators struggle to accurately distinguish certain categories, resulting in mislabeling. Furthermore, in practical applications, it can be difficult to separate mislabeled sample images from the image dataset. Summary of the Invention
[0003] The embodiments of the present application provide an image classification method, apparatus, electronic device, and computer-readable storage medium, which can improve the accuracy of image classification.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] An embodiment of the present application provides an image classification method, which includes: obtaining an image to be classified and an original label of the image to be classified; performing image classification on the image to be classified through a first image model to obtain a first classification label; in response to the first classification label being different from the original label, determining a first confidence of the image to be classified under different image categories through the first image model, and determining a second confidence of the image to be classified under the different image categories through a second image model; the first image model is a pre-trained image classification model, and the second image model is an image classification model to be trained; based on the first confidence and the second confidence, performing image classification on the image to be classified to obtain an image classification result.
[0006] An embodiment of the present application provides an image classification device, comprising: a data acquisition module for acquiring an image to be classified and an original label of the image to be classified; a first image classification module for performing image classification on the image to be classified through a first image model to obtain a first classification label; a confidence determination module for determining, in response to the first classification label being different from the original label, a first confidence of the image to be classified under different image categories through the first image model, and determining a second confidence of the image to be classified under the different image categories through a second image model; the first image model is a pre-trained image classification model, and the second image model is an image classification model to be trained; and a second image classification module for performing image classification on the image to be classified based on the first confidence and the second confidence to obtain an image classification result.
[0007] In the above scheme, the confidence determination module is also used to perform image transformation processing on the image to be classified to obtain at least one transformed image; summarize the image to be classified and the at least one transformed image to obtain a first image set; and determine the first confidence of the images in the first image set under the different image categories through the first image model.
[0008] In the above scheme, the confidence determination module is also used to determine, for each of the image categories, the prediction confidence of each image in the first image set under the image category through the first image model; average the prediction confidences of all images in the first image set to obtain a confidence mean; and determine the confidence mean as the first confidence of the images in the first image set under the image category.
[0009] In the above scheme, the second image classification module is also used to determine the maximum first confidence of the image to be classified under the different image categories; determine the maximum second confidence of the image to be classified under the different image categories; and perform image classification on the image to be classified based on the maximum first confidence and the maximum second confidence to obtain an image classification result.
[0010] In the above scheme, the second image classification module is also used to divide the image to be classified into a first image set in response to the maximum first confidence being less than the preset first confidence threshold, and the maximum second confidence being less than the preset second confidence threshold; the first image set includes images whose original labels match the image information; in response to the maximum first confidence being greater than or equal to the preset first confidence threshold, and / or the maximum second confidence being greater than or equal to the preset second confidence threshold, obtain the first image features output by the first data processing layer of the first image model and the second image features output by the second data processing layer of the second image model; perform image classification on the image to be classified based on the first image features and the second image features to obtain an image classification result.
[0011] In the above scheme, the second image classification module is also used to determine the feature similarity between the first image feature and the second image feature; in response to the feature similarity being greater than a preset similarity threshold, the image to be classified is divided into the first image set; in response to the feature similarity being less than or equal to the preset similarity threshold, the image to be classified is divided into the second image set; the second image set includes images whose original labels do not match the image information.
[0012] In the above scheme, the second image classification module is also used to obtain the current number of training rounds, the preset first hyperparameter and the preset second hyperparameter of the second image model; based on the current number of training rounds and the first hyperparameter, determine the second confidence threshold at the current moment; based on the current number of training rounds and the second hyperparameter, determine the similarity threshold at the current moment.
[0013] In the above solution, the first image classification module is further used to classify the image to be classified into a third image set in response to the first classification label being the same as the original label; the third image set includes images whose original labels match the image information.
[0014] An embodiment of the present application provides an electronic device, comprising: a memory for storing computer-executable instructions or computer programs; and a processor for implementing the image classification method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0015] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the image classification method provided in the embodiment of the present application when executed by a processor.
[0016] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the image classification method provided in the embodiment of the present application is implemented.
[0017] The embodiments of the present application have the following beneficial effects:
[0018] When classifying an image to be classified, first, the image to be classified is classified using a pre-trained image classification model to obtain a first classification label for the image to be classified. When the first classification label is different from the original label of the image to be classified, a first confidence level and a second confidence level of the image to be classified under different image categories are further determined, thereby further classifying the image to be classified based on the first confidence level and the second confidence level. In other words, by comparing whether the first classification label is the same as the original label, a preliminary rough classification of the image to be classified can be performed. Thereafter, when the first classification label is different from the original label, the first confidence level and the second confidence level of the image to be classified under different image categories can be determined using the pre-trained image classification model and the image classification model to be trained, respectively. Finally, the image to be classified is classified based on the first confidence level and the second confidence level to obtain an image classification result. In this way, since the pre-trained image classification model can provide more reliable confidence information, when further image classification is performed based on the confidence determined by the two models, the first confidence obtained by the pre-trained image classification model can be used as a reference indicator, so that based on this reference indicator, a more accurate classification result of the image to be classified can be determined, thereby improving the accuracy of image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is an optional flowchart of the image classification method provided in the embodiment of the present application;
[0020] Figure 2 This is another optional flowchart of the image classification method provided in the embodiment of the present application;
[0021] Figure 3 This is a schematic diagram of an implementation flow for determining a first confidence level provided in an embodiment of the present application;
[0022] Figure 4 This is a schematic diagram of the implementation process of image classification provided in the embodiment of the present application;
[0023] Figure 5 Schematic diagram of the principle of the image classification method provided in the embodiment of the present application;
[0024] Figure 6 This is a structural block diagram of an image classification device provided in an embodiment of the present application;
[0025] Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0027] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0028] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0029] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0030] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0031] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0032] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0033] 1) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0034] 2) Human-computer interaction interface, which is used to provide an interface for human-computer interaction functions / an interface for displaying image classification information.
[0035] For example, graphical user interface (GUI) display, such as augmented reality (AR) interface, virtual reality (VR) interface, voice user interface (VUI), interactive projection interface (using projection technology to display information on a plane), eye movement detection interface (interface controlled by detecting the user's line of sight), holographic interface (three-dimensional hologram formed by projecting images through holographic projection technology, so that stereoscopic images can be seen without wearing special glasses), multimodal interface (interface that combines multiple interaction methods, such as touch, vision, hearing, etc.), brain-machine interface (BMI) interface, etc.
[0036] In image classification tasks, the performance of image classification models degrades due to the presence of noise in the labels of sample data in the dataset (e.g., incorrect or inaccurate labeling). When fine-tuning image classification models on datasets with noisy labels, they are prone to overfitting and poor performance. Noisy labels cause the image classification model to learn incorrect information, which affects the accuracy of the image classification model for new images and reduces its generalization ability. However, related technologies have difficulty accurately distinguishing noisy samples from clean samples.
[0037] Based on at least one problem existing in the related art, the embodiments of the present application provide an image classification method, apparatus, electronic device and computer-readable storage medium. When classifying an image to be classified, first, the image to be classified is classified by a pre-trained image classification model to obtain a first classification label of the image to be classified. When the first classification label is different from the original label of the image to be classified, the first confidence and second confidence of the image to be classified under different image categories are further determined, so that the image to be classified is further classified based on the first confidence and the second confidence. In other words, by comparing whether the first classification label is the same as the original label, a preliminary rough classification of the image to be classified can be performed. Afterwards, when the first classification label is different from the original label, the first confidence and second confidence of the image to be classified under different image categories can be determined by the pre-trained image classification model and the image classification model to be trained respectively. Finally, the image to be classified is classified according to the first confidence and the second confidence to obtain an image classification result. In this way, since the pre-trained image classification model can provide more reliable confidence information, when further image classification is performed based on the confidence determined by the two models, the first confidence obtained by the pre-trained image classification model can be used as a reference indicator, so that based on this reference indicator, a more accurate classification result of the image to be classified can be determined, thereby improving the accuracy of image classification.
[0038] The image classification method provided in the embodiment of the present application can be applied to electronic devices such as laptops, tablet computers, and desktop computers. The embodiment of the present application does not impose any restrictions on the specific type of electronic device.
[0039] The image classification method provided in each embodiment of the present application can be executed by an electronic device, wherein the electronic device can be a server or a terminal, that is, the image classification method provided in each embodiment of the present application can be executed by a server, or by a terminal, or can be executed through interaction between a server and a terminal.
[0040] The image classification method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0041] Figure 1 This is an optional flow chart of the image classification method provided in the embodiment of the present application, which can be applied to electronic devices. The following will be explained by taking the electronic device as an example. Figure 1 As shown, the method includes the following steps S101 to S104:
[0042] Step S101: Obtain an image to be classified and its original label.
[0043] In the embodiments of the present application, it should be noted that the image to be classified here refers to the image in the classification task of classifying the image as a whole, rather than the image in the classification task of classifying the objects contained in the image. The original label of the image to be classified refers to the label of the object in the image in the classification task of classifying the objects contained in the image. For example, in a certain image data set, each sample image has a corresponding label, which is the original label, but there may be a situation where the label is inconsistent with the object contained in the image, which is a noise sample image. The embodiment of the present application can classify each sample image as a noise sample image.
[0044] There are many ways to obtain images to be classified: obtaining them from public data sets and collecting them from actual application scenarios, etc., which are not limited in the embodiments of the present application. For example, ImageNet is a public image classification data set. There are more than 1,000 categories in ImageNet, and each category has a large number of images. The original labels of these images are category names, such as "cat", "dog" and "car". For example, to classify an object in a scene, use an imaging sensor to capture an image of the object, and then use an automatic annotation tool to annotate the image. The annotation result can be used as the name of the image file when the image is saved, such as naming the image file "rose".
[0045] Step S102: classify the image to be classified using a first image model to obtain a first classification label.
[0046] Here, the first image model is a pre-trained image classification model.
[0047] In an embodiment of the present application, a pre-trained image classification model can be a model that has been trained with a large number of sample images and can be trained on a large-scale dataset (such as ImageNet). The pre-trained image classification model can recognize a variety of common image categories. The pre-trained image classification model can assign input images to predefined categories. For example, the pre-trained image classification model can classify images into categories such as "cat," "dog," and "car."
[0048] The image to be classified is fed into a pre-trained image classification model. The pre-trained image classification model uses the features and patterns learned during training to determine the most likely category of the input image. The output is the category label predicted by the pre-trained image classification model, which is also the first classification label.
[0049] For example, a pre-trained residual network model (Residual Network, ResNet) performs image classification on an image to be classified, the file name is "dog.jpg" (for example, a picture containing a dog). The "dog.jpg" image is input into the ResNet model. The ResNet model can extract the features of the image and perform calculations through the neural network structure inside the ResNet model. Ultimately, the ResNet model outputs a probability distribution indicating the probability that the image belongs to each category. For example, in the probability distribution output by the ResNet model, the probability of the "dog" category is the highest (for example, 98%). Then, "dog" is the first classification label.
[0050] Step S103 , in response to the first classification label being different from the original label, determining a first confidence of the image to be classified in different image categories through the first image model, and determining a second confidence of the image to be classified in different image categories through the second image model.
[0051] Here, the second image model is an image classification model to be trained.
[0052] In an embodiment of the present application, when the first classification label differs from the original label, the image to be classified can be classified using a first image model to obtain a first confidence level for the image to be classified under different image categories, and the image to be classified can be classified using a second image model to obtain a second confidence level for the image to be classified under different image categories. The first classification label is the category label predicted by the first image model. If the first classification label differs from the original label, it may indicate that the original label of the image to be classified may be inaccurate, that is, the original label does not match the image information.
[0053] For example, the original label of an image is "dog". When the first image model classifies the image to be classified, it can output the confidence of each category. The confidence can represent the probability that the image predicted by the image model when classifying the image is the classification result. It is usually a value between 0 and 1, and the category corresponding to the maximum confidence value is the classification label. For example, the first image model can classify the three categories of "cat", "dog" and "car". When the first image model classifies the image to be classified, the confidence obtained can be 0.5 for "cat", 0.4 for "dog" and 0.1 for "car", that is, the first image model predicts that the image to be classified has a 50% probability of being a "cat", a 40% probability of being a "dog", and only a 10% probability of being a "car".
[0054] Similarly, the second image model will also classify the image to be classified and output a classification label and confidence level (i.e., the second confidence level). The second image model is an image classification model to be trained, so the prediction results of the second image model may be different from those of the first image model. For example, the second image model can classify the three categories of "cat", "dog", and "car" and obtain confidence levels of 0.65 for "cat", 0.25 for "dog", and 0.1 for "car".
[0055] Step S104: performing image classification on the image to be classified based on the first confidence level and the second confidence level to obtain an image classification result.
[0056] In the embodiment of the present application, the image to be classified can be classified according to the first confidence level and the second confidence level to obtain an image classification result. It should be noted that the image classification here refers to a classification task that classifies the entire image.
[0057] According to the embodiment of the present application, when performing image classification on an image to be classified, first, the image to be classified is classified by using a pre-trained image classification model to obtain a first classification label for the image to be classified. When the first classification label is different from the original label of the image to be classified, the first confidence and second confidence of the image to be classified under different image categories are further determined, thereby further classifying the image to be classified based on the first confidence and second confidence. In other words, by comparing whether the first classification label is the same as the original label, a preliminary rough classification of the image to be classified can be performed. Afterwards, when the first classification label is different from the original label, the first confidence and second confidence of the image to be classified under different image categories can be determined by using the pre-trained image classification model and the image classification model to be trained respectively. Finally, based on the first confidence and the second confidence, the image to be classified is classified to obtain an image classification result. In this way, since the pre-trained image classification model can provide more reliable confidence information, when further image classification is performed based on the confidence determined by the two models, the first confidence obtained by the pre-trained image classification model can be used as a reference indicator, so that based on this reference indicator, a more accurate classification result of the image to be classified can be determined, thereby improving the accuracy of image classification.
[0058] Figure 2 This is another optional flow chart of the image classification method provided in the embodiment of the present application, such as Figure 2 As shown, the method includes the following steps S201 to S210:
[0059] Step S201: The terminal receives an image classification operation input by a user.
[0060] The image classification operation input by the user includes a selection operation or an input operation. The selection operation is used to select an image to be classified, or the input operation is used to input an image identifier of the image to be classified.
[0061] In step S202 , the terminal encapsulates the image identifier of the image to be classified into an image classification request.
[0062] In the embodiment of the present application, the image classification request is used to request the server to perform image classification on the image to be classified.
[0063] Step S203: The terminal sends an image classification request to the server.
[0064] In the embodiment of the present application, the terminal sends an image classification request to the server to request the server to perform image classification on the image to be classified.
[0065] Step S204: The server obtains the image to be classified and the original label of the image to be classified in response to the image classification request.
[0066] In step S205 , the server classifies the image to be classified using the first image model to obtain a first classification label.
[0067] It should be noted that steps S204 to S205 are the same as the above-mentioned steps S101 to S102, and the implementation details of steps S204 to S205 are not repeated in this embodiment of the application.
[0068] Step S206 : In response to the first classification label being the same as the original label, the server divides the image to be classified into a third image set.
[0069] Here, the third image set includes images whose original labels match the image information.
[0070] In this embodiment of the present application, when the first classification label is the same as the original label, it can be considered that the original label of the image to be classified matches the image information, and the image to be classified is classified into the third image set. If the first classification label is the same as the original label, it can be indicated that the original label of the image to be classified is accurate, that is, the original label matches the image information.
[0071] For example, the first image model predicts that an image is "dog." The first classification label is "dog." This first classification label "dog" is the same as the original label "dog." This indicates that the image's original label matches the image content, and the image can be classified into the third image set.
[0072] Through step S206, by verifying whether the first classification label is the same as the original label, it can be verified whether the label of the image matches the image content. If the label of the image matches the image content, the image with the matching label and image content can be divided into a third image set. The third image set can be used for subsequent high-quality dataset construction, model training or verification tasks, which helps to improve dataset quality and model performance.
[0073] In step S207 , in response to the first classification label being different from the original label, the server determines a first confidence of the image to be classified under different image categories through the first image model, and determines a second confidence of the image to be classified under different image categories through the second image model.
[0074] Here, the first image model is a pre-trained image classification model, and the second image model is an image classification model to be trained.
[0075] In some embodiments, see Figure 3 , Figure 3 It is shown that the step S207 of "determining the first confidence of the image to be classified under different image categories by using the first image model" can be implemented by the following steps S2071 to S2073:
[0076] Step S2071: performing image transformation processing on the image to be classified to obtain at least one transformed image.
[0077] In the embodiments of the present application, image transformation may refer to performing a series of processing operations on an image to generate a new image. Image transformation methods may include: geometric transformation (such as rotation, scaling, translation, flipping, etc.), color transformation (such as adjusting brightness, contrast, saturation, etc.), noise addition (such as Gaussian noise, salt and pepper noise, etc.), and filtering operations (such as Gaussian blur, sharpening, etc.). Image transformation can increase the diversity of images.
[0078] For example, if the image to be classified is a dog image (the original image), the following transformations can be performed: rotating the image 45 degrees, reducing the brightness of the image, and adding salt and pepper noise to the image. Through these transformations, three transformed images are obtained: the rotated image "dog01.jpg", the reduced brightness image "dog02.jpg", and the noisy image "dog03.jpg".
[0079] Step S2072: Aggregate the image to be classified and at least one transformed image to obtain a first image set.
[0080] In the embodiment of the present application, the original image to be classified and all transformed images are aggregated to form an image set. This set can be referred to as a first image set. The first image set includes the original image and all transformed images.
[0081] For example, a first set of images may include "dog.jpg," "dog01.jpg," "dog02.jpg," and "dog03.jpg."
[0082] Step S2073: Determine first confidences of images in the first image set under different image categories using the first image model.
[0083] In the embodiment of the present application, the first image model is a pre-trained image classification model, such as a ResNet, MobileNet, etc. The first image model can classify the input image and output the confidence level corresponding to each category.
[0084] For example, a first image model can classify the three categories of "cat," "dog," and "car." When the first image model classifies the to-be-classified image "dog.jpg," the confidence scores obtained may be 0.5 for "cat," 0.4 for "dog," and 0.1 for "car." When the first image model classifies the transformed image "dog01.jpg," the confidence scores obtained may be 0.6 for "cat," 0.3 for "dog," and 0.1 for "car." When the first image model classifies the transformed image "dog02.jpg," the confidence scores obtained may be 0.6 for "cat," 0.4 for "dog," and 0.0 for "car." When the first image model classifies the transformed image "dog03.jpg," the confidence scores obtained may be 0.5 for "cat," 0.4 for "dog," and 0.1 for "car." The confidence information for all images can be summarized to form a confidence matrix or list. Afterwards, the maximum confidence corresponding to each category can be used as the first confidence of the images in the first image set under the image category; the mean of the confidence corresponding to each category can also be used as the first confidence of the images in the first image set under the image category; the weighted sum of the confidence corresponding to each category can also be used as the first confidence of the images in the first image set under the image category. The embodiment of the present application does not limit this, and the method of determining the first confidence of the images in the first image set under the image category can be selected according to actual conditions.
[0085] In some embodiments, step S2073 can be implemented by performing the following processing: first, for each image category, the prediction confidence of each image in the first image set under the image category is determined through the first image model; then, the prediction confidence of all images in the first image set is averaged to obtain the confidence mean; finally, the confidence mean is determined as the first confidence of the images in the first image set under the image category.
[0086] In an embodiment of the present application, first, for each image category, each image in the first image set is classified by the first image model, and the prediction confidence of each image in each image category can be obtained. For example, the prediction confidence of the original image "dog.jpg" is 0.5 for "cat", 0.4 for "dog", and 0.1 for "car". The prediction confidence of the transformed image 1 "dog01.jpg" is 0.6 for "cat", 0.3 for "dog", and 0.1 for "car". The prediction confidence of the transformed image 2 "dog02.jpg" is 0.6 for "cat", 0.4 for "dog", and 0.0 for "car". The prediction confidence of the transformed image 3 "dog03.jpg" is 0.5 for "cat", 0.4 for "dog", and 0.1 for "car". Then, the prediction confidence of all images in the first image set is averaged to obtain the confidence mean. The prediction confidence of each category can be collected and analyzed. For the "cat" class: the prediction confidence of "dog.jpg" is 0.5, the prediction confidence of "dog01.jpg" is 0.6, the prediction confidence of "dog02.jpg" is 0.6, and the prediction confidence of "dog03.jpg" is 0.5. For the "dog" class: the prediction confidence of "dog.jpg" is 0.4, the prediction confidence of "dog01.jpg" is 0.3, the prediction confidence of "dog02.jpg" is 0.4, and the prediction confidence of "dog03.jpg" is 0.4. For the "car" category: the prediction confidence of "dog.jpg" is 0.1, the prediction confidence of "dog01.jpg" is 0.1, the prediction confidence of "dog02.jpg" is 0.0, and the prediction confidence of "dog03.jpg" is 0.0. Next, calculate the mean confidence of each category, for the "cat" category: confidence mean = (0.5+0.6+0.6+0.5) / 4 = 0.55. For the "dog" category: confidence Mean = (0.4 + 0.3 + 0.4 + 0.4) / 4 = 0.375. For the "car" class: Mean confidence = (0.1 + 0.1 + 0.0 + 0.1) / 4 = 0.075. Finally, the first confidence for the images in the first image set in the "cat" class is 0.55. The first confidence for the images in the first image set in the "dog" class is 0.375. The first confidence for the images in the first image set in the "car" class is 0.075.
[0087] Through the above processing, transformed images can be introduced, which increases the diversity of image data and enables the first image model to evaluate the classification results of the image from multiple angles, thereby improving the accuracy of classification. By comprehensively considering the classification results of the original image and the transformed image, accidental errors are reduced, the adaptability and reliability of the first image model in the face of different scenes are enhanced, and the classification accuracy of the first image model is improved.
[0088] Through steps S2071 to S2073, the image to be classified can be transformed to generate a diversified image set (first image set), which increases the diversity of image data. The pre-trained first image model is used to classify the image, and the confidence level of each category is calculated. The final first confidence level is determined through the confidence information of the original image and the transformed image, thereby improving the reliability of the classification results of the first image model.
[0089] In step S208 , the server performs image classification on the image to be classified based on the first confidence level and the second confidence level to obtain an image classification result.
[0090] In some embodiments, see Figure 4 , Figure 4 It shows that step S208 can be implemented by the following steps S2081 to S2083:
[0091] Step S2081: Determine the maximum first confidence level of the image to be classified under different image categories.
[0092] For example, the prediction confidence of each category, for the "cat" category: the prediction confidence of "dog.jpg" is 0.5, the prediction confidence of "dog01.jpg" is 0.6, the prediction confidence of "dog02.jpg" is 0.6, and the prediction confidence of "dog03.jpg" is 0.5. For the "dog" category: the prediction confidence of "dog.jpg" is 0.4, the prediction confidence of "dog01.jpg" is 0.3, the prediction confidence of "dog02.jpg" is 0.4, and the prediction confidence of "dog03.jpg" is 0.4. For the "car" category: the prediction confidence of "dog.jpg" is 0.1, the prediction confidence of "dog01.jpg" is 0.1, the prediction confidence of "dog02.jpg" is 0.0, and the prediction confidence of "dog03.jpg" is 0.0. The maximum first confidence is 0.6.
[0093] Step S2082: Determine the maximum second confidence level of the image to be classified under different image categories.
[0094] For example, the second confidence level may be 0.65 for "cat", 0.25 for "dog", and 0.1 for "car". The maximum second confidence level is 0.65.
[0095] Step S2083: Based on the maximum first confidence level and the maximum second confidence level, perform image classification on the image to be classified to obtain an image classification result.
[0096] In some embodiments, step S2083 can be implemented by performing the following processing: first, in response to the maximum first confidence being less than a preset first confidence threshold, and the maximum second confidence being less than a preset second confidence threshold, the image to be classified is divided into a first image set; the first image set includes images whose original labels match the image information; second, in response to the maximum first confidence being greater than or equal to the preset first confidence threshold, and / or the maximum second confidence being greater than or equal to the preset second confidence threshold, the first image features output by the first data processing layer of the first image model and the second image features output by the second data processing layer of the second image model are obtained; finally, the image to be classified is classified based on the first image features and the second image features to obtain an image classification result.
[0097] In the embodiment of the present application, a first confidence threshold can be pre-set. The first confidence threshold can be set according to actual conditions, and the embodiment of the present application does not limit this. Since the first image model is a pre-trained image classification model, in order to evaluate the accuracy of the first image model classification, a fixed first confidence threshold can be set. If the maximum first confidence is greater than the first confidence threshold, it can be considered that the accuracy of the first image model classification is high. If the maximum first confidence is less than or equal to the first confidence threshold, it can be considered that the accuracy of the first image model classification is low.
[0098] The second confidence threshold can be set in advance, and the second confidence threshold can be a fixed value or a variable value. When the second confidence threshold is a variable value, the current number of training rounds of the second image model and the preset first hyperparameter can be obtained first; then, based on the current number of training rounds and the first hyperparameter, the second confidence threshold at the current moment is determined. For example, the first hyperparameter can include an initial threshold and a decay rate. The current number of training rounds, the initial threshold and the decay rate can be used to construct a calculation formula for the second confidence threshold. As the number of training rounds changes, the second confidence threshold will also change, thereby determining the second confidence threshold corresponding to the current round. By dynamically adjusting the confidence threshold, the screening process of the classification results can be more flexibly controlled, thereby improving the accuracy and reliability of the classification results in subsequent applications.
[0099] When the maximum first confidence is less than the preset first confidence threshold, and the maximum second confidence is less than the preset second confidence threshold, the image to be classified can be considered to have complex image semantics and the first image model cannot accurately classify the objects in the image. Therefore, the image to be classified can be considered to be an image in which the original label matches the image information, and the image to be classified can be classified into the first image set. For example, the maximum first confidence is 0.6, and the maximum second confidence is 0.65. The first confidence threshold and the second confidence threshold are both 0.8. The maximum first confidence of 0.6 is less than the first confidence threshold of 0.8, and the maximum second confidence of 0.65 is less than the second confidence threshold of 0.8. Therefore, the image to be classified "dog.jpg" can be classified into the first image set.
[0100] When the maximum first confidence is greater than or equal to a preset first confidence threshold, and / or the maximum second confidence is greater than or equal to a preset second confidence threshold, the first image feature output by the first data processing layer of the first image model and the second image feature output by the second data processing layer of the second image model can be obtained. For example, the maximum first confidence is 0.6, and the maximum second confidence is 0.65. The first confidence threshold and the second confidence threshold are both 0.6. If the maximum first confidence of 0.6 is equal to the first confidence threshold of 0.6, and the maximum second confidence of 0.65 is greater than the second confidence threshold of 0.6, the first image feature output by the first data processing layer of the first image model and the second image feature output by the second data processing layer of the second image model can be obtained. The first data processing layer can be a layer in the first image model, such as a convolutional layer or a feature extraction layer. The second data processing layer can be a layer in the second image model, such as a convolutional layer or a feature extraction layer. The first image feature is a feature output by the first data processing layer of the first image model. The second image feature is a feature output by the second data processing layer of the second image model.
[0101] After obtaining the first image feature and the second image feature, the image to be classified can be classified based on the first image feature and the second image feature to obtain an image classification result. First, the feature similarity between the first image feature and the second image feature can be determined; then, in response to the feature similarity being greater than a preset similarity threshold, the image to be classified is classified into a first image set; or, in response to the feature similarity being less than or equal to the preset similarity threshold, the image to be classified is classified into a second image set; the second image set includes images whose original labels do not match the image information.
[0102] In an embodiment of the present application, a similarity threshold can be set in advance, and the similarity threshold can be a fixed value or a variable value. When the similarity threshold is a variable value, the current number of training rounds of the second image model and the preset second hyperparameter can be obtained; based on the current number of training rounds and the second hyperparameter, the similarity threshold at the current moment is determined. For example, the second hyperparameter may include an initial threshold and a decay rate. The current number of training rounds, the initial threshold and the decay rate can be used to construct a calculation formula for the similarity threshold. As the number of training rounds changes, the similarity threshold will also change, thereby determining the similarity threshold corresponding to the current round. By dynamically adjusting the similarity threshold, the screening process of the classification results can be more flexibly controlled, thereby improving the accuracy and reliability of the classification results in subsequent applications.
[0103] There are many ways to determine the feature similarity between the first image feature and the second image feature: calculating the cosine similarity between the first image feature and the second image feature as the feature similarity, calculating the Euclidean distance between the first image feature and the second image feature as the feature similarity, calculating the Manhattan distance between the first image feature and the second image feature as the feature similarity, and calculating the Hamming distance between the first image feature and the second image feature as the feature similarity, etc. You can choose according to actual conditions, and the embodiments of the present application are not limited to this.
[0104] When the feature similarity is greater than a preset similarity threshold, the image to be classified can be classified into the first image set. Alternatively, when the feature similarity is less than or equal to the preset similarity threshold, the image to be classified can be classified into the second image set; the second image set includes images whose original labels do not match the image information.
[0105] Through the above processing, two confidence thresholds can be set to screen and process the image classification results. When the maximum confidence of the two models is lower than their respective confidence thresholds, it can be considered that the original label of the image matches the content, and the image is divided into the first image set; when at least one confidence is higher than the threshold, the features of the two models are further extracted and the similarity is calculated. When the similarity is greater than the similarity threshold, it can be considered that the original label of the image matches the content, and the image is divided into the first image set. Alternatively, when the similarity is less than or equal to the similarity threshold, it can be considered that the original label of the image does not match the content, and the image is divided into the second image set. The division method based on multi-condition judgment can achieve fine-grained classification of images, effectively improve the accuracy and reliability of image classification, and by dynamically adjusting the similarity threshold and confidence threshold, the screening process of the classification results can be more flexibly controlled, thereby improving the accuracy and reliability of the classification results in subsequent applications.
[0106] Through steps S2081 to S2083, the maximum confidence levels predicted by the two models for the image to be classified can be determined first, and then a first confidence threshold and a second confidence threshold can be set to filter and process the image classification results. When the maximum confidence levels of the two models are both lower than their respective confidence thresholds, it can be assumed that the original label of the image matches the content, and the image is classified into the first image set. When at least one confidence level is equal to or higher than the threshold, features of the two models are further extracted and further classified. By dynamically adjusting the confidence threshold and similarity threshold, images whose labels may not match the image information can be effectively filtered out, and feature similarity is calculated to improve the accuracy and reliability of classification.
[0107] Step S209: The server sends the image classification result to the terminal.
[0108] Step S210: The terminal displays the image classification result.
[0109] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0110] An embodiment of the present application proposes an image classification method, which can be based on image noise processing technology with external guidance (a large image model or a first image model) and a sample separation strategy to solve the problem of noise labels affecting the performance of image classification models in image classification tasks, thereby improving the accuracy and robustness of image classification models.
[0111] See also Figure 5 , Figure 5 This is a schematic diagram of the principle of the image classification method provided in the embodiment of the present application. First, a suitable large-scale image model (i.e., the first image model) is selected. Depending on the specific image task and the characteristics of the image dataset, a deep learning model pre-trained on a large-scale image dataset can be selected. The deep learning model can have image feature extraction capabilities, which can provide a basis for subsequent sample analysis and model fine-tuning.
[0112] After selecting a suitable large-scale image model, the confidence and features are obtained. The image in the image data 501 can be input as a sample image (i.e., the image to be classified) into the selected large-scale image model 502 to obtain the predicted probability of the large-scale image model for the sample image belonging to each category (i.e., different image categories), and the predicted probability of each category is used together as the confidence value of the sample image. At the same time, the image features output by the large-scale image model in the middle layer can be extracted, i.e., the intermediate features 508 of the large-scale image model (i.e., the first image features). The intermediate features of the large-scale image model contain rich semantic and structural information of the sample image, which can be used for subsequent sample separation and learning strategy design.
[0113] After obtaining the confidence and features, coarse-grained separation is performed. The predicted label 504 (i.e., the first classification label) predicted by the large image model for the sample image can be compared with the original label 503 (i.e., the original label) of the sample image to see if they are consistent, and the sample images can be divided into an easy-to-clean set (a possible clean sample set, i.e., the third image set) and an inconsistent set. Based on the consistency between the category predicted by the large image model and the original annotation, the sample images are preliminarily divided into two categories. For example, if the category predicted by the large image model is consistent with the original annotation, the sample image is regarded as a possible clean sample (which can be included in the easy-to-clean set); if not, it is regarded as a possible noise sample (which can be included in the inconsistent set).
[0114] After coarse-grained separation, data augmentation is performed to enhance the confidence of large-scale image model generation. To make the confidence of large-scale image model generation more robust, M different augmentation methods (i.e., image transformation processing, such as geometric transformation, color transformation, Gaussian noise, and cropping) can be applied to each input image while keeping the semantics of the augmented samples unchanged, resulting in M augmented samples (i.e., at least one transformed image).
[0115] By predicting M types of enhanced samples separately through a large image model, M confidence levels can be obtained. (i.e., prediction confidence), and the confidence Aggregate into a comprehensive confidence 506 (i.e., the first confidence level) is to calculate the comprehensive confidence level for each category. For details, see formula (1):
[0116]
[0117] in, is the confidence obtained for the large image model, is the confidence score generated by the large image model for each augmented sample, and M is the number of augmented samples.
[0118] Figure 5 The confidence level p505 (i.e., the second confidence level) in is the confidence level obtained by the training model when predicting the sample image.
[0119] After data augmentation, fine-grained separation is performed. For the resulting set of possible noise samples (inconsistent set), the dynamic information of the training image model during training (such as changes in predicted probability) and other characteristic information of the sample images (such as texture and color distribution) are further utilized to design more refined separation rules, dividing the sample images into a hard-cleaned set (which may be clean samples mistakenly identified as noise, i.e., the first image set) or a true noise set (i.e., the second image set), thereby more accurately identifying different types of samples.
[0120] After fine-grained separation, threshold setting can be performed. For the confidence generated by large image models, a fixed confidence threshold can be used. (i.e., the first confidence threshold) is used for screening. As for the confidence generated by the training model, since the training model will change during the fine-tuning process, an adaptive confidence threshold τ(t) (i.e., the second confidence threshold) can be used. This threshold increases as the fine-tuning progresses, but will not be too high. The calculation formula of the adaptive confidence threshold τ(t) is shown in formula (2):
[0121] τ(t)=τ-exp(-ωt) (2)
[0122] Among them, τ and ω are hyperparameters that control the threshold, which can be set to τ = 0.6 and ω = 0.5, and t represents the number of rounds of current fine-tuning.
[0123] In addition, a feature similarity threshold Φ(t) (i.e., similarity threshold) can also be set. The calculation formula of the feature similarity threshold Φ(t) is shown in formula (3):
[0124] Φ(t)=Φ-exp(-ωt) (3)
[0125] Where Φ and ω are hyperparameters that control the threshold, and t represents the number of rounds of current fine-tuning.
[0126] When using a fixed confidence threshold After screening by the first confidence threshold, the hard cleaning set and the true noise set can be selected. The sample images with max(p)<τ(t) are taken as the hard cleaning set. The sample images in the hard cleaning set show lower confidence in the predictions of both the large image model and the training model, which may be because the sample images in the hard cleaning set have semantic difficulty, but are considered to be hard cleaning samples, that is, although difficult to classify, the original label is correct. In addition, the similarity between the intermediate feature 507 (i.e., the second image feature) of the training model and the intermediate feature 508 of the large image model is calculated. When the similarity is greater than the similarity threshold Φ(t), the sample image can be considered to be a hard cleaning sample. When the similarity is less than or equal to the similarity threshold Φ(t), the sample can be classified as the true noise set. The sample images in the true noise set show higher confidence in the predictions of the large image model and the training model, but because the sample images in the true noise set are inconsistently predicted in the coarse-grained separation, they are considered to be noise samples.
[0127] The embodiment of the present application can filter out sample data with noisy labels (such as incorrect or inaccurate labeling) in the training data set, and then correct the label information of the sample data, and use the corrected sample data to retrain the image classification model to be trained, thereby improving the performance of the image classification model.
[0128] Based on the image classification method described in the above embodiment, Figure 6 A structural block diagram of an image classification device provided in an embodiment of the present application is shown. The image classification device 100 can be a device in an electronic device (for example, a server). The image classification device can be implemented in software, which can be software in the form of programs and plug-ins, etc., including the following software modules: a data acquisition module 101, a first image classification module 102, a confidence determination module 103 and a second image classification module 104. These modules are logical and can therefore be arbitrarily combined or further split according to the functions implemented.
[0129] Among them, the data acquisition module 101 is used to obtain the image to be classified and the original label of the image to be classified; the first image classification module 102 is used to perform image classification on the image to be classified through a first image model to obtain a first classification label; the confidence determination module 103 is used to determine the first confidence of the image to be classified under different image categories through the first image model in response to the first classification label being different from the original label, and to determine the second confidence of the image to be classified under the different image categories through the second image model; the first image model is a pre-trained image classification model, and the second image model is an image classification model to be trained; the second image classification module 104 is used to perform image classification on the image to be classified based on the first confidence and the second confidence to obtain an image classification result.
[0130] In some embodiments, the confidence determination module 103 is further used to perform image transformation processing on the image to be classified to obtain at least one transformed image; summarize the image to be classified and the at least one transformed image to obtain a first image set; and determine the first confidence of the images in the first image set under the different image categories through the first image model.
[0131] In some embodiments, the confidence determination module 103 is further used to determine, for each of the image categories, the prediction confidence of each image in the first image set under the image category through the first image model; average the prediction confidences of all images in the first image set to obtain a confidence mean; and determine the confidence mean as the first confidence of the images in the first image set under the image category.
[0132] In some embodiments, the second image classification module 104 is further used to determine the maximum first confidence of the image to be classified under the different image categories; determine the maximum second confidence of the image to be classified under the different image categories; and perform image classification on the image to be classified based on the maximum first confidence and the maximum second confidence to obtain an image classification result.
[0133] In some embodiments, the second image classification module 104 is further used to, in response to the maximum first confidence being less than a preset first confidence threshold, and the maximum second confidence being less than a preset second confidence threshold, divide the image to be classified into a first image set; the first image set includes images whose original labels match the image information; in response to the maximum first confidence being greater than or equal to the preset first confidence threshold, and / or the maximum second confidence being greater than or equal to the preset second confidence threshold, obtain the first image features output by the first data processing layer of the first image model and the second image features output by the second data processing layer of the second image model; perform image classification on the image to be classified based on the first image features and the second image features to obtain an image classification result.
[0134] In some embodiments, the second image classification module 104 is further used to determine the feature similarity between the first image feature and the second image feature; in response to the feature similarity being greater than a preset similarity threshold, the image to be classified is divided into the first image set; in response to the feature similarity being less than or equal to a preset similarity threshold, the image to be classified is divided into the second image set; the second image set includes images whose original labels do not match the image information.
[0135] In some embodiments, the second image classification module 104 is further used to obtain the current number of training rounds, a preset first hyperparameter, and a preset second hyperparameter of the second image model; based on the current number of training rounds and the first hyperparameter, determine the second confidence threshold at the current moment; based on the current number of training rounds and the second hyperparameter, determine the similarity threshold at the current moment.
[0136] In some embodiments, the first image classification module 102 is further configured to, in response to the first classification label being the same as the original label, divide the image to be classified into a third image set; the third image set includes images whose original labels match the image information.
[0137] It should be noted that the description of the device embodiment of the present application is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment, so it will not be repeated. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present application for understanding.
[0138] An embodiment of the present application provides an electronic device, Figure 7 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 7 As shown, the electronic device 130 includes: at least one processor 131 ( Figure 7 Only one is shown in the figure), a memory 132, and computer executable instructions 133 stored in the memory 132 and executable on at least one processor 131. When the processor 131 executes the computer executable instructions 133, the steps in any of the above-mentioned image classification method embodiments are implemented.
[0139] The electronic device may include but is not limited to a processor 131 and a memory 132. It will be understood by those skilled in the art that Figure 7 This is merely an example of the electronic device 130 and does not constitute a limitation on the electronic device 130 . The electronic device 130 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0140] The processor 131 may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPG), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0141] In some embodiments, the memory 132 may be an internal storage unit of the electronic device 130, such as a hard disk or memory of the electronic device 130. In other embodiments, the memory 132 may also be an external storage device of the electronic device 130, such as a plug-in hard disk equipped on the electronic device 130, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card, etc. Furthermore, the memory 132 may include both an internal storage unit of the electronic device 130 and an external storage device. The memory 132 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory 132 may also be used to temporarily store data that has been output or is about to be output.
[0142] The present invention provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image classification method described in the present invention.
[0143] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the image classification method provided by the embodiment of the present application, for example, Figure 1 The image classification method shown.
[0144] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0145] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0146] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0147] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0148] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. An image classification method, characterized in that: The method comprises: Obtaining an image to be classified and an original label of the image to be classified; Performing image classification on the image to be classified using a first image model to obtain a first classification label; In response to the first classification label being different from the original label, determining a first confidence level of the image to be classified in a different image category using the first image model, and determining a second confidence level of the image to be classified in the different image category using a second image model; the first image model is a pre-trained image classification model, and the second image model is an image classification model to be trained; Based on the first confidence level and the second confidence level, image classification is performed on the image to be classified to obtain an image classification result.
2. The method according to claim 1, characterized in that The determining, by using the first image model, a first confidence level of the image to be classified in different image categories includes: Performing image transformation processing on the image to be classified to obtain at least one transformed image; Aggregating the image to be classified and the at least one transformed image to obtain a first image set; First confidences of images in the first image set under the different image categories are determined using the first image model.
3. The method according to claim 2, characterized in that Determining, by using the first image model, first confidences of images in the first image set under the different image categories includes: For each of the image categories, determining, by using the first image model, a prediction confidence score of each image in the first image set under the image category; averaging the prediction confidences of all images in the first image set to obtain a confidence mean; The confidence mean is determined as a first confidence of the images in the first image set under the image category.
4. The method according to claim 1, wherein The performing image classification on the image to be classified based on the first confidence level and the second confidence level to obtain an image classification result includes: Determining a maximum first confidence level of the image to be classified under the different image categories; Determining a maximum second confidence level of the image to be classified under the different image categories; Based on the maximum first confidence level and the maximum second confidence level, image classification is performed on the image to be classified to obtain an image classification result.
5. The method according to claim 4, characterized in that The performing image classification on the image to be classified based on the maximum first confidence level and the maximum second confidence level to obtain an image classification result includes: In response to the maximum first confidence being less than a preset first confidence threshold, and the maximum second confidence being less than a preset second confidence threshold, the image to be classified is divided into a first image set; the first image set includes images whose original labels match the image information; In response to the maximum first confidence being greater than or equal to a preset first confidence threshold, and / or the maximum second confidence being greater than or equal to a preset second confidence threshold, obtaining a first image feature output by a first data processing layer of the first image model and a second image feature output by a second data processing layer of the second image model; The image to be classified is classified based on the first image feature and the second image feature to obtain an image classification result.
6. The method according to claim 5, characterized in that The performing image classification on the image to be classified based on the first image feature and the second image feature to obtain an image classification result includes: determining a feature similarity between the first image feature and the second image feature; In response to the feature similarity being greater than a preset similarity threshold, classifying the image to be classified into the first image set; In response to the feature similarity being less than or equal to a preset similarity threshold, the image to be classified is divided into a second image set; the second image set includes images in which the original labels do not match the image information.
7. The method according to claim 6, characterized in that The method further comprises: Obtaining a current number of training rounds, a preset first hyperparameter, and a preset second hyperparameter of the second image model; Determining the second confidence threshold at a current moment based on the current number of training rounds and the first hyperparameter; The similarity threshold at a current moment is determined based on the current number of training rounds and the second hyperparameter.
8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: In response to the first classification label being the same as the original label, the image to be classified is divided into a third image set; the third image set includes images whose original labels match the image information.
9. An image classification device, characterized in that: The device comprises: A data acquisition module is used to acquire the image to be classified and the original label of the image to be classified; A first image classification module is used to classify the image to be classified using a first image model to obtain a first classification label; a confidence determination module configured to, in response to the first classification label being different from the original label, determine a first confidence of the image to be classified in a different image category using the first image model, and determine a second confidence of the image to be classified in the different image category using a second image model; the first image model is a pre-trained image classification model, and the second image model is an image classification model to be trained; The second image classification module is used to perform image classification on the image to be classified based on the first confidence level and the second confidence level to obtain an image classification result.
10. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the image classification method according to any one of claims 1 to 8 when executing the computer-executable instructions or computer programs stored in the memory.
11. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the image classification method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image classification device, image classification method, and image classification program
US20220343632A1