Image classification method and electronic equipment

By employing a dual classification mechanism on mobile devices, combining object detection and overall feature judgment in the image classification method, the problem of performance degradation in mobile devices when dealing with complex backgrounds or small foreground objects is solved, achieving efficient and accurate image classification.

CN120808011APending Publication Date: 2025-10-17LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900963.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

When performing high-precision image classification on mobile devices, existing technologies are limited by computing power and memory, making it difficult to effectively handle image classification tasks with complex backgrounds or small foreground objects, resulting in a decline in classification performance.

Method used

A dual classification mechanism is adopted. The first classification process detects and classifies target objects in the image, and the second classification process judges the overall image features. Through confidence mapping and label fusion, conflicting results are eliminated, achieving efficient and accurate image classification.

Benefits of technology

It improves the accuracy and robustness of image classification, especially in scenarios with complex backgrounds and small and diverse foreground targets. It can maintain efficient operation on mobile devices with limited resources and adapt to the local computing power limitations of mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808011A_ABST
    Figure CN120808011A_ABST
Patent Text Reader

Abstract

The invention discloses an image classification method and electronic equipment. The electronic equipment acquires first classification label information obtained by performing first classification processing on a target image; obtaining second classification label information obtained by performing second classification processing on the target image; verifying the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verifying the first classification label information based on the second classification label information to obtain a classification label information verification result; and fusing the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an image classification method and an electronic device. BACKGROUND

[0002] Image classification is one of the important tasks in the field of computer vision, and is widely used in album management, content recommendation and other scenarios. With the popularity of mobile devices and the increasing awareness of user data privacy protection, more and more applications tend to complete image processing tasks locally rather than relying on cloud services. However, mobile devices have obvious limitations in computing power and memory, which puts higher requirements on high-precision image classification.

[0003] In related technologies, in order to adapt to the hardware conditions of mobile terminals, the original image is usually compressed to a smaller size, and a multi-label classifier is used to identify multiple targets in the picture. Although this method can improve the running efficiency of the model to a certain extent, the classification performance is significantly reduced when dealing with complex backgrounds or small foreground objects. SUMMARY

[0004] The embodiments of the present application provide an image classification method and an electronic device.

[0005] The technical solutions of the embodiments of the present application are as follows:

[0006] In a first aspect, the embodiments of the present application provide an image classification method, comprising:

[0007] obtaining first classification label information obtained by performing first classification processing on a target image;

[0008] obtaining second classification label information obtained by performing second classification processing on the target image;

[0009] verifying the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verifying the first classification label information based on the second classification label information to obtain a classification label information verification result;

[0010] fusing the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.

[0011] In the above method, the first classification processing is used to determine the classification of the target object in the target image as a class in the first classification set;

[0012] The second classification processing is used to determine the classification of the target image as a class in the second classification set.

[0013] In the above method, obtaining the first classification label information obtained by performing first classification processing on the target image comprises:

[0014] performing same-class overlap box suppression processing on the candidate detection boxes in the image respectively used for marking target objects belonging to the categories to obtain first detection boxes corresponding to each category included in the first classification set;

[0015] performing cross-class overlap box suppression processing on the first detection boxes of each category to obtain second detection boxes of each category;

[0016] performing deduplication processing on the second detection boxes of each category to obtain target boxes of each category;

[0017] obtaining first classification label information based on the categories to which the target objects included in the target boxes belong.

[0018] In the above method, the second classification label information obtained by performing second classification processing on the target image comprises:

[0019] obtaining a feature vector corresponding to the target image;

[0020] determining the category to which the target image belongs in the second classification set based on the probability that the feature vector belongs to the vector space corresponding to each category in the second classification set, to obtain the second classification label information.

[0021] In the above method, the first classification label information and the second classification label information are fused based on the classification label information verification result to obtain target classification label information of the target image, comprising:

[0022] if the classification label information verification result represents that the first classification label information and the second classification label information exist mutually exclusive first label and second label, determining a target label included in the target classification label information corresponding to the target image based on a first confidence corresponding to the first label and a first confidence corresponding to the second label;

[0023] wherein the first label belongs to the first classification label information, the second label belongs to the second classification label information, and the target label is the first label or the second label.

[0024] In the above method, the target label included in the target classification label information corresponding to the target image is determined based on the first confidence corresponding to the first label and the first confidence corresponding to the second label, comprising:

[0025] mapping the first confidence to a third confidence corresponding to the second classification processing;

[0026] if the third confidence is higher than the second confidence, determining the first label as the target label included in the target classification label information corresponding to the target image;

[0027] If the third confidence is not higher than the second confidence, the second label is determined as the target label included in the target classification label information corresponding to the target image.

[0028] Or

[0029] The second confidence is mapped to a fourth confidence corresponding to the first classification processing.

[0030] If the first confidence is higher than the fourth confidence, the first label is determined as the target label included in the target classification label information corresponding to the target image.

[0031] If the first confidence is not higher than the fourth confidence, the second label is determined as the target label included in the target classification label information corresponding to the target image.

[0032] In the above method, the first confidence is mapped to a third confidence corresponding to the second classification processing, comprising:

[0033] The first confidence is converted according to a data form of the confidence corresponding to the second classification processing to obtain the third confidence; the data form of the third confidence is the same as that of the second confidence.

[0034] Correspondingly, the second confidence is mapped to a fourth confidence corresponding to the first classification processing, comprising:

[0035] The second confidence is converted according to a data form of the confidence corresponding to the first classification processing to obtain the fourth confidence; the data form of the fourth confidence is the same as that of the first confidence.

[0036] In the above method, the method further comprises:

[0037] If the target classification label information includes one label, the one label is label-extended according to a semantic relationship of the one label to obtain a plurality of extended labels.

[0038] In the above method, the method further comprises:

[0039] The image data in the album is subjected to the first classification processing and the second classification processing to obtain first classification label information and second classification label information of the image data.

[0040] The target classification label information of the image data is determined based on the first classification label information and the second classification label information, and the image data is classified and stored according to the target classification label information.

[0041] In a second aspect, an embodiment of the present application provides an electronic device, which comprises a processor and a memory storing processor-executable instructions; when the instructions are executed by the processor, the following steps are implemented:

[0042] obtain first classification label information obtained by performing first classification processing on the target image;

[0043] obtain second classification label information obtained by performing second classification processing on the target image;

[0044] verify the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verify the first classification label information based on the second classification label information to obtain a classification label information verification result;

[0045] fuse the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image. BRIEF DESCRIPTION OF DRAWINGS

[0046] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not intended to limit the scope of the present application.

[0047] Figure 1 An implementation flowchart of the image classification method proposed by the embodiments of the present application;

[0048] Figure 2 An implementation schematic of the image classification method proposed by the embodiments of the present application Figure One ;

[0049] Figure 3 An implementation schematic of the redundant detection frame removal proposed by the embodiments of the present application;

[0050] Figure 4 An implementation schematic of the label inference proposed by the embodiments of the present application;

[0051] Figure 5 An implementation schematic of the image classification method proposed by the embodiments of the present application Figure Two ;

[0052] Figure 6 An implementation schematic of the label extension proposed by the embodiments of the present application;

[0053] Figure 7 An implementation schematic of the electronic device proposed by the embodiments of the present application. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the related application, and not to limit the application. In addition, it should be noted that, for the convenience of description, only the parts related to the application are shown in the drawings.

[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application.

[0056] An embodiment of the present application provides an image classification method, as shown in the figure, the image classification method of the electronic device can include the following steps: Figure 1

[0057] Step 101, obtaining first classification label information obtained by performing first classification processing on a target image.

[0058] In the embodiments of the present application, the electronic device can obtain first classification label information obtained by performing first classification processing on a target image.

[0059] In the embodiments of the present application, the electronic device can be any device with communication, storage and computing functions, such as tablet computers, notebook computers, mobile phones, displays, all-in-one machines, game consoles and other electronic devices.

[0060] In the embodiments of the present application, the first classification processing is used to determine the classification of the target object in the target image into a category in the first classification set.

[0061] In the embodiments of the present application, the target object refers to an object or scene element with clear semantic or functional meaning in the image, such as a specific entity such as a person, an animal, an article, etc. in the target image; the first classification set can include specific and distinguishable object categories, such as cat, dog, flower, person, etc.

[0062] In the embodiments of the present application, the first classification processing can be understood as detecting and classifying the foreground target in the target image, that is, the electronic device can detect and classify the foreground object with clear boundaries in the target image, such as a person, an animal, an article, etc. through the first classification processing.

[0063] In the embodiments of the present application, the first classification processing can be implemented using a target detection model (such as YOLO, Faster R-CNN, etc.); through the first classification processing, the corresponding classification label information of each target object in the target image can be generated, and the first classification label information can be obtained according to the classification label information of each object; for example, in a picture containing three types of target objects of dog, flower and person, the target detection model will detect the three targets respectively and output the respective classification label information and confidence; this processing method can effectively extract complex multi-target information in the image and reduce the influence of background complexity on the classification result.

[0064] Step 102, obtaining second classification label information obtained by performing second classification processing on the target image. ​

[0065] In an embodiment of the present application, the electronic device can obtain second classification label information obtained by performing second classification processing on the target image.

[0066] In an embodiment of the present application, the second classification processing is used to determine the classification of the target image into a category in the second classification set.

[0067] In some embodiments of the present application, the classifier implementing the second classification processing can be different from the target detection model implementing the first classification processing.

[0068] In some embodiments of the present application, the same classifier can be used to implement the first classification processing and the second classification processing; including performing first classification processing on the target image according to the trained classifier, and then performing second classification processing on the target image using the trained classifier.

[0069] In an embodiment of the present application, the second classification processing can be a process of classifying the entire image based on the classifier; unlike the first classification processing, the second classification processing does not focus on the specific target position in the image, but judges the category to which the entire image belongs according to the global features of the entire image; for example, for a landscape picture, the second classification processing can classify it into the category of landscape, without separately identifying subcategories such as mountains, water or sky in it.

[0070] That is, in an embodiment of the present application, the first classification processing mainly identifies and classifies each specific target object in the image, while the second classification processing classifies based on the overall features of the entire image, and the two complement each other to form a dual mechanism for image classification; the result of the first classification processing can be used to refine the understanding of the image content, while the result of the second classification processing can help to quickly judge the general type of the image, and after the combination of the two, the system not only has stronger classification ability, but also can efficiently run on devices with limited resources.

[0071] As can be seen, by performing different classification processing on the image, including first classification processing and second classification processing, embodiments of the present application can obtain classification results of different dimensions; further based on the mutual verification between classification results of different dimensions, conflicting or inconsistent classification results can be effectively excluded, thereby improving the accuracy of classification; and by using the mutual verification relationship between classification results, the robustness of classification can be enhanced while maintaining the classification efficiency, which is particularly suitable for scenes with complex background, small and diversified foreground targets.

[0072] Step 103, verifying the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verifying the first classification label information based on the second classification label information to obtain a classification label information verification result.

[0073] In embodiments of the present application, after obtaining the first classification label information and the second classification label information, the electronic device can verify the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verify the first classification label information based on the second classification label information to obtain a classification label information verification result.

[0074] In some embodiments of the present application, the classification label information verification result can include labels of the same category existing in the first classification label information and the second classification label information, and labels of mutual exclusion existing in the first classification label information and the second classification label information.

[0075] In some embodiments of the present application, for the case that labels of the same category exist in the first classification label information and the second classification label information, the labels of the same category can be directly merged; and for the case that labels of mutual exclusion exist in the first classification label information and the second classification label information, the label with high confidence in the labels of mutual exclusion is used as a reference to evaluate whether the label with low confidence is reasonable, so as to determine the final target label.

[0076] In some embodiments of the present application, whether the labels in the first classification label information and the second classification label information are mutually exclusive can be determined based on a preset category mutual exclusion table; the preset category mutual exclusion table can be used to indicate categories that are mutually exclusive with each category.

[0077] Step 104: fusing the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.

[0078] In embodiments of the present application, after obtaining the classification label information verification result, the electronic device can fuse the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.

[0079] In embodiments of the present application, the target classification label information refers to a set of image classification results obtained after verifying and fusing two different classification results after the first classification processing of the foreground target detection and the second classification processing of the single-label classification processing; the target classification label information is usually represented in the form of a label, and each label corresponds to a specific category in the image, such as cat, dog, or beach.

[0080] In some embodiments of the present application, in addition to fusing different labels based on the confidence of the labels after obtaining the classification label information verification result, the confidence space of the labels in the first classification label information and the confidence space of the labels in the second classification label information can be aligned first, and then the confidence of each label in the aligned confidence space is used for verification of the classification label information.

[0081] In some embodiments of the present application, when the electronic device obtains the first classification label information obtained by performing the first classification processing on the target image, the following steps can be included:

[0082] Step 101a, performing same-class overlap box suppression processing on the candidate detection boxes in the image for marking the target objects belonging to the same class to obtain the first detection box corresponding to each class included in the first classification set.

[0083] It should be noted that multiple candidate detection boxes of the same class can have partial overlap, which can introduce redundant information and affect the accuracy of the final classification result. Therefore, the optimal detection box in each class is selected through the same-class overlap box suppression processing to reduce the impact on the classification accuracy. The same-class overlap box suppression processing can be realized by calculating the Intersection over Union (IoU) of the overlapping region, retaining the detection box with the highest confidence, and suppressing the remaining detection boxes with high overlap. For example, when multiple candidate detection boxes of cats are detected in an image, only the one with the highest confidence and the smallest overlap with other candidate detection boxes is retained as the first detection box of the class. This can effectively reduce repeated detection and improve the precision and robustness of the model.

[0084] Step 101b, performing cross-class overlap box suppression processing on the first detection boxes of each class to obtain the second detection box of each class.

[0085] It should be noted that although the same-class overlap box suppression processing can remove redundant detection boxes of the same class, the detection boxes of different classes can still have cross-coverage, i.e., a foreground target can be detected by multiple classes at the same time. Therefore, to further improve the accuracy of the detection result, the cross-class overlap box suppression processing is performed, which considers the relevance between different classes. By setting a reasonable IoU threshold, the detection boxes with high overlap and low confidence with other classes are further removed to obtain the second detection box of each class. For example, if there is a large overlap between the detection boxes of dogs and cats, it is determined which one is more reasonable according to the context semantics, so as to retain the detection box with high confidence and more consistent semantics. This can improve the overall consistency of the detection result without sacrificing the richness of the classes. In addition, the cross-class overlap box suppression processing not only depends on the geometric overlap degree, but also can make a comprehensive judgment based on the context semantics and the logical relationship between targets. For example, when the detection boxes of a car and a wheel overlap, the detection box of the car can be retained according to the structural relationship in the common scene, and the wheel is considered as one of its components, thereby avoiding misjudgment.

[0086] Step 101c, performing a deduplication processing on the second detection boxes of each category to obtain a target box of each category.

[0087] In the embodiments of the present application, the deduplication processing refers to further performing a secondary screening on the detection boxes within each category after completing the overlapping box suppression processing across categories, to ensure that only the most accurate detection box, i.e., the target box, is retained for each category.

[0088] Step 101d, obtaining first classification label information based on the category to which the target object contained in the target box belongs.

[0089] In the embodiments of the present application, after the overlapping box suppression processing within the same category, the overlapping box suppression processing across categories, and the deduplication processing, the features corresponding to each category can be extracted according to the finally determined target box, and the features of each category can be classified to obtain the first classification label information of the target image.

[0090] In the embodiments of the present application, when determining the first classification label information based on the target box, since each category has been simplified to an optimal detection box through multiple rounds of deduplication processing of detection boxes, the target detection model can directly perform efficient classification on a single target without needing to deal with complex multi-label or multi-category interference problems; for example, if a dog is contained in the finally retained target box of a certain category, the classifier will output the dog category and the corresponding confidence as part of the first classification label information; this way not only improves the classification efficiency, but also enhances the explainability and accuracy of the classification results.

[0091] As can be seen, through the hierarchical redundant detection box suppression, including the overlapping box suppression processing within the same category, the overlapping box suppression processing across categories, and the deduplication processing, the present application can effectively reduce redundant detection boxes, improve the accuracy and robustness of target detection, which can ensure the data quality of subsequent classification tasks, thereby improving the overall classification effect, and thus better adapting to the scene where the local computing power of mobile terminals is limited.

[0092] In some embodiments of the present application, when the electronic device obtains the second classification label information obtained by performing a second classification processing on the target image, the following steps can be included:

[0093] Step 102a, obtaining a feature vector corresponding to the target image.

[0094] In the embodiments of the present application, the feature vector refers to converting the target image into a set of numerical representations by a neural network or other feature extraction method, which is used to describe the main visual features of the target image. These features can include color histograms, edge information, texture features, semantic embeddings, etc. The feature vector can compress complex image information into a low-dimensional space, facilitating subsequent classification or retrieval operations. For example, in deep learning, the last layer output of a convolutional neural network (CNN) is usually used as the feature vector of the target image.

[0095] In step 102b, the classification to which the target image belongs in the second classification set is determined based on the probability of the feature vector belonging to the vector space corresponding to each classification in the second classification set, and the second classification label information is obtained.

[0096] In the embodiments of the present application, the second classification set can be understood as a pre-set category set, each category of which has its corresponding vector space distribution; the electronic device can calculate the probability distribution of the feature vector in the vector space of each classification, and select the classification with the highest feature vector belonging probability as the final classification result, thereby determining the second classification label information. For example, if the second classification set contains landscape, animal, and person classifications, the probability distribution of the target image belonging to the vector space of these three classifications can be calculated respectively, and the classification corresponding to the maximum probability is selected as the final label. Through the above method, pictures with complex backgrounds or blurred foregrounds in the album can be processed more effectively, and more intelligent localized image management and retrieval functions can be realized.

[0097] It should be noted that through the second classification processing, the optimal decision can be made among multiple possible classifications, improving the accuracy and robustness of the classification. Especially when facing multi-label, complex background or small foreground objects, the main classification can be more reliably identified.

[0098] As can be seen, by constructing a global feature vector and mapping it to the vector space of different classifications, the embodiments of the present application can realize semantic understanding of the entire image, thereby more comprehensively capturing the high-level features of the image. This method can supplement the background information or overall context that may be ignored in the image classification task, enhancing the completeness and diversity of the classification.

[0099] In some embodiments of the present application, when the electronic device fuses the first classification label information and the second classification label information based on the verification result of the classification label information to obtain the target classification label information of the target image, the following steps can be included:

[0100] Step 104a, if the classification label information verification result represents that the first classification label information and the second classification label information exist mutually exclusive first label and second label, determining a target label contained in the target classification label information corresponding to the target image based on the first confidence corresponding to the first label and the first confidence corresponding to the second label; wherein the first label belongs to the first classification label information, the second label belongs to the second classification label information, and the target label is the first label or the second label.

[0101] In the embodiments of the present application, the mutually exclusive first label and the second label represent that the first label and the second label logically conflict with each other, for example, the first label is "train", and the second label is "beach", and it is not likely that "train" appears on "beach", that is, the first label and the second label are mutually exclusive.

[0102] Exemplarily, the first label in the first classification label information is "train", the second label in the second classification label information is "beach", and the first label and the second label are mutually exclusive, so the target label can be determined based on the confidence corresponding to "train" and the confidence corresponding to "beach", and the target label is one of "train" and "beach".

[0103] As can be seen, when there is a contradiction between the classification labels, that is, there is mutual exclusion, the embodiments of the present application can effectively solve the classification conflict problem by introducing the confidence as a decision basis; the rationality and credibility of the classification result is ensured by comparing the classification results from different sources and their confidences, and the reliability of the classification result is improved.

[0104] In some embodiments of the present application, when the electronic device determines the target label contained in the target classification label information corresponding to the target image based on the first confidence corresponding to the first label and the first confidence corresponding to the second label, the following steps can be included:

[0105] Step 201, mapping the first confidence to a third confidence corresponding to the second classification processing.

[0106] In the embodiments of the present application, the confidence mapping is a technical means for converting confidence values of different data types to the same comparable range.

[0107] Step 202, if the third confidence is higher than the second confidence, determining the first label as the target label contained in the target classification label information corresponding to the target image.

[0108] In the embodiments of the present application, it is assumed that the confidence corresponding to "train" is 81%, and the confidence corresponding to "beach" after conversion is 66%, and since the confidence corresponding to "train" is higher than the confidence corresponding to "beach" after conversion, "train" can be determined as the target label.

[0109] Step 203: If the third confidence is not higher than the second confidence, the second label is determined as the target label included in the target classification label information corresponding to the target image.

[0110] It can be understood that when the third confidence after mapping is greater than the first confidence output by the classifier corresponding to the second classification processing, it indicates that the first classification processing has a higher confidence degree for the category, and at this time, the first label can be preferentially adopted as the final classification result. If the first confidence corresponding to the first classification processing is low, it indicates that the recognition result can be unreliable, and at this time, the second label output by the second classification processing is used as the final classification result. This strategy can effectively improve the classification accuracy, and in particular in the case of complex background and multiple targets, it can avoid the error judgment of the classifier due to the background interference.

[0111] Step 204: The second confidence is mapped to a fourth confidence corresponding to the first classification processing.

[0112] In some embodiments of the present application, in addition to mapping the first confidence to the third confidence corresponding to the second classification processing to compare the confidences based on the third confidence and determine the target label, the electronic device can also map the second confidence to the fourth confidence corresponding to the first classification processing to compare the confidences based on the fourth confidence and determine the target label.

[0113] Step 205: If the first confidence is higher than the fourth confidence, the first label is determined as the target label included in the target classification label information corresponding to the target image.

[0114] Step 206: If the first confidence is not higher than the fourth confidence, the second label is determined as the target label included in the target classification label information corresponding to the target image.

[0115] As can be seen, by mapping the confidences corresponding to different classification processes into a unified evaluation space, the embodiments of the present application can realize the comparison and fusion of the results of different classification processes. This strategy can effectively solve the problem of inconsistent dimensions of classification results, improve the consistency and stability of classification results, and is particularly suitable for multi-modal classification tasks, and improves the applicability of image classification methods.

[0116] In some embodiments of the present application, when the electronic device maps the first confidence to the third confidence corresponding to the second classification processing, the following steps can be included:

[0117] Step 201a: The first confidence is converted according to the data form of the confidence corresponding to the second classification processing to obtain the third confidence. The data form of the third confidence is the same as that of the second confidence.

[0118] Correspondingly, in some embodiments of the present application, the electronic device can include the following steps when mapping the second confidence to the fourth confidence corresponding to the first classification processing:

[0119] Step 204a, converting the second confidence according to the data form of the confidence corresponding to the first classification processing to obtain the fourth confidence; the data form of the fourth confidence is the same as that of the first confidence.

[0120] For example, the confidence corresponding to "train" is 81%, and the confidence corresponding to "beach" is 6.6. To unify the data types of the confidences of the two labels, the confidence 6.6 corresponding to "beach" can be converted into a value of the same data form as 81%, for example, the third confidence obtained after conversion is 66%; or the confidence 81% can be converted into a value of the same data form as 6.6, for example, the fourth confidence obtained after conversion is 8.1.

[0121] In some embodiments of the present application, the image classification method of the electronic device can further include the following steps:

[0122] Step 105, if the target classification label information includes one label, performing label expansion on the one label according to the semantic relationship of the one label to obtain a plurality of expanded labels.

[0123] In embodiments of the present application, when the target classification label information only contains one label, it means that the image is preliminarily judged to belong to a single category, but in order to more comprehensively describe the image content, semantic expansion can be performed based on the one label in the target classification label information.

[0124] In some embodiments of the present application, the semantic relationship can include a subordinate relationship and a cross relationship; wherein the subordinate relationship can reflect the hierarchical inclusion relationship between labels, such as "dog animal", which can reflect the logical concept of superior and inferior; the cross relationship mainly reflects the intersection of labels in different classification dimensions, such as "dog mammal" and "dog carnivore", which can reflect the overlapping attributes of the same entity in a multi-dimensional classification system.

[0125] In embodiments of the present application, the label expansion process is based on the semantic relationship of the label in the target classification label information, recursively generates more related labels from bottom to top, thereby generating a more abundant label set, which can supplement the deficiencies of the original classification result, making the image classification more comprehensive and accurate; for example, if the preliminary classification result is the label of cat, the labels of pet and animal can be further added through the label expansion process, thereby more completely expressing the image content.

[0126] It can be understood that in an actual application scenario of album image classification, the label expansion can improve the semantic coverage of image classification and enhance the performance of image search and album management in the album; for example, when a user uses an animal keyword to search for pictures, even if the original classification only labels a cat label, due to the generation of the animal label in the label expansion process, the image containing the cat can still be correctly retrieved, thereby improving the practicability and user experience of the system.

[0127] Therefore, according to the label recursive expansion mechanism from bottom to top, the embodiment of the application can start from a specific classification result, infer a more general semantic category from one label, thereby enriching the classification hierarchy and improving the expression ability and practicability of the classification result.

[0128] In some embodiments of the application, for image classification in an album, the image classification method of the electronic device can include the following steps:

[0129] Step 301, performing first classification processing and second classification processing on image data in the album to obtain first classification label information and second classification label information of the image data.

[0130] In the embodiments of the application, for the scenario of image classification in an album, by performing first classification processing and second classification processing on images in the album, the images can be classified at different granularities, which not only ensures the accuracy of classification, but also improves the comprehensiveness of classification, thereby better meeting the user's demand for album image classification.

[0131] It can be understood that by the first classification processing and the second classification processing, the images can be classified at different granularities, which not only ensures the accuracy of classification, but also improves the comprehensiveness of classification, thereby better meeting the user's demand for album image classification.

[0132] Step 302, determining target classification label information of the image data based on the first classification label information and the second classification label information, and storing the image data according to the target classification label information.

[0133] In the embodiments of the application, the classification storage refers to storing the pictures into the corresponding classification folder or database according to the target classification label information. For example, all pictures labeled as "cat" are stored under the "cat" folder; if a "pet" label is expanded based on the single "cat" label, the pictures can also be stored under the "pet" folder to facilitate multi-level retrieval; in this way, the user's search efficiency can be greatly improved, and more flexible classification management can be supported.

[0134] It can be seen that the embodiment of the present application can realize an efficient and accurate album picture classification method. This method can not only adapt to the limitation of mobile terminal local computing power, but also effectively deal with classification difficulties such as complex background and small foreground objects, realize efficient and flexible album picture classification management, and significantly improve user experience and system performance. The embodiment of the present application provides an image classification method. An electronic device obtains first classification label information obtained by performing first classification processing on a target image; obtains second classification label information obtained by performing second classification processing on the target image; verifies the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verifies the first classification label information based on the second classification label information to obtain a classification label information verification result; and fuses the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image. It can be seen that by performing first classification processing and second classification processing on the image respectively, multi-dimensional classification results can be obtained, and by mutual verification between the two classification results, conflicting or inconsistent classification results can be effectively excluded, thereby improving the accuracy of classification. At the same time, by using the mutual verification relationship between the classification results, the robustness of the classification can be enhanced while maintaining the classification efficiency, which is especially suitable for complex background, small and diversified foreground scenes, greatly improving the performance and efficiency of image classification.

[0135] Based on the above embodiments, in another embodiment of the present application, an exemplary mobile terminal foreground target detection and single-label classification combined album picture classification method is proposed, wherein the electronic device can be a mobile terminal device, such as a mobile phone. The embodiment can effectively improve the album classification ability based on local computing power, mainly including decomposing the composite category and multi-label picture into specific and single categories, making the detection task and classification task simpler, and fusing the classification labels after verifying the results of the two models. Finally, the specific labels are recursively expanded from bottom to top. The above method mainly includes the following contents: foreground target detection (first classification processing), which can use the advantages of target detection to detect specific categories, separate the foreground target from the frame, and more easily detect multi-target, complex background target, and small foreground target. The output detection frame selects the best boundary between categories to reduce interference between different categories; single-label classification (second classification processing), which can classify the picture into specific categories, mainly for single classification task of scene category, background feature single or some large size foreground target picture, and decomposes the wide category and multi-label picture to make the data features more uniform and reduce the task difficulty. Since the second classification processing is mainly used for classifying the main scene category of the whole image, to ensure the efficiency of image classification, the final classification label can be one label, so it belongs to single-label classification. The results of detection and classification can be fused according to the confidence priority principle, and the categories with higher confidence in the foreground target detection and single-label classification can be used as the evaluation standard. The classification results with lower confidence are reasonably evaluated, and the classification results that do not meet the standard are excluded. The label recursion from bottom to top can recursively expand the specific category to a wider category, such as outputting the classification of cat, and the upward classification of pet and animal.

[0136] In some embodiments of the present application, since the confidence priority principle is used to determine the final classification result, i.e., the category with the highest confidence is used as the standard, if there are multiple categories with confidence, only the category with the highest confidence can be obtained, and it is judged whether the category with the highest confidence meets the confidence threshold. Other categories do not need to be judged based on the confidence threshold, which can reduce the difficulty of image classification and improve the classification efficiency.

[0137] In some embodiments of the present application, during the single-label classification process, multiple category labels about the scene of the image can still be obtained; for example, during the model training phase, the label of the image can be "decomposed downward" to obtain multiple labels with more fine-grained texture, and then during the model inference phase after the model training is completed, the generated labels can be "recursively expanded upward" to obtain more labels.

[0138] It should be noted that foreground target detection is conducive to the feature capture of foreground targets, and also reduces the mutual exclusion problem between targets of different categories. Compared with the classifier classifying the entire image into one category, it has higher accuracy and reduces the difficulty of classifying images with complex backgrounds. For single-label classification, one image can get one classification label, and the image features in the training set are relatively consistent. Compared with multi-label networks, data organization is simple, the model is easy to converge, the model is more accurate, and the workload of data organization, data set construction and data analysis is much smaller; verifying the two results of foreground target detection and single-label classification can improve the reliability of the final classification results; this application can reduce the workload of each classification task, make the task simpler, and increase the reliability by breaking down the complex classification task into single tasks such as foreground target detection and single-label classification.

[0139] For example, Figure 2 As shown in the figure, when classifying an image, foreground target detection and single-label classification can be performed first, and then the results of foreground target detection and single-label classification are verified and fused based on the principle of confidence priority. Finally, the obtained labels are expanded from bottom to top to obtain the final image label.

[0140] In some embodiments of the present application, during foreground target detection, redundant detection frames can be removed through different levels of non-maximum suppression (NMS) operations, including overlapping frame suppression processing within the same category, overlapping frame suppression processing across categories, and deduplication processing for each category; by removing redundant detection frames, the most representative target frames can be retained, thereby improving the quality of the detection results; this process not only improves the detection accuracy, but also reduces the risk of misjudgment due to target overlap or interference in subsequent classification tasks, thereby improving the overall reliability of the classification.

[0141] In some embodiments of the present application, the purpose of foreground target detection is to detect multiple targets, targets with complex backgrounds, and small foreground targets, such as specific target objects such as cats, dogs, and flowers.

[0142] For example, the foreground target can be framed and separated from the background, and then multiple targets can be classified separately to reduce the difficulty of classifying multiple targets, targets with complex backgrounds, and small foreground targets; Figure 3 As shown, a picture includes many possible detection frames of "cat" and "dog", and redundant frames can be removed. After obtaining the candidate detection frames, overlapping frames of the same category are suppressed to obtain the first detection frames of "cat" and "dog", and then overlapping frames of different categories are suppressed to further remove the overlapping frames in the first detection frames of "cat" and "dog". Finally, the second detection frames of the two categories of "cat" and "dog" are deduplicated to obtain the final target frames of the two categories of "cat" and "dog".

[0143] For example, Figure 4 As shown in the figure, in order to make the classification task simpler, the classification task uses more specific and single data types to train the model. In the model training stage, for the training data samples, large-size pictures containing only scene categories are used. These pictures have single foreground objects and background features, and each picture has only one label. For the annotation of pictures with multiple labels, start with broad categories and use the top-down recursive principle to decompose them, so that the data features of each category are more unified and the task difficulty is reduced. For example, a picture of a beach has two labels, "beach" and "scenery". Because "beach" belongs to "scenery", the picture is labeled as "beach"; after the model training is completed, for the model inference stage, the label can be expanded from the bottom to the top. For example, the picture can be identified as a beach and then expanded to the label of "scenery".

[0144] In some embodiments of the present application, if the label of a sample image is finally decomposed and contains multiple different labels, then the sample image can be used in foreground target detection or discarded. For example, if a sample image contains both a person and a cat, then this sample image will be placed in the sample library of foreground target detection for training the target detection model.

[0145] In some embodiments of the present application, in addition to fusing different labels based on the confidence of the labels after obtaining the classification label information verification results, the confidence space of the labels in the first classification label information and the confidence space of the labels in the second classification label information can also be aligned first, and then the classification label information can be verified based on the confidence of each label in the aligned confidence space.

[0146] Exemplarily, the labels in the first classification label information include "person" and "train", the labels in the second classification label information include "beach", the confidence of "person" is 95%, the confidence of "train" is 81%, the confidence of "beach" is 6.6, the maximum value of the confidence in the confidence space corresponding to the first classification label information is 100%, the confidence threshold is 80%, the maximum value of the confidence in the confidence space corresponding to the second classification label information is 7.0, and the confidence threshold is 5.0; it is assumed that the confidence space is aligned based on the respective confidence thresholds of the two kinds of classification processing, for example, the confidence space corresponding to the first classification label information is aligned to the confidence space corresponding to the second classification label information, and one implementable manner is: dconf = conf_thres + (dconf - d_thres) / (d_MAX - d_thres) × (conf_MAX - conf_thres), where dconf represents the aligned confidence space of the first classification label information, d_thres represents the confidence threshold in the confidence space corresponding to the first classification label information, d_MAX represents the maximum value of the confidence in the confidence space corresponding to the first classification label information, conf_thres represents the confidence threshold in the confidence space corresponding to the second classification label information, and conf_MAX represents the maximum value of the confidence in the confidence space corresponding to the second classification label information; the above formula is only an exemplary implementation manner, and other implementation manners can also be used to align the confidence in the implementation process, which are not listed here.

[0147] Exemplarily, as shown in FIG. 6, after the foreground target detection and the single-label classification are performed on the picture, the labels in the first classification label information include "train" and "person", the confidence of "person" is 95%, and the confidence of "train" is 81%; the labels in the second classification label information include "beach", and the confidence of "beach" is 6.6; after the confidence mapping, the confidence of "person" is converted to 6.5, and the confidence of "train" is converted to 5.1; taking the class with the highest confidence "beach" as a reference, it is determined that "train" and "beach" are mutually exclusive (it is not likely that a train appears on a beach) by using the preset class mutual exclusion table, and thus the label of "train" is excluded, and finally the output label includes "person" and "beach". Figure 5 Exemplarily, as shown in FIG. 7, it is assumed that the final label output by the foreground target detection and the single-label classification is "cat", and the label of the picture can be recursively expanded in a bottom-up manner, and the specific class can be used to push up the combined class, which can include "pet" and "animal".

[0148] Figure 6 Exemplarily, as shown in FIG. 8, it is assumed that the final label output by the foreground target detection and the single-label classification is "cat", and the label of the picture can be recursively expanded in a bottom-up manner, and the specific class can be used to push up the combined class, which can include "pet" and "animal".

[0149] ​Exemplarily, in the scene category, a bottom-up classification principle is adopted, if some pictures exist a composite multi-classification condition, for example, a picture of a landscape, the picture is classified as "beach", because "beach" belongs to "landscape", the picture can be expanded according to the bottom-up recursive expansion principle to obtain "landscape", so the picture has two labels: "beach" and "landscape".

[0150] To sum up, the mobile terminal device in the embodiment of the application combines foreground target detection with single-label classification, and uses result fusion and label recursive expansion mechanism, to realize an efficient album image classification method; in this way, the classification accuracy and processing efficiency under the local computing power of the mobile terminal can be significantly improved, so as to meet the needs of users for privacy protection and fast classification, and thus the technology can be promoted to be widely applied in more intelligent terminal scenarios.

[0151] In the embodiment of the application, further, as shown in Figure 7 The electronic device 1 can further include a processor 11, a memory 12 storing processor 11 executable instructions; further, the electronic device 1 can further include a communication interface 13, and a bus 14 for connecting the processor 11, the memory 12 and the communication interface 13.

[0152] In the embodiment of the application, the above-mentioned processor 11 can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices used to realize the functions of the above-mentioned processor can also be others, and the embodiments of the application do not make specific limitations. The electronic device 1 can further include a memory 12, which can be connected with the processor 11, wherein the memory 12 is used to store executable program codes, the program codes include computer operation instructions, the memory 12 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least two disk memories.

[0153] In the embodiments of the present application, the bus 14 is used to connect the communication interface 13, the processor 11 and the memory 12 and the mutual communication among these devices.

[0154] In the embodiments of the present application, the memory 12 is used to store instructions and data.

[0155] Further, in the embodiments of the present application, the above-mentioned processor 11 can be used to obtain first classification label information obtained by performing first classification processing on a target image;

[0156] obtain second classification label information obtained by performing second classification processing on the target image;

[0157] verify the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verify the first classification label information based on the second classification label information to obtain a classification label information verification result;

[0158] fuse the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.

[0159] In actual application, the above-mentioned memory 12 can be a volatile memory (such as a Random-Access Memory (RAM)), or a non-volatile memory (such as a Read-Only Memory (ROM), a flash memory, a Hard Disk Drive (HDD) or a Solid-State Drive (SSD)), or a combination of the above-mentioned kinds of memories, and provides instructions and data to the processor 11.

[0160] Specifically, the program instructions corresponding to the image classification method in the embodiments can be stored on a storage medium such as an optical disc, a hard disk, a U disk, etc., and when the program instructions corresponding to the image classification method in the storage medium are read by the processor or executed, the following steps are included:

[0161] obtain first classification label information obtained by performing first classification processing on a target image;

[0162] obtain second classification label information obtained by performing second classification processing on the target image;

[0163] verify the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verify the first classification label information based on the second classification label information to obtain a classification label information verification result;

[0164] fuse the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.

[0165] When the program instruction corresponding to an image classification method in the storage medium is read by the processor or executed, the following steps can also be included:

[0166] obtain first classification label information obtained by performing first classification processing on the target image;

[0167] obtain second classification label information obtained by performing second classification processing on the target image;

[0168] verify the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verify the first classification label information based on the second classification label information to obtain a classification label information verification result;

[0169] fuse the first classification label information and the second classification label information based on the classification label information verification result to obtain target classification label information of the target image.

[0170] In summary, by performing first classification processing and second classification processing on the image respectively, multi-dimensional classification results can be obtained, and by mutual verification between the two classification results, conflicting or inconsistent classification results can be effectively excluded, thereby improving the accuracy of classification. At the same time, by using the mutual verification relationship between the classification results, the robustness of the classification can be enhanced while maintaining the classification efficiency, which is especially suitable for scenes with complex background, small and diversified foreground targets, greatly improving the performance and efficiency of image classification.

[0171] In addition, each functional module in the embodiment can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional module.

[0172] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium based on such understanding. The technical solutions of the embodiments essentially or the parts that contribute to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the embodiments. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0173] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer usable program codes.

[0174] The present application is described with reference to the implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure One The device that implements the functions specified in one flow or multiple flows and / or blocks Figure One The device that implements the functions specified in one block or multiple blocks.

[0175] These computer program instructions can also be stored in a computer readable storage medium that can guide the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure One The device that implements the functions specified in one flow or multiple flows and / or blocks Figure One The device that implements the functions specified in one block or multiple blocks.

[0176] These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are generated to realize the computer-implemented process, and the instructions executed on the computer or other programmable devices provide a process for implementing the functions specified in the flowchart or block diagram Figure One flowchart or block diagram. Figure One flowchart or block diagram.

[0177] The above embodiments are only preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Any equivalent replacement or transformation made by those skilled in the art based on the present application shall fall within the protection scope of the present application.

Claims

1. An image classification method, comprising: Obtaining first classification label information obtained by performing a first classification process on the target image; Obtaining second classification label information obtained by performing a second classification process on the target image; Verifying the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verifying the first classification label information based on the second classification label information to obtain a classification label information verification result; The first classification label information and the second classification label information are fused based on the classification label information verification result to obtain target classification label information of the target image.

2. The image classification method according to claim 1, The first classification process is used to determine whether the target object in the target image is classified into a category in the first classification set; The second classification process is used to determine whether the target image is classified into a category in the second classification set.

3. The image classification method according to claim 2, wherein obtaining first classification label information obtained by performing a first classification process on the target image comprises: performing overlapping frame suppression processing of the same category on candidate detection frames in the image, respectively used to mark target objects of the corresponding category, to obtain first detection frames corresponding to the categories included in the first classification set; performing cross-category overlapping frame suppression processing on the first detection frames of the respective categories to obtain second detection frames of the respective categories; Deduplication processing is performed on the second detection frame of each category to obtain a target frame of each category; The first category label information is obtained based on the category to which the target object contained in the target frame belongs.

4. The image classification method according to claim 2, wherein obtaining second classification label information obtained by performing a second classification process on the target image comprises: Obtaining a feature vector corresponding to the target image; Based on the probability that the feature vector belongs to the vector space corresponding to each category in the second category set, the category to which the target image belongs in the second category set is determined, and the second category label information is obtained.

5. The image classification method according to claim 1, wherein the step of fusing the first classification label information and the second classification label information based on the classification label information verification result to obtain the target classification label information of the target image comprises: If the classification label information verification result indicates that the first classification label information and the second classification label information have mutually exclusive first labels and second labels, determining a target label included in the target classification label information corresponding to the target image based on a first confidence level corresponding to the first label and a first confidence level corresponding to the second label; The first label belongs to the first classification label information, the second label belongs to the second classification label information, and the target label is the first label or the second label.

6. The image classification method according to claim 5, wherein determining the target label included in the target classification label information corresponding to the target image based on the first confidence level corresponding to the first label and the first confidence level corresponding to the second label comprises: Mapping the first confidence level to a third confidence level corresponding to the second classification process; If the third confidence level is higher than the second confidence level, determining the first label as the target label included in the target classification label information corresponding to the target image; If the third confidence level is not higher than the second confidence level, determining the second label as the target label included in the target classification label information corresponding to the target image; or Mapping the second confidence level to a fourth confidence level corresponding to the first classification process; If the first confidence level is higher than the fourth confidence level, determining the first label as the target label included in the target classification label information corresponding to the target image; If the first confidence level is not higher than the fourth confidence level, the second label is determined as the target label included in the target classification label information corresponding to the target image.

7. The image classification method according to claim 6, wherein mapping the first confidence level to a third confidence level corresponding to the second classification process comprises: Converting the first confidence level according to the data format of the confidence level corresponding to the second classification process to obtain the third confidence level; The data format of the third confidence level is the same as the data format of the second confidence level; Accordingly, mapping the second confidence level to a fourth confidence level corresponding to the first classification process includes: Converting the second confidence level according to the data format of the confidence level corresponding to the first classification process to obtain the fourth confidence level; The data format of the fourth confidence level is the same as that of the first confidence level.

8. The image classification method according to any one of claims 1 to 7, further comprising: If the target classification label information includes one label, label expansion is performed on the one label according to the semantic relationship of the one label to obtain a plurality of expanded labels.

9. The image classification method according to any one of claims 1 to 7, further comprising: Performing a first classification process and a second classification process on the image data in the album to obtain first classification label information and second classification label information of the image data; Target classification label information of the image data is determined based on the first classification label information and the second classification label information, and the image data is classified and stored according to the target classification label information.

10. An electronic device comprising a processor and a memory storing instructions executable by the processor; when the instructions are executed by the processor, the following steps are implemented: Obtaining first classification label information obtained by performing a first classification process on the target image; Obtaining second classification label information obtained by performing a second classification process on the target image; Verifying the second classification label information based on the first classification label information to obtain a classification label information verification result, and / or verifying the first classification label information based on the second classification label information to obtain a classification label information verification result; The first classification label information and the second classification label information are fused based on the classification label information verification result to obtain target classification label information of the target image.